SELF-HOSTED • PRIVATE • BUILT FOR LEGACY

The Origin Story
of Leone Cineplex

A personal system powering media, knowledge, and AI for my family — explained simply.

AI
MEDIA
FAM
Built for my sons • Privacy-first • No cloud lock-in

The Origin of Leone Cineplex

Long before Netflix existed, I had a simple dream: every movie, song, audiobook, and book I loved, available instantly at home — on my terms, with no subscription and no one else deciding what I could watch.

It started with VHS tapes that took over floors and walls. Then came DVDs, which eventually demanded entire walls of shelving. When the DVD burning era arrived, I saw the future and made the jump to digital files stored on ZIP disks and hard drives. That decision, more than twenty years ago, set me on the path of building something most people now pay monthly for.

Today, that vision lives on LeoneNAS 7 (our 73 TB media vault) coupled with LeoneMini-2 (our dedicated AV1 media engine) — a complete personal media empire for my family and close friends. Thousands of movies, TV shows, albums, audiobooks, comics, and books, perfectly organized and instantly streamable to any device, anywhere. No algorithms. No ads. No data harvested.

The same mindset that drove me from VHS to a self-hosted media library — own it, understand it, improve it — is exactly why I built LeoneNAS 9 for AI. These machines have become both a playground and a classroom, not just for me, but for my sons.

30,000+ movies
1,000+ TV series
275 audiobooks
20,000+ albums
9,100+ authors
20,000+ books

The Path to Local AI

While most of the world rushed to use AI through cloud services, I chose to build it myself. Over the last year and a half, I’ve been building, breaking, and rebuilding AI servers because I wanted to understand these systems from the ground up — not just consume them.

Artificial intelligence is the most significant technological shift since the internet. Those who understand how it works at the infrastructure level will have a real advantage.

LeoneNAS 9 is my dedicated AI laboratory. By running powerful models locally instead of sending everything to corporate clouds, I maintain complete sovereignty over my data, my conversations, and my knowledge.

The deepest purpose is my sons. Both already show strong technical aptitude. I want them to grow up with an intimate, hands-on understanding of AI — how to run it, modify it, and trust it. This project gives them a living classroom in systems engineering, open-source architectures, and local AI orchestration.

It’s about building something that belongs to us — a foundation of knowledge and independence they can inherit and build upon for the rest of their lives.

132 TOPS Total AI Capacity
36 TOPS NVIDIA A2 Discrete
78 TB RAIDZ1 Storage Pool
64 GB LPDDR5X-8400 RAM
MEDIA VAULT & TRANSCODING STACK

Media Stack

LeoneNAS-7 & LeoneMini-2

This dual-node engine is the heart of our family media world. LeoneNAS 7 holds the 73 TB media vault and runs core backup and Caddy proxy services, while LeoneMini-2 provides dedicated, ultra-efficient AV1 hardware transcoding for high-concurrency streaming.
CURRENT ROLES
Plex (Lunar Lake AV1)
Navidrome + Audiobookshelf
Calibre + Kavita
Caddy Proxy + Backup Vault

Media Infrastructure Specs

LeoneNAS 7 — Storage Vault & Proxy (QNAP TS-664)
CPU
Intel Celeron N5105
4 Cores / 4 Threads • QTS 5.2.9
MEMORY
32 GB DDR4
Dual-channel Non-ECC
STORAGE POOL
73 TB Usable RAID 5
6-bay • Seagate Enterprise HDDs
NETWORK
2 × 2.5 GbE (LACP)
Intel i226-V Bonded
LeoneMini-2 — Dedicated Media Engine (GMKtec NucBox K13)
CPU & AV1 GPU
Intel Core Ultra 7 256V
Lunar Lake • Arc 140V (Hardware AV1)
MEMORY & OS
16 GB LPDDR5X-8533
Proxmox VE 9.2.5
SYSTEM NVME
1 TB PCIe 4.0 NVMe
>7,000 MB/s + Open Expansion Slot
NETWORK
1 × 5 GbE High-Speed
Realtek RTL8126

Why This Architecture Works

  • LeoneMini-2 handles high-efficiency hardware AV1/4K transcoding so everyone can stream on any device without buffering or high energy cost.
  • LeoneNAS 7 keeps 73 TB of high-capacity vault media safe on Seagate Enterprise drives with mirrored NVMe boot drives.
  • Dedicated high-speed 5 GbE and bonded 2.5 GbE links keep playback smooth even with simultaneous family streams.
DEDICATED AI INFERENCE SERVER

LeoneNAS 9

The AI Brain

This is my dedicated AI powerhouse (UGREEN NASync iDX6011 Pro) and the home of "Alister" — my private AI assistant and digital twin. Every large language model, every personal knowledge base, and every heavy generative task runs here on bare-metal hardware under Proxmox VE 9.2.5.
I spent over a year and a half testing, configuring, and optimizing local AI infrastructure because I wanted to understand this technology from the inside out. I want my sons to grow up knowing how these systems actually work at the hardware and model architecture levels.
132 TOPS TOTAL AI CAPACITY

Hardware Specs

CPU + iGPU / NPU
Intel Core Ultra 7 255H
16 Cores • 96 TOPS NPU + Arc
DISCRETE GPU
NVIDIA A2 16 GB GDDR6
36 TOPS INT8 • Inference Engine
MEMORY
64 GB LPDDR5X-8400
High-speed unified memory
STORAGE & NETWORKING
78 TB RAIDZ1 • 2× 10 GbE
Proxmox VE 9.2.5 • Dual TB4

Why This Hardware Excels at AI

  • 132 Combined TOPS (NVIDIA A2 + Intel Arc NPU) dramatically speeds up local LLM inference and token generation so responses feel natural.
  • 64 GB LPDDR5X-8400 RAM provides massive memory bandwidth required to load large models with high context windows.
  • Proxmox VE 9.2.5 virtualizes Ollama, OpenWebUI, AnythingLLM, and local RAG search pipelines cleanly with hardware GPU passthrough.
HOMELAB INFRASTRUCTURE FLEET

Full Fleet At A Glance

Purpose-built hardware nodes operating on Proxmox VE 9.2.5 and QTS 5.2.9

LeoneMini-1
Pulcro TurnKey
i3-1215U • 24 GB DDR4
Home Automation (HAOS VM) & Edge Services
LeoneMini-2
GMKtec K13
Ultra 7 256V • 16 GB • 5GbE
Dedicated Media Server & AV1 Hardware Transcode
LeoneNAS-7
QNAP TS-664
N5105 • 32 GB • 73 TB Usable
Media Vault Storage Pool & Caddy Reverse Proxy
LeoneNAS-8
SuperMicro 2U
2× E5-2690 v3 • 256 GB ECC
Test Lab & Sandbox Environment
LeoneNAS-9
UGREEN iDX6011 Pro
Ultra 7 255H • A2 16GB • 78 TB
Primary AI Inference Node (Alister RAG)
EDUCATIONAL DEEP DIVE

How It All Works

Click any topic to expand. Everything starts closed.

What is it? Large Language Models are AI systems trained on enormous amounts of text. “Inference” is the process of the model actually thinking and generating a response in real time.

Running inference locally on LeoneNAS 9 means all of this computation happens in my house. Nothing is sent to OpenAI, Google, or Anthropic. Your conversations stay completely private.

Ollama + Llama Models

Llama is a family of powerful open-source AI models from Meta. Ollama makes it easy to download and run them locally with excellent reasoning and conversation quality.

Open WebUI

A private, feature-rich chat interface that connects directly to Ollama. It supports voice, document processing, multi-user access, and RAG — making advanced local AI easy for the whole family to use.

Normal LLMs have no knowledge of your specific life or documents. RAG fixes this by searching your private files first, then feeding the most relevant pieces to the model before answering. The result is grounded, accurate, and personal.

Meet "Alister"

My personal AI twin running on LeoneNAS 9. Powered by Ollama and OpenWebUI. It knows my projects, business docs, family stories, and homelab knowledge — and can answer questions about them accurately.

Traditional search looks for exact word matches. Semantic search understands meaning and intent. On LeoneNAS 9, the GPU accelerates creating “embeddings” so searches across thousands of documents happen almost instantly — even when the exact words don’t match.

Beyond answering questions, generative AI can create images, artwork, event graphics, and more. On LeoneNAS 9 we run local models (Stable Diffusion, Flux, etc.) so everything stays private and under our control — no Midjourney accounts or data leaving the house.

This infrastructure represents years of learning, building, and a commitment to ownership and privacy. It’s not just servers — it’s a foundation I’m building for my sons and anyone who cares about controlling their own digital life.