GUIDES

AI infrastructure guides

In-depth, vendor-neutral reference guides for choosing, building and running LLMs locally and on-premise. Written for engineers, architects and decision-makers.

📥 Free resources
📄 Hardware cheat-sheet 🛡️ EU AI Act checklist

Hardware & cost

🎮
Best GPUs for local LLMs VRAM tiers, value picks and what each card runs.
📐
How much VRAM for Llama 70B The sizing formula, quantization and KV-cache.
🧮
KV-cache & context memory math Why long context eats VRAM — the real formula.
💸
Cost of running LLMs locally Full TCO, worked example and the break-even.
RunPod vs Vast.ai GPU cloud compared: price, reliability, production.
🗜️
LLM quantization explained GGUF vs GPTQ vs AWQ; quality vs size.

Deploy & methods

🛠️
Local LLM software stack Ollama vs LM Studio vs vLLM: prototype to production.
⚖️
Ollama vs LM Studio The head-to-head: which local LLM runner, and when.
🏎️
vLLM vs llama.cpp vs TGI vs SGLang The serving showdown: throughput, batching, which to pick.
🧩
RAG vs fine-tuning Knowledge vs behavior — which to use, and when to combine.
🏢
Private ChatGPT for your company Architecture, model, RAG, hardware and security.

On-premise & compliance

⚖️
On-premise vs cloud AI Cost, control, compliance and the hybrid model.
🛡️
EU AI Act & on-premise Risk tiers, GPAI rules and a compliance checklist.
🇮🇹
On-premise LLMs for the Italian SMB GDPR, budgets (€3k–€30k) and build-vs-API break-even.

Step-by-step tutorials

🖥️
Install Ollama + Open WebUI on Windows Your private ChatGPT-style assistant in 20 minutes. BEGINNER · 20 MIN
🏭
Serve your team with vLLM in Docker Production API + chat UI on one 24GB GPU. INTERMEDIATE · 40 MIN
⌨️
Local coding assistant in VS Code Continue + Cline + Ollama: private Copilot, zero subscription. BEGINNER · 30 MIN
💬
Build a private RAG chatbot FastAPI + ChromaDB — the exact stack behind this site's Ask. INTERMEDIATE · 90 MIN
🧬
Fine-tune with QLoRA on 24GB Dataset → training → GGUF → Ollama, with honest evaluation. ADVANCED · 2 H

Looking for a specific model, or the full enterprise picture?

🚀 Run models locally → 💾 On-premise LLM observatory →