DeepSeek V4 Flash in a Basement: ECC RAM Changes the On-Premise Equation
An Epyc 7663, 256 GB ECC RAM and one RTX 5090 run a 151 GB MoE LLM at 24 tokens/s with Q8_K_XL quantization. The case challenges the assumption that aggregate VRAM must match model weights: balancing CPU, memory bandwidth and PCIe becomes the real co...