An EPYC server with twelve cards offering 64 GB of VRAM each, 256 GB of RAM and a €18,000 investment. Not long ago this would have been an enviable setup for running the most capable open source models locally, away from cloud lock-in. Today the owner of that machine is wondering whether he is stuck with hardware that frontier models are already leaving behind.

His experience starts from a concrete point: among the most capable open-weight models available today, GLM 5.3 appears to be the only one that realistically fits within 768 GB of VRAM. But the arrival of Astra, as he describes it, risks making that model quickly outdated. And the next GLM 6, according to the expectations cited in the source, could be twice or even three times as large. Qwen-max and Kimmi are already out of reach. Even DeepSeek V4 Pro is too big for that machine.

The stated constraint is 4-bit quantization. Going below that threshold, the author explains, introduces obvious errors. Fixing a minimum quality level turns the problem from simple memory availability into a question of practical usefulness: a machine that can only load a model with overly aggressive quantization no longer serves the purpose it was built for. If GLM 6 really surpasses two trillion parameters, handling 4-bit weights alone, together with execution overhead, puts the available 768 GB under serious strain. The point is not that the server stops working, but that it stops being a frontier model platform and becomes a machine for flash models.

This is where the structural fracture emerges. Frontier open source models are moving toward scales that assume multi-node clusters, data center power and cooling, not a single EPYC server. The user in question is not a company: he wanted to use this infrastructure to start a business. But if small organizations hit the same wall, the promise of self-hosted as an alternative to the cloud loses ground precisely where data control and sovereignty matter most. The winners are data center hardware vendors and cloud providers, which can aggregate demand and amortize machines that individual buyers cannot justify. The losers are the long tail of operators and hobbyists who invested in high-VRAM machines, now exposed to rapid technological depreciation.

The possible exit suggested by the author himself is to sell the excess GPUs and fall back on flash models with lower costs. That is not necessarily surrender: for many workloads, a smaller and well-integrated model can get close to the results of larger models, but it requires more work in evaluation, integration and fine-tuning. The cost shifts from hardware to engineering. For those choosing between local hardware and the compromises of smaller models, the trade-offs depend on workload, privacy constraints and amortization horizon. AI-RADAR offers analytical frameworks at /llm-onpremise to help evaluate these aspects, not ready-made answers.

The question this story leaves open is not only about a single builder. Does it still make sense to invest in a self-hosted machine if the next generation of models can make it obsolete? The open source model market is moving faster than the hardware meant to run it.