Tomorrow, unexpectedly, Kimi K3 goes open weight. The announcement, leaked on Reddit, arrives with the excitement of seeing another frontier LLM step out of the cathedral of closed APIs. But one detail dampens the mood: «I can’t run it, or even a model a hundred times smaller,» writes the user who broke the news. It’s not a provocation; it’s a snapshot of a structural problem.
The exact size of Kimi K3 isn’t official yet, but the comment speaks volumes. If a model one hundred times smaller is still unmanageable on personal hardware, we’re dealing with an LLM that requires dozens of GPUs with hundreds of GB of VRAM just for inference. That’s far beyond any on-premise deployment that isn’t a corporate datacenter with an extreme hardware budget. Self-hosting—so celebrated when Meta released Llama—hits a wall here: some models are born cloud-native, and open weights don’t change physics.
The inference provider trap
The Reddit user adds a revealing note: they are most looking forward to «the new inference providers that will hopefully open up.» That’s where the short circuit lies. In theory, an open-weight model allows anyone to fine-tune, study the weights, customize behavior. But if you have to rely on a third-party provider to actually use it, independence shrinks. It’s not very different from a proprietary API: inference leaves your perimeter, data flows on someone else’s machines, and sovereignty—the core of the on-premise argument—crumbles.
This scenario reshapes incentives. On one side, cloud infrastructure providers (and new specialists in open-weight LLMs) gain a high-margin service; on the other, organizations that bet on self-hosting for compliance or GDPR reasons find themselves locked out, unless they make colossal investments in accelerators. Kimi K3 thus becomes a consumer product for hyperscalers, while smaller labs and companies with ordinary IT budgets remain spectators.
What this story tells us
Open weights are no longer synonymous with hardware accessibility. For years, we’ve linked the release of weights to genuine democratization—download the model, put it on a server with a few GPUs, and get to work. Today, the gap between model growth and the capacity of typical on-premise machines is so wide that democratization risks turning into a fiction: the weights are there, but the key to using them is held by a few.
Kimi K3 also signals an interesting market evolution. If inference becomes the bottleneck, we’ll see a proliferation of providers offering APIs for open-weight models, with margins built on hardware optimization (quantization, sharding, optimized serving). It’s a business model that bridges open source and managed services, but for those with data residency obligations or audit requirements, it remains a tough compromise.
Ultimately, the question is not just “how good is Kimi K3?” but “how runnable is it without selling your soul to the cloud?” The whole self-hosting debate shifts from weight availability to the availability of affordable hardware for large-scale inference—a topic AI-RADAR tracks closely, because this is where the real architectural choices are made for those unwilling to surrender data or control.
💬 Comments (0)
🔒 Log in or register to comment on articles.
No comments yet. Be the first to comment!