The price of the frontier has shrunk to a unit of time: four months and four days. According to Mozilla's State of Open Source AI report, published on September 15 and shared with Ars ahead of publication, the performance gap between the best Chinese open-weights models and closed models from US tech companies has fallen to 4.4 months. Translated into budget terms, the message is blunt: paying for a frontier model no longer buys an unbridgeable lead, but a head start of just over a quarter at five times the cost.

The report points to Kimi K3 from Moonshot AI, which scores only three points lower on the Artificial Analysis Intelligence Index than Fable 5, the closed Anthropic model named in the survey, while costing about 30 percent less. Raffi Krikorian, Mozilla's CTO, frames the dividing line: closed models earn their premium in a few places — expert professional work, high-intensity retrieval, and long context. The decision to pay for closed, he says, is workload-specific rather than organization-specific.

That shift has structural consequences. If the gap is measured in months, using closed APIs for routine work becomes a convenience, not a technical necessity. The marginal per-token cost of a frontier model should be compared with the TCO of an open-weights model served in-house: hardware, VRAM, quantization, orchestration, and maintenance. It is not free, but for repetitive workloads the break-even point can move quickly.

For organizations with data residency constraints or requirements such as GDPR, the picture changes further. An open model running on self-hosted infrastructure avoids sending prompts and context outside the corporate perimeter. It does not remove security and compliance responsibilities, but it restores control over where and how data is processed.

The Mozilla report does not declare the end of closed models. It redefines their market. The three workloads mentioned by Krikorian — expert professional work, high-intensity retrieval, and long context — are exactly those where reliability, latency, and output quality are worth more than price. They are also the hardest to absorb with modest local hardware: long contexts require more VRAM, intensive retrieval requires fast pipelines and coherent outputs.

The narrowing time advantage also signals a shift in the competitive geography: Chinese open models are closing the gap with US tech companies. For those evaluating on-premise deployments, there are trade-offs between flexibility and operational burden; AI-RADAR offers analytical frameworks at /llm-onpremise to separate real costs from perceived ones. The point is not whether open models will overtake closed ones, but when the residual premium will no longer cover the difference in control.