Open Weights vs Frontier Models Is a 4.4 Month Call
The open weights vs frontier models gap has closed to 4.4 months at roughly 3x the cost. Here is the math that decides which workloads justify the premium.

In this article
- 1.What the 4.4 Month Figure Actually Measures
- 2.Why a Head Start Is a Depreciating Asset
- 3.Scoring Open Weights vs Frontier Models in Four Questions
- 4.Workloads Where the Premium Compounds
- 5.Multi-step agent loops
- 6.Product-ceiling reasoning
- 7.Time-boxed competitive bets
- 8.Workloads That Never Needed a Frontier Model
- 9.Classification, extraction, and routing
- 10.Standard chat and summarization
- 11.Running the Math on Your Own Contract
- 12.What the Lag Number Cannot Tell You
The open weights vs frontier models gap has narrowed to about 4.4 months, and Mozilla's open source AI report puts its flagship pairing, Kimi K3 versus Fable 5, at roughly 3x, with the open model matching the frontier score at about 30 percent of the cost to run. Put those two numbers side by side and the question stops being which camp deserves your contract. What that premium buys is a temporary score lead with a published expiry date rather than a permanently better model, and the open-weight field has been closing that lead in about the 4.4-month window the report measures.
Stay in the loop.
Get the latest posts and exclusive content delivered to your inbox.
Join 5 readers. No spam. Unsubscribe in one click, anytime.
That reframing matters at contract time. Paying roughly 3x for durable superiority is a capability decision, and capability decisions can be argued forever. Paying roughly 3x for a lead that decays on a known schedule is a financing decision, and financing decisions have math. This piece supplies that math: what the lag figure actually measures, a four-question rule for scoring any workload, five archetypes with verdicts, a breakeven formula you can rerun every quarter, and the honest limits of trusting any benchmark parity claim.
What the 4.4 Month Figure Actually Measures
So how far behind are open weight models? Per Mozilla's analysis, published September 15, about 4.4 months, meaning the best open-weight release, in the report's framing the strongest of the Chinese open-weight models, matches a given closed US frontier model's benchmark performance roughly four and a half months after that frontier model shipped. The report's flagship comparison makes the point concrete: Moonshot AI's Kimi K3 lands within three points of Anthropic's Fable 5 on the Artificial Analysis Intelligence Index composite while costing around 30 percent as much to run. That 30 percent comparison, about 3x, is the only cost multiple the report puts on the table; treat anything rounder as a planning assumption you replace with your own quotes, not a report finding.
This style of measurement has a specific shape. You take a composite benchmark score a closed model posts at release, then find the date the open-weight field posts an equivalent score, and the elapsed time is the lag. It is the same family of analysis Epoch AI's open-closed gap data tracks, and its strengths and blind spots come from the method.
Three things the number is not:
- It is not a product parity score. The metric compares benchmark results near release time. It says nothing about latency, throughput, uptime, or serving reliability.
- It is not a compatibility score. Context window behavior, multimodal coverage, structured output stability, and tool-calling consistency are invisible to a composite score.
- It is not a commercial score. Licensing terms, hosting constraints, data residency options, and indemnification differ across providers and across open-weight families, and none of that shows up in a lag figure.
Benchmark parity on this metric is a strong signal that open weights have reached the capability neighborhood. It is not evidence that the two options are interchangeable products.
Why a Head Start Is a Depreciating Asset

The frontier model cost premium buys time, and time is the only thing it buys. If the open-weight field matches a frontier score in roughly 4.4 months, then any quality advantage you paid roughly 3x for has a half-life set by someone else's release calendar. Procurement language for this already exists: you are leasing a lead, not owning an edge.
Run the amortization once and the shape of the mistake becomes obvious. Suppose you sign a 12-month commitment priced on the assumption that the frontier model stays ahead. On the report's clock, the genuine advantage lasts about 4.4 months. The remaining 7.6 months of the term, roughly two-thirds of it, you are paying roughly 3x for what is, on the measured dimension, a tie. The premium bought a quarter and a half of differentiation, then spent the rest of the term as a loyalty tax.
The second problem compounds the first: the asset is depreciating faster over time. The direction of travel across recent model cycles has been toward a narrower gap, the same convergence trend Stanford's AI Index data tracks across benchmark families. Treat the 4.4 months as a reading from one report at one moment, not a physical constant. If the next cycle measures 3 months, the back end of your contract gets worse, not better, and the workloads that justified the premium shrink along with the window.
This is why the decision belongs at the workload level. An advantage that decays in months is still enormously valuable if your value capture also completes in months. It is nearly worthless if your value accrues steadily over years regardless of who holds the benchmark lead this quarter.
Scoring Open Weights vs Frontier Models in Four Questions
Mozilla CTO Raffi Krikorian put the report's conclusion in one line: the decision to pay for closed is "workload-specific rather than organization-specific." Here is that principle turned into something you can run in a planning meeting. For each workload, ask four questions.
| # | Question | What it tests | What a "no" means |
|---|---|---|---|
| 1 | Is this task at or beyond the open-weight ceiling today? | Whether head start is even for sale | Stop. Open weights already cover it. |
| 2 | Does quality compound through the task? | Whether small per-step deltas become large outcomes | The premium buys points you cannot feel. |
| 3 | Can you convert head start into value inside the lag window? | Whether the expiring asset is monetized before expiry | You are buying depreciation. |
| 4 | Does that value exceed the cumulative premium? | Whether the purchase clears breakeven | The deal is underwater on arrival. |
A workload earns frontier spend only with four yes answers. One no is disqualifying, and the most common failure pattern is answering question 1 honestly but never asking question 3: the task genuinely needs frontier quality, but nothing about the business captures value in under five months, so the premium quietly converts to parity spend.
Krikorian's named zones where closed "earns its premium," namely expert professional work, high-intensity retrieval, and long context, map cleanly onto questions 2 and 3. Those are settings where the model sits at the top of the value chain and where head start converts directly into shippable output.
The rule also exposes the most common LLM procurement error: signing one contract at the org level for a workload-level problem. The same team should rationally pay the premium for its agent orchestration and refuse it for its ticket routing, and bundling both under a single enterprise commitment means one of the two is mispriced. Score the workloads separately, then buy separately.
Workloads Where the Premium Compounds

The question of when to pay for frontier models over open weights resolves to one mechanism: compounding. Three archetypes clear all four questions.
Multi-step agent loops
Per-step error rates multiply, so tiny quality deltas become large end-to-end deltas. Illustrative arithmetic: a 10-step agent at 95 percent per-step success completes the full task about 60 percent of the time (0.95 to the 10th power). At 97 percent per step, it completes about 74 percent. Two points of benchmark head start bought you 14 points of task success, and task success is what your product sells. This is where the roughly 3x premium has a fighting chance of paying for itself.
Product-ceiling reasoning
Some products are only as good as the model inside them: expert drafting, hard synthesis, novel reasoning where the model's ceiling is the product's ceiling. Here every point of composite score is directly sellable differentiation, and the frontier lead is the feature. Krikorian's expert professional work zone lives here.
Time-boxed competitive bets
A launch timed to a window, a demo that decides a deal, a capability a rival cannot match for two quarters. The value is realized inside the lag window by design, which is exactly what question 3 demands. The bet is that you can ship before the lead stops being ahead, not that the model stays ahead.
Workloads That Never Needed a Frontier Model
Mozilla's recommendation goes one step further than the lag figure: run open models as the default for the majority of your work, because most of what organizations actually run never touches the capability frontier. Taken seriously, that position implies a waste ledger. For tasks below the open-weight ceiling, the frontier premium has been margin loss since parity arrived in an earlier cycle, not a head start at any point, and the ledger entry writes itself: monthly token spend times the premium multiple times the months since parity. At the report's flagship pairing of roughly 3x, a task moving 60 million tokens a month at illustrative prices of $2 open and $6 frontier bleeds $240 a month for as long as nobody re-tenders the contract.
Classification, extraction, and routing
Discrete outputs, clear quality floors, high volume. Open-weight models cleared these floors cycles ago; the residual score gap is worth fractions of a cent per million tokens, while the premium applies to every call. When latency, hosting control, or data residency matter, self-hosting open weights is often the strictly dominant option anyway, because an open license lets you place the weights wherever compliance requires them.
Standard chat and summarization
Routine support conversations and document compression saturate fast. Once a model family is comfortably past the quality floor, additional composite points buy nothing the customer notices, and cost per token is the only live variable left in the decision.
The five archetypes in one view:
| Workload | Compounds? | Value inside window? | Verdict |
|---|---|---|---|
| Multi-step agent loop | Yes, per-step | Often | Frontier can pay |
| Product-ceiling reasoning | Yes, model is ceiling | Often | Frontier can pay |
| Time-boxed competitive bet | Yes, by design | Yes, by definition | Frontier can pay |
| Classification, extraction, routing | No | No | Open weights |
| Standard chat, summarization | No | No | Open weights |
Running the Math on Your Own Contract
The breakeven condition is one line: the premium is rational when the value of roughly N months of head start exceeds the cumulative cost delta across those months. Everything else is filling in four inputs.
| Input | Where it comes from | Illustrative value |
|---|---|---|
| Monthly token volume | Gateway logs | 60M tokens |
| Open-weight all-in cost | API quote, or an open weight TCO calculator if self-hosting | $2 per million tokens |
| Frontier price | Current provider quote | $6 per million (illustrative, roughly 3x open) |
| Advantage window | Current lag figure, re-checked quarterly | 4.4 months |
Worked example, all numbers illustrative. At 60M tokens per month, the open-weight route costs $120 per month and the frontier route $360, a delta of $240. Across a 4.4-month advantage window, the premium totals roughly $1,050. That is the price of the head start. Now ask the value side: does having frontier-grade quality for those months produce more than $1,050 for this specific workload? An agent loop driving a paid product where each failed run costs support and churn, plausibly yes, many times over. An internal classifier, no, and not close.
Two adjustments make the model honest. First, if you are weighing self-hosted open weights, price the GPU amortization, engineering, and failover into that $2, rather than comparing sticker API prices; self-hosting changes both the cost base and the premium multiple, sometimes in your favor, sometimes badly against it. Second, redo this table whenever the lag figure updates. A shrinking window lowers the premium side of the inequality automatically, and workloads migrate from the pay column to the open column without anyone changing their mind.
What the Lag Number Cannot Tell You
Even a perfectly accurate 4.4-month figure is a benchmark statement, and benchmarks carry known failure modes: contamination in training data, selection effects in which benchmarks get tracked, and distribution mismatch between the benchmark suite and your traffic. Parity on a composite index is not parity on your workload, and no report can substitute for measuring your own.
Before any contract decision, run a calibration pass:
- Pull a few hundred to a few thousand prompts from real production traffic, not from leaderboards, and run both candidate models blind.
- Score failures by cost, not by count. A 2 percent error rate that corrupts downstream state can outweigh a 10 percent rate that triggers a retry.
- Read the actual license. Open-weight families range from permissive standards like Apache 2.0 to community licenses with usage restrictions, and your legal exposure differs accordingly.
- Check the dimensions the lag metric ignores: context handling at your real document lengths, latency and throughput against your SLA, modality coverage, and safety posture for your domain.
- Re-check the lag figure each cycle. The number that justified this quarter's decision will not justify next quarter's.
The 4.4 months will keep shrinking, and as it does, the frontier column gets shorter while the open-weight column gets longer. The four questions do not change. Ask them per workload, per quarter, and let the contract follow the math rather than the narrative.
Stay in the loop.
Get the latest posts and exclusive content delivered to your inbox.
Join 5 readers. No spam. Unsubscribe in one click, anytime.
About the author
Rachel Brennan
AI Research Editor
Rachel tracks AI research so the rest of us don't have to. With a background in NLP and a habit of reproducing papers, she turns new models and methods into ideas you can actually use.
Related Posts
NVIDIA Hugging Face Acquisition Ends the Neutral Hub
The NVIDIA Hugging Face acquisition turns the Hub into vendor infrastructure. Here is how builders mirror weights, pin revisions, and cut lock-in.
OpenAI Navier-Stokes, a Reported $40M Lesson in Verification
The OpenAI Navier-Stokes run reportedly burned $40M and 130 billion tokens yet produced no verified proof. Verification, not generation, now binds.
ChatGPT DSA Designation Sweeps In Every AI Search Engine
The ChatGPT DSA designation reportedly makes it the EU's first AI-native very large online search engine. Here is the test deciding which AI tools follow.


