Code generation: a model at 1/271st the price of the best scorer, at 97.8% of its quality.

We spent two weeks trying to beat single models with model combinations. Single models won.

In plain words

We route each request to the cheapest model that is measured to clear a quality bar for that kind of work. An obvious question is whether we should go further: run several models on one request and pick the best answer. We tested that idea five ways, with pre-registered targets and hard budget caps. Every test lost to the best single model, and this issue says exactly why.

The strongest model on our hardest code exam never missed — twice, across every item. But a model at roughly a twenty-third of its price came within about one point of it. That gap is the entire space a combination of models could win, and it is smaller than the cost of refereeing the combination.

This week's frontiers

One row per kind of work. Routed pick is the model Potion currently sends that work to. Stored quality is its exam score when it was measured in full, give or take the margin. Canary is this week's small re-check: a few fresh tasks, scored the same way, to catch a model that has got worse. Held means the re-check landed inside the margin.

kind of workverdictrouted pickstored qualitycanary
agentic-tool-useinconclusiveor-ling-3.0-flash0.886 ± 0.136 n=0
classificationheldname withheld0.976 ± 0.0331.000 n=4
code-genheldname withheld0.979 ± 0.0190.944 n=4
code-reviewheldor-deepseek-v4-flash-07310.869 ± 0.0890.750 n=4
creativeheldor-deepseek0.850 ± 0.0710.800 n=4
extractionheldname withheld0.962 ± 0.0171.000 n=4
multi-step-reasoningheldor-ling-3.0-flash0.960 ± 0.0551.000 n=4
rag-answerheldname withheld0.907 ± 0.0780.750 n=4
rewrite-editheldor-gpt-mini0.866 ± 0.0460.975 n=4
summarizationheldor-ling-3.0-flash0.900 ± 0.0500.925 n=4

Nine frontiers held their positions. Zero moved to a different model. One check was inconclusive. A canary is a small weekly re-check of four examples that detects whether a model has collapsed, not whether it has shifted by a single percentage point. Every quality figure comes with a margin of error (a 95% confidence interval, meaning we are 95% sure the true score falls within that range). This week or-deepseek-v4-flash-0731 held code-review work at 0.869, give or take 0.089. Or-deepseek held creative writing at 0.850, give or take 0.071. Or-ling-3.0-flash held multi-step reasoning at 0.960, give or take 0.055, and summarization at 0.900, give or take 0.050.

Auditions

An audition is a newly released model's first exam, on the kind of work it looks suited to. It earns a place only by beating the model already doing that work on quality or price.

Three newly released models received auditions (a first exam to see if they beat the current frontier pick). Or-ling-2-6-flash was tested on rewrite-edit work in the multilingual lane and did not beat the incumbent. Or-l3-lunaris-8b and or-nex-n2-mini were both tested on classification work in the small/cheap lane. Neither beat the incumbent.

Mixing

Mixing means using two or three cheaper models together in a particular way instead of one expensive one. We score combinations by replaying results we already have, so this costs nothing to explore. How a combination works is part of the product and is not published; what it achieves is.

Five adjudications, one verdict. Having a referee model pick between two answers realizes about a quarter of the theoretical gain, because referees mispick. Picking by executing tests fixes the mispicks but pays for every model on every request — at the expensive end the combination cost more than the champion it chased, and at the cheap end a second model multiplies the bill six-fold against a margin inside measurement noise. Escalating from a cheap model only when it seems unsure fails differently: models are confident exactly when they are wrong, so the gate misses the errors and the escalations it does fire often trade a right answer away. Where combinations DO win is across requests, not within one: sending each kind of work to the model measured best for its price. That is what this product does.

What it means for you

You almost certainly do not need the most expensive model for most of your work, and you do not need exotic model combinations either. The measured gap between the best model and the best value model is about a point on our hardest exams, and the price gap is up to twenty-three times. Routing each kind of work to the cheapest model that clears your bar captures that spread with receipts. We publish our failures so you can trust that claim: settled for the task shapes we measure, at the current model market — and our weekly saturation checks will say so publicly if the market changes shape.

How the numbers are made

Every experiment was pre-registered: the target, the shapes, and the success bar were written to an internal ledger before any measurement ran. Comparisons are like-for-like — same run, same items, fresh baselines, never cached aggregates. All spend is capped and ledgered; two runs were refused by our own budget machinery and re-scoped. Combination mechanics (how selectors and gates work) are part of the product and are not published; what they achieved, and failed to achieve, is.

Numbers

10
canaries run
9
frontiers held
0
frontiers drifted
36
items graded
326
listings screened
3
new models measured
1
inconclusive
$1.42
measurement spend

Questions

Does this mean model combinations are useless?
Within one request, at today’s model market, on the task shapes we measure: yes, we could not make them pay. Across requests they are the whole product — each kind of work goes to the model measured best for its price. If the market changes shape, our weekly checks will catch it and we will re-open the question publicly.
Why publish negative results?
Because the product is trust in measurement. A router that only ever reports wins is indistinguishable from a router that does not measure. Five honest losses for about thirteen dollars is the cheapest credibility we will ever buy.
What should I do with this?
Ask what your current model bill would be if every kind of work went to the cheapest model that clears your quality bar. That number is measurable on your own traffic, with a receipt on every answer.

Glossary. A frontier is the short list of models that are the best deal at their level of quality: nothing else is both better and cheaper. A floor is the lowest exam score you are willing to accept. A margin (or interval) is how far the true score could sit from the measured one, because an exam is a sample. A canary is a small weekly re-check. See the docs and the evidence.

What does this look like on my workload?

Measure this on your workload →
We spent two weeks trying to beat single models with model combinations. Single models won. · Frontier Notes