We spent two weeks trying to beat single models with model combinations. Single models won.
We route each request to the cheapest model that is measured to clear a quality bar for that kind of work. An obvious question is whether we should go further: run several models on one request and pick the best answer. We tested that idea five ways, with pre-registered targets and hard budget caps. Every test lost to the best single model, and this issue says exactly why.
The strongest model on our hardest code exam never missed — twice, across every item. But a model at roughly a twenty-third of its price came within about one point of it. That gap is the entire space a combination of models could win, and it is smaller than the cost of refereeing the combination.
This week's frontiers
One row per kind of work. Routed pick is the model Potion currently sends that work to. Stored quality is its exam score when it was measured in full, give or take the margin. Canary is this week's small re-check: a few fresh tasks, scored the same way, to catch a model that has got worse. Held means the re-check landed inside the margin.
| kind of work | verdict | routed pick | stored quality | canary |
|---|---|---|---|---|
| agentic-tool-use | inconclusive | or-ling-3.0-flash | 0.886 ± 0.136 | — n=0 |
| classification | held | name withheld | 0.976 ± 0.033 | 1.000 n=4 |
| code-gen | held | name withheld | 0.979 ± 0.019 | 0.944 n=4 |
| code-review | held | or-deepseek-v4-flash-0731 | 0.869 ± 0.089 | 0.750 n=4 |
| creative | held | or-deepseek | 0.850 ± 0.071 | 0.800 n=4 |
| extraction | held | name withheld | 0.962 ± 0.017 | 1.000 n=4 |
| multi-step-reasoning | held | or-ling-3.0-flash | 0.960 ± 0.055 | 1.000 n=4 |
| rag-answer | held | name withheld | 0.907 ± 0.078 | 0.750 n=4 |
| rewrite-edit | held | or-gpt-mini | 0.866 ± 0.046 | 0.975 n=4 |
| summarization | held | or-ling-3.0-flash | 0.900 ± 0.050 | 0.925 n=4 |
Nine frontiers held their positions. Zero moved to a different model. One check was inconclusive. A canary is a small weekly re-check of four examples that detects whether a model has collapsed, not whether it has shifted by a single percentage point. Every quality figure comes with a margin of error (a 95% confidence interval, meaning we are 95% sure the true score falls within that range). This week or-deepseek-v4-flash-0731 held code-review work at 0.869, give or take 0.089. Or-deepseek held creative writing at 0.850, give or take 0.071. Or-ling-3.0-flash held multi-step reasoning at 0.960, give or take 0.055, and summarization at 0.900, give or take 0.050.
Auditions
An audition is a newly released model's first exam, on the kind of work it looks suited to. It earns a place only by beating the model already doing that work on quality or price.
Three newly released models received auditions (a first exam to see if they beat the current frontier pick). Or-ling-2-6-flash was tested on rewrite-edit work in the multilingual lane and did not beat the incumbent. Or-l3-lunaris-8b and or-nex-n2-mini were both tested on classification work in the small/cheap lane. Neither beat the incumbent.
- or-ling-2-6-flash · multilingual · rewrite-edit · did not beat the incumbent
- or-l3-lunaris-8b · small/cheap · classification · did not beat the incumbent
- or-nex-n2-mini · small/cheap · classification · did not beat the incumbent
Mixing
Mixing means using two or three cheaper models together in a particular way instead of one expensive one. We score combinations by replaying results we already have, so this costs nothing to explore. How a combination works is part of the product and is not published; what it achieves is.
Five adjudications, one verdict. Having a referee model pick between two answers realizes about a quarter of the theoretical gain, because referees mispick. Picking by executing tests fixes the mispicks but pays for every model on every request — at the expensive end the combination cost more than the champion it chased, and at the cheap end a second model multiplies the bill six-fold against a margin inside measurement noise. Escalating from a cheap model only when it seems unsure fails differently: models are confident exactly when they are wrong, so the gate misses the errors and the escalations it does fire often trade a right answer away. Where combinations DO win is across requests, not within one: sending each kind of work to the model measured best for its price. That is what this product does.
- code work · judge-picked pair · quality 0.940 · premium · n=84
- code work · execution-picked pair · quality 0.994 · premium · n=84
- writing work · picked pair · quality 0.936 · premium · n=28
- code work · confidence-escalated · quality 0.981 · mid · n=336
- extraction work · structure-checked pair · quality 0.940 · mid · n=144
What it means for you
You almost certainly do not need the most expensive model for most of your work, and you do not need exotic model combinations either. The measured gap between the best model and the best value model is about a point on our hardest exams, and the price gap is up to twenty-three times. Routing each kind of work to the cheapest model that clears your bar captures that spread with receipts. We publish our failures so you can trust that claim: settled for the task shapes we measure, at the current model market — and our weekly saturation checks will say so publicly if the market changes shape.
How the numbers are made
Every experiment was pre-registered: the target, the shapes, and the success bar were written to an internal ledger before any measurement ran. Comparisons are like-for-like — same run, same items, fresh baselines, never cached aggregates. All spend is capped and ledgered; two runs were refused by our own budget machinery and re-scoped. Combination mechanics (how selectors and gates work) are part of the product and are not published; what they achieved, and failed to achieve, is.
- Every quality figure is a mean over a retrieval-hostile suite with a bootstrap 95% interval; two points whose intervals overlap are reported as tied.
- Code-generation and code-review quality is scored by executing the code; other clusters are scored by a rubric against a reference answer.
- A canary is a small weekly sample (four items) against the stored measurement; it detects collapse, not one-point movement.
- Combinations of models are replayed from stored item-level results; agreement between models is modelled conservatively, so the combination figures understate rather than overstate.
- Free-tier listings are excluded from auditions: their quality is not stable enough to measure.
- The mixing verdict is scoped: settled for measured task shapes at the current model market; the weekly saturation alarm owns the reopening condition.
- Frontier and audition facts on this page are the 2026-W35 weekly readings, unchanged.
Numbers
Questions
- Does this mean model combinations are useless?
- Within one request, at today’s model market, on the task shapes we measure: yes, we could not make them pay. Across requests they are the whole product — each kind of work goes to the model measured best for its price. If the market changes shape, our weekly checks will catch it and we will re-open the question publicly.
- Why publish negative results?
- Because the product is trust in measurement. A router that only ever reports wins is indistinguishable from a router that does not measure. Five honest losses for about thirteen dollars is the cheapest credibility we will ever buy.
- What should I do with this?
- Ask what your current model bill would be if every kind of work went to the cheapest model that clears your quality bar. That number is measurable on your own traffic, with a receipt on every answer.
Glossary. A frontier is the short list of models that are the best deal at their level of quality: nothing else is both better and cheaper. A floor is the lowest exam score you are willing to accept. A margin (or interval) is how far the true score could sit from the measured one, because an exam is a sample. A canary is a small weekly re-check. See the docs and the evidence.
What does this look like on my workload?
Measure this on your workload →