Code generation: a model at 1/271st the price of the best scorer, at 97.8% of its quality.

All 10 routing frontiers held this week; l3-lunaris-8b earned a slot on classification.

In plain words

Potion keeps a scoreboard of AI models: how well each one does a kind of work, and what it costs. This week every model on the scoreboard was re-checked and none had got worse. 3 newly released models were tested; 2 made the scoreboard. No combination of cheaper models beat the best single model this week.

Week 37 of 2026: 10 drift canaries re-checked every routing frontier, 329 new catalogue listings were screened and 3 were measured. No frontier moved. Nothing in the mixing lane beat the best single model this week.

This week's frontiers

One row per kind of work. Routed pick is the model Potion currently sends that work to. Stored quality is its exam score when it was measured in full, give or take the margin. Canary is this week's small re-check: a few fresh tasks, scored the same way, to catch a model that has got worse. Held means the re-check landed inside the margin.

kind of workverdictrouted pickstored qualitycanary
agentic-tool-useheldling-3.0-flash0.886 ± 0.1360.750 n=4
classificationheldname withheld0.976 ± 0.0331.000 n=4
code-genheldname withheld0.979 ± 0.0190.975 n=4
code-reviewhelddeepseek-v4-flash-07310.869 ± 0.0890.750 n=4
creativehelddeepseek0.850 ± 0.0710.775 n=4
extractionheldname withheld0.962 ± 0.0170.969 n=4
multi-step-reasoningheldling-3.0-flash0.960 ± 0.0551.000 n=4
rag-answerheldname withheld0.907 ± 0.0781.000 n=4
rewrite-editheldgpt-mini0.866 ± 0.0460.950 n=4
summarizationheldling-3.0-flash0.900 ± 0.0500.875 n=4

Every routed pick reproduced its stored quality inside its interval. agentic-tool-use: 0.750 observed against 0.886 ± 0.136 stored; classification: 1.000 observed against 0.976 ± 0.033 stored; code-gen: 0.975 observed against 0.979 ± 0.019 stored.

Auditions

An audition is a newly released model's first exam, on the kind of work it looks suited to. It earns a place only by beating the model already doing that work on quality or price.

l3-lunaris-8b (small/cheap) was measured on classification and earned a frontier slot; nex-n2-mini (small/cheap) was measured on classification and earned a frontier slot; gpt-oss-20b (small/cheap) was measured on classification and did not beat the incumbent.

Mixing

Mixing means using two or three cheaper models together in a particular way instead of one expensive one. We score combinations by replaying results we already have, so this costs nothing to explore. How a combination works is part of the product and is not published; what it achieves is.

The mixing lane replays combinations of measured models over stored item-level results; nothing this week was both cheaper and as good as the best single model.

What it means for you

If you pay for AI by the request, the gap between the best model and the cheapest model that is good enough is where your money goes. Routing each request to the cheapest model that passed the exam is how you keep the quality and stop paying for the rest.

How the numbers are made

How these numbers are made, in plain words. We sort requests into kinds of work: sorting text into categories, pulling fields out of documents, writing code, answering from a set of documents, and so on. For each kind we keep a private exam of tasks the models have never seen. Every model sits the same exam. Code is marked by running it; other answers are marked against a reference answer. A model's quality is its average mark, and because an exam is a sample we also give a margin of error: two models whose margins overlap are called a tie. Cost is what a thousand requests would cost at the provider's public prices. A frontier is the short list of models that are the best deal at their level of quality, meaning nothing else is both better and cheaper. Each week we re-check every model on that list with a few fresh tasks to catch any that have got worse, and we give newly released models the exam for the work they look suited to.

Numbers

10
canaries run
10
frontiers held
0
frontiers drifted
40
items graded
329
listings screened
3
new models measured
0
inconclusive
$0.87
measurement spend

Questions

What is a routing frontier?
The set of model options for one kind of work that nothing else beats on quality, cost and latency at the same time. A routing policy picks from it: cheapest above a quality floor, fastest, or best quality.
Did any model's quality change this week?
No. All 10 frontier picks reproduced their stored quality inside the 95% interval on a 4-item canary.
Which measured model held its frontier pick this week?
ling-3.0-flash on agentic-tool-use (0.886 ± 0.136); deepseek-v4-flash-0731 on code-review (0.869 ± 0.089); deepseek on creative (0.850 ± 0.071). Some picks are withheld by name; their numbers are published.

Glossary. A frontier is the short list of models that are the best deal at their level of quality: nothing else is both better and cheaper. A floor is the lowest exam score you are willing to accept. A margin (or interval) is how far the true score could sit from the measured one, because an exam is a sample. A canary is a small weekly re-check. See the docs and the evidence.

What does this look like on my workload?

Measure this on your workload →
All 10 routing frontiers held this week; l3-lunaris-8b earned a slot on classification. · Frontier Notes