Code generation: a model at 1/271st the price of the best scorer, at 97.8% of its quality.
Measured model routing

Your inference is unique. Your router should be too.

Potion builds around your actual workload, quality bar, and economics — and serves every request at the lowest price the evidence allows.

  • Lower bills, measured not promised. Each request goes to the cheapest model measured good enough.
  • Your rule: cost, quality, or speed. Picked from the measured Pareto frontier.
  • New models earn their place. Measured against your bar before they ever serve.
  • One line of code. A receipt with every answer.
one request · two roads
A GATEWAY · ONE MODEL FOR EVERYTHINGPOTION · THE MEASURED FIELDsorting textwriting codeextracting fieldsanswering from documentsmulti-step reasoningwriting with flairthe onemodelEVERY TIME · $0.3580 / 1kREQUESTsorting textCOST →Q ↑floor 0.95or-██████ · $0.0041 / 1k · scores 0.97cheapest point above the floor · 98% under the one modelCOST / 1K$0.3580QUALITY1.00COST / 1K$0.0041QUALITY0.9798%
tallies are means over the requests served so far · same work, same quality bar, two roads
Figure 1. The same requests through a gateway (one model for everything) and through Potion (the measured field for that kind of work, a quality floor, the cheapest point that clears it). Every point and price is a committed measurement; the tallies are means over the requests served so far.
01
The obvious question

Isn't this what OpenRouter does?

Less than it used to be — gateways now ship auto-routers, and that is exactly the point. Prediction is becoming free. Others predict which model should work. Potion measures what actually clears your bar — and signs the receipt.

You get
gateway · Every model, one API
The measured pick for each kind of work
Who chooses
gateway · You, or a predictive auto-router
Held-out measurement, under your stated bar
Based on
gateway · Leaderboards and predictions
Held-out tests of your kind of work
Quality
gateway · Whatever the prediction picked
Only points measured above your floor are served
After the answer
gateway · Tokens and a price
A receipt: what served it, and why
A new model ships
gateway · You re-evaluate by hand
Measured first, adopted only if it earns it
Figure 2. What a gateway answers and what Potion answers, row by row. The gateway column is factual about what gateways do well; the difference is judgment per request, backed by held-out measurement.

Access stopped being scarce the day gateways shipped. Judgment — measured per kind of work, stated with its error bars, enforced as a floor — is the scarce layer. That layer is Potion, and it works the same over any gateway or provider underneath.

02
What you get

A router with your name on it.

Potion compiles your router from three things: your traffic’s actual kinds of work, your quality bar, and the measured frontiers. It ships as a model id — potion/your-org — and as a document you can read: every kind of work, what it routes to, the measured quality and price behind the choice.

It is versioned. When a new model ships or a measurement moves, Potion recompiles, mints the next version, and writes what changed on it — and every receipt names the version that served it. Roll it back, audit it, set boundaries on it.

You never hand-assign a model. You change what you want — the floor, the ceiling, a ban — and Potion recompiles. The intelligence stays on our side of the API.

potion/your-orgv4
kind of workroutes toquality$/1K
extraction·small0.977$0.05
code-gen·mid0.917$0.55
summarization·small0.850$0.01
rag-answer·mid0.941$0.12

v3 → v4 · a new model cleared the bar on extraction — estimated at your mix: saves another 11%

the artifact, schematically — strategies masked, numbers illustrative. Real ones are compiled per organization, from live measurements, on your Router page.
03 · Measured, not claimed

The same work.
A 271× price range.

Five models writing code to specification — scored by running their code, not by opinion. The quality difference across this table is 2.1 points in a hundred. The price difference is 271-fold.

This is why routing pays: most kinds of work are served from the bottom row, a few genuinely need the top one, and only a measurement can tell them apart.

measured 2026-08-21 · frontier v4 · retrieval-hostile suite · scored by execution · error bars on the full table in the docs

or-grok-4.6xAIquality 1.000 · $6.2568/1k
or-gemini-flashGooglequality 0.996 · $0.5506/1k
or-gpt-miniOpenAIquality 0.990 · $0.2509/1k
or-deepseekDeepSeekquality 0.980 · $0.1904/1k
or-████████████name withheldquality 0.979 · $0.0231/1k

↑ the routed pick — 97.8% of the top row's quality at 1/271st the price. The name? That's the product.

Figure 3. The code-gen-hard frontier, measured 2026-08-21, 30 items scored by execution (90 graded runs per point). Bars show cost per 1,000 requests · teal = what a 0.95 quality floor actually buys · 2 further frontier points measured below that floor and are not drawn

04
The research engine

The frontier is measured weekly, by a machine. Negatives included.

Every model Potion considers routing to must first earn its place through measurement — the same private exams, per kind of work, with confidence intervals and dates on every point. New models are auditioned the week they ship. Only what dominates on quality, cost and speed at once is published to the frontier your requests are routed from.

The engine also tests the clever ideas, so you never pay for one that does not work. It spent two weeks trying to beat single-model routing with multi-model combinations — pre-registered, budget-capped — and every attempt lost to the best single model. We published all five losses. If the model market ever changes shape so a combination pays, the same machinery will find it, measure it, and only then serve it.

Why this is hard to copy: the measured corpus, the live receipts, and the replay engine live in one place, and the corpus compounds every week — including the negatives. A router that only reports wins is indistinguishable from a router that does not measure.

the research loop · runs weekly · no one in it
measurereplayauditionpublishroutereceipts feed the next measurement
what compounds: item-level results, every model and combination, weekly, with intervals and dates on every point · a combination is one hash — what it is made of is not published
Figure 4. The weekly loop, with no one in it: measure the models that could earn a route; replay candidates against stored results at no cost; audition new models the week they ship; publish what dominates — and publish what lost; route live traffic with receipts, which feed the next measurement.
05
How you know we are not making this up

The actual map. Drive it yourself.

Measured options for one kind of work, error bars included. Pick a rule, drag the slider, and you are running the same selection the router runs in production — including its refusal to answer when nothing measured qualifies.

multi-step-reasoning · v3 · n = 50 per point
0.40.50.60.70.80.91.0$0.01$0.1$1COST PER 1,000 REQUESTS · LOGMEASURED QUALITYfloor 0.60or-████████████ — q 0.50 ±0.14 · $0.0093/1k · p95 7,160 msor-deepseek-v4-flash-0731 — q 0.98 ±0.04 · $0.0304/1k · p95 9,702 msor-nemotron-3.5-lightning — q 0.76 ±0.12 · $0.0845/1k · p95 2,739 msor-gpt-full — q 0.56 ±0.14 · $0.1482/1k · p95 1,126 msor-gemini-flash — q 0.66 ±0.13 · $0.1656/1k · p95 2,293 msor-inkling-small — q 0.92 ±0.08 · $0.1906/1k · p95 4,009 msor-gemini-3.7-flash — q 1.00 ±0.00 · $0.2929/1k · p95 3,242 msor-inkling — q 0.94 ±0.07 · $0.5657/1k · p95 2,449 msor-kimi-k3 — q 0.96 ±0.05 · $1.65/1k · p95 2,422 msor-deepseek-v4-flash-07310.98 ± 0.04 · $0.0304/1k · p95 9,702 mson the frontierdominated95% interval
selected or-deepseek-v4-flash-0731quality 0.98 ± 0.04cost $0.0304 / 1kp95 9,702 ms
x-frontier-trace: cluster=multi-step-reasoning;strategy=1fd419ee;frontier=v3;policy=min_cost;fallback=0;provenance=live

Real measured points, quoted from the committed frontier; hover any point for its name and numbers. The cheapest row costs under a cent per 1k and measures 0.50, a coin flip, which is why the router will not send reasoning work there: cheap only wins where the measurement clears your floor. Try latency_bound at 2,500 ms; the answer changes.

Figure 5. The multi-step-reasoning frontier, version 3: nine measured options, 50 items each, 95% intervals drawn. The selection you run here is the selection the router runs in production, including its refusal when nothing qualifies.
06
The engineering

Built to refuse before it is built to answer.

A router that spends your money has to be trustworthy before it is clever. So the serving path is written to fail closed: the server will not start if a frontier names a model it cannot serve; a frontier cannot be republished if it regresses; a request is refused before a budget is crossed, not after; and nothing is ever served from a number that was not measured.

Priced like the incentives should be: model costs pass through at cost, and Potion earns a share of the savings your own receipts verify — if it saves you nothing, it earns nothing above cost. No minimum. You bring no provider accounts and no keys: Potion buys from every provider at once, which is also what lets it route across the whole market rather than the one account you happened to open.

what is true of every request, today
A receipt on every answer

Kind of work, strategy, policy, provenance: what answered and why, on the response itself.

Fails closed

Boot refuses a frontier that names an unservable model. A regression cannot be published. No measurement, no route.

Budgets that stop the request

A hard cap is enforced before the money is spent, within seconds of being crossed, never after.

Intervals and dates on every point

Each frontier point carries its sample size, its 95% interval, and the day it was measured.

Re-checked every week

Every routed pick is re-measured on fresh tasks; a model that drifts is caught before it costs you.

One line, the whole protocol

OpenAI chat completions, including streaming and tool calls. Point your client at Potion; keep everything else.

Second opinion

Have an assistant read this site and argue with it.

One click opens a chat with the question already written: what Potion does, and what you should check before believing any of it.

The question is copied to your clipboard too, since Gemini will not take it from a link.

Change one line. Keep the receipts.

Point a client at Potion and watch the routing decisions arrive with the answers.