Code generation: a model at 1/271st the price of the best scorer, at 97.8% of its quality.

How the numbers are made

Every number Potion publishes — on the answers pages, in the weekly issues, on a customer's receipts — comes from one measurement discipline. This page is that discipline, stated plainly, including where it fails.

instruments

Held-out suites, hardened when they saturate

Each kind of work has a held-out suite: fixed task items with known correct answers or executable checks, authored privately so they cannot be memorized from the public internet. Models are measured on the same items under the same scoring. When multiple models ace a suite repeatedly, the suite has stopped discriminating — an automatic saturation alarm triggers authoring of harder items, because a perfect score on a saturated instrument is a statement about the instrument, not the model.

scoring

Deterministic first; judges only with calibration receipts

Wherever objective truth exists, scoring is deterministic: generated code executes against tests in a sandbox; extractions match fields against references; classifications match labels exactly; tool calls match expected functions and arguments. LLM judges are used only where a task is genuinely subjective — and every judge is calibrated against tasks with deterministic truth before its scores count. The calibrations are published, including the failures: reference-free judging of extractions is measurably blind to omissions, so it is not used.

uncertainty

Intervals that stay honest at the boundary

Every quality number carries a sample size and a 95% interval. Intervals use the generalized Jeffreys method, which stays honest at the boundary: a model that scores 42/42 is reported as “no failures observed in 42 tasks” — a lower bound with real width — never as certainty. A champion's crown is a ≥-bound.

cost

Cost measured from real usage, not list price

Cost per 1,000 requests comes from measured token usage on real provider calls at current prices — including hidden reasoning tokens, which make some models far more expensive per request than their per-token price suggests. Scoring spend is accounted separately so a measurement's own cost never contaminates a model's.

freshness

Re-measured weekly; drift caught by canaries

The frontiers behind these pages are re-checked weekly: small fresh canary runs on every routed pick, full re-measurement when the market moves, auditions for newly listed models. A published page updates when its measurement does, and shows its measurement date. A model that silently degrades behind an API is caught by the canary, not by a customer.

negatives

Negatives published with the same prominence

When a hypothesis fails, the failure is published: the multi-model mixing program closed with five pre-registered negative results, in public. A measurement lab that only reports wins is a marketing department.

scope

What these numbers are not

Suite measurements are a prior, not a verdict on your traffic. Production behavior is verified per customer: a learning period measures consenting customers' own requests, picks are re-verified on their data, and guarantees ride those measurements — not these pages. Names that are part of the product are withheld from public pages; their numbers are not.

the current answers: /answers · the weekly issues: /research · the product: withpotion.com

How the numbers are made — measurement methodology · Frontier Notes