Code generation: a model at 1/271st the price of the best scorer, at 97.8% of its quality.

When do you need a reasoning model?

the measured answer

As of August 24, 2026, the cheapest live-measured option within 90% of top quality on multi-step reasoning is ling-3.0-flash at $0.0158 per 1,000 requests — 0.960 measured quality vs 1.000 at the top, with a 19× price spread across the measured frontier (frontier v4).

re-measured weekly · how these numbers are made

Multi-step reasoning is the workload frontier pricing was built on — and the one where the "do I need it" question has a measurable answer. The suite scores multi-step problems with deterministic ends, so a model’s reasoning is graded on reaching the right result, not on showing appealing work.

The measured frontier

Every row is a live measurement on Potion's held-out multi-step reasoning suite — same items, same scoring, per option. Names that are part of the product are withheld; their numbers are not.

optionvendormeasured quality$ / 1K requestsp95 latency
Potion's routed pick (name withheld)0.481$0.00917160 ms
ling-3.0-flashinclusionai0.960$0.015812194 ms
deepseek-v4-flash-0731deepseek0.982$0.02599702 ms
gemini-3.7-flashgoogle1.000$0.29383242 ms

frontier v4 · measured 2026-08-24 · live provider calls only — simulated evidence never appears on this page

Head to head

deepseek-v4-flash-0731 vs ling-3.0-flashgemini-3.7-flash vs ling-3.0-flashdeepseek-v4-flash-0731 vs gemini-3.7-flash

Questions

When is a reasoning model worth it?

When the measured gap on your problem class exceeds the interval — which happens on genuinely multi-step, error-compounding work, and largely does not on retrieval, extraction, or classification dressed up as reasoning.

How is reasoning scored without a judge?

By final-answer correctness on problems with known results, plus journey-grain suites where each step feeds the next and only the end artifact is scored.

Do reasoning tokens change the cost math?

Substantially — reasoning models spend hidden tokens, so their cost per request is measured from real usage, not list price per token. The frontier below prices what requests actually cost.

the routed alternative

Potion routes each request to the cheapest option measured at your quality bar — with a receipt on every answer and this page's evidence behind every pick. Get an API key or read the docs.

When do you need a reasoning model? Measured August 2026 · Frontier Notes