Code generation: a model at 1/271st the price of the best scorer, at 97.8% of its quality.

What is the best model for summarization?

the measured answer

As of August 24, 2026, the cheapest live-measured option within 90% of top quality on summarization is ling-3.0-flash at $0.0345 per 1,000 requests — 0.900 measured quality vs 0.983 at the top, with a 297× price spread across the measured frontier (frontier v4).

re-measured weekly · how these numbers are made

Summarization quality is subjective at the margins, which makes it exactly the workload where judge-based scores need calibration receipts. The measured suite anchors judging against references and calibrates judges against tasks with deterministic truth, publishing the correlation rather than asking for trust.

The measured frontier

Every row is a live measurement on Potion's held-out summarization suite — same items, same scoring, per option. Names that are part of the product are withheld; their numbers are not.

optionvendormeasured quality$ / 1K requestsp95 latency
ling-3.0-flashinclusionai0.900$0.03452651 ms
Potion's routed pick (name withheld)0.933$0.040310006 ms
muse-glimmer-30bmeta0.978$1.3420724 ms
kimi-k3moonshotai0.983$10.2651602 ms

frontier v4 · measured 2026-08-24 · live provider calls only — simulated evidence never appears on this page

Head to head

ling-3.0-flash vs muse-glimmer-30bkimi-k3 vs ling-3.0-flashkimi-k3 vs muse-glimmer-30b

Questions

Can summarization quality be measured at all?

Yes, carefully: reference-anchored judging (the judge sees a reference summary) calibrates near-perfectly against deterministic truth; reference-free judging does not, and is not used for these numbers.

Do I need a frontier model to summarize?

For routine document and thread summarization, measured mid-tier models are frequently within the interval of frontier ones. Long, dense, or high-stakes material shifts the answer — measure on your own traffic before trusting anyone’s general number.

What matters more, model or prompt?

Both, but the frontier is measured with each model given the same instruction under the same scoring — the ranking isolates the model. Your prompt can move absolute quality; it rarely reorders a wide measured gap.

the routed alternative

Potion routes each request to the cheapest option measured at your quality bar — with a receipt on every answer and this page's evidence behind every pick. Get an API key or read the docs.

What is the best model for summarization? Measured August 2026 · Frontier Notes