What is the best model for summarization?
As of August 24, 2026, the cheapest live-measured option within 90% of top quality on summarization is ling-3.0-flash at $0.0345 per 1,000 requests — 0.900 measured quality vs 0.983 at the top, with a 297× price spread across the measured frontier (frontier v4).
re-measured weekly · how these numbers are made
Summarization quality is subjective at the margins, which makes it exactly the workload where judge-based scores need calibration receipts. The measured suite anchors judging against references and calibrates judges against tasks with deterministic truth, publishing the correlation rather than asking for trust.
The measured frontier
Every row is a live measurement on Potion's held-out summarization suite — same items, same scoring, per option. Names that are part of the product are withheld; their numbers are not.
| option | vendor | measured quality | $ / 1K requests | p95 latency |
|---|---|---|---|---|
| ling-3.0-flash | inclusionai | 0.900 | $0.0345 | 2651 ms |
| Potion's routed pick (name withheld) | — | 0.933 | $0.0403 | 10006 ms |
| muse-glimmer-30b | meta | 0.978 | $1.34 | 20724 ms |
| kimi-k3 | moonshotai | 0.983 | $10.26 | 51602 ms |
frontier v4 · measured 2026-08-24 · live provider calls only — simulated evidence never appears on this page
Head to head
ling-3.0-flash vs muse-glimmer-30bkimi-k3 vs ling-3.0-flashkimi-k3 vs muse-glimmer-30b
Questions
Can summarization quality be measured at all?
Yes, carefully: reference-anchored judging (the judge sees a reference summary) calibrates near-perfectly against deterministic truth; reference-free judging does not, and is not used for these numbers.
Do I need a frontier model to summarize?
For routine document and thread summarization, measured mid-tier models are frequently within the interval of frontier ones. Long, dense, or high-stakes material shifts the answer — measure on your own traffic before trusting anyone’s general number.
What matters more, model or prompt?
Both, but the frontier is measured with each model given the same instruction under the same scoring — the ranking isolates the model. Your prompt can move absolute quality; it rarely reorders a wide measured gap.
Potion routes each request to the cheapest option measured at your quality bar — with a receipt on every answer and this page's evidence behind every pick. Get an API key or read the docs.