Code generation: a model at 1/271st the price of the best scorer, at 97.8% of its quality.

kimi-k3 vs ling-3.0-flash for summarization

the measured answer

As of August 24, 2026, kimi-k3 measures 0.983 on summarization vs 0.900 for ling-3.0-flash (+0.083), at 297.1× the price ($10.26 vs $0.0345 per 1K requests).

same suite, same items, same scoring — frontier v4, measured 2026-08-24 · method

modelvendormeasured quality$ / 1K requestsp95 latency
kimi-k3moonshotai0.983$10.2651602 ms
ling-3.0-flashinclusionai0.900$0.03452651 ms

Summarization quality is subjective at the margins, which makes it exactly the workload where judge-based scores need calibration receipts. The measured suite anchors judging against references and calibrates judges against tasks with deterministic truth, publishing the correlation rather than asking for trust.

These two are part of a larger measured frontier — what is the best model for summarization shows every measured option for this workload, and other workloads rank these models differently: a model that wins here can lose on another kind of work, which is the whole argument for routing per workload rather than picking one model for everything.

the routed alternative

Potion routes each request to the cheapest option measured at your quality bar — including picks this public page does not name. Get an API key or read the docs.

kimi-k3 vs ling-3.0-flash for summarization — measured · Frontier Notes