Code generation: a model at 1/271st the price of the best scorer, at 97.8% of its quality.

inkling-small vs ling-3.0-flash for RAG answering

the measured answer

As of August 24, 2026, inkling-small measures 0.982 on rag question answering vs 0.963 for ling-3.0-flash (+0.019), at 14.0× the price ($0.0937 vs $0.0067 per 1K requests).

same suite, same items, same scoring — frontier v4, measured 2026-08-24 · method

modelvendormeasured quality$ / 1K requestsp95 latency
inkling-smallthinkingmachines0.982$0.09371653 ms
ling-3.0-flashinclusionai0.963$0.00671660 ms

In retrieval-augmented answering, the model’s job narrows to faithful synthesis over supplied context — a different skill from open-ended knowledge, and one where paying for a frontier model’s world knowledge is often paying twice. The measured suite scores answers against references derived from the supplied documents.

These two are part of a larger measured frontier — what is the best model for rag answers shows every measured option for this workload, and other workloads rank these models differently: a model that wins here can lose on another kind of work, which is the whole argument for routing per workload rather than picking one model for everything.

the routed alternative

Potion routes each request to the cheapest option measured at your quality bar — including picks this public page does not name. Get an API key or read the docs.

inkling-small vs ling-3.0-flash for RAG answering — measured · Frontier Notes