Code generation: a model at 1/271st the price of the best scorer, at 97.8% of its quality.

What is the best model for RAG answers?

the measured answer

As of August 24, 2026, the cheapest live-measured option within 90% of top quality on rag question answering is Potion's routed pick (name withheld) at $0.0058 per 1,000 requests — 0.907 measured quality vs 0.982 at the top, with a 16× price spread across the measured frontier (frontier v4).

re-measured weekly · how these numbers are made

In retrieval-augmented answering, the model’s job narrows to faithful synthesis over supplied context — a different skill from open-ended knowledge, and one where paying for a frontier model’s world knowledge is often paying twice. The measured suite scores answers against references derived from the supplied documents.

The measured frontier

Every row is a live measurement on Potion's held-out rag question answering suite — same items, same scoring, per option. Names that are part of the product are withheld; their numbers are not.

optionvendormeasured quality$ / 1K requestsp95 latency
Potion's routed pick (name withheld)0.907$0.00584864 ms
ling-3.0-flashinclusionai0.963$0.00671660 ms
inkling-smallthinkingmachines0.982$0.09371653 ms

frontier v4 · measured 2026-08-24 · live provider calls only — simulated evidence never appears on this page

Head to head

inkling-small vs ling-3.0-flash

Questions

Does a bigger model help when the context contains the answer?

Less than expected: the measured gap narrows sharply when answers must come from supplied context. Faithfulness, not knowledge, dominates — and is measurable.

How is faithfulness measured?

Against document-derived references: an answer scores on agreement with what the supplied context supports, so a fluent hallucination scores as the miss it is.

Should retrieval and answering use the same model?

They are different workloads; route them separately. Embedding and answering frontiers are measured independently.

the routed alternative

Potion routes each request to the cheapest option measured at your quality bar — with a receipt on every answer and this page's evidence behind every pick. Get an API key or read the docs.

What is the best model for RAG answers? Measured August 2026 · Frontier Notes