Code generation: a model at 1/271st the price of the best scorer, at 97.8% of its quality.

What is the best model for data extraction?

the measured answer

As of August 24, 2026, the cheapest live-measured option within 90% of top quality on information extraction is Potion's routed pick (name withheld) at $0.0238 per 1,000 requests — 0.962 measured quality vs 0.985 at the top, with a 24× price spread across the measured frontier (frontier v4).

re-measured weekly · how these numbers are made

Extraction — pulling fields out of documents, emails, and records into structured JSON — is scored field-by-field against known references, so the numbers below are deterministic, not judged. It is also the workload where reference-free LLM judging fails measurably (judges cannot see omissions), which is why extraction claims anywhere should be treated with suspicion unless the scoring method is stated.

The measured frontier

Every row is a live measurement on Potion's held-out information extraction suite — same items, same scoring, per option. Names that are part of the product are withheld; their numbers are not.

optionvendormeasured quality$ / 1K requestsp95 latency
Potion's routed pick (name withheld)0.962$0.023814119 ms
deepseek-chat-v3.1deepseek0.962$0.335513033 ms
inkling-smallthinkingmachines0.985$0.574715156 ms

frontier v4 · measured 2026-08-24 · live provider calls only — simulated evidence never appears on this page

Head to head

deepseek-chat-v3.1 vs inkling-small

Questions

How is extraction quality measured?

Field-match against a known reference: every extracted field either matches or it does not, with fractional credit per document. No judge model is involved; a missing field cannot hide.

Can a small model really do production extraction?

On measured suites, the gap between the best cheap model and the best frontier model is often inside the confidence interval. Real documents are messier than any suite — which is why the measured answer is a starting point and customer-traffic verification is the finish line.

What about PDFs, scans, and images?

Vision extraction is measured separately, on rendered documents. The frontier below is text extraction; vision numbers appear in the weekly issues as they are re-measured.

the routed alternative

Potion routes each request to the cheapest option measured at your quality bar — with a receipt on every answer and this page's evidence behind every pick. Get an API key or read the docs.

What is the best model for data extraction? Measured August 2026 · Frontier Notes