Code generation: a model at 1/271st the price of the best scorer, at 97.8% of its quality.

claude-opus-4.6 vs deepseek-chat-v3.1 for extraction

the measured answer

As of September 7, 2026, claude-opus-4.6 measures 0.962 on information extraction vs 0.950 for deepseek-chat-v3.1 (+0.012), at 19.9× the price ($3.62 vs $0.1824 per 1K requests).

same suite, same items, same scoring — frontier v5, measured 2026-09-07 · method

modelvendormeasured quality$ / 1K requestsp95 latency
claude-opus-4.6anthropic0.962$3.622723 ms
deepseek-chat-v3.1deepseek0.950$0.182411996 ms

Extraction — pulling fields out of documents, emails, and records into structured JSON — is scored field-by-field against known references, so the numbers below are deterministic, not judged. It is also the workload where reference-free LLM judging fails measurably (judges cannot see omissions), which is why extraction claims anywhere should be treated with suspicion unless the scoring method is stated.

These two are part of a larger measured frontier — what is the best model for data extraction shows every measured option for this workload, and other workloads rank these models differently: a model that wins here can lose on another kind of work, which is the whole argument for routing per workload rather than picking one model for everything.

the routed alternative

Potion routes each request to the cheapest option measured at your quality bar — including picks this public page does not name. Get an API key or read the docs.

claude-opus-4.6 vs deepseek-chat-v3.1 for extraction — measured · Frontier Notes