What is the best model for tool calling?
As of August 24, 2026, the cheapest live-measured option within 90% of top quality on agentic tool use is ling-3.0-flash at $0.1094 per 1,000 requests — 0.886 measured quality vs 0.972 at the top, with a 404× price spread across the measured frontier (frontier v4).
re-measured weekly · how these numbers are made
Agent stacks live or die on tool calls: the right function, the right arguments, the discipline not to invent either. The measured suite scores tool-call validity deterministically. Whole-journey agent completion — plan, call, read, recover — is measured separately, because call-level scores can all look healthy while the journey fails.
The measured frontier
Every row is a live measurement on Potion's held-out agentic tool use suite — same items, same scoring, per option. Names that are part of the product are withheld; their numbers are not.
| option | vendor | measured quality | $ / 1K requests | p95 latency |
|---|---|---|---|---|
| ling-3.0-flash | inclusionai | 0.886 | $0.1094 | 48468 ms |
| Potion's routed pick (name withheld) | — | 0.950 | $0.1449 | 39997 ms |
| gemini-3.7-flash | 0.950 | $3.63 | 12489 ms | |
| gpt-5.6-terra-pro | openai | 0.972 | $44.15 | 34671 ms |
frontier v4 · measured 2026-08-24 · live provider calls only — simulated evidence never appears on this page
Head to head
gemini-3.7-flash vs ling-3.0-flashgpt-5.6-terra-pro vs ling-3.0-flashgemini-3.7-flash vs gpt-5.6-terra-pro
Questions
How is tool calling measured?
Deterministically: full credit for the expected tool with the expected arguments, partial for the right tool misconfigured, zero for prose or a wrong tool. No judge.
Is the best chat model also the best agent model?
Measurably not always — function-calling reliability separates from prose quality, which is why agents deserve a routed pick rather than inheriting the chat default.
What about the whole agent journey?
End-to-end journeys (multi-step, output feeding the next step, only the final artifact scored) are a separate instrument — the methodology page describes it, and journey results appear in the weekly issues.
Potion routes each request to the cheapest option measured at your quality bar — with a receipt on every answer and this page's evidence behind every pick. Get an API key or read the docs.