Code generation: a model at 1/271st the price of the best scorer, at 97.8% of its quality.

What is the best model for tool calling?

the measured answer

As of August 24, 2026, the cheapest live-measured option within 90% of top quality on agentic tool use is ling-3.0-flash at $0.1094 per 1,000 requests — 0.886 measured quality vs 0.972 at the top, with a 404× price spread across the measured frontier (frontier v4).

re-measured weekly · how these numbers are made

Agent stacks live or die on tool calls: the right function, the right arguments, the discipline not to invent either. The measured suite scores tool-call validity deterministically. Whole-journey agent completion — plan, call, read, recover — is measured separately, because call-level scores can all look healthy while the journey fails.

The measured frontier

Every row is a live measurement on Potion's held-out agentic tool use suite — same items, same scoring, per option. Names that are part of the product are withheld; their numbers are not.

optionvendormeasured quality$ / 1K requestsp95 latency
ling-3.0-flashinclusionai0.886$0.109448468 ms
Potion's routed pick (name withheld)0.950$0.144939997 ms
gemini-3.7-flashgoogle0.950$3.6312489 ms
gpt-5.6-terra-proopenai0.972$44.1534671 ms

frontier v4 · measured 2026-08-24 · live provider calls only — simulated evidence never appears on this page

Head to head

gemini-3.7-flash vs ling-3.0-flashgpt-5.6-terra-pro vs ling-3.0-flashgemini-3.7-flash vs gpt-5.6-terra-pro

Questions

How is tool calling measured?

Deterministically: full credit for the expected tool with the expected arguments, partial for the right tool misconfigured, zero for prose or a wrong tool. No judge.

Is the best chat model also the best agent model?

Measurably not always — function-calling reliability separates from prose quality, which is why agents deserve a routed pick rather than inheriting the chat default.

What about the whole agent journey?

End-to-end journeys (multi-step, output feeding the next step, only the final artifact scored) are a separate instrument — the methodology page describes it, and journey results appear in the weekly issues.

the routed alternative

Potion routes each request to the cheapest option measured at your quality bar — with a receipt on every answer and this page's evidence behind every pick. Get an API key or read the docs.

What is the best model for tool calling? Measured August 2026 · Frontier Notes