Code generation: a model at 1/271st the price of the best scorer, at 97.8% of its quality.

The last 8.7 points of agentic tool use quality cost 404×.

Potion Research

On Potion's measured agentic-tool-use suite, one model scored 0.972 at $44.15 per thousand requests while another scored 0.886 at $0.1094. Both figures come from the same held-out agentic tool use suite, 14 scored items, measured the same way — so the comparison is like for like.

The premium buys 8.7 quality points on this suite. Whether that is worth 404× the price depends on what a failure costs you: at high volume with a cheap retry, it usually is not; on work that ships unreviewed, it can be.

Someone is paying the difference. If that is you, the question worth asking is which of these two numbers your workload actually needs — and this is the measurement that answers it.

3 measurement cycles ran in the last 24 hours, 3 newly listed models were measured, no recipe reached a frontier.

The current numbers live on the measured answers; the method is public.

The last 8.7 points of agentic tool use quality cost 404×. · Frontier Notes