Code generation: a model at 1/271st the price of the best scorer, at 97.8% of its quality.

gemini-3.7-flash vs inkling-small for multi-step reasoning

the measured answer

As of September 8, 2026, gemini-3.7-flash and inkling-small measure within 0.005 quality of each other on multi-step reasoning (1.000 vs 1.000) — but inkling-small costs $0.3450/1K requests vs $1.13, 3.3× less for the same measured result.

same suite, same items, same scoring — frontier v6, measured 2026-09-08 · method

modelvendormeasured quality$ / 1K requestsp95 latency
gemini-3.7-flashgoogle1.000$1.133887 ms
inkling-smallthinkingmachines1.000$0.34505181 ms

Multi-step reasoning is the workload frontier pricing was built on — and the one where the "do I need it" question has a measurable answer. The suite scores multi-step problems with deterministic ends, so a model’s reasoning is graded on reaching the right result, not on showing appealing work.

These two are part of a larger measured frontier — when do you need a reasoning model shows every measured option for this workload, and other workloads rank these models differently: a model that wins here can lose on another kind of work, which is the whole argument for routing per workload rather than picking one model for everything.

the routed alternative

Potion routes each request to the cheapest option measured at your quality bar — including picks this public page does not name. Get an API key or read the docs.

gemini-3.7-flash vs inkling-small for multi-step reasoning — measured · Frontier Notes