gemini-2.5-flash vs grok-4.6 for code generation
As of August 24, 2026, grok-4.6 measures 1.000 on code generation vs 0.994 for gemini-2.5-flash (+0.006), at 12.4× the price ($6.85 vs $0.5537 per 1K requests).
same suite, same items, same scoring — frontier v5, measured 2026-08-24 · method
| model | vendor | measured quality | $ / 1K requests | p95 latency |
|---|---|---|---|---|
| gemini-2.5-flash | 0.994 | $0.5537 | 2372 ms | |
| grok-4.6 | x-ai | 1.000 | $6.85 | 46055 ms |
Code generation is the easiest workload to measure honestly — generated code either passes executable tests or it does not — and the hardest to generalize about, because "coding" spans one-shot functions to repository-scale agents. The frontier below is measured by executing every answer against held-out test suites: fractional credit per test, no judge, references gated on passing their own tests.
These two are part of a larger measured frontier — what is the best model for code generation shows every measured option for this workload, and other workloads rank these models differently: a model that wins here can lose on another kind of work, which is the whole argument for routing per workload rather than picking one model for everything.
Potion routes each request to the cheapest option measured at your quality bar — including picks this public page does not name. Get an API key or read the docs.