What is the best model for text classification?
As of August 24, 2026, the cheapest live-measured option within 90% of top quality on classification is Potion's routed pick (name withheld) at $0.0042 per 1,000 requests — 0.976 measured quality vs 1.000 at the top, with a 85× price spread across the measured frontier (frontier v5).
re-measured weekly · how these numbers are made
Classification — routing tickets, labeling intents, applying rules to text — is the workload where model choice matters most and is judged least. The task has a right answer, so quality is measurable exactly, and the measured gap between the cheapest adequate model and a frontier model is routinely the widest of any workload. Paying frontier prices here is the most common form of silent AI overspend.
The measured frontier
Every row is a live measurement on Potion's held-out classification suite — same items, same scoring, per option. Names that are part of the product are withheld; their numbers are not.
| option | vendor | measured quality | $ / 1K requests | p95 latency |
|---|---|---|---|---|
| Potion's routed pick (name withheld) | — | 0.976 | $0.0042 | 3489 ms |
| gemini-2.5-flash | 0.988 | $0.0270 | 1080 ms | |
| gemini-3.7-flash | 1.000 | $0.3580 | 3367 ms |
frontier v5 · measured 2026-08-24 · live provider calls only — simulated evidence never appears on this page
Head to head
gemini-2.5-flash vs gemini-3.7-flash
Questions
Do I need a frontier model for classification?
Usually not. On measured classification suites, small models regularly match frontier models within the confidence interval. The honest answer depends on your label set and edge cases — which is why these numbers are re-measured weekly rather than asserted once.
How is classification quality measured here?
Deterministically: each suite item has a known correct label, and a model scores by exact agreement — no LLM judge anywhere in the loop. Scores carry sample sizes and 95% intervals.
What if my labels are domain-specific?
Generic benchmarks are a prior, not a verdict. Potion measures consenting customers’ own traffic during a learning period and re-verifies the pick on their labels before it guarantees anything.
Potion routes each request to the cheapest option measured at your quality bar — with a receipt on every answer and this page's evidence behind every pick. Get an API key or read the docs.