What is the best model for creative writing?
As of August 24, 2026, the cheapest live-measured option within 90% of top quality on creative writing is deepseek-chat-v3.1 at $0.8281 per 1,000 requests — 0.850 measured quality vs 0.900 at the top, with a 9× price spread across the measured frontier (frontier v4).
re-measured weekly · how these numbers are made
Creative work is where measurement is hardest and overclaiming easiest. These numbers come from judge-scored suites whose judges are themselves calibrated against tasks with deterministic truth — and the calibration results are published, including the failures.
The measured frontier
Every row is a live measurement on Potion's held-out creative writing suite — same items, same scoring, per option. Names that are part of the product are withheld; their numbers are not.
| option | vendor | measured quality | $ / 1K requests | p95 latency |
|---|---|---|---|---|
| Potion's routed pick (name withheld) | — | 0.743 | $0.0611 | 25054 ms |
| deepseek-chat-v3.1 | deepseek | 0.850 | $0.8281 | 22494 ms |
| claude-sonnet-4.5 | anthropic | 0.900 | $7.66 | 9464 ms |
frontier v4 · measured 2026-08-24 · live provider calls only — simulated evidence never appears on this page
Head to head
claude-sonnet-4.5 vs deepseek-chat-v3.1
Questions
Can creative quality be measured honestly?
Within limits, and the limits are stated: calibrated judges, rubric-hashed scoring, intervals on every number. Where judging is unreliable, the methodology says so rather than shipping a confident number.
Why does the ranking differ from vibes online?
Public sentiment measures memorable peaks; suites measure consistency across many briefs under one rubric. Both are real; only one has intervals.
Should chat products route creative work separately?
Yes — it is one of the clearest cases for per-workload routing, since the best creative model is rarely the best extraction or classification model, and is rarely priced like them.
Potion routes each request to the cheapest option measured at your quality bar — with a receipt on every answer and this page's evidence behind every pick. Get an API key or read the docs.