Code generation: a model at 1/271st the price of the best scorer, at 97.8% of its quality.

claude-sonnet-4.5 vs deepseek-chat-v3.1 for creative writing

the measured answer

As of August 24, 2026, claude-sonnet-4.5 measures 0.900 on creative writing vs 0.850 for deepseek-chat-v3.1 (+0.050), at 9.2× the price ($7.66 vs $0.8281 per 1K requests).

same suite, same items, same scoring — frontier v4, measured 2026-08-24 · method

modelvendormeasured quality$ / 1K requestsp95 latency
claude-sonnet-4.5anthropic0.900$7.669464 ms
deepseek-chat-v3.1deepseek0.850$0.828122494 ms

Creative work is where measurement is hardest and overclaiming easiest. These numbers come from judge-scored suites whose judges are themselves calibrated against tasks with deterministic truth — and the calibration results are published, including the failures.

These two are part of a larger measured frontier — what is the best model for creative writing shows every measured option for this workload, and other workloads rank these models differently: a model that wins here can lose on another kind of work, which is the whole argument for routing per workload rather than picking one model for everything.

the routed alternative

Potion routes each request to the cheapest option measured at your quality bar — including picks this public page does not name. Get an API key or read the docs.

claude-sonnet-4.5 vs deepseek-chat-v3.1 for creative writing — measured · Frontier Notes