What is the best model for rewriting and editing text?
As of August 24, 2026, the cheapest live-measured option within 90% of top quality on rewriting & editing is gpt-4.1-mini at $1.71 per 1,000 requests — 0.866 measured quality vs 0.950 at the top, with a 18× price spread across the measured frontier (frontier v5).
re-measured weekly · how these numbers are made
Rewrite-and-edit work — tone changes, tightening, constraint-preserving edits — punishes models that "improve" text by discarding requirements. The measured suite scores whether stated constraints survive the edit, which is where cheap and expensive models genuinely separate.
The measured frontier
Every row is a live measurement on Potion's held-out rewriting & editing suite — same items, same scoring, per option. Names that are part of the product are withheld; their numbers are not.
| option | vendor | measured quality | $ / 1K requests | p95 latency |
|---|---|---|---|---|
| Potion combination (composition withheld) | — | 0.836 | $0.3439 | 61110 ms |
| Potion's routed pick (name withheld) | — | 0.821 | $1.29 | 9472 ms |
| ling-3.0-flash | inclusionai | 0.804 | $1.41 | 6288 ms |
| gpt-4.1-mini | openai | 0.866 | $1.71 | 2558 ms |
| gpt-4.1 | openai | 0.864 | $2.00 | 2174 ms |
| claude-sonnet-4.5 | anthropic | 0.906 | $3.65 | 6010 ms |
| claude-opus-5-fast | anthropic | 0.950 | $29.93 | 7591 ms |
frontier v5 · measured 2026-08-24 · live provider calls only — simulated evidence never appears on this page
Head to head
gpt-4.1-mini vs ling-3.0-flashgpt-4.1 vs ling-3.0-flashclaude-sonnet-4.5 vs ling-3.0-flashclaude-opus-5-fast vs ling-3.0-flashgpt-4.1 vs gpt-4.1-miniclaude-sonnet-4.5 vs gpt-4.1-miniclaude-opus-5-fast vs gpt-4.1-miniclaude-sonnet-4.5 vs gpt-4.1claude-opus-5-fast vs gpt-4.1claude-opus-5-fast vs claude-sonnet-4.5
Questions
What does quality mean for a rewrite?
Constraint survival: the edit is scored on preserving required facts, references, and instructions while making the requested change — checked against the item’s stated requirements.
Why do models fail rewrites?
The measured failure mode is silent constraint-dropping: the output reads better and quietly loses a requirement. This also erodes across multi-step pipelines, which is why journeys are measured end-to-end separately.
Is this the same as creative writing?
No — creative work is measured on its own suite with judge calibration. Rewriting is closer to extraction: there is a checkable contract.
Potion routes each request to the cheapest option measured at your quality bar — with a receipt on every answer and this page's evidence behind every pick. Get an API key or read the docs.