Code generation: a model at 1/271st the price of the best scorer, at 97.8% of its quality.

What is the best model for rewriting and editing text?

the measured answer

As of August 24, 2026, the cheapest live-measured option within 90% of top quality on rewriting & editing is gpt-4.1-mini at $1.71 per 1,000 requests — 0.866 measured quality vs 0.950 at the top, with a 18× price spread across the measured frontier (frontier v5).

re-measured weekly · how these numbers are made

Rewrite-and-edit work — tone changes, tightening, constraint-preserving edits — punishes models that "improve" text by discarding requirements. The measured suite scores whether stated constraints survive the edit, which is where cheap and expensive models genuinely separate.

The measured frontier

Every row is a live measurement on Potion's held-out rewriting & editing suite — same items, same scoring, per option. Names that are part of the product are withheld; their numbers are not.

optionvendormeasured quality$ / 1K requestsp95 latency
Potion combination (composition withheld)0.836$0.343961110 ms
Potion's routed pick (name withheld)0.821$1.299472 ms
ling-3.0-flashinclusionai0.804$1.416288 ms
gpt-4.1-miniopenai0.866$1.712558 ms
gpt-4.1openai0.864$2.002174 ms
claude-sonnet-4.5anthropic0.906$3.656010 ms
claude-opus-5-fastanthropic0.950$29.937591 ms

frontier v5 · measured 2026-08-24 · live provider calls only — simulated evidence never appears on this page

Head to head

gpt-4.1-mini vs ling-3.0-flashgpt-4.1 vs ling-3.0-flashclaude-sonnet-4.5 vs ling-3.0-flashclaude-opus-5-fast vs ling-3.0-flashgpt-4.1 vs gpt-4.1-miniclaude-sonnet-4.5 vs gpt-4.1-miniclaude-opus-5-fast vs gpt-4.1-miniclaude-sonnet-4.5 vs gpt-4.1claude-opus-5-fast vs gpt-4.1claude-opus-5-fast vs claude-sonnet-4.5

Questions

What does quality mean for a rewrite?

Constraint survival: the edit is scored on preserving required facts, references, and instructions while making the requested change — checked against the item’s stated requirements.

Why do models fail rewrites?

The measured failure mode is silent constraint-dropping: the output reads better and quietly loses a requirement. This also erodes across multi-step pipelines, which is why journeys are measured end-to-end separately.

Is this the same as creative writing?

No — creative work is measured on its own suite with judge calibration. Rewriting is closer to extraction: there is a checkable contract.

the routed alternative

Potion routes each request to the cheapest option measured at your quality bar — with a receipt on every answer and this page's evidence behind every pick. Get an API key or read the docs.

What is the best model for rewriting and editing text? Measured August 2026 · Frontier Notes