weekly issue2026-08-25
Five pre-registered experiments asked whether combining models beats the best single model on the same work. It does not, at the current model market — and the reasons are now measured, published, and good news for anyone paying for inference.
weekly issue2026-08-24
We keep a scoreboard of AI models: how well each does a particular kind of work, and what it costs per thousand requests. This week every model on the scoreboard was re-checked with fresh examples. None had deteriorated.
weekly issue2026-08-22
Every week Potion tests AI models on ten kinds of work to find which ones give the best value. This week all ten top picks stayed stable. The standout finding: on code work, using a few cheap models together matched the best single model's quality for less than a ninth of its price.
daily ledger2026-09-05
On Potion's measured agentic-tool-use suite, one model scored 0.972 at $44.15 per thousand requests while another scored 0.886 at $0.1094.
daily ledger2026-09-04
On Potion's measured code-gen suite of 94 tasks, grok-4.6 scored 1.000 while a cheaper model scored 0.979, and the price gap was 296× — from $6.85 to $0.0232 per thousand requests.