Unbiased

Blog

Opus 4.8 vs Opus 5: same price, different bills

People search this comparison for capability. The answer nobody gives them is about cost: the two models share a rate card and do not share a bill. We ran an identical six-task battery, twice, on both. Here is the measured gap and what to do about it.

By the Unbiased Team · published · measured July 30, 2026 (12 runs per model); prices verified July 28

$5 / $25
both models' rate card, identical to the cent, per million tokens in and out
+77%
Opus 5's bill on the identical six-task battery: $0.2018 vs $0.1140. All of it is token count
1.48 vs 0.92
their Verbosity Index scores: the newest flagship writes the most, its predecessor writes below median

Claude Opus 5 and Claude Opus 4.8 both cost $5 per million input tokens and $25 per million output. On an identical six-task battery run twice on each model (July 30, 2026), Opus 5 billed $0.2018 and Opus 4.8 billed $0.1140, a 77% gap produced entirely by output token count: Opus 5 writes more for the same work.

The measured head-to-head

Task (identical prompts)Opus 4.8 tokensOpus 5 tokensOpus 4.8 billOpus 5 bill
Code generation (dashboard)2,7725,305$0.0698$0.1331
Quantitative analysis526998$0.0138$0.0256
Email drafting474782$0.0123$0.0200
Short factual answer108273$0.0029$0.0070
Summarization194226$0.0060$0.0068
JSON extraction319319$0.0093$0.0093
Six-task battery total$0.1140$0.2018

Mean of two runs per cell, default settings, no length guidance. Raw data in the Verbosity Index feed. Note the exception that proves the mechanism: on rigidly-structured JSON extraction, where the output length is dictated by the data, the two models tie exactly.

What this is and is not saying

This page measures cost, not capability. Opus 5 is the current flagship and Anthropic positions it as the stronger model; on tasks where its extra reasoning pays, the extra tokens may be exactly what you want. What the data says is narrower and, for budgeting, more useful: an upgrade from 4.8 to 5 is a price increase nobody announced, roughly 77% on this battery, arriving through verbosity rather than the rate card. If you upgrade blindly and your bill jumps, this is why.

If you are on Opus 4.8 today

Three moves, in order. Measure before migrating: run your own eval set on both and compare bills, not just quality; the OpenAI-compatible surface makes it an afternoon. Cap the verbosity: explicit conciseness instructions and max-token limits reclaim much of the gap if you do move. Reconsider the premise: if 4.8 clears your quality bar, its 0.92 verbosity score at frontier quality makes it quietly the best cost-per-task flagship on the market right now, for as long as Anthropic keeps serving it.

The bigger pattern

This gap is one edition of a measurement we now run monthly across every model we can meter: the Verbosity Index. Rate cards are the price of tokens; verbosity is the price of models. Our calculator applies both, and the model card carries the benchmark-grade version of cost-per-task, where the answer to "which model should serve this request" stops being a manual decision at all.

The questions people search

Is Opus 5 more expensive than Opus 4.8?

On the rate card, no: both are $5/$25 per million tokens. On measured bills, yes: 77% more across our identical six-task battery, entirely through higher output token counts. Same prices, different bills.

Is Opus 5 better than Opus 4.8?

Anthropic positions Opus 5 as the stronger flagship, and on hard reasoning the extra tokens can be earning their keep. We measure cost, not capability; run your own failure cases before deciding the 77% is buying you anything on your traffic.

Why does Opus 5 write more tokens?

Newer models tend to elaborate more by default: fuller explanations, more caveats, richer code comments. Elaboration is billed. The one task where they tied, rigid JSON extraction, is the one where the output length is dictated by the data instead of the model.

Should I stay on Opus 4.8?

If it clears your quality bar, its below-median verbosity at frontier quality makes it the quiet cost-per-task bargain of the current ladder. Watch deprecation timelines and abstract your model string either way.

Or let the routing decide per request. $100 to verify.
Buy Pareto API credits