Pareto
Model Card
Pareto is the first model available through Unbiased. It is a blended model: multiple LLMs are engaged in parallel on every request, with results synthesized dynamically based on the task. You send one request and get one answer. Here is how it measures against Claude Opus 4.8, head-to-head, in the same harness.
Results
| Benchmark | Pareto | Opus 4.8 | Pareto $/task | Opus $/task | On the Opus dollar |
|---|---|---|---|---|---|
| Humanity's Last Exam | 32.80% | 31.10% | $0.0068 | $0.064 | 11¢ |
| GPQA-Diamond | 86.40% | 83.80% | $0.005 | $0.011 | 45¢ |
| MMMU-Pro · vision | 76.80% | 63.30% | $0.0009 | $0.005 | 18¢ |
| ArXivMath | 37.50% | 30.00% | $0.0217 | $0.08 | 27¢ |
| DRACO | 67.4% | 65.4% | $0.012 | $1.097 | 1¢ |
| SWE-Bench Verified | 75.00% | 75.00% | $0.26 | $1.29 | 20¢ |
| Terminal-Bench 2.1 | 80.70% | 69.00% | $0.95 | $4.07 | 23¢ |
Scores are each benchmark's standard metric. "On the Opus dollar" is Pareto's cost per task divided by Opus 4.8's, same harness, same runs. SWE-Bench Verified is a tie on quality; the bill is not.
Vision
Pareto takes image inputs. On MMMU-Pro, the multimodal benchmark, Pareto scores 76.80% against Opus 4.8's 63.30%, at $0.0009 per task versus $0.005.
The bill, drawn to scale
How we measure
One harness, both models, same day
We never compare our numbers against someone else's benchmark run. Scores depend heavily on the harness — the prompts, retries, tooling, and grading — which is why published numbers for the same model routinely disagree with each other.
So every comparison on this page comes from one place: Pareto and Opus 4.8 through the same harness, same configuration, same run window, with both sides reported. If our Opus numbers differ from ones you have seen elsewhere, that is the harness difference at work — and it applies equally to both columns.
The public evaluation harness includes the pinned dependencies and configs used to validate these results. If a number looks wrong to you, run it.
Identical on both sides
- The same task prompts, verbatim
- The same harness, configuration, and grading
- The same run window
- Costs measured per completed task, not estimated from token math
Neither side received
- Benchmark-specific prompt tuning
- Retries beyond the harness's standard policy
- Hand-picked task subsets
Latency
On agentic tasks, including coding, Pareto adds no latency compared to Opus 4.8. On very difficult reasoning tasks, Pareto can run up to 3x slower: the blend gates its answer before it speaks, so hard requests arrive in one late burst rather than a stream. That is the one drawback we have measured, and we would rather state it than have you discover it.
The baseline, exactly
- Baseline model
- Claude Opus 4.8, public list price ($5 input / $25 output per million tokens, Anthropic published list, July 2026). No negotiated discounts on either side.
- Cost basis
- Measured per completed task, not estimated from token math.
- Modalities
- Text and vision.
- Change policy
- Any change to the blend's routing or member models is published before it serves member traffic. Numbers on this page carry their run date once the harness publishes.