Unbiased

Pareto

Model Card

Pareto is the first model available through Unbiased. It is a blended model: multiple LLMs are engaged in parallel on every request, with results synthesized dynamically based on the task. You send one request and get one answer. Here is how it measures against Claude Opus 4.8, head-to-head, in the same harness.

7 / 7
benchmarks matched or beaten against Opus 4.8, including multimodal MMMU-Pro
1–45¢
cost per task, on the Opus dollar, across all seven benchmarks
+0 latency
on agentic tasks including coding

Results

Benchmark Pareto Opus 4.8 Pareto $/task Opus $/task On the Opus dollar
Humanity's Last Exam32.80%31.10%$0.0068$0.06411¢
GPQA-Diamond86.40%83.80%$0.005$0.01145¢
MMMU-Pro · vision76.80%63.30%$0.0009$0.00518¢
ArXivMath37.50%30.00%$0.0217$0.0827¢
DRACO67.4%65.4%$0.012$1.097
SWE-Bench Verified75.00%75.00%$0.26$1.2920¢
Terminal-Bench 2.180.70%69.00%$0.95$4.0723¢

Scores are each benchmark's standard metric. "On the Opus dollar" is Pareto's cost per task divided by Opus 4.8's, same harness, same runs. SWE-Bench Verified is a tie on quality; the bill is not.

Vision

Pareto takes image inputs. On MMMU-Pro, the multimodal benchmark, Pareto scores 76.80% against Opus 4.8's 63.30%, at $0.0009 per task versus $0.005.

The bill, drawn to scale

SWE-Bench Verified · same 75.00% score
Opus$1.29
Pareto$0.26
Terminal-Bench 2.1 · Pareto scores higher
Opus$4.07
Pareto$0.95
Humanity's Last Exam · Pareto scores higher
Opus$0.064
Pareto$0.0068

Bars are cost per completed task. Each Opus bar is normalized to 100%; Pareto shows its share of that cost. Quality is shown beside each pair.

How we measure

One harness, both models, same day

We never compare our numbers against someone else's benchmark run. Scores depend heavily on the harness — the prompts, retries, tooling, and grading — which is why published numbers for the same model routinely disagree with each other.

So every comparison on this page comes from one place: Pareto and Opus 4.8 through the same harness, same configuration, same run window, with both sides reported. If our Opus numbers differ from ones you have seen elsewhere, that is the harness difference at work — and it applies equally to both columns.

The public evaluation harness includes the pinned dependencies and configs used to validate these results. If a number looks wrong to you, run it.

Identical on both sides

  • The same task prompts, verbatim
  • The same harness, configuration, and grading
  • The same run window
  • Costs measured per completed task, not estimated from token math

Neither side received

  • Benchmark-specific prompt tuning
  • Retries beyond the harness's standard policy
  • Hand-picked task subsets

Latency

On agentic tasks, including coding, Pareto adds no latency compared to Opus 4.8. On very difficult reasoning tasks, Pareto can run up to 3x slower: the blend gates its answer before it speaks, so hard requests arrive in one late burst rather than a stream. That is the one drawback we have measured, and we would rather state it than have you discover it.

The baseline, exactly

Baseline model
Claude Opus 4.8, public list price ($5 input / $25 output per million tokens, Anthropic published list, July 2026). No negotiated discounts on either side.
Cost basis
Measured per completed task, not estimated from token math.
Modalities
Text and vision.
Change policy
Any change to the blend's routing or member models is published before it serves member traffic. Numbers on this page carry their run date once the harness publishes.
Test it against your own workload.
Buy Pareto API credits