unbiased ai

Pareto 26.8

Technical Details

Pareto is the first model available through Unbiased. It is a blended model: multiple LLMs are engaged in parallel on every request, with results synthesized dynamically based on the task. You send one request and get one answer. Here is how it measures against Kimi K3, Claude Opus 5, GPT-5.6 Sol, Claude Fable 5, and DeepSeek V4F-0731, in the same harness.

Running several models and keeping one answer sounds like it should cost more, not less. Two things make it cheaper. The models in the blend are far cheaper per token than the frontier model you would otherwise send every request to, so several of them together still come in under one frontier call. And the figures below are not estimates: every cost on this page is measured per completed task, blend included. Whatever the blend spends to reach an answer is already inside the number you are reading.

Results

Seven benchmarks · Frontier models

Composite score against the bill.

Unweighted means on both axes: seven scores, measured $ per task.

Composite score against the bill: Pareto 26.8 against four frontier models Composite eval score (unweighted mean of seven public benchmarks) versus measured cost per task in dollars, averaged over the same seven benchmarks. The score axis begins at 65, not zero. Pareto 26.8: 73.5 at $0.10 per task. Opus 5: 73.4 at $0.61. GPT 5.6 Sol: 72.3 at $0.23. Fable 5: 69.4 at $0.64. Kimi K3: 66.5 at $0.51. DeepSeek V4F-0731 appears in the table below but publishes no MMMU-Pro result, so it has no seven-benchmark composite. The full data table follows this figure on the page. 66 68 70 72 74 Composite eval score $0 $0.10 $0.20 $0.30 $0.40 $0.50 $0.60 $0.70 Measured cost per task, $ · mean of the seven benchmarks ↖ better GPT 5.6 Sol Kimi K3 Opus 5 Fable 5 Pareto 26.8 $0.10 / task · 73.5 Composite score against the bill: Pareto 26.8 against four frontier models Composite eval score (unweighted mean of seven public benchmarks) versus measured cost per task in dollars, averaged over the same seven benchmarks. The score axis begins at 65, not zero. Pareto 26.8: 73.5 at $0.10 per task. Opus 5: 73.4 at $0.61. GPT 5.6 Sol: 72.3 at $0.23. Fable 5: 69.4 at $0.64. Kimi K3: 66.5 at $0.51. DeepSeek V4F-0731 appears in the table below but publishes no MMMU-Pro result, so it has no seven-benchmark composite. The full data table follows this figure on the page. 65 70 75 $0 $0.35 $0.70 Measured $ per task ↖ better GPT 5.6 Sol Kimi K3 Opus 5 Fable 5 Pareto 26.8 $0.10 / task · 73.5

Benchmarked August 2026

7 / 7

benchmarks where no model we tested scores higher at a lower measured cost

$0.10

mean measured cost per task across all seven benchmarks; the frontier models on this slate run $0.23–$0.64

73.5

composite score at $0.10 per task, the highest composite we measured, at the lowest mean cost

Details

Terminal-Bench 2.1

90 65 $0 $1.30 Kimi K3: 71.9 at $0.3045 / task Opus 5: 83.1 at $0.7118 / task GPT 5.6 Sol: 81.8 at $0.3446 / task Fable 5: 69.0 at $1.2051 / task DeepSeek V4F: 69.7 at $0.0248 / task Pareto 26.8: 86.0 at $0.2690 / task 86.0 · $0.2690

SWE-Bench Verified

95 70 $0 $2.40 Kimi K3: 86.0 at $2.3701 / task Opus 5: 92.0 at $1.9953 / task GPT 5.6 Sol: 74.0 at $0.7981 / task Fable 5: 88.0 at $1.6203 / task DeepSeek V4F: 77.1 at $0.0154 / task Pareto 26.8: 86.0 at $0.2708 / task 86.0 · $0.2708

GPQA-Diamond

95 80 $0 $0.04 Kimi K3: 88.0 at $0.0378 / task Opus 5: 91.4 at $0.0283 / task GPT 5.6 Sol: 91.9 at $0.0169 / task Fable 5: 90.0 at $0.0396 / task DeepSeek V4F: 84.2 at $0.0013 / task Pareto 26.8: 90.4 at $0.0093 / task 90.4 · $0.0093

Humanity's Last Exam

55 30 $0 $0.40 Kimi K3: 44.9 at $0.2763 / task Opus 5: 51.7 at $0.2529 / task GPT 5.6 Sol: 48.4 at $0.0761 / task Fable 5: 46.0 at $0.3728 / task DeepSeek V4F: 33.7 at $0.0053 / task Pareto 26.8: 47.0 at $0.0261 / task 47.0 · $0.0261

arXivMath

80 15 $0 $0.80 Kimi K3: 35.9 at $0.2906 / task Opus 5: 51.3 at $0.6225 / task GPT 5.6 Sol: 75.0 at $0.1220 / task Fable 5: 59.0 at $0.7900 / task DeepSeek V4F: 20.0 at $0.0079 / task Pareto 26.8: 69.2 at $0.0377 / task 69.2 · $0.0377

DRACO

65 50 $0 $0.70 Kimi K3: 57.9 at $0.2603 / task Opus 5: 61.4 at $0.6575 / task GPT 5.6 Sol: 55.7 at $0.2221 / task Fable 5: 60.3 at $0.4517 / task DeepSeek V4F: 52.0 at $0.0155 / task Pareto 26.8: 58.0 at $0.0773 / task 58.0 · $0.0773

MMMU-Pro

85 70 $0 $0.07 Kimi K3: 80.7 at $0.0644 / task Opus 5: 82.9 at $0.0182 / task GPT 5.6 Sol: 79.5 at $0.0120 / task Fable 5: 73.4 at $0.0265 / task Pareto 26.8: 77.9 at $0.0039 / task 77.9 · $0.0039

Pareto 26.8

Kimi K3, Opus 5, GPT 5.6 Sol, Fable 5, DeepSeek V4F

Frontier intelligence: at or above the weakest frontier score on each chart

Benchmark score vs. measured cost per task · DeepSeek V4F shown for comparison
Benchmark Kimi K3 Opus 5 GPT 5.6 Sol Fable 5 DeepSeek V4F Pareto 26.8
Terminal-Bench 2.1 71.9$0.304583.1$0.711881.8$0.344669.0$1.205169.7$0.024886.0$0.2690
SWE-Bench Verified 86.0$2.370192.0$1.995374.0$0.798188.0$1.620377.1$0.015486.0$0.2708
GPQA-Diamond 88.0$0.037891.4$0.028391.9$0.016990.0$0.039684.2$0.001390.4$0.0093
Humanity's Last Exam 44.9$0.276351.7$0.252948.4$0.076146.0$0.372833.7$0.005347.0$0.0261
arXivMath 35.9$0.290651.3$0.622575.0$0.122059.0$0.790020.0$0.007969.2$0.0377
DRACO 57.9$0.260361.4$0.657555.7$0.222160.3$0.451752.0$0.015558.0$0.0773
MMMU-Pro 80.7$0.064482.9$0.018279.5$0.012073.4$0.0265n/a77.9$0.0039
Composite, mean of 7 66.573.472.369.456.1*73.5
Mean $ / task, measured $0.51$0.61$0.23$0.64$0.012*$0.10

best score

bold lowest bill

settled

*DeepSeek V4F: no MMMU-Pro; its composite and mean $/task cover the 6 benchmarks it ran, not comparable with the seven-benchmark figures

Vision

Pareto takes image inputs. On MMMU-Pro, the multimodal benchmark, Pareto scores 77.9% at $0.0039 per task, the lowest bill on the row. Claude Opus 5 scores higher, 82.9% at $0.0182 per task, and Kimi K3 scores 80.7% at $0.0644. DeepSeek V4F-0731 publishes no MMMU-Pro result.

The bill, drawn to scale

SWE-Bench Verified · same 86.0 score
Kimi K3$2.37
Pareto$0.27
Terminal-Bench 2.1 · Pareto scores higher
Opus 5$0.71
Pareto$0.27
arXivMath · Pareto scores higher
Fable 5$0.79
Pareto$0.04

Bars are cost per completed task. Each competitor bar is normalized to 100%; Pareto shows its share of that cost. Quality is shown beside each pair.

These bills are reproducible.
Test them on your workload

How we measure

Every comparison on this page runs through one public harness: same prompts, same configuration, same run window, for every model, with every side reported. Read the full methodology, or go straight to the public evaluation harness and run it yourself.

Latency

Latency was last measured in July 2026, against Claude Opus 4.8, the prior baseline. We haven't measured latency against this slate yet; when we do, the numbers go here. In that measurement, Pareto added no latency on agentic tasks, including coding. On very difficult reasoning tasks, Pareto can run up to 3x slower: the blend gates its answer before it speaks, so hard requests arrive in one late burst rather than a stream. That is the one drawback we have measured, and we would rather state it than have you discover it.

Production volume

In the ten days ending August 21, 2026, the Pareto gateway served just over 900 million tokens, with a single-day peak near 290 million on August 20. Most of that traffic is one production customer running real workloads; most of the remainder is our own internal use. We state the composition because a volume number without it is an easy place to mislead. The figures are read from Metronome, the meter that prices every request, so the count published here is the same count that produces the bills.

The comparison set, exactly

Comparison models
Kimi K3 ($3 input / $15 output per million tokens, Moonshot AI list, verified Aug 1, 2026), Claude Opus 5 ($5 / $25, Anthropic), Claude Fable 5 ($10 / $50, Anthropic), GPT-5.6 Sol ($5 / $30, OpenAI), and DeepSeek V4F-0731, for which we publish measured per-task bills only; it carries no list-price entry in /data/prices.json. List prices as of Aug 1, 2026, machine-readable in /data/prices.json. No negotiated discounts on any side.
Why this set
These are the frontier models you would otherwise send every request to (the models the blend has to beat on your bill), plus DeepSeek V4F-0731, the lowest-billed model on the slate. It is one slate, not a survey of every model.
Cost basis
Measured per completed task, not estimated from token math.
Modalities
Text and vision.
Members of the blend, and our change policy
We don't publish a roster (it changes too often to stay accurate), but every change is published before it serves member traffic. Read why on How it works.
Test it against your own workload.
Get started with Unbiased