Running several models and keeping one answer sounds like it should cost more, not less. Two things make it cheaper. The models in the blend are far cheaper per token than the frontier model you would otherwise send every request to, so several of them together still come in under one frontier call. And the figures below are not estimates: every cost on this page is measured per completed task, blend included. Whatever the blend spends to reach an answer is already inside the number you are reading.
Results
Seven benchmarks · Frontier models
Composite score against the bill.
Unweighted means on both axes: seven scores, measured $ per task.
Benchmarked August 2026
benchmarks where no model we tested scores higher at a lower measured cost
mean measured cost per task across all seven benchmarks; the frontier models on this slate run $0.23–$0.64
composite score at $0.10 per task, the highest composite we measured, at the lowest mean cost
Details
Terminal-Bench 2.1
SWE-Bench Verified
GPQA-Diamond
Humanity's Last Exam
arXivMath
DRACO
MMMU-Pro
Pareto 26.8
Kimi K3, Opus 5, GPT 5.6 Sol, Fable 5, DeepSeek V4F
Frontier intelligence: at or above the weakest frontier score on each chart
| Benchmark | Kimi K3 | Opus 5 | GPT 5.6 Sol | Fable 5 | DeepSeek V4F | Pareto 26.8 |
|---|---|---|---|---|---|---|
| Terminal-Bench 2.1 | 71.9$0.3045 | 83.1$0.7118 | 81.8$0.3446 | 69.0$1.2051 | 69.7$0.0248 | 86.0$0.2690 |
| SWE-Bench Verified | 86.0$2.3701 | 92.0$1.9953 | 74.0$0.7981 | 88.0$1.6203 | 77.1$0.0154 | 86.0$0.2708 |
| GPQA-Diamond | 88.0$0.0378 | 91.4$0.0283 | 91.9$0.0169 | 90.0$0.0396 | 84.2$0.0013 | 90.4$0.0093 |
| Humanity's Last Exam | 44.9$0.2763 | 51.7$0.2529 | 48.4$0.0761 | 46.0$0.3728 | 33.7$0.0053 | 47.0$0.0261 |
| arXivMath | 35.9$0.2906 | 51.3$0.6225 | 75.0$0.1220 | 59.0$0.7900 | 20.0$0.0079 | 69.2$0.0377 |
| DRACO | 57.9$0.2603 | 61.4$0.6575 | 55.7$0.2221 | 60.3$0.4517 | 52.0$0.0155 | 58.0$0.0773 |
| MMMU-Pro | 80.7$0.0644 | 82.9$0.0182 | 79.5$0.0120 | 73.4$0.0265 | n/a | 77.9$0.0039 |
| Composite, mean of 7 | 66.5 | 73.4 | 72.3 | 69.4 | 56.1* | 73.5 |
| Mean $ / task, measured | $0.51 | $0.61 | $0.23 | $0.64 | $0.012* | $0.10 |
best score
bold lowest bill
settled
*DeepSeek V4F: no MMMU-Pro; its composite and mean $/task cover the 6 benchmarks it ran, not comparable with the seven-benchmark figures
Vision
Pareto takes image inputs. On MMMU-Pro, the multimodal benchmark, Pareto scores 77.9% at $0.0039 per task, the lowest bill on the row. Claude Opus 5 scores higher, 82.9% at $0.0182 per task, and Kimi K3 scores 80.7% at $0.0644. DeepSeek V4F-0731 publishes no MMMU-Pro result.
The bill, drawn to scale
How we measure
Every comparison on this page runs through one public harness: same prompts, same configuration, same run window, for every model, with every side reported. Read the full methodology, or go straight to the public evaluation harness and run it yourself.
Latency
Latency was last measured in July 2026, against Claude Opus 4.8, the prior baseline. We haven't measured latency against this slate yet; when we do, the numbers go here. In that measurement, Pareto added no latency on agentic tasks, including coding. On very difficult reasoning tasks, Pareto can run up to 3x slower: the blend gates its answer before it speaks, so hard requests arrive in one late burst rather than a stream. That is the one drawback we have measured, and we would rather state it than have you discover it.
Production volume
In the ten days ending August 21, 2026, the Pareto gateway served just over 900 million tokens, with a single-day peak near 290 million on August 20. Most of that traffic is one production customer running real workloads; most of the remainder is our own internal use. We state the composition because a volume number without it is an easy place to mislead. The figures are read from Metronome, the meter that prices every request, so the count published here is the same count that produces the bills.
The comparison set, exactly
- Comparison models
- Kimi K3 ($3 input / $15 output per million tokens, Moonshot AI list, verified Aug 1, 2026), Claude Opus 5 ($5 / $25, Anthropic), Claude Fable 5 ($10 / $50, Anthropic), GPT-5.6 Sol ($5 / $30, OpenAI), and DeepSeek V4F-0731, for which we publish measured per-task bills only; it carries no list-price entry in /data/prices.json. List prices as of Aug 1, 2026, machine-readable in /data/prices.json. No negotiated discounts on any side.
- Why this set
- These are the frontier models you would otherwise send every request to (the models the blend has to beat on your bill), plus DeepSeek V4F-0731, the lowest-billed model on the slate. It is one slate, not a survey of every model.
- Cost basis
- Measured per completed task, not estimated from token math.
- Modalities
- Text and vision.
- Members of the blend, and our change policy
- We don't publish a roster (it changes too often to stay accurate), but every change is published before it serves member traffic. Read why on How it works.