Pareto is a blended model. It composites output from multiple frontier and open-source models to produce frontier-quality results at a fraction of frontier prices. Today, Pareto 26.8 moves from private to public beta.
On our benchmark slate, 26.8 (we version by year.month) beats Fable 5 and GPT 5.6 Sol at 75–85% less than their per-token price. Every number below can be reproduced with one command from our pareto-evals repo.
Thanks to the many builders who hammered on the private beta. We use Pareto as our own daily driver. It reviewed our expenses, built our website, helped write this post, and researched and implemented improvements that shipped in 26.8. Yes, it's now improving itself.
What Pareto isn't
Pareto is not a “model router.” Routers read the prompt, estimate its complexity, and pick a single model — if a reliable way to do that existed, we'd use it. Routers also break prompt caching when they switch models mid-conversation, and at scale, the cache misses erase the savings. Pareto performs optimally on easy prompts and hard ones alike, and the savings hold across long trajectories.
Performance
All data below comes from our own reproductions, apples-to-apples in pareto-evals. Replicate any number with one command.
Graduate-level math and science
GPQA-Diamond
Humanity's Last Exam
arXivMath
Deep research
DRACO
Visual reasoning
MMMU-Pro
Agentic tasks, including coding
SWE-Bench Verified
Terminal-Bench 2.1
Summary
| Benchmark | Pareto 26.8 | Kimi K3 | Opus 5 | Fable 5 | GPT-5.6 Sol | DeepSeek V4F-0731 |
|---|---|---|---|---|---|---|
| GPQA-Diamond | 90.4%$0.0093 | 88.0%$0.0378 | 91.4%$0.0283 | 90.0%$0.0396 | 91.9%$0.0169 | 84.2%$0.0013 |
| Humanity's Last Exam | 47.0%$0.0261 | 44.9%$0.2763 | 51.7%$0.2529 | 46.0%$0.3728 | 48.4%$0.0761 | 33.7%$0.0053 |
| arXivMath | 69.2%$0.0377 | 35.9%$0.2906 | 51.3%$0.6225 | 59.0%$0.79 | 75.0%$0.122 | 20.0%$0.0079 |
| DRACO | 58.0%$0.0773 | 57.9%$0.2603 | 61.4%$0.6575 | 60.3%$0.4517 | 55.7%$0.2221 | 52.0%$0.0155 |
| MMMU-Pro | 77.9%$0.0039 | 80.7%$0.0644 | 82.9%$0.0182 | 73.4%$0.0265 | 79.5%$0.012 | n/a (not multimodal) |
| SWE-Bench Verified | 86.0%$0.2708 | 86.0%$2.3701 | 92.0%$1.9953 | 88.0%$1.6203 | 74.0%$0.7981 | 77.1%$0.0154 |
| Terminal-Bench 2.1 | 86.0%$0.269 | 71.9%$0.3045 | 83.1%$0.7118 | 69.0%$1.2051 | 81.8%$0.3446 | 69.7%$0.0248 |
Score % over measured $ per task.
Pricing
$1.25/Mtok input, $0.15/Mtok cached input, $6.25/Mtok output. Twelve and a half cents on Fable 5's dollar; twenty-five on GPT 5.6 Sol's.
When the frontier improves, Pareto improves; when frontier prices fall, you should benefit — manually updating by workload would be a peculiar artisanal hobby.
| Kimi K3 | Opus 5 | Fable 5 | GPT-5.6 Sol | Pareto 26.8 | |
|---|---|---|---|---|---|
| Input $/Mtok | $3.00 | $5.00 | $10.00 | $5.00 | $1.25 |
| Cached input $/Mtok | $0.30 | $0.50 | $1.00 | $0.50 | $0.15 |
| Output $/Mtok | $15.00 | $25.00 | $50.00 | $30.00 | $6.25 |
| % of Fable 5 price | 30% | 50% | 100% | 50–60% | 12.5–15% |
Until now, overpaying for inference was nobody's fault — there was no way to check. We'd rather you check.
Access
Public beta is capacity-limited. We anticipate most people want to do one of the following:
- Verify. A free eval allocation for anyone who asks. Run the full slate in pareto-evals, or your own suite. Rate-limited, expires in seven days.
- Evaluate. Bring your heaviest workload. 72 hours of access, side by side with your current model. You leave with two numbers: the quality delta and the bill. Apply.
- Commit. A $5,000 credit block. Onboarding within 24 hours, priority capacity, a direct line to the team.
There's no $10 plan yet. If a lab subscription isn't serving you, tell us what you need and we'll see if one of our alpha products is a good match.
Access broadens later this month. We've also applied to list Pareto on OpenRouter; when it clears, it will run there for a limited window.