unbiased

Announcement

Introducing Pareto 26.8

Frontier-quality results at 75–85% less than frontier per-token prices. Today, Pareto moves from private to public beta.

Pareto is a blended model. It composites output from multiple frontier and open-source models to produce frontier-quality results at a fraction of frontier prices. Today, Pareto 26.8 moves from private to public beta.

On our benchmark slate, 26.8 (we version by year.month) beats Fable 5 and GPT 5.6 Sol at 75–85% less than their per-token price. Every number below can be reproduced with one command from our pareto-evals repo.

Thanks to the many builders who hammered on the private beta. We use Pareto as our own daily driver. It reviewed our expenses, built our website, helped write this post, and researched and implemented improvements that shipped in 26.8. Yes, it's now improving itself.

What Pareto isn't

Pareto is not a “model router.” Routers read the prompt, estimate its complexity, and pick a single model — if a reliable way to do that existed, we'd use it. Routers also break prompt caching when they switch models mid-conversation, and at scale, the cache misses erase the savings. Pareto performs optimally on easy prompts and hard ones alike, and the savings hold across long trajectories.

Performance

All data below comes from our own reproductions, apples-to-apples in pareto-evals. Replicate any number with one command.

Graduate-level math and science

GPQA-Diamond

80859095$0$0.01$0.02$0.03$0.04Score %Cost per task (USD)Kimi K3 — 88% at $0.0378 per taskKimi K3Opus 5 — 91.4% at $0.0283 per taskOpus 5Fable 5 — 90% at $0.0396 per taskFable 5GPT-5.6 Sol — 91.9% at $0.0169 per taskGPT-5.6 SolDeepSeek V4F-0731 — 84.2% at $0.0013 per taskDeepSeek V4F-0731Pareto 26.8 — 90.4% at $0.0093 per taskPareto 26.8

Humanity's Last Exam

303540455055$0$0.10$0.20$0.30$0.40Score %Cost per task (USD)Kimi K3 — 44.9% at $0.2763 per taskKimi K3Opus 5 — 51.7% at $0.2529 per taskOpus 5Fable 5 — 46% at $0.3728 per taskFable 5GPT-5.6 Sol — 48.4% at $0.0761 per taskGPT-5.6 SolDeepSeek V4F-0731 — 33.7% at $0.0053 per taskDeepSeek V4F-0731Pareto 26.8 — 47% at $0.0261 per taskPareto 26.8

arXivMath

10305070$0$0.20$0.40$0.60$0.80Score %Cost per task (USD)Kimi K3 — 35.9% at $0.2906 per taskKimi K3Opus 5 — 51.3% at $0.6225 per taskOpus 5Fable 5 — 59% at $0.79 per taskFable 5GPT-5.6 Sol — 75% at $0.122 per taskGPT-5.6 SolDeepSeek V4F-0731 — 20% at $0.0079 per taskDeepSeek V4F-0731Pareto 26.8 — 69.2% at $0.0377 per taskPareto 26.8

Deep research

DRACO

50556065$0$0.10$0.20$0.30$0.40$0.50$0.60$0.70Score %Cost per task (USD)Kimi K3 — 57.9% at $0.2603 per taskKimi K3Opus 5 — 61.4% at $0.6575 per taskOpus 5Fable 5 — 60.3% at $0.4517 per taskFable 5GPT-5.6 Sol — 55.7% at $0.2221 per taskGPT-5.6 SolDeepSeek V4F-0731 — 52% at $0.0155 per taskDeepSeek V4F-0731Pareto 26.8 — 58% at $0.0773 per taskPareto 26.8

Visual reasoning

MMMU-Pro

70758085$0$0.01$0.02$0.03$0.04$0.05$0.06$0.07Score %Cost per task (USD)Kimi K3 — 80.7% at $0.0644 per taskKimi K3Opus 5 — 82.9% at $0.0182 per taskOpus 5Fable 5 — 73.4% at $0.0265 per taskFable 5GPT-5.6 Sol — 79.5% at $0.012 per taskGPT-5.6 SolPareto 26.8 — 77.9% at $0.0039 per taskPareto 26.8DeepSeek V4F-0731 — not multimodal

Agentic tasks, including coding

SWE-Bench Verified

707580859095$0$0.50$1$1.5$2$2.5Score %Cost per task (USD)Kimi K3 — 86% at $2.3701 per taskKimi K3Opus 5 — 92% at $1.9953 per taskOpus 5Fable 5 — 88% at $1.6203 per taskFable 5GPT-5.6 Sol — 74% at $0.7981 per taskGPT-5.6 SolDeepSeek V4F-0731 — 77.1% at $0.0154 per taskDeepSeek V4F-0731Pareto 26.8 — 86% at $0.2708 per taskPareto 26.8

Terminal-Bench 2.1

657075808590$0$0.25$0.50$0.75$1$1.25Score %Cost per task (USD)Kimi K3 — 71.9% at $0.3045 per taskKimi K3Opus 5 — 83.1% at $0.7118 per taskOpus 5Fable 5 — 69% at $1.2051 per taskFable 5GPT-5.6 Sol — 81.8% at $0.3446 per taskGPT-5.6 SolDeepSeek V4F-0731 — 69.7% at $0.0248 per taskDeepSeek V4F-0731Pareto 26.8 — 86% at $0.269 per taskPareto 26.8

Summary

Benchmark scores and measured cost per task for each model. Same data as the graphs.
BenchmarkPareto 26.8Kimi K3Opus 5Fable 5GPT-5.6 SolDeepSeek V4F-0731
GPQA-Diamond90.4%$0.009388.0%$0.037891.4%$0.028390.0%$0.039691.9%$0.016984.2%$0.0013
Humanity's Last Exam47.0%$0.026144.9%$0.276351.7%$0.252946.0%$0.372848.4%$0.076133.7%$0.0053
arXivMath69.2%$0.037735.9%$0.290651.3%$0.622559.0%$0.7975.0%$0.12220.0%$0.0079
DRACO58.0%$0.077357.9%$0.260361.4%$0.657560.3%$0.451755.7%$0.222152.0%$0.0155
MMMU-Pro77.9%$0.003980.7%$0.064482.9%$0.018273.4%$0.026579.5%$0.012n/a (not multimodal)
SWE-Bench Verified86.0%$0.270886.0%$2.370192.0%$1.995388.0%$1.620374.0%$0.798177.1%$0.0154
Terminal-Bench 2.186.0%$0.26971.9%$0.304583.1%$0.711869.0%$1.205181.8%$0.344669.7%$0.0248

Score % over measured $ per task.

Pricing

$1.25/Mtok input, $0.15/Mtok cached input, $6.25/Mtok output. Twelve and a half cents on Fable 5's dollar; twenty-five on GPT 5.6 Sol's.

When the frontier improves, Pareto improves; when frontier prices fall, you should benefit — manually updating by workload would be a peculiar artisanal hobby.

Per-token list prices in USD per million tokens, and each model's price as a percentage of Fable 5's.
Kimi K3Opus 5Fable 5GPT-5.6 SolPareto 26.8
Input $/Mtok$3.00$5.00$10.00$5.00$1.25
Cached input $/Mtok$0.30$0.50$1.00$0.50$0.15
Output $/Mtok$15.00$25.00$50.00$30.00$6.25
% of Fable 5 price30%50%100%50–60%12.5–15%

Until now, overpaying for inference was nobody's fault — there was no way to check. We'd rather you check.

Access

Public beta is capacity-limited. We anticipate most people want to do one of the following:

  • Verify. A free eval allocation for anyone who asks. Run the full slate in pareto-evals, or your own suite. Rate-limited, expires in seven days.
  • Evaluate. Bring your heaviest workload. 72 hours of access, side by side with your current model. You leave with two numbers: the quality delta and the bill. Apply.
  • Commit. A $5,000 credit block. Onboarding within 24 hours, priority capacity, a direct line to the team.

There's no $10 plan yet. If a lab subscription isn't serving you, tell us what you need and we'll see if one of our alpha products is a good match.

Access broadens later this month. We've also applied to list Pareto on OpenRouter; when it clears, it will run there for a limited window.