The claim, stated plainly
Our headline claim is frontier performance for less: roughly 8.8x more compute for the same spend. Here’s the same claim, with the receipts. For all seven benchmarks we publish, no model we tested scores higher at a lower measured cost.
You should not take that at face value. When one product claims the same output as another at a fraction of the bill, something has usually been left out: the quality, or the honesty of the measurement. The only way to answer that suspicion is to show the work, in full, and then hand you the tools to repeat it. That is what the rest of this page does.
Here’s where those numbers come from: Pareto 26.8’s published rate, and the measured bill per task across the same seven-benchmark slate.
- Input
- $1.25/Mtok
- Cached input
- $0.15/Mtok
- Output
- $6.25/Mtok
12.5–15% of Fable 5’s per-token price, depending on the token mix.
Fable 5’s measured bill runs 6.4× ours; GPT 5.6 Sol’s runs 2.3×.
DeepSeek V4F’s published mean is $0.012, but it covers only the 6 benchmarks it ran, not comparable with the seven-benchmark figures above.
The testing behind the numbers
Every score and every dollar figure we publish comes out of one harness. The same prompts go to every model, verbatim. The same configuration, the same run window. Every side of every comparison is reported, including the rows we lose. There is no benchmark-specific prompt tuning, no extra retries, and no hand-picked subsets. The harness is public at pareto-evals, with pinned configs, and any number below can be reproduced with one command.
A “task” is one benchmark problem carried to completion: one SWE-Bench issue fixed, one GPQA question answered, one terminal job finished. The dollar attached to it is not an estimate from token math. It is the measured bill for that completed task, read from the run itself, for every model in the comparison.
The slate spans graduate-level math and science, deep research, visual reasoning, and agentic work including coding. Here is Pareto 26.8's row, score and measured bill per task:
- SWE-Bench Verified (agentic coding) — 86.0$0.2708 / task
- GPQA-Diamond (graduate-level science) — 90.4$0.0093 / task
Two of the seven. Every benchmark, every model's score and bill, on the technical details page. Same harness, same runs.
One row is worth pausing on, because it is the sharpest form of the claim. On SWE-Bench Verified, Pareto 26.8 scores 86.0. Kimi K3 ties it exactly, at 86.0. Kimi's measured bill for the task is $2.3701; ours is $0.2708. The same score, at 11.4% of the bill. We do not call that row a win on quality, because it is not one. It is a tie on quality. The difference is the invoice.
This testing is not a launch-week exercise. We spend thousands of hours a week testing models so you don't have to. The model also carries our own production load: we use Pareto as our own daily driver; it reviewed our expenses, built our website, and researched and implemented improvements that shipped in 26.8. See the production volume numbers on the technical details page.
The mechanism
Pareto is our blended AI model, available through the Unbiased AI platform. Multiple LLMs are engaged in parallel on every request, and their results are synthesized dynamically based on the task. You send one request and get one answer. The blend includes frontier and open models.
It is worth being precise about what it is not. Pareto is not a “model router.” Routers read the prompt, estimate its complexity, and pick a single model. If a reliable way to do that existed, we'd use it. Routers also break prompt caching when they switch models mid-conversation, and at scale the cache misses erase the savings. Pareto works the same way on easy prompts and hard ones, and the savings hold across long trajectories.
The model composition will keep changing as we test new models that drop and if the old ones can't keep earning their place. What binds instead is a change policy: any change to the blend's routing or member models is published before it serves member traffic.
The mechanism has a measured drawback, and we would rather state it than have you discover it. On very difficult reasoning tasks, the blend gates its answer before it speaks, so hard requests arrive in one late burst rather than a stream, and Pareto can run up to 3× slower. In the same measurement, taken in July 2026 against Claude Opus 4.8, Pareto added no latency on agentic tasks, including coding. We have not yet re-measured latency against the current slate; when we do, the numbers go on the model card.
Why the price is possible
Running several models and keeping one answer sounds like it should cost more, not less. Two things make it cheaper.
First, the models in the blend are far cheaper per token than the frontier model you would otherwise send every request to, so several of them together still come in under one frontier call. That is the whole arithmetic of the price. There is no third trick behind it.
Second, nothing the blend spends is hidden from the figure you see. Every cost we publish is measured per completed task, blend included. Whatever the blend spends to reach an answer is already inside the number you are reading. When the model card says a SWE-Bench task cost $0.2708, that is the bill after the whole blend has run, not the price of one lucky call.
The savings also survive real workloads, not just single prompts, because the blend never switches models mid-conversation and so never breaks prompt caching. That is why the per-task figures hold across long trajectories, where router-style approaches give the savings back in cache misses.
And the price itself is checkable, not folklore. The token rates are published in machine-readable form at /data/prices.json, last verified against provider pricing pages on August 21, 2026. They are verified weekly by automation; drift opens a public review within one day.
Where we lose, and what we tell you then
A page like this is only believable if the losses are on it. They are, at full size, in the same table as the wins.
Opus 5 beats Pareto 26.8 on four of the seven benchmarks we publish; GPT 5.6 Sol beats it on two. Our composite, the unweighted mean across all seven, is the highest we measured, at the lowest mean cost, but the row-by-row record includes real losses on score, and we print each one at the same size as the ties and the wins. See every row on the technical details page.
The same policy runs our Delta board. You pick the model you actually run, and it tells you whether something as good costs less and what switching would save, with every number dated and public sources linked where they exist. If what you run is already the better fit for your work, that is the answer the board gives you, and the right move is to keep running it. A comparison page that can only ever recommend its owner is an advertisement; this one is a measurement, and measurements sometimes come back against us.
If your workload lives where Opus 5 or Sol wins, and the quality gap matters more to you than the bill, staying put is the correct decision and we will say so. What we ask is only that the decision be made on measured numbers rather than on list prices and reputation.
Check it yourself
Nothing above requires trust. Four standing offers exist so that you can replace this page with your own numbers:
- Rerun ours. The harness is public at pareto-evals: reproducible head-to-head runs, accuracy and cost per task, with cost read from the API's own response where it is reported, against contamination-checked datasets. It points at any OpenAI-compatible endpoint, including the one you use today. A free eval allocation is available to anyone who asks: run the full slate or your own suite. Rate-limited, expires in seven days.
- Run yours. Bring your heaviest workload. 72 hours of access, side by side with your current model. You leave with two numbers: the quality delta and the bill. Apply; a person replies within 24 hours.
- Start small: Purchase pay-as-you-go credits. We’re still in beta, so every signup is reviewed and approved by hand before your API access turns on.
- Get an hour of help. A free live onboarding hour with the team comes with any purchase: a real workload, tuned live. Optional.
If the numbers hold on your workload, you will know exactly what you are paying for and why it costs what it does. If they don't, you will have found that out for free, on your own harness, and you should stay where you are. Either way, the question “am I overpaying?” stops being unanswerable. Questions to info@unbiased.ai; a person reads every one.