unbiased ai

Blog

Pareto 26.10 Preview

Frontier intelligence. Zero data retention. Open-source pricing.

By the Unbiased AI Team ·

Two weeks ago we put Pareto 26.9 on OpenRouter, OpenCode, and Cloudflare AI Gateway under a made-up name. A few thousand developers ran it hard, told us what broke, and told us what it cost them. We took the notes.

Pareto 26.10 is what we did with those notes. It’s in preview as of today.

Frontier intelligence. Zero data retention. Open-source pricing. The rest of this post is the evidence.

What’s different

Cheaper. 26.10 is priced at $0.80 per million input tokens, $0.03 cached, $3.20 output. Compared with 26.9’s $2.50 / $0.25 / $7.50 rates, that is 68% less for input, 88% less for cached input, and 57% less for output.

GPT-6.1 Sol is $2 / $0.10 / $10. Claude Fable 5.1 is $10 / $0.25 / $50. Prices are per million input / cached input / output tokens. See Pareto pricing.

Better. We ran 26.10 on four benchmarks. In these preliminary results, it matches or sets the Pareto frontier on all of them:

Pareto 26.10 Preview · October 1, 2026 runs
BenchmarkScoreCost / task
GPQA-DiamondScience reasoning92.4%$0.004
Humanity’s Last ExamText-only49.9%$0.008
DeepSWE v1.1Agentic coding69.9%$0.24
Terminal-Bench 4.0Agentic terminal use50.8%$0.48

Score is percent correct; cost is reported mean USD per task, not a token rate. These are preliminary numbers from a serving stack that’s still settling; the model card will update as they firm up.

Performance × cost

Benchmarks

DeepSWE

Cost per task, USD · log scale

DeepSWE v1.1 · 3 reported results · Score (%) and USD mean per task
Pareto 26.10 PreviewOct 1, 202669.9%$0.24
Fable 5Aug 15, 202670.0%$13.50
Pareto 26.9Sep 20, 202670.0%$0.29

Bars use one linear cost scale per benchmark.

Evaluation details
  • Preliminary comparison. Pareto 26.10 Preview results are from October 1, 2026 runs and may change before final publication. Other results are transcribed from a September 21, 2026 comparison and are not independently validated. Denominators and cost methodology require confirmation.
  • Run dates: October 1, 2026 for Pareto 26.10 Preview; September 20, 2026 for Pareto 26.9; August 15, 2026 for Fable 5. The Pareto 26.9 and Fable 5 DeepSWE results use a 30-task slice, not the full DeepSWE v1.1 set. Pareto 26.9’s HLE result was reported as Humanity’s Last Exam, without a text-only qualifier.
  • Each point is one reported result; no performance curves are interpolated. Costs are reported mean USD per task, not token prices. Missing results stay missing and never plot as zero.
  • Axes fit the results shown: cost on a log scale, score over at least 30 points. The orange step line shows the best score reached at or below each cost.
  • Grey points are other labs’ published results as of October 1, 2026, from the DeepSWE and Terminal-Bench leaderboards, Artificial Analysis, Epoch AI, OpenAI’s GPT-6 Sol and Luna and GPT-6.1 Sol announcements, and Anthropic’s Claude Sonnet 5.5 announcement and system card. They use their own task sets and harnesses, so they are not directly comparable with our runs. Results without a published cost sit on the rail beside the chart, at their score.
  • Pareto 26.9 is the previous release, shown for comparison. Fable 5 is a historical comparator, not a model in the published scores. We have not run GPT-6.1 Sol, GPT-6 Luna or Claude Sonnet 5.5 ourselves. Their points are published results from their labs and independent evaluators. Router results from these runs are not shown.
  • Published competitor results come from separate evaluators, not our preliminary runs. Each competitor score is paired only with its source’s task cost, never Pareto’s task cost.

Faster. Lower latency across the board; the biggest gains are on long agentic runs, where 26.9 users felt it most.

Zero data retention, attested

26.10 adds a full zero-data-retention tier. All of our OpenRouter and Cloudflare traffic runs on this tier.

ZDR means Zero Data Retention. Period.

Frontier intelligence, zero data retention, and open-source pricing are usually a pick-two. This is all three.

Preview means preview

The model will keep improving over the next week or so, and we won’t freeze it to hit a date.

What won’t change: the price won’t go up, and the preview slug won’t break under you. When we lock a stable release we’ll say so at the top of this post, and the slug graduates in place.

When it graduates, 26.10 becomes the default for subscribers. At that point (not during preview) weekly token allowances double for existing and new subscribers: Personal goes from 25 million to 50 million tokens at the same $10 per month, and Personal Max becomes Max, going from 300 million to 600 million tokens at the same $100 per month.

Try it

Show us what you build.

The Unbiased AI Team