unbiased ai

What will you do with 5x the intelligence?

Make every AI model compete for your traffic.

Plans from $10/month · Compare plans ↓

Benchmarks

DeepSWE

Cost per task, USD · log scale

DeepSWE (n=30) · 3 reported results · Score (%) and USD mean per task
Pareto 26.9Sep 20, 202670.0%$0.29
Fable 5Aug 15, 202670.0%$13.50
GPT-5.6 SolAug 15, 202660.0%$0.52

Bars use one linear cost scale per benchmark.

Evaluation details
  • Preliminary comparison transcribed from the supplied September 21, 2026 screenshot, not independently validated. Denominators and cost methodology require confirmation.
  • Run dates: September 20, 2026 for Pareto 26.9; August 15, 2026 for the other results. DeepSWE and ArXivMath use a 30-task slice, not the full benchmark.
  • Each point is one reported result; no performance curves are interpolated. Costs are reported mean USD per task, not token prices. Missing results stay missing and never plot as zero.
  • Axes fit the results shown: cost on a log scale, score over at least 30 points. The orange step line shows the best score reached at or below each cost.
  • Grey points are other labs’ published results as of September 22, 2026, from the DeepSWE and Terminal-Bench leaderboards, MathArena (ArXivMath June 2026), Artificial Analysis, and OpenAI’s GPT-6 Sol and Luna announcement. They use their own task sets and harnesses, so they are not directly comparable with our runs. Only results with a published cost per task are shown.
  • Fable 5 and GPT-5.6 Sol are historical comparators, not the models in the published scores. Router results from these runs are not shown.
  • These runs are separate from the published score slate; do not pair its 74 DeepSWE score with these task costs.

PARETO 26.9 / PUBLISHED SCORES

Published individual-model scores

Pareto 26.9 against Fable 5.1, GPT 6 Astra and DeepSeek 4.1 Flash, from a separate evaluation of individual models.

Scores out of 100 on each benchmark · higher is better
BenchmarkPareto 26.9Fable 5.1GPT 6 AstraDeepSeek 4.1 Flash
DeepSWEAgentic coding74677474
Terminal-Bench 4.0Agentic terminal use51565831
MMMU-ProMultimodal78818777
HLE (no tools)Expert reasoning49555439
ArXivMathResearch math88729128

Task costs: Not reported. These scores come from a separate evaluation and are not plotted on the cost chart, which is why Pareto 26.9 shows 74 on DeepSWE here and 70 on the preliminary 30-task slice. Tinted cells mark the best score on each benchmark, including ties.

Results and methodology on the model card →

Subscriptions & API access

Explore plans

Subscribe for personal use, or pay as you go for your team.

For heavier personal use

Personal Max

$100/month

More included usage for bigger projects.

  • For one person
  • 300M tokens a week
  • 60 requests per minute
  • Billed monthly
Choose Personal Max

For teams

Team

Pay as you go

Prepaid credits for usage across your team.

  • Unlimited team members
  • No weekly token cap
  • 300 requests per minute
  • Prepaid credits, recharge as needed
Choose pay as you go

Prices in USD. Subscriptions include a weekly token allowance, not unlimited usage. Enter your email on the platform to choose your plan.

Questions? Contact us

01Why this exists

Nobody's job is to keep you on the most effective model.

A new frontier model ships every few weeks: ahead on one , behind on another, priced differently again. Most teams pick one model and stop checking, because checking is a full-time job nobody was hired to do. And every vendor you’d ask to check for you has a reason to recommend themselves.

Unbiased does that checking for you, continuously, and says so plainly when the honest answer is a model that isn’t ours.

Measured against models from

  • Anthropic
  • OpenAI
  • Google
  • SpaceXAI
  • DeepSeek
  • Moonshot AI
  • Z.ai
  • Alibaba Qwen

02How Pareto works

One solution, powered by multiple models.

Unbiased is our platform. Pareto is our own blended AI model, available through the API. One model string, one bill. Under the hood it runs several models on your request and keeps the best answer. Bring-your-own-key traffic splitting is on the roadmap, not sold today.

Pareto runs a mix of and open source models against each other on every request, checks which one earns the answer, and keeps the best result for less.

Your code makes one API call and gets one response back, with no per-model integrations to build or maintain.

Unlike , Pareto never switches models mid-conversation, so your prompt cache, and the savings, stay intact.

Dig deeper into how Pareto works →

03Who's behind this

We believe AI should be accessible to everyone.

Our mission is to make AI accessible to everyone, not just those who can afford frontier prices. Unbiased and Pareto are built and maintained by Circuit & Chisel, a remote-first team across the US and Canada: people making tools for other people. Reach out and talk to us.

Meet the team →

04Questions

FAQ

Is Pareto right for every task?

No. The point of Unbiased is to say plainly when Pareto earns the work and when another model is a better fit. Until the public evidence is complete, test Pareto on your own representative workload.

How should I test Pareto?

Start with a real, representative workload and a clear success criterion. Compare quality, reliability, latency, tool behavior, and cost against the model you normally use.

How is Unbiased different from OpenRouter?

OpenRouter offers broad access to third-party models, provider routing, and fallbacks. Unbiased is starting with Pareto and a buyer-side thesis: determine which intelligence earns each workload and be honest when Pareto does not. If you need a large model catalog or provider failover today, OpenRouter may be the right tool.

What happens to my prompts and responses?

Prompt and response content can pass through Circuit & Chisel infrastructure to operate the service. The Data Policy explains current uses, security practices, and open retention questions. Contact the team to confirm which commitments apply to your use before sending sensitive data.

All questions →

Put Pareto on a real workload.

One model string, one bill. Plans from $10/month.