Unbiased

Blog

The case for being pinned to the frontier

Routers were the right first step: one API, every model. The second step is noticing the job they created. Someone at your company is now the model hunter, and the market makes the hunt permanent.

By the Unbiased Team · published · fees and prices verified July 28, 2026; measured data dated in place

9 months
three flagship-tier price events: a 67% cut, a new premium tier, a scheduled +50%
77%
the billing gap between two same-priced models on our battery. Rate cards cannot see it
14x
what right-model assignment moved versus what any gateway fee cost, measured

Catalog routers solved LLM integration and created a standing job: model hunting. Prices repriced three times at the flagship tier in nine months, same-priced models bill 77% apart through verbosity, and every event reopens a 400-model choice. Being pinned to the frontier means delegating that choice per request, so the market's motion stops being your maintenance.

Routers were the right first step

Credit first: aggregators fixed a real problem. One key instead of five, one API surface instead of five SDKs, the whole market a dropdown away, and list prices passed through honestly. We measured the hop at +68ms median with p95 better than direct; the plumbing objections do not survive contact with data. If you are early, exploring, or breadth-hungry, a router is the correct purchase and we have said so, by name, across this site.

But notice what the catalog did to the decision. It did not make model choice easier; it made model choice permanent.

The hunt never ends, measurably

Here is nine months of the job: November 2025, the flagship price falls 67% overnight, every routing table in the industry is suddenly wrong. July 2026, a premium tier opens above the flagship. September 1, a scheduled 50% increase on the mid-tier's introductory rate. Between events, models that share a rate card bill 77% apart on identical work through token count alone, which no pricing page will ever show you.

Each event reopens the question for every route you maintain. Hunting and pecking through new models toward the frontier is real work, it never finishes, and it compounds: our measured workload showed right-versus-wrong assignment moving 14x more money than every fee in the category combined. The catalog gave you the market and quietly made you its analyst.

Start from your constraint

Pick the constraint that is actually binding.

Honest by design: two of the four answers are not us. Fees verified 2026-07-28; the deeper walkthrough of each trade is above.

Pinned, defined precisely

The alternative is not freezing on one model; freezing is just losing slowly, overpaying on routine traffic while the ladder moves. Pinned to the frontier means the selection decision runs inside every request: several LLMs engage, the result synthesizes into one answer, and when the market reprices, your bill improves without your config changing. That is the blended model, and ours bills the models at cost.

The claim is checkable, which is the point: seven public benchmarks against Opus 4.8 on the same harness, metered task bills, and the disclosure that hardest-reasoning tasks can run up to 3x slower, all on the model card. The router was step one. Step two is retiring the hunt.

Questions, answered

What does 'pinned to the frontier' mean?

Selection runs per request instead of per quarter: several LLMs engage on every call and one synthesized answer returns, so market repricing and new launches update your outcomes without updating your job.

Are routers a bad choice, then?

No, they are the right first step and remain right for breadth, exploration, and teams that want selection in-house. The argument is about the standing job catalogs create, not against the category.

Can't I just automate model selection myself?

You can write routing rules; teams do. The rules know what you knew when you wrote them, and the market moved three times at one tier in nine months. Self-maintained automation is the hunt with extra steps.

How do I check the blended-model claim?

The way you would check anyone: run your own tasks, meter the bills, run your eval set. Our methodology, harness, and per-benchmark results are public; $100 of credits covers a serious evaluation.

Compare any two models

VS

Rates verified 2026-07-28. "Measured task" = our identical dashboard-generation prompt, metered where marked ✓ and list-math otherwise. Verbosity from the Verbosity Index, Edition 1. Data: prices.json.

Run the comparison on your own traffic. $100 to verify.
Buy Pareto API credits