unbiased ai

Blog

All of a Sudden, Elon Cares About Routing

The “best model for the job” idea just went mainstream. Here’s what you should want from it.

By Louis Amira, CEO, Unbiased AI ·

This morning Elon Musk posted a short note about Grok Bot:

Going forward, @SpaceX will use the best back end model for any given task, including Claude Opus 5.5, MidJourney, Suno and other leading APIs. Whatever is most likely to give you the best outcome.

Read that again. One of the biggest names in AI just said, in public, that no single model should do everything. Use whatever works best for the task in front of you.

That’s the whole idea behind Unbiased AI. It didn’t even take 24 hours for “model routing” to become the main story.

The category just won

For a while, choosing models across vendors sounded like a niche infrastructure topic. Most teams pick one model, wire it in, and stop checking. Checking is a full-time job nobody got hired to do.

Then the person running one of the most visible AI products said they’re doing it too. Plenty of people saw this coming. Farzad described the idea of model-agnostic orchestration back in August: an AI that knows which models to use, at what cost and speed, so people don’t have to keep track of it all themselves.

Pretty soon, nearly every AI product will work this way.

The obvious question: how?

“Use the best model for any given task” is easy to say. Two questions come right after it:

  1. Who decides what “best” means, and how? That’s the actual product. It’s a hard problem, and you won’t solve it with a lookup table.
  2. Who pays for it? How can anyone afford Opus for everything? Sending every task to Opus gets expensive fast. As the routing layer improves, there’s less reason to pay frontier prices for work that a cheaper model can do just as well.

There’s a third question, too: privacy. If your prompts fan out across Claude, MidJourney, Suno, and “other leading APIs,” where does your data end up, and who keeps it?

That gives you three requirements: smart, cheap, private. You want all three in one model string.

A quick word on “router”

On X today I’ve been saying “auto router,” because that’s the term people use. To be precise: Pareto isn’t a classic router.

A typical router guesses how hard your prompt is and picks one model. When it switches models mid-conversation, the new model may not be able to reuse the previous model’s prompt cache. At scale, those cache misses can eat into the savings. Pareto blends instead. It engages several models on each request, works the same way on easy and hard prompts, and doesn’t switch models mid-conversation. Whether it saves you money is something to measure on your workload.

Whatever you call it, the job is the same: give you the best answer for the task at the best price, and stay honest when a different model would actually serve you better.

Now show the work

Elon’s post puts the idea in front of a much bigger audience. What’s left to figure out is who does it well, at a price that makes sense, without hoarding your data.

We don’t ask you to trust us on that. Rerun our numbers, or bring your heaviest real workload and compare Pareto with whatever you use now. If Pareto doesn’t hold up on your workload, stay where you are. Either way, you’ll find out whether you’re overpaying.

And Grok team, if you’re reading: we’d love to see what Grok does with Pareto behind it.