Blog
Choosing a model on OpenRouter: the job that never finishes
Four hundred models is a feature until someone has to pick one. Nine months of price events say the picking never ends; here is the evidence, and the three ways teams cope.
Model selection on a catalog is a standing job because the market moves: flagship prices fell 67% overnight in November 2025, Fable 5 opened a premium tier in July 2026, Sonnet 5 rises 50% on September 1, and same-priced models differ 77% in what they bill through verbosity. Every event reopens the choice for all 400+ options.
The evidence that it never settles
Take the last nine months as the sample. November 2025: the flagship tier repriced 67% overnight, from $15/$75 to $5/$25 per million tokens. July 2026: Fable 5 opened a premium tier above the flagship at $10/$50. September 1: Sonnet 5's introductory rate ends with a scheduled 50% increase. Each event changed the right answer for whole classes of traffic.
Prices are the visible half. The invisible half is verbosity: Opus 5 and Opus 4.8 share a rate card and billed 77% apart on our measured battery, purely through token count. The rate card cannot tell you this; only measurement can, and it changes with every model swap.
What choosing well actually requires
Done properly, model selection means: metering your real tasks on candidate models (rate cards mislead; our same prompt billed $0.0162 on Haiku and $0.1462 on Fable 5, both working), re-running that measurement at every price event and major launch, tracking verbosity so same-price swaps do not quietly move the bill, and, on open-weights catalog models, watching provider variance too.
That is a real discipline, and teams that treat it as one do fine on a catalog. The failure mode is the middle path: choosing once, in launch week, from vibes, and then never again while the market moves twelve times.
Start from your constraint
Honest by design: two of the four answers are not us. Fees verified 2026-07-28; the deeper walkthrough of each trade is above.
The three coping strategies
Own it: make selection someone's explicit job with a measurement loop, using tools like our calculator and the Verbosity Index; this is the right answer for teams where model choice is a competency. Freeze it: pin one safe frontier model and accept overpaying on routine traffic; honest, common, and expensive, up to 9x on the routine slice by our measurements.
Delegate it: move selection into the request itself. That is the blended-model category, ours: Pareto runs several LLMs on every request, synthesizes one answer, and bills at cost, so repricing events change your bill without changing your job. The receipts, and the honest latency caveat, live on the model card.
Questions, answered
What is the best model on OpenRouter?
The question has no stable answer; it depends on the task and the month. The durable method: meter your own top tasks on 3-4 candidates, compare bills for working output, and re-run at every price event.
How often do LLM prices change?
At the flagship tier, three major events in the last nine months: a 67% cut, a new premium tier, and a scheduled 50% increase. Assume several reprices a year across the ladder, plus new-model launches monthly.
Do same-priced models cost the same?
Measurably no: identical rate cards billed 77% apart on our battery through token count alone. Verbosity, not the price sheet, decides what a swap does to your bill.
Is there a way to stop re-choosing models?
Two honest options: staff the job with a measurement loop, or delegate selection per request to a blended model. Freezing on one model without measurement is the expensive non-choice.
Compare any two models
Rates verified 2026-07-28. "Measured task" = our identical dashboard-generation prompt, metered where marked ✓ and list-math otherwise. Verbosity from the Verbosity Index, Edition 1. Data: prices.json.