Pareto 26.9’s published benchmark scores are on the model card. Measured task costs and a composite score have not been published for this release.
The ranking, by constraint
| If your constraint is | Best pick | Fee | The honest catch |
|---|---|---|---|
| Data locality / VPC | LiteLLM | $0 + ops time | free costs ~$1,000/mo in engineer-hours without a platform team |
| Governance, audit, compliance | Portkey | $49/mo flat | visibility does not lower token prices |
| Trying many models, low spend | OpenRouter | 5.5% on credits | the fee scales with spend; catalog choice stays your job |
| Cache-heavy repetitive traffic | Requesty | 5% markup | same percentage-scaling issue past ~$900/mo |
| The bill itself | Pareto | models at cost | up to 3x latency on hardest reasoning, documented |
Pareto 26.9’s published benchmark scores are on the model card. Measured task costs and a composite score have not been published for this release.
Why constraint-first beats feature-first
Feature grids make every gateway look interchangeable. The dollars do not: at $40,000/month of spend, a 5.5% fee costs about $26,400 a year against $588 for a flat control plane, and neither number touches the bigger lever. On our measured four-task workload, sending the work to the right models saved fourteen times what the gateway fee cost. Pick for your constraint, then check the fee, then remember the fee is not the money.
The numbers behind this page
Everything here traces to a measurement you can re-run: the full comparison (fees, user-feedback record, latency method and script), the gateway guide (the four types explained), the fee calculator (your spend, every fee structure), and the model card (the benchmark receipts). External references: OpenRouter fee breakdown, Portkey pricing, Requesty pricing, OpenRouter's own provider-variance data, and the Series B coverage.
How to decide in one afternoon
Take the Stack Finder for a constraint-first recommendation (a few quick questions, no account). Then verify: every option on this page speaks the OpenAI-compatible surface, so a real side-by-side on your own traffic is a base-URL swap and an afternoon, not a migration. Whatever you pick, demand published evidence; vendors who publish harnesses can be checked.
The questions people actually search
What is the best LLM gateway in 2026?
Constraint-dependent: LiteLLM for data locality, Portkey for governance, OpenRouter for multi-model exploration, Requesty for cache-heavy traffic, and Pareto when the goal is the lowest bill per completed task without hand-managing model selection. There is no single winner, and any page that names one without asking your constraint is selling something.
Which LLM gateway is cheapest?
Depends on fee basis and volume: compare OpenRouter's 5.5% per credit purchase (min $0.80) and Requesty's 5% token markup with Portkey's $49/month base, which includes 100k logs and bills $9 per additional 100k. Cheapest overall is self-hosted LiteLLM if your ops time is genuinely free.
Is OpenRouter the best gateway?
For prototyping across many models with one key, it is genuinely the best at that job. For production cost control it is a 5.5% fee on top of list prices and model selection stays your job, which is where most of the money actually moves.
Do LLM gateways add latency?
Measurably little for the hop itself: +68ms median time-to-first-token in our published test, with p95 better through the gateway than direct. Heavy per-request routing reasoning on hard tasks is the real cost to ask about, and vendors should publish theirs.
What is the best OpenRouter alternative?
By direction of departure: LiteLLM if you need self-hosting, Portkey if you outgrew the fee at scale, Pareto if the goal is the bill per completed task. Our full comparison covers all of them with fees and user-feedback records.
Compare any two models
List rates and dated competitor measurements: prices and measured bills. Pareto 26.9 measured task costs are not published. Verbosity: Edition 2.