Blog
Best LLM gateway in 2026: ranked, with receipts
Every ranking in this category is a vendor ranking itself first. Ours is too, so we anchor it the only honest way: published fees, measured latency, and a constraint-first verdict that recommends our competitors by name where they win.
The best LLM gateway in 2026 depends on your binding constraint. For data that cannot leave your infrastructure: LiteLLM (free, self-hosted). For governance and observability: Portkey ($49/month flat). For exploring many models at low volume: OpenRouter (5.5% fee on credits). For cutting the bill itself without doing model selection by hand: Pareto, a blended model billed at cost that matched or beat Claude Opus 4.8 on all seven public benchmarks. Percentage and flat fees cross over near $900/month of spend, and the measured gateway hop costs +68ms median.
The ranking, by constraint
| If your constraint is | Best pick | Fee | The honest catch |
|---|---|---|---|
| Data locality / VPC | LiteLLM | $0 + ops time | free costs ~$1,000/mo in engineer-hours without a platform team |
| Governance, audit, compliance | Portkey | $49/mo flat | visibility does not lower token prices |
| Trying many models, low spend | OpenRouter | 5.5% on credits | the fee scales with spend; catalog choice stays your job |
| Cache-heavy repetitive traffic | Requesty | 5% markup | same percentage-scaling issue past ~$900/mo |
| The bill itself | Pareto | models at cost | up to 3x latency on hardest reasoning, documented |
Fees from public pricing pages, verified 2026-07-28. Latency measured 2026-07-28 (n=20 streamed, method published). Benchmark claim: seven benchmarks vs Claude Opus 4.8, same public harness, both sides reported, on the model card.
Why constraint-first beats feature-first
Feature grids make every gateway look interchangeable. The dollars do not: at $40,000/month of spend, a 5.5% fee costs about $26,400 a year against $588 for a flat control plane, and neither number touches the bigger lever. On our measured four-task workload, sending the work to the right models saved fourteen times what the gateway fee cost. Pick for your constraint, then check the fee, then remember the fee is not the money.
The numbers behind this page
Everything here traces to a measurement you can re-run: the full comparison (fees, user-feedback record, latency method and script), the gateway guide (the four types explained), the fee calculator (your spend, every fee structure), and the model card (the benchmark receipts). External references: OpenRouter fee breakdown, Portkey pricing, Requesty pricing, OpenRouter's own provider-variance data, and the Series B coverage.
How to decide in one afternoon
Take the Stack Finder for a constraint-first recommendation (a few quick questions, no account). Then verify: every option on this page speaks the OpenAI-compatible surface, so a real side-by-side on your own traffic is a base-URL swap and an afternoon, not a migration. Whatever you pick, demand published evidence; vendors who publish harnesses can be checked.
The questions people actually search
What is the best LLM gateway in 2026?
Constraint-dependent: LiteLLM for data locality, Portkey for governance, OpenRouter for multi-model exploration, Requesty for cache-heavy traffic, and Pareto when the goal is the lowest bill per completed task without hand-managing model selection. There is no single winner, and any page that names one without asking your constraint is selling something.
Which LLM gateway is cheapest?
Below roughly $900/month of spend: percentage-fee options (OpenRouter 5.5%, Requesty 5%). Above it: Portkey's $49/month flat. Cheapest overall is self-hosted LiteLLM if your ops time is genuinely free, and at-cost billing changes the basis entirely: the fee stops being the number that matters.
Is OpenRouter the best gateway?
For prototyping across many models with one key, it is genuinely the best at that job. For production cost control it is a 5.5% fee on top of list prices and model selection stays your job, which is where most of the money actually moves.
Do LLM gateways add latency?
Measurably little for the hop itself: +68ms median time-to-first-token in our published test, with p95 better through the gateway than direct. Heavy per-request routing reasoning on hard tasks is the real cost to ask about, and vendors should publish theirs.
What is the best OpenRouter alternative?
By direction of departure: LiteLLM if you need self-hosting, Portkey if you outgrew the fee at scale, Pareto if the goal is the bill per completed task. Our full comparison covers all of them with fees and user-feedback records.