Unbiased

Blog

Every gateway fee structure, priced at your spend

Four structures cover the whole category: a percentage on credits, a percentage on tokens, a flat fee, and free-plus-ops. They rank differently at every spend level; here is the whole table and the crossovers.

By the Unbiased Team · published · fees and prices verified July 28, 2026; measured data dated in place

4 structures
5.5% on credits (OpenRouter), 5% on tokens (Requesty), $49 flat (Portkey), $0+ops (LiteLLM)
~$900/mo
where flat overtakes both percentages; ~$18k/mo where honest ops costs beat 5.5%
14x
what model selection moved versus what fees cost, on our measured workload

LLM gateway fees come in four structures: OpenRouter's 5.5% on credit purchases, Requesty's 5% token markup, Portkey's $49/month flat, and LiteLLM's $0 plus ops time (~$1,000/month honestly counted). Percentages win below roughly $900/month of spend; flat wins above it; self-hosting wins with spare ops capacity. Selection, not fees, moves the real money.

The four structures at five spend levels

At $100 a month: the percentages cost about $5.50 and $5, the flat fee costs $49, ops costs $1,000; percentages win easily. At $900: the structures converge, this is the crossover. At $5,000: percentage fees run $3,300 and $3,000 a year against $588 flat. At $18,000: even an honest $1,000/month ops bill beats 5.5%. At $40,000: the percentage costs $26,400 a year, and everything else on the table beats it.

Slide your own number in the explorer below, or use the full fee calculator; both run the same math this paragraph does.

What each fee buys, so the comparison is fair

The 5.5% buys the category's widest catalog and zero ops. The 5% markup includes caching, worth 2.6x on our metered agentic session, and EU residency. The $49 buys governance: SSO, audit logs, spend controls, tokens billed direct. The $0 buys sovereignty and a second job. Ranking fees without their inclusions is how teams pick a cheaper fee and a more expensive month.

Pareto sits outside the fee table by construction: a blended model billing models at cost has no fee layer to rank. What it charges for is the product itself; the receipts say whether that trade earns its keep.

Every fee structure at your spend

OpenRouter 5.5% on credit top-ups; Requesty 5% token markup; Portkey $49/month flat; LiteLLM $0 licence with ops estimated at ~$1,000/month (two loaded engineer-hours a week) unless the box is ticked. Fees verified 2026-07-28. The crossover between 5.5% and $49 flat sits near $891/month.

The proportion that matters more than any fee

On our measured four-task workload, sending each task to the right model saved 14x what the platform fee cost. Fees are the legible number on the pricing page and the small number on the invoice; model-to-task assignment is the big one, and every option in the fee table except the blended model leaves that assignment with you.

So price the fee, then spend your real attention where the money is: metering your top tasks, tracking verbosity, and re-checking assignment when the ladder moves. It moved three times in the last nine months.

Questions, answered

Which LLM gateway has the lowest fees?

Below ~$900/month of spend: the percentage structures (5% and 5.5%). Above it: Portkey's $49 flat. With genuinely spare ops capacity: self-hosted LiteLLM at $0. The ranking is a function of your spend, not a fixed list.

Do any gateways mark up token prices?

Requesty does explicitly (5%, with caching included). OpenRouter passes list prices and fees credits instead. Portkey and self-hosted options leave token billing direct with providers.

Are gateway fees negotiable?

At enterprise spend, structures often are; below that, take the published numbers as real. The cheaper lever at any scale is the model-selection layer, which moved 14x the fee on our measurements.

What fees does Pareto charge?

None as a layer: Pareto is a blended model, not a gateway you pay rent on. Models run at cost inside every request; the model card carries the receipts and the honest latency disclosure.

Compare any two models

VS

Rates verified 2026-07-28. "Measured task" = our identical dashboard-generation prompt, metered where marked ✓ and list-math otherwise. Verbosity from the Verbosity Index, Edition 1. Data: prices.json.

Run the comparison on your own traffic. $100 to verify.
Buy Pareto API credits