unbiased ai

Blog

Shadow mode

Part of the Unbiased glossary: the vocabulary of LLM pricing and infrastructure, defined with receipts instead of adjectives.

Shadow mode runs a candidate model on a copy of real production traffic without serving its outputs to users, so its quality and cost can be measured against the incumbent on identical inputs, at zero user risk. It is the standard technique for evaluating a model switch honestly.

Why it matters

Benchmarks are somebody else's workload. Shadow mode is yours: the same prompts, documents, and edge cases your system actually sees, scored side by side. It converts a model decision from a leaderboard argument into a measured comparison on the traffic that pays your bills.

The at-home version, measured

Pareto 26.9’s published benchmark scores are on the model card. Measured task costs and a composite score have not been published for this release.

Common questions

Does shadow mode double my costs?

It adds the candidate's inference on mirrored traffic for the duration of the test, so budget the comparison window. Sampling (shadowing 5-10% of traffic) keeps the bill small while the comparison stays honest.

How long should a shadow test run?

Long enough to cover your traffic's real variety, typically one to two representative weeks. A day of easy traffic proves nothing about the tail.

What do I measure?

Task success on your own criteria, cost per completed task, and latency, on identical inputs. If you only measure vibes, shadow mode is theater.

Compare any two models

VS

List rates and dated competitor measurements: prices and measured bills. Pareto 26.9 measured task costs are not published. Verbosity: Edition 2.

See the receipts behind the definitions.
Get started with Unbiased