unbiased ai

Kimi K3 vs Fable 5

The rate cards are a 3.3x gap: Kimi K3 at $3 in / $15 out per million tokens against Fable 5's $10 / $50. Whether the outputs justify the gap is exactly the kind of question nobody answers with receipts. So we are putting it through the harness.

Kimi K3 list rates per Moonshot, July 2026 ($0.30 per million on cache hits). Fable 5 per Anthropic list.

or match your spend in credits

Kimi K3

API live Jul 16, 2026, per Moonshot's list

$3 / $15 in / out, per Mtok

Cache-hit input
$0.30 /MtokMoonshot lists cache-hit input at $0.30 per million tokens; misses bill at $3. The card is flat across context lengths.
This task, metered
queuedK3 is slotted into the same harness that produced every bill on this page. No number appears here until the run completes; we do not publish estimates.
Output over input
5x

Fable 5

The frontier model, API-only since Jul 7

$10 / $50 in / out, per Mtok

This task, metered
$0.1462Same fixed prompt on every model, run the same day, priced from the provider's own bill, not token estimates. Raw outputs are in the Fable 5 pricing breakdown.
Access
API usage creditsFable 5 left the Claude subscription plans on Jul 7, 2026. Subscriptions had carried weekly usage limits since summer 2025, metered in 5-hour session windows.
Output over input
5x

The tale of the tape, receipts included. Any number with an i button carries its method note; hover it, tap it, or tab to it. One record line is still blank on purpose.

Where the run is

One harness, both models, the same day: identical prompts, standard retries, no per-model tuning, and cost measured per completed task rather than estimated from token math. Both sides publish, including where the cheaper model loses. That is the same method behind the Pareto model card, and the pinned configs are public.

  1. Queued

    Where this matchup is now. K3 is slotted, configs pinned and public, nothing run yet.

  2. Running

    Same prompts, same grading, standard retries. Bills metered per completed task, not token math.

  3. Published

    Scores, bills, and raw outputs, losses included. The list reads it before this page changes.

The results land by email first, then the page updates. Jump to the signup to be in the first group.

The short version, today

about 3xthe rate-card gap, input and output, card against card$3 against $10 on input, $15 against $50 on output. Both gaps sit near 3.3x; the honest rounding is about 3x.
$0.30per million on input when Kimi's cache hits; misses bill at $3
5xoutput rate over input rate on both cardsKimi K3 prices $15 out over $3 in. Fable 5 prices $50 out over $10 in. Both exactly five to one.
Jul 16the day the K3 API went live, at these rates

The rate duel, today

No metered Kimi bill exists yet, so these bars are card rates only. Flip to cache hits to see the sliver Moonshot charges for cached input.

Input, Kimi K3$3
Input, Fable 5$10
Output, Kimi K3$15
Output, Fable 5$50

All bars share one scale, $50 at full width. List rates per provider, July 2026.

Input, Kimi K3, cache hit$0.30
Input, Fable 5, list$10
Output, Kimi K3$15
Output, Fable 5$50

A cache hit reprices Kimi input from $3 to $0.30 per million, a tenth of list. Output is unchanged. Fable cache rates are not part of this page's data, so its bar stays at list.

Same task, on the meter

Fable 5$0.1462
Kimi K3no bill yet
queued, publishes with the run

Fable's bill is real and published. Kimi's line stays blank until the harness fills it. No estimates, no token math.

Price a workload on both cards

20
5
$135.00Kimi K3, this workload
$450.00Fable 5, this workload
$315.00the gap, card against card

At 20M tokens in and 5M out, the cards price this at $135.00 on Kimi K3 and $450.00 on Fable 5.

Card arithmetic only: your token counts times the list rates. It says nothing about output quality or token appetite. That is what the harness run is for.

A million tokens is about 750k words of English, so the sliders above cover real working volumes, not toy ones.

The whole field, July 2026

List rates per provider. Metered bills: the same fixed task, same day, priced from provider bills. Sort any column.

July 2026 rate cards and metered task bills across models
Model Notes
Fable 5$10$50$0.1462Left Claude subscription plans Jul 7; API usage credits
GPT-5.6 Sol$5$30not runOutput prices at 6x input on this card
Claude Opus 5$5$25$0.1158Launched Jul 24; same card as Opus 4.8
Claude Opus 4.8$5$25$0.0724The previous flagship
Kimi K3$3$15queuedAPI live Jul 16; cache-hit input $0.30; flat across context
Terra$2.50$15not run
Sonnet 5$2$10$0.0410Intro card to Aug 31; then $3 / $15
Luna$1$6not run
Haiku 4.5off-listoff-list$0.0162Metered in the same run; its card is not tracked on this page
Pareto, the blendcreditscredits$0.0196Blended per request; receipts on the model card

Rates move. Dated changes are on the timeline below, and the biggest scheduled one has a countdown.

How the prices moved

Jul 16: Kimi K3 API launches

Moonshot's Kimi K3 API went live at $3 / $15, cache-hit input at $0.30, flat across context lengths. Two weeks later it is in our harness.

Next scheduled move: Sonnet 5 intro pricing ends Sep 1, 2026.

Method and caveats

How the metered bills are made

One harness, every model, the same day: identical prompts, standard retries, no per-model tuning. Cost is read off the provider's bill per completed task, never estimated from token math. The model card publishes seven same-harness benchmarks with both bills, losses included.

Why the rate card alone misleads

Cards price tokens; bills price behavior. On the identical task, Opus 5 billed 4,610 tokens where Opus 4.8 billed 2,877, about 60% more, on the same rates. Output tokens run 5 to 6x input across the July cards, so verbosity compounds fast. That is why Kimi's slot above stays blank until it is measured.

The second task we meter

A fixed game build, metered the same way: Fable 5 billed $0.2269, the blend billed $0.0220. Task shape moves bills more than rate cards do, which is why we publish per-task receipts instead of one blended average.

What routing adds in latency

Twenty streamed runs, gateway against direct: median time to first token was 68ms slower through the gateway, and p95 through the gateway beat direct. The routing tax is measurable, small, and published.

3/4of an English word, roughly one token; a million tokens is about 750k words
5 to 6xoutput rate over input rate across the July 2026 cards
2.6xcheaper: one real coding session, priced with and without prompt cachingA real Claude Code session ledger: $0.5887 total across five API calls, 78x more tokens read than written. Full breakdown is on the Opus 4.8 page of this series.

The results land by email first

This matchup is queued in our harness. Scores, bills, and raw outputs go to the list before the page updates.

One request in, one answer out: the blend picks per request, so you skip this choice entirely.

Get your credits matched