Skip the read: get a measured recommendation in a few quick questions. Run the Stack Finder
GPT-6 Luna costs $0.10 per million input tokens and $0.50 per million output as of October 1, 2026, OpenAI's volume tier. Claude Haiku 4.5, the closest Anthropic rival, lists at $1 and $5, ten times Luna on both sides. Luna's cached input bills at $0.01 per million, a tenth of its input rate.
How much does GPT-6 Luna cost?
$0.10 per million input tokens, $0.50 per million output, and $0.01 per million for cached input (OpenAI pricing). These are Standard rates for short-context requests, verified October 1, 2026; the API model id is gpt-6-luna. At these prices the cache discount is a fraction of a fraction of a cent per request, which is the point: this is the tier where you stop thinking per request and start thinking per million requests.
Luna vs Haiku 4.5: the bake-off that matters
On the rate card this is not close. Take a short extraction request with 2,000 input tokens and 300 output tokens: on Luna that is $0.0002 of input plus $0.00015 of output, about $0.00035. The same token counts on Haiku 4.5 cost $0.0035. Per million such requests: about $350 versus $3,500. That is arithmetic on list prices, not a measurement; we have not metered Luna. The caveat runs both directions: at this tier, quality cliffs cost more than rates do, and each family fails differently. Run both on your worst hundred cases before believing either rate card, and remember a volume-tier failure that escalates to a frontier retry can cost more than the task you tried to save on.
The same two tasks on every model we could meter
We sent an identical dashboard-generation prompt and an identical invoice-extraction prompt to every model we could call directly, on July 28, 2026, and recorded the actual bills. No estimates in the measured rows. GPT-6 Luna and Claude Sonnet 5.5 came out after this run and are not in it.
| Model | Rate $/M in / out | Dashboard task, measured | Extraction task, measured |
|---|---|---|---|
| Claude Fable 5 | $10 / $50 | $0.1462 | $0.0148* |
| Claude Opus 5 | $5 / $25 | $0.1158 | $0.0114 |
| Claude Opus 4.8 | $5 / $25 | $0.0724 | $0.0093 |
| Claude Sonnet 5.5 | $2 / $10 | not measured | not measured |
| Claude Sonnet 5 | $2 / $10 | $0.0410 | $0.0037 |
| Claude Haiku 4.5 | $1 / $5 | $0.0162 | $0.0016 |
Measured 2026-07-28, one shot each, default settings, list prices. *Fable extraction is same-tokens list math from our earlier run; every other figure is a metered bill. Claude Sonnet 5.5 was released after this run. Prompts published verbatim in the Fable 5 breakdown. All outputs completed the task; quality-per-dollar comparisons need the published benchmark scores, not this table alone.
The volume tier's real job
Luna and Haiku exist because most production traffic is routine: classify this, extract that, summarize this thread. The craft is in the split: send the routine 80% to the volume tier, escalate the hard 20%, and never pay frontier rates for boilerplate. Making that split per request, with receipts, is the entire product we sell: which is exactly why this page tells you the volume tiers are genuinely good. They are what good routing routes to.
GPT-6 Luna pricing questions, answered
Is GPT-6 Luna or Claude Haiku 4.5 cheaper?
Luna, on list price: $0.10 input and $0.50 output per million tokens against Haiku's $1 and $5. On the same token counts that is a tenth of the cost. At a million short tasks a month, list math puts Luna near $350 and Haiku near $3,500, if quality is equal on your traffic, which is the thing to verify.
Is Luna good enough for coding?
It is worth testing for boilerplate, snippets, and simple transformations. We have not measured Luna on code; our July runs show even Haiku 4.5 producing a working one-shot dashboard, which is the bar to check it against. For multi-file agentic work, failure-and-retry costs erase volume-tier savings fast; that traffic belongs a tier or two up.
What does Luna's cached input rate mean at this price?
Cached input bills at $0.01 per million tokens, a tenth of Luna's $0.10 input rate. A 100,000-token context read from cache costs $0.001 instead of $0.01. Cheap to read does not mean good at reasoning over long inputs; quality on long-context tasks is the thing to test, not the price.
When does the volume tier become a false economy?
When failure rates climb: every failed cheap run that escalates to a frontier retry bills you twice, and review time is a cost too. Measure cost per successful task, not cost per request; the gap between those two numbers is where volume tiers quietly lose.
Not sure which model fits?
The Stack Finder asks a few quick questions about your workload and gives you a straight recommendation. No account required.
Compare any two models
List rates and dated competitor measurements: prices and measured bills. Pareto 26.10 Preview benchmark costs per task are on the model card. Verbosity: Edition 2.
Don’t act on this yourself. Hand it to your agent and let it do the switching math for you.