unbiased ai

Blog

GPT-6 Luna pricing: the volume tier at a tenth of Haiku's rates

List prices only. We have not metered GPT-6 Luna. Every Luna figure on this page is OpenAI's published Standard rate, verified October 1, 2026, or arithmetic on it, labeled as such. The Claude task bills are dated July 28, 2026 measurements. See current rates. The comparison widget uses current rates and separately dated measurements.

Luna is OpenAI's volume tier at $0.10 input / $0.50 output per million tokens: the price where classification, extraction, and short summaries live. Claude Haiku 4.5 lists at $1/$5, so on the rate card Luna costs a tenth as much. Whether it does the job as well is the part to test.

Skip the read: get a measured recommendation in a few quick questions. Run the Stack Finder

$0.10 / $0.50
GPT-6 Luna per million tokens in / out, OpenAI Standard rates for short-context requests
$1 / $5
Claude Haiku 4.5, Anthropic's volume tier: ten times Luna's rates on both sides
$0.0016
our measured Haiku 4.5 bill for a real invoice-extraction task on July 28, 2026. Volume-tier work is this cheap done right

GPT-6 Luna costs $0.10 per million input tokens and $0.50 per million output as of October 1, 2026, OpenAI's volume tier. Claude Haiku 4.5, the closest Anthropic rival, lists at $1 and $5, ten times Luna on both sides. Luna's cached input bills at $0.01 per million, a tenth of its input rate.

How much does GPT-6 Luna cost?

$0.10 per million input tokens, $0.50 per million output, and $0.01 per million for cached input (OpenAI pricing). These are Standard rates for short-context requests, verified October 1, 2026; the API model id is gpt-6-luna. At these prices the cache discount is a fraction of a fraction of a cent per request, which is the point: this is the tier where you stop thinking per request and start thinking per million requests.

Want this priced against your own workload? Run the Stack Finder or start with $100 in credits.

Luna vs Haiku 4.5: the bake-off that matters

On the rate card this is not close. Take a short extraction request with 2,000 input tokens and 300 output tokens: on Luna that is $0.0002 of input plus $0.00015 of output, about $0.00035. The same token counts on Haiku 4.5 cost $0.0035. Per million such requests: about $350 versus $3,500. That is arithmetic on list prices, not a measurement; we have not metered Luna. The caveat runs both directions: at this tier, quality cliffs cost more than rates do, and each family fails differently. Run both on your worst hundred cases before believing either rate card, and remember a volume-tier failure that escalates to a frontier retry can cost more than the task you tried to save on.

The same two tasks on every model we could meter

We sent an identical dashboard-generation prompt and an identical invoice-extraction prompt to every model we could call directly, on July 28, 2026, and recorded the actual bills. No estimates in the measured rows. GPT-6 Luna and Claude Sonnet 5.5 came out after this run and are not in it.

ModelRate $/M in / outDashboard task, measuredExtraction task, measured
Claude Fable 5$10 / $50$0.1462$0.0148*
Claude Opus 5$5 / $25$0.1158$0.0114
Claude Opus 4.8$5 / $25$0.0724$0.0093
Claude Sonnet 5.5$2 / $10not measurednot measured
Claude Sonnet 5$2 / $10$0.0410$0.0037
Claude Haiku 4.5$1 / $5$0.0162$0.0016

Measured 2026-07-28, one shot each, default settings, list prices. *Fable extraction is same-tokens list math from our earlier run; every other figure is a metered bill. Claude Sonnet 5.5 was released after this run. Prompts published verbatim in the Fable 5 breakdown. All outputs completed the task; quality-per-dollar comparisons need the published benchmark scores, not this table alone.

The volume tier's real job

Luna and Haiku exist because most production traffic is routine: classify this, extract that, summarize this thread. The craft is in the split: send the routine 80% to the volume tier, escalate the hard 20%, and never pay frontier rates for boilerplate. Making that split per request, with receipts, is the entire product we sell: which is exactly why this page tells you the volume tiers are genuinely good. They are what good routing routes to.

GPT-6 Luna pricing questions, answered

Is GPT-6 Luna or Claude Haiku 4.5 cheaper?

Luna, on list price: $0.10 input and $0.50 output per million tokens against Haiku's $1 and $5. On the same token counts that is a tenth of the cost. At a million short tasks a month, list math puts Luna near $350 and Haiku near $3,500, if quality is equal on your traffic, which is the thing to verify.

Is Luna good enough for coding?

It is worth testing for boilerplate, snippets, and simple transformations. We have not measured Luna on code; our July runs show even Haiku 4.5 producing a working one-shot dashboard, which is the bar to check it against. For multi-file agentic work, failure-and-retry costs erase volume-tier savings fast; that traffic belongs a tier or two up.

What does Luna's cached input rate mean at this price?

Cached input bills at $0.01 per million tokens, a tenth of Luna's $0.10 input rate. A 100,000-token context read from cache costs $0.001 instead of $0.01. Cheap to read does not mean good at reasoning over long inputs; quality on long-context tasks is the thing to test, not the price.

When does the volume tier become a false economy?

When failure rates climb: every failed cheap run that escalates to a frontier retry bills you twice, and review time is a cost too. Measure cost per successful task, not cost per request; the gap between those two numbers is where volume tiers quietly lose.

Not sure which model fits?

The Stack Finder asks a few quick questions about your workload and gives you a straight recommendation. No account required.

Try the Stack Finder

Compare any two models

VS

List rates and dated competitor measurements: prices and measured bills. Pareto 26.10 Preview benchmark costs per task are on the model card. Verbosity: Edition 2.

The volume tier, picked for you request by request. $100 to verify.
Get started with Unbiased

Don’t act on this yourself. Hand it to your agent and let it do the switching math for you.