AI/Comparison

Claude Haiku 5.5 vs GPT-6 Luna: Price and Benchmarks

Anthropic says Haiku 5.5 beats GPT-6 Luna on its own benchmarks. Here is the real pricing, context, and effort-level breakdown for both cheap-tier models.

Claude Haiku 5.5 announcement graphic on an orange Anthropic background
Claude Haiku 5.5 launch graphic. Image: Anthropic.

Claude Haiku 5.5 is the cheaper of the two budget models on paper, and Anthropic's own benchmark table backs that up on the three tests it chose to publish. For prompts under 100,000 tokens, Haiku 5.5 costs $0.10 per million input tokens and $0.50 per million output tokens, matching GPT-6 Luna's $0.10 input rate and tying it on output at the short-context tier. The real gap shows up past 100K tokens, where Haiku 5.5 jumps to $0.50 input and $2.50 output while Luna only doubles to $0.20 input and $0.75 output. On the benchmarks Anthropic ran itself — OSWorld 2.1's offline subset, GDPval-AA v2.1, and Terminal-Bench 4.0 — Haiku 5.5 beat Luna on all three. Those numbers come from Anthropic's own release page, not an independent lab, so treat them as a vendor's self-reported comparison rather than neutral ground truth.

The bottom line

Both companies now sell a "cheap tier" model priced to make per-token cost nearly irrelevant for most apps, and both released theirs within the same few weeks: Anthropic shipped Claude Haiku 5.5 on October 7, 2026, and OpenAI's GPT-6 Luna had already been live for about two weeks at that point. The pitch from both sides is nearly identical — fast, disposable-feeling, good enough for classification, extraction, routing, and high-volume agent work where a frontier model would be overkill and overpriced. The question this piece answers is whether "cheap" actually means the same thing at both companies, and whether Anthropic's claim of beating Luna on benchmarks holds up when you read the fine print.

Quick facts
  • Claude Haiku 5.5: $0.10 / $0.50 per million input/output tokens (≤100K context), $0.50 / $2.50 above 100K; 1M token context window; 128K max output; released October 7, 2026.
  • GPT-6 Luna: $0.10 / $0.50 per million input/output tokens (short context), $0.20 / $0.75 above 272K tokens; roughly 1.05M token context window (922K max input); 128K max output.
  • On Anthropic's own benchmark table, Haiku 5.5 beats Luna on OSWorld 2.1 offline (72.4% vs 48.9%), GDPval-AA v2.1 (1620 vs 1437 Elo), and Terminal-Bench 4.0 (39.2% vs 16.4%).
  • Those are Anthropic-run numbers at the model's "max" effort setting, not an independent benchmark.

Pricing, side by side

Pulled directly from Anthropic's pricing page and OpenAI's developer pricing page, here is the full rate card for both models' standard API tier:

MetricClaude Haiku 5.5GPT-6 Luna
Input, short context$0.10 / MTok (≤100K tokens)$0.10 / MTok (≤272K tokens)
Output, short context$0.50 / MTok$0.50 / MTok
Input, long context$0.50 / MTok (>100K tokens)$0.20 / MTok (>272K tokens)
Output, long context$2.50 / MTok$0.75 / MTok
Cached input read$0.01 / MTok (short) / $0.05 (long)$0.01 / MTok (short) / $0.02 (long)
Cache write$0.125 / MTok (short) / $0.625 (long)$0.125 / MTok (short) / $0.25 (long)
Batch discount50% offFlex/Batch tier: $0.05 input / $0.25 output
Context window1,000,000 tokens~1,050,000 tokens (922K max input)
Max output128,000 tokens128,000 tokens

The headline rates are identical at the short-context tier: a dime per million tokens in, fifty cents per million out. That symmetry isn't a coincidence — both companies have spent the back half of 2026 racing each other's list prices down, and Luna itself replaced an earlier $0.20/$1.20 "Luna" tier on September 22. Where the two diverge is long-context pricing. Anthropic's jump from 100K tokens is steep: input triples and output goes up 5x. OpenAI's jump from 272K tokens is gentler — input only doubles and output only rises 50%. If your workload regularly crosses into long-context territory (large codebases, long documents, big retrieval dumps), Luna's rate card is friendlier even though its threshold is higher to begin with.

What Claude Haiku 5.5 actually is

Haiku 5.5 is the small model in Anthropic's 5.5 generation, sitting below Sonnet 5.5 and Opus 5.5. According to Anthropic's own model documentation, it carries a 1 million token context window and a 128,000 token maximum output on the synchronous Messages API (300K on the batch API with a beta header). Its training/knowledge cutoff is listed as June 2026. Anthropic frames it, in its own words, as "the cheapest, fastest, and most capable small model we've ever released," and it's built for "high-volume, latency-sensitive tasks such as classification, extraction, and routing."

The headline new feature is effort control: Haiku 5.5 is the first Haiku-generation model that lets a developer pick how hard the model thinks before answering, with levels of low, medium, high, xhigh, and max, and medium as the API default. That single parameter is doing a lot of the benchmark lifting discussed below — Anthropic's published scores are run at max effort, not the default a typical integration would ship with.

What GPT-6 Luna actually is

Luna is OpenAI's small, fast tier within the GPT-6 family, sitting below GPT-6 Sol and GPT-6 Astra on both price and, per Anthropic's chart, benchmark scores. Per OpenAI's own model page, Luna has a roughly 1.05 million token context window (with a 922,000 token cap on input specifically) and a 128,000 token maximum output — functionally matching Haiku 5.5's output ceiling and landing in the same context-window neighborhood. OpenAI describes it as "our most efficient model for focused, high-volume tasks," with a knowledge cutoff of May 18, 2026. The two companies are explicitly chasing the same customer: whoever is running thousands of cheap, fast, repetitive model calls a day and doesn't want frontier-model pricing attached to each one.

The benchmark comparison, and its asterisk

Anthropic's own release page for Haiku 5.5 publishes a comparison table that puts GPT-6 Luna directly alongside Haiku 5.5, Haiku 4.5, and Sonnet 5.5 "for reference." These are the three results Anthropic chose to show:

BenchmarkHaiku 5.5Haiku 4.5GPT-6 LunaSonnet 5.5 (reference)
OSWorld 2.1, offline subset (computer use)72.4%15.7%48.9%83.9%
GDPval-AA v2.1 (Elo, professional knowledge work)162073514371840
Terminal-Bench 4.0 (agentic terminal tasks)39.2%0.0%16.4%70.6%

Three things are worth flagging before you treat this as settled. First, it's a vendor comparison: Anthropic ran all four models through its own evaluation harness and published the results on its own marketing page, which is standard practice in the industry but is not the same as an independent benchmark lab publishing a neutral leaderboard. Second, Anthropic's numbers for Haiku 5.5 are reported at its highest effort setting. Third-party coverage of the launch has noted that Haiku 5.5's GDPval-AA score drops considerably at the medium effort level that ships as the API default, so the gap over Luna you'd see in a default, out-of-the-box integration is almost certainly smaller than the max-effort table suggests — Anthropic itself hasn't published a medium-effort column in its public comparison. Third, no independent lab had published its own Haiku 5.5 vs. Luna numbers at the time of writing, so there's no outside check on any of these figures yet.

None of that means the comparison is wrong. It means it's Anthropic grading its own homework against a competitor, on tests Anthropic picked, at a setting that costs more than the default. That's a normal way for a model vendor to make its case, and the direction of the result — Haiku 5.5 ahead of Luna on agentic and computer-use tasks — is plausible given how much ground Haiku 4.5 made up in a single generation. But "Anthropic says it wins" and "it wins" are different claims, and this piece can only verify the first one.

Effort levels change the real-world price, not just the quality

The number that matters most for a budget-tier model usually isn't the per-token list price, it's the total cost of a finished task, and that's where Haiku 5.5's effort dial complicates a simple price comparison. Running at max effort to match Anthropic's benchmark numbers means generating more thinking tokens before the model answers, and thinking tokens are billed as output tokens at the same $0.50 (or $2.50) per-million rate. A classification job run at low effort might cost a fraction of a cent; the same job pushed to max effort to squeeze out Anthropic's published accuracy could cost meaningfully more, even though the list price per token never changed. Luna doesn't publish an equivalent effort dial on its model page, so its listed price is closer to a flat rate regardless of how OpenAI tunes reasoning internally. If you're comparing the two purely on sticker price, Haiku 5.5's number is a floor, not a fixed cost, in a way Luna's currently isn't.

Context window and output limits, in practice

On paper the two are close: both top out at 128,000 output tokens, and both advertise roughly a million-token context window. The small print differs slightly. Anthropic states Haiku 5.5's 1M context plainly, while OpenAI's own model page for Luna specifies a combined context of about 1.05 million tokens but caps the input portion specifically at 922,000 tokens, with the rest presumably reserved for output and reasoning within that same window. For most real workloads — a long support thread, a sizeable codebase, a multi-document retrieval batch — neither limit will bind before your actual use case does, and the more relevant decision point is the pricing tier breakpoint (100K for Haiku 5.5, 272K for Luna) discussed above, since that's what decides which row of the rate card you're actually paying.

Which cheap model should you actually use

If your workload stays under roughly 100K tokens per request — most classification, extraction, routing, and short-agent-step use cases — the two are priced identically on paper, so the decision comes down to which model's accuracy and latency you prefer for your specific task, which argues for running your own small eval rather than trusting either vendor's chart. If your workload regularly spans long documents or large codebases past that threshold, Luna's gentler long-context pricing curve gives it a real cost edge over Haiku 5.5 regardless of benchmark scores. And if the task genuinely benefits from more reasoning depth — multi-step computer use, agentic terminal work, the kind of thing OSWorld and Terminal-Bench are built to test — Anthropic's own data suggests Haiku 5.5 at higher effort levels is the stronger of the two cheap options, though you're spending some of your price advantage to get there. For teams already standardized on one ecosystem, this is also a good moment to revisit how Haiku 5.5 slots in against Anthropic's own larger models; our Haiku 5.5 vs. Sonnet 5.5 vs. Opus 5.5 comparison covers when it's worth paying more within the same family, and our look at GPT-6 Sol vs. Astra pricing and access does the same on OpenAI's side for teams that might outgrow Luna.

What next

Treat this comparison as a starting point, not a verdict. The pricing numbers are directly verifiable on both companies' own pages and are accurate as of this writing, but list prices change quickly in this market — both of these models already replaced a predecessor at a different price within the last month, and another cut is plausible before year's end. The benchmark table is Anthropic's own and hasn't been independently replicated, so if a decision of real consequence rides on the Haiku-vs-Luna gap, run both models against a sample of your actual prompts at the effort level you'd actually ship, rather than relying on anyone's marketing page, including this one. Watch Anthropic's Haiku 5.5 announcement and OpenAI's Luna model page directly for pricing changes, since aggregator sites tend to lag the official numbers by days.

Frequently asked questions

Is Claude Haiku 5.5 cheaper than GPT-6 Luna?

At the short-context tier, they're priced identically: $0.10 per million input tokens and $0.50 per million output tokens on both models' standard API rate. Above each model's context breakpoint — 100K tokens for Haiku 5.5, 272K for Luna — Luna's long-context rates rise less steeply, making it the cheaper option for very long prompts.

What is the context window for Claude Haiku 5.5 and GPT-6 Luna?

Claude Haiku 5.5 has a 1 million token context window, per Anthropic's documentation. GPT-6 Luna's own model page lists a roughly 1.05 million token context window with a 922,000 token cap on input specifically. Both cap output at 128,000 tokens.

What is effort control on Claude Haiku 5.5?

Effort control is a parameter that lets developers choose how much the model reasons before responding, with levels of low, medium, high, xhigh, and max. It's the first Haiku-generation model to expose this setting, and medium is the API default — Anthropic's published benchmark scores versus GPT-6 Luna were run at max effort.

Did Anthropic's benchmarks against GPT-6 Luna come from an independent lab?

No. The OSWorld 2.1, GDPval-AA v2.1, and Terminal-Bench 4.0 scores comparing Haiku 5.5 and GPT-6 Luna come from Anthropic's own release page and evaluation harness, not a third-party benchmark lab. No independently run comparison of the two models had been published at the time of writing.

Which model should I use for a high-volume, low-cost workload?

If your prompts stay under about 100,000 tokens, the two are priced the same, so pick based on your own accuracy and latency testing. For workloads that regularly exceed that length, GPT-6 Luna's gentler long-context pricing curve gives it a cost edge.

What did Claude Haiku 5.5 and GPT-6 Luna replace?

Haiku 5.5 replaced Haiku 4.5 in Anthropic's lineup on October 7, 2026. GPT-6 Luna replaced an earlier, more expensive Luna tier on September 22, 2026, when OpenAI cut its small-model pricing.

Sources

More on Claude →Claude Haiku 5.5GPT-6 LunaAnthropicOpenAIAI pricing
Theo Park
Written byTheo Park

Theo Park runs the AI desk at Pandromeda. He follows model launches from the frontier labs and the open-weight community, tracks the assistants and developer tools built on them, and explains what each release changes on pricing, capability and safety. His reporting leans on primary sources: model cards, technical reports, API documentation and the companies' own announcements.

More from AI

See all