AI/Comparison

Gemini 3.8 Flash vs GPT-6 vs Claude Opus 5.5: Benchmarks and Price

Three flagship AI models shipped in the last month. Here is how their prices, context windows and independent benchmark scores actually compare, use case by use case.

Abstract yellow and charcoal textured collage artwork from Anthropic's Claude Opus 5.5 launch page
Collage artwork from Anthropic's Claude Opus 5.5 announcement. Image: Anthropic.

Claude Opus 5.5 currently tops independent benchmark rankings and OpenAI and Anthropic's own coding evaluations, but it is also the most expensive of the three to run. GPT-6 Sol lands within single digits of Opus 5.5 on several of the same tests while costing roughly a fifth as much per task, and its sibling GPT-6 Luna and Google's Gemini 3.8 Flash both price input tokens under a dollar per million for high-volume work. All three shipped within the last four weeks — Gemini 3.8 Flash on September 2, GPT-6 Sol and Luna and Claude Opus 5.5 both on September 22 — so this is the first time all three have been directly comparable on price and benchmarks at once. Here is what each one actually costs, what it scores, and which one to reach for depending on the job.

Quick facts

  • Claude Opus 5.5 (Anthropic, Sept 22, 2026): $4 / $20 per million input/output tokens, 1M context, 128K max output, tops Anthropic's and Artificial Analysis's own benchmark boards.
  • GPT-6 Sol (OpenAI, Sept 22, 2026): $2 / $10 per million tokens, 1.05M context, 128K max output, OpenAI's mid-tier agentic/coding model.
  • GPT-6 Luna (OpenAI, Sept 22, 2026): $0.10 / $0.50 per million tokens, same 1.05M context and 128K output as Sol, built for high-volume tasks.
  • Gemini 3.8 Flash (Google DeepMind, Sept 2, 2026): $0.75 / $3.75 per million tokens through Dec 31, 2026 (then $1.50 / $7.50), 1M input context, 64K max output.

Two other names will come up in the benchmark charts below and are worth flagging upfront so they don't get confused with the three models this piece is actually comparing. GPT-6 Astra is OpenAI's most capable and most expensive GPT-6-tier model, positioned above both Sol and Luna. Claude Fable 5.1 is Anthropic's other current flagship line, which Anthropic says Opus 5.5 now matches "on most work" at a much lower cost. Neither is covered head-to-head here — see Pandromeda's GPT-6 Sol and Luna launch coverage, Claude Opus 5.5 launch coverage and Gemini 3.8 Flash launch coverage for the full specs on each individual model.

Pricing compared: cost per million tokens

On list price alone, the gap between these models is enormous — a 40x spread between the cheapest and most expensive input token. Anthropic's Claude Opus 5.5 costs $4 per million input tokens and $20 per million output tokens on standard API access, a 20% cut from Opus 5, with cached input reads discounted 90% to $0.20 per million and cache writes at $5 per million. Anthropic also offers a Fast mode inside Claude Code and the Claude Platform that runs up to 2.5x faster for roughly double the price, at $8 / $40 per million tokens.

OpenAI cut GPT-6 Sol and Luna's prices by 50% versus their GPT-5.6 promotional pricing at launch. Sol now costs $2 per million input tokens and $10 per million output tokens, with cached reads at $0.20 and cache writes at $2.50. Luna is $0.10 input / $0.50 output, with cached reads at $0.01 and cache writes at $0.125 — cheap enough that running the same workload on Luna instead of Opus 5.5 cuts the input-token bill by roughly 97%.

Google prices Gemini 3.8 Flash at an introductory $0.75 per million input tokens and $3.75 per million output tokens, a rate Google has said holds through December 31, 2026, after which it rises to $1.50 / $7.50. That puts Flash between Luna and Sol on cost, though its promotional window means the effective price will climb by 2026's end.

ModelInput / 1M tokensOutput / 1M tokensContext windowMax outputKnowledge cutoff
Claude Opus 5.5$4.00$20.001M tokens128K tokensJune 2026
GPT-6 Sol$2.00$10.001.05M tokens128K tokensApr 20, 2026
GPT-6 Luna$0.10$0.501.05M tokens128K tokensMay 18, 2026
Gemini 3.8 Flash$0.75 (until Jan 2027)$3.75 (until Jan 2027)1M tokens (input)64K tokensMarch 2026

Context window and output limits

All four models cluster around the 1-million-token mark for context, which is the more meaningful number for feeding a model a large codebase, a long research corpus, or a full contract set. Claude Opus 5.5 and both GPT-6 tiers all accept roughly 1M–1.05M tokens of input and can generate up to 128,000 tokens of output in a single response (Anthropic also offers an extended 300K-token output ceiling in its Batch API). Gemini 3.8 Flash matches the ~1M-token input ceiling but caps output at 64,000 tokens — half of what the other three allow — which matters for tasks that need the model to write very long documents or large diffs in one pass rather than for reading long inputs.

Intelligence benchmarks head-to-head

Each company publishes its own benchmark comparisons, and — because these models launched three weeks apart — those self-reported numbers don't all compare against the same rival generation. Google's Gemini 3.8 Flash launch benchmarks (published Sept 2) compare it against Claude Opus 5, Claude Sonnet 5 and GPT-5.6, since Opus 5.5 and GPT-6 Sol/Luna didn't exist yet. OpenAI's Sol/Luna launch benchmarks (Sept 22) compare against Claude Opus 5 and Claude Fable 5.1, not the newer Opus 5.5, which shipped the same day. For a single apples-to-apples read across all three current models, the clearest source is Artificial Analysis, a third-party benchmark aggregator that re-runs every model through the same evaluation suite.

On Artificial Analysis's rebuilt Intelligence Index (v4.3.2, which combines ten evaluations including Humanity's Last Exam, GDPval-AA and Terminal-Bench 4.0), Claude Opus 5.5 currently holds the top score the index has recorded for any model. GPT-6 Sol at its highest reasoning effort sits meaningfully behind it, with Gemini 3.8 Flash and GPT-6 Luna further back but still competitive for their price tier.

ModelAA Intelligence Index (v4.3.2)AA Coding Agent Index
Claude Opus 5.5 (max effort)58 — highest recorded66 — #1, at $13.04/task
GPT-6 Sol (max effort)4857, at $2.99/task
Gemini 3.8 Flash41not separately indexed by AA at time of writing
GPT-6 Luna (max effort)3741

Source: Artificial Analysis's GPT-6 Sol and Luna cost-efficiency writeup, a third-party benchmark index, not a Pandromeda-run test.

Coding and agentic tool-use performance

For coding agents specifically, Anthropic says Opus 5.5 scores 66.4% on Terminal-Bench 4.0 at its highest reasoning effort, against 52.3% for the previous Opus 5 and 57.9% for OpenAI's GPT-6 Astra as reported by OpenAI. On CursorBench 4.0, Anthropic puts Opus 5.5 at 57.8% versus 46.6% for Opus 5. Anthropic also reports Opus 5.5 completing a 680,000-line code migration in under a day during internal testing, and succeeding 39 of 40 times on a load-time optimization task where Opus 5 sometimes altered app behavior.

OpenAI's own DeepSWE v1.1 numbers (a benchmark for complex, real-codebase software engineering) put GPT-6 Sol at max effort at 68.8%, within 1.1 points of Claude Fable 5's best score (69.9%) at roughly 80% lower cost per task, according to OpenAI. GPT-6 Luna at max effort scores 66.6% on the same benchmark — comparable to Claude Opus 5 and Fable 5 at medium effort — while costing 93% less per task than Opus 5 and 96% less than Fable 5, per OpenAI's figures. Google, meanwhile, says Gemini 3.8 Flash "outperforms most larger frontier models" on DeepSWE v1.1 and scores 90.8% on Terminal-Bench 2.1 — a different benchmark version than the Terminal-Bench 4.0 figures Anthropic and Artificial Analysis use, which is a useful reminder that "Terminal-Bench" scores aren't always comparable across vendors unless the version number matches.

For agentic, multi-tool workflows, OpenAI's AutomationBench figures (which test 47 tools across business functions) show GPT-6 Sol at high reasoning effort scoring 33.2% at $0.27 per task, versus 26.9% for Claude Opus 5 at 11.1x the cost — though that comparison again predates Opus 5.5. On OSWorld 2.0, a computer-use benchmark, OpenAI reports GPT-6 Sol reaching 60.5% versus 60.3% for Claude Opus 5 at roughly 80% lower cost.

Long-context research and knowledge work

For tasks that lean on the full context window — digesting a long research corpus, a multi-file legal discovery set, or a large enterprise knowledge base — Anthropic's GDPval-AA v2.1 results put Opus 5.5 at 1,846 Elo, ahead of Opus 5 (1,708) and GPT-6 Astra (1,542) on Anthropic's benchmark. Google reports Gemini 3.8 Flash scoring highest among the models it tested on two professional-domain agent benchmarks: 61.4% on Vals' Finance Agent v2 (ahead of Claude Opus 5's 58.6% and GPT-5.6 Sol's 53.8%) and 10.0% on Harvey's Legal Agent Benchmark (ahead of Opus 5's 6.7%). On HLE-Verified, a multi-step reasoning test, Google puts Gemini 3.8 Flash marginally ahead at 54.9%, with GPT-5.6 Sol at 54.5% and Claude Opus 5 at 54.4% — a near-tie that predates both GPT-6 Sol and Opus 5.5. Since Gemini 3.8 Flash's output cap of 64K tokens is half the other models' 128K, it's better suited to reading and reasoning over long inputs than to producing very long single-pass outputs.

Where the vendors' own numbers disagree with each other

Because Anthropic, OpenAI and Google each publish comparisons against whichever rival models existed at their own launch date, none of the three companies' own blog posts contain a true three-way, same-day comparison — that only exists via a neutral third party like Artificial Analysis, which retests every model itself rather than relying on vendor-reported figures. Readers should also note that Google's own launch chart shows Gemini 3.8 Flash slightly ahead of Claude Opus 5 (not 5.5) and GPT-5.6 Sol (not GPT-6 Sol) on several evaluations — comparisons against superseded model generations that will look different once run against Opus 5.5 and GPT-6 Sol directly. Where Artificial Analysis's rebuilt Intelligence Index and each vendor's own self-reported scores disagree, this piece has used the vendor's own figures for that specific benchmark and cited whose number it is, and used Artificial Analysis's index for the one true side-by-side ranking.

Which model wins for which use case

For hard, high-stakes coding or agentic tasks where accuracy matters more than per-call cost — a large migration, a complex agent workflow, or work you'll only run once — Claude Opus 5.5 is the strongest performer on both Anthropic's own benchmarks and Artificial Analysis's independent index, and its price premium is easier to justify on a small number of expensive tasks.

For teams running agentic or coding workloads at volume, where cost per task compounds quickly, GPT-6 Sol is the more efficient choice: OpenAI and Artificial Analysis both show it landing within single digits of Opus 5.5-class performance on several coding and agentic benchmarks at roughly a fifth of the cost per task.

For long-context research, document-heavy knowledge work, and reasoning tasks where Google's own comparisons show it leading on price-adjusted performance, Gemini 3.8 Flash is a strong pick, with the caveat that its 64K output ceiling makes it less suited to generating very long single-pass documents or diffs.

For high-volume, cost-sensitive use — classification, extraction, simple agent steps repeated thousands of times a day — GPT-6 Luna is the cheapest of the four by a wide margin, at $0.10 / $0.50 per million tokens, while still scoring competitively on DeepSWE v1.1 against last generation's mid-tier models.

Availability and how to access each model

Claude Opus 5.5 is available now to Claude Pro, Max, Team and Enterprise subscribers, through the Claude Platform API as claude-opus-5-5, and via Amazon Web Services, Google Cloud and Microsoft Azure. GPT-6 Sol and GPT-6 Luna are live in ChatGPT Work and Codex for Plus, Pro, Business, Enterprise and Edu users, with Luna also available to Free and Go users in the desktop app; both are available in the OpenAI API as gpt-6-sol and gpt-6-luna. Gemini 3.8 Flash is available in the Gemini app for Pro and Ultra subscribers, in Google AI Studio and the Gemini API for developers, in Gemini Enterprise, and inside Google Search's AI Mode and Google Sheets.

Frequently asked questions

Which of the three models is cheapest to run?

GPT-6 Luna, at $0.10 per million input tokens and $0.50 per million output tokens. That undercuts Gemini 3.8 Flash's introductory $0.75/$3.75 rate and GPT-6 Sol's $2/$10, and is roughly 97% cheaper on input tokens than Claude Opus 5.5's $4/$20.

Which model has the largest context window?

GPT-6 Sol and GPT-6 Luna both accept up to 1.05 million tokens of input, slightly ahead of Claude Opus 5.5 and Gemini 3.8 Flash's 1 million. For output, Opus 5.5 and both GPT-6 tiers allow up to 128,000 tokens per response, versus 64,000 for Gemini 3.8 Flash.

Which model scores highest on independent benchmarks?

On Artificial Analysis's rebuilt Intelligence Index (v4.3.2), a third-party benchmark aggregator that retests every model on the same evaluation suite, Claude Opus 5.5 currently holds the highest score the index has recorded, ahead of GPT-6 Sol, Gemini 3.8 Flash and GPT-6 Luna in that order.

Can I trust the benchmark comparisons each company publishes on its own blog?

Treat them as directional rather than a fair three-way test. Gemini 3.8 Flash launched September 2 and its own launch benchmarks compare it to Claude Opus 5 and GPT-5.6, not the newer Opus 5.5 or GPT-6 Sol, which didn't exist yet. GPT-6 Sol and Luna's launch benchmarks (September 22) compare to Claude Opus 5 and Fable 5.1, not the same-day Opus 5.5. A neutral index like Artificial Analysis is the only source in this piece that retests all three current models under identical conditions.

Which model is best for coding agents specifically?

Claude Opus 5.5 posts the highest scores on Anthropic's own coding benchmarks and on Artificial Analysis's Coding Agent Index, but GPT-6 Sol lands within a few points of similar results at roughly a fifth of the cost per task, according to OpenAI's and Artificial Analysis's figures — making Sol the more efficient pick for high-volume coding-agent workloads.

Is Gemini 3.8 Flash's pricing going to change?

Yes. Google has said the $0.75 input / $3.75 output per-million-token introductory pricing holds through December 31, 2026, after which it rises to $1.50 input / $7.50 output.

Sources

More on AI Model Benchmarks →Gemini 3.8 FlashGPT-6 SolGPT-6 LunaClaude Opus 5.5AI benchmarksLLM pricing
Theo Park
Written byTheo Park

Theo Park runs the AI desk at Pandromeda. He follows model launches from the frontier labs and the open-weight community, tracks the assistants and developer tools built on them, and explains what each release changes on pricing, capability and safety. His reporting leans on primary sources: model cards, technical reports, API documentation and the companies' own announcements.

More from AI

See all