AI/Comparison

Gemini 4 Argon vs GPT-6 Astra vs Claude Opus 5.5: Price and Benchmarks

Google priced Gemini 4 Argon below GPT-6 Astra and Claude Opus 5.5, but access, context window, and benchmark disclosures differ sharply across all three.

Key art for Google's Gemini 4 Argon announcement
Image: Google.

Gemini 4 Argon, Google's new frontier model, launches at $2 per million input tokens and $10 per million output tokens with a headline 1 million-token output limit, undercutting GPT-6 Astra's $10/$50 per-million pricing and landing well below Claude Opus 5.5's $4/$20 rate — but Argon is currently restricted to a small group of cybersecurity testers, while Astra and Opus 5.5 are already generally available through their makers' APIs. Here is what each company has actually published, with the gaps in Google's disclosure called out explicitly.

The short version
  • Gemini 4 Argon: $2/$10 per million tokens (input/output, introductory), 1M-token output limit, rolling out only to Google's "Fairwind Program" cyber-defense testers as of this week.
  • GPT-6 Astra: $10/$50 per million tokens, 1,050,000-token context window, generally available on ChatGPT and the OpenAI API since September 3, 2026.
  • Claude Opus 5.5: $4/$20 per million tokens, context window Anthropic lists as "up to 1M," generally available since September 22, 2026 on the Claude Platform, AWS Bedrock, Google Cloud, and Azure.

What Gemini 4 Argon actually changes

Google's own announcement is narrower than the "Gemini 4" hype cycle that preceded it. In its post introducing the model, Google says it is "significantly expanding the model's output token limit to an industry-leading 1M tokens, up from the previous 64K tokens" — that is an output ceiling, not the input context window most buyers care about when they say "context window." Google's post does not publish an input token limit for Argon at all. If you've seen figures of 2 million or 10 million input tokens floating around, those trace back to unconfirmed leaks and LMArena checkpoint-spotting from mid-September, not to anything Google has published. Treat them as rumor, not spec.

What Google did publish: Argon is launching first through a limited-access "Fairwind Program" aimed at cyber defenders, with the company saying it is releasing the model "without cyber guardrails" for that vetted group so they can use its full offensive-and-defensive security capability. Paid API customers and Google AI Ultra subscribers are next in line, with broader developer, enterprise, and consumer access promised "as soon as possible" — no date attached. For a fuller walkthrough of the rollout, see our explainer on Gemini 4 Argon.

Pricing compared: Argon vs Astra vs Opus 5.5

All three companies price per million tokens, but the structures differ enough that a straight input/output comparison undersells the real cost. Astra and its cheaper sibling GPT-6 Sol both double input/cache pricing and apply a 1.5x output multiplier once a request crosses 272,000 input tokens — a detail that is easy to miss when just quoting the headline rate.

ModelInput $/MOutput $/MCached input $/MNotes
Gemini 4 Argon$2.00$10.0095% off input rate (~$0.10)Introductory pricing; Google hasn't said how long it holds
GPT-6 Astra$10.00$50.00$1.00Rates double (input) and rise 1.5x (output) above 272K input tokens
GPT-6 Sol$2.00$10.00$0.20OpenAI's mid-tier model, same 272K long-context surcharge rule
Claude Opus 5.5$4.00$20.00$0.20 (read) / $5 (write)Fast mode: $8/$40; US-only inference at 1.1x

On sticker price, Argon's introductory rate matches GPT-6 Sol almost exactly and comes in at a fifth of Astra's cost per token. Anthropic's Opus 5.5 sits in between: twice Argon's and Sol's rate, but less than half of Astra's. None of this accounts for how many tokens each model actually needs to finish a task, which is where benchmark efficiency claims (below) start to matter more than the rate card.

Context window and output limits

This is the category where the three companies are hardest to compare directly, because they disclose different numbers.

  • Gemini 4 Argon: Google confirms a 1,000,000-token output limit (up from 64K in prior models) and does not publish an input context window figure for Argon in its launch post.
  • GPT-6 Astra: OpenAI's model documentation lists a 1,050,000-token total context window, with a maximum of 922,000 input tokens and 128,000 output tokens.
  • Claude Opus 5.5: Anthropic's pricing page lists context window as "up to 1M," varying by model, without a separate published maximum-output figure for Opus 5.5 specifically.

Net: Astra is the only one of the three with a fully itemized input/output split on the record. Argon's standout number is a 1M-token output ceiling, aimed at letting the model write much longer single responses (think full codebases or long reports in one pass) rather than read more in a single prompt.

Benchmark results, side by side

Each company reports its own benchmark suite, and the overlap between them is small — a recurring problem when comparing frontier models. One benchmark both Google and Anthropic happened to report is AutomationBench, a business-workflow-automation test:

BenchmarkGemini 4 ArgonGPT-6 AstraClaude Opus 5.5
AutomationBench (business workflows)51.3% (Google says #1)not published by OpenAI40.0%
DeepSWE v1.1 (software engineering)77.9%not publishednot published
FrontierMath Tier 4not published98%not published
ARC-AGI-3not published99.9% on a custom "provider adapter" rig; 62.7% on the standard semi-private rignot published
Terminal-Bench (agentic coding)not published on this benchmarknot published on this exact version66.4% (Terminal-Bench 4.0)
Computer-use (OSWorld family)not publishedcited as a comparison point in Argon's own post81.8% (OSWorld 2.1)

Two things worth flagging on honesty grounds. First, Google's own post says Argon "matches GPT-6 Astra's score for its Intelligence Index at 60 percent of the cost per task at current discounted prices" — that's Google's characterization of a third-party aggregate index, not an independent test, and OpenAI hasn't published a response to it. Second, OpenAI's own ARC-AGI-3 figure for Astra needs its asterisk read in full: the 99.9% came from a custom test harness built around OpenAI's native context-management tools, while the same model scored 62.7% on the standard semi-private rig most other labs use. That's a huge gap, and it means the 99.9% number isn't directly comparable to how other labs typically report ARC-AGI-3.

Where each model wins, by the numbers that exist

Going only off what's been published:

  • Cheapest frontier access: Gemini 4 Argon and GPT-6 Sol tie at $2/$10 per million tokens — both a fifth of Astra's rate.
  • Longest documented context window: GPT-6 Astra, at 1.05M tokens total with a clearly itemized 922K input / 128K output split.
  • Longest single-response output: Gemini 4 Argon, at 1M output tokens — more than 7x Astra's 128K output cap, on paper.
  • Agentic-coding benchmark disclosed with the most detail: Claude Opus 5.5, which published scores across nine named benchmarks (Terminal-Bench 4.0, FrontierCode v1.1, CursorBench 4.0, GDPval-AA v2.1, Humanity's Last Exam, OSWorld 2.1, Chartography, AutomationBench, and Terminal-Bench-Science) in one release.
  • General availability today: Astra and Opus 5.5 both ship now through standard commercial APIs; Argon does not — it's gated to Google's cyber-defender program.

For a deeper look at how Opus 5.5 stacks up against Anthropic's own smaller model, see our Opus 5.5 vs Sonnet 5.5 comparison.

Availability and API access

This is arguably the biggest practical difference between the three models right now, and it's easy to miss if you only read pricing tables.

ModelStatusWhere to get it
Gemini 4 ArgonLimited accessFairwind Program (vetted cyber defenders) only; API/consumer rollout not yet dated
GPT-6 AstraGenerally availableChatGPT Plus/Pro/Business/Enterprise, OpenAI API, Microsoft Azure, AWS Bedrock
GPT-6 SolGenerally availableOpenAI API (Chat Completions, Responses), replaced GPT-5.6 Sol on September 22, 2026
Claude Opus 5.5Generally availableClaude Platform (model ID claude-opus-5-5), AWS Bedrock, Google Cloud, Microsoft Azure

If you need a frontier-class model you can actually call in production today, the practical shortlist is GPT-6 Astra or Claude Opus 5.5 — not Argon, regardless of how its price card compares. That gating detail also explains why so many "Argon vs Astra" and "Argon vs Opus 5.5" comparisons circulating right now lean on leaked benchmark screenshots rather than hands-on API testing: outside of Google's vetted cyber-defender cohort, nobody with a credit card can actually run Argon side by side against the other two yet. Any claim you see about Argon's real-world coding speed, latency, or agentic reliability compared to Astra or Opus 5.5 should be read as unverified until Google opens broader access and independent benchmarking groups get a turn at it.

Gemini 4 Argon vs GPT-6 Sol, the closer price match

Because Argon's introductory rate ($2/$10 per million) lines up almost exactly with OpenAI's cheaper GPT-6 Sol tier rather than flagship Astra, the more useful "budget" comparison is Argon vs Sol, not Argon vs Astra. Both share the 272K-token threshold rule that doubles input pricing on longer requests (Sol's structure mirrors Astra's here), and OpenAI has already said Sol "outscores Claude Opus 5 on a business workflow benchmark at 9% of the cost" — a claim made before Opus 5.5 existed, so it's now dated. Note also that OpenAI's developer docs point users toward an even newer "GPT-6.1 Sol" as the current version; we cover that update in our OpenAI DevDay 2026 roundup.

How this lines up against Claude Fable 5.1 and Gemini 3.8 Flash

Two other names keep coming up in the same searches as this comparison. Anthropic has said Opus 5.5 performs at roughly Claude Fable 5.1's level on most tasks while costing about 40% less to run — Fable remains Anthropic's pricier, separately positioned model, not something Opus 5.5 replaces outright. On the Google side, Argon's benchmark table is explicitly framed as an improvement over Gemini 3.8 Flash Cyber on vulnerability-discovery and cybersecurity tasks specifically, not as a general-purpose comparison. If you want the numbers on the model Argon is actually measured against in Google's own post, our Gemini 3.8 Flash vs GPT-6 vs Claude Opus 5.5 comparison covers the pricing and benchmarks for that generation.

What we don't know yet

Being precise about the gaps matters more than filling them in with guesses:

  • Google has not published Argon's input context window, a general-availability date, or a post-introductory price.
  • Google has not published how Argon performs on the same benchmarks OpenAI and Anthropic used (Terminal-Bench, FrontierMath, OSWorld), so none of the three vendors' headline scores line up on a single shared test.
  • OpenAI's 99.9% ARC-AGI-3 figure for Astra used a custom test rig; the comparable standard-rig score is 62.7%, and it's not clear which rig Google or Anthropic would use if they ran the same test on Argon or Opus 5.5.
  • Anthropic has not published a specific maximum output-token figure for Opus 5.5 the way OpenAI did for Astra and Sol.
  • None of the three companies has published independent, third-party-audited scores for this specific trio of models on the same day — every number in the tables above is self-reported by the model's own maker.
  • Pricing for Argon is explicitly labeled "introductory" by Google, which typically signals a price increase once the model exits limited access; Astra's and Opus 5.5's listed rates are their standing, non-promotional prices.

That last point matters for budgeting: teams that build around Argon's $2/$10 rate today should assume it is a launch discount, not a durable price floor, the same way GPT-6 Sol replaced GPT-5.6 Sol at half its predecessor's price only to have OpenAI later point developers toward a revised GPT-6.1 Sol release.

What to do next

If you need a model in production today, the decision isn't really a three-way tie: Argon is not generally purchasable yet, so the real choice right now is between GPT-6 Astra (longest documented context window, highest published cost) and Claude Opus 5.5 (cheaper, with the most detailed public benchmark disclosure of the three). Teams specifically doing cybersecurity defense work can apply to Google's Fairwind Program for early Argon access. Everyone else should watch for Google to publish an input-context figure and a general-availability date — until then, Argon's pricing is a preview, not a product you can build on. Check back on this page; we'll update the tables above as each company publishes more.

Frequently asked questions

Is Gemini 4 Argon available to the public yet?

No. As of this week Google is only granting access through its "Fairwind Program" for vetted cybersecurity defenders. Google says paid API customers and Google AI Ultra subscribers will get access next, with broader developer, enterprise, and consumer availability "as soon as possible" but no date has been set.

What is Gemini 4 Argon's context window?

Google has only published an output-token limit of 1 million tokens (up from 64K previously). Google's launch post does not state an input context window figure for Argon. Reports of a 2 million or 10 million token input context window are unconfirmed leaks, not official Google specifications.

How does Gemini 4 Argon's price compare to GPT-6 Astra and Claude Opus 5.5?

Argon launches at $2 per million input tokens and $10 per million output tokens, an introductory rate from Google. GPT-6 Astra costs $10/$50 per million tokens. Claude Opus 5.5 costs $4/$20 per million tokens. Argon is the cheapest of the three on paper, but it is not yet generally purchasable.

Which model has the longest context window: Argon, Astra, or Opus 5.5?

Based on official disclosures, GPT-6 Astra has the most clearly documented context window at 1,050,000 tokens total (922,000 input, 128,000 output). Anthropic lists Claude Opus 5.5's context window only as "up to 1M." Google has not published an input context window figure for Gemini 4 Argon.

Is GPT-6 Astra's 99.9% ARC-AGI-3 score comparable to other models?

Not directly. OpenAI reported that score using a custom "provider adapter" test rig built around its own context-management tools. On the standard semi-private test rig most labs use, OpenAI reported Astra scoring 62.7% on the same benchmark.

What is GPT-6 Sol, and how is it different from GPT-6 Astra?

GPT-6 Sol is OpenAI's lower-cost GPT-6 tier, priced at $2 per million input tokens and $10 per million output tokens, versus Astra's $10/$50 flagship pricing. Both share the same 1,050,000-token context window structure, but Astra is positioned as OpenAI's most capable model while Sol targets cost-sensitive coding and agentic workflows.

Sources

More on AI Model Benchmarks →Gemini 4 ArgonGPT-6 AstraClaude Opus 5.5AI benchmarksAI pricing
Theo Park
Written byTheo Park

Theo Park runs the AI desk at Pandromeda. He follows model launches from the frontier labs and the open-weight community, tracks the assistants and developer tools built on them, and explains what each release changes on pricing, capability and safety. His reporting leans on primary sources: model cards, technical reports, API documentation and the companies' own announcements.

More from AI

See all