AI/News

Gemini 4 Argon: Google's New Flagship AI Model Explained

Google's new flagship model adds a 1-million-token output limit, a staged Fairwind cyber-defense rollout, and introductory pricing of $2/$10 per million tokens.

Gemini 4 Argon key art from Google DeepMind's announcement
Gemini 4 Argon key art. Image: Google.

Google DeepMind introduced Gemini 4 Argon on September 30, 2026, positioning it as the company's new frontier flagship model for long, complex workflows in software engineering, enterprise knowledge work, and cybersecurity defense. Argon ships with an industry-leading 1-million-token output limit, introductory API pricing of $2 per million input tokens and $10 per million output tokens, and a staged rollout that starts with vetted cyber defenders in Google's Fairwind Program before reaching paid API customers and Google AI Ultra subscribers.

Key facts
  • Announced: September 30, 2026, by Google DeepMind
  • Output context: 1,000,000 tokens, up from 64K on prior Gemini models
  • Introductory price: $2 / million input tokens, $10 / million output tokens; cached input tokens at a 95% discount
  • Standard price (after intro period): $4 / million input tokens, $20 / million output tokens
  • First access: Trusted cyber defenders via the Fairwind Program, without cyber guardrails
  • Next access: Paid API customers and Google AI Ultra subscribers, no fixed date given
  • Headline benchmarks: DeepSWE v1.1 (77.9%), AutomationBench #1 (51.3%), CWE-bench v1 tied-first (68%), LVBench state of the art (91.7%)

What is Gemini 4 Argon?

Gemini 4 Argon is Google DeepMind's successor to the Gemini 3.x line, built specifically to sustain deep reasoning across long-horizon tasks rather than single-turn prompts. In its announcement post, Google DeepMind SVP and Chief AI Architect Koray Kavukcuoglu described Argon as delivering "frontier performance in complex workflows across real-world software engineering, enterprise knowledge work like legal and finance, and cybersecurity defense." Google says the model is already running inside its own engineering organization, where staff have used it for everyday debugging, large-scale codebase migrations, and algorithm design.

Among the internal use cases Google has disclosed: a 40% improvement over a published baseline on a quantum algorithmic optimization problem, more than 300 TiB of memory freed across internal systems (with projected savings in the 500 TiB-to-1 PiB range), a 2.7x speed-up on a SIMD-optimized video codec library (libgav1), and large migrations of C and C++ code to Rust spanning more than 800,000 lines, including components of the Fuchsia operating system's Zircon kernel.

Google frames these as evidence that Argon can sustain the kind of multi-hour, multi-file engineering work that previously required a human to break a project into smaller chunks and stitch results back together across many separate model calls. The company says the model is also being used internally for financial research, legal drafting, and chart analysis — tasks that mix long-document reasoning with structured output rather than pure code generation, which is part of why Google is positioning Argon as a general knowledge-work model rather than a coding-only release.

Pricing and availability

Google is launching Argon at an introductory rate of $2 per million input tokens and $10 per million output tokens, with cached input tokens discounted 95% off the standard input price. Once the introductory period ends, pricing steps up to $4 per million input tokens and $20 per million output tokens — a detail Google disclosed directly in the model's announcement post rather than on a separate pricing page, since Argon has not yet reached general API availability.

That rollout is deliberately staggered. Argon is currently available only to a limited set of trusted cyber defenders through the Fairwind Program, Google's application-only access tier for governments, healthcare providers, telecommunications operators, and other critical-infrastructure organizations, which Google has also detailed in a separate post on proactive cyber defense. Google says it is simultaneously participating in the United States government's voluntary pre-release review process for advanced models, and that it will expand access "to developers, enterprises, and consumers as soon as possible," starting with paid API customers and Google AI Ultra subscribers once that review and early-tester feedback are folded into the model's guardrails.

TierInput (per 1M tokens)Output (per 1M tokens)Notes
Introductory$2$10Cached input tokens: 95% off input price
Standard (post-introductory)$4$20Applies once the introductory period expires

Why cyber defenders get it first

Google trained Argon specifically to be capable at defensive cybersecurity work, and says the model can autonomously find, validate, and patch critical software vulnerabilities. For trusted Fairwind partners and Google's own internal security teams, Argon is being released "without cyber guardrails," giving those vetted users its full frontier-level offensive-adjacent capability for legitimate defense work — a deliberate trade-off Google frames as necessary so defenders are not operating with a weaker tool than attackers might eventually access elsewhere.

Fairwind members can pair Argon with CodeMender, Google's automated vulnerability-patching system, to move from discovery straight to a deployment-ready fix. Google points to one early example from security firm Wiz, which used Argon through its Scan for Good initiative — a free program that hunts for high-risk exposures in critical public infrastructure — to uncover a critical vulnerability exposing sensitive personal data in hospital software used worldwide, a flaw Google says "previous frontier models had missed." The Fairwind Program itself, which Google introduced in September 2026 alongside Gemini 3.8 Flash Cyber, says it already counts several hundred partner organizations globally.

The benchmark claims

Google's announcement leans heavily on a set of third-party and internal evaluations to argue Argon is its strongest model yet for agentic, long-horizon work. On DeepSWE v1.1, which measures real-world, long-horizon software engineering performance, Google says Argon sets a new state of the art with a score of 77.9%. On the Vals Index, which weights economic impact across finance, coding, legal, and tax work by each sector's contribution to U.S. GDP, Google describes Argon as the leading model, alongside similarly strong results on Vals' Finance Agent v2 and Harvey's Legal Agent Benchmark.

On AutomationBench, a Zapier-built benchmark measuring end-to-end execution across core business functions, Google says Argon ranks #1 with a score of 51.3%. For long-video understanding, Google cites a state-of-the-art score of 91.7% on LVBench. On the security side, Argon ties for first place on CWE-bench v1 — which evaluates a model's ability to remediate security vulnerabilities — with a score of 68%, building on the frontier result Google says its earlier Gemini 3.8 Flash Cyber model posted on CWE-bench v0. The DeepSWE benchmark itself is maintained independently at deepswe.datacurve.ai, where its leaderboard and task methodology are published. Google also says Argon leads on Gray Swan's Indirect Prompt Injection (IPI) benchmark, a measure of resistance to a common agentic-AI attack technique, and that it outperforms 3.8 Flash Cyber on Google's internal vulnerability-discovery benchmark and on Wiz's black-box penetration-testing benchmark.

All of these figures are Google's own reported results from its launch materials; Pandromeda has not independently reproduced them, and Google's post does not name rival models such as OpenAI's or Anthropic's systems directly when presenting the scores.

BenchmarkWhat it measuresArgon's reported result
DeepSWE v1.1Real-world, long-horizon software engineering77.9% (new state of the art, per Google)
AutomationBench (Zapier)End-to-end business-process execution#1 rank, 51.3%
CWE-bench v1Security vulnerability remediationTied for first, 68%
LVBenchLong video understanding91.7% (state of the art, per Google)
Vals IndexEconomic-impact-weighted finance/coding/legal/tax workLeading model, per Google
Gray Swan IPIIndirect prompt-injection robustnessLeading performance, per Google

What the 1-million-token output limit changes

Argon's most distinctive technical change is a jump in output context from the 64,000-token ceiling on earlier Gemini models to 1 million tokens. Google frames this as giving the model "headroom to think deeply and generate hundreds of thousands of tokens in a single trajectory," which it says adds a new level of depth to reasoning on hard, multi-step problems rather than forcing the model to truncate or split long coding sessions, large document rewrites, or multi-stage agentic workflows across several separate calls. For teams already comparing context-window strategies across Gemini's recent Flash-tier releases, Argon represents a parallel, heavier-weight track aimed at the hardest agentic tasks rather than everyday chat latency.

Safety and misuse mitigations

Because Argon is designed to operate with unusually few restrictions for vetted cyber-defense users, Google says it layered in several additional safety measures before any release: refusal training intended to block misuse while preserving legitimate security research, internal activation monitoring meant to flag misuse patterns, robustness testing from both internal and external red teams, chain-of-thought and action monitoring aimed at catching misalignment before a task goes beyond what a user intended, sandboxed execution environments in line with Google's broader agent-control roadmap, and automated red-teaming specifically targeting prompt-injection attacks.

How Argon fits against GPT-6 Astra and Claude Opus 5.5

Argon arrives into a frontier-model field that has moved fast in 2026. OpenAI's GPT-6 Astra and Anthropic's Claude Opus 5.5 both launched as their respective companies' top-tier flagship models earlier this year, and both are frequently benchmarked against each other and against Google's Gemini line by outside reviewers — Pandromeda's own comparison of Claude Opus 5.5 pricing and benchmarks and its three-way look at Gemini, GPT-6, and Claude Opus 5.5 lay out how the three labs' flagship and mid-tier models have been stacking up on price and throughput. Google's own Argon announcement does not name either rival model directly or publish head-to-head scores against them; its benchmark claims are framed only against Google's prior Gemini models (such as 3.8 Flash Cyber) and against third-party leaderboards like DeepSWE, AutomationBench, and the Vals Index, where Argon's reported results are presented as leading or state-of-the-art without naming which other systems were tested on the same boards.

What is clear from Google's materials is the strategic positioning: Argon's 1-million-token output window and its specific tuning for coding, legal/finance reasoning, and cybersecurity defense are aimed squarely at the same long-horizon, agentic-workflow use cases that OpenAI and Anthropic have both emphasized with their most recent flagship releases. Pricing-wise, Argon's introductory $2/$10 per-million-token rate undercuts the rate it will settle into once the introductory window closes ($4/$20), a pattern Google has also used with earlier Gemini launches to seed early adoption ahead of standard pricing.

What this means for developers and enterprises right now

For most developers and businesses, Argon is not yet something to build against today. Access is currently restricted to the Fairwind Program's vetted cyber-defense partners, and Google has not published a firm date for when paid API customers or Google AI Ultra subscribers will get access. Teams planning around Gemini should treat the published $2/$10 introductory pricing and $4/$20 standard pricing as the numbers to budget against once general availability arrives, and should expect the model's 1-million-token output ceiling to matter most for long agentic sessions, large codebase-scale refactors, and multi-document knowledge work rather than short conversational queries, where the extra headroom offers little benefit over existing Gemini models.

Organizations that already run security or DevOps teams eligible for the Fairwind Program have the clearest near-term path to hands-on access, since Google is explicitly prioritizing governments, healthcare providers, telecommunications operators, and other critical-infrastructure operators in this phase. Everyone else — including teams evaluating Argon purely as a coding or knowledge-work model rather than a security tool — will need to wait for the paid-API and Google AI Ultra rollout Google has promised but not yet scheduled. In the meantime, the 95% cached-token discount is worth noting for planning purposes: workflows that repeatedly reuse large context (a long codebase, a legal corpus, a knowledge base) will cost meaningfully less per request than the headline per-token rate suggests once Argon reaches general availability.

What's next

Google says it will keep gathering feedback from Fairwind's early cyber-defender cohort and keep iterating on Argon's guardrails before widening access. The next milestone to watch is general availability through the paid Gemini API and Google AI Ultra, which Google has committed to doing "as soon as possible" without giving a calendar date. Until then, expect comparisons to GPT-6 Astra and Claude Opus 5.5 to be driven largely by third-party testing rather than head-to-head numbers from Google itself, since the current announcement keeps its benchmark claims framed against Google's own prior models and independent leaderboards rather than named competitors.

Frequently asked questions

What is Gemini 4 Argon?

Gemini 4 Argon is Google DeepMind's new frontier flagship AI model, announced September 30, 2026, built for long-horizon software engineering, enterprise knowledge work, and cybersecurity defense tasks, with a 1-million-token output limit.

How much does Gemini 4 Argon cost?

Google's introductory pricing is $2 per million input tokens and $10 per million output tokens, with cached input tokens discounted 95%. After the introductory period, pricing rises to $4 per million input tokens and $20 per million output tokens.

Who can use Gemini 4 Argon right now?

Access is currently limited to trusted cyber defenders through Google's Fairwind Program. Google says it will expand access to paid API customers and Google AI Ultra subscribers next, though no firm date has been given.

What is the Fairwind Program?

Fairwind is Google's application-only early-access tier that gives governments, healthcare providers, telecommunications operators, and other critical-infrastructure organizations early access to advanced AI models for cyber defense before general release.

What benchmarks does Google cite for Gemini 4 Argon?

Google cites a new state of the art on DeepSWE v1.1 (77.9%), the #1 rank on AutomationBench (51.3%), a tied-first score on CWE-bench v1 (68%), and a state-of-the-art score on LVBench (91.7%), among other results.

How does Gemini 4 Argon compare to GPT-6 Astra and Claude Opus 5.5?

Google's announcement does not name OpenAI's GPT-6 Astra or Anthropic's Claude Opus 5.5 directly or publish head-to-head scores against them; Argon's benchmark claims are framed against Google's own prior models and independent third-party leaderboards.

Sources

More on Gemini 4 Argon →Gemini 4 ArgonGoogle DeepMindAI modelsFairwind ProgramGemini pricing
Theo Park
Written byTheo Park

Theo Park runs the AI desk at Pandromeda. He follows model launches from the frontier labs and the open-weight community, tracks the assistants and developer tools built on them, and explains what each release changes on pricing, capability and safety. His reporting leans on primary sources: model cards, technical reports, API documentation and the companies' own announcements.

More from AI

See all