AI/News

Gemini 3.8 Flash Release Date, Price and Benchmarks Explained

Google's third Flash release in six weeks keeps 3.7 Flash pricing while posting higher agentic and reasoning benchmark scores, Google says.

Google's announcement graphic for Gemini 3.8 Flash and Gemini 3.8 Flash Cyber
Announcement graphic for Gemini 3.8 Flash and 3.8 Flash Cyber. Image: Google.

Gemini 3.8 Flash is Google's third Flash-tier model release in six weeks, launched on September 2, 2026, and it plugs straight into the same $0.75-per-million-input / $3.75-per-million-output pricing that Gemini 3.7 Flash already carried, while posting higher scores on Google's own agentic, coding and reasoning benchmarks. It ships alongside a locked-down sibling, Gemini 3.8 Flash Cyber, built specifically for vulnerability discovery and patching.

Gemini 3.8 Flash at a glance

  • Released September 2, 2026 — the third Flash update in six weeks, after 3.6 and 3.7 Flash
  • API price: $0.75/1M input tokens, $3.75/1M output tokens through December 31, 2026 (then $1.50/$7.50)
  • 1,048,576-token context window, 65,536-token output limit, text/image/video/audio/PDF input
  • Free in Google AI Studio; in the Gemini app, AI Mode and Sheets for AI Pro/Ultra subscribers
  • Scores 54.9% on HLE-Verified and beats 3.7 Flash on Vals Finance Agent V2 and Harvey's Legal Agent benchmark

Gemini 3.8 Flash release date

Google published "Introducing Gemini 3.8 Flash and 3.8 Flash Cyber" on its official blog on September 2, 2026, credited to Tulsee Doshi, Senior Director of Product Management, and Raluca Ada Popa, Gemini Security Lead at Google DeepMind. The post frames the release as building "on the momentum of 3.7 Flash from three weeks ago," making 3.8 Flash the third Flash-branded model Google has shipped in only six weeks, following 3.6 Flash and 3.7 Flash. That cadence is unusually fast even by the standards Google set earlier in 2026, and it means the Flash line has effectively become Google's testbed for shipping smaller, cheaper reasoning gains between the less frequent Pro model updates.

The model went live simultaneously across developer and consumer surfaces rather than rolling out gradually — the Gemini API, Google AI Studio, Vertex AI's Gemini Enterprise Agent Platform, the Gemini app, AI Mode in Search, Gemini in Sheets, Android Studio, Google Antigravity and the UI-generation tool Stitch all listed the model on day one, according to Google's own release notes.

What's new versus Gemini 3.7 Flash and earlier Flash models

Google is explicit that 3.8 Flash is not a from-scratch model so much as a deliberately "harder-working" version of the same Flash-class architecture. The company's blog post describes the core design choice behind the release: on complex tasks, 3.8 Flash "exhibits greater diligence — executing extra reasoning steps, and calling tools iteratively," which can mean it burns more tokens than 3.7 Flash to reach a better answer, particularly at higher effort levels. Developers who need to hold token spend flat can dial the effort level down, or simply keep using 3.7 Flash, which Google says "remains fully supported for efficiency-first workloads."

Google's own comparison data, published on the Gemini model card, shows measurable gains over 3.7 Flash on domain-specific agent benchmarks rather than just generic leaderboard bumps: 3.8 Flash reaches 61.4% on the Vals Finance Agent V2 benchmark versus 59.0% for 3.7 Flash, and 10.0% on Harvey's Legal Agent Benchmark versus 8.8% for 3.7 Flash. Google also cites an early evaluation from workplace-search company Glean, which found 3.8 Flash "completed more than three times as many tasks as Gemini 3.7 Flash" on document-heavy enterprise workflows — a proxy for the kind of long-horizon, multi-step agent work the model is aimed at.

Compared with the Gemini 3.1 generation — which is now mostly represented in Google's current lineup by 3.1 Flash-Lite, 3.1 Flash Image ("Nano Banana 2") and 3.1 Flash Live Preview rather than a standalone "3.1 Flash" text model — 3.8 Flash represents several release cycles of iteration on reasoning depth, tool-calling reliability and prompt-injection robustness, while staying inside the same Flash cost tier. Google says the 3.8 generation also made "a significant leap in prompt injection robustness as measured by Gray Swan," protecting users from malicious prompt-injection attacks, and ships with mitigations against misuse in chemical, biological, radiological and nuclear (CBRN) domains and cyber offense, in line with its Frontier Safety Framework.

On the technical spec sheet, nothing about the context window changed: 3.8 Flash keeps the 1,048,576-token input window and a 65,536-token output ceiling that recent Flash models share, and it accepts text, image, video, audio and PDF input while outputting text only. It supports the same broad tool surface as other Gemini 3 models — caching, code execution, a preview computer-use mode, file search, function calling, Google Maps and Search grounding, structured outputs, and three selectable thinking levels (low, medium, high).

Gemini 3.8 Flash benchmarks Google published

Google did not release a single giant benchmark table for 3.8 Flash the way it sometimes does for Pro-tier launches. Instead, the blog post and DeepMind's model page call out a handful of specific evaluations, several of them comparative against 3.7 Flash or "larger frontier models":

BenchmarkWhat it measuresGemini 3.8 FlashGemini 3.7 Flash
HLE-VerifiedMulti-step reasoning across STEM, humanities and professional fields54.9%Not published by Google for direct comparison
Vals Finance Agent V2Agentic financial analysis and reporting tasks61.4%59.0%
Harvey's Legal Agent BenchmarkLegal-domain agentic reasoning10.0%8.8%
DeepSWE v1.1Long-horizon autonomous software engineeringOutperforms most larger frontier models, Google says, "at a fraction of the cost"Not directly compared by Google
Glean enterprise task completionDocument-heavy agentic workflows (third-party evaluation)"More than three times as many tasks completed"Baseline

Google frames the overall positioning as 3.8 Flash "often approaching the performance of higher-cost frontier models" while keeping Flash-tier speed and pricing — a claim aimed squarely at developers deciding whether a task needs a Pro-tier model or can run on Flash. The company has not published head-to-head numbers against GPT or Claude models for 3.8 Flash specifically, so any such comparisons circulating online are third-party, not Google-sourced.

For teams comparing agentic coding assistants more broadly, it's worth reading how Gemini's agent capabilities stack up against other tool-calling models in our explainer on how AI agents actually work, and how rival labs are positioning their own reasoning models in our Grok 4.7 benchmarks breakdown.

Gemini 3.8 Flash API pricing

Gemini 3.8 Flash launched at the exact same introductory rate as 3.7 Flash and 3.6 Flash: $0.75 per million input tokens and $3.75 per million output tokens (including thinking tokens) on the standard tier, according to Google's Gemini Developer API pricing page. That introductory pricing is explicitly time-limited — it runs through December 31, 2026, after which the standard rate doubles to $1.50 per million input tokens and $7.50 per million output tokens starting January 1, 2027.

Google also publishes Batch, Flex and Priority pricing tiers for 3.8 Flash. Batch and Flex both run at half the standard input/output rate ($0.375/$1.875 through the end of 2026), while Priority inference costs 1.8x standard ($1.35/$6.75 through the end of 2026) in exchange for faster, more consistent response times. Context caching is priced at $0.075 per million tokens (rising to $0.15) plus a $0.50-per-million-tokens-per-hour storage fee (rising to $1.00).

ModelStandard input (per 1M tokens)Standard output (per 1M tokens)Notes
Gemini 3.8 Flash$0.75 → $1.50 (Jan 2027)$3.75 → $7.50 (Jan 2027)Newest Flash model, Sept 2, 2026
Gemini 3.7 Flash$0.75 → $1.50 (Jan 2027)$3.75 → $7.50 (Jan 2027)Same introductory pricing as 3.8
Gemini 3.6 Flash$0.75 → $1.50 (Jan 2027)$3.75 → $7.50 (Jan 2027)Same introductory pricing as 3.8
Gemini 3.5 Flash$1.50 flat$9.00 flatPredates the introductory-pricing scheme
Gemini 3.1 Flash-Lite$0.25 (text/image/video)$1.50Cheaper, lighter-weight tier

Google AI Studio itself remains free to use in all supported regions regardless of which model is selected, with the free API tier offering "limited access to certain models," free input and output tokens, and the caveat that content submitted on the free tier may be used to improve Google's products — a trade-off developers on the paid tier avoid.

Consumer access: Gemini app, AI Mode and Sheets

For everyday users rather than developers, Gemini 3.8 Flash is not a separate app or subscription — it's a model swapped in behind existing products. Google's blog post states plainly that "3.8 Flash is available to Google AI Pro and Ultra subscribers across the Gemini app, AI Mode in Google Search and Gemini in Google Sheets." There is no indication in Google's materials that the free tier of the Gemini app gets 3.8 Flash directly; free users typically get routed to a lighter model or a rate-limited version of the flagship, consistent with how Google has gated its more capable Flash releases in the past.

If you're deciding between Gemini's paid tiers and competing subscriptions from OpenAI or Anthropic, our running comparison of ChatGPT, Claude and Gemini plans and prices breaks down what each subscription tier actually unlocks.

Developer and enterprise availability: AI Studio, Vertex AI, API

On the developer side, Google lists 3.8 Flash under the model ID gemini-3.8-flash in the Gemini API, reachable through Google AI Studio for prototyping and through the Gemini API directly for production traffic. The model also appears in Google Cloud's Gemini Enterprise Agent Platform — the rebranded successor to what was previously called Vertex AI's Model Garden — alongside 3.7 Flash, 3.6 Flash and 3.5 Flash, giving enterprise customers the same model access with Google Cloud's provisioned throughput, dedicated support and volume-based discounts layered on top.

Google is also pushing 3.8 Flash as the default engine behind several of its newer agentic tools: Google Antigravity (an agent-first coding environment), Android Studio's built-in Gemini integration, and Stitch, Google's UI-generation tool. The company's launch post includes several demos built entirely inside Antigravity from single prompts — a 3D castle exploration game using textures generated with Nano Banana, a playable DOS-style version of Google Maps with working directions and Street View, and an interactive hardware teardown visualizer built in Google AI Studio — intended to showcase the model's long-horizon, multi-step coding ability rather than one-shot code generation.

Gemini 3.8 Flash Cyber: the security-focused sibling

Announced in the same post, Gemini 3.8 Flash Cyber is described as Google's "most capable cybersecurity model," built on the same foundational intelligence as 3.8 Flash but tuned specifically for vulnerability detection and automated patching. Unlike the standard model, Cyber is not broadly available — it's distributed only through Google's new Fairwind Program, which Google says gives "trusted government authorities, as well as critical infrastructure operators and software maintainers" prioritized access, with applications handled through Google directly.

Google published several Cyber-specific results: on CyberGym, an industry-standard benchmark for autonomous vulnerability discovery, 3.8 Flash Cyber "surpasses both 3.5 Flash Cyber as well as significantly larger frontier models." On an internal benchmark spanning 20 programming languages, Google says the model reaches "a success rate exceeding 70%." On CWE-Bench, an external patching benchmark run by Collinear, 3.8 Flash Cyber scores a 47.2% pass@1, close to a leading frontier model's 47.8% but, Google says, "at a significantly lower cost." Google also cites internal deployments: its Chrome Security team reportedly found the model produced "2.6 times more correct patches" than larger commercial models, security firm Wiz measured "7.5-9.7% higher recall" at 2.3-5.2x lower cost on its penetration-testing benchmark, and Google's own Cloud Vulnerability Research team says it used the model to find a critical vulnerability in under two hours on a task that "usually takes months." Because Cyber ships with more permissive safety mitigations than the general-purpose 3.8 Flash, Google restricts it to vetted defenders rather than the general public.

How Gemini 3.8 Flash is being positioned against the market

Google's messaging leans hard on cost-efficiency: the pitch throughout the launch post is that 3.8 Flash closes ground on "higher-cost frontier models" without leaving the Flash price tier, rather than claiming outright benchmark supremacy over any specific competing model. That's a deliberate contrast with how Google talks about its Pro-tier releases, and it reflects Flash's role as the workhorse tier developers reach for when a task doesn't need — or can't afford — a full Pro-class model on every call.

Because Google ships Flash updates roughly every three weeks right now (3.6, then 3.7, then 3.8 inside six weeks), the more relevant comparison for most developers is version-to-version rather than against outside models: is the jump from 3.7 to 3.8 worth re-testing prompts and re-benchmarking latency and cost for your specific workload? Given that pricing is identical between the two, the answer mostly depends on whether your use case benefits from the extra reasoning depth 3.8 Flash is willing to spend tokens on. For agent-heavy, document-heavy or multi-step coding workloads, Google's own Vals, Harvey and Glean figures suggest real gains; for simple, latency-sensitive single-turn tasks, 3.7 Flash — or the cheaper 3.1 Flash-Lite — may still be the better fit. Teams weighing coding-assistant options more broadly can compare tool-specific pricing in our look at Cursor, GitHub Copilot, Claude Code and Codex pricing.

What's next

Given Google's current release cadence — three Flash updates in six weeks — a 3.9 Flash or equivalent iteration inside the same broad pricing tier would not be a surprise before the end of 2026, particularly once the January 1, 2027 price increase takes effect and Google may look to justify the higher cost with further reasoning gains. The bigger open question is the Fairwind Program: how many governments and infrastructure operators get access to 3.8 Flash Cyber, and whether Google eventually loosens access as confidence in the model's patching safety record grows. For now, standard Gemini 3.8 Flash is the one most developers and consumers will actually touch, available today through AI Studio, the Gemini API, the Gemini app and Google Cloud's Gemini Enterprise Agent Platform.

Frequently asked questions

When did Google release Gemini 3.8 Flash?

Google published its Gemini 3.8 Flash announcement on September 2, 2026, calling it the third Flash-tier model release in six weeks after Gemini 3.6 Flash and Gemini 3.7 Flash.

How much does Gemini 3.8 Flash cost through the API?

Standard API pricing is $0.75 per million input tokens and $3.75 per million output tokens through December 31, 2026. Google says both rates double to $1.50 and $7.50 per million tokens starting January 1, 2027.

Is Gemini 3.8 Flash free to use?

Google AI Studio access is free in all supported regions with rate limits on the free API tier. For everyday consumers, 3.8 Flash is available in the Gemini app, AI Mode in Search and Gemini in Sheets only to Google AI Pro and Ultra subscribers.

What's new in Gemini 3.8 Flash compared to 3.7 Flash?

Google says 3.8 Flash 'works harder' on complex tasks by running extra reasoning steps and calling tools more iteratively, producing measurable gains on benchmarks like Vals Finance Agent V2 (61.4% vs 59.0%) and Harvey's Legal Agent Benchmark (10.0% vs 8.8%), at the same price as 3.7 Flash.

Where can I access Gemini 3.8 Flash?

It's available through the Gemini API and Google AI Studio, Google Cloud's Gemini Enterprise Agent Platform (formerly Vertex AI), Google Antigravity, Android Studio, Stitch, the Gemini app, AI Mode in Google Search, and Gemini in Google Sheets.

What is Gemini 3.8 Flash Cyber?

It's a cybersecurity-focused sibling model built on the same foundation as 3.8 Flash, tuned for vulnerability discovery and automated patching. It's distributed only to vetted governments, infrastructure operators and software maintainers through Google's Fairwind Program.

Sources

More on Gemini 3.8 Flash →Gemini 3.8 FlashGoogle GeminiGoogle AIGemini APIGoogle DeepMind
Theo Park
Written byTheo Park

Theo Park runs the AI desk at Pandromeda. He follows model launches from the frontier labs and the open-weight community, tracks the assistants and developer tools built on them, and explains what each release changes on pricing, capability and safety. His reporting leans on primary sources: model cards, technical reports, API documentation and the companies' own announcements.

More from AI

See all