Gemini 3.8 Flash pricing, plans and limits

Sits below Gemini 3.1 Pro in Google's lineup as the fast, cost-efficient workhorse for long-horizon coding agents and enterprise document work.

  • ga
  • proprietary
  • multimodal
  • Gemini 3 Flash family
checked

Gemini 3.8 Flash is Google DeepMind's September 2026 Flash-tier release, further trained from Gemini 3.7 Flash with a 1M-token input window and 64K-token output ceiling. It replaces the prior Flash release for long-horizon coding and finance-document agents, though its slow time to first token rules it out for live chat or voice products.

Gemini 3.8 Flash is Google DeepMind's Gemini 3 Flash-tier model, released September 2, 2026, scoring 90.8% on Terminal-Bench 2.1, ahead of GPT-5.6 Terra's 87.4% and Claude Sonnet 5's 80.4% in the same third-party comparison. It takes text, image, video, audio, and PDF input across a 1M-token context window, built for long-horizon coding agents.

Provider: Google DeepMind · Family: Gemini 3 Flash

More about Google DeepMind on HokAI

Context window: 1,048,576 tokens · Max output: 65,536

Input modalities: text, image, video, audio, pdf, tool-calls · Output: text, tool-calls

About Gemini 3.8 Flash

Gemini 3.8 Flash is Google DeepMind's fourth Flash-tier release in the Gemini 3 family, shipped on September 2, 2026, three weeks after Gemini 3.7 Flash and roughly four months after the line's prior refresh. It is a sparse Mixture-of-Experts transformer that Google's own model card describes as further-trained on top of Gemini 3.7 Flash rather than a new base model; Google defers the Architecture, Training Dataset, and Hardware sections of the 3.8 Flash model card to the 3.7 Flash card for that reason. Exact parameter counts are not disclosed, consistent with every Gemini release to date. Google positions it as its most intelligent workhorse model yet for coding and agents, aimed at long-horizon software engineering, autonomous agent loops, and document-heavy enterprise work, sitting below Gemini 3.1 Pro in the lineup but priced and tuned for high-volume, latency-tolerant workloads. Independent testing by DataCamp shows gains across every disclosed benchmark versus its predecessor: SWE-Bench Pro rose from 60.4% to 61.6%, SWE-Atlas from 48.0% to 51.9%, tau3-bench Banking from 30.9% to 38.1%, and the multimodal CharXiv benchmark from 84.5% to 86.2%; its terminal-agent benchmark improved as well, with the exact figure and rival comparison covered in the benchmarks FAQ. Google's own model page separately reports 54.9% on HLE-Verified, a DeepSWE v1.1 success rate that clears 70% on its long-horizon coding task suite, and a leading 61.4% on the Vals Finance Agent v2 benchmark. Artificial Analysis independently scored the high-reasoning configuration at 59 on its Intelligence Index and measured 326.9 output tokens per second, the fastest in its comparison cohort. The model keeps the 1M-input, 64K-output token class Google introduced with Gemini 3.6 Flash and carried through 3.7 Flash. Google has not published a dedicated long-context recall evaluation specific to 3.8 Flash, so buyers evaluating retrieval accuracy above roughly 100K tokens should run their own needle-in-haystack test rather than assume parity with the flagship Pro model. Input modalities cover text, image, video, audio, and PDF; output is text only, with no native audio or image generation. The model supports function calling, structured outputs, code execution, search grounding, URL-context grounding, Google Maps grounding, and a preview computer-use control loop for driving a browser or desktop. Three thinking levels (low, medium, high) let a caller trade latency and token spend for deeper reasoning; a minimal setting, offered on some sibling models, is not supported here. Google priced the launch at parity with its immediate predecessor, with a temporary introductory rate on the paid tier that expires at the end of 2026 and roughly doubles at the start of 2027; see the pricing FAQ for the exact per-token figures. Google also offers a no-cost tier for experimentation, and context caching, batch processing, and a Flex mode each cut costs further for repeat-context or asynchronous workloads. On a representative workload, summarizing a long research document as a single call costs a small fraction of a dollar, and a full day of a coding agent moving roughly a million tokens through the model lands around a dollar and a half. Gemini 3.8 Flash is available through the Gemini API and Google AI Studio, Google Vertex AI as part of the Gemini Enterprise Agent Platform, the consumer Gemini app for Pro and Ultra subscribers, Gemini in Google Sheets, AI Mode in Search, and Google's own Antigravity IDE, which the company is positioning as a direct answer to Cursor and GitHub Copilot. Third-party access includes Cursor, which added the model shortly after launch. Google documents official SDKs for Python, JavaScript/TypeScript, Go, and Java alongside the REST API; per-region availability and rate limits are set inside each account's AI Studio or Vertex console rather than published as one public table. Google's safety write-up for 3.8 Flash largely defers to the Gemini 3.7 Flash model card too: because the company's Frontier Safety Framework evaluation found no meaningful new capabilities and no Tracked or Critical Capability Levels reached, it carried the 3.7 Flash assessment forward rather than re-running the full suite. Manual red teaming was still performed by a specialist team outside the model's development group, though Google does not name external red-team partners in the public card. The model card reports a small regression in multilingual safety and a small rise in unjustified refusals compared with the prior release, alongside otherwise stable automated safety metrics. It suits teams building autonomous coding agents, long-running document or finance workflows, and terminal-based agent tasks where total throughput matters more than instant responsiveness; the agentic and coding gains over its predecessor are real and independently verified. It is a poor fit for latency-sensitive, user-facing chat or voice products, since it visibly reasons before it starts responding, and for teams on a fixed token budget who have not yet tested its heavier-than-average output verbosity against their own prompts. Teams that only need the prior Flash release's capability level and want to avoid any verbosity or safety-regression risk can stay on that version at the same price. Training data specifics are not published separately for 3.8 Flash; the model card points to the 3.7 Flash card, which describes a mix of public web documents, code, and licensed and synthetically generated data. The stated knowledge cutoff is March 2026, though Google notes some domains reflect information only as current as January 2025. Data handling splits by tier: paid usage is excluded from training under Google's terms, while free-tier submissions can feed back into product development; the data-handling FAQ below has the full policy. This is Google's third Flash release in six weeks and the fourth in under four months, following Gemini 3.6 Flash and Gemini 3.7 Flash on August 13, 2026. Alongside the general-purpose release, Google shipped a restricted cybersecurity sibling, Gemini 3.8 Flash Cyber, built on the same underlying model but gated behind a case-by-case Fairwind Program for vetted defenders rather than opened through the standard API.

Pricing

Paid-tier pricing is $0.75 per 1M input tokens and $3.75 per 1M output tokens (including thinking tokens) through December 31, 2026, rising to $1.50 and $7.50 on January 1, 2027. Cached input tokens are $0.075 per 1M through 2026 (then $0.15), plus $0.50 per 1M tokens per hour for cache storage (then $1.00). Batch and Flex modes cut standard input/output rates by half. A free tier is available through Google AI Studio and unpaid API quota.

Key Features

  • 1M-Token Multimodal Context: Reads text alongside images, video, audio, and PDFs in a single request, with room for large codebases or long filings in one call.
  • Configurable Thinking Levels: Low, medium, and high reasoning-effort settings trade latency and token spend for deeper multi-step reasoning on a per-request basis.
  • Long-Horizon Agentic Coding: Scores 61.6% on SWE-Bench Pro and 51.9% on SWE-Atlas, both gains over its predecessor, with iterative tool calling and roadblock navigation built in.
  • Context Caching: Repeated system prompts or long reference documents can be cached at a steep discount off the standard input rate, cutting cost on repeat-context workloads.
  • Computer Use and Grounding (Preview): Ships a preview computer-use control loop plus built-in search, URL-context, and Google Maps grounding alongside standard function calling and structured outputs.

Pros

  • Beats its own predecessor and two named frontier competitors on the same agentic terminal-task benchmark, per independent third-party disclosures.
  • 326.9 tokens/second output throughput ranks first in its Artificial Analysis cohort, so long autonomous runs finish fast once generation starts.
  • Context caching cuts input costs by roughly 90% on repeated system prompts or reference documents, per Google's published pricing page.
  • A free tier through Google AI Studio and unpaid API quota needs no credit card to start experimenting.

Cons

  • Time-to-first-token averages close to 13 seconds in independent testing (Artificial Analysis), a poor fit for live chat or voice products.
  • Runs meaningfully more verbose than the median comparable model on identical benchmark tasks, which can offset its per-token price advantage on some workloads.
  • No native audio or image output; every response comes back as text regardless of what modalities went in.

Benchmarks

  • charxiv: 86.2
  • swe atlas: 51.9
  • hle verified: 54.9
  • swe bench pro: 61.6
  • tau3 bench banking: 38.1
  • terminal bench 2 1: 90.8
  • humanitys last exam: 45.4
  • vals finance agent v2: 61.4
  • deepswe v1 1 min success rate: 70
  • artificial analysis intelligence index: 59
  • artificial analysis speed tokens per sec: 326.9

Frequently Asked Questions

What does Gemini 3.8 Flash cost to run?

On the paid tier, input runs $0.75 per 1M tokens and output $3.75 per 1M tokens through the end of 2026 (thinking tokens bill as output); the standard rate doubles to $1.50 and $7.50 once the introductory window closes. Cached input drops to $0.075 per 1M tokens over the same period, and both the Batch and Flex APIs cut standard rates by half. Google AI Studio and unpaid API quota stay free, with no charge for input, output, or context caching.

How does Gemini 3.8 Flash perform against GPT-5.6 and Claude Sonnet 5?

Gemini 3.8 Flash leads Terminal-Bench 2.1 at 90.8%, compared with 87.4% for GPT-5.6 Terra and 80.4% for Claude Sonnet 5 in the same third-party comparison. Google's own testing separately puts it at 54.9% on HLE-Verified and above 70% on DeepSWE v1.1 for long-horizon software engineering, figures it does not break out for GPT or Claude directly. It generally wins on agentic terminal and coding tasks; broader general-reasoning leaderboards were not part of the disclosed comparison.

Is Gemini 3.8 Flash open source?

No. Gemini 3.8 Flash is proprietary and API-only, licensed under Google's Gemini API Additional Terms of Service, with no downloadable weights. It is available through the Gemini API, Google AI Studio, and Google Vertex AI, alongside Google's own Antigravity IDE and the consumer Gemini app.

Does Gemini 3.8 Flash train on the data you send it?

On the paid Gemini API and Vertex AI, Google states it does not use prompts or responses to train its models, retaining them only briefly for abuse detection, safety, and legal compliance. On the free Google AI Studio tier and unpaid API quota, submitted content may be used to improve Google's products and may be reviewed by people. Users in the EEA, Switzerland, and the UK get the paid-tier data terms applied even on free usage.

Who should use Gemini 3.8 Flash, and who should avoid it?

It fits teams running autonomous coding agents, terminal-based agent tasks, and long, document-heavy finance or enterprise workflows, where its coding and agentic benchmark gains over its predecessor are real. It is a weak choice for latency-sensitive chat or voice products, given its multi-second time to first token, and for cost-capped, high-volume simple tasks, since it tends to use more output tokens than comparable models on the same task. Teams already satisfied with the prior Flash release's capability level have no cost reason to switch immediately.

More AI Models on HokAI

Visit Gemini 3.8 Flash Official Page