Gemini 3.6 Flash

Google DeepMind's cost-efficient Flash model for agentic coding and computer use

Gemini 3.6 Flash: 1M Context & 58.7% SWE-Bench (2026)

Gemini 3.6 Flash is Google DeepMind's July 2026 Flash-tier model with a 1 million token context window, computer use, and lower output pricing than before.

Gemini 3.6 Flash replaces the previous Flash generation as Google's cost-efficient workhorse for agentic coding and desktop automation, finishing tasks with 17% fewer output tokens along the way. Its knowledge cutoff advances a full 14 months to March 2026, making it the default pick for teams that outgrew the old Flash tier but don't need a full Pro-tier model.

Gemini 3.6 Flash is Google DeepMind's multimodal Flash-tier model, released in July 2026, with a 1 million token context window and computer use built into the Gemini API. It out-benchmarks its predecessor across agentic coding, computer use and long-horizon engineering tasks, and costs less per output token as well.

Provider: Google DeepMind · Family: Gemini 3.6

More about Google DeepMind on HokAI

Context window: 1,000,000 tokens · Max output: 64,000

Input modalities: text, image, video, audio, pdf · Output: text

About Gemini 3.6 Flash

Gemini 3.6 Flash is a proprietary multimodal large language model built by Google DeepMind, released July 21, 2026 alongside Gemini 3.5 Flash-Lite and the cybersecurity-focused Gemini 3.5 Flash Cyber. It replaces Gemini 3.5 Flash, which launched May 19, 2026 at Google I/O, as Google's mid-tier workhorse model, sitting below Gemini 3.1 Pro in the lineup but built for the same agentic coding and long-horizon reasoning workloads at a fraction of the cost. Google has not disclosed whether the architecture is dense or mixture-of-experts, or the parameter count; the vendor frames the release around token efficiency rather than raw size. On SWE-Bench Pro, an agentic coding benchmark, Gemini 3.6 Flash scores 58.7%, up from 55.1% for Gemini 3.5 Flash. On DeepSWE v1.1, a long-horizon software engineering test, it scores 49%, versus 37% for its predecessor. OSWorld-Verified, a computer-use benchmark, climbed from 78.4% to 83.0%. MLE-Bench rose from 49.7% to 63.9%, and the GDPval-AA v2 composite moved from 1349 to 1421. Artificial Analysis lists a composite Intelligence Index score of 50 for the model. The model accepts a 1 million token input context window and produces up to 64,000 output tokens. On GDM-MRCR, Google's internal long-context recall test, it holds 91.8% accuracy at 128,000 tokens but drops to 54.0% at the full 1 million token window, a reminder that headline context size and usable recall are not the same measurement. Input modalities are text, image, video, audio and PDF; output is text only, with no native image or audio generation. It supports function calling, search as a tool, code execution, and computer use as a built-in client-side tool through the Gemini API and Gemini Enterprise. Reasoning effort and tool budgets are configurable per request, and the model supports parallel tool calls across multi-step agent workflows. Pricing fell on the output side compared with the previous Flash generation even as every benchmark improved, an unusual combination in a single release. A Batch tier cuts cost for asynchronous, non-real-time jobs, context caching sharply lowers the price of repeat system prompts, and a Priority tier trades a real premium for guaranteed low latency; exact per-token rates are in the pricing table below. Gemini 3.6 Flash is live on the Gemini API through Google AI Studio, Vertex AI, the Gemini Enterprise Agent Platform, Google Antigravity, the consumer Gemini app, and GitHub Copilot as of its launch day. The model ID string is gemini-3.6-flash. Gemini API rate limits run on a rolling 10-minute spend window, capping at $10 for Tier 1 accounts and $200 for Tier 2 and 3, before returning a 429 RESOURCE_EXHAUSTED error. Google reports broadly positive safety movement against its predecessor: improved text-to-text and multilingual refusal accuracy, unchanged image-to-text safety, and a small regression in refusal tone and unjustified refusals. Frontier Safety mitigations were strengthened specifically against CBRN and cyber-offense misuse, and Google says the model is substantially more resistant to jailbreaks than the prior release. The best fit is teams already on Gemini who want agentic coding and computer-use gains without paying full Pro-tier rates, plus high-volume workloads where fewer output tokens per completed task compound into real savings. Teams needing top-tier raw reasoning should look at a full Pro-tier model or Claude Opus 4.7 instead, and teams needing open weights should look at Llama 4 or Qwen3, since Gemini 3.6 Flash ships proprietary and API-only. Knowledge cutoff moved from January 2025 to March 2026, a 14-month jump, the largest single-generation cutoff advance in the Flash line. On the paid Gemini API and Vertex AI, Google does not use prompt or response content to train its models; the free Google AI Studio tier does unless Gemini Apps Activity is turned off. Gemini 3.6 Flash launched the same day as Gemini 3.5 Flash-Lite, built for high-throughput, low-latency work, and Gemini 3.5 Flash Cyber, a cybersecurity-tuned variant. Google has publicly teased Gemini 4 pretraining without giving a release date.

Pricing

Standard usage costs $1.50 per 1M input tokens against $7.50 per 1M output tokens. Batch drops both by half for async workloads, and Priority adds roughly an 80% premium for guaranteed low latency, landing near $2.70 and $13.50.

Key Features

Pros

Cons

Benchmarks

Frequently Asked Questions

How much does Gemini 3.6 Flash cost per 1M tokens?

The standard tier charges $1.50 for every 1M input tokens and $7.50 for every 1M output tokens, cheaper on output than the $9.00 rate on the prior Flash generation. Batch mode halves both rates, and cached input runs $0.15 per 1M, a 90% saving. Priority guarantees low latency for about $2.70 and $13.50.

How does Gemini 3.6 Flash compare on benchmarks vs its predecessor?

Gemini 3.6 Flash beats its predecessor on every published benchmark: SWE-Bench Pro rose from 55.1% to 58.7% and DeepSWE from 37% to 49%. Computer-use and long-horizon engineering scores improved as well, and it reaches those results using fewer output tokens per task.

Is Gemini 3.6 Flash open source or proprietary?

Gemini 3.6 Flash is proprietary and API-only under Google's standard commercial terms; there are no downloadable weights. It is available via the Gemini API, Google AI Studio, Vertex AI, and the Gemini Enterprise Agent Platform.

Does Gemini 3.6 Flash train on user data?

Paid usage through the Gemini API and Vertex AI is not used to train Google's models; only the no-cost AI Studio quota feeds training unless Gemini Apps Activity is switched off. Eligible enterprise customers can get zero-data-retention terms through a Data Processing Addendum amendment.

Who is Gemini 3.6 Flash best for and who should avoid it?

It fits agentic coding, desktop automation and high-volume workloads that want Pro-tier-style gains at Flash pricing. For raw reasoning depth, a full Pro-tier model or GPT-5 is the better call, and anyone relying on full-window retrieval accuracy should chunk documents rather than lean on the entire context window.

More AI Models on HokAI

Visit Gemini 3.6 Flash Official Page