Gemini 3.6 Flash: 1M Context & 58.7% SWE-Bench (2026)
Gemini 3.6 Flash is Google DeepMind's July 2026 Flash-tier model with a 1 million token context window, computer use, and lower output pricing than before.
Gemini 3.6 Flash replaces the previous Flash generation as Google's cost-efficient workhorse for agentic coding and desktop automation, finishing tasks with 17% fewer output tokens along the way. Its knowledge cutoff advances a full 14 months to March 2026, making it the default pick for teams that outgrew the old Flash tier but don't need a full Pro-tier model.
Gemini 3.6 Flash is Google DeepMind's multimodal Flash-tier model, released in July 2026, with a 1 million token context window and computer use built into the Gemini API. It out-benchmarks its predecessor across agentic coding, computer use and long-horizon engineering tasks, and costs less per output token as well.
Provider: Google DeepMind · Family: Gemini 3.6
More about Google DeepMind on HokAI
Context window: 1,000,000 tokens · Max output: 64,000
Input modalities: text, image, video, audio, pdf · Output: text
About Gemini 3.6 Flash
Gemini 3.6 Flash is a proprietary multimodal large language model built by Google DeepMind, released July 21, 2026 alongside Gemini 3.5 Flash-Lite and the cybersecurity-focused Gemini 3.5 Flash Cyber. It replaces Gemini 3.5 Flash, which launched May 19, 2026 at Google I/O, as Google's mid-tier workhorse model, sitting below Gemini 3.1 Pro in the lineup but built for the same agentic coding and long-horizon reasoning workloads at a fraction of the cost. Google has not disclosed whether the architecture is dense or mixture-of-experts, or the parameter count; the vendor frames the release around token efficiency rather than raw size. On SWE-Bench Pro, an agentic coding benchmark, Gemini 3.6 Flash scores 58.7%, up from 55.1% for Gemini 3.5 Flash. On DeepSWE v1.1, a long-horizon software engineering test, it scores 49%, versus 37% for its predecessor. OSWorld-Verified, a computer-use benchmark, climbed from 78.4% to 83.0%. MLE-Bench rose from 49.7% to 63.9%, and the GDPval-AA v2 composite moved from 1349 to 1421. Artificial Analysis lists a composite Intelligence Index score of 50 for the model. The model accepts a 1 million token input context window and produces up to 64,000 output tokens. On GDM-MRCR, Google's internal long-context recall test, it holds 91.8% accuracy at 128,000 tokens but drops to 54.0% at the full 1 million token window, a reminder that headline context size and usable recall are not the same measurement. Input modalities are text, image, video, audio and PDF; output is text only, with no native image or audio generation. It supports function calling, search as a tool, code execution, and computer use as a built-in client-side tool through the Gemini API and Gemini Enterprise. Reasoning effort and tool budgets are configurable per request, and the model supports parallel tool calls across multi-step agent workflows. Pricing fell on the output side compared with the previous Flash generation even as every benchmark improved, an unusual combination in a single release. A Batch tier cuts cost for asynchronous, non-real-time jobs, context caching sharply lowers the price of repeat system prompts, and a Priority tier trades a real premium for guaranteed low latency; exact per-token rates are in the pricing table below. Gemini 3.6 Flash is live on the Gemini API through Google AI Studio, Vertex AI, the Gemini Enterprise Agent Platform, Google Antigravity, the consumer Gemini app, and GitHub Copilot as of its launch day. The model ID string is gemini-3.6-flash. Gemini API rate limits run on a rolling 10-minute spend window, capping at $10 for Tier 1 accounts and $200 for Tier 2 and 3, before returning a 429 RESOURCE_EXHAUSTED error. Google reports broadly positive safety movement against its predecessor: improved text-to-text and multilingual refusal accuracy, unchanged image-to-text safety, and a small regression in refusal tone and unjustified refusals. Frontier Safety mitigations were strengthened specifically against CBRN and cyber-offense misuse, and Google says the model is substantially more resistant to jailbreaks than the prior release. The best fit is teams already on Gemini who want agentic coding and computer-use gains without paying full Pro-tier rates, plus high-volume workloads where fewer output tokens per completed task compound into real savings. Teams needing top-tier raw reasoning should look at a full Pro-tier model or Claude Opus 4.7 instead, and teams needing open weights should look at Llama 4 or Qwen3, since Gemini 3.6 Flash ships proprietary and API-only. Knowledge cutoff moved from January 2025 to March 2026, a 14-month jump, the largest single-generation cutoff advance in the Flash line. On the paid Gemini API and Vertex AI, Google does not use prompt or response content to train its models; the free Google AI Studio tier does unless Gemini Apps Activity is turned off. Gemini 3.6 Flash launched the same day as Gemini 3.5 Flash-Lite, built for high-throughput, low-latency work, and Gemini 3.5 Flash Cyber, a cybersecurity-tuned variant. Google has publicly teased Gemini 4 pretraining without giving a release date.
Pricing
Standard usage costs $1.50 per 1M input tokens against $7.50 per 1M output tokens. Batch drops both by half for async workloads, and Priority adds roughly an 80% premium for guaranteed low latency, landing near $2.70 and $13.50.
Key Features
- 1M-Token Context Window: Accepts up to 1,000,000 input tokens with a 64,000-token max output, enough for full codebases or hour-long transcripts in a single call.
- Built-in Computer Use: Ships computer use as a client-side tool via the Gemini API and Gemini Enterprise, scoring 83.0% on OSWorld-Verified, up from 78.4% on the prior Flash release.
- Fewer Reasoning Steps Per Task: Reaches the same task outcomes using measurably fewer reasoning steps and tool calls than the previous Flash generation, per Artificial Analysis.
- March 2026 Knowledge Cutoff: Training data extends to March 2026, well past the cutoff on the previous Flash generation.
- Configurable Reasoning Effort: Lets you set a tool-call and reasoning budget per request instead of a fixed depth, useful for tuning cost and latency independently.
Pros
- MLE-Bench score climbed from 49.7% to 63.9% versus the previous Flash generation, the largest single engineering-benchmark jump in this release.
- The GDPval-AA v2 composite, Google's broadest task-quality benchmark, rose from 1349 to 1421 in the same release.
- Output pricing fell even as every benchmark improved versus the prior generation, an unusual price-and-quality combination.
Cons
- Long-context recall thins out near the top of the window: GDM-MRCR shows only 54.0% accuracy at the full context size versus 91.8% at a quarter of that depth.
- No native image or audio output; multimodal support only goes one direction.
- Google's own model card flags a slight regression in refusal tone and unjustified refusals versus the prior release.
Benchmarks
- mle bench: 63.9
- gdm mrcr 1m: 54
- deepswe v1 1: 49
- gdpval aa v2: 1421
- gdm mrcr 128k: 91.8
- swe bench pro: 58.7
- charxiv no tools: 85.2
- osworld verified: 83
- charxiv with tools: 89.4
- artificial analysis intelligence index: 50
- artificial analysis price blended per m: 1.16
- artificial analysis speed tokens per sec: 271
Frequently Asked Questions
How much does Gemini 3.6 Flash cost per 1M tokens?
The standard tier charges $1.50 for every 1M input tokens and $7.50 for every 1M output tokens, cheaper on output than the $9.00 rate on the prior Flash generation. Batch mode halves both rates, and cached input runs $0.15 per 1M, a 90% saving. Priority guarantees low latency for about $2.70 and $13.50.
How does Gemini 3.6 Flash compare on benchmarks vs its predecessor?
Gemini 3.6 Flash beats its predecessor on every published benchmark: SWE-Bench Pro rose from 55.1% to 58.7% and DeepSWE from 37% to 49%. Computer-use and long-horizon engineering scores improved as well, and it reaches those results using fewer output tokens per task.
Is Gemini 3.6 Flash open source or proprietary?
Gemini 3.6 Flash is proprietary and API-only under Google's standard commercial terms; there are no downloadable weights. It is available via the Gemini API, Google AI Studio, Vertex AI, and the Gemini Enterprise Agent Platform.
Does Gemini 3.6 Flash train on user data?
Paid usage through the Gemini API and Vertex AI is not used to train Google's models; only the no-cost AI Studio quota feeds training unless Gemini Apps Activity is switched off. Eligible enterprise customers can get zero-data-retention terms through a Data Processing Addendum amendment.
Who is Gemini 3.6 Flash best for and who should avoid it?
It fits agentic coding, desktop automation and high-volume workloads that want Pro-tier-style gains at Flash pricing. For raw reasoning depth, a full Pro-tier model or GPT-5 is the better call, and anyone relying on full-window retrieval accuracy should chunk documents rather than lean on the entire context window.