Gemini 3.5 Flash-Lite: 1M Context & 86.9% GPQA (2026)
Released July 2026, Gemini 3.5 Flash-Lite scores 36 on Artificial Analysis's Intelligence Index, runs 350 tokens/sec, built for high-volume AI subagent tasks.
Gemini 3.5 Flash-Lite carries a 1,048,576-token context window and replaces Gemini 2.5 Flash and simpler Gemini 3 Flash workloads for teams running high-volume subagents, document parsing, and agentic search. It fits best where thousands of daily calls make small per-token savings add up fast, more than where peak reasoning quality matters.
Gemini 3.5 Flash-Lite is Google DeepMind's lightweight multimodal model, released July 21, 2026, scoring 86.9% on GPQA Diamond while running as the fastest model in the current Gemini 3.5 lineup. It reads text, image, video, audio, and PDF input and replies with text output only, aimed at high-throughput subagent work rather than deep reasoning.
Provider: Google DeepMind · Family: Gemini 3.5
More about Google DeepMind on HokAI
Context window: 1,048,576 tokens · Max output: 65,536
Input modalities: text, image, video, audio, pdf · Output: text
About Gemini 3.5 Flash-Lite
Gemini 3.5 Flash-Lite is a proprietary multimodal model built by Google DeepMind, released July 21, 2026 alongside Gemini 3.6 Flash and Gemini 3.5 Flash Cyber. It replaces Gemini 3.1 Flash-Lite, Google's March 2026 budget tier, as the cheapest and fastest model in the current Gemini lineup, built for high-volume subagent calls, document parsing, and agentic search rather than frontier reasoning depth. Google has not disclosed parameter count or whether the architecture is dense or mixture-of-experts, consistent with its practice across every Gemini generation to date. On SWE-bench Pro, Flash-Lite scores 54.2%, ahead of Gemini 3 Flash's 49.6% on the same benchmark. GPQA Diamond sits at 86.9%. Terminal-Bench 2.1 improves to 54% from 31% on the prior generation, OSWorld-Verified reaches 74.0% versus 65.1%, and GDM-MRCR v2 long-context recall hits 72.2% against 60.1%. On Artificial Analysis's composite Intelligence Index it scores 36, up 11 points from Gemini 3.1 Flash-Lite's 25, still well behind Gemini 3.5 Flash and Gemini 3.5 Pro but ahead of similarly priced competitors from other vendors. The context window is 1,048,576 tokens, matching every current Gemini 3.x model, with a maximum output of 65,536 tokens. Google has not published a needle-in-haystack score specific to the Flash-Lite tier; the GDM-MRCR v2 figure above is the closest public proxy for how well it holds onto information deep in a long prompt. Flash-Lite reads text, image, video, audio, and PDF as input and writes text back out. It supports function calling, structured output, and remote MCP tool calls. The API defaults to a minimal thinking level tuned for classification, routing, and JSON extraction; Google's own developer docs recommend raising that level to medium or high for subagents that write code, run terminal commands, or call external APIs, since the low default trades accuracy for speed. Flash-Lite undercuts every other model in the 3.5 family on both input and output token rates, positioned as the fixed price floor of Google's Gemini lineup regardless of prompt length. Exact per-token rates and worked cost examples are covered in the pricing FAQ on this page rather than repeated here. It is available directly through Google AI Studio and the Gemini API, and through Vertex AI on the Gemini Enterprise Agent Platform. Rate limiting is spend-based on a rolling 10-minute window rather than fixed requests-per-minute: Tier 1 accounts cap at $10 of spend in that window before requests start returning a 429 RESOURCE_EXHAUSTED error. Third-party inference platforms have not confirmed hosting the model yet; Google has kept it exclusive to its own AI Studio and Vertex surfaces at launch. Knowledge cutoff is March 2026. Most training-data and safety-evaluation detail for this tier lives in the older Gemini 3.1 Flash-Lite documentation rather than a standalone Flash-Lite system card, since the newer version builds on that same base. Two API behaviors stand out from OpenAI- and Anthropic-style APIs: the model silently drops custom temperature, top-K, and top-P values, and it throws an error if you set a custom frequency or presence penalty instead of just ignoring it. Flash-Lite fits organizations firing thousands of subagent calls a day where raw throughput and floor-level pricing matter more than topping a reasoning leaderboard: agentic search, bulk document parsing, and classification or routing subagents inside larger multi-agent systems. Teams that need frontier reasoning depth should look at Gemini 3.5 Pro, GPT-5.5, or Claude Opus 4.7 instead; teams that need faster agentic coding without dropping to the absolute cost floor should compare against Gemini 3.5 Flash. The model is proprietary and closed-weights, available only under Google's standard Gemini API terms. There is no open-weights release or self-hosting option for this tier, and none is expected given how Google has handled every prior Flash-Lite generation. Flash-Lite launched the same day as the larger Gemini 3.6 Flash and the coding-focused Gemini 3.5 Flash Cyber, positioning it as the budget floor of a three-model refresh rather than a standalone release. Its direct predecessor, Gemini 3.1 Flash-Lite, has a confirmed API shutdown date of May 7, 2027, giving developers roughly ten months from launch to migrate before the older model stops serving requests.
Pricing
Standard pricing runs $0.30 per 1M input tokens and $2.50 per 1M output tokens, a flat rate that does not change with prompt length. Google has not disclosed a batch-API discount or cached-input rate for this tier.
Key Features
- 1M-Token Context Window: 1M-token input window with a 65,536-token maximum output, shared across the current Gemini 3.x family.
- Fastest in the 3.5 Series: The quickest model Google has shipped this generation, per Artificial Analysis speed benchmarking.
- Multimodal Input: Accepts image, video, audio, and PDF alongside text in a single request, natively, no separate transcription step.
- Configurable Thinking Level: Defaults to minimal thinking for routing and extraction; can be raised to medium or high for agentic coding subagents.
- Remote MCP Tool Calls: Supports function calling, structured output, and calling external services directly via Model Context Protocol.
Pros
- Beats Gemini 3 Flash on agentic coding: 54.2% on SWE-bench Pro versus 49.6%.
- Fastest model Google has shipped in the Gemini 3.5 series, per Artificial Analysis benchmarking.
- 1M-token context window at the lowest price point in the current Gemini lineup.
- Skips a separate transcription step for images, video, or audio, unlike text-only rivals in its price tier.
Cons
- Defaults to minimal thinking level, trading accuracy for speed unless manually raised for coding or tool-calling subagents.
- Ignores custom temperature, top-K, and top-P settings, and rejects custom frequency or presence penalty values with an error.
- No confirmed availability on third-party inference platforms outside Google's own AI Studio and Vertex AI at launch.
Benchmarks
- mmlu: 89.2
- mmlu pro: 83
- gdm mrcr v2: 72.2
- gdpval aa v2: 1140
- gpqa diamond: 86.9
- swe bench pro: 54.2
- osworld verified: 74
- terminal bench 2 1: 54
- artificial analysis intelligence index: 36
- artificial analysis speed tokens per sec: 350
Frequently Asked Questions
How much does Gemini 3.5 Flash-Lite cost per 1M tokens?
At launch, Gemini 3.5 Flash-Lite charges $0.30 for every 1M input tokens and $2.50 for every 1M output tokens generated, a flat rate that undercuts every other model in the Gemini 3.5 family. No batch or cached-input discount has been published for this tier so far.
How does Gemini 3.5 Flash-Lite compare to Gemini 3 Flash on benchmarks?
It pulls ahead on agentic tasks, reaching 74.0% on OSWorld-Verified against Gemini 3 Flash's 65.1%, and 54% on Terminal-Bench 2.1 versus 31% for the prior generation. It still trails the larger Gemini 3.5 Flash and Gemini 3.5 Pro on Artificial Analysis's composite Intelligence Index.
Is Gemini 3.5 Flash-Lite open source or proprietary?
Gemini 3.5 Flash-Lite is proprietary and closed-weights. It is available only through Google's own Gemini API, Google AI Studio, and Vertex AI on the Gemini Enterprise Agent Platform, with no self-hosting or open-weights option.
Does Gemini 3.5 Flash-Lite train on user data?
Google has not published a standalone system card for this tier; training-data and retention detail is documented under the older Gemini 3.1 Flash-Lite release instead, since Flash-Lite builds on that base. Enterprise use through Vertex AI follows Google Cloud's standard data-processing terms rather than the consumer Gemini app's defaults.
Who should use Gemini 3.5 Flash-Lite and who should avoid it?
It suits high-volume workloads like agentic search, document parsing, and routing subagents inside larger multi-agent systems, where speed and per-call cost outweigh reasoning depth. Teams chasing frontier reasoning should pick Claude Opus 4.7, GPT-5.5, or Gemini 3.5 Pro instead, and teams needing more coding headroom should compare against Gemini 3.5 Flash.
Top Alternatives
- Gemini 3.1 Flash-Lite: Pick Gemini 3.5 Flash-Lite for higher agentic benchmark scores at the same price; Gemini 3.1 Flash-Lite is being retired in May 2027 anyway.
- Gemini 3.5 Flash: Pick Gemini 3.5 Flash if your subagents need more reasoning headroom than Flash-Lite's minimal default thinking level provides.
- GPT-4o mini: Pick Gemini 3.5 Flash-Lite for a far larger context window; GPT-4o mini's ceiling is a fraction of the size.
- DeepSeek-V4 Flash: Pick Gemini 3.5 Flash-Lite for native multimodal input; pick DeepSeek-V4 Flash if open-weights access matters more than raw speed.