All models

Gemini 3.5 Proreview, pricing and limits

by Google

Google DeepMind's frontier multimodal Pro model targeting a 2M-token context window and Deep Think reasoning, currently in limited Vertex AI enterprise preview.

previewproprietarymultimodalGemini 3.5 family
checked
Context
2.0M tokens
In stacks
0

Gemini 3.5 Pro is for enterprise teams on Vertex AI who need more context or reasoning power than Gemini 3.1 Pro offers today. It is not yet generally available, so teams needing a production model now should use Gemini 3.5 Flash or 3.1 Pro instead. Announced May 19, 2026; GA is targeted for late June.

A 2M-token context window anchors Gemini 3.5 Pro, Google DeepMind's Pro-tier multimodal model, paired with a Deep Think mode built for hard scientific and coding reasoning. It processes text, image, audio, and video in one API call. Access is currently limited to Vertex AI enterprise preview customers.

Provider: Google · Family: Gemini 3.5

More about Google on HokAI

Context window: 2,000,000 tokens

Input modalities: text, image, audio, video, tool-calls · Output: text, tool-calls

About Gemini 3.5 Pro

Gemini 3.5 Pro is Google DeepMind's frontier multimodal model, announced at Google I/O on May 19, 2026. It is the Pro tier of the Gemini 3.5 family, positioned above Gemini 3.5 Flash to absorb the hardest reasoning, deepest multimodal, and longest-context workloads previously routed to Google's Ultra tier. Architecture details have not been disclosed; prior Gemini Pro releases used dense Transformer architectures with Google's native multimodal training pipeline, and 3.5 Pro is built on the same lineage. Access is currently limited to select Vertex AI enterprise customers through a preview program, with general availability targeted for late June 2026. The model is not yet reachable through Google AI Studio, the public Gemini API, AWS Bedrock, Azure, Together, or Fireworks, and authentication runs through Google Cloud IAM on Vertex. Published benchmarks for Gemini 3.5 Pro do not exist yet, since the model has not been released for independent evaluation. Its predecessor Gemini 3.1 Pro, the current GA Pro tier, remains Google's benchmark reference point from its February 2026 launch. Gemini 3.5 Pro accepts text, image, audio, and video in a single native API call, continuing the joint-reasoning multimodal design from Gemini 3.1 Pro. Its Deep Think mode adds extended chain-of-thought for the hardest reasoning tasks. Function calling, structured output, tool use, and Google Search grounding are expected based on Gemini 3.5 Flash's confirmed capabilities, though computer use and code execution remain unconfirmed for the Pro tier. Gemini 3.5 Pro has no published model card yet; the Gemini 3.5 Flash model card is live at deepmind.google/models/model-cards/gemini-3-5-flash/. Prior Gemini Pro safety training filtered for CSAM, violent content, and other harmful outputs under Google's AI Principles, and Deep Think reasoning is expected to inherit the controls applied to Gemini 3.1 Deep Think, which outperformed Gemini 2.5 Pro on safety and tone metrics while keeping refusals low. Gemini 3.5 Pro's training data cutoff has not been published; Gemini 3.1 Pro's cutoff was February 2026 as a reference point. Google's pipeline applies deduplication, quality filtering, and the same CSAM and safety filtering used in prior generations, and API inputs are not used to train production models under Google's standard data policy. GDPR compliance and Cloud data-residency options apply on Vertex; no system card URL is published yet. Gemini 3 Pro was deprecated on March 26, 2026, with Gemini 3.1 Pro Preview becoming the active Pro tier in the interim. Gemini 3.5 Pro is positioned to succeed 3.1 Pro Preview as the primary Pro tier once it reaches GA; Gemini Spark launched the same day as a separate, smaller sibling.

Pricing

Official pricing not yet announced. Analyst estimate based on Flash-to-Pro ratio: approximately $15/$60 per 1M tokens. Gemini 3.5 Flash (the GA sibling) is $1.50/$9 per 1M. Gemini 3.1 Pro (the prior Pro tier) is $2/$12 per 1M. Confirm at official launch.

Key Features

  • 2M-Token Context Window: The largest announced context window of any production frontier model as of June 2026. Handles entire codebases, multi-document research corpora, or book-length texts in one API call.
  • Deep Think Reasoning Mode: Extended chain-of-thought reasoning for hard scientific, mathematical, and coding problems. Prior Deep Think tier scored 84.6% on ARC-AGI-2 and achieved gold-medal performance at international olympiads.
  • Native Multimodal Joint Reasoning: Text, image, audio, and video processed in one API request. No separate model routing for different modalities, consistent with Gemini 3.1 Pro architecture.
  • Vertex AI Enterprise Integration: Deep integration with Google Cloud services including BigQuery, Cloud Storage, and Workspace. Enterprise data residency, IAM-based access control, and provisioned throughput available.
  • Function Calling and Structured Output: Native function calling and JSON structured output, consistent with Gemini 3.5 Flash's confirmed capabilities. Enables agentic workflows without custom orchestration.

Pros

  • The context window is double what Gemini 3.5 Flash and GPT-5.5 currently offer, letting teams skip retrieval chunking for large codebases entirely.
  • Deep Think reasoning carries forward the prior tier's gold-medal olympiad-level performance, a differentiator no competing Pro-tier model has matched yet.
  • One API call handles every modality at once, so teams don't need to route images, audio, or video through separate specialist endpoints.

Cons

  • Not generally available yet: production deployments are blocked until Google opens access beyond the Vertex AI enterprise preview.
  • Official benchmarks and pricing have not been published; any circulating estimate should be treated as unconfirmed until GA.
  • Closed weights, API-only: no self-hosting, fine-tuning of base weights, or air-gapped deployment.

Benchmarks

    Frequently Asked Questions

    How much does Gemini 3.5 Pro cost per 1M tokens?

    Google has not announced official pricing for Gemini 3.5 Pro. Based on historical Pro-to-Flash pricing ratios, analysts estimate roughly $15 input and $60 output per 1M tokens, though this is unconfirmed. For reference, the GA sibling Gemini 3.5 Flash costs $1.50 input / $9 output per 1M, and Gemini 3.1 Pro costs $2 input / $12 output per 1M. Confirmed pricing will appear on the Vertex AI pricing page at GA.

    Is Gemini 3.5 Pro open source or proprietary?

    Gemini 3.5 Pro is proprietary and closed-weight, reachable only through the Vertex AI API. Google's open-weight line is the separate Gemma family, which is a different product from the Gemini Pro series. There is no self-hosting, fine-tuning, or air-gapped deployment path for Gemini 3.5 Pro.

    What are the best alternatives to Gemini 3.5 Pro?

    Gemini 3.1 Pro is the current GA Pro-tier model from the same family and the closest drop-in choice until 3.5 Pro reaches general availability. Gemini 3.5 Flash, announced the same day, is already GA at a fraction of the estimated Pro cost for teams that don't need the full context window. Outside Google, GPT-5.5 and Claude Opus 4.8 are the main frontier competitors targeting similar enterprise reasoning workloads.

    How does Gemini 3.5 Pro compare to GPT-5.5?

    GPT-5.5 has published benchmark scores (88.7% SWE-bench Verified, 92.4% MMLU), while Gemini 3.5 Pro has released none yet since it remains in limited preview. Gemini 3.1 Pro, Google's current GA reference point, scored 94.3% on GPQA Diamond and 80.6% on SWE-bench Verified at its February 2026 launch. A direct comparison between GPT-5.5 and Gemini 3.5 Pro is only possible once the latter reaches general availability.

    How do you get started with Gemini 3.5 Pro?

    Access is currently restricted to select enterprise accounts inside Google's Vertex AI preview program; standard Gemini API keys and Google AI Studio do not grant access to it. Apply for access through the Vertex AI console, or start with Gemini 3.5 Flash, which is already GA, as a drop-in for non-Pro-tier tasks while waiting.

    Top Alternatives

    • Gemini 3.1 Pro: Pick Gemini 3.1 Pro if you need a Pro-tier model with published, independently verified benchmarks available today; pick Gemini 3.5 Pro once GA lands if you need the larger context window and Deep Think reasoning.
    • Gemini 3.5 Flash: Pick Gemini 3.5 Flash if you need a GA model now at confirmed, lower per-token pricing; pick Gemini 3.5 Pro for the larger context window and Deep Think reasoning once it reaches general availability.
    • GPT-5.5: Pick GPT-5.5 if you need a frontier model with published benchmarks available now; pick Gemini 3.5 Pro if its 2M-token context window is the deciding factor once GA lands.
    • Claude Opus 4.8: Pick Claude Opus 4.8 if you need a GA frontier model with a proven track record today; pick Gemini 3.5 Pro if the context window is the deciding factor once it exits preview.

    More AI Models on HokAI

    Visit Gemini 3.5 Pro Official Page