Gemini 3.8 Live review, pricing and limits

Google's cost-efficient default for real-time voice agents, one step below Gemini 3.8 Live Extended Thinking on background reasoning depth.

  • ga
  • proprietary
  • multimodal
  • Gemini 3.8 family
checked

Gemini 3.8 Live is built for teams that need a voice agent talking and acting at once, not a text chatbot with a voice bolted on. It replaces Google's older Live defaults with a session that holds up to 65,536 output tokens, auto-switches across 97 languages, and stays on the line for 15 minutes before an audio-only call needs to reconnect.

Google DeepMind shipped Gemini 3.8 Live in September 2026 as the default native speech-to-speech model for the Gemini Live API. It accepts text, image, video and audio input and answers in synthesized audio, grounding replies in live visual context and running tool calls in the background without interrupting the conversation.

Where it sits

  • $1.69/M$ per 1M tokensBlended price (3:1)Lower is better#30 / 61peer median $1.71/Mvendor price, checked by HokAI
  • --tokens/sOutput speedHigher is better-- / 36peer median 90 tok/scited: Artificial Analysis
  • --% solvedSWE-bench VerifiedHigher is better-- / 26peer median 78.3%per source, see benchmark scores
  • --% correctGPQA DiamondHigher is better-- / 43peer median 87.8%per source, see benchmark scores

Priced around the middle of the 61 GA models with a published price (rank 30), and one of 21 that document a zero-data-retention option.

Ranks are against GA models on HokAI that publish the same figure; ties share a rank.

Provider: Google DeepMind · Family: Gemini 3.8

More about Google DeepMind on HokAI

Context window: 131,072 tokens · Max output: 65,536

Input modalities: text, image, video, audio · Output: text, audio

About Gemini 3.8 Live

Gemini 3.8 Live is Google DeepMind's default native speech-to-speech model, released September 15, 2026 alongside a higher-reasoning sibling, Gemini 3.8 Live Extended Thinking. Both replace the prior Gemini 2.5 Flash and Gemini 3.1 Flash Live Preview defaults inside the Gemini Live API. Where most of Google DeepMind's model lineup, including Gemini 3.8 Flash, is built for text and code, Gemini 3.8 Live is audio-to-audio: it takes microphone audio, images and video as input and streams synthesized speech back over a persistent WebSocket connection. Google positions it as the workhorse tier "built for scale and cost efficiency," reserving Extended Thinking for calls that need visible multi-step reasoning mid-conversation. It lands in a rapid Gemini 3.8 release cycle that also includes Gemini 3.8 Flash Cyber and predecessors Gemini 3.7 Flash and Gemini 3.6 Flash, shipped within weeks of one another.

On Artificial Analysis' Speech to Speech Quality Index, the independent scoreboard for conversational voice models, Gemini 3.8 Live scores 76.0, built from a 92% Speech Reasoning sub-score, a 96.1% Conversational Dynamics sub-score and a 30.1% Agentic Performance sub-score. Google has not published SWE-bench, GPQA or MMLU-style text-reasoning scores for either Live model, since neither targets text or coding; the Speech to Speech Index is the only third-party benchmark available at launch.

A session holds 131,072 input tokens and can generate up to 65,536 output tokens, roughly double the 32,000-token ceiling Google documents for its older, non-native-audio Live models. Session length is capped separately: an audio-only call runs 15 minutes before needing a reconnect, and a call that also streams video drops to 2 minutes, though Google says both can be extended with session-management techniques. That token ceiling places it in HokAI's 128K-400K context tier for models.

Input covers text, images, video (up to 1 frame per second) and 16-bit PCM audio; output is limited to audio and text, with no image generation. The model detects and switches between 97 supported languages mid-conversation, and Google markets it for reading confirmation codes and claim numbers aloud accurately, a known weak spot for earlier voice models. Function calling runs asynchronously by default: the model keeps talking while a background tool call completes. It supports search grounding, but Google marks context caching, code execution, structured outputs and the Batch API unsupported, unlike text models such as Gemini Omni 1.1 Flash. HokAI's audio-capable model slice shows how the input mix compares with other speech-capable models shipped this cycle.

Gemini 3.8 Live bills audio input and output as two separate per-minute meters rather than one blended per-token rate, on top of standard per-token text pricing; exact figures are in the pricing table below. A no-cost tier runs through Google AI Studio and unpaid API quota for testing. Google has not published a separate price for Extended Thinking, so choosing between the two models is a capability decision, not a cost one.

The model is reachable through the Gemini API and Google AI Studio at launch, with enterprise access via Gemini Enterprise and listing on Vertex AI. It is hosted and closed-weight only, so there is no self-hosting path, unlike on-device Gemini Nano. Google lists first-party integrations for Pipecat, LiveKit, LangChain, Agora, Fishjam, Vercel and Vision Agents. It ships at general availability rather than preview, matching the proprietary, closed-weight status of the rest of Google's hosted lineup, alongside still-active predecessor Gemini 3.5 Flash.

Every audio clip is watermarked with SynthID, Google's imperceptible audio watermark, so downstream systems can flag AI-generated speech. Google had not published a dedicated model card or red-teaming partner list for Gemini 3.8 Live specifically as of launch; the concurrent Gemini 3.8 Flash card describes only routine safety changes for the text side of the same generation, and that assessment does not cover the Live variant.

Gemini 3.8 Live fits customer-support and field-service voice agents that need to talk, listen and trigger a backend action in the same breath: insurance claims intake, banking verification reads and multilingual support lines are the use cases Google leads with. Teams building a text-only chatbot get nothing from the Live variant over Gemini 3.8 Flash or a cheaper tier like Gemini 3.5 Flash-Lite; teams needing narrated step-by-step reasoning should reach for the Extended Thinking sibling, and anyone needing offline voice inference should look at Gemini Nano. Anyone comparing video-grounded models more broadly can browse HokAI's video-input model slice.

Google has not published a training data cutoff for Gemini 3.8 Live specifically. Under Google's standard Gemini API terms, paid-tier prompts and responses are not used to train models, with logs kept up to 30 days solely to detect abuse; free-tier usage may be used to improve Google's products, the same policy that applies across the rest of the Gemini API.

Pricing

On the Gemini API, Gemini 3.8 Live prices text at $0.75 per 1M input tokens and $4.50 per 1M output tokens, audio input at $3.00 per 1M tokens ($0.005 per minute) and audio output at $12.00 per 1M tokens ($0.018 per minute), and image or video input at $1.00 per 1M tokens ($0.002 per minute); there is no image or video output to price. Context caching is not supported on this model, so no cached-input discount applies. Development and testing are free of charge through Google AI Studio's unpaid quota.

What a real job costs

JobInputOutputTotal
Summarise a 20-page PDF$0.022$0.0045$0.027
Support reply$0.0015$0.0013$0.0029
One coding agent run$0.150$0.090$0.240

Budgets: 20-page PDF = 30k in / 1k out · Support reply = 2k in / 300 out · Coding agent run = 200k in / 20k out. Computed from the vendor's per-token prices at render time; cached-input discounts are not applied.

Key Features

  • Asynchronous Function Calling: Executes API and tool calls in the background while streaming audio, so a lookup or booking action does not force the model to go silent mid-reply.
  • Live Visual Grounding: Accepts image and video frames alongside audio input, so replies can reference what the model currently sees, not only what it hears.
  • Mid-Conversation Language Switching: Detects and switches languages automatically mid-call, without the caller re-selecting one or being transferred to a different line.
  • Search Grounding: Can ground spoken answers in live search results during a call, rather than relying only on parametric knowledge from training.
  • SynthID Audio Watermarking: Every generated audio clip carries Google's imperceptible SynthID watermark, so downstream systems can flag it as AI-generated.

Pros

  • Bills voice traffic by direction, with independent per-minute rates for what callers say and what the model speaks back, and a no-cost tier for pre-production testing.
  • Its Speech Reasoning component measured 92% in the same third-party voice-quality breakdown, the strongest of the model's three sub-scores.
  • Executes tool and API calls asynchronously in the background so the conversation keeps flowing instead of going silent while a lookup completes.
  • Ships first-party support for LiveKit, Pipecat, Agora, LangChain, Vercel, Fishjam and Vision Agents, covering most real-time voice-agent frameworks already in use.

Cons

  • A session that also streams video is capped at 2 minutes before it needs a reconnect, far shorter than an audio-only session.
  • No context caching, code execution, structured outputs, or Batch API support, unlike Google's text-first Gemini models.
  • Google has not published a dedicated model card, training data cutoff, or red-teaming disclosure for the Live variant at launch.

Frequently Asked Questions

What does Gemini 3.8 Live actually cost?

Google meters Gemini 3.8 Live on three separate rates: $0.75 and $4.50 per 1M text input and output tokens, $3.00 and $12.00 per 1M audio input and output tokens (equivalent to $0.005 and $0.018 per minute), and $1.00 per 1M image or video input tokens ($0.002 per minute). There is no charge for image or video output, since the model does not generate either.

Can you use Gemini 3.8 Live without paying?

Yes, at no cost through Google AI Studio or the unpaid Gemini API quota, with no published rate-limit table for that tier. The tradeoff is data handling: traffic sent this way can feed Google's own product development, a policy that does not apply once billing is turned on.

What are Gemini 3.8 Live's closest competitors?

OpenAI's GPT-Live-1, xAI's Grok Voice, and Amazon's Nova Sonic all target the same real-time, speech-to-speech voice-agent category. GPT-Live-1 pairs with a separate reasoning model behind it for tool use and decision-making, where Gemini 3.8 Live keeps reasoning and speech in one model.

How does Gemini 3.8 Live compare to GPT-Live-1 in 2026?

OpenAI shipped GPT-Live-1 on September 10, 2026 at a flat $0.05 per minute for the voice layer, billing whatever reasoning model sits behind it separately. Gemini 3.8 Live instead meters audio input and output at two different per-minute rates and keeps reasoning inside the same model, so switching to GPT-Live-1 means paying twice: once for the voice layer, once for the model doing the thinking.

How do you set up Gemini 3.8 Live?

Grab a Gemini API key in Google AI Studio, open a WebSocket session against the gemini-3.8-live Live API endpoint, and stream 16-bit PCM audio in while handling the audio and text events that come back. Google lists first-party support for LiveKit, Pipecat, Agora, LangChain, Fishjam, Vercel and Vision Agents if you would rather build on an existing real-time voice framework than the raw WebSocket protocol.

Top Alternatives

  • Gemini 3.8 Flash: Pick Gemini 3.8 Flash if your workload is text and code generation rather than live spoken dialogue.
  • Gemini Omni 1.1 Flash: Pick Gemini Omni 1.1 Flash if you need turn-based multimodal chat rather than a persistent low-latency voice session.
  • Gemini Nano: Pick Gemini Nano if you need on-device, offline inference; Gemini 3.8 Live requires a live cloud WebSocket connection.

More AI Models on HokAI

Visit Gemini 3.8 Live Official Page