Aimed at audiobooks, podcasts, game characters and dubbing where delivery matters more than cost, Gemini 3.8 Flash TTS is the fidelity-first sibling of Gemini 3.8 Flash-Lite TTS. Each request accepts at most 8,192 input tokens and 2 speakers, so long scripts and larger casts are split across calls.
What changed
Gemini 3.8 Flash TTS is Google DeepMind's flagship text-to-speech model, released September 22, 2026, that converts text into expressive audio in 130 languages. It supports one or two speakers, 30 prebuilt voices and custom voices made through Voice design or Voice replication, and returns audio only.
Where it sits
- $2.63/M$ per 1M tokensBlended price (3:1)Lower is better#44 / 79peer median $1.71/Mvendor price, checked by HokAI
Priced around the middle of the 79 GA models with a published price (rank 44), and one of 31 that document a zero-data-retention option.
Ranks are against GA models on HokAI that publish the same figure; ties share a rank.
Provider: Google DeepMind · Family: Gemini 3.8
More about Google DeepMind on HokAI
Context window: 8,192 tokens · Max output: 16,384
Input modalities: text · Output: audio
About Gemini 3.8 Flash TTS
Gemini 3.8 Flash TTS (model code gemini-3.8-flash-tts) is the text-to-speech model that Google DeepMind made generally available in the Gemini API on September 22, 2026, together with the cheaper Gemini 3.8 Flash-Lite TTS. Google's launch post is dated September 23. The model card places it in the Gemini 3.8 Audio group, built on Gemini 3 Pro with a January 2025 knowledge cutoff, and it follows Gemini 3.1 Flash TTS Preview from April 2026. It belongs to the audio line of Google DeepMind, next to Gemini 3.8 Live for spoken conversation and Gemini 3.5 Transcribe for the reverse direction, and unlike Gemini 3.8 Flash it only reads text in and writes audio out.
The independent evidence is a listening test, not a benchmark table. On the Artificial Analysis Speech Arena, where listeners vote blind between two clips of the same text, the model held an Elo of 1272 (95% interval 1257 to 1287, 2,677 samples) when we read the leaderboard on October 11, 2026. That is about 28 points above its Flash-Lite sibling (1244) and 60 above Gemini 3.1 Flash TTS (1212), but still short of ElevenLabs. Google adds its own figures: 71.4 on Hume AI's Voice Design Benchmark and 60.8 on accent modeling, plus the claim that this model and Flash-Lite lead Hume's overall quality index. Hume's public TTS leaderboard, read the same day, lists Google's Gemini 3.8 entries at 0.92 and 0.91 on its overall score. Treat the Hume numbers as vendor-reported until you listen yourself.
Access is the part most likely to trip up a team. Developers get the model today through the Gemini API and Google AI Studio, and consumers meet it inside Gemini Notebook. Enterprise access runs through Google Cloud's Gemini Enterprise Agent Platform, where the model page still carries a Preview label, a September 28, 2026 release date and a single global location (HokAI's Google Vertex AI listing covers the hosting side). Google names LiveKit, Pipecat, Vercel and Agora as developer platforms that wire it in, and Figma, HeyGen, Linguana, Wondercraft, 99.co and Ollang as companies integrating it for dubbing, localization and voice agents.
Voice handling carries the most editorial weight. Replication needs a verbal consent recording that matches the reference speaker, every clip carries a SynthID watermark, and Google says replicated voices also get C2PA credentials. In AI Studio, voice replication is unavailable in Illinois, Texas, the EEA, the UK, Switzerland and India. One thing Google has not settled is the size of the Extended Voice Library: the launch post and Cloud docs say more than 2,000 curated voices, the Gemini API guide says hundreds, and the API changelog says 150 or more are queryable, so verify the count against your own account before promising a client a specific voice.
Where it fits: choose it over ElevenLabs or Cartesia when a long, multilingual, budget-sensitive job matters more than the last increment of English naturalness, and compare it with finished apps such as Murf AI, Speechify and Fish Audio if you want an editor instead of an API. Skip it for live conversation (the Live API does not serve it), for casts of more than 2 voices in one request, and for scripts past the 8,192-token input limit. More options sit in the AI voice assistants category and the Google DeepMind model directory.
Pricing
Gemini API list prices, USD per 1M tokens, read from Google's pricing page on October 11, 2026. Standard: $0.50 text input and $9.00 audio output through December 31, 2026, then $1.00 and $18.00 from January 1, 2027. Batch and Flex: $0.25 and $4.50 through 2026, then $0.50 and $9.00. Priority: $0.90 and $16.20, then $1.80 and $32.40. Google counts 25 audio tokens per second, so 10 seconds of speech is 250 tokens, about $0.00225 now and $0.0045 after the change. Google's table lists the free tier as free of charge on Standard and Priority and not available on Batch or Flex, and context caching is not supported. For other voice products see [ElevenLabs](/hub/tools/elevenlabs) and [Resemble AI](/hub/tools/resemble-ai).
What a real job costs
| Job | Input | Output | Total |
|---|---|---|---|
| Summarise a 20-page PDF | $0.015 | $0.0090 | $0.024 |
| Support reply | $0.0010 | $0.0027 | $0.0037 |
| One coding agent run | $0.100 | $0.180 | $0.280 |
Budgets: 20-page PDF = 30k in / 1k out · Support reply = 2k in / 300 out · Coding agent run = 200k in / 20k out. Computed from the vendor's per-token prices at render time; cached-input discounts are not applied.
Key Features
- Two-speaker scenes in one request: Configures 2 prebuilt-voice speakers per call, with each turn tagged by speaker and an optional style, and the REST API adds a conversational mode for natural turn-taking.
- Style field plus inline vocal tags: Sustained delivery such as pace or whispering goes in speech_metadata.style, while moments like <laugh>, <sigh> and <short pause> sit inline, and |reaction| segments create overlapping backchannels.
- Voice design: Creates a persistent voice from a written description through POST /v1beta/voices, returning a voice ID and a WAV preview, with room for 200 stored voices per project that expire a year after last use.
- Voice replication: Recreates a voice from reference and consent audio, stored as a voice ID by default or as a stateless voicekey that lasts 7 days.
- Telephony-ready output formats: Returns 24 kHz mono WAV for single requests and raw PCM chunks when streaming, with mu-law and A-law options at sample rates down to 8,000 Hz.
- Automatic language detection: Detects the input language itself across 130 languages and scripts, including Cantonese, Swahili and several Indian regional languages.
Pros
- Artificial Analysis lists its per-character price at a fraction of ElevenLabs Eleven v4, which matters for audiobook-length jobs.
- Google says voice identity, timbre and room tone hold across hours of continuous audio, the claim to test first for long narration.
- It shares one API schema with Flash-Lite TTS, so moving between the two tiers is a one-parameter change.
Cons
- Blind listener votes put it behind ElevenLabs Eleven v4 and Eleven v4 Turbo, so top English polish still favors the rival.
- Text goes in and audio comes out, with no Live API, function calling or caching, so a voice agent needs an LLM in front of it.
- Voice replication through AI Studio is blocked in several regions, so some teams need another route.
Frequently Asked Questions
What does Gemini 3.8 Flash TTS actually cost?
On the paid Gemini API tier it charges $0.50 per 1M text input tokens and $9.00 per 1M audio output tokens until the end of 2026, then it doubles to $1.00 and $18.00 on January 1, 2027. Audio output is metered at 25 tokens per second, which puts an hour of generated speech near $0.81 of output today and $1.62 after the change. Batch and Flex halve both rates, so schedule bulk narration there before the price step.
How does Gemini 3.8 Flash TTS compare to ElevenLabs Eleven v4 in 2026?
In Artificial Analysis blind votes Eleven v4 sits at 1323 and Eleven v4 Turbo at 1327, roughly 50 Elo points above Gemini 3.8 Flash TTS, so [ElevenLabs](/hub/tools/elevenlabs) still wins on raw listener preference. Artificial Analysis lists Gemini at $16.5 per 1M characters against $80.0 for Eleven v4, and Gemini adds built-in voice design. Pick Gemini for long, cost-sensitive or multilingual work, and pick Eleven v4 when English polish decides the project; audition both on your own script.
Is Gemini 3.8 Flash TTS open source or proprietary?
It is proprietary and hosted by Google only, so there are no downloadable weights, no self-hosting and no fine-tuning of the base model. You reach it through the Gemini API, Google AI Studio, the enterprise platform preview and Google's own apps. Use falls under the Gemini API Additional Terms of Service and Google's Generative AI Prohibited Use Policy.
Does Gemini 3.8 Flash TTS train on your data?
Google's pricing page marks free-tier content as used to improve its products and paid-tier content as not used, so scripts you want kept private belong on a paid project. The page does not state a retention period, so check Google's data logging policy before sending sensitive text. Voice replication adds its own consent step, where the voice owner records spoken approval.
Who should use Gemini 3.8 Flash TTS and who should pick something else?
It suits audiobooks, podcasts, game characters and dubbing, where acting and long-run voice stability matter more than latency. Pick Gemini 3.8 Flash-Lite TTS for high-volume agent cascades and read-aloud features, and [Gemini 3.8 Live](/hub/models/gemini-3.8-live) when one session must listen and speak. Plan on one call per turn if your cast is bigger than 2 or mixes designed voices.
Top Alternatives
- Gemini 3.8 Live: Pick Gemini 3.8 Live if the agent must listen and answer in one real-time voice session; pick Gemini 3.8 Flash TTS for scripted narration you want to direct line by line.
- Gemini 3.5 Transcribe: Pick Gemini 3.5 Transcribe if you are turning speech into text; Gemini 3.8 Flash TTS goes the other way and only turns text into audio.
- Gemini Omni 1.1 Flash: Pick Gemini Omni 1.1 Flash if you need video with synchronized audio; pick Gemini 3.8 Flash TTS when audio alone is the deliverable and voice control matters.