Gemini 3.8 text-to-speech blocks voice cloning in EEA
Google shipped two TTS models with 2,000+ voices across 100+ languages, yet voice replication stays blocked in the EEA, UK and Switzerland.
In short
Google released two text-to-speech models on September 23, 2026 — Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS — and withheld their 30-second voice replication feature in six named territories, among them the EEA, the UK and Switzerland.
At a glance
- Two models: Flash TTS for directed performance, Flash-Lite TTS for high-volume workloads.
- Voice library: 2,000+ production-ready voices spanning more than 100 languages and dialects.
- Replication needs a 30-second sample plus a matching verbal consent recording from the owner.
- Hume AI benchmark: 71.4 on voice design and 60.8 on accent modeling, first place in both.
- Every generated audio file carries a SynthID watermark and C2PA credentials.
Google introduced Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS on September 23, 2026. Both turn text into speech and both stamp an inaudible watermark into the output. The headline capability — rebuilding an existing voice from a 30-second sample — is switched off in six named territories.
What each model is for
Flash TTS is pitched at directed performance: pacing, emotion and acting cues can be set line by line, and two speakers can be staged in a conversation that takes turns naturally. Flash-Lite TTS targets volume and cost. Both are live through the Gemini API and Google AI Studio, while the enterprise endpoint is announced but not yet open.
Consumer surfaces split the two apart: Flash TTS powers Gemini Notebook, Flash-Lite TTS powers Google Vids.
A catalog that starts at 30 voices
The library opens with 30 original voices and, by Google's account, expands to more than 2,000 production-ready options across over 100 languages and dialects. Regional variants are part of the pitch, including Mexican Spanish, Quebec French and Scots English. On top of that, a plain-language prompt can define a fresh voice by role, accent and timbre.
Consent first, then a map of exclusions
Copying a real voice requires a verbal consent recording from the voice owner that matches the reference speaker. Even so, the feature never turns on in the EEA, the UK, Switzerland, India, Illinois or Texas. The post gives no reason for the list, though it lines up closely with places that regulate biometric identifiers tightly.
For teams in those markets the practical rule is narrow: design a voice, do not clone one.
Where the models rank
On the Hume AI voice design benchmark Flash TTS scores 71.4 and takes first place; on accent modeling it scores 60.8, again first. The two models hold the top two slots on the overall quality index. On the Voice Arena leaderboard they lead for Japanese, Brazilian Portuguese, Vietnamese, Modern Standard Arabic, Mexican Spanish and Hindi.
What the post leaves out
No per-character or per-token price appears anywhere in the announcement, and no latency figure does either, so neither is confirmed. Exact API model identifiers are also missing. The separately listed DeepMind link redirects to this same post, leaving a single primary source for now.
FAQ
Which Gemini 3.8 text-to-speech models did Google launch?
Two of them: Gemini 3.8 Flash TTS for directed, character-driven audio and Gemini 3.8 Flash-Lite TTS for high-volume, cost-sensitive work. Both are reachable through the Gemini API and Google AI Studio.
Is Gemini 3.8 voice cloning available in the EU or the UK?
No. Voice replication is unavailable in the EEA, the UK, Switzerland, India, Illinois and Texas. Designing an entirely new synthetic voice from a prompt is not affected by that restriction.
How much does Gemini 3.8 text-to-speech cost?
The announcement lists no pricing and no latency numbers. Anyone budgeting for it has to wait for the Gemini API price sheet; the blog post supports no estimate either way.