Gemini 3.5 Transcribe: Google's new speech-to-text model
Google reports a 2.6 percent word error rate for the new transcription model and a 70 percent faster time to a final transcript than Chirp 3.
Illustration: a studio microphone in dim light, with sound waves dissolving into floating particles above it.
Gemini 3.5 Transcribe is Google's new speech recognition model, which the company says covers more than 85 languages and reaches a 2.6 percent word error rate in its non-streaming mode.
At a glance
- Two variants: Gemini 3.5 Transcribe for recordings, Gemini 3.5 Transcribe Live for streaming with sub-second latency.
- Word error rate per Google: 2.6 percent non-streaming, 4.0 percent when streaming.
- On the multilingual FLEURS benchmark Google cites 5.04 percent non-streaming and 5.50 percent streaming.
- Speaker separation with timestamps for up to three voices; beyond three speakers is labeled experimental.
- Public preview through the Gemini API, AI Studio and Antigravity. The post lists no pricing.
Google introduced Gemini 3.5 Transcribe on August 26, 2026, a model that does more than turn speech into text: it tidies the text as it goes. Filler sounds are dropped, a mid-sentence correction is applied rather than transcribed twice, and formatting is handled by the model. Teams can register domain terms through a custom vocabulary.
One model, two modes
Finished audio goes to Gemini 3.5 Transcribe; live conversation goes to Gemini 3.5 Transcribe Live. Google describes the live path as bidirectional streaming with latency under one second, able to switch languages mid-conversation. Function calling lets the model hand subtasks to other Gemini models. Speakers are separated with timestamps, cleanly for up to three voices; past three, Google marks the capability experimental.
The numbers Google published
Google puts the word error rate at 2.6 percent without streaming and 4.0 percent with it. On the multilingual FLEURS benchmark the figures are 5.04 percent and 5.50 percent respectively. Time to a final transcript is said to be 70 percent faster than the previous Chirp 3 model. The company says the model automatically detects and transcribes more than 85 languages, including regional accents and dialects.
Where it already ships
Developers get it as a public preview through the Gemini API in Google AI Studio and through Google Antigravity, exposed as gemini-3.5-transcribe-live in the Live API and gemini-3.5-transcribe in the Interactions API. Enterprises get a preview on the Gemini Enterprise Agent Platform. On the consumer side it powers the Gemini app on macOS in English and the Rambler feature in Gboard on Android in selected countries, with Chrome announced but not shipped. Seven platforms have already wired up the Live API: Agora, Fishjam, LangChain, LiveKit, Pipecat, Vercel and Vision Agents.
What we could not check
Every figure above comes from Google's own announcement. At the time of writing we had no second, independent newsroom confirming or contextualizing it, and no independent re-measurement of the error rates exists so far. Google does not describe the test conditions or the audio quality behind the numbers. Pricing, quotas and latency guarantees are absent from the post, as are the specific countries and languages where Rambler is live and any date for Chrome.
FAQ
What is Gemini 3.5 Transcribe?
A Google speech recognition model that converts audio to text while removing filler words, applying spoken self-corrections and formatting the output. It comes in one variant for recordings and one for real-time streaming.
How accurate is Gemini 3.5 Transcribe?
Google reports a 2.6 percent word error rate without streaming and 4.0 percent with streaming, and 5.04 percent versus 5.50 percent on the multilingual FLEURS benchmark. None of these figures has been independently verified so far.
Where can I use Gemini 3.5 Transcribe?
As a public preview through the Gemini API in Google AI Studio and through Google Antigravity, for companies via the Gemini Enterprise Agent Platform, and for consumers in the Gemini app on macOS and the Rambler feature in Gboard on Android.