Gemini 3.8 Live gains a real-time video avatar
Google pairs its live voice sessions with an animated on-screen face, lip-synced across 97 languages, and ships it to enterprise customers first.
In short
As of September 24, 2026, Gemini 3.8 Live can answer through an animated avatar inside Gemini Enterprise and over the Agent Platform API, with SynthID watermarks embedded in both the audio and the video.
At a glance
- Announced September 24, 2026: Gemini 3.8 Live with Live Avatar, available in Gemini Enterprise the same day.
- Developer access runs through the Gemini Enterprise Agent Platform; custom avatars require allowlisting.
- Speech-to-speech output is lip-synced across 97 languages, which Google says holds without drifting.
- Audio and video are stamped with SynthID so generated material stays machine-detectable.
- The announcement lists no pricing, no rate limits and no date for consumer accounts.
Google put a face on its live voice model. Gemini 3.8 Live with Live Avatar, announced on September 24, 2026, streams an animated character whose mouth moves in step with the model's speech, and it lands in Gemini Enterprise first, with API access through the Gemini Enterprise Agent Platform.
A face bolted onto a voice pipeline
The system takes audio and video in at the same time and manages turn-taking during a conversation. Google says lip movement adapts to whichever language is being spoken and does not drift over the course of a session. The company puts the count for direct speech-to-speech output at 97 languages.
The less visible addition matters more for how this feels: tool calls can now run asynchronously, so the model fetches data in the background while talking. A talking head that freezes mid-sentence while a lookup completes is worse than no talking head at all.
Enterprise first, custom faces gated
At launch the feature is limited to Gemini Enterprise, with developers reaching it through the Agent Platform studio. Preset avatars are available to those customers, and custom avatars can be built from reference images, but that path sits behind allowlisting rather than being open to anyone. Google publishes no pricing, no rate limits and no consumer timeline.
Every clip carries a SynthID mark
Output is watermarked with SynthID, which Google describes as “imperceptible” and woven into the audio and video themselves. The goal is provenance: a recording should still be identifiable as model output after it leaves the product.
That is a meaningful precaution for this specific feature, because a face and a voice buy trust that a text box never does. What the announcement does not address is how well the mark survives re-encoding or a phone pointed at a screen.
What we could not verify
Only Google's own announcement was reachable for this report; a second independent write-up could not be retrieved. Everything beyond the company's wording is therefore unconfirmed, including measured latency, how the lip-sync holds up outside a controlled demo, and whether workers actually want a face answering them.
FAQ
What is Gemini 3.8 Live with Live Avatar?
It extends Gemini's live conversations with an animated avatar that streams video lip-synced to the model's speech, while taking audio and video in at the same time.
Who can use the Gemini Live Avatar right now?
Gemini Enterprise customers, plus developers working through the Gemini Enterprise Agent Platform. Building a custom avatar from reference images requires allowlisting, and Google gives no consumer release date.
How can you tell a Gemini avatar video is AI-generated?
Google watermarks the audio and video with SynthID, which it calls imperceptible. You cannot spot it by eye; detecting it takes a tool built to read the watermark.