Google is adding a face to its live voice models. Gemini 3.8 Live with Live Avatar, introduced today, pairs near real-time video generation with speech so its native live dialogue models can listen, see, and respond through a visual persona. The feature is available starting today in Gemini Enterprise, and Google did not announce a consumer rollout or a price.
The pitch is aimed at companies rather than individuals. Google says Live Avatar is meant to help enterprises expand their virtual offerings, naming customer service and interactive walkthroughs as examples of what the visual persona can handle. The company describes the experience as one that listens, sees, and speaks, with precise lip-syncing, natural expressions, and fluid turn-taking. Those are Google's characterizations of the feature, not independently measured results.
Google says the feature builds on the Gemini 3.8 Live release it announced last week. That earlier wave of voice work was aimed squarely at developers and enterprises building near real-time voice agents, with Gemini 3.8 Live positioned for scale and cost efficiency and a companion Extended Thinking model targeting harder, multi-step jobs. Google disclosed no pricing then either, and left availability questions open for anyone planning to build on the models.
The developer-facing side of that push arrived in the Gemini API and Google AI Studio, where Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking shipped alongside Gemini 3.5 Transcribe, a speech-to-text model supporting more than 85 languages. Google said at the time that Gemini 3.8 Live could hold a conversation while carrying out tasks, and that Extended Thinking ranked first on Artificial Analysis' leaderboard, a claim the company made rather than an independent finding.
What the new notice does not say is how Live Avatar is priced, which languages it supports, or when it might reach developers outside Gemini Enterprise. Google also gives no independent evaluation of the lip-syncing or turn-taking quality it describes. For enterprises weighing a virtual agent, the practical question of cost per interaction remains unanswered.
The announcement lands in a busy stretch for Google's voice stack. Text-to-speech models Gemini 3.8 Flash TTS and 3.8 Flash-Lite TTS were introduced a day earlier, and the company has also been pushing video creation tools into Google Vids. Read together, the releases suggest Google is treating voice and visual presence as a product line of its own rather than a feature attached to its text models.













