A Google DeepMind bemutatta a Gemini 3.8 Live with Live Avatar funkciót, amely valós idejű vizuális megjelenést biztosít a csevegő AI-nak. Az új eszköz az alacsony késleltetésű videógenerálást ötvözi a beszéddel, így az avatarok képesek egyszerre figyelni, látni és beszélni.
Az avatarok pontos szájmozgással, természetes arckifejezésekkel és folyamatos beszélgetésvezetéssel működnek, ráadásul zökkenőmentesen váltanak 97 különböző nyelv között. A rendszer képes a háttérben aszinkron módon adatokat lekérni és külső eszközöket kezelni, miközben a felhasználóval folytatott beszélgetés egy pillanatra sem szakad meg.
A funkció mától érhető el a Gemini Enterprise csomagban, ahol a fejlesztők egyetlen referenciafotó alapján saját, márkára szabott karaktereket is létrehozhatnak. A visszaélések elkerülése érdekében a Google DeepMind minden generált hangot és videót láthatatlan SynthID vízjellel lát el.
Az eredeti szöveg (Google DeepMind)
Gemini 3.8 Live with Live Avatar brings real-time visual presence to Gemini’s conversational AI. By natively coupling our live dialogue capabilities with low-latency streaming video, Live Avatar enables a more natural and intuitive conversational experience for enterprises and their users.
Software Engineer, on behalf of the Gemini Audio Team
Building on the momentum of last week's Gemini 3.8 Live launch, today we are excited to introduce Gemini 3.8 Live with Live Avatar — bringing near real-time visual presence to our native live dialogue models. By pairing near real-time video generation with speech, the Live Avatar feature creates an experience that listens, sees, and speaks with a dynamic visual persona.
With precise lip-syncing, natural expressions, and fluid turn-taking, Live Avatar enables enterprises to expand their virtual offerings more interactively. Whether providing engaging customer service or delivering interactive walkthroughs, it transforms digital exchanges into richer, more accessible experiences.
Starting today, Gemini 3.8 Live with Live Avatar is available in Gemini Enterprise.
See how Gemini 3.8 Live with Live Avatar supports a wide range of characters, each with a distinct look, voice, and expressive presence.
Conversation is inherently multimodal: we listen, look, speak, and use facial expressions to communicate. Live Avatar brings these capabilities to enterprise agents. By processing visual and audio inputs simultaneously, it generates enriching conversations for a more comprehensive experience.
Watch how Gemini 3.8 Live with Live Avatar takes in what it sees and hears in near real time, responding with expressive audio and video for a more natural conversation.
Beyond visual presence, the feature is backed by Gemini’s advanced reasoning. With asynchronous tool calling, Live Avatar can trigger tool calls and fetch data in the background while continuing active dialogue, handling complex tasks while ensuring an uninterrupted conversational flow.
See how Gemini 3.8 Live with Live Avatar handles complex tasks like checking in a guest at a hotel. Calling tools in the background while the dialogue continues uninterrupted.
Conversational presence should feel natural and not be limited by languages. Live Avatar features native multilingual speech-to-speech synchronization. The feature dynamically adapts its lip-sync and expressions and can seamlessly transition across 97 languages without degrading video fidelity or introducing visual drift.
Watch how Gemini 3.8 Live Avatar switches between languages mid-conversation, with lip-sync and expressions adapting seamlessly across 97 languages.
Organizations often need distinct visual identities to fit their brand. In addition to a library of diverse, preset avatars, organizations can customize their Live Avatars. From a high-quality reference image, developers can generate a fully animated, responsive avatar while preserving reference likeness, brand styling, or character identity. Custom avatar creation is currently available only through enterprise allowlisting.
We built Live Avatar with strict safeguards designed to respect identity, and keep AI-generated content transparent. All output generated by our AI products is watermarked with SynthID. This imperceptible watermark is woven directly into the audio and video output, helping to ensure AI-generated content remains detectable to help minimise misinformation and misattribution. To explore our comprehensive approach to safety and responsible deployment, read our model card.
Gemini 3.8 Live with Live Avatar is available in Gemini Enterprise. Explore the API documentation to get started.