Skip to content
Five.Reviews
Menu

How-To & Tutorials

How to Use Gemini 3.8 Live With Live Avatar

Analytics dashboard on a laptop screen used to represent evidence-led software evaluation
Free browser-based audio. No tracking or paid API required.

This guide covers both paths: the no-code Google Cloud console workflow and the developer workflow through the Gemini Live API.

Voice assistants can answer questions, but they rarely give users a face to talk to. Gemini 3.8 Live with Live Avatar is Google’s approach to real-time conversational AI with synchronized avatar video. Google announced it on September 24, 2026, and it is now generally available in Gemini Enterprise. Below you will find what it is, where to access it, how to configure prebuilt and custom avatars, how camera and screen input work, how to build with the API, and which limits to plan for. Product details come from Google’s announcements and documentation and may change.

Quick answer: 

Gemini 3.8 Live with Live Avatar pairs Google’s real-time speech-to-speech model, gemini-3.8 Live, with synchronized avatar video. To try it, open the Google Cloud console, go to Agent Platform > Studio > Stream realtime, select gemini-3.8-live, choose Live Avatar, pick an avatar and voice, and click Start Session. Developers connect through the Gemini Live API over WebSocket.

Key takeaways

Which setup should you use?

If you want to…Use
Test the experienceGoogle Cloud Agent Platform
Use a prebuilt avatarLive Avatar configuration
Build an applicationGemini Live API
Add a custom likenessCustom avatar, if allowlisted
Process camera or screen contextGemini Live multimodal input

What Is Gemini 3.8 Live With Live Avatar?

Gemini 3.8 Live with Live Avatar combines Gemini’s real-time conversational model with synchronized avatar video, so an AI agent can listen, understand multimodal input, reply in speech, and present that reply through a talking digital avatar.

What Gemini 3.8 Live does

Google describes Gemini 3.8 Live as a real-time model for bidirectional voice and video interaction. It is native speech-to-speech, so audio streams in and out without a traditional speech-to-text pipeline, and Google says it recovers from interruptions without losing context. It also accepts text, live camera feeds and screen shares, runs tool calls in the background, and, per Google, speaks 97 languages with automatic language detection.

What Live Avatar adds

Live Avatar is the presentation layer. Google says it natively couples live dialogue with low-latency streaming video, producing lip-synced, expressive video at 24 FPS. The reasoning stays with the model. The avatar is how that reasoning shows up on screen.

Gemini 3.8 Live vs. Gemini 3.8 Live With Live Avatar

Live Avatar is not a separate intelligence model. It adds synchronized visual output to the Gemini 3.8 Live experience.

FeatureGemini 3.8 LiveGemini 3.8 Live + Live Avatar
Real-time voiceYesYes
Audio inputYesYes
Camera inputYesYes
Screen inputYesYes
Tool callingYesYes
Audio responseYesYes
Avatar videoNoYes
Lip synchronizationNo avatarYes
Prebuilt avatarsNoYes
Custom avatarNot applicableAllowlist only
Live APIYesYes

Read More: Gemini 3.8 Flash: Benchmarks, Pricing, Features & Performance 

How to Set Up Gemini 3.8 Live With Live Avatar

What you need

Before starting the Live Avatar session, make sure you have:

For a first test, you do not need to configure the API or build a WebSocket connection. The Google Cloud Studio workflow lets you configure and start a Live Avatar session directly.

Step 1: Open Agent Platform Studio

Sign in to the Google Cloud Console and open:

Agent Platform > Studio > Stream real-time

This opens Google’s real-time model testing interface. Make sure you are working in the Google Cloud project that has access to the Gemini Live features you want to use.

If you do not see Stream realtime, or the available options differ from the steps below, check your project selection and account access before continuing.

Step 2: Select Gemini 3.8 Live

In the Studio interface, click Switch model and select:

gemini-3.8-live

Once the model loads, select Live Avatar from the main configuration panel.

Selecting the Live model first is important because the avatar configuration is tied to the real-time Gemini experience. The Live Avatar option adds synchronized avatar video to the model’s real-time spoken interaction.

Step 3: Select an avatar

Open the Avatar selector and choose one of the available built-in avatars.

The selected avatar becomes the visual character that presents Gemini’s responses during the live session. For your first test, use a prebuilt avatar rather than trying to configure a custom avatar, since custom avatars require separate enterprise access.

You can change the avatar before starting the session, so you can test different available options without changing the underlying model.

Step 4: Select a voice

Open the Voice selector and choose a supported prebuilt voice.

The voice determines how Gemini speaks, while the avatar determines how those spoken responses are visually presented. Google documents that prebuilt avatars can be paired with prebuilt HD voices available through the Live API.

For a natural first test, choose a voice that matches the role you gave the assistant. A support assistant, for example, may benefit from a clear conversational voice rather than an overly expressive one.

Step 5: Add system instructions (optional)

Use the system instructions field to define how Gemini should behave during the conversation.

You can specify the assistant’s role, tone, response length, and how it should handle unclear requests. This gives the Live Avatar a more consistent personality and interaction pattern than simply starting with a blank session.

For example:

You are a concise customer support assistant.

Answer naturally and keep responses brief.

Acknowledge the user’s request before providing a solution.

Ask one clarifying question when the request is ambiguous.

Keep the instructions focused on behavior rather than adding a long prompt. You can refine them after testing the first session.

Step 6: Start the session

Before starting, optionally enable camera input if you want Gemini to see live visual information. If you only want a voice conversation, leave the camera disabled.

Then click Start Session.

When your browser asks for microphone or camera access, select Allow for the devices you have enabled. Once the session connects, speak into the microphone and wait for the avatar to respond.

For the first test, use a simple prompt such as:

Explain what you can help me with in three short sentences.

After confirming that the avatar responds correctly, try a follow-up question to test the live conversation. If camera input is enabled, you can also ask Gemini about something visible to the camera.

The response should combine spoken output with synchronized avatar video, giving you the complete Gemini 3.8 Live with Live Avatar experience.

Quick setup checklist: Google Cloud project selected, Agent Platform Studio opened, gemini-3.8-live selected, Live Avatar enabled, avatar and voice chosen, system instructions added if needed, microphone permission allowed, optional camera input configured, and the session started.

How to Use a Custom Gemini Avatar

Custom avatars are not universally available. Google states that custom avatar creation is currently limited to enterprise allowlisting. Google’s Cloud announcement says to contact your Google Cloud sales representative about allowlisting.

The flow is: reference image → avatar configuration → Gemini 3.8 Live → synchronized avatar video. Google says developers can generate a fully animated avatar from a high-quality reference image while preserving likeness, brand styling, or character identity. In the API, you set avatar_config.customized_avatar with the base64 image data and its MIME type.

Prebuilt avatarCustom avatar
AvailabilityEnterprise customersAllowlist only
SetupPick from a listSupply a reference image
VoicePrebuilt HD voicesPrebuilt or custom voice

Google’s documented image guidance:

Consent matters. Google’s overview describes custom avatars generated from authorized employee likenesses or paid talent with executed corporate releases. Never use someone’s likeness without written authorization.

How Gemini 3.8 Live Handles Camera and Screen Input

Gemini 3.8 Live can process live camera feeds and screen shares alongside audio at the same time, which lets an agent see what the user sees.

Camera input suits moments where showing beats describing: pointing at a damaged product, demonstrating a fault, or asking the agent to identify something visible. Google’s own demo has a claims intake agent that watches damage on camera while the conversation continues.

Screen sharing suits software troubleshooting, product walkthroughs, and guided onboarding. Cox Automotive’s Autotrader assistant, which Google cites, uses live screen highlighting and tool calling to guide shoppers.

The data flow is: camera or screen → visual frames → Gemini 3.8 Live → reasoning → speech response → Live Avatar.

Video is not streamed as continuous full-frame footage. Google’s documentation says the model takes 1 FPS JPEG frames, with 768×768 as the optimal resolution. That makes it a poor fit for fast-changing scenes. The console page documents an optional camera toggle. I found no documented screen-share toggle there, so treat screen sharing as an API-level capability.

How to Build Gemini 3.8 Live With Live Avatar Using the API

Developers use the Gemini Live API, which supports the Google Gen AI SDK and direct WebSocket connections.

User

  ↓

Microphone / Camera / Screen

  ↓

Application

  ↓

Gemini Live API

  ↓

Gemini 3.8 Live

  ↓

Audio + Avatar Video + Tool Calls

  ↓

User

Expert insight: The important distinction is not that Gemini can talk through a face. Real-time multimodal input, speech-to-speech interaction, tool calling, and synchronized video output together let you build an interactive agent rather than a prerecorded avatar.

WebSocket architecture

The native Live API is a stateful WebSocket (WSS) connection. Your app opens a socket, sends a setup message, then streams input and receives output in both directions. WebRTC is not the native transport. Third-party platforms may wrap the API with WebRTC, but that is their layer.

Input audio is 16 kHz, 16-bit, little-endian, mono PCM, sent in chunks of 20 to 100 ms. Output audio is 24 kHz.

Prebuilt avatar configuration

Per Google’s documentation, set response_modalities to [“VIDEO”], name the avatar in avatar_config, and set the voice in speech_config. Replace the placeholders with values from the docs and confirm exact field placement against Google’s full sample.

python

SETUP_MESSAGE = {

    “setup”: {

        “model”: “gemini-3.8-live”,

        “generation_config”: {

            “response_modalities”: [“VIDEO”],

            “speech_config”: {

                “voice_config”: {

                    “prebuilt_voice_config”: {“voice_name”: “VOICE_NAME”}

                }

            },

        },

        “avatar_config”: {“avatar_name”: “AVATAR_NAME”},

        “system_instruction”: {

            “parts”: [{“text”: “You are a concise support assistant.”}]

        },

    }

}

Custom avatar configuration

Swap the avatar block for the reference image. This works only if your project is allowlisted.

python

“avatar_config”: {

    “customized_avatar”: {

        “image_data”: avatar_b64,   # base64-encoded reference image

        “image_mime_type”: “png”,

    }

}

Authentication

Authenticate with Google Cloud OAuth 2.0 bearer tokens in the WebSocket headers. Never expose credentials in client-side production code. Route browser traffic through your own backend, or use a short-lived credential mechanism that Google documents for your setup.

Tool calling

Asynchronous tool calling lets the agent keep talking while work happens in the background. A user says, “Check whether my order has shipped.” Gemini replies, “I’ll check that for you,” your backend runs the lookup, and the conversation continues until the result arrives. You define the tools; the model does not gain backend abilities on its own.

How to Reduce Latency and Improve Gemini Live Avatar Quality

Google does not publish a guaranteed latency figure in the sources reviewed, so measure your own.

Audio: send the exact sample rate above, keep chunks within 20 to 100 ms, and avoid extra processing before transmission.

Video: set media_resolution to LOW, MEDIUM, or HIGH based on how much detail the task needs. Google says this balances per-frame token use against visual detail. Send 768×768 frames at 1 FPS, not larger or faster.

Conversation: keep system instructions focused, ask for concise replies, run tools asynchronously, and handle interruptions cleanly.

Custom avatar quality: follow the image guidance above. Sharp, well-framed, front-facing images with plain backgrounds give the model the best input.

Gemini 3.8 Live Avatar Availability, Limitations, and Considerations

Best Use Cases for Gemini 3.8 Live With Live Avatar

Common Mistakes When Setting Up Gemini Live Avatar

Conclusion

Gemini 3.8 Live with Live Avatar gives you real-time conversational intelligence plus synchronized visual presence. Test it in the Google Cloud console through Agent Platform > Studio > Stream real-time, then build with the Gemini Live API over WebSocket. Custom avatars stay restricted, while camera, screen understanding, and background tool calling make it a working agent, not just a talking avatar. Your next step: run one session with a prebuilt avatar, then ask your Google Cloud representative about allowlisting if branding needs it.

Frequently Asked Questions 

What is Gemini 3.8 Live with Live Avatar?

Gemini 3.8 Live with Live Avatar combines real-time conversation with a synchronized, lip-synced avatar that speaks responses aloud. The avatar streams at 24 FPS with human-like expressions, making interactions more natural for support bots and conversational AI.

How do I use Gemini Live Avatar in the console?

In Google Cloud Console, go to Agent Platform > Studio > Stream real-time, select gemini-3.8-live, and enable Live Avatar. Choose an avatar and voice, configure instructions, and start a session; developers can use the Live API with WebSockets and OAuth 2.0 or service-account authentication.

Where can I access Gemini 3.8 Live with Live Avatar?

It is available through Google Cloud’s Gemini Enterprise plan and the Gemini Live API with consumption-based pricing. Google AI Studio does not currently support Live Avatar, so Cloud access is required.

Is it available in Google AI Studio?

No. Google AI Studio supports Gemini Live for text and voice, but Live Avatar is available through Google Cloud’s Agent Platform for avatar configuration and enterprise API management.

Can I create a custom avatar?

Yes, but custom avatars require enterprise allowlisting rather than self-service access. You can upload a portrait PNG meeting Google’s requirements, then request access through your Google Cloud sales representative.

Does it support camera input?

Yes. The API accepts JPEG camera frames at 1 FPS with configurable resolutions, making it suitable for periodic visual understanding such as object and scene recognition rather than continuous video analysis.

Can the avatar understand screen sharing?

Yes. You can send screen captures as video frames at 1 FPS, allowing the avatar to reference documents, dashboards, or code. This is useful for technical support, training, and collaborative problem-solving.

Why WebSockets instead of WebRTC?

The Gemini Live API uses WebSockets for stateful, bidirectional audio and video streaming without WebRTC’s peer-to-peer negotiation. This simplifies implementation and suits cloud-based, server-side agent architectures.

Which API does it use?

It uses the Gemini Live API with model ID gemini-3.8-live. Developers configure response_modalities as [“VIDEO”] and avatar_config, with authentication through OAuth 2.0 or Google Cloud service accounts.

How are speech and video synchronized?

Google generates speech and avatar video together at the API level, keeping the avatar’s facial movements synchronized with spoken responses. This avoids the timing issues common when audio and video are generated separately.

What’s the minimum latency I should expect?

Gemini Live typically delivers around 200–600 ms of end-to-end latency, depending on network, model load, and video resolution. Actual performance can vary, so testing with your specific setup is recommended.

Are there usage limits or quotas?

Google Cloud applies quotas based on billing tier and region, with usage metered through tokens and video output duration. Check the Google Cloud quota dashboard and pricing information for current limits and request increases through Support if needed.