Muse Avatar, OpenAI Dots, Gemini Live Avatar vs FaceMode
Compare Muse Avatar, OpenAI Dots, Gemini Live Avatar, and FaceMode: platform agents versus a talking avatar API for your own voice stack.
Muse Avatar, OpenAI Dots, and Gemini Live Avatar put different kinds of visual presence around AI agents. For a developer, the first question is not which looks best: it is whether the product supplies an agent experience or a talking avatar API you can attach to the voice agent you already run.
Meta's Muse Realtime Avatar research announcement appeared on September 23, 2026. Google's Gemini 3.8 Live with Live Avatar announcement followed on September 24, and OpenAI introduced Dots on September 29. They are related developments, but they are not interchangeable integration options.
This is FaceMode's documentation-based category comparison, checked on October 2, 2026. We have not run an independent visual-quality or latency benchmark across these products. The differences below describe documented product boundaries, not a claim that one provider wins every workload.
The short version
| Product | What the reviewed sources describe | Developer integration question |
|---|---|---|
| Meta Muse Avatar | Avatar rendering coupled to Muse Realtime Voice; images can drive expressive characters | The research post does not document an external PCM-to-avatar API for your room |
| OpenAI Dots | Persistent assistants with cloud computers, connected apps, messaging, and voice calls | An assistant product, not a documented lip-synced video output API |
| Gemini Live Avatar | Avatar capability natively paired with Gemini 3.8 Live, available in Gemini Enterprise | Evaluate as part of the Gemini Live conversation stack |
| FaceMode | Audio-driven avatar API and plugins, publishing into your LiveKit or Daily room | Evaluate as an added visual layer for your existing audio pipeline |
The table separates categories before comparing features. A hosted assistant may be the right product for an employee who wants work delegated. A video-generation layer may be the right component for a team building a customer-facing voice application. Both needs are real; they do not share the same acceptance criteria.
Meta Muse Realtime Avatar: shared speech tokens and video
Meta's research post describes Muse Realtime Voice and Muse Realtime Avatar as a single streaming system. The voice model produces speech tokens carrying both content and delivery. An audio decoder turns those tokens into speech, while the avatar model consumes the same stream to generate the visual performance.
The avatar model is an audio-driven Diffusion Transformer conditioned on the speech-token stream, reference media, and recent video context. It produces video in short causal chunks and carries recent generated context into the next chunk. That is the technical mechanism Meta describes for maintaining continuity during a live conversation.
Meta reports 448×768 portrait video at 25fps and approximately 870ms from the end of a user's turn to the first byte of a synchronized voice-and-video response. This is a vendor-reported end-of-turn response metric, not the time taken to render an isolated audio frame and not a measured browser playback delay from our own test.
The post also describes company-run live-call preference tests against Runway Characters and HeyGen LiveAvatar. Raters preferred Muse overall, while the mannerism comparison with Runway was not statistically distinguishable from parity. That result should be attributed to Meta's evaluation and its tested conditions, not turned into a universal ranking of every avatar API.
For integration planning, the important limit is what the post actually documents. It explains Muse's coupled voice-and-avatar system and notes that its examples do not all reflect avatars available in the Muse app. It does not provide a public recipe for sending your own PCM audio to Muse Avatar or publishing its output into your LiveKit room. That is a limit of the reviewed documentation, not proof that Meta will never offer such a product.
OpenAI Dots: an assistant with a visual identity
OpenAI's Dots announcement describes always-on agents powered by GPT-6 Astra. They have cloud computers, can use a browser, connect to more than 4,000 apps through plugins, and work on goals over time. Users can message or call them through ChatGPT, with Slack and Teams also described as communication channels.
This is a persistent-assistant architecture. Its value is task continuity: a user gives the agent context and goals, then reviews work or decisions rather than supplying every next instruction. The announcement includes examples such as investigating software feedback, updating research analysis, and preparing content for approval.
The cartoon-like identity shown around Dots is not equivalent to a streaming avatar renderer. The launch article does not document an endpoint that accepts a developer's audio stream and returns a lip-synced video participant. Calling it an “avatar” in news coverage does not make it a substitute for a real-time avatar API.
OpenAI says rollout begins across Pro, Business Premium, and Enterprise plans in eligible markets. Those access details matter if you want a personal or organizational assistant. They do not answer whether your product can stream its own TTS into a custom face; that is a different capability and is not documented in the reviewed announcement.
Choose Dots for the assistant workflow it actually offers. Do not evaluate it as though its cloud computer, connected apps, and background tasks were features that a narrowly scoped audio-to-video renderer ought to provide. FaceMode does not replace that agent harness or perform those tasks on your behalf.
Gemini Live Avatar: a native conversational video capability
Google's announcement presents Gemini 3.8 Live with Live Avatar as near-real-time visual presence natively coupled to Gemini's live dialogue model. It includes lip sync, expressions, multimodal conversation, and asynchronous tool execution while dialogue continues.
Google says the feature can transition across 97 languages, adapting speech-to-speech synchronization and expressions. It also describes a preset avatar library and custom avatars created from reference images, with custom creation available through enterprise allowlisting. These are Google's published capabilities, not independently tested quality claims from this article.
The announcement says Live Avatar is available in Gemini Enterprise and links developers to API documentation. It also describes SynthID watermarking for audio and video output. For organizations already evaluating Gemini's live conversational stack, those capabilities belong in the procurement and technical review.
What the announcement does not establish is a drop-in renderer for an arbitrary external TTS stream. Its architecture is described as paired with Gemini 3.8 Live. That is narrower than saying “all your infrastructure must be Google,” which the launch post alone does not prove. Review the actual enterprise API and deployment terms before drawing conclusions about external clients, networking, or room integration.
A team adopting Gemini for conversation generation is making a different decision from a team that already has a working STT/LLM/TTS pipeline and wants only video output. The distinction is about the integration boundary, not whether either architecture is inherently better.
FaceMode: a talking avatar API for existing voice agents
FaceMode's product summary describes a focused audio-to-video layer. Your agent retains its speech recognition, language model, voice synthesis, and turn detection. FaceMode receives the audio it produces, renders a lip-synced avatar, and publishes media as a participant in a LiveKit or Daily room you own.
The integration overview covers LiveKit Agents plugins for Python and TypeScript, a Python Pipecat integration, and the direct REST/WebSocket path. The direct API is useful when your application produces audio outside those frameworks, provided you can implement the documented session lifecycle and PCM negotiation.
That does not mean every hosted voice service or conferencing application is already supported natively. Verify that your platform exposes the audio you need and that your intended output can use the documented room transport. “Can stream audio” is an input capability, not proof of an end-to-end integration with every product bearing an API label.
FaceMode documents a front-facing portrait upload workflow and five library styles: realism, semi-realism, 3D animation, anime, and flat vector. Its published latency target is approximately 960ms from audio to rendered frame. That target is not a guarantee of total conversational response time or instantaneous interruption cutoff.
The quickstart describes private-beta access, 3,600 free credits equivalent to 60 minutes, one credit per second of streamed video, ten-minute sessions, and one concurrent session per account. Paid top-ups are not open during beta. Those operating limits should be part of the decision, especially if your next step is production deployment rather than a proof of concept.
Latency claims are not directly comparable
Meta's approximately 870ms number begins at the end of the user's turn and ends at the first byte of synchronized response. FaceMode's approximately 960ms target begins with supplied audio and concerns its rendering path. Comparing the two numbers as a speed ranking would combine different timing boundaries and a reported result with a product target.
The reviewed Dots announcement is not an avatar-rendering benchmark. Google's Live Avatar announcement does not publish a directly comparable numeric latency figure. Missing numbers should remain missing, rather than being inferred from words such as “real-time” or “near-real-time.”
For an application test, record session startup, end-of-turn response onset, media synchronization, and interruption recovery separately. Use the same voice settings, client region, network conditions, and utterances where the architectures permit it. Keep cold-start observations separate from ongoing turns.
This also makes failures easier to diagnose. A delayed LLM response, slow avatar startup, and browser audio playback restriction can all look like “the avatar is slow” while requiring different fixes. A comparison article should not flatten those into one unexplained millisecond figure.
Where each product is a useful fit
- Muse Avatar: investigate it when you care about Meta's integrated consumer-facing embodiment research, shared speech-token architecture, or expressive character generation. Its research demos provide technical context, but the reviewed post is not an external integration guide.
- OpenAI Dots: evaluate it when you want a persistent assistant working across connected applications with review and permission controls. It addresses delegated work, not just a visual response to supplied speech.
- Gemini Live Avatar: evaluate it when Gemini's live dialogue stack and enterprise access fit your application. The announced multilingual synchronization, tools, and custom-avatar pathway are genuine reasons to include it in that review.
- FaceMode: test it when your voice agent already exists and its audio needs a visual layer in a documented LiveKit or Daily room. Account for private-beta limits and test media readiness, not only API responses.
FaceMode is not the only provider in the last category. Other vendors also offer bring-your-own-stack real-time avatar components. For a more direct comparison, read Synthesia Interactive Avatar vs FaceMode, which examines shared architecture as well as differences in runtimes, transport, pricing, and access.
Adding a face to an existing LiveKit voice agent
If you choose the component route, attach an avatar session inside the worker you already run. This FaceMode excerpt is adapted from the Python integration README. It assumes agent_session and a connected ctx.room already exist, and that environment credentials are configured on the server:
import os
from livekit_plugins_facemode import AvatarSession
avatar = AvatarSession(
api_key=os.environ["FACEMODE_API_KEY"],
avatar_id=os.environ["FACEMODE_AVATAR_ID"],
)
await avatar.start(agent_session, ctx.room, room={
"type": "livekit",
"url": os.environ["LIVEKIT_URL"],
"token": os.environ["LIVEKIT_TOKEN"],
})The complete flow also starts the agent with direct audio output disabled, waits for avatar media, and handles shutdown. The LiveKit avatar plugin tutorial covers those steps and the browser subscription path. Use the direct API quickstart if you are not using a framework plugin.
Keep credentials separate: the avatar worker receives a server-side publisher token, while the browser gets its own room token. A read-only viewer and a user publishing a microphone have different permissions. An API key or worker token should never be the shortcut used to make a browser demo connect.
Questions to answer before choosing
Do you need an assistant or a renderer? If the requirement is autonomous work across apps, compare agent products and their controls. If it is visual output for existing speech, compare real-time avatar APIs and transports.
Can you keep the voice pipeline you already have? FaceMode documents that path explicitly. Google's launch describes native pairing with Gemini Live. Meta's research describes shared voice tokens. Dots is a packaged assistant. Check documentation for the specific integration boundary rather than assuming the marketing term “avatar” implies the same contract.
What will you actually test? Start with a short real conversation: join, speak, interrupt, pause, resume, and end. Inspect video, audio, and session cleanup. Then check pricing and concurrency with the account you will deploy, not only the public demo.
The goal is to choose the right layer for your application. If FaceMode's component model fits, request beta access or read the FaceMode docs before wiring it into your agent.
Primary sources
[ Try it yourself ]
Give your voice agent a face.
Beta testers get 60 free minutes of avatar video. Keep your voice stack, keep your room - just add the face.