BlogOctober 2, 202610 min read

LiveKit Avatar Plugin: Add a Face to Your Voice Agent

Use the FaceMode LiveKit avatar plugin to add lip-synced video to an existing voice agent. Python and TypeScript setup, room tokens, and pitfalls.

Use the FaceMode LiveKit avatar plugin to turn your existing voice agent's speech into lip-synced video in the room you already run. This guide covers the Python and TypeScript attachment paths, then shows how a browser subscribes to the avatar's media.

Your STT, LLM, TTS, and conversation logic stay in your agent. The avatar layer receives its audio and publishes synchronized audio/video as another room participant, following the model described in LiveKit's virtual avatar overview.

The examples below are integration excerpts validated against FaceMode's plugin READMEs and integration docs. They assume you already have a working LiveKit agent; they are not standalone worker applications.

What you need

Four things, all server-side:

  1. A FaceMode account. Sign-up is waitlist-gated during the private beta at console.facemode.io/signup. Approved testers get 3,600 free credits - 60 minutes of avatar video, no card required.
  2. A FaceMode API key (fm_...) from the dashboard.
  3. An avatar ID for the face that should speak. Pick from the library across five art styles, or upload one front-facing portrait and make it your own.
  4. A LiveKit room you own, plus a publisher token. The token must grant the FaceMode worker room_join, can_publish, and can_subscribe on your room.

The access and beta limits come from the FaceMode quickstart. Use a separate participant identity for the avatar worker and for each viewer. Reusing a token identity for multiple connections is not a way to give every participant the same permissions.

Set the credentials as environment variables on your backend:

export FACEMODE_API_KEY="fm_your_key_here"
export FACEMODE_AVATAR_ID="your_avatar_id"
export LIVEKIT_URL="wss://your-project.livekit.cloud"
export LIVEKIT_TOKEN="<publisher-token>"

Keep every token in this guide on the server. The browser only ever receives the room URL and its own subscribe-only viewer token - never the API key, never the publisher token.

The 30-second mental model

FaceMode is not another agent framework. It is one job done well: turn the audio your agent already produces into a lip-synced avatar, published into the room you already run.

Your agent (STT -> LLM -> TTS)
        |
        v  (the TTS audio you already have)
    FaceMode  --  lip-synced video frames, ~960ms audio-to-frame target
        |
        v
  Your LiveKit room  --  the avatar joins as an ordinary participant
        |
        v
  Your frontend  --  subscribes with the livekit-client SDK you already ship

Nothing sits between your app and your room. FaceMode joins as a participant, publishes its video track, and leaves when the session ends. If you removed FaceMode tomorrow, your agent would keep working exactly as before.

Step 1A: Python LiveKit avatar integration

If your agent runs on LiveKit Agents for Python, install the plugin:

pip install livekit-plugins-facemode

The FaceMode Python integration guide requires Python 3.13+ and LiveKit Agents 1.7+. In your existing asynchronous worker entrypoint, ctx is its job context and agent_session is the AgentSession you have already configured with speech and model providers. The attachment excerpt is:

import os
 
from livekit.agents import Agent, AgentSession, room_io
from livekit_plugins_facemode import AvatarSession
 
avatar = AvatarSession(
    api_key=os.environ["FACEMODE_API_KEY"],
    avatar_id=os.environ.get("FACEMODE_AVATAR_ID", ""),
)
 
await ctx.connect()
await avatar.start(
    agent_session,
    ctx.room,
    room={
        "type": "livekit",
        "url": os.environ["LIVEKIT_URL"],
        "token": os.environ["LIVEKIT_TOKEN"],
    },
)
await agent_session.start(
    agent=Agent(instructions="You are a helpful assistant."),
    room=ctx.room,
    room_options=room_io.RoomOptions(audio_output=False),
)
await avatar.wait_for_join()

Three details worth noticing:

  • audio_output=False. The avatar publishes synchronized audio and video. Disabling the agent's separate audio path avoids duplicate playback; it is not a guarantee against every possible synchronization issue.
  • The room object. You hand the plugin the same LiveKit URL and publisher token you already manage. FaceMode never creates or replaces your room.
  • wait_for_join(). A one-line confirmation that the avatar participant is actually in the room before you consider the session live.

Success means wait_for_join() resolves and the avatar's video track is available in the room, not merely that the REST request succeeded. The Python integration docs distinguish session creation, WebSocket negotiation, and media readiness.

Under the hood, avatar.start() creates a FaceMode session over REST, opens the worker WebSocket, negotiates the audio format, and forwards TTS frames as PCM. After a started session loses its transport, the plugin attempts up to two reconnects with fresh one-time tokens. Exhaustion surfaces an error; reconnect is not a guarantee that every outage is invisible.

Optionally, pin the upstream input provider so the backend knows which TTS is driving the audio:

avatar = AvatarSession(
    api_key=os.environ["FACEMODE_API_KEY"],
    avatar_id=os.environ["FACEMODE_AVATAR_ID"],
    input_provider="elevenlabs",  # deepgram, gemini, gnani, elevenlabs, openai, cartesia, sarvam, custom
)

Any TTS works. If it can stream audio, it can drive a face.

Step 1B: TypeScript LiveKit avatar integration

Use @facemode/agents-plugin-facemode with Node 20+ and @livekit/agents 1.x, as specified in the TypeScript integration guide. This is an alternative to the Python route, not a second worker you need to run:

npm install @facemode/agents-plugin-facemode

Use this inside your worker entrypoint with an already-configured agentSession. The session-start portion is adapted from the TypeScript docs so the excerpt includes the audio-output setting as well as the avatar attachment.

import { voice } from "@livekit/agents";
import { AvatarSession } from "@facemode/agents-plugin-facemode";
 
const avatar = new AvatarSession({
  apiKey: process.env.FACEMODE_API_KEY!,
  avatarId: process.env.FACEMODE_AVATAR_ID,
  inputProvider: "deepgram", // optional
});
 
await ctx.connect();
await avatar.start(agentSession, ctx.room, {
  room: {
    type: "livekit",
    url: process.env.LIVEKIT_URL!,
    token: process.env.LIVEKIT_TOKEN!,
  },
});
await agentSession.start({
  agent: new voice.Agent({ instructions: "You are a helpful voice assistant." }),
  room: ctx.room,
  outputOptions: { audioEnabled: false },
});
await avatar.waitForJoin();

The plugin replaces the agent session's audio tail: TTS frames flow to FaceMode as binary pcm_s16le, interruptions send a cancel_utterance control, and reconnects are handled for you. Sample rates from 8kHz to 48kHz are accepted, with 8, 11.025, 12, 16, 22.05, 24, 32, 44.1, and 48kHz as the preferred set covering common TTS outputs.

Step 2: render the LiveKit avatar in the browser

This is the part developers expect to be hard and is actually the easiest. The avatar is just a participant in your room, so your existing LiveKit client code subscribes to it like any other video track.

Mint a separate, subscribe-only viewer token for a read-only viewer and pass it down with the room URL. In the browser, register the subscription listener before connecting and insert the elements returned by track.attach() into the DOM. This excerpt is adapted from LiveKit's track subscription guide:

import { Room, RoomEvent } from "livekit-client";
 
const room = new Room();
const media = document.createElement("div");
document.body.appendChild(media);
room.on(RoomEvent.TrackSubscribed, (track) => {
  if (track.kind === "video" || track.kind === "audio") {
    media.appendChild(track.attach()); // render the avatar
  }
});
await room.connect(roomUrl, viewerToken); // your own viewer token

Success is visible video plus audible speech, not just a connected room. In a multi-participant app, filter tracks by the avatar participant identity instead of rendering every subscribed participant. Your real app should also detach media on unsubscribe and disconnect the room when its viewer component unmounts.

This is a read-only viewer example. A user who talks to the agent needs microphone publication in your existing voice UI and appropriate grants on their own token. Do not give them the worker's publisher token. Browser autoplay policies can require a user gesture to enable audio; use LiveKit's playback controls in that interaction rather than interpreting silent playback as an avatar failure.

Interruptions, silence, and the in-between moments

Real conversations are messy, and the avatar behaves accordingly:

  • Barge-in. The plugin docs say an agent interruption sends cancel_utterance. That control is not evidence of zero-delay playback cancellation; test the visible and audible result in your actual client.
  • Silence between utterances. The FaceMode product summary describes crossfading into an idle loop between utterances. Check that this remains visible when your voice agent is waiting on a long-running tool call.
  • Startup budgets. The Python docs allow up to 240 seconds for readiness polling and WebSocket opening, but use a separate 15-second protocol acknowledgement budget and a 30-second default media-join timeout. Those are timeout settings, not a measured cold-start duration or a guaranteed successful startup.

A short pitfalls checklist

Before your first session, check the five things that bite most often:

  1. Token grants. The publisher token must allow join, publish, and subscribe. If it pins a fixed participant identity, make the room policy allow that identity.
  2. Secrets stay server-side. API key and publisher token never reach the browser. Only the room URL and a subscribe-only viewer token do.
  3. audio_output=False on the agent session, or you will hear double.
  4. Declared sample rate must match the actual PCM. The plugin locks the format at negotiation and rejects a mismatch mid-stream.
  5. Beta limits. Sessions auto-end after 10 minutes, one concurrent session per account, and a session needs at least 60 credits of remaining balance to start.

Step 3: validate media, failure recovery, and shutdown

A LiveKit avatar integration crosses several boundaries: your voice agent, the FaceMode API, its ingestion WebSocket, the avatar participant, and browser playback. Diagnose the boundary that failed rather than treating all five as one connection.

Start with a short utterance whose beginning and end are easy to recognize. Confirm that only one audio path is audible and that the visible avatar is the participant you intended to subscribe to. Then interrupt a longer sentence. Observe whether the agent stops generating speech, whether cancellation reaches the avatar, and whether already-buffered client media continues playing. This separates agent behavior from transport and playback behavior.

Next, test a pause with no speech. Idle animation should not be mistaken for proof that the agent is ready to answer, and a connected room should not be mistaken for proof that the ingestion socket is healthy. Keep the voice agent's own state indicator alongside the video so users can distinguish listening from waiting.

The Python lifecycle guide documents await avatar.stop() or aclose() for shutdown. The TypeScript guide documents the equivalent methods. Wire these into your existing call cleanup path, including failures; an abandoned browser tab should not be your only cleanup mechanism. The canonical ended acknowledgement confirms the protocol message, not that every backend cleanup action has finished.

If you test reconnect behavior, do it with disposable development sessions. The plugin only attempts recovery after negotiation has completed; intentional shutdown and fatal protocol errors do not trigger the same recovery path. Do not write code that retries a consumed one-time WebSocket token or replays previously sent audio.

What the latency target does and does not mean

FaceMode publishes an approximately 960ms audio-to-rendered-frame target. Your total voice-agent response time also includes turn detection, transcription, LLM generation, TTS, transport, and client playback. A sub-second target for the avatar layer does not imply the whole conversation responds within one second.

When comparing providers, measure the same utterance from the same boundary. Keep session startup separate from steady-state response latency. Also keep interruption recovery separate: a pipeline can begin speaking quickly and still feel poor if it cannot stop naturally when the user cuts in.

Record observations from a real room rather than using a console-only voice test. Console mode can validate conversation logic without exercising the video subscription and browser playback paths. These are suggested acceptance checks, not benchmarks we ran for this article.

Choosing a plugin or the direct talking avatar API

Use the plugin when LiveKit Agents already owns your speech pipeline. It captures the framework's audio frames and lifecycle events, so you do not have to reproduce its WebSocket protocol handling. Use the direct talking avatar API when your application produces audio outside that framework and you can manage session creation, PCM negotiation, and shutdown yourself.

That choice does not automatically prove compatibility with every hosted voice-agent platform. A provider that cannot expose its generated audio, or whose calls are phone-only, may need an additional transport or product integration. FaceMode's documented room output is LiveKit or Daily, not a claim of native support for every conferencing system.

For a comparison with another bring-your-own-stack provider, see Synthesia Interactive Avatar vs FaceMode. Both preserve existing conversation models; their documented runtimes, transport paths, and access limits are the useful differences.

LiveKit avatar plugin: short version

Inside your existing asynchronous Python worker, connect the room, attach the configured avatar, start the configured agent session with direct audio disabled, and wait for video. This recap uses the room binding and objects configured in Step 1A; it is not a standalone script:

await ctx.connect()
await avatar.start(agent_session, ctx.room, room=room)
await agent_session.start(
    agent=Agent(instructions="You are a helpful assistant."),
    room=ctx.room,
    room_options=room_io.RoomOptions(audio_output=False),
)
await avatar.wait_for_join()

Here room is the server-side dictionary containing type, url, and the worker token from the earlier example. Stop the avatar through your call's cleanup path once the conversation ends.

From zero to a talking face

The whole flow is: install the plugin, hand it your room, keep your agent code. If you want the raw HTTP and WebSocket protocol instead - any language that can hold a socket open works - the direct API quickstart walks it end to end.

For the broader category distinction, read Muse Avatar, OpenAI Dots, Gemini Live Avatar vs FaceMode. It separates platform agent experiences from avatar components you integrate into a product.

Ready to see your agent talk? Join the beta - 60 free minutes, no card - or skim the docs first.

Back to all posts

[ Try it yourself ]

Give your voice agent a face.

Beta testers get 60 free minutes of avatar video. Keep your voice stack, keep your room - just add the face.