This page builds the mental model for the Interactive Avatar API and defines the vocabulary the rest of the docs use.
Scope of this page: this describes the LiveKit plugin path. If you're calling the REST endpoints directly, the primitives are the same but you mint the room token yourself — see Start an interactive avatar session.
The mental model
An Interactive Avatar isn't a video you generate and download. It's a live participant that Synthesia renders in real time and places into a LiveKit room you already control.
The integration is a plugin you attach to your LiveKit agent. Your agent produces speech the way it always does; the plugin routes that audio to Synthesia's hosted avatar worker, which generates the matching video and broadcasts lip-synced video and audio back to every participant in the room.
You own the conversation; Synthesia owns the rendered avatar. No GPU runs in your agent process; rendering happens on the hosted worker. For the full client/agent/LiveKit/Synthesia responsibility split, see the Overview.
Core primitives
Your LiveKit agent is the conversational program you run (Python 3.10+, livekit-agents >= 1.8.2). It orchestrates your stack and produces the audio the avatar speaks. Node.js agents are not supported.
The plugin (synthesia) is the Synthesia component. Install it with pip:
pip install "livekit-plugins-synthesia~=1.8"Import it as from livekit.plugins import synthesia. It captures your agent's synthesized audio and forwards it to the avatar worker.
AvatarSession is the object you attach to your agent. It authenticates, dispatches the worker, and wires audio. You configure it with an AvatarConfig(avatar_ids=[...]). It needs an API key, passed as api_key or read from the SYNTHESIA_API_KEY environment variable, and you can optionally override api_url or join_timeout.
AvatarConfig is the configuration you pass to AvatarSession. It carries avatar_ids: one to five gallery ids of avatars available to your workspace, each prefixed av_…. The first id is rendered; the rest are precomputed so swap_avatar() can switch to them mid-session. avatar_ids must be a list, even for one avatar. Passing a bare string raises ValueError.
Finding your avatar ID
Only synthetic and personal avatars work. Stock actor-based avatars (for example Ryan or Ada) can't be used as interactive avatars.
For an eligible avatar your plan gives you access to, copy its Interactive ID from the Avatars page: open the ••• menu and select Copy Interactive ID.
A repeat request for the same source returns the existing interactive avatar rather than starting a new conversion, so it's safe to call again if you're not sure whether it's been converted already. Trying to convert a stock avatar, or one outside your plan, returns an error rather than an interactive avatar.
Testing this in the Try It panel? Select Request Example from the dropdown above the cURL preview before hitting Try It. The default view only pre-fills
regenerate; it omits the requiredsourceIdfield, which can look like a missing parameter. Request Example shows the full request body,sourceIdincluded.
The hosted avatar worker is Synthesia's GPU render service. It turns your agent's audio into a lip-synced avatar and joins your room as a participant under a default identity you can configure. The Plugin reference has the default, and the constraint that it must be unique per concurrent avatar in a room.
The LiveKit room is the real-time session your user and agent are already connected to. The worker joins it too, and all audio and video is exchanged there.
Streams. Your agent's speech reaches the worker over a LiveKit data stream; the rendered avatar is published back into the room as normal video and audio tracks your client subscribes to.
Token. The avatar worker authenticates to your room with a LiveKit room token. On the plugin path, the plugin mints this token inside your process, exactly like any other participant, and your LiveKit secret never leaves your process. On the REST path you mint it yourself and pass it as livekitToken. There is no separate Synthesia session token.
Session. One live conversation, corresponding to one AvatarSession in one room. Concurrency limits are counted in sessions.
Events. The plugin emits none. A session ends either cleanly, when the room disconnects, or unexpectedly, when the avatar stops publishing mid-session. Both are logged, not raised as events you subscribe to.
Turn-taking. Listening, speaking, and interruption are handled by your LiveKit
AgentSessionand your model. The plugin has no turn-taking surface of its own.
The session lifecycle
An AvatarSession moves through three phases:
- Start. You attach the avatar with
await avatar.start(...). It authenticates to Synthesia and dispatches the worker, which joins the room as a participant, and audio routing is wired up.start()returns once the avatar has joined and published its video track. - Active. Your agent's speech is routed to the worker, and the avatar renders and speaks in the room.
- End. The room disconnects cleanly, or the avatar drops unexpectedly. Shutdown restores normal audio routing.
Two mistakes are easy to make here, and neither one raises an error:
- Attaching the avatar after
session.start(). It doesn't work, and the avatar simply never appears. - Reassigning
session.output.audioafter attaching. The plugin owns that routing, and overriding it breaks lip-sync while leaving the avatar visible.
See the Overview and the Quickstart guides for full integration detail.
Testing your integration
Run your agent with dev (or connect --room <name>) against a real LiveKit room.
Two different things are called "Console," and only one of them works:
| What it is | Use it? | |
|---|---|---|
python agent.py console | A terminal run mode that uses a mock room | No. The avatar silently never appears, and no error is raised. |
| LiveKit Cloud Agent Console | A real hosted room in the LiveKit dashboard (Agents → Console) | Yes. The fastest way to test an avatar agent. |
Stuck?The Claude Code / Cursor skill can scan your codebase and suggest fixes.
Glossary
- Agent: your Python LiveKit program (STT + LLM + TTS orchestration, or a realtime model).
- Plugin (
synthesia): the Synthesia component, installed viapip install "livekit-plugins-synthesia~=1.8". AvatarSession: the object that authenticates, dispatches the worker, and wires audio.AvatarConfig: configuration forAvatarSession; carriesavatar_ids(one to five gallery ids).- Avatar worker: Synthesia's hosted GPU render service; joins your room as a participant.
- Room: the LiveKit real-time session shared by your user, your agent, and the avatar.
- Participant: any member of a room; the rendered avatar joins as one.
- Session: one live conversation instance; the unit of concurrency.
- Token: the LiveKit room token used to authenticate the avatar worker to your room.
- BYO: Bring Your Own STT / LLM / TTS.