This page shows how the pieces fit together in a Bring-Your-Own (BYO) application: the architecture, how data flows, who is responsible for what, and what to plan for in production. For the vocabulary, see Concepts.
Architecture
You bring a Python LiveKit voice agent and attach a single component, synthesia.AvatarSession, to it. That component does three things:
- Authenticates with Synthesia using your workspace API key.
- Dispatches a hosted avatar worker into your LiveKit room, where it joins as a normal participant with the identity
synthesia-avatar-agent. - Re-routes your agent's audio output to the worker over a LiveKit data stream, so the avatar drives lip-sync from whatever speech your agent produces.
No GPU or video code runs in your agent process. Rendering is fully hosted.
[your Python LiveKit Agent + synthesia.AvatarSession]
│ 1. authenticate (API key)
│ 2. dispatch worker
▼
[Synthesia API gateway] ──▶ [Avatar worker: GPU render + lip-sync]
│ │
│ 3. worker joins your room as 'synthesia-avatar-agent'
│ 4. your agent's speech audio ──▶ worker (LiveKit data stream)
│ 5. lip-synced video + audio ──▶ published to the room
▼
[LiveKit room] ◀──────────────▶ [end user: any web/mobile LiveKit client]The avatar is a regular LiveKit participant, so any web or mobile LiveKit client renders it with no Synthesia-specific frontend code.
Data flow
A single turn through the system:
- The end user speaks in your LiveKit room (or your realtime model listens directly to their audio).
- Your agent does what it always does: runs your realtime model or your STT → LLM → TTS pipeline, and produces speech audio.
- The plugin intercepts that audio (it replaces
session.output.audio) and forwards it to the avatar worker over a LiveKit data stream, instead of publishing it straight to the room. - The avatar worker renders lip-synced video and publishes both video and audio back into the room as the
synthesia-avatar-agentparticipant. - The end user's client subscribes to that participant like any other and displays the talking avatar.
The audio the user hears comes from the avatar worker, not directly from your agent, which is why the avatar and the voice stay in sync.
Who owns what
| Layer | Responsibility | Owner |
|---|---|---|
| Conversation logic, knowledge sources, workflow | Deciding what the avatar says | You |
| Speech generation | STT + LLM + TTS, or a realtime model | You (BYO) |
| The LiveKit room and transport | Provisioning the LiveKit Cloud project and room | You / LiveKit |
| Room token for the avatar | Minted in-process by the plugin using your LiveKit secret | You (via the plugin) |
| The frontend | Any LiveKit client; renders the avatar as a normal participant | You |
| Authentication + worker dispatch | Verifying the workspace key, launching the worker | Synthesia (via the plugin) |
| GPU render + lip-sync | Turning your audio into avatar video | Synthesia |
| Publishing the avatar | Joining the room and publishing video + audio tracks | Synthesia |
In short, you own the conversation and the room, and Synthesia owns the rendered avatar. This suits teams that already have an LLM, agent, or conversational logic and want Synthesia only for the avatar layer.
Audio routing
When you attach the avatar, the plugin replaces session.output.audio, so the avatar lip-syncs to whatever your agent says without any extra code on your side. Two things follow from this:
- Attach the avatar before
session.start(). Attaching afterwards won't work. - Don't reassign
session.output.audioafter attaching the avatar, or lip-sync will break.
Because the re-routing works on session.output.audio, the avatar attaches the same way whether you run a realtime model or a component pipeline.
Choosing an integration path
The avatar attaches the same way regardless of how your agent generates speech, so pick whichever speech stack suits you:
| Voice-to-voice realtime | Component pipeline | |
|---|---|---|
| What it is | A single realtime model handles VAD, STT, LLM, and TTS internally (e.g. OpenAI Realtime) | A classic VAD + STT + LLM + TTS stack (e.g. Silero VAD, Deepgram STT, an OpenAI LLM, Cartesia TTS) |
| Support | Validated path | Identical wiring; flag any issues to your Synthesia contact |
| You control | Fewer moving parts; voice/behavior set by the realtime model | Each stage independently: swap STT/LLM/TTS providers freely |
| Attach code | await avatar.start(session, room=ctx.room) before session.start() | Same |
Because TTS is BYO, Synthesia voices are not available through this API. Pick any TTS provider compatible with your LiveKit pipeline, and match the voice to your avatar's persona so audio and face don't clash.
See the Quickstart guides for complete, runnable code.
Production considerations
- Keep secrets server-side. The Synthesia API key, LiveKit API key/secret, and model provider keys belong in your agent's environment or a secret manager, never in frontend code. The plugin mints the avatar's room token in-process, so your LiveKit secret never leaves your agent.
- Deploy with the dependency declared. When deploying to LiveKit Cloud, install the plugin with
uv add "livekit-plugins-synthesia~=1.8"so it's recorded inpyproject.toml; an environment-only install won't be present in the Cloud build. (See the Quickstart guides.) Your agent worker is the only thing you deploy; rendering is always Synthesia-hosted. - Handle cold starts.
avatar.start()waits up tojoin_timeout(default 30s) for the worker to join; raise it if the worker cold-starts slowly. A timeout surfaces asSynthesiaErrorwithtype=ErrorType.TIMEOUT. - Handle mid-session drops. The plugin logs an unexpected drop and tears the session down; there are no events to subscribe to, and there is no built-in reconnect. Treat recovery as a new
start()call. - Distinguish retryable from terminal failures. Every failure raises a single
SynthesiaError. Branch on itsretryableattribute rather than inferring from the failure type, and honorretry_afterwhen it's set. Full error-type list in the Plugin reference; HTTP-level codes in the Errors reference. - Video quality is set on the client. The hosted worker fixes the avatar's resolution and the plugin has no quality setting, so to keep it crisp, have your client subscribe at full resolution rather than downscaling to a small tile.
- Concurrency is capped. Interactive Avatar sessions count against your plan's concurrent-session limit: 1 for Freemium, 100 for all paid plans. A separate minute-based cap also applies; its specific threshold isn't published. Exceeding the concurrency cap raises
SynthesiaErrorwithtype=ErrorType.CONCURRENCY_LIMIT; an exhausted quota raisestype=ErrorType.QUOTA_EXCEEDED. Called directly,POST /api/interactive-avatars/sessionsreturns429 concurrency_limitor402 quota_exceededrespectively.