Interactive Avatar: Integration Overview

This page shows how the pieces fit together in a Bring-Your-Own (BYO) application: the architecture, how data flows, who is responsible for what, and what to plan for in production. For the vocabulary, see Concepts.

Architecture

You bring a Python LiveKit voice agent and attach a single component, synthesia.AvatarSession, to it. That component does three things:

  1. Authenticates with Synthesia using your workspace API key.
  2. Dispatches a hosted avatar worker into your LiveKit room, where it joins as a normal participant with the identity synthesia-avatar-agent.
  3. Re-routes your agent's audio output to the worker over a LiveKit data stream, so the avatar drives lip-sync from whatever speech your agent produces.

No GPU or video code runs in your agent process. Rendering is fully hosted.

[your Python LiveKit Agent + synthesia.AvatarSession]
        │  1. authenticate (API key)
        │  2. dispatch worker
        ▼
[Synthesia API gateway] ──▶ [Avatar worker: GPU render + lip-sync]
        │                                   │
        │  3. worker joins your room as 'synthesia-avatar-agent'
        │  4. your agent's speech audio ──▶ worker (LiveKit data stream)
        │  5. lip-synced video + audio ──▶ published to the room
        ▼
[LiveKit room]  ◀──────────────▶  [end user: any web/mobile LiveKit client]

The avatar is a regular LiveKit participant, so any web or mobile LiveKit client renders it with no Synthesia-specific frontend code.

Data flow

A single turn through the system:

  1. The end user speaks in your LiveKit room (or your realtime model listens directly to their audio).
  2. Your agent does what it always does: runs your realtime model or your STT → LLM → TTS pipeline, and produces speech audio.
  3. The plugin intercepts that audio (it replaces session.output.audio) and forwards it to the avatar worker over a LiveKit data stream, instead of publishing it straight to the room.
  4. The avatar worker renders lip-synced video and publishes both video and audio back into the room as the synthesia-avatar-agent participant.
  5. The end user's client subscribes to that participant like any other and displays the talking avatar.

The audio the user hears comes from the avatar worker, not directly from your agent, which is why the avatar and the voice stay in sync.

Who owns what

LayerResponsibilityOwner
Conversation logic, knowledge sources, workflowDeciding what the avatar saysYou
Speech generationSTT + LLM + TTS, or a realtime modelYou (BYO)
The LiveKit room and transportProvisioning the LiveKit Cloud project and roomYou / LiveKit
Room token for the avatarMinted in-process by the plugin using your LiveKit secretYou (via the plugin)
The frontendAny LiveKit client; renders the avatar as a normal participantYou
Authentication + worker dispatchVerifying the workspace key, launching the workerSynthesia (via the plugin)
GPU render + lip-syncTurning your audio into avatar videoSynthesia
Publishing the avatarJoining the room and publishing video + audio tracksSynthesia

In short, you own the conversation and the room, and Synthesia owns the rendered avatar. This suits teams that already have an LLM, agent, or conversational logic and want Synthesia only for the avatar layer.

Audio routing

When you attach the avatar, the plugin replaces session.output.audio, so the avatar lip-syncs to whatever your agent says without any extra code on your side. Two things follow from this:

  • Attach the avatar before session.start(). Attaching afterwards won't work.
  • Don't reassign session.output.audio after attaching the avatar, or lip-sync will break.

Because the re-routing works on session.output.audio, the avatar attaches the same way whether you run a realtime model or a component pipeline.

Choosing an integration path

The avatar attaches the same way regardless of how your agent generates speech, so pick whichever speech stack suits you:

Voice-to-voice realtimeComponent pipeline
What it isA single realtime model handles VAD, STT, LLM, and TTS internally (e.g. OpenAI Realtime)A classic VAD + STT + LLM + TTS stack (e.g. Silero VAD, Deepgram STT, an OpenAI LLM, Cartesia TTS)
SupportValidated pathIdentical wiring; flag any issues to your Synthesia contact
You controlFewer moving parts; voice/behavior set by the realtime modelEach stage independently: swap STT/LLM/TTS providers freely
Attach codeawait avatar.start(session, room=ctx.room) before session.start()Same

Because TTS is BYO, Synthesia voices are not available through this API. Pick any TTS provider compatible with your LiveKit pipeline, and match the voice to your avatar's persona so audio and face don't clash.

See the Quickstart guides for complete, runnable code.

Production considerations

  • Keep secrets server-side. The Synthesia API key, LiveKit API key/secret, and model provider keys belong in your agent's environment or a secret manager, never in frontend code. The plugin mints the avatar's room token in-process, so your LiveKit secret never leaves your agent.
  • Deploy with the dependency declared. When deploying to LiveKit Cloud, install the plugin with uv add "livekit-plugins-synthesia~=1.8" so it's recorded in pyproject.toml; an environment-only install won't be present in the Cloud build. (See the Quickstart guides.) Your agent worker is the only thing you deploy; rendering is always Synthesia-hosted.
  • Handle cold starts. avatar.start() waits up to join_timeout (default 30s) for the worker to join; raise it if the worker cold-starts slowly. A timeout surfaces as SynthesiaError with type=ErrorType.TIMEOUT.
  • Handle mid-session drops. The plugin logs an unexpected drop and tears the session down; there are no events to subscribe to, and there is no built-in reconnect. Treat recovery as a new start() call.
  • Distinguish retryable from terminal failures. Every failure raises a single SynthesiaError. Branch on its retryable attribute rather than inferring from the failure type, and honor retry_after when it's set. Full error-type list in the Plugin reference; HTTP-level codes in the Errors reference.
  • Video quality is set on the client. The hosted worker fixes the avatar's resolution and the plugin has no quality setting, so to keep it crisp, have your client subscribe at full resolution rather than downscaling to a small tile.
  • Concurrency is capped. Interactive Avatar sessions count against your plan's concurrent-session limit: 1 for Freemium, 100 for all paid plans. A separate minute-based cap also applies; its specific threshold isn't published. Exceeding the concurrency cap raises SynthesiaError with type=ErrorType.CONCURRENCY_LIMIT; an exhausted quota raises type=ErrorType.QUOTA_EXCEEDED. Called directly, POST /api/interactive-avatars/sessions returns 429 concurrency_limit or 402 quota_exceeded respectively.