Interactive Avatar: Minimal Quickstart

The smallest useful Synthesia Interactive Avatar app: a Python LiveKit agent you can talk to, rendered as a photoreal, lip-synced avatar in your browser. Drop in three API keys and go.

How it works: the agent listens with Cartesia Ink-2 STT, thinks with OpenAI Realtime used as a text-only streaming brain, and speaks with Cartesia Sonic TTS. Both Cartesia models run through LiveKit Inference, billed to your LiveKit account, so you don't need a Cartesia key. The Synthesia plugin reroutes that speech to a hosted avatar worker, which joins the LiveKit room as a regular participant publishing lip-synced video. The frontend is a plain LiveKit client with no Synthesia-specific code.

Latency: realtime_preflight.py hides ~0.3–0.5 s per turn by starting the Realtime reply on the eager end-of-turn transcript. Look for Adopting realtime preflight speculation in the logs; disable with REALTIME_PREFLIGHT_ENABLED=false.

Prerequisites

Run it

python3 -m venv .venv && source .venv/bin/activate
pip install -r requirements.txt

cp .env.example .env   # then fill in the three keys

python agent.py dev    # terminal 1: the agent worker
python server.py       # terminal 2: frontend at http://localhost:8080

Open http://localhost:8080, click Start, and allow the microphone. You join the room, the agent connects, the avatar joins and publishes video (cold starts can take longer), and then it greets you. Say hello.

Don't use python agent.py console. It runs a mock room and the avatar will silently never appear. Always test through a real room, like this frontend.

The integration, in three lines

Everything avatar-specific in agent.py is:

from livekit.plugins import synthesia

avatar = synthesia.AvatarSession(synthesia.AvatarConfig(avatar_ids=[AVATAR_ID]))
await avatar.start(session, room=ctx.room)   # before session.start() — order matters

The agent produces speech exactly as it always does; the plugin intercepts session.output.audio and the avatar lip-syncs it. The plugin ships on PyPI as livekit-plugins-synthesia; see the LiveKit plugin docs for its full API. That means you can swap any STT/LLM/TTS providers in AgentSession and the avatar lines never change.

Make it yours

All in agent.py:

  • Personality: edit INSTRUCTIONS.
  • Voice: set CARTESIA_VOICE_ID to any Cartesia library voice (a free account is enough to browse; usage bills via LiveKit Inference).
  • Custom / cloned voices: LiveKit Inference only serves Cartesia's public library, so a voice cloned in your own Cartesia account needs the direct plugin instead: get a Cartesia API key, add CARTESIA_API_KEY= to .env, change requirements.txt to livekit-agents[cartesia,openai,silero], and swap the tts= line to cartesia.TTS(model="sonic-3.6", voice=CARTESIA_VOICE_ID) (adding cartesia to the livekit.plugins import). TTS then bills to your Cartesia account rather than LiveKit.
  • Avatar: set AVATAR_ID to the Interactive ID of a synthetic or Personal avatar your workspace can use. In Synthesia Studio, open the avatar's ••• menu and choose Copy Interactive ID. Stock actor-based avatars such as Ryan or Ada can't be used as interactive avatars on any plan.
  • Models: swap the stt= / llm= / tts= lines for any LiveKit-supported provider. Preflight is a no-op for non-Realtime models.

server.py mints room tokens with sync_streams=True (keeps the avatar's audio and video in sync in the browser) and a RoomAgentDispatch matching the worker's agent_name. Dispatch is explicit, so without it the agent never joins, and without agent_name the worker would join every room in your LiveKit project. Keep both when you build your own token endpoint, and keep SYNTHESIA_API_KEY server-side: it's a workspace-bound secret that must never reach frontend code.

Troubleshooting

SymptomLikely cause / fix
SynthesiaError with type AUTHAPI key invalid or expired.
SynthesiaError with type FEATURE_NOT_IN_PLANYour Synthesia workspace's plan doesn't include Interactive Avatars.
SynthesiaError with type LIVEKIT_CREDENTIALS_REJECTEDLIVEKIT_URL and LIVEKIT_API_KEY/LIVEKIT_API_SECRET aren't from the same LiveKit project.
SynthesiaError with type UNKNOWN_AVATARAVATAR_ID isn't a usable Interactive ID: it's a stock actor-based avatar, it hasn't finished converting to an interactive avatar, or it isn't in your plan.
SynthesiaError with type QUOTA_EXCEEDEDMinute or concurrent-session cap hit.
SynthesiaError with type TIMEOUTCold start took too long. Retry; the agent already waits 60 s and retries transient errors.
Avatar never appears, no errorYou're in console mode, avatar.start() ran after session.start(), or the token lacks the RoomAgentDispatch room config.
Avatar joins but doesn't lip-syncSomething reassigned session.output.audio after the avatar attached.

Source for this recipe: interactive-avatar-quickstarts/minimal.