Build real-time, conversational avatar experiences on your own stack. You bring the LLM and speech-to-text; Synthesia renders the avatar, with optional built-in text-to-speech if you don't want to bring your own.
You'll need Python 3.10+ and a Synthesia API key with Interactive Avatar access to get started.
Get oriented
Start with the Minimal quickstart — the smallest working real-time avatar session.
Building something more specific?
Or integrate directly
Not using the Python plugin? Call the API yourself — you still connect to a LiveKit room, but you mint the token and manage the session yourself instead of the plugin doing it for you.