Interactive Avatar: RAG Quickstart

A Synthesia Interactive Avatar that answers spoken questions from a knowledge base. It speaks only from retrieved source material, never from the model's training knowledge. Built on the minimal quickstart: a Python LiveKit agent rendered as a photoreal, lip-synced avatar in your browser.

Out of the box it retrieves from the public Wikipedia API (KB_SOURCE=wikipedia), which needs no key and covers any topic, so you can hear grounding work in minutes with no cloud setup. When you're ready to use your own private corpus, switch to KB_SOURCE=bedrock and point it at an AWS Bedrock knowledge base. The worked example below grounds an Aristotle avatar in his own works; see Grounding your own corpus with Bedrock.

Why grounding matters

An LLM will happily produce plausible but invented facts. For a photoreal avatar, especially one with a real person's likeness, that means the avatar stating something the person never said, which is a reputational and potentially legal risk. So the knowledge base has to be the single source of truth. Substance comes only from retrieved, approved content, and the model's job is to phrase it. Applying documented principles to a new question is fine; inventing specific facts or opinions about people or things not in the corpus is not, even in character.

RAG doesn't train the model. Each answer is assembled live from freshly retrieved passages, so editing the knowledge base changes behaviour immediately, and nothing is retained between sessions.

This is prompt-based grounding, so it's statistical rather than a hard guarantee. For real-person likenesses, add a post-generation check that verifies each claim against the retrieved passages before it's spoken.

Prerequisites

Run it

python3 -m venv .venv && source .venv/bin/activate
pip install -r requirements.txt

cp .env.example .env   # fill in the three keys — no AWS needed for the default

python agent.py dev    # terminal 1: the agent worker
python server.py       # terminal 2: frontend at http://localhost:8080

Open http://localhost:8080, click Start, and allow the microphone. You join the room, the agent connects, the avatar joins and publishes video (cold starts can take longer), and then it greets you. Ask it anything; the default Wikipedia source covers any topic.

The avatar does not announce that it's AI by default. How (and whether) to declare that is left to you. For real deployments, especially avatars with a real person's likeness, add a disclosure (see the greeting in agent.py).

Don't use python agent.py console. It runs a mock room and the avatar will silently never appear. Always test through a real room, like this frontend.

How a turn flows

The livekit-agents framework's AgentSession run loop (livekit/agents/voice/agent_activity.py) owns the pipeline order. This quickstart supplies the components, one retrieval hook, and one ordering flag. A full turn:

mic audio ─► STT (Cartesia Ink-2, streaming)
          ─► turn detection  ── user stops speaking
          ─► on_user_turn_completed(turn_ctx, new_message)   ← the only seam we own
                 ├─ retrieve passages from the KB (Wikipedia / Bedrock)
                 └─ inject them into turn_ctx as system context
          ─► LLM (GPT-4o) generates from the augmented context
          ─► TTS (Cartesia Sonic) speaks the reply
          ─► Synthesia avatar lip-syncs the audio

Everything RAG-specific in agent.py lives in that one hook:

class GroundedAgent(Agent):
    async def on_user_turn_completed(self, turn_ctx, new_message):
        passages = await self._retrieve(new_message.text_content)   # Wikipedia or Bedrock
        turn_ctx.add_message(role="system", content="Knowledge base context:\n" + passages)

Two things guarantee retrieval runs before generation on every turn:

  1. The hook is awaited. The framework won't start the LLM until on_user_turn_completed returns, so retrieved passages are always in context before generation.
  2. Preemptive generation is off (turn_handling in agent.py). With it on, the framework would speculatively start the LLM on the partial transcript, racing retrieval and risking an ungrounded reply.

So the sequence is STT → KB → LLM → TTS → avatar, with the KB step in the one gap the framework opens between transcription and generation. That ordering has a latency cost. On the Wikipedia path each turn runs a query rewrite (a small-model call) plus the search and page fetches before generation can start, so expect roughly a second more per turn than the minimal quickstart. Bedrock's semantic search skips the rewrite step.

_retrieve dispatches on KB_SOURCE; adding another backend is just another retriever behind the same hook. When retrieval returns nothing, the model is told to say so and decline rather than answer from general knowledge.

Seeing the tool chain (the status pill)

A small pill in the browser shows what the pipeline is doing each turn, from Listening… through Querying knowledge base… and Thinking… to Speaking…. You can watch the KB query fire before the model generates, on every turn.

The agent publishes each stage on a LiveKit data topic (pipeline-status) and the frontend renders it, with no polling or extra service. The KB and Thinking stages come from the retrieval hook, so they're correctly ordered within a turn; Speaking and Listening come from the session's own agent_state_changed events. It's also a debugging aid: if you never see "Querying knowledge base…", retrieval isn't wired up.

Grounding your own corpus with Bedrock

KB_SOURCE selects where passages come from. Both sources sit behind the same retrieval hook; only the retriever changes.

KB_SOURCECorpusSetupUse it for
wikipedia (default)Public Wikipedia, any topicNone (no key, no cloud)Trying the quickstart cold; general-knowledge demos
bedrockYour private docs in S3AWS account + a Bedrock managed KBGrounding in a controlled, private corpus (real customer PoCs)

1. Build the knowledge base

  1. Create an S3 bucket; upload your approved source docs under a material/ prefix. Prefer smaller, semantically coherent chunks (one per file) over whole documents; it keeps retrieval tight and fast.
  2. Create a Bedrock managed knowledge base whose data source is scoped to that material/ prefix only. Let it ingest/sync. Note the KB id.
  3. Check what you ingest. Mis-attributed content corrupts every answer, since the KB is the avatar's only source of truth. Strip licence text and boilerplate (e.g. Project Gutenberg headers) before ingest, or it gets retrieved as if it were the author.

The worked example is an Aristotle avatar grounded in his own works (Nicomachean Ethics, Politics, Poetics, Categories): public-domain texts, chunked one theme per file, with PERSONA_NAME=Aristotle giving first-person answers whose substance comes only from those texts.

2. Point the agent at it

The Bedrock retriever is already in the code, so switching only needs environment changes:

pip install boto3          # optional dep, only the Bedrock path needs it
# in .env
KB_SOURCE=bedrock
BEDROCK_KB_ID=your-kb-id       # required — the managed KB you just built
AWS_REGION=eu-west-1           # run the agent in the same region as the KB
AWS_ACCESS_KEY_ID=...          # or set AWS_PROFILE instead of these three
AWS_SECRET_ACCESS_KEY=...
AWS_SESSION_TOKEN=...          # only for temporary / SSO credentials

boto3 uses the standard credential chain: keys in .env, or leave them blank and set AWS_PROFILE. SSO / temporary credentials expire every few hours and need refreshing; in a deployed setup attach an IAM role instead. Restart the agent, and the persona and greeting switch automatically from Kenji to Aristotle.

Sanity-check retrieval from the CLI before touching the agent:

aws bedrock-agent-runtime retrieve \
  --knowledge-base-id "$BEDROCK_KB_ID" \
  --retrieval-query '{"text":"what is virtue?"}' \
  --retrieval-configuration '{"managedSearchConfiguration":{}}' \
  --region "$AWS_REGION"

If passages come back, you're good to go. If you get an expired-token error, refresh your AWS credentials first. Note the managedSearchConfiguration: a managed KB rejects the vectorSearchConfiguration you'd use for a custom KB, and if you pass the wrong one retrieval silently returns nothing. The agent's _retrieve_bedrock already passes the right one.

Proving it's grounded (the canary test)

A plausible answer doesn't prove retrieval fired, since the model may already know it from training. On the Bedrock path you can prove grounding by planting a fact the model couldn't know:

  1. Add a short document to your S3 source prefix with an invented, clearly fictional fact (e.g. "This is the Indigo Reference Edition, compiled by Prospero Quill."), then re-sync the data source.
  2. Ask the avatar about it. If it repeats the planted fact, the answer can only have come from the KB. Watch the agent log for the matching [KB] injected N passages line.
  3. Delete the canary from S3 and re-sync before anyone sees it.

On the Wikipedia path you can't plant facts, but the [KB] injected N passages log line (and the pill's Querying knowledge base… stage) still confirm retrieval fired on every turn; ask about very recent events to see it ground on content past the model's training cutoff.

Make it yours

All in agent.py (or via .env):

  • Persona: PERSONA_NAME sets the manner (defaults to "Kenji" for Wikipedia, matching the default avatar face, and "Aristotle" for Bedrock); substance always comes from the KB. build_instructions() picks the base prompt by source: a generic grounded assistant for Wikipedia, a first-person embodiment for Bedrock.
  • Retrieved passages: RETRIEVAL_TOP_N controls how many best-first passages are injected. Passages are kept in retrieval rank order (the batched Wikipedia extracts are re-sorted to search rank; Bedrock managed search already returns best-first). For production-grade retrieval, add a reranker (e.g. Cohere Rerank or a cross-encoder) between retrieve and inject.
  • Query rewriting (Wikipedia path): spoken questions are rewritten into a clean search query by a small model (QUERY_REWRITE_MODEL, default gpt-4o-mini) before hitting Wikipedia's keyword search, so "tell me about Latvia" searches Latvia rather than the song "Santa Tell Me". A heuristic cleaner (clean_query) is the fallback. Bedrock's semantic search doesn't need this.
  • Voice: CARTESIA_VOICE_ID (Cartesia voices).
  • Avatar: SYNTHESIA_AVATAR_ID must be the Interactive ID of a synthetic or Personal avatar your workspace can use. In Synthesia Studio, open the avatar's ••• menu and choose Copy Interactive ID. Stock actor-based avatars such as Ryan or Ada can't be used as interactive avatars on any plan.
  • Models: swap the stt= / llm= / tts= lines for any LiveKit-supported provider; the RAG hook is provider-independent.

server.py mints room tokens with sync_streams=True (avatar audio/video sync) and a RoomAgentDispatch matching the worker's agent_name (see Security & production). Keep both when you build your own token endpoint, and keep SYNTHESIA_API_KEY and your AWS credentials server-side. They must never reach frontend code.

Security & production

Demo only. This is a local quickstart. Don't deploy it as-is.

  • The /token endpoint is unauthenticated. The demo server binds to localhost and mints 15-minute, room-scoped tokens, so exposure is limited to your machine. But anyone who can reach the endpoint can dispatch an agent worker, which costs money, so a production endpoint must sit behind your app's authentication.
  • Agent dispatch is explicit. The worker sets agent_name and only joins rooms whose token requests it via RoomAgentDispatch (server.py), so keep the two names in sync. Don't remove agent_name: an unnamed worker auto-joins every new room in the LiveKit project.
  • Keep secrets server-side. SYNTHESIA_API_KEY, OPENAI_API_KEY, and AWS credentials live only in .env (git-ignored) or your secrets manager, never in the frontend.

Troubleshooting

SymptomLikely cause / fix
Avatar answers but ignores the KB / makes things upRetrieval returned nothing. Check the CLI retrieve works; confirm you're using managedSearchConfiguration, the right BEDROCK_KB_ID, and the right AWS_REGION.
[KB] retrieve failed in the logAWS creds missing/expired, wrong region, or the role lacks bedrock:Retrieve on this KB.
ExpiredTokenException / InvalidSignatureExceptionTemporary AWS credentials lapsed. Refresh them.
SynthesiaError with type AUTHSynthesia API key invalid or expired.
SynthesiaError with type FEATURE_NOT_IN_PLANYour Synthesia workspace's plan doesn't include Interactive Avatars.
SynthesiaError with type LIVEKIT_CREDENTIALS_REJECTEDLIVEKIT_URL and LIVEKIT_API_KEY/LIVEKIT_API_SECRET aren't from the same LiveKit project.
SynthesiaError with type UNKNOWN_AVATARSYNTHESIA_AVATAR_ID isn't a usable Interactive ID: it's a stock actor-based avatar, it hasn't finished converting to an interactive avatar, or it isn't in your plan.
SynthesiaError with type QUOTA_EXCEEDEDSynthesia minute or concurrent-session cap hit.
Avatar never appears, no errorYou're in console mode, avatar.start() ran after session.start(), or the token lacks the RoomAgentDispatch room config.
Avatar joins but doesn't lip-syncSomething reassigned session.output.audio after the avatar attached.

Source for this recipe: interactive-avatar-quickstarts/rag.