LiveKit Voice Memory
Long-term memory for LiveKit Agents voice AI — callers your agent remembers across calls, injected under a hard latency budget so a reply is never late. Background prefetch makes the common case a zero-network cache hit, and every caller turn is sent for fact extraction with idempotency guarantees.
What the package provides
| Piece | What it does |
|---|---|
MemorySyncMemory (composition) | Attach memory to YOUR Agent subclass: one call in on_user_turn_completed injects recalled context under a hard budget (default 1.2s); attach(session) sends the caller’s turns for fact extraction as they finalize. No inheritance demanded. |
MemorySyncAgent (drop-in) | An Agent subclass with recall and capture pre-wired, for greenfield agents. |
create_memory_search_tool | A function_tool the LLM calls to search memory on demand — the right recall path for speech-to-speech realtime models. Errors return as readable strings, never raise into the model. |
| Failure contract | Slow backend → reply proceeds without memories. Dead backend, quota exhaustion → same. A memory outage can never stall or break a call, by test. |
| Mem0 | Supermemory | Zep | MemorySync | |
|---|---|---|---|---|
| LiveKit integration | △ docs recipe, no package | ✗ none | ✓ zep-livekit package | ✓ livekit-memorysync package |
| Recall latency budget | ✗ unbounded await in the reply path | — | ✗ unbounded await | ✓ hard timeout, default 1.2s, tested against a 5s-slow backend |
| Prefetch | ✗ | — | ✗ | ✓ next recall warmed in the background — common case is a cache hit |
| Composition or inheritance | copy-paste recipe | — | must subclass ZepUserAgent | ✓ composable engine — keep your own Agent class (drop-in also available) |
| Injected context re-stored as memory | unguarded | — | unguarded | ✓ injection is turn-only and capture-excluded, tested |
Requirements
| You need | Where to get it | Used for |
|---|---|---|
| MemorySync API key | Dashboard → API Keys (read + write scopes) | Recall and capture |
| LiveKit server or Cloud project | livekit.io — Cloud free tier works; note LIVEKIT_URL, LIVEKIT_API_KEY, LIVEKIT_API_SECRET | The realtime room your agent joins |
| An LLM / STT / TTS provider | Any combination LiveKit Agents supports — the single-key path is OpenAI for all three | YOUR agent’s voice loop — its brain, ears and voice. Not used by MemorySync. |
Python 3.10+ with livekit-agents 1.0+ | pip install "livekit-agents[openai,silero]" | The agent runtime |
LIVEKIT_URL=wss://your-project.livekit.cloudLIVEKIT_API_KEY=your_livekit_keyLIVEKIT_API_SECRET=your_livekit_secret# your agent's own voice loop (LLM/STT/TTS) - never sent to MemorySyncOPENAI_API_KEY=sk-...# the only key MemorySync needsMEMORYSYNC_API_KEY=ms_...
No MemorySync-side setup is needed beyond the key — tenants, embeddings and fact extraction are provisioned automatically on first write.
Install and wire it in
pip install livekit-memorysync
from livekit.agents import Agent, AgentSessionfrom livekit_memorysync import MemorySyncMemorymemory = MemorySyncMemory(api_key="ms_...", # or MEMORYSYNC_API_KEY env varuser_id="caller-42", # stable end-user idthread_id="room-123", # optional: scope to this room)class Assistant(Agent):def __init__(self) -> None:super().__init__(instructions="You are a helpful voice assistant.")async def on_user_turn_completed(self, turn_ctx, new_message):# Inject memories for THIS turn only (turn-scoped, not persisted)await memory.on_user_turn(turn_ctx, new_message)session = AgentSession(...) # your STT / LLM / TTS choicesmemory.attach(session) # send the caller's turns as they finalizeawait session.start(agent=Assistant(), ...)
# What capture performs on the wire — one call per finalized caller turn.# The server extracts the durable facts in it and stores only those,# grouped under the livekit:: session scope. The agent's replies are# not sent.curl --request POST https://api.memorysync.io/v1/memory/add_turn \--header "X-API-Key: $MEMORYSYNC_API_KEY" \--header "Content-Type: application/json" \--data '{"tenant_id":"acme","user_id":"caller-42","source":"livekit","text":"I always take the window seat","role":"user","speaker":"human@livekit::room-123#h<content-hash>","metadata":{"session_id":"livekit::room-123","item_id":"<livekit-item-id>"}}'
Prefer zero wiring? MemorySyncAgent(instructions=..., api_key=..., user_id=...) is the same engine pre-attached. Every sent turn carries a deterministic idempotency seed, so the server recognises reconnects and retries and never extracts a turn twice.
Complete runnable example
import osfrom dotenv import load_dotenvfrom livekit.agents import Agent, AgentSession, JobContext, WorkerOptions, clifrom livekit.plugins import openai, silerofrom livekit_memorysync import MemorySyncMemoryload_dotenv()class Assistant(Agent):def __init__(self, memory: MemorySyncMemory) -> None:super().__init__(instructions=("You are a warm, concise voice assistant. ""Use anything you remember about the caller naturally."))self._memory = memoryasync def on_user_turn_completed(self, turn_ctx, new_message):await self._memory.on_user_turn(turn_ctx, new_message)async def entrypoint(ctx: JobContext):await ctx.connect()memory = MemorySyncMemory(user_id="caller-42", # your stable end-user id (see Identity)thread_id=ctx.room.name, # group this call's facts by room)session = AgentSession(vad=silero.VAD.load(),stt=openai.STT(),llm=openai.LLM(model="your-model"),tts=openai.TTS(voice="ash"),)memory.attach(session) # send the caller's turns as they finalizeawait session.start(agent=Assistant(memory), room=ctx.room)await session.generate_reply(instructions="Greet the caller briefly.")if __name__ == "__main__":cli.run_app(WorkerOptions(entrypoint_fnc=entrypoint))
Say "My name is Alex and my favorite color is teal", hang up, reconnect — the agent greets you with what it remembers. console mode is the fastest way to see the loop working before you deploy a room.
Understanding user_id and thread_id
| Field | What it scopes | Choose it like this |
|---|---|---|
user_id | The person. All recall and all facts attach to this id — it is what makes the caller recognized across calls, rooms, and even other MemorySync integrations. | A stable id from YOUR system: account id, CRM id, or the SIP caller number for telephony. Never a room name, never a random per-session value. |
thread_id | The conversation. Groups the facts distilled from this call under one session scope (livekit::<thread>). | The room name or call id. Optional — omit it and the facts still store under the user. |
What runs when
| Moment | What happens |
|---|---|
| User turn completes | The block prefetched after the previous turn is used when it is ready — it was recalled for the previous utterance, so there is no network wait; otherwise recall runs now under the hard budget. The memory block lands in the turn context with a guard line — background information, not instructions. |
| LLM reply | Proceeds on time regardless: with memories when recall answered, without them when it did not. |
| Conversation item finalizes | A caller turn is sent in the background for fact extraction, with its LiveKit item_id (and interrupted: true when LiveKit marks it interrupted) as metadata. The agent’s replies are not sent, and injected memory blocks are excluded — recalled context never re-enters storage. |
| After each turn | A recall for this turn’s text runs in the background (with prefetch=True, the default), so the next injection is usually a zero-network cache hit. That recall is metered whether or not the next turn uses it. |
| Backend slow / dead / over quota | Empty context, silent skip — the call continues. Free-tier quota exhaustion is silent by design; evaluation keys surface strict 429s instead. |
On-demand recall: the memory search tool
from livekit_memorysync import MemorySyncMemory, create_memory_search_toolmemory = MemorySyncMemory(api_key="ms_...", user_id="caller-42")search_memory = create_memory_search_tool(memory)class Assistant(Agent):def __init__(self) -> None:super().__init__(instructions="Search your memory when the caller references the past.",tools=[search_memory],)
The tool returns matches as a readable string and never raises into the model — a memory outage becomes "no results", not a broken tool call. It composes with turn injection: injection covers the ambient "remember me" baseline, the tool covers pointed questions like "what did I order last time?".
Configuration
| Parameter | Default | Meaning |
|---|---|---|
api_key | MEMORYSYNC_API_KEY env | MemorySync API key |
user_id | required | Stable end-user identity — memories follow this id across calls |
thread_id | default | Group the facts from one room or call thread (scope livekit::<thread_id>) |
recall_timeout | 1.2 | Hard budget in seconds for recall injection |
top_k | 6 | Memories injected per turn |
persist_injection | False | True writes the memory block into the session context instead of turn-only |
prefetch | True | Warm the next recall in the background |
Example: what the caller experiences
| Call 1 — caller | "Hi! I’m Emma. I’m vegetarian and I always book aisle seats." |
| Agent | "Nice to meet you, Emma! Noted — vegetarian meals and aisle seats. How can I help today?" |
| *(server, seconds later)* | Extracted facts land: *Is vegetarian* · *Prefers aisle seats* — not the caller’s words, and nothing from the agent’s reply |
| Call 2, next week — caller | "Can you book my usual flight setup?" |
| Agent | "Of course — aisle seat as always, and I’ll flag the vegetarian meal, Emma." |
What happens server-side
| Concern | Behaviour |
|---|---|
| What is stored | The caller’s turns go to fact extraction, and only the durable facts in them ("Prefers aisle seats") are stored — never the turn text, and nothing from the agent’s replies. It is the same pipeline every MemorySync integration uses. Filler turns ("ok", "hmm") store nothing. |
| Billing | Each sent caller turn counts one add request. Each recall counts one retrieval request, plus one more when it returns no context and the plain semantic query fallback runs. A turn injection that is not served from the prefetch, the background prefetch after each turn and each memory-tool search are all recalls. Deletes are free. |
| Over quota | Production keys degrade silently — sends are accepted and skipped (processing_status: "skipped"), recalls answer empty — so your agent never speaks a billing error to a caller. Evaluation keys get a strict 429 instead, so coding agents see the truth. |
| Duplicates | Deterministic idempotency seeds — reconnects, retries, and replays are recognised server-side and extracted once. |
Verify it is working
# list what the agent remembers about a callercurl "https://api.memorysync.io/v1/memory/acme/caller-42/list?limit=20" \--header "X-API-Key: $MEMORYSYNC_API_KEY"
Or open Memory Explorer in the dashboard and filter by the user id — facts appear within seconds of each turn (extraction is asynchronous, typically under 10s). During development, use a fresh user_id per take so old test facts never leak into a demo.
Troubleshooting
| Symptom | Cause and fix |
|---|---|
| Agent never remembers anything | Almost always a user_id that changes between calls — pin it to a stable id. Second check: the key needs write scope. |
| Remembers within a call, not across calls | Extraction is asynchronous — give it ~10s after the call. If facts exist in Memory Explorer but recall misses them, raise top_k. |
| Replies feel slow on a home network | Raise recall_timeout (e.g. 6.0) during development — the default 1.2s budget is tuned for datacenter latency and will skip injection rather than wait. |
| Injection lands after the model starts speaking | Realtime speech-to-speech pipeline — use the memory search tool instead (see above). |
| Nothing stores and nothing errors | Quota exhausted — production keys degrade silently by design. Check the Usage page; the dashboard shows the silent-skip events. |
Privacy and data safety
Only extracted facts persist — the caller’s turns are processed and discarded, and the agent’s replies are never sent, so chit-chat, filler and mishears never become permanent records. The injected memory block is turn-scoped and capture-excluded: recalled context can never re-enter storage or compound. Deletion is a first-class API (/memory/forget) and dashboard action, per end user. The memory engine sees text only — audio never reaches MemorySync.
Supported versions
| Surface | Requires | Verified on |
|---|---|---|
livekit-memorysync 1.1.1 | livekit-agents 1.0+, Python 3.10+ | 17 CI checks driven through real AgentSession machinery on the latest livekit-agents: budget enforcement against a 5s-slow backend, prefetch instant-hit, injection exclusion, caller-turn-only capture (replies never sent), idempotency seeds, thread scoping, quota modes, dead-server survival |