MemorySync
Integrations

LiveKit Voice Memory

Long-term memory for LiveKit Agents voice AI — callers your agent remembers across calls, injected under a hard latency budget so a reply is never late. Background prefetch makes the common case a zero-network cache hit, and every caller turn is sent for fact extraction with idempotency guarantees.

LiveKit takes prefetched recall at the turn boundary, answers on time, then sends finalized caller turns for fact extraction.

What the package provides

PieceWhat it does
MemorySyncMemory (composition)Attach memory to YOUR Agent subclass: one call in on_user_turn_completed injects recalled context under a hard budget (default 1.2s); attach(session) sends the caller’s turns for fact extraction as they finalize. No inheritance demanded.
MemorySyncAgent (drop-in)An Agent subclass with recall and capture pre-wired, for greenfield agents.
create_memory_search_toolA function_tool the LLM calls to search memory on demand — the right recall path for speech-to-speech realtime models. Errors return as readable strings, never raise into the model.
Failure contractSlow backend → reply proceeds without memories. Dead backend, quota exhaustion → same. A memory outage can never stall or break a call, by test.
Mem0SupermemoryZepMemorySync
LiveKit integration△ docs recipe, no package✗ none✓ zep-livekit package✓ livekit-memorysync package
Recall latency budget✗ unbounded await in the reply path—✗ unbounded await✓ hard timeout, default 1.2s, tested against a 5s-slow backend
Prefetch✗—✗✓ next recall warmed in the background — common case is a cache hit
Composition or inheritancecopy-paste recipe—must subclass ZepUserAgent✓ composable engine — keep your own Agent class (drop-in also available)
Injected context re-stored as memoryunguarded—unguarded✓ injection is turn-only and capture-excluded, tested

Requirements

You needWhere to get itUsed for
MemorySync API keyDashboard → API Keys (read + write scopes)Recall and capture
LiveKit server or Cloud projectlivekit.io — Cloud free tier works; note LIVEKIT_URL, LIVEKIT_API_KEY, LIVEKIT_API_SECRETThe realtime room your agent joins
An LLM / STT / TTS providerAny combination LiveKit Agents supports — the single-key path is OpenAI for all threeYOUR agent’s voice loop — its brain, ears and voice. Not used by MemorySync.
Python 3.10+ with livekit-agents 1.0+pip install "livekit-agents[openai,silero]"The agent runtime
BASH
LIVEKIT_URL=wss://your-project.livekit.cloud
LIVEKIT_API_KEY=your_livekit_key
LIVEKIT_API_SECRET=your_livekit_secret
# your agent's own voice loop (LLM/STT/TTS) - never sent to MemorySync
OPENAI_API_KEY=sk-...
# the only key MemorySync needs
MEMORYSYNC_API_KEY=ms_...

No MemorySync-side setup is needed beyond the key — tenants, embeddings and fact extraction are provisioned automatically on first write.

Install and wire it in

BASH
pip install livekit-memorysync
PYTHON
from livekit.agents import Agent, AgentSession
from livekit_memorysync import MemorySyncMemory
memory = MemorySyncMemory(
api_key="ms_...", # or MEMORYSYNC_API_KEY env var
user_id="caller-42", # stable end-user id
thread_id="room-123", # optional: scope to this room
)
class Assistant(Agent):
def __init__(self) -> None:
super().__init__(instructions="You are a helpful voice assistant.")
async def on_user_turn_completed(self, turn_ctx, new_message):
# Inject memories for THIS turn only (turn-scoped, not persisted)
await memory.on_user_turn(turn_ctx, new_message)
session = AgentSession(...) # your STT / LLM / TTS choices
memory.attach(session) # send the caller's turns as they finalize
await session.start(agent=Assistant(), ...)
BASH
# What capture performs on the wire — one call per finalized caller turn.
# The server extracts the durable facts in it and stores only those,
# grouped under the livekit:: session scope. The agent's replies are
# not sent.
curl --request POST https://api.memorysync.io/v1/memory/add_turn \
--header "X-API-Key: $MEMORYSYNC_API_KEY" \
--header "Content-Type: application/json" \
--data '{"tenant_id":"acme","user_id":"caller-42","source":"livekit","text":"I always take the window seat","role":"user","speaker":"human@livekit::room-123#h<content-hash>","metadata":{"session_id":"livekit::room-123","item_id":"<livekit-item-id>"}}'

Prefer zero wiring? MemorySyncAgent(instructions=..., api_key=..., user_id=...) is the same engine pre-attached. Every sent turn carries a deterministic idempotency seed, so the server recognises reconnects and retries and never extracts a turn twice.

Complete runnable example

import os
from dotenv import load_dotenv
from livekit.agents import Agent, AgentSession, JobContext, WorkerOptions, cli
from livekit.plugins import openai, silero
from livekit_memorysync import MemorySyncMemory
load_dotenv()
class Assistant(Agent):
def __init__(self, memory: MemorySyncMemory) -> None:
super().__init__(instructions=(
"You are a warm, concise voice assistant. "
"Use anything you remember about the caller naturally."
))
self._memory = memory
async def on_user_turn_completed(self, turn_ctx, new_message):
await self._memory.on_user_turn(turn_ctx, new_message)
async def entrypoint(ctx: JobContext):
await ctx.connect()
memory = MemorySyncMemory(
user_id="caller-42", # your stable end-user id (see Identity)
thread_id=ctx.room.name, # group this call's facts by room
)
session = AgentSession(
vad=silero.VAD.load(),
stt=openai.STT(),
llm=openai.LLM(model="your-model"),
tts=openai.TTS(voice="ash"),
)
memory.attach(session) # send the caller's turns as they finalize
await session.start(agent=Assistant(memory), room=ctx.room)
await session.generate_reply(instructions="Greet the caller briefly.")
if __name__ == "__main__":
cli.run_app(WorkerOptions(entrypoint_fnc=entrypoint))

Say "My name is Alex and my favorite color is teal", hang up, reconnect — the agent greets you with what it remembers. console mode is the fastest way to see the loop working before you deploy a room.

Understanding user_id and thread_id

FieldWhat it scopesChoose it like this
user_idThe person. All recall and all facts attach to this id — it is what makes the caller recognized across calls, rooms, and even other MemorySync integrations.A stable id from YOUR system: account id, CRM id, or the SIP caller number for telephony. Never a room name, never a random per-session value.
thread_idThe conversation. Groups the facts distilled from this call under one session scope (livekit::<thread>).The room name or call id. Optional — omit it and the facts still store under the user.

What runs when

MomentWhat happens
User turn completesThe block prefetched after the previous turn is used when it is ready — it was recalled for the previous utterance, so there is no network wait; otherwise recall runs now under the hard budget. The memory block lands in the turn context with a guard line — background information, not instructions.
LLM replyProceeds on time regardless: with memories when recall answered, without them when it did not.
Conversation item finalizesA caller turn is sent in the background for fact extraction, with its LiveKit item_id (and interrupted: true when LiveKit marks it interrupted) as metadata. The agent’s replies are not sent, and injected memory blocks are excluded — recalled context never re-enters storage.
After each turnA recall for this turn’s text runs in the background (with prefetch=True, the default), so the next injection is usually a zero-network cache hit. That recall is metered whether or not the next turn uses it.
Backend slow / dead / over quotaEmpty context, silent skip — the call continues. Free-tier quota exhaustion is silent by design; evaluation keys surface strict 429s instead.

On-demand recall: the memory search tool

PYTHON
from livekit_memorysync import MemorySyncMemory, create_memory_search_tool
memory = MemorySyncMemory(api_key="ms_...", user_id="caller-42")
search_memory = create_memory_search_tool(memory)
class Assistant(Agent):
def __init__(self) -> None:
super().__init__(
instructions="Search your memory when the caller references the past.",
tools=[search_memory],
)

The tool returns matches as a readable string and never raises into the model — a memory outage becomes "no results", not a broken tool call. It composes with turn injection: injection covers the ambient "remember me" baseline, the tool covers pointed questions like "what did I order last time?".

Configuration

ParameterDefaultMeaning
api_keyMEMORYSYNC_API_KEY envMemorySync API key
user_idrequiredStable end-user identity — memories follow this id across calls
thread_iddefaultGroup the facts from one room or call thread (scope livekit::<thread_id>)
recall_timeout1.2Hard budget in seconds for recall injection
top_k6Memories injected per turn
persist_injectionFalseTrue writes the memory block into the session context instead of turn-only
prefetchTrueWarm the next recall in the background

Example: what the caller experiences

Call 1 — caller"Hi! I’m Emma. I’m vegetarian and I always book aisle seats."
Agent"Nice to meet you, Emma! Noted — vegetarian meals and aisle seats. How can I help today?"
*(server, seconds later)*Extracted facts land: *Is vegetarian* · *Prefers aisle seats* — not the caller’s words, and nothing from the agent’s reply
Call 2, next week — caller"Can you book my usual flight setup?"
Agent"Of course — aisle seat as always, and I’ll flag the vegetarian meal, Emma."

What happens server-side

ConcernBehaviour
What is storedThe caller’s turns go to fact extraction, and only the durable facts in them ("Prefers aisle seats") are stored — never the turn text, and nothing from the agent’s replies. It is the same pipeline every MemorySync integration uses. Filler turns ("ok", "hmm") store nothing.
BillingEach sent caller turn counts one add request. Each recall counts one retrieval request, plus one more when it returns no context and the plain semantic query fallback runs. A turn injection that is not served from the prefetch, the background prefetch after each turn and each memory-tool search are all recalls. Deletes are free.
Over quotaProduction keys degrade silently — sends are accepted and skipped (processing_status: "skipped"), recalls answer empty — so your agent never speaks a billing error to a caller. Evaluation keys get a strict 429 instead, so coding agents see the truth.
DuplicatesDeterministic idempotency seeds — reconnects, retries, and replays are recognised server-side and extracted once.

Verify it is working

BASH
# list what the agent remembers about a caller
curl "https://api.memorysync.io/v1/memory/acme/caller-42/list?limit=20" \
--header "X-API-Key: $MEMORYSYNC_API_KEY"

Or open Memory Explorer in the dashboard and filter by the user id — facts appear within seconds of each turn (extraction is asynchronous, typically under 10s). During development, use a fresh user_id per take so old test facts never leak into a demo.

Troubleshooting

SymptomCause and fix
Agent never remembers anythingAlmost always a user_id that changes between calls — pin it to a stable id. Second check: the key needs write scope.
Remembers within a call, not across callsExtraction is asynchronous — give it ~10s after the call. If facts exist in Memory Explorer but recall misses them, raise top_k.
Replies feel slow on a home networkRaise recall_timeout (e.g. 6.0) during development — the default 1.2s budget is tuned for datacenter latency and will skip injection rather than wait.
Injection lands after the model starts speakingRealtime speech-to-speech pipeline — use the memory search tool instead (see above).
Nothing stores and nothing errorsQuota exhausted — production keys degrade silently by design. Check the Usage page; the dashboard shows the silent-skip events.

Privacy and data safety

Only extracted facts persist — the caller’s turns are processed and discarded, and the agent’s replies are never sent, so chit-chat, filler and mishears never become permanent records. The injected memory block is turn-scoped and capture-excluded: recalled context can never re-enter storage or compound. Deletion is a first-class API (/memory/forget) and dashboard action, per end user. The memory engine sees text only — audio never reaches MemorySync.

Supported versions

SurfaceRequiresVerified on
livekit-memorysync 1.1.1livekit-agents 1.0+, Python 3.10+17 CI checks driven through real AgentSession machinery on the latest livekit-agents: budget enforcement against a 5s-slow backend, prefetch instant-hit, injection exclusion, caller-turn-only capture (replies never sent), idempotency seeds, thread scoping, quota modes, dead-server survival

Where to go next

Was this page helpful?