ElevenLabs Voice Memory
Long-term memory for ElevenLabs Agents — returning callers are recognized on the web and on the phone. Three integration tiers in one package: a zero-code memory tool, session-start context with post-call fact capture from the caller’s turns, and a memory-injecting Custom LLM proxy with a hard latency budget.
Three tiers, one package
| Tier | What runs | Memory freshness | Code needed |
|---|---|---|---|
| 0 — Zero code | An ElevenLabs webhook tool calls the MemorySync REST API directly | When the LLM decides to look | None — dashboard only |
| 1 — No proxy | fetch_memory_variables at session start + the post-call webhook receiver | Start of call | ~5 lines |
| 2 — Full proxy | create_proxy_app as the agent’s Custom LLM | Every turn, under a hard budget | ~3 lines |
| Mem0 | Supermemory | Zep | MemorySync | |
|---|---|---|---|---|
| ElevenLabs integration | △ client-tools recipe — recall only if the LLM decides to call | ✗ none at all | △ proxy example, copy-paste | ✓ elevenlabs-memorysync package, three tiers |
| Recall latency budget | ✗ unbounded await | — | ✗ unbounded await | ✓ hard 1.2s default, tested against a 5s-slow backend |
| Phone-caller identity | ✗ static env var (single user!) | — | ✗ browser localStorage only | ✓ prompt-tag reads {{system__caller_id}} — zero client code |
| Upstream LLM choice | n/a | — | ✗ OpenAI hard-coded | ✓ any OpenAI-compatible provider |
| Post-call webhook capture | ✗ unused | — | ✗ unused | ✓ HMAC-verified; a turn already sent live is recognised and never extracted twice |
| Signature verification | n/a | — | shared secret only | ✓ pinned byte-for-byte to the official SDK scheme, constant-time compare |
Requirements
| You need | Where to get it | Used for |
|---|---|---|
| MemorySync API key | Dashboard → API Keys (read + write scopes) | Recall and capture |
| ElevenLabs account + agent | elevenlabs.io → Agents (free tier works for evaluation) | The hosted voice agent |
| An OpenAI-compatible LLM key | OpenAI, Azure, Groq, Gemini compat, LiteLLM… (Tier 2 only) | The model YOUR agent speaks with — the proxy forwards to it. Not used by MemorySync for memory work. |
| A public HTTPS URL | Your host, or a tunnel while developing (cloudflared tunnel --url http://127.0.0.1:8013) | ElevenLabs must reach the proxy / webhook (Tiers 1–2) |
Tier 0 needs only the first two rows — no server, no code. Python 3.10+ for Tiers 1–2.
Tier 2 — the memory proxy
pip install elevenlabs-memorysync
# server.py — deploy anywhere ElevenLabs can reachfrom elevenlabs_memorysync import create_proxy_appapp = create_proxy_app(api_key="ms_...", # MemorySync (or MEMORYSYNC_API_KEY)upstream_api_key="sk-...", # your LLM key (or OPENAI_API_KEY)# upstream_base_url="https://api.groq.com/openai/v1", # any providerproxy_api_key="a-long-random-secret",)# uvicorn server:app --host 0.0.0.0 --port 8013
# What the proxy's capture performs on the wire — one call per caller# turn. The server extracts the durable facts in it and stores only# those, grouped under the elevenlabs:: session scope. The model's# replies are not sent.curl --request POST https://api.memorysync.io/v1/memory/add_turn \--header "X-API-Key: $MEMORYSYNC_API_KEY" \--header "Content-Type: application/json" \--data '{"tenant_id":"acme","user_id":"+15551234567","source":"elevenlabs","text":"I always take the window seat","role":"user","speaker":"human@elevenlabs::+15551234567#h<content-hash>","metadata":{"session_id":"elevenlabs::conv_abc"}}'
In the agent: LLM → Custom LLM, Server URL = your deployment, Model ID = the id of your chosen upstream model, API key = your proxy_api_key. Every turn is enriched with recalled memories under the budget, and the caller’s turn is sent for fact extraction with an idempotency seed — a slow or dead memory backend degrades to an unenriched turn, never a delayed reply. The model’s replies are relayed to ElevenLabs but not sent to MemorySync.
Who is calling? The identity ladder
| Rung | Mechanism | Works for |
|---|---|---|
| 1 | customLlmExtraBody: { user_id } at session start (enable *Custom LLM extra body* in the agent’s Security tab) | Web apps and SDKs where your code knows the user |
| 2 | One line in the agent’s system prompt: memorysync-user: {{system__caller_id}} — the proxy extracts the interpolated value and strips the line before the model sees it | Phone calls (Twilio/SIP), the widget, everything — zero client code |
| 3 | default_user_id= fallback | Single-user kiosks and demos |
How a turn flows through the proxy
| # | Step |
|---|---|
| 1 | ElevenLabs POSTs the conversation to your proxy’s /v1/chat/completions — exactly as it would to OpenAI. |
| 2 | The proxy resolves the caller (extra body → prompt tag → fallback) and strips the tag line so the model never sees it. |
| 3 | The caller’s newest turn is sent for fact extraction, fire-and-forget with an idempotency seed — before the upstream call, and nothing in the reply path waits on it. |
| 4 | Recall races the hard budget (default 1.2s). On time → a guarded memory block joins the system message. Too slow → the request forwards unenriched, on time. |
| 5 | The upstream reply streams back chunk-for-chunk — tool-call deltas (end_call, transfers) pass through untouched. The reply is not sent to MemorySync. |
| 6 | Server-side, the durable facts in the caller’s turn are stored under the caller’s id — the turn text itself is not kept. |
Tier 1 — session-start context, no proxy
from elevenlabs_memorysync import fetch_memory_variables# When your app starts a session (web SDK / widget / SIP):variables = await fetch_memory_variables(user_id="customer-7")# {"memory_summary": "Prefers aisle seats. Vegetarian. ..."}conversation = Conversation(client, AGENT_ID,conversation_initiation_client_data={"dynamic_variables": variables,},)
Nothing runs in the voice path — memory is fetched once at session start and interpolated into the prompt. Pair it with the post-call webhook below for capture, and the memory is one call stale at worst: perfect when you cannot host a proxy but want recognized callers.
Post-call capture (tiers 1 and 2)
from elevenlabs_memorysync import create_webhook_appapp = create_webhook_app(api_key="ms_...",webhook_secret="wsec_...", # or ELEVENLABS_WEBHOOK_SECRET)# Configure the URL under Agents → Settings → Post-call webhooks
Signatures are verified with the exact official scheme (t=…,v0=HMAC-SHA256, 30-minute tolerance — pinned against the ElevenLabs SDK in CI) plus a constant-time compare. The caller’s turns from the transcript are sent to fact extraction with the same idempotency seeds the proxy uses, so running both gives live memory plus a sweep that catches anything missed — a turn the proxy already sent is recognised and never extracted twice, by test. The agent’s turns are not sent. The receiver answers with a summary of the call’s transcript entries — sent, stored (the same count), already, skipped, errors — and no memory ids, since facts are extracted asynchronously and get their own. A poison payload gets a 200 skip (the webhook can never be auto-disabled by one bad event); only a fully-failed delivery returns 500 so ElevenLabs redelivers.
Tier 0 — zero code, dashboard only
| Field | Value |
|---|---|
| Tool type / name | Webhook · search_memory |
| Method + URL | POST https://api.memorysync.io/memory/query |
Header X-API-Key | {{secret__memorysync_api_key}} (workspace secret — never reaches the LLM) |
Header X-End-User-ID | {{user_id}} (dynamic variable) |
| Body schema | query (string — what to look up), k (integer, default 5) |
No server at all — the agent calls MemorySync directly when the model decides to look something up. That trade-off (LLM-gated recall) is exactly why the proxy tier exists; start here, graduate up the ladder.
Configuration (proxy)
| Parameter | Default | Meaning |
|---|---|---|
upstream_base_url | https://api.openai.com/v1 | Any OpenAI-compatible provider |
proxy_api_key | none | Bearer secret ElevenLabs must present — set it in production |
recall_timeout | 1.2 | Hard recall budget in seconds |
top_k | 5 | Memories injected per turn |
min_prompt_chars | 8 | Skip recall for trivial utterances |
buffer_words | none | e.g. "One moment… " — spoken filler emitted only when the turn is already slow |
buffer_after_ms | 900 | How slow is “slow” before buffer words are used |
default_user_id | none | Identity fallback for single-user deployments |
Example: what the caller experiences
| Call 1 — caller | "My name is Sam and I run a dive shop in Sydney." |
| Agent | "Great to meet you, Sam! How can I help the shop today?" |
| *(server, seconds later)* | Extracted facts land under the caller’s id: *Name is Sam* · *Runs a dive shop* · *Lives in Sydney* |
| Call 2, days later — caller | "Hi, it’s me again — do you remember me?" |
| Agent | "Of course, Sam! You run a dive shop in Sydney. How’s everything going?" |
What happens server-side
| Concern | Behaviour |
|---|---|
| What is stored | The caller’s turns go to fact extraction, and only the durable facts in them are stored — never the turn text, and nothing from the agent’s replies. It is the pipeline every MemorySync integration shares. Filler turns store nothing. |
| Billing | Each sent caller turn counts one add request. Each enrichment — a proxy turn, or the session-start variables fetch — is a recall: one retrieval request, plus one more when it returns no context and the plain semantic query fallback runs. Deletes are free. |
| Over quota | Production keys degrade silently — sends are accepted and skipped (processing_status: "skipped"), recalls answer empty — an agent never speaks a billing error to a caller. Evaluation keys get a strict 429 instead, so coding agents see the truth. |
| Duplicates | Proxy capture and webhook capture share the same idempotency seeds — a turn sent by both is recognised server-side and extracted once, by test. |
Verify it is working
# list what the agent remembers about a callercurl "https://api.memorysync.io/v1/memory/acme/+15551234567/list?limit=20" \--header "X-API-Key: $MEMORYSYNC_API_KEY"
Or open Memory Explorer and filter by the caller id. This whole path is verified against a real ElevenLabs agent in production — agent created via the platform API, Custom LLM pointed at the proxy, two live conversations: facts landed under the prompt-tag identity, the second call answered from memory, and the tag was never spoken.
Troubleshooting
| Symptom | Cause and fix |
|---|---|
| Agent gives generic replies, no memory | The agent’s LLM is not the proxy — re-check Custom LLM Server URL (must include /v1) and that the deployment is reachable over HTTPS. |
401 in proxy logs | The agent’s Custom LLM API key must equal your proxy_api_key. |
| Memories exist but recall misses them | Identity mismatch — the caller resolved to a different user_id than the one you checked. Confirm the prompt-tag line interpolates (Rung 2) or the extra body carries user_id (Rung 1). |
| Phone callers all share one memory | You set default_user_id and no per-caller rung matched — add the prompt-tag line with {{system__caller_id}}. |
| Webhook receiver never fires | Post-call webhooks are a workspace-level setting and fire for all agents — configure the URL under Agents → Settings, and remember ElevenLabs auto-disables after 10 consecutive failures. |
| Nothing stores and nothing errors | Quota exhausted — production keys degrade silently by design. Check the Usage page. |
Privacy and data safety
Only extracted facts persist — the caller’s turns are processed and discarded, and the agent’s replies are never sent. The injected memory block is capture-excluded, so recalled context never re-enters storage. Workspace secrets (secret__*) never reach the LLM; the prompt tag is stripped before the upstream call so an identity can never be spoken aloud. Deletion is a first-class API (/memory/forget) and dashboard action, per end user. Audio never reaches MemorySync — text only.
Supported versions
| Surface | Requires | Verified on |
|---|---|---|
elevenlabs-memorysync 1.1.0 | Python 3.10+, FastAPI; an ElevenLabs agent with Custom LLM and/or post-call webhooks | 38 CI checks driven over the real ASGI apps: SSE relay with tool-call deltas split at awkward byte boundaries, hard budget vs a 5s-slow backend, prompt-tag extraction/stripping, caller-turn-only sending (replies never sent), both quota modes, webhook HMAC cross-checked against the official elevenlabs SDK at latest, and the sweep proof that a turn sent live is not extracted twice |