MemorySync
Integrations

ElevenLabs Voice Memory

Long-term memory for ElevenLabs Agents — returning callers are recognized on the web and on the phone. Three integration tiers in one package: a zero-code memory tool, session-start context with post-call fact capture from the caller’s turns, and a memory-injecting Custom LLM proxy with a hard latency budget.

ElevenLabs can recall through a webhook tool, fetch session variables, or route every turn through create_proxy_app.

Three tiers, one package

TierWhat runsMemory freshnessCode needed
0 — Zero codeAn ElevenLabs webhook tool calls the MemorySync REST API directlyWhen the LLM decides to lookNone — dashboard only
1 — No proxyfetch_memory_variables at session start + the post-call webhook receiverStart of call~5 lines
2 — Full proxycreate_proxy_app as the agent’s Custom LLMEvery turn, under a hard budget~3 lines
Mem0SupermemoryZepMemorySync
ElevenLabs integration△ client-tools recipe — recall only if the LLM decides to call✗ none at all△ proxy example, copy-paste✓ elevenlabs-memorysync package, three tiers
Recall latency budget✗ unbounded await—✗ unbounded await✓ hard 1.2s default, tested against a 5s-slow backend
Phone-caller identity✗ static env var (single user!)—✗ browser localStorage only✓ prompt-tag reads {{system__caller_id}} — zero client code
Upstream LLM choicen/a—✗ OpenAI hard-coded✓ any OpenAI-compatible provider
Post-call webhook capture✗ unused—✗ unused✓ HMAC-verified; a turn already sent live is recognised and never extracted twice
Signature verificationn/a—shared secret only✓ pinned byte-for-byte to the official SDK scheme, constant-time compare

Requirements

You needWhere to get itUsed for
MemorySync API keyDashboard → API Keys (read + write scopes)Recall and capture
ElevenLabs account + agentelevenlabs.io → Agents (free tier works for evaluation)The hosted voice agent
An OpenAI-compatible LLM keyOpenAI, Azure, Groq, Gemini compat, LiteLLM… (Tier 2 only)The model YOUR agent speaks with — the proxy forwards to it. Not used by MemorySync for memory work.
A public HTTPS URLYour host, or a tunnel while developing (cloudflared tunnel --url http://127.0.0.1:8013)ElevenLabs must reach the proxy / webhook (Tiers 1–2)

Tier 0 needs only the first two rows — no server, no code. Python 3.10+ for Tiers 1–2.

Tier 2 — the memory proxy

BASH
pip install elevenlabs-memorysync
PYTHON
# server.py — deploy anywhere ElevenLabs can reach
from elevenlabs_memorysync import create_proxy_app
app = create_proxy_app(
api_key="ms_...", # MemorySync (or MEMORYSYNC_API_KEY)
upstream_api_key="sk-...", # your LLM key (or OPENAI_API_KEY)
# upstream_base_url="https://api.groq.com/openai/v1", # any provider
proxy_api_key="a-long-random-secret",
)
# uvicorn server:app --host 0.0.0.0 --port 8013
BASH
# What the proxy's capture performs on the wire — one call per caller
# turn. The server extracts the durable facts in it and stores only
# those, grouped under the elevenlabs:: session scope. The model's
# replies are not sent.
curl --request POST https://api.memorysync.io/v1/memory/add_turn \
--header "X-API-Key: $MEMORYSYNC_API_KEY" \
--header "Content-Type: application/json" \
--data '{"tenant_id":"acme","user_id":"+15551234567","source":"elevenlabs","text":"I always take the window seat","role":"user","speaker":"human@elevenlabs::+15551234567#h<content-hash>","metadata":{"session_id":"elevenlabs::conv_abc"}}'

In the agent: LLM → Custom LLM, Server URL = your deployment, Model ID = the id of your chosen upstream model, API key = your proxy_api_key. Every turn is enriched with recalled memories under the budget, and the caller’s turn is sent for fact extraction with an idempotency seed — a slow or dead memory backend degrades to an unenriched turn, never a delayed reply. The model’s replies are relayed to ElevenLabs but not sent to MemorySync.

Who is calling? The identity ladder

RungMechanismWorks for
1customLlmExtraBody: { user_id } at session start (enable *Custom LLM extra body* in the agent’s Security tab)Web apps and SDKs where your code knows the user
2One line in the agent’s system prompt: memorysync-user: {{system__caller_id}} — the proxy extracts the interpolated value and strips the line before the model sees itPhone calls (Twilio/SIP), the widget, everything — zero client code
3default_user_id= fallbackSingle-user kiosks and demos

How a turn flows through the proxy

#Step
1ElevenLabs POSTs the conversation to your proxy’s /v1/chat/completions — exactly as it would to OpenAI.
2The proxy resolves the caller (extra body → prompt tag → fallback) and strips the tag line so the model never sees it.
3The caller’s newest turn is sent for fact extraction, fire-and-forget with an idempotency seed — before the upstream call, and nothing in the reply path waits on it.
4Recall races the hard budget (default 1.2s). On time → a guarded memory block joins the system message. Too slow → the request forwards unenriched, on time.
5The upstream reply streams back chunk-for-chunk — tool-call deltas (end_call, transfers) pass through untouched. The reply is not sent to MemorySync.
6Server-side, the durable facts in the caller’s turn are stored under the caller’s id — the turn text itself is not kept.

Tier 1 — session-start context, no proxy

from elevenlabs_memorysync import fetch_memory_variables
# When your app starts a session (web SDK / widget / SIP):
variables = await fetch_memory_variables(user_id="customer-7")
# {"memory_summary": "Prefers aisle seats. Vegetarian. ..."}
conversation = Conversation(
client, AGENT_ID,
conversation_initiation_client_data={
"dynamic_variables": variables,
},
)

Nothing runs in the voice path — memory is fetched once at session start and interpolated into the prompt. Pair it with the post-call webhook below for capture, and the memory is one call stale at worst: perfect when you cannot host a proxy but want recognized callers.

Post-call capture (tiers 1 and 2)

from elevenlabs_memorysync import create_webhook_app
app = create_webhook_app(
api_key="ms_...",
webhook_secret="wsec_...", # or ELEVENLABS_WEBHOOK_SECRET
)
# Configure the URL under Agents → Settings → Post-call webhooks

Signatures are verified with the exact official scheme (t=…,v0=HMAC-SHA256, 30-minute tolerance — pinned against the ElevenLabs SDK in CI) plus a constant-time compare. The caller’s turns from the transcript are sent to fact extraction with the same idempotency seeds the proxy uses, so running both gives live memory plus a sweep that catches anything missed — a turn the proxy already sent is recognised and never extracted twice, by test. The agent’s turns are not sent. The receiver answers with a summary of the call’s transcript entries — sent, stored (the same count), already, skipped, errors — and no memory ids, since facts are extracted asynchronously and get their own. A poison payload gets a 200 skip (the webhook can never be auto-disabled by one bad event); only a fully-failed delivery returns 500 so ElevenLabs redelivers.

Tier 0 — zero code, dashboard only

FieldValue
Tool type / nameWebhook · search_memory
Method + URLPOST https://api.memorysync.io/memory/query
Header X-API-Key{{secret__memorysync_api_key}} (workspace secret — never reaches the LLM)
Header X-End-User-ID{{user_id}} (dynamic variable)
Body schemaquery (string — what to look up), k (integer, default 5)

No server at all — the agent calls MemorySync directly when the model decides to look something up. That trade-off (LLM-gated recall) is exactly why the proxy tier exists; start here, graduate up the ladder.

Configuration (proxy)

ParameterDefaultMeaning
upstream_base_urlhttps://api.openai.com/v1Any OpenAI-compatible provider
proxy_api_keynoneBearer secret ElevenLabs must present — set it in production
recall_timeout1.2Hard recall budget in seconds
top_k5Memories injected per turn
min_prompt_chars8Skip recall for trivial utterances
buffer_wordsnonee.g. "One moment… " — spoken filler emitted only when the turn is already slow
buffer_after_ms900How slow is “slow” before buffer words are used
default_user_idnoneIdentity fallback for single-user deployments

Example: what the caller experiences

Call 1 — caller"My name is Sam and I run a dive shop in Sydney."
Agent"Great to meet you, Sam! How can I help the shop today?"
*(server, seconds later)*Extracted facts land under the caller’s id: *Name is Sam* · *Runs a dive shop* · *Lives in Sydney*
Call 2, days later — caller"Hi, it’s me again — do you remember me?"
Agent"Of course, Sam! You run a dive shop in Sydney. How’s everything going?"

What happens server-side

ConcernBehaviour
What is storedThe caller’s turns go to fact extraction, and only the durable facts in them are stored — never the turn text, and nothing from the agent’s replies. It is the pipeline every MemorySync integration shares. Filler turns store nothing.
BillingEach sent caller turn counts one add request. Each enrichment — a proxy turn, or the session-start variables fetch — is a recall: one retrieval request, plus one more when it returns no context and the plain semantic query fallback runs. Deletes are free.
Over quotaProduction keys degrade silently — sends are accepted and skipped (processing_status: "skipped"), recalls answer empty — an agent never speaks a billing error to a caller. Evaluation keys get a strict 429 instead, so coding agents see the truth.
DuplicatesProxy capture and webhook capture share the same idempotency seeds — a turn sent by both is recognised server-side and extracted once, by test.

Verify it is working

BASH
# list what the agent remembers about a caller
curl "https://api.memorysync.io/v1/memory/acme/+15551234567/list?limit=20" \
--header "X-API-Key: $MEMORYSYNC_API_KEY"

Or open Memory Explorer and filter by the caller id. This whole path is verified against a real ElevenLabs agent in production — agent created via the platform API, Custom LLM pointed at the proxy, two live conversations: facts landed under the prompt-tag identity, the second call answered from memory, and the tag was never spoken.

Troubleshooting

SymptomCause and fix
Agent gives generic replies, no memoryThe agent’s LLM is not the proxy — re-check Custom LLM Server URL (must include /v1) and that the deployment is reachable over HTTPS.
401 in proxy logsThe agent’s Custom LLM API key must equal your proxy_api_key.
Memories exist but recall misses themIdentity mismatch — the caller resolved to a different user_id than the one you checked. Confirm the prompt-tag line interpolates (Rung 2) or the extra body carries user_id (Rung 1).
Phone callers all share one memoryYou set default_user_id and no per-caller rung matched — add the prompt-tag line with {{system__caller_id}}.
Webhook receiver never firesPost-call webhooks are a workspace-level setting and fire for all agents — configure the URL under Agents → Settings, and remember ElevenLabs auto-disables after 10 consecutive failures.
Nothing stores and nothing errorsQuota exhausted — production keys degrade silently by design. Check the Usage page.

Privacy and data safety

Only extracted facts persist — the caller’s turns are processed and discarded, and the agent’s replies are never sent. The injected memory block is capture-excluded, so recalled context never re-enters storage. Workspace secrets (secret__*) never reach the LLM; the prompt tag is stripped before the upstream call so an identity can never be spoken aloud. Deletion is a first-class API (/memory/forget) and dashboard action, per end user. Audio never reaches MemorySync — text only.

Supported versions

SurfaceRequiresVerified on
elevenlabs-memorysync 1.1.0Python 3.10+, FastAPI; an ElevenLabs agent with Custom LLM and/or post-call webhooks38 CI checks driven over the real ASGI apps: SSE relay with tool-call deltas split at awkward byte boundaries, hard budget vs a 5s-slow backend, prompt-tag extraction/stripping, caller-turn-only sending (replies never sent), both quota modes, webhook HMAC cross-checked against the official elevenlabs SDK at latest, and the sweep proof that a turn sent live is not extracted twice

Where to go next

Was this page helpful?