Respan Memory Observability
MemorySync memory + Respan (formerly Keywords AI) gateway observability: recall, LLM call, and persist as one fully traced, cost-tracked waterfall. Respan documents that the incumbent’s memory operations cannot be traced at the SDK level — MemorySync ships the first instrumentor that can.
How it works
- Respan routes your LLM traffic through one OpenAI-compatible gateway and turns every call into traces, spend metrics, and evals.
- Respan’s tracing stack is OpenTelemetry underneath — the same pipeline our instrumentor emits into.
- Instrument the MemorySync SDK once, and the memory spans land in the SAME waterfall as the gateway’s LLM spans: recall latency, model cost, and persist idempotency side by side.
Before you start
- A Respan account and API key (the gateway lives at
https://api.respan.ai/api/). - A MemorySync API key. Create one in the dashboard under Settings → API Keys, inside a project.
- Python 3.9+ with the
memorysyncSDK; any OpenAI-compatible SDK for the gateway side.
Install
pip install opentelemetry-instrumentation-memorysync openai
The pattern: recall → gateway → persist
Step 1 — instrument and point the clients
One instrumentor line; the LLM client simply points at Respan’s gateway:
from openai import OpenAIfrom memorysync import MemorySyncClientfrom opentelemetry.instrumentation.memorysync import instrument_memorysyncinstrument_memorysync()ms = MemorySyncClient(api_key="ms_...", base_url="https://api.memorysync.io")llm = OpenAI(api_key=RESPAN_API_KEY, base_url="https://api.respan.ai/api/")
Step 2 — one traced turn
Recall becomes span 1, the gateway traces the model call, the persist becomes span 2 — one waterfall:
ctx = ms.recall(tenant_id=t, user_id=u, prompt=question, k=6) # span 1reply = llm.chat.completions.create(model="your-model", messages=[{"role": "system", "content": f"Memory:\n{ctx.get('context', '')}"},{"role": "user", "content": question},]) # gateway-tracedms.add_turn(tenant_id=t, user_id=u,text=question, role="user") # span 2
add_turn sends the user’s turn to fact extraction: only the durable facts it contains are stored, as memories, so span 2 reports the extraction receipt (processing_status, request_id) rather than a memory id. Assistant replies are not stored as memories, so the reply is not sent.
Memory content stays out of the spans by default — counts, scores, ids, and latencies only. See the AgentOps guide’s privacy section for the capture-content opt-in; it is the same instrumentor.
The gap this closes
Respan’s own Mem0 integration page states verbatim: *"Mem0 doesn’t expose its own SDK-level tracing instrumentor; route all calls through the Respan gateway to get full observability automatically."* Routing Mem0’s internal LLM calls through the gateway shows those LLM calls — but the memory operations themselves (the add, the search) never appear as spans. A slow search, a failed add, or a quota rejection is invisible.
| Mem0 + Respan | MemorySync + Respan | |
|---|---|---|
| LLM calls traced | ✓ via gateway | ✓ via gateway |
| Memory operations as spans | ✗ invisible (no SDK instrumentor) | ✓ first-class memorysync.* spans |
| Recall latency visible | ✗ | ✓ with server-side latency attribute |
| Failed persists visible | ✗ | ✓ ERROR spans with HTTP status |
| Memory content in traces | n/a | never, unless explicitly opted in |
Troubleshooting
- LLM spans appear but memory spans don’t —
instrument_memorysync()must run in the same process before the memory calls; confirm Respan’s tracing is initialized so a global OTel provider exists. - Spans in two separate traces — start the turn under one root span (your framework’s request/turn span) so recall, LLM call, and persist share a trace id.
- Costs missing on the LLM call — that side comes from the Respan gateway; confirm the client’s
base_urlpoints athttps://api.respan.ai/api/with your Respan key.
Supported versions
| Surface | Requires | Verified on |
|---|---|---|
opentelemetry-instrumentation-memorysync 1.1.0 | Python 3.9+; any OpenAI-compatible SDK for the gateway side | 74 CI checks (shared with the AgentOps pairing): real-SDK spans including add_turn extraction receipts and the history and state spans, privacy default, pass-through fidelity, error spans, lifecycle, late-binding provider seam |