AutoGen Memory
Give Microsoft AutoGen agents long-term memory in three lines. The adapter implements AutoGen’s own Memory protocol: recalled context is injected before every model call under a hard latency budget, and one extra call sends the user’s side of each exchange for fact extraction — outage-proof and safe to retry.
How it works
- Install
autogen-memorysyncfrom PyPI. - Create a
MemorySyncMemorywith your API key and the end user’s id. - Pass it to any
AssistantAgentviamemory=[memory]— recall now happens automatically before every model call. - After each exchange, call
add_turn_paironce: the user’s text is sent for fact extraction and the reply is not stored. That’s the whole loop.
Before you start
- A MemorySync API key. Create one in the dashboard under Settings → API Keys, inside a project.
- Python 3.10+ with
autogen-agentchat/autogen-core0.4 or newer.
Install
One package, zero framework lock-in:
pip install autogen-memorysync
Quickstart
Step 1 — create the memory
The memory object is where you say whose memories these are. user_id is required — it keeps every customer’s memories isolated. session_id separates conversations (support ticket vs onboarding chat): every distilled fact is tagged with it, and clear() is bounded by it.
from autogen_memorysync import MemorySyncMemorymemory = MemorySyncMemory(api_key="ms_...", # or set MEMORYSYNC_API_KEY in the environmentuser_id="customer-42", # required — the end user these memories belong tosession_id="support", # optional — tags this conversation's facts; bounds clear())
Step 2 — attach it to your agent
Hand the memory to the agent. From now on, AutoGen calls the adapter before every model call: relevant memories are recalled, injected as a SystemMessage, and announced as a MemoryQueryEvent — the framework’s own eventing, nothing custom.
from autogen_agentchat.agents import AssistantAgentagent = AssistantAgent("assistant", model_client=..., memory=[memory])result = await agent.run(task="Which seat should I book for Alex?")
Step 3 — capture the exchange
By AutoGen’s own design, nothing is captured automatically — that is the application’s job. This one call sends the user’s text to fact extraction: the server stores only the durable facts in it (filler such as “ok” yields none), and the assistant’s reply is not stored. It is safe to retry: a replay is recognised server-side and not extracted twice.
await memory.add_turn_pair(user_text, result.messages[-1].content)
Optional: memory as agent tools
For tool-equipped agents, two ready-made FunctionTools let the model search and save memory on its own. search_memory performs the server’s semantic query; save_memory goes through server-side extraction, so the server — not the model — decides whether the text is durable enough to keep.
from autogen_memorysync import create_memory_toolstools = create_memory_tools(memory) # search_memory + save_memoryagent = AssistantAgent("assistant", model_client=...,memory=[memory], tools=tools)
Built-in safety guarantees
| Guarantee | How |
|---|---|
| A reply is never late | Recall waits at most recall_timeout (default 1.2s); on a miss the agent answers without memories |
| An outage never breaks the agent | update_context catches everything, logs a warning, injects nothing |
| A retry is not extracted twice | Seeds derived from role + session + content hash, recognised server-side |
| Only the user’s words are sent | add() sends user (and role-less) content as plain text with role: "user"; content whose metadata["role"] is assistant, system or tool is not sent, and add_turn_pair sends only the user text |
| Your metadata dict is safe | Copied before reading — never mutated |
clear() cannot nuke a customer | Session-scoped deletion by default; clear_scope="user" is an explicit, documented opt-in |
Configuration
| Parameter | Default | Meaning |
|---|---|---|
user_id | — (required) | End user the memories belong to — never auto-generated (a random namespace strands data) |
session_id | default | Session scope autogen::<session> — tags this conversation’s facts; bounds clear() |
top_k | 5 | Memories considered per turn |
recall_timeout | 1.2 | Hard recall budget in seconds |
min_prompt_chars | 8 | Skip recall for trivial prompts |
context_template | built-in | {context} placeholder; brace-safe .replace rendering |
clear_scope | session | clear() blast radius; "user" opt-in wipes everything |
Choosing a user_id
The user_id is a stable string you pick to identify whose memories these are: your app’s internal user ID, an email address, or a UUID. Use the same value when storing and recalling, or recall returns nothing. It also powers per-user isolation — one customer’s memories can never reach another’s agent.
Quotas and plan limits
Hitting a monthly plan limit never breaks the agent. On free and paid plans, over-limit writes are accepted without storing and reads return empty — the conversation continues. Evaluation keys instead surface a truthful 429, so you find out during testing, not in production.
Troubleshooting
- `pip install` can’t find the package — upgrade pip (
python -m pip install -U pip); the package requires Python 3.10+. - `ModuleNotFoundError: autogen_agentchat` — the framework itself isn’t installed:
pip install autogen-agentchat. - Recall returns nothing right after storing — extraction is asynchronous; give it a moment. Also confirm store and recall use the same
user_id. - The agent answered without memories once — that’s the latency budget working: a slow lookup is skipped rather than delaying the reply. The next turn recalls normally.
- `clear()` deleted less than expected — by design it clears only the current session. Pass
clear_scope="user"explicitly to wipe a user.
Supported versions
| Surface | Requires | Verified on |
|---|---|---|
autogen-memorysync 1.1.0 | autogen-agentchat / autogen-core 0.4+ (Python 3.10+) | 43 CI checks against the latest autogen-agentchat: a REAL AssistantAgent driven via the official ReplayChatCompletionClient (SystemMessage injection in the true turn order, MemoryQueryEvent emission), the role-aware retrieval regression, user-turn-only capture (assistant, system and tool content never sent), the 1.2s budget vs a slow backend, session-scoped clear, seed idempotency, metadata non-mutation, and both monthly-quota server modes |
How it compares
| Mem0 (`autogen-ext[mem0]`) | Zep (`zep-autogen`) | Supermemory | MemorySync | |
|---|---|---|---|---|
| Async correctness | ✗ sync client inside async methods — blocks the event loop | ✓ | — no adapter at all | ✓ async httpx throughout |
| Recall latency budget | ✗ none | ✗ none | — | ✓ hard 1.2s default, tested against a slow backend |
| Retrieval query | ✗ messages[-1] — assistant text after a tool result | last context message | — | ✓ last user message, role-aware |
| Explicit query() on failure | ✗ silently returns empty — outage looks like amnesia | logged | — | ✓ raises; only the hot path fails open |
clear() blast radius | ✗ the entire user | ✗ the entire user | — | ✓ session-scoped by default; whole-user wipe is an explicit opt-in |
| Retry safety | ✗ | ✗ | — | ✓ deterministic idempotency seeds |
| Docs for the current API | △ docs page shows the old 0.2 recipe | ✗ integration pages 404 | — | ✓ this page |