AG2 Memory
An automatic memory loop for AG2’s classic ConversableAgent framework: one attach call sends the user’s turns for fact extraction and injects recalled context before every reply — duplicate-proof in multi-agent chats, deadlock-free under asyncio, zero framework dependencies.
How it works
- Install
ag2-memorysyncfrom PyPI. - Create a
MemorySyncCapabilitywith your API key and the end user’s id. - Attach it to any
ConversableAgentwith one call:memory.add_to_agent(assistant). - Done — from now on the durable facts in every user message are remembered and relevant context is recalled before each answer. The agent’s own replies are not stored. No per-turn code.
Before you start
- A MemorySync API key. Create one in the dashboard under Settings → API Keys, inside a project.
- Python 3.10+ with the classic ConversableAgent framework (
pip install autogen). The rewrittenpip install ag2v1 has no hook system yet — a future adapter target.
Install
One package, zero framework dependencies:
pip install ag2-memorysync
Quickstart
Step 1 — create the capability
The capability is where you say whose memories these are. user_id is required; session_id separates conversations — every distilled fact is tagged with it.
from ag2_memorysync import MemorySyncCapabilitymemory = MemorySyncCapability(api_key="ms_...", # or set MEMORYSYNC_API_KEY in the environmentuser_id="customer-42", # required — the end user these memories belong tosession_id="support", # optional — tags this conversation's facts)
Step 2 — attach it to your agent
One call wires the whole loop. AG2-classic has no memory protocol — register_hook is its only per-turn seam — so the capability registers two hooks for you and everything else is automatic.
from autogen import ConversableAgentassistant = ConversableAgent("assistant", llm_config=...)memory.add_to_agent(assistant) # done
The automatic loop, precisely
| Hook | Fires | What the capability does |
|---|---|---|
process_last_received_message | inside generate_reply, before the LLM | Sends the ORIGINAL incoming user text for fact extraction (fire-and-forget), recalls context under the budget, returns the text with a context prefix — AG2 hands the prefixed text to the LLM but never writes it back to the stored conversation |
process_message_before_send | on send/a_send, after the reply is generated | Sends the outgoing message only when the sending agent speaks for a person (fire-and-forget); returns the message untouched. An assistant reply is not sent |
Tool and function messages are never sent. Each user turn goes to MemorySync as plain text with role: "user" under the ag2::<session> session scope, with deterministic idempotency seeds; the server extracts the durable facts it contains and stores only those — the same shared user memories as every other MemorySync surface. Filler such as “ok” yields none, and the turn text itself is not stored.
The deadlock-free bridge
All async work rides one persistent background event loop per process. Hooks submit coroutines with run_coroutine_threadsafe and block — bounded by the recall budget — on a plain future. The caller’s thread and event loop are never touched, so a chat driven from inside asyncio.run(...) completes (a test proves it). Persistence is fire-and-forget off the hot path; memory.flush() waits for in-flight writes at shutdown, and memory.close() flushes and releases the HTTP client.
Optional: memory as agent tools
If you want the model itself to search or save memory, register the two ready-made tools. Both are synchronous — they ride the same bridge, so they work in plain initiate_chat without event-loop concerns. The caller agent needs an llm_config, as usual for AG2 tool registration.
from ag2_memorysync import register_memory_tools# caller decides to use tools; executor runs them (AG2's standard split)register_memory_tools(memory, caller=assistant, executor=user_proxy)
Configuration
| Parameter | Default | Meaning |
|---|---|---|
user_id | — (required) | End user the memories belong to |
session_id | default | Session scope ag2::<session> — tags this conversation’s facts |
top_k | 5 | Memories considered per turn |
recall_timeout | 1.2 | Hard recall budget in seconds |
min_prompt_chars | 8 | Skip recall for trivial messages |
context_template | built-in | {context} placeholder; brace-safe .replace rendering |
capture | both | received / sent to capture one side only. The sent side carries only what an agent speaking for a person says, so sent belongs on the person’s UserProxyAgent; on an assistant it sends nothing |
Choosing a user_id
The user_id is a stable string you pick to identify whose memories these are: your app’s internal user ID, an email address, or a UUID. Use the same value across sessions or recall returns nothing. It also powers per-user isolation — one customer’s memories can never reach another’s agent.
Quotas and plan limits
Hitting a monthly plan limit never breaks a conversation. On free and paid plans, over-limit writes are accepted without storing and reads return empty — the chat continues. Evaluation keys instead surface a truthful 429, so you find out during testing, not in production.
Troubleshooting
- `ModuleNotFoundError: autogen` — install the classic framework:
pip install autogen. The rewrittenpip install ag2v1 has no hooks yet and is not supported. - Two agents in one chat, each user message sent once — that’s correct: the dedup registry prevents the double-store bug zep-ag2 documents.
- Nothing is stored for the assistant’s replies — by design: assistant replies are not stored as memories. Only the durable facts in what the person says are kept.
- A chat inside `asyncio.run(...)` hangs with other adapters — not this one: the bridge never touches your event loop. If you see a hang, it isn’t the memory layer.
- Recall seems empty right after a message — extraction is asynchronous, and trivial messages (under
min_prompt_chars) skip recall by design. Sameuser_idacross turns is required. - Writes seem missing after a short script exits — sending is fire-and-forget; call
memory.flush()before exiting so in-flight sends land.
Supported versions
| Surface | Requires | Verified on |
|---|---|---|
ag2-memorysync 1.1.0 | The classic ConversableAgent framework (pip install autogen, Python 3.10+). The rewritten pip install ag2 v1 has no hook system yet — a future adapter target. | 31 CI checks against the latest ag2-classic: REAL two-agent initiate_chat conversations (user turns sent with seeds, assistant replies never sent, context reaching the LLM but never the transcript), the zep-ag2 double-store reproduction (each utterance handled once), a UserProxyAgent + assistant chat with the capability on both sides (the person’s words sent once as the user’s turn), a chat driven from inside asyncio.run() (no deadlock), the recall budget vs a slow backend, and both monthly-quota server modes |
How it compares
| Mem0 | Zep (`zep-ag2`) | Supermemory | MemorySync | |
|---|---|---|---|---|
| AG2 adapter exists | ✗ docs show an AutoGen-0.2 recipe with placeholder model names | ✓ (5 weeks old) | ✗ nothing | ✓ |
| Multi-agent duplication | — | ✗ documented bug: every utterance stored twice with conflicting roles | — | ✓ dedup registry + seeds — each utterance handled once, and an agent’s reply is never sent as the person’s words, with a reproduction test |
| Sync→async bridge | — | ✗ per-call loop spin; documented deadlock caveat under asyncio | — | ✓ one persistent bridge loop; asyncio-driven chats pass, by test |
| Recall latency budget | — | ✗ none | — | ✓ hard 1.2s default |
| Framework coupling | — | ✗ pins ag2<1 — breaks on the v1 rewrite | — | ✓ zero framework dependencies (duck-typed attach) |
| Extra LLM calls per turn | Teachability needs an analyzer-LLM call every turn | none | — | none |