Microsoft Agent Framework Memory
Give Microsoft Agent Framework agents long-term memory in three lines. The provider rides the framework’s own ContextProvider seam: relevant memories are recalled into the instructions layer before every run, and the user’s messages are sent for fact extraction automatically afterwards — without ever being able to crash the agent.
How it works
- Install
agent-framework-memorysyncfrom PyPI. - Create a
MemorySyncContextProviderwith your API key and the end user’s id. - Pass it to your
Agentviacontext_providers=[provider]. - Done — before every run the agent recalls what it knows about this user; after every run the user’s messages are sent for fact extraction, and the agent’s replies are not stored. No per-turn code.
Before you start
- A MemorySync API key. Create one in the dashboard under Settings → API Keys, inside a project.
- Python 3.10+ with
agent-framework-core1.8 or newer.
Install
One package:
pip install agent-framework-memorysync
Quickstart
Step 1 — create the provider
The provider is where you say whose memories these are. user_id is required; session_id separates conversations — every distilled fact is tagged with it.
from agent_framework_memorysync import MemorySyncContextProviderprovider = MemorySyncContextProvider(api_key="ms_...", # or set MEMORYSYNC_API_KEY in the environmentuser_id="customer-42", # required — the end user these memories belong tosession_id="support", # optional — tags this conversation's facts)
Step 2 — attach it to your agent
Hand the provider to the agent. The framework now calls it around every run: recall lands in the instructions layer (never disguised as a fake user message), and the user’s messages are sent to MemorySync after the run completes. The server extracts the durable facts they contain and stores only those; filler such as “ok” yields none.
from agent_framework import Agentagent = Agent(client=..., context_providers=[provider])result = await agent.run("Which seat should I book?")
Per-run identity — the documented gap, solved
Both competitor adapters bind identity at construction and document it as a framework limitation: multi-tenant servers must build one provider per user. This provider adds a three-rung ladder, resolved fresh on every run and pinned in the provider’s session-state slice so recall and capture can never diverge within a run:
| Rung | Mechanism | Use case |
|---|---|---|
| 1 | await agent.run(..., options={"memorysync_user_id": user}) | Multi-tenant servers: one agent, per-request identity |
| 2 | user_id_resolver=lambda session: ... | Identity derived from your session store |
| 3 | Constructor user_id | Single-user agents and simple deployments |
Built-in safety guarantees
| Guarantee | How |
|---|---|
| A run is never stalled | Recall waits at most recall_timeout (default 1.2s; the very first call of a fresh process gets a one-time 3s grace for connection setup); on a miss the agent runs without memories |
| A memory outage never crashes the agent | after_run catches everything and logs — the run’s success is never converted into a failure (the Mem0 provider does exactly that) |
| Tool loops capture once | One send per user turn, not per model round-trip |
| A retry is not extracted twice | Deterministic seeds from role + session + content hash, recognised server-side |
| Only the user’s words are sent | User input messages go as plain text with role: "user"; the agent’s replies are not sent |
| Sessions stay serializable | The provider writes only JSON-native values to its state slice — AgentSession.to_dict() keeps working |
Configuration
| Parameter | Default | Meaning |
|---|---|---|
user_id | — (required) | End user the memories belong to |
session_id | default | Session scope agent-framework::<session> — tags this conversation’s facts |
top_k | 5 | Memories considered per run |
recall_timeout | 1.2 | Hard recall budget in seconds |
min_prompt_chars | 8 | Skip recall for trivial prompts |
context_template | built-in | {context} placeholder; brace-safe .replace rendering |
capture | True | Send the run’s user messages for fact extraction after each run |
expose_search_tool | False | Register search_memory + save_memory tools each run |
user_id_resolver | None | Callable for dynamic identity |
Choosing a user_id
The user_id is a stable string you pick to identify whose memories these are: your app’s internal user ID, an email address, or a UUID. Use the same value across runs or recall returns nothing. One user_id drives both storage and retrieval — there is no separate search_user_id to forget.
Quotas and plan limits
Hitting a monthly plan limit never breaks a run. On free and paid plans, over-limit writes are accepted without storing and reads return empty — the agent keeps working. Evaluation keys instead surface a truthful 429, so you find out during testing, not in production.
Troubleshooting
- `ModuleNotFoundError: agent_framework` — the framework itself isn’t installed:
pip install agent-framework. - Recall returns nothing right after storing — extraction is asynchronous; give it a moment. Also confirm store and recall use the same
user_id. - The agent ran without memories once — that’s the latency budget working: a slow lookup is skipped rather than delaying the run. The next run recalls normally.
- Multi-tenant identity — don’t build one provider per user; pass
options={"memorysync_user_id": ...}per run (rung 1 of the ladder). - A crash after a successful run — not from this provider: capture is entirely fail-open by design. Check the model client.
Supported versions
| Surface | Requires | Verified on |
|---|---|---|
agent-framework-memorysync 1.1.0 | agent-framework-core 1.8+ (Python 3.10+) | 26 CI checks against the latest agent-framework-core: real SessionContext/AgentSession objects and a REAL Agent run via a stub chat client — instructions-layer injection, the two-phase recall budget vs a slow backend, the per-run identity ladder, fail-open once-per-turn capture (the Mem0 crash regression), user-message-only capture (the agent’s replies never sent), seed idempotency, JSON-safe session state, and both monthly-quota server modes |
How it compares
| Mem0 (`agent-framework-mem0`) | Zep (`zep-ms-agent-framework`) | Supermemory | MemorySync | |
|---|---|---|---|---|
| Injection layer | ✗ fabricates a role="user" message the model believes the user wrote | ✓ instructions | — no adapter at all | ✓ instructions |
| Recall latency budget | ✗ none | ✗ none | — | ✓ hard 1.2s default |
| Capture failure | ✗ after_run errors crash the agent after a successful run | swallowed | — | ✓ entirely fail-open, logged |
| Scope model | ✗ storage ≠ retrieval scopes; forget search_user_id and the agent is silently memoryless | single | — | ✓ one user_id drives both |
| Per-run identity | ✗ construction-only | ✗ construction-only (documented) | — | ✓ run options → resolver → constructor |
| Write dedup | ✗ re-adds every turn | ✗ | — | ✓ deterministic idempotency seeds |
| Release status | △ beta | 0.2.1 | — | ✓ stable 1.x |
| Python floor | 3.10 | ✗ 3.11 only | — | ✓ 3.10 (matches the framework) |