NeMo Agent Toolkit Memory
Give NVIDIA NeMo Agent Toolkit workflows long-term memory with one YAML block. Installing the package registers a real MemoryEditor plugin: the toolkit’s built-in memory tools — or the fully automatic auto_memory_agent — recall per-fact scored memories under a hard latency budget and learn new facts from the user’s turns as the workflow runs.
How it works
- Install
nat-memorysyncfrom PyPI — this registers_type: memorysync_memorywith the toolkit automatically. - Add a
memory:block to your workflow YAML naming that type. - Wire the toolkit’s built-in
add_memory/get_memorytools to it — or wrap your workflow inauto_memory_agentfor zero-tool automatic memory. - Done — the workflow saves facts and recalls them in later conversations. No custom code.
Before you start
- A MemorySync API key. Create one in the dashboard under Settings → API Keys, inside a project.
- Python 3.11+ with
nvidia-nat-core1.5 or newer (the toolkit’s own floor). - Set the key as an environment variable so your workflow YAML stays credential-free:
set MEMORYSYNC_API_KEY=ms_...
Install
One package:
pip install nat-memorysync
That’s the whole plugin story: installing registers _type: memorysync_memory through the toolkit’s nat.components entry point. The toolkit discovers it automatically — nothing to import, nothing to submit anywhere.
Quickstart
Step 1 — declare the memory block
In your workflow YAML, add a memory: section. The name (saas_memory here) is yours to choose; the _type is what the plugin registered. The API key comes from the MEMORYSYNC_API_KEY environment variable.
memory:saas_memory:_type: memorysync_memory
Step 2 — give the agent the memory tools
Wire the toolkit’s built-in add_memory and get_memory tool functions to your memory block, and hand them to the agent. The tool descriptions tell the model when to save and when to recall.
functions:add_memory:_type: add_memorymemory: saas_memorydescription: Save any user preference or fact for later conversations.get_memory:_type: get_memorymemory: saas_memorydescription: Recall previously saved user preferences and facts.workflow:_type: react_agenttool_names: [add_memory, get_memory]llm_name: my_llm
Fully automatic memory — no tools at all
Prefer memory that just happens? The toolkit ships an auto_memory_agent wrapper (in nvidia-nat-langchain) that sends every user message to fact extraction and enriches every prompt automatically — no tools, no prompt changes. Assistant replies are not stored as memories, so they are not sent. Point it at the same memory block:
memory:saas_memory:_type: memorysync_memoryworkflow:_type: auto_memory_agentaugmented_fn: my_actual_workflowmemory: saas_memory
In this mode search runs on every response, which is exactly why the hard recall budget matters: a slow lookup fails open to “no memories this turn” instead of stalling the workflow. Writes are fail-open too, and deterministic idempotency seeds let the server recognise a RetryMixin retry and extract the turn once.
Using it from Python (Builder API)
Everything the YAML does is also available programmatically — useful in tests and embedded workflows. This is the same editor object the tools drive:
from nat.builder.workflow_builder import WorkflowBuilderfrom nat_memorysync import MemorySyncMemoryConfigasync with WorkflowBuilder() as builder:await builder.add_memory_client("saas_memory", MemorySyncMemoryConfig())editor = await builder.get_memory_client("saas_memory")items = await editor.search("seat preference", top_k=5, user_id="customer-42",)for item in items:print(item.memory, item.similarity_score)
Search returns one MemoryItem per fact with similarity_score populated — not a joined text blob, and never with the scores thrown away.
Choosing a user_id
The user_id is a stable string identifying whose memories these are: your app’s internal user ID, an email address, or a UUID. The toolkit’s memory tools pass it per call, and every stored row is keyed by it — one customer’s memories can never reach another’s workflow. Use the same value across conversations or recall returns nothing.
Deletes that cannot nuke a customer
remove_items refuses to guess. Called with no arguments it raises a ValueError (the in-repo Mem0 editor silently does nothing — you believe data was deleted and it was not). Deleting a whole user requires an explicit opt-in:
| Call | What is deleted |
|---|---|
remove_items() | Nothing — raises ValueError (refuses to guess) |
remove_items(user_id="u1") | The facts extracted from this adapter’s current session for u1 only |
remove_items(user_id="u1", memory_id="m1") | One specific memory row |
remove_items(user_id="u1", scope="user") | All of u1’s memories — explicit opt-in |
Configuration
| YAML field | Default | Meaning |
|---|---|---|
api_key | MEMORYSYNC_API_KEY env var | Keep it in the env var so workflow YAML stays credential-free |
base_url | https://api.memorysync.io | Override for staging |
project_id | — | Optional X-Project-ID header |
top_k | 5 | Default memories per search |
recall_timeout | 1.2 | Hard recall budget in seconds |
min_query_chars | 8 | Skip recall for trivial queries |
source | nat | Source label on the facts this client’s writes produce |
num_retries, retry_on_status_codes, … | toolkit defaults | RetryMixin knobs — safe here because a retried turn is recognised and extracted once |
Quotas and plan limits
Hitting a monthly plan limit never breaks a workflow. On free and paid plans, over-limit writes are accepted without storing and reads return empty — the agent keeps working. Evaluation keys instead surface a truthful 429, so you find out during testing, not in production.
Troubleshooting
- `_type: memorysync_memory` not found — the plugin isn’t installed in the same environment as the toolkit:
pip install nat-memorysync, then restart the workflow. - `search()` raises `ValueError` about `user_id` — that’s deliberate: the editor names the missing kwarg instead of the in-repo editors’ bare
KeyError. Passuser_id="..."on every search. - Recall returns nothing right after storing — extraction is asynchronous; give it a moment. Also confirm store and recall use the same
user_id. - A turn ran without memories once — that’s the latency budget working: a slow lookup is skipped rather than delaying the turn. The next turn recalls normally.
- `remove_items()` raised instead of deleting — also deliberate: it refuses to guess the blast radius. Say what to delete (see the deletes table).
Supported versions
| Surface | Requires | Verified on |
|---|---|---|
nat-memorysync 1.1.0 | nvidia-nat-core 1.5+ (Python 3.11+, the toolkit’s own floor) | 31 CI checks against the latest nvidia-nat-core: the REAL WorkflowBuilder round-trip and NVIDIA’s real add_memory/get_memory tool functions driving this editor end to end — user-turn-only sending (assistant, system and tool turns never sent), per-fact scored results, caller-mutation safety, the loud ValueError contracts, session-scoped vs scope="user" deletes, bleed-safe user_id keying, the 1.2s budget, seed idempotency, and both monthly-quota server modes |
How it compares
The toolkit’s own Mem0 and Zep editors live in NVIDIA’s repo as example code — not installable packages — and both have sharp edges this editor is engineered against:
| Mem0 (in-repo editor) | Zep (in-repo editor) | Supermemory | MemorySync | |
|---|---|---|---|---|
| Packaging | ✗ example code in NVIDIA’s repo | ✗ example code in NVIDIA’s repo | — nothing at all | ✓ pip-installable entry-point plugin |
search() without user_id | ✗ bare KeyError | — thread-scoped | — | ✓ ValueError naming the kwarg |
| Multi-user isolation | per-call user_id | ✗ everyone shares `"default_zep_thread"` when no conversation id is set | — | ✓ rows always keyed by item user_id — bleed impossible |
Your metadata dict after add_items | ✗ mutated — keys popped out | ✓ | — | ✓ untouched (copy-first, tested) |
| Search result shape | ✗ items, scores discarded | ✗ one joined text blob | — | ✓ one MemoryItem per fact, similarity_score populated |
remove_items() with no kwargs | ✗ silent no-op | deletes the current thread | — | ✓ raises — refuses to guess |
| Delete blast radius | ✗ whole user | ✗ whole thread | — | ✓ session-scoped by default; whole user is an explicit scope="user" opt-in |
| Slow or down backend | ✗ blocks the turn | ✗ blocks the turn | — | ✓ 1.2s recall budget, fail-open both directions |