AWS Bedrock Memory
A runnable memory reference architecture for AWS Bedrock: a memory-augmented Converse loop, plus the FIRST Bedrock Agents action-group memory integration from any vendor — recall under a hard budget, idempotent persistence, and a multi-tenant identity ladder.
Two runnable patterns
- Pattern 1 — the Converse loop. For apps that call
bedrock-runtimedirectly: relevant memories are recalled before eachconversecall and injected as a system block; after the reply, the user’s message is sent for fact extraction. - Pattern 2 — the Agents action group. For managed Bedrock Agents: deploy one Lambda and any agent gains four curated memory functions it can call on its own.
Both ride the same MemorySync account — a fact learned in the Converse loop is recallable by a Bedrock Agent, a LangChain app, or a voice agent.
Before you start
- A MemorySync API key. Create one in the dashboard under Settings → API Keys, inside a project.
- Python 3.10+ with
boto31.34+ (the Converse API), and AWS credentials with Bedrock model access. - For Pattern 2 only: the SAM CLI and permission to deploy a Lambda.
Install
pip install bedrock-memorysync
Pattern 1 — the Converse loop
Step 1 — build the loop
You own the boto3 client — region, credentials, retry policy stay yours:
import boto3from bedrock_memorysync import MemoryConverseLooploop = MemoryConverseLoop(bedrock_client=boto3.client("bedrock-runtime"),model_id="anthropic.claude-3-haiku-20240307-v1:0",user_id="customer-42", # requiredsystem_prompt="You are a helpful travel assistant.",)
Step 2 — chat, then prove the memory
loop.chat("I always prefer window seats on long flights")# a new session, days later:loop.new_session("support-2")loop.chat("which seat should I book for the Oslo flight?") # remembers
Before each converse call the user’s relevant memories are recalled under a hard 1.2s budget and injected as a system block. After the reply, the user’s message goes to fact extraction with a deterministic seed: MemorySync stores the durable facts in it (preferences, personal details, decisions), not the text itself, and recognises a retried turn so it is processed once. Assistant replies are not stored as memories, so the loop does not send them. A slow or down memory service degrades to a memory-free turn — memory can never delay or break the model call.
Pattern 2 — the Bedrock Agents action group
No memory vendor ships this — this is the first.
Step 1 — deploy the Lambda
The package includes a SAM template:
sam deploy --guided --parameter-overrides MemorySyncApiKey=ms_...
Step 2 — attach the action group
Use the importable function schema — no hand-written JSON:
from bedrock_memorysync import FUNCTION_SCHEMAbedrock_agent.create_agent_action_group(agentId=agent_id, agentVersion="DRAFT",actionGroupName="memorysync-memory",actionGroupExecutor={"lambda": function_arn},functionSchema={"functions": FUNCTION_SCHEMA},)
What the agent can now do
| Function | What it does |
|---|---|
remember | Send what the user said to fact extraction — only the durable facts in it are stored; agent retries are recognised and not processed twice. An optional role (user, the default, or assistant) says who said it: an assistant turn is not stored as a memory, but like every call it counts one add |
recall_context | A prompt-ready block of the user’s relevant memories |
search_memories | Scored JSON list (≤25 results) |
forget_memory | Delete exactly ONE memory by id, loudly — a failed delete never looks deleted |
remember tells the agent what happened: status: "accepted" with stored_as: "facts" when the text went to fact extraction, or status: "skipped" with the reason in processing_status and message — an assistant turn (not stored as a memory), nothing worth remembering, the monthly quota, or already_exists: true when the same text was already sent. It never reports a memory id: facts are extracted asynchronously and get their own ids.
Who the memories belong to
Identity ladder (multi-tenant safe): sessionAttributes.memorysync_user_id (your app sets it when invoking the agent) → a user_id parameter → the Bedrock sessionId. The model can never talk its way into another user’s memory. Every failure returns a structured responseState: FAILURE body the agent can react to — never an unhandled Lambda error surfaced as a wall of text.
Quotas and plan limits
Hitting a monthly plan limit never breaks a turn. Over-limit writes are accepted without storing and recalls return empty — Converse loops keep answering and agents keep running. Evaluation keys instead surface a truthful 429, so limits show up in testing, not production.
Troubleshooting
- `ValidationException` from Bedrock — your boto3 predates the Converse API. Upgrade to 1.34+.
- A turn ran without memories — that’s the 1.2s fail-open budget: recall is skipped rather than delaying the model call. Check connectivity if it persists.
- Agent retried and you feared duplicates — deterministic seeds absorb retries: the same text is recognised server-side (
already_exists: true) and not extracted twice. - Lambda returns `responseState: FAILURE` — the body carries the reason (auth, quota, bad id). The agent sees it as structured output and can react; check the Lambda logs for the HTTP status.
Supported versions
| Surface | Requires | Verified on |
|---|---|---|
bedrock-memorysync 1.1.0 | Python 3.10+; boto3 1.34+ (Converse API); Lambda runtime python3.12 | 41 CI checks: the REAL boto3 client under botocore’s Stubber with verbatim request assertions (the recall system block, multi-turn accumulation, only the user turn sent after the reply, the no-memory-block shape when recall times out) plus the documented Bedrock Agents event shapes — identity ladder, idempotent retries, honest remember reports (accepted, assistant turn not stored, nothing to remember, quota), FAILURE response states, result caps, body truncation, the no-delete_all schema assertion, extracted facts searched and recalled like any other with family-sweep deletes, org-connector-row filtering, honest not-found deletes, and quota modes |
How it compares
| AWS AgentCore Memory | Mem0 | Zep / Supermemory / Letta | MemorySync | |
|---|---|---|---|---|
| Bedrock Agents action group | n/a (separate runtime) | ✗ none | ✗ nothing at all | ✓ first and only |
| Setup | control-plane resource + strategy activation polling | ✗ OSS sidecar: self-hosted Mem0 + an OpenSearch domain you operate | — | ✓ one API key |
| Extraction control | ✗ black-box strategies | your own infra to tune | — | ✓ managed pipeline with server-side gating |
| Cross-platform recall | ✗ AWS-only (actorId inside one account) | ⚠ self-hosted API | — | ✓ the same memories in LangChain, Flowise, Dify, voice agents… |
| Portability | ✗ total AWS lock-in | ⚠ | — | ✓ any cloud + Bedrock |