Resume agents without snapshots
The problem
Our agents kept losing their work. One would be halfway through a long task, the pod would get rescheduled or the spot instance reclaimed, and it would come back with no idea what it was doing or how far it got. New image rollouts and scale-to-zero did the same.
Snapshotting the pod with CRIU or a volume snapshot worked, but it brought back the old image and model after an upgrade, losing anything done since the last snapshot, and gave the process memory nobody could read.
Introducing Recall Resume()
Most of what a snapshot saves can be recovered some other way. The image rebuilds the OS, the repo and tickets live elsewhere, and caches can be thrown away. What's left is the agent's own state: its goal, plan, progress, decisions and last few turns. That is small.
Resume keeps only that. The agent writes it to Polign Recall as it works, Recall stores it in your bucket, and the next process to open the same agent id starts from it.
| Container snapshot | Recall resume | |
|---|---|---|
| What is saved | The whole process: memory, files, connections | The agent's state, pointers to its work, its turns |
| When | Each time you take one | Every turn |
| Size in our test | 15 MB | 134 KB |
| After an upgrade | Old image and model come back | New image and model carry on |
| Readable | No, and it holds your secrets | Yes |
How it works
What gets recorded
- Working state: goal, plan, progress, decisions and open questions. The model updates it with a tool call when its plan changes.
- Pointers: where the work lives, like a git branch, a bucket object or a ticket.
- Turns: every turn, word for word.
One process at a time
Resume takes a lease on the agent id so two processes never work for the same agent. The lease is a
chain of objects in the bucket written with conditional writes, so only one process can claim it. A
process that gets taken over sees lease_lost instead of carrying on. A clean exit hands
the lease over right away; after a crash it expires in 60 seconds by default.
The briefing
On restart, the model gets a briefing of up to 8,000 tokens: working state and pointers first, then recent turns and relevant memories. Tool outputs too big to fit are stored whole and fetched by reference when needed. Voice agents can't wait for a lease to expire, so the LiveKit package reads the records right away and takes the lease in the background.
What about checkpointers and durable execution?
Resume sits beside LangGraph's checkpointer, not in place of it. A checkpointer like
PostgresSaver saves the graph's state so a thread can continue, but it doesn't stop two
processes from working the same thread at once, and it hands the model the whole message history
back. Resume adds the lease, so only one process owns the agent, and a briefing sized to a token
budget instead of a full replay.
Durable execution tools like Temporal, DBOS and Restate replay a workflow's steps exactly, so the code has to stay compatible with its history. Resume doesn't replay anything. A new process, image or model reads the records and carries on.
Configuring it
This uses LangGraph. The LangGraph page has every option, and the LiveKit page covers voice agents.
Step 1: install
pip install recall-langgraph langchain-openai
This also installs the polign_db server and CLI, so there is nothing else to download.
Step 2: wrap the agent
from langchain_openai import ChatOpenAI
from langgraph.prebuilt import create_react_agent
from recall_langgraph import RecallResume, recall_tools
with RecallResume("billing-migrator", local_dir="./recall-data") as resume:
agent = create_react_agent(
ChatOpenAI(model="gpt-4.1"),
tools=[*your_tools, *recall_tools(resume)],
prompt="Keep your working state current with update_working_state when your plan changes.",
pre_model_hook=resume.pre_model_hook,
post_model_hook=resume.post_model_hook,
)
agent.invoke({"messages": [("user", "Carry on with the migration.")]})
The with block takes the lease and hands it back on exit. The hooks give the model its
briefing and record each reply. recall_tools adds update_working_state,
milestone, fetch_output and set_pointer.
Step 3: kill it and run it again
python agent.py # kill -9 it partway
python agent.py # picks up where the first run stopped
Within 60 seconds you'll get lease_held, since the killed process still holds the lease.
Give it a minute.
Step 4: keep the records in your bucket
local_dir keeps records on one machine. If your agents move around, run
polign-server on your S3, GCS or Azure bucket (Get started)
and point resume at it:
RecallResume(
"billing-migrator",
env={"POLIGN_URL": "http://memory.internal:23000", "POLIGN_API_KEY": key},
)
What we measured
Does a resumed agent finish?
Before shipping this, I wanted a simple bar. If I kill an agent partway through, it should still finish the job, and resuming shouldn't cost more than 10% extra tokens over a run that never got interrupted.
I gave the agent a code migration: 16 files in a 21-file repo that need the change, and 5 decoys it has to leave alone. It ran on gpt-4.1-mini at temperature 0. I let 10 runs go start to finish, and killed 20 others with SIGKILL, anywhere from the second model call to 90% of the way in. Each kill lands right after a model reply, before its tool calls run, so the reply is lost but no tool is left half done. Mid-tool kills are the reconcile case under Limitations. Each killed agent came back in a fresh process.
| Killed agents that finished | 20 of 20 |
| Extra tokens | 4.4% on average (bar: under 10%) |
| Mean tokens per run | 105k resumed, 101k uninterrupted |
One uninterrupted run hit an OpenAI rate limit and was left out. Harness:
python/recall-langgraph/bench/gate.py.
Against a CRIU checkpoint
Then I wanted to see how it holds up against CRIU, which is what most people reach for to checkpoint
a container. I also added an upgrade, since that's where snapshots usually hurt. agent:v1
runs on gpt-4.1-mini and gets killed at model call 20. It should come back as agent:v2, a
new image on gpt-4.1. I ran this on Podman 4.9 with runc, CRIU 4.2.1 and Ubuntu 24.04 on GCP.
| CRIU checkpoint | Recall resume | |
|---|---|---|
| Came back as | v1 on gpt-4.1-mini | v2 on gpt-4.1 |
| Saved state | 15 MB of memory, API key included | 134 KB of records |
| Work redone | 8 model calls | None |
| Model calls, tokens | 45, 93k | 40, 96k |
| Setup changes needed | 3: runc instead of crun, --tcp-established, container-managed DNS |
None |
Both agents finished and passed the tests. But CRIU came back as the old agent on the old model, and
its 15 MB checkpoint had the API key sitting in it. Getting CRIU to work at all took three changes to
a plain setup. Recall came back as the upgraded agent, picked up at the next file, and redid nothing.
Token use ended up about the same, because every call after the restart carries the briefing. The
scripts and recordings are in
examples/resume-vs-criu.
Replay of the recorded run
The same coding agent migrates 16 files, one tool call at a time. It is killed at call 20, then upgraded. Left: recovered from a container snapshot. Right: resumed from its Recall records.
Limitations
- No reconcile step yet. If a process dies after a tool ran but before its result was recorded, the resumed agent may run it again. Stick to read-only, idempotent or easy-to-check actions until reconcile hooks land. The tests above don't cover this case.
- Unwritten reasoning is lost. The briefing only has what the agent recorded.
- Small sample. One task, one model family, one CRIU run.
Availability
Available today on PyPI:
- polign_db 0.8.1: lease API and
polign mcp -agent. - polign-recall 0.4.1: Python client,
client.resume(...). - recall-langgraph 0.1.1: LangGraph agents.
- recall-livekit 0.3.0: LiveKit voice agents,
RecallAgent(..., resume="room-name").
The Python packages are Apache-2.0 in github.com/Polign/polign. polign_db is free to self-host.
Try it
Running long-lived agents? We'll help you wire resume into one, in your own cloud.