Blog · Launch

Resume agents without snapshots

· Anup Talwalkar · Polign

The problem

Our agents kept losing their work. One would be halfway through a long task, the pod would get rescheduled or the spot instance reclaimed, and it would come back with no idea what it was doing or how far it got. New image rollouts and scale-to-zero did the same.

Snapshotting the pod with CRIU or a volume snapshot worked, but it brought back the old image and model after an upgrade, losing anything done since the last snapshot, and gave the process memory nobody could read.

Introducing Recall Resume()

Most of what a snapshot saves can be recovered some other way. The image rebuilds the OS, the repo and tickets live elsewhere, and caches can be thrown away. What's left is the agent's own state: its goal, plan, progress, decisions and last few turns. That is small.

Resume keeps only that. The agent writes it to Polign Recall as it works, Recall stores it in your bucket, and the next process to open the same agent id starts from it.

Container snapshot Recall resume
What is saved The whole process: memory, files, connections The agent's state, pointers to its work, its turns
When Each time you take one Every turn
Size in our test 15 MB 134 KB
After an upgrade Old image and model come back New image and model carry on
Readable No, and it holds your secrets Yes

How it works

What gets recorded

  • Working state: goal, plan, progress, decisions and open questions. The model updates it with a tool call when its plan changes.
  • Pointers: where the work lives, like a git branch, a bucket object or a ticket.
  • Turns: every turn, word for word.

One process at a time

Resume takes a lease on the agent id so two processes never work for the same agent. The lease is a chain of objects in the bucket written with conditional writes, so only one process can claim it. A process that gets taken over sees lease_lost instead of carrying on. A clean exit hands the lease over right away; after a crash it expires in 60 seconds by default.

The briefing

On restart, the model gets a briefing of up to 8,000 tokens: working state and pointers first, then recent turns and relevant memories. Tool outputs too big to fit are stored whole and fetched by reference when needed. Voice agents can't wait for a lease to expire, so the LiveKit package reads the records right away and takes the lease in the background.

What about checkpointers and durable execution?

Resume sits beside LangGraph's checkpointer, not in place of it. A checkpointer like PostgresSaver saves the graph's state so a thread can continue, but it doesn't stop two processes from working the same thread at once, and it hands the model the whole message history back. Resume adds the lease, so only one process owns the agent, and a briefing sized to a token budget instead of a full replay.

Durable execution tools like Temporal, DBOS and Restate replay a workflow's steps exactly, so the code has to stay compatible with its history. Resume doesn't replay anything. A new process, image or model reads the records and carries on.

Configuring it

This uses LangGraph. The LangGraph page has every option, and the LiveKit page covers voice agents.

Step 1: install

pip install recall-langgraph langchain-openai

This also installs the polign_db server and CLI, so there is nothing else to download.

Step 2: wrap the agent

from langchain_openai import ChatOpenAI
from langgraph.prebuilt import create_react_agent
from recall_langgraph import RecallResume, recall_tools

with RecallResume("billing-migrator", local_dir="./recall-data") as resume:
    agent = create_react_agent(
        ChatOpenAI(model="gpt-4.1"),
        tools=[*your_tools, *recall_tools(resume)],
        prompt="Keep your working state current with update_working_state when your plan changes.",
        pre_model_hook=resume.pre_model_hook,
        post_model_hook=resume.post_model_hook,
    )
    agent.invoke({"messages": [("user", "Carry on with the migration.")]})

The with block takes the lease and hands it back on exit. The hooks give the model its briefing and record each reply. recall_tools adds update_working_state, milestone, fetch_output and set_pointer.

Step 3: kill it and run it again

python agent.py    # kill -9 it partway
python agent.py    # picks up where the first run stopped

Within 60 seconds you'll get lease_held, since the killed process still holds the lease. Give it a minute.

Step 4: keep the records in your bucket

local_dir keeps records on one machine. If your agents move around, run polign-server on your S3, GCS or Azure bucket (Get started) and point resume at it:

RecallResume(
    "billing-migrator",
    env={"POLIGN_URL": "http://memory.internal:23000", "POLIGN_API_KEY": key},
)

What we measured

Does a resumed agent finish?

Before shipping this, I wanted a simple bar. If I kill an agent partway through, it should still finish the job, and resuming shouldn't cost more than 10% extra tokens over a run that never got interrupted.

I gave the agent a code migration: 16 files in a 21-file repo that need the change, and 5 decoys it has to leave alone. It ran on gpt-4.1-mini at temperature 0. I let 10 runs go start to finish, and killed 20 others with SIGKILL, anywhere from the second model call to 90% of the way in. Each kill lands right after a model reply, before its tool calls run, so the reply is lost but no tool is left half done. Mid-tool kills are the reconcile case under Limitations. Each killed agent came back in a fresh process.

Killed agents that finished 20 of 20
Extra tokens 4.4% on average (bar: under 10%)
Mean tokens per run 105k resumed, 101k uninterrupted

One uninterrupted run hit an OpenAI rate limit and was left out. Harness: python/recall-langgraph/bench/gate.py.

Against a CRIU checkpoint

Then I wanted to see how it holds up against CRIU, which is what most people reach for to checkpoint a container. I also added an upgrade, since that's where snapshots usually hurt. agent:v1 runs on gpt-4.1-mini and gets killed at model call 20. It should come back as agent:v2, a new image on gpt-4.1. I ran this on Podman 4.9 with runc, CRIU 4.2.1 and Ubuntu 24.04 on GCP.

CRIU checkpoint Recall resume
Came back as v1 on gpt-4.1-mini v2 on gpt-4.1
Saved state 15 MB of memory, API key included 134 KB of records
Work redone 8 model calls None
Model calls, tokens 45, 93k 40, 96k
Setup changes needed 3: runc instead of crun, --tcp-established, container-managed DNS None

Both agents finished and passed the tests. But CRIU came back as the old agent on the old model, and its 15 MB checkpoint had the API key sitting in it. Getting CRIU to work at all took three changes to a plain setup. Recall came back as the upgraded agent, picked up at the next file, and redid nothing. Token use ended up about the same, because every call after the restart carries the briefing. The scripts and recordings are in examples/resume-vs-criu.

Replay of the recorded run

Recorded run · replay

The same coding agent migrates 16 files, one tool call at a time. It is killed at call 20, then upgraded. Left: recovered from a container snapshot. Right: resumed from its Recall records.

Files migrated 0 / 16Saved state none
Waiting to start
Files migrated 0 / 16Saved state 0 turns
Waiting to start

Limitations

  • No reconcile step yet. If a process dies after a tool ran but before its result was recorded, the resumed agent may run it again. Stick to read-only, idempotent or easy-to-check actions until reconcile hooks land. The tests above don't cover this case.
  • Unwritten reasoning is lost. The briefing only has what the agent recorded.
  • Small sample. One task, one model family, one CRIU run.

Availability

Available today on PyPI:

  • polign_db 0.8.1: lease API and polign mcp -agent.
  • polign-recall 0.4.1: Python client, client.resume(...).
  • recall-langgraph 0.1.1: LangGraph agents.
  • recall-livekit 0.3.0: LiveKit voice agents, RecallAgent(..., resume="room-name").

The Python packages are Apache-2.0 in github.com/Polign/polign. polign_db is free to self-host.

Try it

Running long-lived agents? We'll help you wire resume into one, in your own cloud.