Blog · Benchmark

Recall on LongMemEval: accuracy, tokens and latency

October 9, 2026 · Anup Talwalkar · Polign

Agent memory is worth having when it puts the right facts in front of the model for a fraction of the tokens of the whole conversation, and answers fast enough to sit on every turn. We measured Recall 0.12 on that: how often the answer is right, how many tokens it sends, how long a read takes, and what it costs to write a user's history into memory.

89.9%
answers correct with a frontier model
vs 83.9% keyword search
96%
fewer tokens than the full history
about 4,400 vs 103,839
0.9 ms
exact recall, p50
4.5 ms for a search
$0.036
to extract one user's history
48 sessions, gpt-4o-mini
What we measured

The user-fact questions of LongMemEval_S: 267 of its 500 questions, each asked over about 104,000 tokens of chat history. They ask what the user said, prefers, changed, and when. We left out questions about what the assistant said, questions that combine facts across many sessions (reported separately below), and 44 questions whose answer comes from a session dated after the question was asked. Recall answers as of the question date, so it correctly cannot see those.

Accuracy depends on the model reading the memory

Every method hands the same kind of model a context and asks it to answer. We ran each one with two answering models: gpt-4o, which the benchmark was published with, and gpt-6-astra, a current frontier model. A gpt-4o judge scored every answer, the same judge for every row.

Answering model Keyword search (BM25) Recall Recall with embeddings Oracle
gpt-4o 73.8% 81.3% 85.0%
gpt-6-astra 83.9% 86.1% 89.9% 96.6%

Keyword search returns the 10 best-matching turns of the conversation. The oracle is handed only the sessions that contain the answer, so it is the ceiling for any retrieval method. "With embeddings" adds OpenAI text-embedding-3-small to Recall's word search.

RecallKeyword search and oracle

With gpt-4o, Recall answers 81.3% of the questions against 73.8% for keyword search, winning 27 questions and losing 7 (sign test p = 0.0008), close to the 85.0% gpt-4o reaches when handed only the right sessions. At that point the answering model is the limit.

A frontier model changes that. gpt-6-astra improves every method, keyword search most of all, and given the right sessions it answers 96.6%. Now retrieval is the limit. Recall with word search alone is only 2.2 points ahead of keyword search (11 won, 5 lost), which is not a significant difference. With an embedding model Recall reaches 89.9%, six points ahead (19 won, 3 lost, p = 0.0009), on the same 4,400 tokens per question. If a frontier model reads your agent's memory, run Recall with an embedding model.

By kind of question

Question Keyword search Recall Recall with embeddings Oracle
Something that changed
uses the current address, plan or owner, not the old one
93.5% 94.8% 94.8% 96.1%
When something happened
"what did I do two weeks ago?"
70.0% 76.7% 83.3% 93.3%
Knowing it does not know
says so instead of inventing an answer
73.3% 86.7% 86.7% 80.0%
Preferences
how the user likes things done
73.3% 70.0% 80.0% 100%
A fact stated once 95.7% 95.7% 97.1% 100%

gpt-6-astra answering, 267 questions. The 15 questions that test whether the model knows it does not know are also counted inside their own kind.

Recall is strongest where memory goes wrong in production: facts that changed, questions about time, and declining to answer when nothing was said. It refuses more reliably than the oracle because it only returns what was believed by the question date. The largest gaps to the ceiling are questions about time (83.3% against 93.3%) and preferences (80% against 100%). Both are retrieval misses, and they are where we are working next.

What changed in Recall 0.12

Two changes took Recall from 77.5% to 82.0% with gpt-4o, with no extra tokens and no extra latency:

Change Accuracy Time questions
Recall 0.10 77.5% 68.9%
A search that names a time looks there first
"two weeks ago", "last Saturday", "in January"
80.5% 76.7%
Every memory says how many days old it is
days_ago on each belief
82.0% 78.9%

gpt-4o answering. A later run of the released 0.12 binary retrieved exactly the same memories for all 267 questions and scored 81.3%: gpt-4o answered four of them differently. Runs with the same retrieval vary by about two questions.

The time window is the bigger of the two. Most of the misses before it were questions like "what gardening did I do two weeks ago?", where word search found gardening from any month. Recall now reads the phrase against the date the question is asked, searches that span first, and still returns everything else after it. Nothing is dropped. A caller can also set the window explicitly with observed_after and observed_before.

Tokens: what memory costs on every turn

The tokens memory adds to each request are a cost an agent pays again on every turn. Recall sends about 4,400 tokens per question, including the original sentence each fact came from. Sending the whole history is about 104,000.

Context per question Tokens Monthly input cost
Whole conversation history 103,839 $519,195
Keyword search, top 10 turns 5,068 $25,340
Recall with embeddings 4,450 $22,250
Recall 4,364 $21,820

Mean over the 267 questions. Monthly cost assumes 10,000 users each making 200 memory reads a month at $2.50 per million input tokens, and counts only the memory context, not the user's message or the reply.

Change the numbers to see the monthly input cost of the memory context at your own volume:

Latency

Recall runs beside your agent and keeps its memory in your own polign_db, so a read makes no call to a third-party service. We timed the released recall binary through its Python client:

Operation p50 p95 p99
Start the server and connect 5.6 ms
Remember a fact 9.0 ms 14.0 ms 18.7 ms
Read right after a write 2.6 ms 3.3 ms 3.4 ms
Exact recall
a known subject and predicate
0.9 ms 2.8 ms 3.5 ms
Search recall 4.5 ms 11.4 ms 12.9 ms

300 rounds on the managed local store that recall setup creates, on a laptop.

Bar runs from p50 to p95, in milliseconds

In the benchmark, where each user's store holds about 550 memories and every result comes back with its source conversation, a recall took 83 ms at the median and 138 ms at p95. The time-window search added nothing measurable. An embedding model adds the time to embed the question: about 350 ms per read through a hosted embedding API, which a local embedding model would avoid.

A write is readable on the very next call, so an agent that records a correction can rely on it in the same conversation. With memory in object storage, the bucket sets the floor:

Where memory lives Warm recall One remember Read right after a write
S3 Express One Zone, same zone 18 to 25 ms 15 to 21 ms 45 to 71 ms
S3 Standard, same region 48 to 72 ms 26 to 49 ms 152 to 273 ms
GCS Standard, same region 82 to 136 ms 57 to 96 ms 256 to 547 ms

Measured September 18, 2026 on a small store, with the server in the bucket's region. Ranges span polign_db's two serving modes.

Writing a user's history into memory

A model has to read new conversation to turn it into typed facts. Recall makes about one extraction call per session, and every fact keeps a link to the sentence it came from. You can also skip extraction: Recall then keeps each turn as a note and searches those.

Write path Model tokens per user history Cost per user history
Typed facts linked to their source 157,000 in, 20,000 out $0.036
Notes only, no model 0 $0

One user history is a LongMemEval_S conversation: 48 sessions, 494 messages, 490 KB. Extraction by gpt-4o-mini at its list price. Storage is your own bucket; an idle Recall node costs a few dollars a month.

How we ran it
  • Data. LongMemEval_S, cleaned 2025/09 release. The 267 questions are its single-session-user, single-session-preference, knowledge-update and temporal-reasoning questions, minus the 44 whose answer session is dated after the question.
  • Recall. Recall 0.12.0 from PyPI with polign_db 0.14.0. Each conversation was written session by session in date order with its session date, facts extracted by gpt-4o-mini with their source sentence, and every turn kept as a note. The stores were written once by a development build of the same engine and read by the released binary. Each question asked Recall for 10 results as of the question date, with sources.
  • Keyword search. BM25 over every user turn and the reply after it, top 10.
  • Prompts and scoring. LongMemEval's own reader and judge prompts, byte for byte. Judge: gpt-4o-2024-08-06. Significance: two-sided sign test on the questions where the two methods disagree.
  • Embeddings. OpenAI text-embedding-3-small at 512 dimensions, fused with word search.
  • Latency. 300 rounds per operation through the Python client against the managed local store, on an Apple silicon laptop.
  • Reproduce. The harness is in the Recall repository under eval/longmemeval; run.py --subset user-facts selects these questions.

Try it

pip install polign-recall brings Recall and the database it keeps memory in. Get started saves and corrects a fact in a few minutes, and Claude Code takes one command.