Get 90% of memory questions right on 4% of the tokens.
On 267 user-fact questions from LongMemEval, Recall 0.12 with an embedding model answered 89.9% with a frontier model, sending about 4,400 tokens instead of the whole 104,000-token history. Keyword search answered 83.9%.