Corpus as of 2026-09-12 · run2 · 995 rows · 973 distinct artifacts

Agent Memory Benchmark Catalogue

Every long-term-memory and long-horizon benchmark found for LLM agents, sorted newest first, then by citations. Each row carries a description in two voices — the authors' own verbatim sentence where one exists, then a plain-language summary. Where a source never characterises its own artifact, the cell says so rather than inventing one. Full data with evidence spans: data/catalog.csv.

External memory records whether the benchmark involves a store outside the model's context window — files, a database, a vector store, a scratchpad — and whose it is: the agent's own memory, the task environment's state, or a fixed corpus it retrieves from. A long context window does not count. Every row here has been verified. A cheap first pass read all of them, then a stronger pass re-read every one and overturned 32% of what it checked; ten hand-read gold items anchor both. What you see is the verified reading.

Benchmark Year What it is — authors' words, then ours Construct One example is… External memory Defines "long horizon" Cites