Corpus as of 2026-09-12 · run2 · 995 rows · 973 distinct artifacts
Agent Memory Benchmark Catalogue
Every long-term-memory and long-horizon benchmark found for LLM agents, sorted
newest first, then by citations. Each row carries a description in two voices — the
authors' own verbatim sentence where one exists, then a plain-language summary. Where a source
never characterises its own artifact, the cell says so rather than inventing one.
Full data with evidence spans: data/catalog.csv.
External memory records whether the benchmark
involves a store outside the model's context window — files, a database, a vector store, a
scratchpad — and whose it is: the agent's own memory, the task environment's state, or a fixed
corpus it retrieves from. A long context window does not count.
Every row here has been verified. A cheap first pass read all of them, then a stronger pass re-read every one and overturned 32% of what it checked; ten hand-read gold items anchor both. What you see is the verified reading.