Reports

Research reports and the data behind them.

Agent Memory Benchmarks

Corpus survey · 2026-09-12

973benchmarks 3,314screened 86%read at full text 62%never define "long horizon"

Every long-term-memory and long-horizon benchmark we could find for LLM agents, with each one described in the authors' own words and ours — and a measurement of how little the field agrees on what those words mean.