Corpus survey · 2026-09-12 · 973 benchmarks

The Long-Horizon Register

A catalogue of every long-term-memory and long-horizon benchmark we could find for LLM agents — and a measurement of how little the field agrees on what those words mean.

973distinct benchmark artifacts
3,314candidates screened, 100% judged
86%read at full text, not abstract
62%never define "long horizon"
~80%estimated corpus recall

01 · The finding

The field builds long-horizon benchmarks without saying what "long horizon" is

Across 995 in-charter evaluation artifacts — every one of them — 619, or 62.2%, never pin down the construct they are built to measure. Not buried in an appendix: absent. We checked bodies precisely because we expected the definitions to be hiding there.

The absence has a characteristic shape. Papers reach for "long-horizon" and "long-term" as motivating vocabulary in the introduction, then define difficulty by some other quantity entirely — or by nothing. Four of the most-cited agent benchmarks in the corpus illustrate it:

The paper never uses "long-term" or "long-horizon" anywhere in the text; it frames its core difficulty via "extremely long contexts" — the size of the codebase a model must process (438K lines on average) — a long-CONTEXT-WINDOW framing.
Extractor note on SWE-bench · 3,690 citations · read at full text
The paper never uses "long-horizon" or "long-term"; it defines task difficulty entirely via annotator-counted "number of steps" and "number of tools" needed to solve a question.
Extractor note on GAIA · 1,188 citations · read at full text
The paper repeatedly names "long-horizon decision-making" and "long-term user-assistant interactions" as challenges posed by real computer environments … but never pins them down.
Extractor note on OSWorld · 1,147 citations · read at full text
The paper "formally defines" the general agent-control task … but never pins down what specifically makes tasks "long-horizon" with a turn/step criterion.
Extractor note on WebArena · 1,975 citations · read at full text

This is not a complaint about sloppy writing. It has a measurable consequence, which is the subject of the next section: when the term goes undefined, the quantity it stands for drifts across seven orders of magnitude without anyone in the field having to notice.

How this was produced Per-paper extraction over the core ring, one record per benchmark, each carrying a verbatim span and an explicit polarity — an "absent" answer is a claim with its own evidence (the span showing what the paper does instead), never a silent null. Five independent workers reported the same pattern unprompted. Limit: extractors used targeted section reads, so a definition stated only in a figure caption could be scored absent.
Frame — view: core ring (in + relevant) · substrate run2 @2026-09-12 · n = 995 rows (100% of the in-charter ring) · evidence 86.2% full text / 13.8% abstract-only

02 · How long is long

"Long horizon" spans 21 years and five orders of magnitude

Every benchmark that stated the length of a single example was recorded in its own units, with no conversion. Two ladders below plot those stated lengths on log axes — one for discrete units (steps, turns, sessions, tasks), one for wall-clock durations. Each mark is one benchmark's stated example length. No benchmark appears on both ladders unless it reported both.

Discrete units — steps, turns, sessions, tasks

592 stated lengths from benchmarks that count actions rather than time. Hover or tap a mark for the benchmark and what the number counts.

Wall-clock time — normalised to hours

161 stated durations, converted to hours only within the time family (minutes, days, months and years are exact multiples of an hour; nothing is converted across families).

The two ladders are not commensurable with each other, and that is the point. A benchmark whose "long horizon" is 30 interactions and one whose "long horizon" is ten expert days are both in the literature under the same phrase. Where papers do fix a threshold, the thresholds disagree by orders of magnitude:

We test long-term interactions of LLM Agents up to 30 interactions.
AgentBoard · 320 citations · operational threshold
We first partition the dataset based on the number of steps in the oracle policy … tasks needing more than 15 steps.
Our benchmark features long-horizon tasks that may require hours to days for a professional software engineer to complete.
It would be extremely long horizon (hundreds of thousands of steps) and difficult for an agent to tackle.
MineDojo · 614 citations · operational threshold

What a benchmark counts, when it counts anything

Share of 1,154 classified quantity mentions by unit family. A benchmark long on more than one axis contributes a mention to each, so these are mention shares, not benchmark shares.

How this was produced Units were open-coded per benchmark, then folded into eight families by a codebook derived bottom-up from the corpus's own codes; 55 mentions (4.5%) resisted classification and ship as a labelled tail in catalog-aggregates.json rather than hidden in an "other" bucket. Dataset size was excluded by instruction and recorded separately — "500 questions" is size, "across 500 sessions" is length.
Frame — core ring · n = 676 benchmarks stating any length · 1,154 classified mentions

03 · Who defines it

Of those that do define it, most do so sideways

How the 376 definitions are given

Counted only over benchmarks with a definition present.

Definition and length, by evidence depth

Percent of benchmarks carrying each field. The two strata are reported separately and never pooled — the gap measures evidence access, not benchmark quality.

109 benchmarks of 995 — 11.0% — state a definition outright. The remaining 267 that define the construct do it by operational threshold (145) or by contrast with a shorter setting (122). A reader looking for a citable definition of "long horizon" in this literature will find one in roughly one paper in ten.

The same gap shows up from a second direction. A controlled ablation in the corpus reports that memory benchmarks under-test the very construct they name:

A controlled ablation on the LoCoMo benchmark shows that Bayesian belief updating alone provides little benefit over naive last-write-wins because existing conversational memory benchmarks rarely contain contradictory or differently reliable evidence.

The same paper measures a 27.5-point discrepancy between strict token-F1 and LLM-as-judge on identical outputs, which is a metric-validity problem in the same family: the construct is under-specified and the measurement of it is unstable.

How this was produced definition_kind is a kind-typed field, so each record had to say which kind of definition it found rather than tick a boolean. Limit: the boundary between "operational threshold" and "contrastive" is a judgement call — a threshold is often stated by contrast. Treat the two oblique kinds as one class if that distinction matters to you.
Frame — core ring · n = 376 rows with a definition present, of 995

04 · Shape of the field

What this corpus is made of

Memory construct targeted

Non-exclusive; a benchmark may target several. Families with n ≥ 18 shown.

Setting

Primary evaluation setting. Families with n ≥ 20 shown; the long tail of open-coded settings is in the catalogue.

What kind of artifact ships

Flags carried

Non-exclusive. These drive how the catalogue is read, not whether a row is included.

304 of 995 artifacts (31%) are method-origin — a benchmark shipped alongside a memory method rather than built as a benchmark-first contribution — and 24 ship no artifact at all, reusing LoCoMo, ALFWorld or an existing dataset. Both are kept and flagged rather than dropped, so benchmark independence stays readable. 168 rows are flagged unverified: the released artifact could not be confirmed from the text.

How this was produced Tallies over per-paper extraction records, deduplicated by corpusId before counting (0 double extractions). Flag definitions were given to extractors verbatim in the worker packet. Limit: papers shipping both a strong benchmark and a method collapse toward method-origin, so 30% is an upper read.
Frame — core ring · n = 995 rows · 973 distinct artifacts after name-collision splitting

05 · The catalogue

All 995 rows, searchable

Sorted by citation count. Search matches name, description, construct, setting and definition text. Every row links to the paper. The full table with evidence spans and warrants is in data/catalog.csv, which also carries the verbatim evidence span for every row.

The "What it is" column gives each benchmark twice: first the authors' own sentence, copied verbatim from the paper — 920 of 995 (92.5%) had one, 849 from the abstract and 78 only findable in the body — and then a plain-language summary written here. Where a source never characterises its own artifact, the cell says so rather than inventing a description; 51 rows fall back to a curated-list line and 12 have nothing but a title. Those weaker cells are marked in a different style.

Benchmark What it is — authors' words, then ours Year Construct One example is… Defines "long horizon" Cites

Showing the first 300 matches here. The full catalogue — all 995 rows, newest first, with year/setting/construct filters — is a separate page: Agent Memory Benchmark Catalogue (also catalog.html in this package). Raw data: data/catalog.csv.

06 · Coverage & boundary

What this corpus does not cover

Estimated recall is about 80%, and 80% is the optimistic read. Two independent modality pairs agree: citation-snowball × human-curated lists estimates a population of 1,236 ± 161; keyword sweeps × human-curated estimates 1,279 ± 123. Against 995 captured, that leaves roughly 250–290 in-charter artifacts uncaptured. Capture-recapture estimates are lower bounds under heterogeneous catchability, so the true recall is likely below 80%, not above.

A third estimator was computed and discarded: S2-sweep × arXiv-sweep returned a population estimate of 728 against 993 observed. An estimate below the observed count is the signature of two samples that are not independent — both sweeps run the same keyword culture over overlapping indexes. It passed the standard overlap gate, which is precisely the case that gate cannot see.

Numbers in this section trace to coverage-verdict.md and scratch/mention-shadow-gaps.json in the run directory.

07 · Method

How the corpus was built

Pipeline

  • Nine acquisition modalities: parametric seeding, curated lists, HuggingFace + vendor pages, survey benchmark tables, closed-denominator proceedings, raw S2 and arXiv sweeps, citation snowball, and a false-negative recovery pass over the screen's own discards.
  • 3,314 candidates, every one relevance-judged against three criteria with graded scores and a verbatim evidence quote. Label coverage 1.00.
  • 995 landed in-charter. 865 paper bodies fetched and cached; all 995 extracted by a 21-shard fleet, with zero double-extractions.
  • Every pooled claim on this page was re-read against the rows behind it and carries a written because, unless and basis note in data/synthesis.json.

Known defects, recorded

  • 158 judged rows were never merged into the corpus by the previous session — 44 of them in-charter. Caught by an integrity gate and repaired; every earlier count was short by that much.
  • A stale clause survived packet derivation, telling full-text workers to copy spans from abstracts. A worker flagged the contradiction and resolved it correctly; recorded because it could have silently capped span quality.
  • One worker died mid-shard. 30 of 48 records survived on disk and the shard resumed from record 31 — the sub-batch discipline cost one batch, not a shard.
  • 11 gold items were planted in all 26 judging shards, so raw judgment lines overcount by 114. All counts here dedupe by paper id first.
Files behind this page data/catalog.csv — all 995 rows with descriptions, spans, warrants and links · data/catalog-aggregates.json — every number quoted above · data/synthesis.json — warrants for each pooled claim · data/charts.json and data/page-data.json — the chart and ladder inputs. Run artifacts (candidates, judgments, observations, coverage verdict, extraction records) live in the run directory alongside this report.
Corpus as of 2026-09-12 · refresh when the deferred gap-closing round runs · substrate run2