Benchmarks, environments and scenarios that make LLM agents juggle many goals

A catalogue of 480 evaluation artifacts published in 2025–2026 whose settings require an agent to pursue and track multiple goals and subgoals — what those goals are, and over what horizon. Software-engineering and other code-centric domains are deliberately excluded.
Complete as-of 2026-09-14. Built from 4,327 candidates, 2,191 of them judged. Expect staleness after roughly 6–8 weeks — see the refresh trigger in the coverage section.

480artifacts catalogued
318from 2026
162from 2025
334state no horizon
291give subgoal credit

Question 1 of 3What kinds of goals and subgoals must be planned and tracked?

The dominant shape is a chain or a dependency graph handed to the agent up front: 160 artifacts present goals as a sequential chain and 125 as a DAG with explicit precedence, while 297 of 480 hand the agent its goals at the start. Settings where the agent generates its own goals remain rare — 15 artifacts.

sequential-chain160DAG-with-precedence125hierarchical80set-of-independent71open-ended31other (singleton)13
How goals relate to one another. n=480 artifacts, one value each.
given-up-front297emitted-by-environment-over-time70mixed61implied-by-constraints19injected-by-user-mid-episode18self-generated-by-agent15
How goals reach the agent. n=480.
how this number was checked

Claim. Goal structure is dominated by sequential chains (160) and DAGs with precedence (125); hierarchical decomposition accounts for 80, sets of independent goals 71, and open-ended goal generation only 31.

Why it holds. Each value is the extractor's kind-typed choice from a fixed vocabulary, grounded in a verbatim span describing how the setting's goals relate.

What would break it. unless 'sequential-chain' is absorbing cases that are really DAGs whose precedence the abstract does not spell out — the two are distinguishable only when a paper describes its dependency structure, which many do not.

480 goal-structure records; decomposition normalized by leading enum token, 13 singletons kept in a labeled tail.

how this number was checked

Claim. 297 of 480 (62%) hand the agent all its goals up front; only 15 (3%) have the agent generate its own goals, and 70 (15%) have the environment emit goals over time.

Why it holds. goal_origin is a kind-typed field with a fixed vocabulary; the mixed class (61) was folded from 'mixed:*' variants rather than dropped.

What would break it. unless 'given-up-front' is over-applied to benchmarks whose task instruction is given up front but whose subgoals genuinely emerge during execution — the field records how GOALS arrive, and extractors may have keyed on the instruction rather than the subgoal stream.

480 goal-structure records; goal_origin normalized to 6 families.

What the goals are varies enormously by domain, and the catalogue below records them per artifact in the extractors’ own kind-typed terms. A few representative shapes, quoted from the papers:

“Agents must balance inventories, place orders, set prices, and handle daily fees - tasks that are each simple but collectively, over long horizons (>20M tokens per run) stress an LLM’s capacity for sustained, coherent decision-making.” — Vending-Bench
“Because terminal win/loss is too sparse a signal in games spanning hundreds of turns and multiple opponents, CivBench trains models on turn-level game state to estimate victory probabilities throughout the match.” — CivBench
How performed. Corpus: the 480-artifact catalogue (the fully-curated ring; borderline and unjudged papers excluded). Method: one kind-typed extraction record per paper under a goal-structure lens, each carrying a verbatim span from the paper; counts are tallies over those records, deduped by paper. Evidence depth: abstract-primary, with full-text escalation where a field was missing. Frame: charter core under thread config v3, n=480. Key limit: single-extractor per paper — three anchor papers repeated across all eight extractors agreed on artifact identity but showed vocabulary drift on the structure label.

Question 2 of 3What time horizon is actually being considered?

The single most consequential finding in this corpus: most of these papers never say. 334 of 480 (70%) give no quantified horizon at all — no step count, no turn count, no wall-clock figure, no simulated duration. Many describe their tasks as “long-horizon” in prose while reporting only task counts, model counts and success rates. One extraction warrant records it plainly: “The paper describes tasks qualitatively as ‘long-horizon’ without ever giving a quantified step/turn/time figure.”

states a quantified horizon146states none334
Whether the artifact quantifies its own horizon. n=480.
how this number was checked

Claim. Of 480 catalogued artifacts, 332 (69%) give no quantified time horizon and 148 (31%) do.

Why it holds. Support composition is 148 own-experiment/benchmark-design records against 332 absence records, and absence here is a CLAIM with a basis span showing what the paper reports instead (task counts, model counts, success rates), not a silent null. Fit is explicit on 299 of the 332 absence rows.

What would break it. unless the absence is an artifact of abstract brevity rather than the papers: only 56 of the 332 absence rows were escalated to full text, and extractor confidence was 1 (not 2) on 258 of 332 — so this is a claim about what these papers FOREGROUND, and the full-paper rate is unmeasured for the other 276.

480 horizon-lens extraction records, one per core paper, deduped by (corpusId, lens); 56 escalated to full text and still yielded no figure.

Among the 146 that do quantify, there is no shared unit. The field measures its horizons in turns, agent steps, actions, simulated days, episodes, sessions, tool calls, wall-clock hours and minutes, human-expert hours, and even simulated years — which makes cross-benchmark horizon comparison essentially impossible today.

“Agents must balance inventories, place orders, set prices, and handle daily fees - tasks that are each simple but collectively, over long horizons (>20M tokens per run) stress an LLM’s capacity for sustained, coherent decision-making.” — Vending-Bench, one of the few artifacts that quantifies its horizon, and it does so in tokens
not recorded334turns26other (singleton)25agent-steps22actions17simulated-days16episodes8sessions8tool-calls7wall-clock-hours6wall-clock-minutes4human-expert-hours3simulated-years3wall-clock-days1
Units used by the artifacts that quantify a horizon, normalized to families with a labelled singleton tail. “not recorded” is kept a distinct class, never folded into other. n=480.
how this number was checked

Claim. Where a horizon IS stated, the 148 papers use 13 mutually incommensurable units — turns (26), agent-steps (22), actions (17), simulated-days (16), episodes (8), sessions (8), tool-calls (7), wall-clock-hours (6), wall-clock-minutes (4), human-expert-hours (3), simulated-years (3), wall-clock-days (1), plus 25 one-off units.

Why it holds. Each unit count comes from a verbatim span in the paper's own words, so the diversity is the field's, not an artifact of coding; 144 of the 148 present-records rest on own-experiment support.

What would break it. unless the unit families are collapsible in ways this coding does not attempt — 'turns' and 'agent-steps' may be the same quantity under two names in some benchmarks, which would reduce the count but not the incommensurability across the wall-clock / simulated-time / step-count groups.

148 horizon records with polarity=present; horizon_unit normalized to families with a labeled singleton tail of 25.

How performed. Corpus: the 480-artifact catalogue. Method: a dedicated horizon lens per paper; a paper with no horizon figure yields an explicit absence record carrying a span of what it reports instead, never a silent blank. Evidence depth: abstract-primary; 56 absence records were escalated to full text and still yielded no figure. Frame: charter core under thread config v3, n=480. Key limit, and it matters: 276 of the 332 absence records were judged from the abstract alone and extractor confidence was 1 rather than 2 on 258 of them — so this is a measured claim about what these papers foreground, not a proof that the number appears nowhere in the PDF.

Question 3 of 3Is progress on subgoals actually scored?

291 of 480 (61%) award credit at the subgoal, checkpoint, milestone or rubric level rather than only at the end; 152 score final outcomes only, and 37 do not record a scheme. That majority matters for anyone studying goal tracking: partial-credit scoring is what makes where an agent broke down observable at all.

“Because terminal win/loss is too sparse a signal in games spanning hundreds of turns and multiple opponents, CivBench trains models on turn-level game state to estimate victory probabilities throughout the match.” — CivBench, on why end-of-episode scoring is not enough
yes291no152unclear37
Whether subgoal-level credit exists. n=480; “unclear” is a distinct class.
subgoal / checkpoint partial credit291other (unclassified)87not recorded37binary final success only32continuous / cumulative reward25programmatic verifier6LLM-judge or expert rubric2
Scoring mechanism kind. n=480; the unclassified tail is labelled, not hidden.
how this number was checked

Claim. 291 of 480 (61%) award subgoal- or checkpoint-level partial credit; 152 (32%) score only final outcomes; 37 (8%) do not record a scoring scheme.

Why it holds. Classification runs over the extractors' scoring field, which was written from a verbatim span per paper; the three classes are kept distinct with 'not recorded' as its own class rather than folded into an 'other' bucket.

What would break it. unless the keyword classifier over-assigns partial credit: it fires on rubric/progress/coverage/trajectory language, which a paper may use for reporting without awarding intermediate credit during evaluation. 87 rows remain in a labeled 'other (unclassified)' tail and are not counted on either side.

480 goal-structure extraction records; scoring classified to answer the asked question (does subgoal-level credit exist) rather than left as 432 raw strings.

How performed. Corpus: the 480-artifact catalogue. Method: the extractors recorded each artifact’s scoring scheme from a verbatim span; those free-text values were then classified to answer the asked question — does subgoal-level credit exist — rather than left as 432 distinct raw strings. Frame: charter core under thread config v3, n=480. Key limit: the classifier keys on rubric/progress/coverage language, which a paper may use when reporting results without awarding intermediate credit during evaluation; 87 rows stay in a labelled unclassified tail and count on neither side.

Shape of the fieldThe landscape

15 domain families, and 2026 outnumbers 2025 by 318 to 162 — this is a field reorganising itself around long-horizon evaluation in real time. Each artifact's family comes from the domain its extraction recorded after reading the paper, not from a keyword in its title; 153 artifacts genuinely straddle two families and carry a second one in the catalogue.

Business, office & enterpr105Personal assistants & long65Embodied & robotics46Multi-agent organisations &amp44Web & GUI agents39Games & interactive fictio34Tool & API use29OS & computer use24Customer service & long di23Travel & constraint planni19Open-ended sandboxes18Scientific discovery14Information seeking & deep10Healthcare & clinical8Education & tutoring220252026
Artifacts per domain family, split by year. n=480.
how this number was checked

Claim. The corpus splits into 15 domain families, led by business office enterprise (105), personal assistant memory (65), embodied household (46), multi agent org (44), web GUI (39); 2026 outnumbers 2025 by 318 to 162 across the corpus.

Why it holds. Family now comes from the kind-typed `domain` field the extractor recorded after reading the paper — 470 of 480 rows, with every `other:<specific>` value routed explicitly rather than pooled into a catch-all. No assignment rests on a title keyword. The exception is information-seeking (10 artifacts), re-derived from each paper's extracted GOALS because the extraction packet's domain vocabulary never offered that option — a gap in the packet, not an extractor judgment; those rows are marked evidence-rederived in family_source.

What would break it. unless the deliberate merges mislead: robotics-sim is folded into embodied-household and text-game-IF into games, because readers of this corpus treat each pair as one literature. 153 artifacts genuinely span two families and carry a recorded secondary family; counting only the primary understates those literatures. And information-seeking is small here for a scope reason, not a field reason: 130 information-seeking artifacts were judged out because a single research question decomposed into subqueries is one goal under this corpus's multi-goal test.

480 catalog rows; family from the extraction records, per-family era counts from the candidate year field. Supersedes an earlier ordered-keyword rollup that let a title word override the extractor (it filed a Slay the Spire testbed under memory because the title said 'Bounded-Memory'); that version over-counted business by 31, assistants by 29 and multi-agent by 25, under-counted tool-use and OS/computer-use by 16 each, and hid open-ended-sandbox entirely.

benchmark350environment/simulator101dataset17scenario-suite8arena/leaderboard3other:deployed-multi-agent-assistant1
What the artifacts call themselves. n=480.
how this number was checked

Claim. 350 of the artifacts describe themselves as benchmarks and 101 as environments or simulators, with 17 datasets, 8 scenario suites and 3 arenas.

Why it holds. artifact_kind is a kind-typed extraction field taken from the paper's own self-description.

What would break it. unless the benchmark/environment distinction is softer than the counts suggest — several artifacts ship an interactive environment AND a fixed task set, and extractors had to pick one label.

480 goal-structure records.

How performed. Corpus: the 480-artifact catalogue. Method: family assignment composes each paper’s kind-typed domain field, the relevance judge’s domain code and the title under one deterministic rule applied to all rows. Frame: charter core under thread config v3, n=480. Key limit: family boundaries are a choice — business/office/enterprise absorbs retail, finance and workplace simulation.

The deliverableThe catalogue — 480 artifacts

Every artifact, grouped by domain, with the goals it demands, what the agent must keep track of, its horizon, and the paper’s own words behind the entry. Filter and search below; each title links to the paper. Citation counts are shown per artifact — with 318 of 480 published in 2026 and 126 sitting at zero, treat that number as a measure of age rather than of quality. A sortable table view of the same 480 rows is available alongside this page.

Business, office & enterprise work 105 artifacts · 40 state a horizon

Less Context, Better Agents: Efficient Context Engineering for Long-Horizon Tool-Using LLM Agents
Goals to track itemize every line item of a hotel expense report completely and accurately using enterprise MCP tools; avoid context overflow / stale-state errors while repeatedly calling verbose enterprise tool APIs across the itemization workflow
Must keep track of The agent must track already-itemized line items, running token/context budget, and must avoid stale or overflowing tool-call history while repeatedly invoking enterprise MCP tools to reach complete, accurate itemization.
structure: set-of-independent · goals arrive: given-up-front · count: 50 tasks; five expense types grouped int · subgoal credit: yes · 1 citations · unit: wall-clock-hours
Horizonfull-context retention: 1,480,996 tokens and 14.56 hours per benchmark (50 tasks x 5 runs); pruning to last 5 tool calls: 535,274 tokens and 5.39 hours; pruning + summari
the paper’s own words
We evaluate four GPT-5 configurations on a 50-task hotel expense benchmark: no user model, full conversation history, context pruned to the last 5 tool call/response pairs, and pruning with automated summarization... The no-user-model baseline achieves only 8.0% complete itemization. Full-context retention improves completion to 71.0%... Adding summarization achieves the best result: 91.6% complete itemization and 99.64% average amount itemized.
Full-context retention improves completion to 71.0%, but consumes 1,480,996 tokens and 14.56 hours per benchmark. Pruning to the last 5 tool calls improves completion to 79.0% while reducing token use to 535,274 and runtime to 5.39 hours. Adding summarization achieves the best result: 91.6% complete itemization and 99.64% average amount itemized, with 553,374 tokens and 5.79 hours.
AD-Bench2026relevant
AD-Bench: A Real-World, Trajectory-Aware Advertising Analytics Benchmark for LLM Agents
Goals to track answer a real user marketing-analysis request via multi-round, multi-tool collaboration; maintain trajectory coverage consistent with an expert tool-call trajectory; produce answers that remain correct/consistent with a continuously evolving production advertising platform (dynamic ground truth); succeed across three stratified difficulty levels (L1-L3) requiring increasing multi-round, multi-tool collaboration
Must keep track of The agent must track which professional tools it has called and in what order across multiple rounds, since evaluation jointly measures end-to-end answer correctness (Pass@k) and how well its trajectory covers the expert reference trajectory, especially as this compounds at higher difficulty levels.
structure: sequential-chain · goals arrive: given-up-front · count: three difficulty levels (L1-L3) stratify · subgoal credit: yes · 3 citations
Horizonthe paper states none
the paper’s own words
a trajectory-aware evaluation that jointly measures end-to-end answer correctness (Pass@k) and trajectory coverage. Requests are stratified into three difficulty levels (L1-L3) to probe multi-round, multi-tool collaboration.
APEX-Agents
Goals to track execute a long-horizon, cross-application task created by investment-banking analysts, management consultants, or corporate lawyers; navigate realistic work environments containing files and tools to produce the required deliverable
Must keep track of Agent must track which files and tool states it has already produced/consulted while navigating a cross-application work environment, using rubrics and gold outputs to determine success (Pass@1).
structure: sequential-chain · goals arrive: given-up-front · count: 480 tasks (n=480) · subgoal credit: unclear · 20 citations
Horizonthe paper states none
the paper’s own words
APEX-Agents requires agents to navigate realistic work environments with files and tools. We test eight agents for the leaderboard using Pass@1.
APEX-Agents requires agents to execute long-horizon, cross-application tasks created by investment banking analysts, management consultants, and corporate lawyers
AeroCopilotBench: A Two-Tier Benchmark for Evaluating LLM Agents as Aviation Copilots in an Interactive Virtual Cockpit Environment
Goals to track diagnose faults and operate aircraft systems through standardized tool interfaces during emergency/abnormal procedures; achieve all task goal-conditions for each of 73 Tier-2 emergency/abnormal tasks; avoid violating any hard safety constraint while progressing toward goals
Must keep track of Agent must interpret cockpit state, diagnose faults, and operate aircraft systems while simultaneously tracking goal-condition progress and hard safety-constraint compliance across the executable procedure.
structure: DAG-with-precedence · goals arrive: given-up-front · count: 73 emergency/abnormal Tier-2 tasks (plus · subgoal credit: yes · 0 citations
Horizonthe paper states none
the paper’s own words
We establish a safety-gated evaluation framework in which a trajectory succeeds only when all task goals are achieved without violating any hard safety constraint, while safe goal progress and trajectory safety are measured separately.
Tier-2 comprises 73 emergency and abnormal tasks derived from the manufacturers'Pilot's Operating Handbooks (POHs) and instantiated in ACOE.
Agentic ERP: Multi-Agent Large Language Model Architecture for Autonomous Enterprise Resource Planning
Goals to track execute end-to-end business workflows across ERP functional boundaries (e.g. procurement, inventory, crisis response); resolve cross-functional crisis tasks via a Planner-Executor-Reflector-Responder orchestration; sustain a full simulated year of ERP operation while avoiding stockouts, compared against rule-based RPA and no-intervention baselines
Must keep track of The agent(s) must track the structured enterprise state (inventory, demand stream, crisis conditions) across a full simulated year, coordinate across role-aligned agents via externalised grading criteria and sprint contracts, and avoid accumulating stockouts as the rule-based baseline does.
structure: hierarchical · goals arrive: emitted-by-environment-over-time · count: a scenario-based task suite; a compariso · subgoal credit: yes · 0 citations · also Multi-agent organisations &amp; societies · unit: simulated-days
Horizona 365-day agent-in-the-loop simulation (a simulated year of ERP operation)
the paper’s own words
the system is evaluated at three levels: a scenario-based task suite, a comprehensive comparison of six orchestration paradigms on cross-functional crisis tasks, and a 365-day agent-in-the-loop simulation against rule-based RPA and no-intervention baselines. Across these levels the proposed multi-agent method is significantly better than the baseline, and the system sustains a simulated year of operation with zero stockouts while the rule-based baseline accumulates hundreds under the same demand stream.
AgenticPay2026relevant
AgenticPay: A Multi-Agent LLM Negotiation System for Buyer-Seller Transactions
Goals to track negotiate to reach agreement given private constraints and product-dependent valuations; maximize feasibility, efficiency, and welfare across a negotiation; successfully complete each of 110+ tasks ranging from bilateral bargaining to many-to-many markets
Must keep track of Agents must track their own private constraints/valuations, the evolving history of natural-language offers and counteroffers across negotiation rounds, and, in many-to-many markets, the state of other concurrent negotiations.
structure: sequential-chain · goals arrive: given-up-front · count: over 110 tasks ranging from bilateral ba · subgoal credit: yes · 13 citations · also Multi-agent organisations &amp; societies
Horizonthe paper states none
the paper’s own words
AgenticPay models markets in which buyers and sellers possess private constraints and product-dependent valuations, and must reach agreements through multi-round linguistic negotiation rather than numeric bidding alone. The framework supports a diverse suite of over 110 tasks ranging from bilateral bargaining to many-to-many markets, with structured action extraction and metrics for feasibility, efficiency, and welfare.
AgenticVBench2026relevant
AgenticVBench: Can AI Agents Complete Real-World Post-Production Tasks?
Goals to track complete each of 100 agentic post-production tasks across 4 task families; compose capabilities across text, image, audio, and video understanding within one task; plan and execute long-horizon multi-step production workflows using appropriate tools
Must keep track of The agent must track progress through a long-horizon, multi-modal production workflow and correctly sequence tool use, since tasks are constructed from real production workflows contributed by industry experts and evaluated jointly by programmatic verifiers and expert rubrics.
structure: other:four-task-families-each-with-real-produc · goals arrive: given-up-front · count: 100 agentic tasks across 4 task families · subgoal credit: yes · 3 citations
Horizonthe paper states none
the paper’s own words
Tasks are paired with evaluation specifications that combine programmatic verifiers and expert rubrics.
they require composite capabilities across text, image, audio, and video understanding, along with long-horizon planning, and tool use
Agents' Last Exam
Goals to track complete a long-horizon, economically valuable, real-world professional task with a verifiable outcome, drawn from a specific occupational sub-field; achieve sustained performance across the task rather than a single-shot correct answer, particularly on the hardest ('last-exam') tier
Must keep track of The agent must sustain performance on long-horizon, economically valuable real-world tasks whose outcomes are verifiable, across an occupational taxonomy spanning 55 sub-fields and 13 industry clusters, with the hardest tier proving far from saturated (average full pass rate below 1%).
structure: set-of-independent · goals arrive: given-up-front · count: 1K+ tasks organized into 55 sub-fields g · subgoal credit: yes · 12 citations
Horizonthe paper states none
the paper’s own words
This paper introduces Agents' Last Exam (ALE), a benchmark designed to evaluate AI agents on long horizon, economically valuable, real world tasks with verifiable outcomes. Developed in collaboration with 250+ industry experts, ALE covers non-physical industries defined with reference to O*NET / SOC 2018 (the U.S. federal occupational taxonomy). It is organized around a task taxonomy with 55 sub fields grouped into 13 industry clusters covering 1K+ tasks. Current results show that the hardest tier remains far from saturated: across mainstream harness and backbone configurations, the average full pass rate is below 1%.
AutomationBench2026relevant
AutomationBench
Goals to track discover the relevant REST API endpoints needed for a cross-application workflow (CRM, inbox, calendar, messaging, etc.); follow layered business-policy rules while writing data; navigate environments containing irrelevant or misleading records without being derailed; get correct data into the right systems by the end of the workflow (end-state correctness)
Must keep track of Agent must track which endpoints it has discovered across multiple applications, which business-policy rules apply to each write, and whether records encountered are relevant or misleading, to ensure the correct final data lands in each system.
structure: DAG-with-precedence · goals arrive: implied-by-constraints · subgoal credit: yes · 1 citations
Horizonthe paper states none
the paper’s own words
Grading is programmatic and end-state only: whether the correct data ended up in the right systems.
tasks span Sales, Marketing, Operations, Support, Finance, and HR domains
BigFinanceBench2026relevant
BigFinanceBench: A Workflow-Grounded Benchmark for Financial-Research Agents
Goals to track produce an auditable financial-research derivation (source choice, period/accounting definition, assumptions, calculation), not just a final answer; satisfy each of many independently checkable rubric points per item (36,241 points across 928 items); correctly complete each of 928 expert-authored open-ended financial-research tasks
Must keep track of The agent must track and justify each auditable component of its derivation (data source, period/accounting definition, assumptions, calculation steps) so that an independent rubric can check each step, not just the final numeric answer.
structure: hierarchical · goals arrive: given-up-front · count: 928 items; 36,241 rubric points total (a · subgoal credit: yes · 5 citations
Horizonthe paper states none
the paper’s own words
We introduce BigFinanceBench, a 928-item expert-authored benchmark of open-ended financial-research tasks in which each item pairs a ground-truth reference answer with a point-weighted rubric that decomposes the derivation into independently checkable steps... Across 36,241 rubric points, the benchmark supports partial-credit evaluation and localization of failures across the analyst workflow.
which source was chosen, which period and accounting definition were used, which assumptions were made, and how the calculation was performed
BlueFin2026relevant
BlueFin: Benchmarking LLM Agents on Financial Spreadsheets
Goals to track synthesize new spreadsheet content/formulas from source data; manipulate existing financial spreadsheet workbooks (edits, restructuring); comprehend and answer questions about spreadsheet content; satisfy each of up to 3,225 granular rubric criteria across 131 tasks
Must keep track of The agent must track and satisfy many granular, LM-judge-scored rubric criteria across synthesis, manipulation, and comprehension actions within one workbook, maintaining correctness as the workbook's formulas/state evolve ('dynamic correctness').
structure: hierarchical · goals arrive: given-up-front · count: 131 tasks containing 3,225 granular rubr · subgoal credit: yes · 1 citations · unit: human-expert-hours
Horizontasks require at least 45+ minutes of work for a human analyst to complete from scratch
the paper’s own words
we curate a set of 131 challenging, complex tasks with real-world relevance in the domain, containing 3,225 granular rubric criteria; notably, our rubric criteria and LM judge evaluations are validated by a team of expert human annotators.
at least 45+ minutes of work for an analyst to complete from scratch
Business Arena: Benchmarking LLM Agents in a Realistic Marketplace
Goals to track infer business opportunities from partial market signals; commit capital under uncertainty when buying from suppliers; set/adjust pricing to sell to buyers profitably; satisfy regulatory obligations before trading legally; sustain profitable operation of a cross-border shop over a long horizon
Must keep track of The agent must track capital, delayed and coupled consequences of sourcing/pricing/recovery decisions, market conditions calibrated from real data, and regulatory obligations across a long-horizon cross-border trading operation.
structure: open-ended · goals arrive: given-up-front · subgoal credit: no · 0 citations
Horizonthe paper states none
the paper’s own words
We introduce Business Arena, a controlled environment where an AI agent runs a cross-border shop, buying from suppliers and selling to buyers over a long horizon. ... action-level attribution identifies the sourcing, pricing, and recovery decisions that create or destroy value.
CEO-Bench2026in
CEO-Bench: Can Agents Play the Long Game?
Goals to track operate a startup for 500 days, managing pricing, marketing, budgeting, and other business aspects; acquire information from noisy, interconnected business databases and translate it into strategy; adapt strategy to a changing world/market over the run; orchestrate many interdependent business decisions toward a coherent long-term financial goal
Must keep track of The agent must track noisy, interconnected business databases (churn regimes, billing timing, customer losses, projected cash) and coordinate many interdependent decisions via code across a 500-day simulated run.
structure: open-ended · goals arrive: given-up-front · subgoal credit: yes · 3 citations · also Games &amp; interactive fiction · unit: simulated-days
Horizon500
the paper’s own words
We introduce CEO-Bench, which evaluates these capabilities together by simulating a representative real-world task: operating a startup for 500 days. An agent manages pricing, marketing, budgeting, and many other aspects of a fictional company through a programmable Python interface.
CEO-Bench2026in
Can LLMs Be CEOs? Benchmarking Strategic Resource Reallocation with Multi-Role Agent Simulation
Goals to track integrate conflicting recommendations from four role-conditioned C-suite advisors (CFO, CTO, COO, CMO) into one allocation plan; redirect capital across business units under information asymmetry and organizational constraints; synthesize advice consistently across multiple rounds with temporal dependencies
Must keep track of Agent (as CEO) must track and reconcile conflicting, privately-signaled recommendations from four C-suite advisors across a multi-round resource-reallocation process, remembering prior decisions for history-sensitive judgment.
structure: DAG-with-precedence · goals arrive: mixed:given-up-front-scenario-with-conflicting · count: 13 scenarios; four role-conditioned advi · subgoal credit: yes · 1 citations
Horizonthe paper states none
the paper’s own words
LLM agents receive conflicting advice from four role-conditioned C-suite advisors (CFO, CTO, COO, CMO), each with private signals and distinct priorities, and must synthesize these into a concrete allocation plan evaluated along four dimensions: role integration, conditional boldness, history-sensitive judgment, and plan validity.
the process of redirecting capital across business units in a multi-round, constraint-rich organizational environment... Experiments across five frontier models on 13 scenarios reveal that all models achieve high structural validity but diverge sharply on strategic calibration
CHI-Bench: Can AI Agents Automate End-to-End, Long-Horizon, Policy-Rich Healthcare Workflows?
Goals to track ground each decision in a large library of medical, insurance, and operational policy rules; play multiple roles within a single task, handing off between them; conduct multilateral multi-turn dialogs (e.g. peer-to-peer review, patient outreach) as intermediate workflow steps; drive a clinical case to a terminal status via tool calls and role-artifact writing
Must keep track of The agent must track its current role and handoffs to other roles, policy compliance against a 1,290+ document handbook, and the clinical case's evolving status across a high-fidelity simulator of 20 apps, until it reaches a terminal status.
structure: DAG-with-precedence · goals arrive: given-up-front · count: long-horizon workflows across 3 domains, · subgoal credit: no · 6 citations · also Healthcare &amp; clinical · unit: other:not-stated
Horizonthe paper states none
the paper’s own words
we introduce $\chi$-Bench, a benchmark of long-horizon healthcare workflows across three domains: provider prior authorization, payer utilization management, and care management. Each task hands the agent a clinical case in a high-fidelity simulator of 20 healthcare apps exposed via 87 MCP tools, which it must drive to a terminal status through tool calls and writing the role's artifacts, guided by a 1,290+ document managed-care operations handbook skill.
executing all tasks in a single session slumps the performance to 3.8%
Claw-Eval-Live2026relevant
Claw-Eval-Live: A Live Agent Benchmark for Evolving Real-World Workflows
Goals to track complete end-to-end units of work across software tools, business services, and local workspaces per release; pass controlled tasks with fixed fixtures/services/workspaces/graders reflecting current public workflow-demand signals; handle HR, management, and multi-system business workflows as well as local workspace repair
Must keep track of Agent must produce verifiable execution traces, audit logs, and consistent service/workspace state across each end-to-end task, since grading uses deterministic checks when evidence is sufficient and structured LLM judging only for semantic dimensions.
structure: set-of-independent · goals arrive: given-up-front · count: 105 tasks per release (ClawHub Top-500 s · subgoal credit: no · 13 citations
Horizonthe paper states none
the paper’s own words
For grading, Claw-Eval-Live records execution traces, audit logs, service state, and post-run workspace artifacts, using deterministic checks when evidence is sufficient and structured LLM judging only for semantic dimensions. The release contains 105 tasks spanning controlled business services and local workspace repair, and evaluates 13 frontier models under a shared public pass rule.
ClawMark2026in
ClawMark: A Living-World Benchmark for Multi-Turn, Multi-Day, Multimodal Coworker Agents
Goals to track carry a professional coworker task forward correctly across multiple working days; detect and adapt to exogenous environment updates (new emails, calendar shifts, KB edits) that occur between turns; satisfy each of the task's deterministic Python checkers over post-execution service state (mean 15.4 checkers/task); coordinate consistent state across five stateful sandboxed services (filesystem, email, calendar,
Must keep track of The agent must track evolving state across five sandboxed services (filesystem, email, calendar, knowledge base, spreadsheet) over multiple working days, detecting and incorporating exogenous updates injected between turns, verified by up to 29 deterministic checkers per task.
structure: sequential-chain · goals arrive: emitted-by-environment-over-time · count: 100 tasks across 13 professional scenari · subgoal credit: yes · 15 citations · unit: simulated-days
Horizon2-6 turns per task (mean 3.6); one turn = one in-universe working day
the paper’s own words
The current release contains 100 tasks across 13 professional scenarios, executed against five stateful sandboxed services (filesystem, email, calendar, knowledge base, spreadsheet) and scored by 1537 deterministic Python checkers over post-execution service state; no LLM-as-judge is invoked during scoring. ... The strongest model reaches 75.8 weighted score, but the best strict Task Success is only 20.0%
Tasks range from two to six turns (mean 3.6) and 6 to 29 checkers (mean 15.4).
ClawsBench2026relevant
ClawsBench: Evaluating Capability and Safety of LLM Productivity Agents in Simulated Workspaces
Goals to track complete single-service productivity tasks (e.g. Gmail, Slack, Calendar, Docs, Drive) correctly and safely; complete cross-service workflows spanning multiple mock services while maintaining consistent state; avoid unsafe/irreversible actions in safety-critical scenarios
Must keep track of Agents must track persistent state across five mock services (Gmail, Slack, Calendar, Docs, Drive), recognize safety-critical constraints to avoid irreversible/unsafe actions, and coordinate behavior across services via a meta-prompt layer.
structure: DAG-with-precedence · goals arrive: given-up-front · count: 44 structured tasks spanning single-serv · subgoal credit: yes · 25 citations
Horizonthe paper states none
the paper’s own words
It includes five high-fidelity mock services (Gmail, Slack, Google Calendar, Google Docs, Google Drive) with full state management and deterministic snapshot/restore, along with 44 structured tasks covering single-service, cross-service, and safety-critical scenarios.
Experiments across 6 models, 4 agent harnesses, and 33 conditions show that with full scaffolding, agents achieve task success rates of 39-64% but exhibit unsafe action rates of 7-33%.
CoffeeBench: Benchmarking Long-Horizon LLM Agents in Heterogeneous Multi-Agent Economies
Goals to track maximize cumulative net income for one's own firm (coffee roaster) over a long simulated run; manage cash, inventory, and pricing on an ongoing basis; communicate and transact with other heterogeneous firms (farmers, roasters, retailers) to secure supply/demand
Must keep track of The agent must track its own cash, inventory, and pricing state day by day over the 90-day simulation, plus the state of its ongoing communications/transactions with other firms, to maximize cumulative net income rather than any single day's outcome.
structure: sequential-chain · goals arrive: given-up-front · subgoal credit: yes · 6 citations · also Multi-agent organisations &amp; societies · unit: simulated-days
Horizon90-day simulation
the paper’s own words
two farmers, two roasters, and two retailers autonomously operate their businesses over a 90-day simulation, each seeking to maximize cumulative net income through communication and transactions while managing cash, inventory, and pricing. The evaluated model controls one coffee roaster, while the remaining firms are controlled by fixed reference agents.
Do AI-Native Biotechs Need Departments? Benchmarking Company World Models for AI-Driven Drug Development
Goals to track maintain and update a persistent asset-to-value state (Live Asset Value Record) across scientific, regulatory, BD, commercial, financial, and execution constraints; resolve 45 retrospective public-information decision cases with hidden outcomes under strict time cutoffs; achieve success via external BD deals, regulatory approval/launch, and revenue discipline
Must keep track of The architecture must maintain a persistent, continuously updated asset-to-value state record across scientific, regulatory, BD, commercial, financial, and execution constraints via Deal/Approval/Revenue/Investment Arbiter loops.
structure: DAG-with-precedence · goals arrive: given-up-front · count: 45 retrospective decision cases · subgoal credit: no · 0 citations
Horizonthe paper states none
the paper’s own words
The benchmark contains 45 retrospective public-information decision cases with strict time cutoffs, hidden outcomes, common schemas, automatic scoring, and blinded pairwise judging... The value-conversion architecture is a prompt-level approximation of a Company World Model: a Live Asset Value Record updated by Deal, Approval, Revenue, and Investment Arbiter loops.
EnterpriseBench Corecraft: Training Generalizable Agents on High-Fidelity RL Environments
Goals to track perform multi-step, domain-specific customer-support work across a simulated enterprise with 2,500+ entities across 14 entity types; satisfy all expert-authored rubric criteria for a given task using 23 available tools
Must keep track of Agent must track entity state across the enterprise simulation and verify each expert-authored rubric criterion is satisfied before the task counts as solved.
structure: DAG-with-precedence · goals arrive: given-up-front · count: expert-authored rubric criteria per task · subgoal credit: yes · 7 citations
Horizonthe paper states none
the paper’s own words
Frontier models such as GPT-5.2 and Claude Opus 4.6 solve fewer than 30% of tasks when all expert-authored rubric criteria must be satisfied.
CoreCraft is a fully operational enterprise simulation of a customer support organization, comprising over 2,500 entities across 14 entity types with 23 unique tools, designed to measure whether AI agents can perform the multi-step, domain-specific work that real jobs demand.
Can LLM Agents Respond to Disasters? Benchmarking Heterogeneous Geospatial Reasoning in Emergency Operations
Goals to track perform disaster perception from heterogeneous geospatial imagery; conduct spatial relational analysis over roads/population/facilities; plan rescue and evacuation operations; reason about temporal evolution of the disaster; synthesize a multi-modal operational report
Must keep track of The agent must compose calls across a 108-tool MCP library over heterogeneous, multi-temporal geospatial data, tracking intermediate perception/analysis outputs and their correctness as the pipeline progresses through the five task dimensions.
structure: DAG-with-precedence · goals arrive: given-up-front · count: 515 expert-authored tasks; gold trajecto · subgoal credit: no · 4 citations · unit: tool-calls
Horizon3,500 tool-call steps total (across 515 tasks, ~6.8/task average)
the paper’s own words
we introduce Disaster Operational Response Agent benchmark (DORA), the first agentic benchmark for end-to-end disaster response: 515 expert-authored tasks across 45 real-world disaster events spanning 10 types, paired with expert-verified, replayable gold trajectories totaling 3,500 tool-call steps. Tasks span five dimensions that cover the operational disaster-response pipeline: disaster perception, spatial relational analysis, rescue and evacuation planning, temporal evolution reasoning, and multi-modal report synthesis.
DocOps2026relevant
DocOps: A Verifiable Benchmark for Autonomous Agents in Complex Document Operations
Goals to track perform atomic document-operation actions correctly (e.g. edits without destructive metadata corruption); complete escalating, more complex, more tightly coupled document-workflow tasks built from those atomic operations; maintain long-term state tracking and global document consistency across the workflow; correctly verify (not just superficially check) that a document edit was semantically achieved
Must keep track of The agent must maintain long-term state tracking of the document's evolving structure/content and verify (rather than superficially assume) that each operation was semantically correct, to preserve global document consistency across highly coupled, long-range tasks.
structure: hierarchical · goals arrive: given-up-front · subgoal credit: yes · 1 citations
Horizonthe paper states none
the paper’s own words
we introduce DocOps, a deterministically verifiable evaluation framework underpinned by a hierarchical taxonomy that deconstructs document operations inspired by real-world practices into atomic dimensions and escalating workflow complexities... a fine-grained analysis of existing agents' manipulation behaviors uncovers 3 key failure modes: long-term state tracking collapse, shallow semantic verification, and destructive editing of structural metadata.
DuMateBench2026relevant
DuMateBench: Evaluating Autonomous Agents in Complex Real-World Workflows
Goals to track complete each of 200 real-session-derived tasks spanning 8 broad scenarios and 17 fine-grained capability categories; coordinate multiple capability categories within a single task; preserve and correctly use persistent configurations and workspace state carried over from prior interaction history; maintain performance under injected real-world environmental complexity (Insufficient, Unstable, Noisy conditions)
Must keep track of The agent must track persistent configurations, workspace state, and pre-solution interaction history carried into each task, while coordinating multiple capability categories under injected Insufficient, Unstable, or Noisy environmental perturbations.
structure: set-of-independent · goals arrive: given-up-front · count: 200 tasks spanning 8 broad scenarios and · subgoal credit: no · 0 citations
Horizonthe paper states none
the paper’s own words
The resulting benchmark comprises 200 tasks spanning 8 broad scenarios and 17 fine-grained capability categories, with most tasks requiring multiple capability coordination. We execute these tasks in isolated Docker containers injected with three forms of real-world environmental complexity: Insufficient, Unstable, and Noisy, and assess performance using a hybrid deterministic and LLM-as-Judge evaluation protocol.
Each task preserves the relevant pre-solution interaction history, persistent configurations, and workspace state, and is then validated through human verification.
Anchor: Mitigating Artifact Drift in Agent Benchmark Generation
Goals to track satisfy explicit procurement-workflow constraints in a production-grade ERP system; satisfy explicit manufacturing-workflow constraints in a production-grade ERP system; reach the solver-certified fully optimal end-state solution for a long-horizon business task, not merely a feasible one
Must keep track of The agent must track state-based verifier conditions tied to the end-state of a production-grade ERP system across a long-horizon procurement or manufacturing workflow, since reward depends solely on end-state business correctness rather than intermediate steps.
structure: other:constraint-optimization-program-with-con · goals arrive: given-up-front · count: 300 long-horizon tasks (ERP-Bench) spann · subgoal credit: no · 1 citations
Horizonthe paper states none
the paper’s own words
we find that generation parameters predict realized difficulty, and that frontier models satisfy explicit task constraints in 26.1% of trials but reach a fully optimal solution in only 17.4% of trials
AI agents are beginning to complete valuable, long-horizon business operations tasks
ERPBench2026in
ERPBench: Evaluating LLM Agents for Enterprise Decision-Making Across Competitive Market Ecologies
Goals to track make coupled decisions each round across pricing, production, procurement, inventory, and finance; maximize valuation/rank while competing against fixed rule-based opponents (Solo) or other evaluated LLM agents (Arena) in a shared market; sustain a coherent business strategy across all six rounds of the ERP simulation
Must keep track of Agent must track its own accumulated valuation, inventory, and finance position across all six rounds, plus (in Arena mode) the shared market state shaped by five other competing LLM agents, to decide coupled pricing/production/procurement actions each round.
structure: sequential-chain · goals arrive: given-up-front · count: 100 fixed problems, each run for 6 round · subgoal credit: yes · 0 citations · unit: turns
Horizon6
the paper’s own words
We introduce ERPBench, an execution-instrumented benchmark for enterprise decision agents in a six-round Enterprise Resource Planning (ERP) simulation with coupled pricing, production, procurement, inventory, finance, and shared-market competition.
EcoGym2026in
EcoGym: Evaluating LLMs for Long-Horizon Plan-and-Execute in Interactive Economies
Goals to track maintain the Vending sub-environment's profitability (net worth, income) as in Vending-Bench; acquire and retain freelance income under budgeted actions (Freelance sub-environment); sustain operational metrics such as DAU under partial observability (Operation sub-environment); maintain long-term strategic coherence across an effectively unbounded, budgeted-action economic horizon
Must keep track of The agent must track budgeted actions, business-relevant outcome metrics (net worth, income, DAU), and partial observability/stochasticity across an effectively unbounded horizon of 1000+ steps (up to 365 day-loops) per evaluation.
structure: open-ended · goals arrive: given-up-front · count: three sub-environments (Vending, Freelan · subgoal credit: no · 1 citations · unit: agent-steps
Horizon1000+ steps over 365 simulated day-loops (evaluation horizon)
the paper’s own words
EcoGym comprises three diverse environments: Vending (adapted from the closed-source Vending-Bench, with full open-source release), Freelance (new), and Operation (new), implemented in a unified decision-making process with standardized interfaces, and budgeted actions over an effectively unbounded horizon (1000+ steps if 365 day-loops for evaluation). The evaluation of EcoGym is based on business-relevant outcomes (e.g., net worth, income, and DAU), targeting long-term strategic coherence and robustness under partial observability and stochasticity.
EntWorld2026relevant
EntWorld: A Holistic Environment and Benchmark for Verifiable Enterprise GUI Agents
Goals to track complete each of 1,756 tasks spanning six representative enterprise domains (CRM, ITIL, ERP, etc.); operate under strict business logic constraints and high-density enterprise UIs; maintain precise, state-consistent information retrieval verified via SQL-based state-transition checks; complete synthesized long-horizon workflows reverse-engineered from database schemas
Must keep track of The agent must track precise, state-consistent information across a high-density enterprise UI and the underlying database schema, since success is verified deterministically via SQL-based state-transition checks rather than visual matching.
structure: sequential-chain · goals arrive: given-up-front · count: 1,756 tasks across six enterprise domain · subgoal credit: no · 5 citations · also Web &amp; GUI agents
Horizonthe paper states none
the paper’s own words
Experimental results demonstrate that state-of-the-art models (e.g., GPT-4.1) achieve 47.61% success rate on EntWorld, substantially lower than the human performance, highlighting a pronounced enterprise gap in current agentic capabilities.
enabling the synthesis of realistic, long-horizon workflows
Can LLM Agents Be CFOs? Benchmarking Long-Horizon Resource Allocation in an Uncertain Enterprise Environment
Goals to track manage liquidity across the firm's operations; close financial books accurately on a regular cycle; gather costly signals about the macro/industry environment before acting; request equity or debt financing appropriately as conditions change; survive (avoid insolvency/failure) across the full 132-month horizon under shifting macroeconomic regimes
Must keep track of The agent must track liquidity, capital structure (equity/debt), and signals about shifting macroeconomic/industry regimes month by month across the 132-month horizon, since consequences of earlier decisions are delayed and only become apparent later.
structure: DAG-with-precedence · goals arrive: given-up-front · count: 132-month simulated horizon; 23 LLMs and · subgoal credit: yes · 2 citations · unit: other:simulated-months
Horizon132-month CFO simulation
the paper’s own words
We introduce EnterpriseArena, a 132-month CFO simulator that evaluates long-horizon resource allocation under uncertainty in a FinTech lending firm. Agents must manage liquidity, close books, gather costly signals, and request equity or debt financing across changing macroeconomic regimes.
EnterpriseLab: A Full-Stack Platform for developing and deploying agents in Enterprises
Goals to track complete complex enterprise workflows spanning IT, HR, sales, and engineering domains; correctly invoke and sequence tool calls across 140+ tools exposed via Model Context Protocol across 15 applications; match frontier-model performance while running as a smaller (8B) privacy-preserving model; generalize/remain robust across diverse enterprise benchmarks (EnterpriseBench, CRMArena)
Must keep track of The trained agent must track which of many interdependent enterprise tools/applications it has invoked and their resulting state, to complete complex multi-tool enterprise workflows without needing frontier-scale model capacity.
structure: DAG-with-precedence · goals arrive: given-up-front · count: 15 applications; 140+ tools across IT, H · subgoal credit: no · 1 citations
Horizonthe paper states none
the paper’s own words
Our results demonstrate that 8B-parameter models trained within EnterpriseLab match GPT-4o's performance on complex enterprise workflows while reducing inference costs by 8-10x, and remain robust across diverse enterprise benchmarks, including EnterpriseBench (+10%) and CRMArena (+10%).
We validate the platform through EnterpriseArena, an instantiation with 15 applications and 140+ tools across IT, HR, sales, and engineering domains.
EnterpriseClawBench: Benchmarking Agents from Real Workplace Sessions
Goals to track read heterogeneous workplace files relevant to a task; invoke the correct tools to accomplish a business objective; deliver a business artifact matching role-specific hard rules and semantic rubrics
Must keep track of The agent must track which files it has read, what tools it has invoked, and whether its eventual delivered artifact satisfies the task's hard rules and semantic rubric, since harness-model combination, artifact delivery, and visual quality are all separately reported.
structure: sequential-chain · goals arrive: given-up-front · count: 852 reproducible tasks, each with recove · subgoal credit: yes · 1 citations
Horizonthe paper states none
the paper’s own words
Starting from a large archive of workplace sessions, the EnterpriseClawBench produces 852 reproducible tasks, each paired with recovered fixtures, rewritten prompts, role classes, skill subclasses, hard rules, and semantic rubrics.
EnterpriseOps-Gym2026relevant
EnterpriseOps-Gym: Environments and Evaluations for Stateful Agentic Planning and Tool Use in Enterprise Settings
Goals to track complete each of 1,150 expert-curated enterprise tasks across eight mission-critical verticals (e.g., Customer Service, HR, IT); plan and act correctly amid persistent state changes and strict access-control protocols; correctly refuse infeasible tasks rather than attempting unintended, potentially harmful side effects
Must keep track of Agent must track persistent state changes across 164 database tables, access-control constraints, and whether the current task is feasible at all, to avoid unintended side effects from attempting an infeasible task.
structure: DAG-with-precedence · goals arrive: given-up-front · count: 1,150 expert-curated tasks across eight · subgoal credit: yes · 11 citations
Horizonthe paper states none
the paper’s own words
EnterpriseOps-Gym features a containerized sandbox with 164 database tables and 512 functional tools to mimic real-world search friction. Within this environment, agents are evaluated on 1,150 expert-curated tasks across eight mission-critical verticals (including Customer Service, HR, and IT).
Evaluating Deep Research Agents on Expert Consulting Work: A Benchmark with Verifiers, Rubrics, and Cognitive Traps
Goals to track produce a multi-document, decision-grade consulting deliverable from SME-authored prompts; pass deterministic binary verifiers (mean 14.9 per task); satisfy each of a five-criterion SME rubric (Data Integrity, Analytical Rigor, Relevance & Focus, Execution Precision, Format & Deliverability); avoid embedded cognitive traps (human-error mimicry, deterministic precision traps) that penalize surface-pattern reas
Must keep track of The agent must track which of the ~14.9 verifiers per task it has satisfied, avoid embedded cognitive traps requiring reconciliation against context, and keep its final deliverable consistent with all five rubric criteria simultaneously.
structure: set-of-independent · goals arrive: given-up-front · count: mean 14.9 binary verifiers per task, plu · subgoal credit: yes · 1 citations · unit: other:not-stated
Horizonthe paper states none
the paper’s own words
we score two complementary layers: deterministic binary verifiers (mean 14.9 per task) and a five-criterion 0--3 SME rubric ... combined into a Verifier-Rubric Score (VRS, 0--100). Acceptance under a joint threshold (rubric mean >= 2.5 and verifier pass rate >= 80%) is uniformly low
a benchmark of 70 SME-authored management consulting prompts, each embedding cognitive traps that penalize surface-pattern reasoning
EvoEnv2026in
The Agent's First Day: Benchmarking Learning, Exploration, and Scheduling in the Workplace Scenarios
Goals to track schedule and prioritize streaming tasks with varying priorities under a dynamic workload; actively seek out information to reduce hallucination before acting (prudent information acquisition); distill and reuse generalized strategies learned from earlier dynamically generated tasks in later ones (continuous evolution)
Must keep track of The agent must track incoming streaming tasks and their priorities, its own accumulated exploration/experience for continual learning, and its confidence in acquired information to avoid hallucination, across a continuously evolving workplace scenario.
structure: hierarchical · goals arrive: emitted-by-environment-over-time · count: 50 dynamic scenarios, each with 2-6 task · subgoal credit: yes · 1 citations · unit: actions
Horizon50 dynamic scenarios, each with 2-6 task instances; observed usage up to ~90 steps and 232 tool calls for the highest-usage evaluated model (Gemini-3-Flash)
the paper’s own words
\method{} evaluates agents along three dimensions: (1) context-aware scheduling for streaming tasks with varying priorities; (2) prudent information acquisition to reduce hallucination via active exploration; and (3) continuous evolution by distilling generalized strategies from rule-based, dynamically generated tasks.
Each scenario encompasses 2 to 6 task instances... While Gemini-3-Flash uses substantially more steps (90) and tool calls (232) than the middle-tier models, this increased activity reflects the necessary complexity
FM-Bench2026in
FM-Bench: A Benchmark for Long-Horizon Management with Competing Agents
Goals to track draft and manage a football squad within a fixed budget shared with rivals; trade players and negotiate contracts across the season; invest in facilities and youth development; set match lineups; maintain board confidence (avoid being fired) while maximizing a final accumulated score over 20 in-game years
Must keep track of The agent must track squad composition, budget/cash flow, ongoing contract-renewal deadlines, facility/youth investments, and the board's confidence in it, across roughly 340-400 decision stops spanning 20 in-game years, since 'the order settles only late in the horizon.'
structure: sequential-chain · goals arrive: given-up-front · count: 20 in-game years; 26 tools; roughly 340- · subgoal credit: yes · 0 citations · unit: other:decision-stops
Horizon20 in-game years; roughly 340 to 400 decision stops per run
the paper’s own words
An LLM agent runs a football club for 20 in-game years through 26 tools and roughly 340 to 400 decision stops. It drafts a squad on the same budget as every rival, trades players, negotiates contracts, invests in facilities and youth, sets lineups, and answers to a board that can fire it, while a deterministic engine accumulates every year into one final score with no LLM judge or human rater.
FORESIGHT-9: Prospective and Process-Aware Evaluation of Adaptive Trading Agents
Goals to track maintain a coherent, active factor-ensemble/trading strategy across staged macro-financial events without silent internal collapse; respond adaptively to joint multi-asset anchors and disclosed in-world-time observations across a multi-year counterfactual worldline; keep portfolio holdings and decision-records mutually consistent (avoid decision-record vs. execution divergence)
Must keep track of The agent must track its adaptive factor library/ensemble state, executed portfolio holdings, and staged event disclosures across a multi-year, ~2,468-trading-day counterfactual worldline, ensuring internal decision-records remain consistent with actual executed holdings.
structure: sequential-chain · goals arrive: emitted-by-environment-over-time · count: nine worldlines x two agent frameworks x · subgoal credit: yes · 0 citations · unit: simulated-days
Horizoneach worldline runs from the 2026-07-15 information boundary through 2035-12-31, approximately 2,468 simulated trading days (~9 years 5 months) per run, across 36 long-ho
the paper’s own words
We evaluate two adaptive trading-agent frameworks with two foundation-model backbones across 36 long-horizon runs. Agent rankings vary substantially across worldlines and backbones, and a fixed equal-weight policy outperforms 31 of 36 runs. Process telemetry exposes failures that terminal returns conceal: in one high-return run, the live factor library collapsed while executed holdings converged to the equal-weight fallback, even though decision records continued to report an active factor ensemble.
each run terminates at 2035-12-31 ... approximately 2,468 daily marks
FinEvo-Bench: A Longitudinal Benchmark for Self-Evolving Agents in Professional Financial Workflows
Goals to track complete each of 120 real-case-grounded financial tasks across 20 business scenes in six financial domains, following institution-provided procedures and constraints; retain and re-apply experience from earlier cases in a scene to later, substantively distinct cases sharing the same professional procedure (self-evolution); maintain financial-compliance quality across each task as scored by a manually reviewed rubric;
Must keep track of The agent (self-evolving scaffold) must retain experience -- via memory, skill distillation, or both -- from earlier cases in the interleaved task stream and apply it to later, procedurally-related but factually distinct cases, while satisfying institution-provided professional-procedure constraints and financial-compliance requirements on each individual task.
structure: other:six-related-cases-per-scene-with-longitu · goals arrive: emitted-by-environment-over-time · count: 120 tasks; 20 business scenes across six · subgoal credit: yes · 1 citations · unit: episodes
Horizon120 real-case-grounded tasks per longitudinal stream (20 scenes x 6 cases each); three independently shuffled, globally interleaved task streams
the paper’s own words
Each scene contains six related but substantively distinct cases that share a professional procedure and a manually reviewed rubric for task quality and financial compliance... Letta achieves the highest evolved score (91.65) and fewest compliance issues (0.09 per task); Codex achieves the largest self-evolution gain (+19.37).
We introduce FinEvo-Bench, a longitudinal benchmark with 120 real-case-grounded tasks, 20 business scenes across six financial domains... We compare four self-evolving agent scaffolds using the same Qwen3.7-Max backbone and three independently shuffled, globally interleaved task streams.
FrontierFinance2026relevant
FrontierFinance: A Long-Horizon Computer-Use Benchmark of Real-World Financial Tasks
Goals to track complete each of 25 complex financial modeling tasks across five core finance models; produce client-ready outputs matching industry-standard financial-modeling workflows; satisfy detailed structured evaluation rubrics per task
Must keep track of The agent must track detailed financial-modeling state (assumptions, formulas, model linkages) across a task requiring on average over 18 hours of equivalent skilled human labor, checked against detailed structured rubrics.
structure: sequential-chain · goals arrive: given-up-front · count: 25 complex financial modeling tasks acro · subgoal credit: yes · 2 citations · unit: human-expert-hours
Horizonaverage of over 18 hours of skilled human labor per task
the paper’s own words
we introduce FrontierFinance, a long-horizon benchmark of 25 complex financial modeling tasks across five core finance models, requiring an average of over 18 hours of skilled human labor per task to complete. Developed with financial professionals, the benchmark reflects industry-standard financial modeling workflows and is paired with detailed rubrics for structured evaluation.
GDPevo2026relevant
GDPevo: Evaluating Agent Self-Evolution on Real Business Tasks
Goals to track update an agent's persistent state from prior training-task experience (self-evolution); apply recombined atomic business rules correctly on held-out test tasks across CRM/ERP/finance/healthcare/legal/data-centric workflows; achieve held-out accuracy gains attributable specifically to training experience rather than data contamination
Must keep track of The evaluation harness must track which atomic business rules were exposed during training versus recombined at test time, per task group, to attribute held-out gains specifically to training experience rather than contamination.
structure: set-of-independent · goals arrive: given-up-front · count: V1: 120 tasks in 12 groups (5 training + · subgoal credit: yes · 0 citations
Horizonthe paper states none
the paper’s own words
Its core mechanism, rule hybridization, decomposes each enterprise workflow into atomic business rules, distributes subsets of these rules across training tasks, and recombines them in held-out test tasks so that test-time gains are attributable.
V1 release contains 120 tasks in 12 groups, with five training and five held-out test tasks per group
H-AdminSim: A Multi-Agent Simulator for Realistic Hospital Administrative Workflows with FHIR Integration
Goals to track process hospital administrative requests (drawn from a workload of 10,000+ requests/day in large hospitals); coordinate across multiple administrative subtasks rather than handling them in isolation; operate correctly against FHIR-integrated, heterogeneous hospital-setting data
Must keep track of The multi-agent simulation must track hospital administrative request state across heterogeneous hospital settings via a unified FHIR-integrated environment, evaluated against detailed rubrics.
structure: other:multi-agent-simulated-administrative-wor · goals arrive: given-up-front · subgoal credit: yes · 0 citations · also Multi-agent organisations &amp; societies
Horizonthe paper states none
the paper’s own words
These tasks are quantitatively evaluated using detailed rubrics, enabling systematic comparison of LLMs.
in large hospitals, process over 10,000 requests per day
HANDBOOK.md: A Benchmark for Long-Context Agentic Instruction Following
Goals to track locate the specific clauses in a long standing-policy document (20-124 pages) that apply to the current situation; carry out routine professional work strictly governed by that policy document across many tool-mediated actions; avoid letting a plausible but unauthorized in-environment request override the standing policy; satisfy every one of many deterministic, programmatic grading criteria (824 total) simultaneousl
Must keep track of The agent must hold the relevant clauses of a long standing-policy document across roughly 17 reasoning steps and 30 tool calls on average, checking each action against required and prohibited criteria rather than losing rule details over the long horizon.
structure: DAG-with-precedence · goals arrive: given-up-front · count: 65 agentic tasks; a rubric of 824 total · subgoal credit: yes · 2 citations · unit: tool-calls
Horizon~17 reasoning steps and ~30 tool calls on average per task
the paper’s own words
We present HANDBOOK_md, a benchmark of 65 agentic tasks modeled on how employees follow company handbooks. Each task places an agent in a self-contained company environment... and instructs it to carry out routine professional work governed by an expert-written standard operating procedure of 20-124 pages... each task carries a rubric of programmatic criteria (824 in total) that check both that required actions occurred and that prohibited actions did not.
Completing a task requires locating the clauses that apply, holding them across a horizon of roughly 17 reasoning steps and 30 tool calls on average.
HealthAdminBench2026relevant
HealthAdminBench: Evaluating Computer-Use Agents on Healthcare Administration Tasks
Goals to track complete a Prior Authorization workflow end-to-end across an EHR and payer portal; complete an Appeals and Denials Management workflow; complete a Durable Medical Equipment (DME) Order Processing workflow; satisfy each of the fine-grained verifiable subtasks a task decomposes into (often 15+ per task)
Must keep track of The agent must track progress across many fine-grained, cross-application subtasks (spanning EHR, payer portals, and fax) per administrative workflow, under a fixed interaction budget, to reach a correct terminal workflow state.
structure: DAG-with-precedence · goals arrive: given-up-front · count: 135 tasks yielding 1,698 evaluation poin · subgoal credit: yes · 8 citations · also Healthcare &amp; clinical · unit: other:not-stated
Horizonthe paper states none
the paper’s own words
Each task is decomposed into fine-grained, verifiable subtasks, yielding 1,698 evaluation points. ... the best-performing agent (Claude Opus 4.6 CUA) achieves only 36.3 percent task success, while GPT-5.4 CUA attains the highest subtask success rate (82.8 percent).
Since many tasks involve 15 or more subtasks ... agents operate 'under a fixed interaction budget' and references 'maximum steps,' but does not specify an exact number.
Herculean2026relevant
Herculean: An Agentic Benchmark for Financial Intelligence
Goals to track complete Trading workflow tasks via a standardized MCP-based skill environment; complete Hedging workflow tasks requiring long-horizon coordination and state consistency; complete Market Insights workflow tasks; complete Auditing workflow tasks requiring structured verification
Must keep track of The agent must track workflow-specific state, tool interactions, and constraints within each MCP-based skill environment, with Hedging and Auditing additionally requiring sustained state consistency and structured verification across long-horizon coordination.
structure: set-of-independent · goals arrive: given-up-front · count: 4 representative financial workflows (Tr · subgoal credit: yes · 3 citations
Horizonthe paper states none
the paper’s own words
We introduce Herculean, the first skilled benchmark for agentic financial intelligence spanning four representative workflows, including Trading, Hedging, Market Insights, and Auditing. Each workflow is instantiated as a standardized MCP-based skill environment with its own tools, interaction dynamics, constraints, and success criteria, enabling consistent end-to-end assessment of heterogeneous agent systems.
where long-horizon coordination, state consistency, and structured verification are critical
JobBench2026relevant
JobBench: Aligning Agent Work With Human Will
Goals to track complete 130 agentic tasks spanning 35 occupations that experts identify as high-priority for delegation; reason through cluttered, heterogeneous reference-file information streams within a professional workspace; satisfy a fact-anchored chain of rubrics averaging 35.6 binary criteria per task
Must keep track of Agent must correctly reason through a workspace of heterogeneous reference files and track satisfaction of an average of 35.6 fact-anchored binary rubric criteria per task.
structure: hierarchical · goals arrive: given-up-front · count: 130 tasks across 35 occupations; averagi · subgoal credit: yes · 4 citations
Horizonthe paper states none
the paper’s own words
Each task is packaged as a workspace of heterogeneous reference files, requiring the agent to reason through the cluttered information streams of real professional work. Outputs are graded by a fact-anchored chain of rubrics, averaging 35.6 binary criteria per task.
JobBench covers 130 agentic tasks across 35 occupations.
LH-Bench2026relevant
LH-Bench: Skill-Grounded Evaluation of Long-Horizon Agents on Subjective Enterprise Tasks
Goals to track produce subjective, context-dependent enterprise work whose quality depends on organizational goals and user intent, not a single correct answer; produce correct intermediate artifacts across long, multi-tool workflows (e.g., chapter-level content, Figma-to-code conversions); satisfy expert-grounded rubrics scoring subjective work quality; align with pairwise human preference judgments as convergent validation
Must keep track of The agent must track and produce correct intermediate artifacts (e.g., each chapter of a course, or each Figma-to-code conversion) across long, multi-tool workflows, since ground-truth artifacts enable stepwise reward signals at that granularity rather than only a single final judgment.
structure: hierarchical · goals arrive: given-up-front · count: two environments: Figma-to-code (33 real · subgoal credit: yes · 2 citations
Horizonthe paper states none
the paper’s own words
The pillars are: (i) expert-grounded rubrics that give LLM judges the domain context needed to score subjective work, (ii) curated ground-truth artifacts that enable stepwise reward signals (e.g., chapter-level annotation for content tasks), and (iii) pairwise human preference evaluation for convergent validation.
Programmatic content (41 courses comprising 183 individually-evaluated chapters on a course platform serving 30+ daily users)
Post-Training on Office Work Improves Software Engineering: A Behavioral Account of Cross-Domain Transfer
Goals to track select goals appropriately within nested and branching office work; construct task-relevant state as work unfolds; maintain fidelity to higher-level objectives across nested/branching sub-work; verify completion against the environment
Must keep track of Agent must repeatedly select goals, construct task-relevant state, maintain fidelity to higher-level objectives, and verify completion against the environment across nested and branching long-horizon office work.
structure: hierarchical · goals arrive: mixed:given-up-front-tasks-with-self-directed- · count: 363 Long-Horizon Multi-Tool Agent (LHMTA · subgoal credit: yes · 0 citations
Horizonthe paper states none
the paper’s own words
We call this capability goal-directed execution (GDE): the repeated application of four behaviors, namely selecting goals, constructing task-relevant state, maintaining fidelity to higher-level objectives, and verifying completion against the environment... We test this by post-training Qwen3.5-122B-A10B on 363 Long-Horizon Multi-Tool Agent (LHMTA) tasks drawn from office workflows.
LongHorizon-Bench2026relevant
Four-Axis Decision Alignment for Long-Horizon Enterprise AI Agents
Goals to track make high-stakes regulated decisions (loan qualification, insurance claims adjudication) under lossy memory and multi-step reasoning; maintain factual precision (FRP) about case facts over the long horizon; maintain reasoning coherence (RCS) across the multi-step decision process; reconstruct compliance/regulatory justification (CRR) for the decision; calibrate abstention (CAR) -- know when to decline to decide rathe
Must keep track of The agent must retain case facts accurately over a long-horizon multi-step review under lossy memory, maintain a coherent reasoning chain, reconstruct which regulatory rules justify the decision, and calibrate when to abstain rather than commit -- all against deterministic ground truth.
structure: set-of-independent · goals arrive: given-up-front · count: four alignment axes (FRP, RCS, CRR, CAR) · subgoal credit: yes · 1 citations
Horizonthe paper states none
the paper’s own words
We propose that long-horizon decision behavior decomposes into four orthogonal alignment axes, each independently measurable and failable: factual precision (FRP), reasoning coherence (RCS), compliance reconstruction (CRR), and calibrated abstention (CAR).
Long-horizon enterprise agents make high-stakes decisions (loan underwriting, claims adjudication, clinical review, prior authorization) under lossy memory, multi-step reasoning, and binding regulatory constraints.
LongMedBench: Benchmarking Medical Agents for Long-Horizon Clinical Decision-Making
Goals to track aggregate evidence across a patient's repeated visits, tests, and evolving treatments to answer fact-based QA; perform temporal reasoning over the patient's event stream, including implicit (not just explicit-timestamped) time inference; make a long-horizon clinical decision that correctly uses historical patient information accumulated over many visits
Must keep track of Agent must integrate time-series clinical events (admission records, notes) across many visits per patient, correctly distinguish explicit timestamps from cases requiring implicit time inference, and carry this forward into long-horizon decision-making.
structure: sequential-chain · goals arrive: emitted-by-environment-over-time · count: 335 patients, avg 19.72 inpatient visits · subgoal credit: unclear · 0 citations · also Healthcare &amp; clinical · unit: sessions
Horizon19.72 inpatient visits per patient (average); 44.91 medical events per visit (average)
the paper’s own words
It comprises 335 patients, with 19.72 inpatient visits per patient on average and 44.91 medical events per visit.
MBABench2026relevant
MBABench: Evaluating LLM Agents on End-to-End Spreadsheet Tasks in Finance
Goals to track construct an entire financial spreadsheet (e.g., financial model, forecast, scenario analysis) end-to-end from a high-level instruction; jointly satisfy Accuracy, Formula, and Format criteria, each with fine-grained sub-criteria reflecting professional standards
Must keep track of Agent must track intermediate calculation dependencies within the spreadsheet, professional formatting/readability conventions, and formula correctness simultaneously while building the deliverable end-to-end.
structure: set-of-independent · goals arrive: given-up-front · count: three top-level evaluation dimensions (A · subgoal credit: yes · 2 citations
Horizonthe paper states none
the paper’s own words
To reflect the multidimensional nature of solution quality, we develop an evaluation taxonomy comprising three dimensions: Accuracy, Formula, and Format, each comprising fine-grained criteria that reflect professional standards.
Evaluating over 18 agents, the benchmark reveals that even the strongest agents fall short of basic professional finance standards, and their performance degrade sharply as the difficulty increases beyond a few chained calculations.
MerchantBench: Benchmarking LLM Agents for Long-Term Coherence in E-Commerce Operations
Goals to track source products and manage upstream supplier events over time; set and adjust listing/pricing to remain competitive and solvent; manage cash-flow across delayed, heterogeneous-latency order outcomes; follow individual order lifecycles end-to-end and revisit earlier decisions as new (delayed) information arrives
Must keep track of The agent must follow individual order lifecycles end-to-end, track upstream supplier events and their promised delayed downstream outcomes, manage cash-flow/net-assets state, and revisit/adapt earlier sourcing and pricing decisions as delayed feedback arrives across a full 365-simulated-day run.
structure: DAG-with-precedence · goals arrive: emitted-by-environment-over-time · count: recurrent decisions across 4 named categ · subgoal credit: yes · 1 citations · unit: simulated-days
Horizona 365-day order-level simulation; 48 runs, each spanning 365 simulated days
the paper’s own words
Seller-side e-commerce provides a suitable setting for this evaluation through recurrent and interdependent decisions over Product Sourcing, Listing and Pricing Control, Cash-Flow Management, and Mixed-Latency Feedback Adaptation. We introduce MerchantBench, a 365-day order-level simulation grounded in 98,843 real e-commerce product records and equipped with 26 tools for agent interaction. MerchantBench couples promptly observable Upstream Supplier Events with delayed Downstream Order Outcomes, requiring agents to follow individual order lifecycles and revisit earlier decisions.
CORPGEN: Simulating Corporate Environments with Autonomous Digital Employees in Multi-Horizon Task Environments
Goals to track manage dozens of concurrent, interleaved long-horizon corporate tasks (45+ tasks, 500-1500+ steps each); handle inter-task dependencies expressed as DAGs rather than simple chains; reprioritize among concurrent tasks as load and context change over a persistent execution context spanning hours
Must keep track of The system must maintain hierarchical goal alignment, isolate sub-agent context to prevent cross-task contamination, and manage tiered (working/structured/semantic) memory with adaptive summarization across a persistent execution context spanning hours.
structure: DAG-with-precedence · goals arrive: mixed:given-up-front-with-reprioritization-emi · count: 45+ concurrent tasks, each requiring 500 · subgoal credit: no · 0 citations · unit: agent-steps
Horizon500-1500+
the paper’s own words
We identify four failure modes that cause baseline CUAs to degrade from 16.7% to 8.7% completion as load scales 25% to 100%, a pattern consistent across three independent implementations. These failure modes are context saturation (O(N) vs O(1) growth), memory interference, dependency complexity (DAGs vs. chains), and reprioritization overhead.
requiring coherent execution across dozens of interleaved tasks (45+, 500-1500+ steps) within persistent execution contexts spanning hours.
OR-Space2026relevant
OR-Space: A Full-Lifecycle Workspace Benchmark for Industrial Optimization Agents
Goals to track construct a solver-ready optimization model from heterogeneous business artifacts (Build); revise an existing model under changing requirements or solver feedback while preserving valid prior logic (Revise); answer grounded questions about solutions, constraints, and business implications using evidence spread across workspace artifacts (Explain)
Must keep track of The agent must track the current state of a persistent multi-artifact workspace (documents, structured data, code, solver outputs) across model construction, revision, and explanation stages, preserving valid prior modeling logic when new requirements or solver feedback arrive.
structure: DAG-with-precedence · goals arrive: mixed:given-up-front-with-new-requirements-inj · count: 3 task modes (Build, Revise, Explain) · subgoal credit: unclear · 2 citations
Horizonthe paper states none
the paper’s own words
OR-Space defines three task modes: Build, where agents construct solver-ready optimization models from heterogeneous artifacts; Revise, where agents modify existing models under changing requirements or solver feedback while preserving valid prior logic; and Explain, where agents answer grounded questions about solutions, constraints, and business implications using evidence spread across workspace artifacts.
persistent multi-artifact workspaces and multi-stage task lifecycles
OSWorkerBench2026relevant
UI-Mate: Advancing Open-Weight Foundation GUI Agents with In-Context Demonstrations
Goals to track complete each of 100 long-horizon office tasks spanning 41 applications; follow demonstrated subtask-level workflows (self-demo or variant-demo) and re-plan from the live interface when needed; make measurable progress toward task completion even when strict end-to-end success is not reached
Must keep track of Agent must track its position in a demonstrated or self-planned subtask sequence, detect when the live interface diverges from the demonstration, and re-plan accordingly across a long-horizon office task.
structure: sequential-chain · goals arrive: given-up-front · count: 100 long-horizon office tasks across 41 · subgoal credit: yes · 0 citations · also Web &amp; GUI agents
Horizonthe paper states none
the paper’s own words
OSWorkerBench Benchmark and Insights: A benchmark of 100 long-horizon office tasks across 41 applications that supports instruction-only and demonstration-guided evaluation.
OccuBench2026relevant
OccuBench: Evaluating AI Agents on Real-World Professional Tasks via Language Environment Simulation
Goals to track complete each of 100 real-world professional task scenarios spanning 65 specialized domains; maintain task completion under controlled fault injection (explicit errors, implicit data degradation, mixed faults)
Must keep track of Agent must track task-completion progress plus signals of environmental robustness (whether tool responses are timing out, truncated, or subtly degraded) within each professional scenario.
structure: set-of-independent · goals arrive: given-up-front · count: 100 real-world professional task scenari · subgoal credit: yes · 4 citations
Horizonthe paper states none
the paper’s own words
OccuBench evaluates agents along two complementary dimensions: task completion across professional domains and environmental robustness under controlled fault injection (explicit errors, implicit data degradation, and mixed faults).
We introduce OccuBench, a benchmark covering 100 real-world professional task scenarios across 10 industry categories and 65 specialized domains, enabled by Language Environment Simulators (LESs) that simulate domain-specific environments through LLM-driven tool response generation.
OmegaUse-OfficeVal2026relevant
OmegaUse-OfficeVal: Benchmarking LLM Agents on Long-Horizon Office-Suite Tasks with Economic Grounding
Goals to track complete 100 practitioner-derived office-suite tasks end-to-end; achieve deliverable quality verified via fine-grained rubric-based code verifiers; remain economically competitive relative to human labor time and task price
Must keep track of Agent must track fine-grained rubric criteria and produce a deliverable matching the practitioner's request, verified via code-based verifiers, implicitly benchmarked against the human labor time (avg. 2.32 hours) needed for the same task.
structure: set-of-independent · goals arrive: given-up-front · count: 100 tasks derived from practitioner offi · subgoal credit: yes · 0 citations · unit: human-expert-hours
Horizon2.32 (average)
the paper’s own words
To support stable evaluation, we develop code-based verifiers from fine-grained rubrics. We evaluate several frontier LLMs together with a human baseline. Although all evaluated LLMs are substantially cheaper and faster than human workers, they have not yet approached human-level deliverable quality.
The benchmark comprises 100 tasks derived from office-suite requests proposed by practitioners and adapted through a privacy-preserving process. On average, these tasks require 2.32 hours of human labor to complete.
OptAgent2026relevant
OptAgent: an Agentic AI framework for Intelligent Building Operations
Goals to track assess how a system/control upgrade changes energy use; assess how the same upgrade changes operating cost; assess how the same upgrade changes thermal comfort; assess how the same upgrade changes flexibility, via coordinated multi-domain multi-agent analytics
Must keep track of The orchestrator must track which specialist agent/tool has been invoked for which sub-domain (thermal dynamics, HVAC, DER), intermediate physics-informed simulation outputs, and how upgrades propagate across energy use, cost, comfort, and flexibility metrics within one workflow.
structure: DAG-with-precedence · goals arrive: given-up-front · count: large-scale benchmark of about 4,000 run · subgoal credit: no · 1 citations · unit: other:not-stated
Horizonthe paper states none
the paper’s own words
an agentic AI layer with 11 specialist agents and 72 Model Context Protocol (MCP) tools that enable end-to-end execution of multi-step energy analytics. A representative case study demonstrates multi-domain, multi-agent coordination for assessing how system and control upgrades affect energy use, operating cost, thermal comfort, and flexibility.
a large-scale benchmark (about 4000 runs) systematically evaluates workflow performance in terms of accuracy, token consumption, execution time, and inference cost
Evidence-Grounded AI for Musculoskeletal Care
Goals to track integrate evolving imaging, laboratory, pathology, and order data as it arrives across visits; produce evidence-based decisions at each stage of care, from admission diagnosis through rehabilitation planning; maintain continuous, individualised management across the full musculoskeletal care pathway rather than isolated per-visit decisions
Must keep track of System must continuously retrieve and integrate real-time imaging, laboratory, pathology, and order data across visits/departments/hospital systems, translating evolving patient state into stage-specific functional goals across the whole care pathway.
structure: sequential-chain · goals arrive: emitted-by-environment-over-time · subgoal credit: yes · 0 citations · also Healthcare &amp; clinical · unit: other:real-world-months-to-years-per-patient-pathway
Horizonmonths to years (per patient pathway); 1,870 cases / 8,240 inpatients across the study
the paper’s own words
Clinicians must repeatedly integrate evolving patient evidence, medical knowledge and stage-specific functional goals, yet evidence is often fragmented across visits, departments and hospital systems, disrupting continuous, individualised management.
recovery, remodelling and degeneration of bones, joints and related tissues unfold over months to years, care requires longitudinal management rather than isolated decisions
POLARIS2026relevant
POLARIS: Typed Planning and Governed Execution for Agentic AI in Back-Office Automation
Goals to track synthesize a type-checked directed acyclic graph (DAG) plan for a back-office document-processing task; select a single compliant plan via rubric-guided reasoning among structurally diverse candidate DAGs; pass validator-gated checks and a bounded repair loop before execution; route or block side effects per compiled policy guardrails, including anomaly routing
Must keep track of System must track the candidate DAG structure, per-node type/validator status, and the full execution trace/audit trail to support bounded repair and decision-grade anomaly routing.
structure: DAG-with-precedence · goals arrive: given-up-front · subgoal credit: yes · 2 citations
Horizonthe paper states none
the paper’s own words
A planner proposes structurally diverse, type checked directed acyclic graphs (DAGs), a rubric guided reasoning module selects a single compliant plan, and execution is guarded by validator gated checks, a bounded repair loop, and compiled policy guardrails that block or route side effects before they occur.
Applied to document centric finance tasks, POLARIS produces decision grade artifacts and full execution traces while reducing human intervention.
PPT-Eval2026relevant
PPT-Eval: A Benchmark for Computer-Use Agents on PowerPoint Tasks
Goals to track complete content-creation and presentation-editing tasks across 120 PowerPoint tasks in 12 files; satisfy task-specific rubric criteria that award partial credit for intermediate steps; avoid unnecessary changes and poor aesthetics while making required edits
Must keep track of Agent must track which task-specific rubric criteria (intermediate steps, aesthetics, unnecessary changes) have been satisfied across a multimodal editing session, since rubrics award partial credit rather than only final binary success.
structure: hierarchical · goals arrive: given-up-front · count: 120 PowerPoint tasks across 12 files, or · subgoal credit: yes · 5 citations · also Web &amp; GUI agents
Horizonthe paper states none
the paper’s own words
we design a robust evaluation framework to help create task-specific rubrics for PowerPoint tasks... These rubrics award partial credit for intermediate steps, penalize unnecessary changes and poor aesthetics, and provide natural language feedback. This nuanced approach proves highly effective, achieving a Kendall's tau-b correlation of 0.77 with human judgments.
We introduce PPT-Eval, a benchmark of 120 PowerPoint tasks across 12 files that cover both content creation and presentation editing scenarios, organized by difficulty.
PhysicianBench: Evaluating LLM Agents in Real-World EHR Environments
Goals to track retrieve relevant clinical data across multiple encounters in the EHR; reason over heterogeneous clinical information (labs, notes, orders) to reach a decision; execute consequential clinical actions (e.g. prescribing, ordering) grounded against the environment; produce clinical documentation reflecting the completed workflow
Must keep track of Agent must track data retrieved across multiple encounters, intermediate clinical reasoning state, and which structured checkpoints have been satisfied, using execution-grounded verification against real patient records via standard EHR APIs.
structure: hierarchical · goals arrive: given-up-front · count: 100 long-horizon tasks; 670 checkpoints · subgoal credit: yes · 14 citations · also Healthcare &amp; clinical · unit: tool-calls
Horizon27 (average)
the paper’s own words
Each task is decomposed into structured checkpoints (670 in total across the benchmark) capturing distinct stages of completion graded by task-specific scripts with execution-grounded verification.
requiring an average of 27 tool calls per task
PolyWorkBench2026relevant
PolyWorkBench: Benchmarking Multilingual Long-Horizon LLM Agents
Goals to track process heterogeneous multilingual inputs correctly within a workflow task; perform iterative reasoning and invoke external tools while maintaining linguistic consistency; produce a structured, correct output for one of five workplace domains (commerce, knowledge work, legal analysis, localization, manufacturing)
Must keep track of Agent must track functional correctness and linguistic consistency simultaneously across a workflow's reasoning and tool-invocation steps, since the two can drift independently as multilingual inputs are processed.
structure: sequential-chain · goals arrive: given-up-front · count: 67 tasks across 5 domains · subgoal credit: yes · 1 citations
Horizonthe paper states none
the paper’s own words
PolyWorkBench consists of 67 tasks across five domains, including commerce, knowledge work, legal analysis, localization, and manufacturing, where agents must process heterogeneous multilingual inputs, perform iterative reasoning, invoke external tools, and produce structured outputs.
PolyWorkBench2026relevant
PolyWorkBench: Benchmarking LLM Agents for Cross-Lingual Long-Horizon Workflows
Goals to track integrate heterogeneous multilingual inputs relevant to a workplace task; execute iterative tool-use trajectories within a target domain (commerce, knowledge work, legal analysis, localization, manufacturing); produce structured domain artifacts as verifiable task output
Must keep track of Agent must track and correctly integrate heterogeneous multilingual information while executing an iterative sequence of tool calls, and produce a structured domain artifact that is later verified.
structure: sequential-chain · goals arrive: given-up-front · count: 67 tasks across five core domains: comme · subgoal credit: yes · 0 citations
Horizonthe paper states none
the paper’s own words
we adopt Grade, a task-specific structural scoring rubric, as our primary ranking metric, and complement it with Pytest for executable state verification and LLM-as-Judge for semantic quality diagnostics.
enterprise workflows inherently require processing multilingual resources across extended trajectories. The interaction between multilinguality and long-horizon execution, however, remains underexplored.
PowerAgentBench-SS2026relevant
PowerAgentBench-SS: A Benchmark for Agentic AI in Power System Steady-State Studies
Goals to track inspect a grid case and select appropriate tools/simulators for the workflow; screen a large space of contingencies within a limited validation budget; propose admissible mitigations for discovered risks and validate their physical validity; produce an auditable evidence trail and submit a bounded, ranked report of top contingencies
Must keep track of The agent must track which contingencies it has already validated (against a fixed budget of 80 out of 1,035), the evidence log supporting each, and whether its running set of top candidates still reflects the best available evidence as it allocates its remaining validation budget.
structure: sequential-chain · goals arrive: given-up-front · count: 1,035 total N-2 contingency cases per in · subgoal credit: yes · 4 citations · unit: actions
Horizonvalidation budget of 80 contingency-case validations per episode (out of 1,035 total N-2 cases), submitting a ranked report of exactly 20 contingencies; hidden dangerous
the paper’s own words
The benchmark exposes public case data, action constraints, a tool API, and a validation budget to an agent, while a hidden evaluator recomputes physical validity and scores the submitted report. We define the agent interface, tool contract, evidence log, and risk-sensitive metrics, including submitted recall, evidence-backed recall, found recall, false-safe penalties, severity regret, residual violation score, action cost, tool-use efficiency, and workflow diagnostics.
The validation budget is B=80 ... The report size is m=20 ... the top 5% of cases by severity
RetailBench: Benchmarking long horizon reasoning and coherent decision making of LLM agents in realistic retail environments
Goals to track manage pricing across the store's product assortment; manage replenishment and supplier selection to keep inventory stocked; manage shelf assortment and inventory aging; respond appropriately to customer feedback and external events; maintain solvency (cash-flow constraints) while maximizing net worth/sales over a long simulated horizon
Must keep track of The agent must track pricing, inventory levels/aging, supplier relationships, customer feedback, external events, and its own cash-flow position day by day across the 180-day (or longer) simulated run, since only a small subset of evaluated agents survive the full evaluation horizon.
structure: sequential-chain · goals arrive: given-up-front · count: 7 contemporary LLMs evaluated under repr · subgoal credit: yes · 6 citations · unit: simulated-days
Horizon180-day evaluation horizon; the simulator supports thousand-day-scale simulations
the paper’s own words
RetailBench models retail management as a partially observable decision process and is designed to support thousand-day-scale simulations. In this environment, agents must manage pricing, replenishment, supplier selection, shelf assortment, inventory aging, customer feedback, external events, and cash-flow constraints. We evaluate seven contemporary LLMs under representative agent frameworks over a 180-day evaluation horizon and compare them with a privileged oracle policy.
RiskWebWorld2026relevant
RiskWebWorld: A Realistic Interactive Benchmark for GUI Agents in E-commerce Risk Management
Goals to track investigate flagged e-commerce risk cases across multiple verification sub-steps on production risk-control pipelines; operate GUI actions on uncooperative websites subject to partial environmental hijacking; complete each of 1,513 tasks spanning 8 core risk-control domains; succeed at long-horizon professional risk-investigation workflows
Must keep track of The agent must track evidence and verification state gathered across multiple sub-steps of a risk investigation on an uncooperative website, while detecting and coping with partial environmental hijacking attempts, across long-horizon professional tasks.
structure: sequential-chain · goals arrive: given-up-front · count: 1,513 tasks sourced from production risk · subgoal credit: no · 0 citations · also Web &amp; GUI agents
Horizonthe paper states none
the paper’s own words
Our evaluation across diverse models reveals a dramatic capability gap: top-tier generalist models achieve 49.1% success, while specialized open-weights GUI models lag at near-total failure.
This highlights that foundation model scale currently matters more than zero-shot interface grounding in long-horizon professional tasks.
SaaS-Bench2026relevant
SaaS-Bench: Can Computer-Use Agents Leverage Real-World SaaS to Solve Professional Workflows?
Goals to track navigate and operate real, deployed SaaS systems to complete a professional workflow; coordinate state and context across multiple applications within the same workflow; apply domain-specific knowledge correctly within the SaaS system; recover from errors and maintain progress over a long-horizon task
Must keep track of The agent must maintain state and context across multiple SaaS applications over long-horizon execution, tracking partial progress against weighted verification checkpoints, since 'agents become stuck acquiring information or manipulating interfaces, confuse source and output devices, or terminate before all conditions are jointly satisfied.'
structure: DAG-with-precedence · goals arrive: given-up-front · count: 106 tasks grounded in realistic work sce · subgoal credit: yes · 3 citations · unit: actions
Horizonaverage of over 100 interaction steps per task
the paper’s own words
SaaS-Bench, a benchmark built on 23 deployable SaaS systems across six professional domains, containing 106 tasks grounded in realistic work scenarios. These tasks require long-horizon execution, cover both text-only and multimodal settings, and are evaluated with weighted verification checkpoints that measure strict task completion and partial progress.
SaaS-Bench introduces long-horizon tasks with an average of over 100 interaction steps
SpreadsheetBench 22026relevant
SpreadsheetBench 2: Evaluating Agents on End-to-End Business Spreadsheet Workflows
Goals to track generate new spreadsheet content/formulas correctly across a large multi-sheet workbook; debug existing incorrect formulas/content within the workbook; produce correct visualizations from the workbook's data
Must keep track of The agent must track and correctly propagate changes across an average of 11.8 interdependent worksheets requiring 593.5 cell modifications per task, correctly identifying target cells under a unified multi-turn agent scaffold.
structure: DAG-with-precedence · goals arrive: given-up-front · count: 321 tasks; each instance averages 11.8 w · subgoal credit: yes · 4 citations · unit: other:cell-modifications
Horizoneach task instance averages 11.8 worksheets and requires 593.5 cell modifications
the paper’s own words
The benchmark contains 321 tasks; each instance averages 11.8 worksheets and requires 593.5 cell modifications, reflecting large multi-sheet workbooks with cross-sheet dependencies.
SupChain-Bench: Benchmarking Large Language Models for Real-World Supply Chain Management
Goals to track orchestrate long-horizon, multi-step supply-chain tool use grounded in standard operating procedures (SOPs); correctly apply supply-chain domain knowledge across a sequence of dependent tool calls
Must keep track of Agent must track SOP-grounded procedural state across a long-horizon sequence of dependent tool calls in order to correctly complete supply-chain management workflows.
structure: sequential-chain · goals arrive: given-up-front · subgoal credit: no · 0 citations
Horizonthe paper states none
the paper’s own words
supply chain workflows require reliable long-horizon, multi-step orchestration grounded in domain-specific procedures, which remains challenging for current models... we introduce SupChain-Bench, a unified real-world benchmark that assesses both supply chain domain knowledge and long-horizon tool-based orchestration grounded in standard operating procedures (SOPs).
Synthetic Computers at Scale for Long-Horizon Productivity Simulation
Goals to track complete each of multiple professional deliverables comprising a computer-specific productivity objective; navigate the synthetic computer's filesystem to ground actions in the user's actual context; coordinate with simulated collaborators as needed to complete the objective; sustain progress toward an objective requiring about a month of simulated human work
Must keep track of The acting agent must track the evolving state of a realistic folder hierarchy and content-rich artifacts (documents, spreadsheets, presentations) across a run spanning over 2,000 turns, coordinating with simulated collaborators toward multiple professional deliverables.
structure: open-ended · goals arrive: self-generated-by-agent · count: 1,000 synthetic computers; each run requ · subgoal credit: no · 1 citations · unit: turns
Horizon>2,000 turns on average (each run also requiring over 8 hours of agent runtime)
the paper’s own words
one agent creates productivity objectives that are specific to the computer's user and require multiple professional deliverables and about a month of human work; another agent then acts as that user and keeps working across the computer -- for example, navigating the filesystem for grounding, coordinating with simulated collaborators, and producing professional artifacts -- until these objectives are completed. ... each run requires over 8 hours of agent runtime and spans more than 2,000 turns on average.
Thinkingbox-bench2026relevant
One Success Isn't Reliability: Thinkingbox, a Sandbox and Benchmark for Agents in Stateful Business Workflows
Goals to track gather missing information over multiple turns before acting; follow domain-specific policies for each of 507 policy-conditioned workflows; coordinate dependent tools correctly; realize exactly the correct persistent backend state transition without collateral effects
Must keep track of Agent must gather missing information across multiple turns, track applicable domain policies, coordinate dependent tool calls, and verify the resulting persistent backend state matches exactly the required final state with no collateral effects.
structure: DAG-with-precedence · goals arrive: given-up-front · count: 507 policy-conditioned workflows across · subgoal credit: no · 0 citations · also Information seeking &amp; deep research
Horizonthe paper states none
the paper’s own words
Each attempt is evaluated by task-specific executable checks that accept valid trajectories while rejecting wrong, missing, or extra effects; designated tasks additionally check required properties of the final response... the strongest model Claude Opus 5 achieves 66.50% pass@1, but only 47.53% pass^20.
Thinkingbox-bench contains 507 policy-conditioned workflows across numerous scenarios, including retail, hospitality, auto insurance, neobank internal IT, and consulting IT/HR support.
Underwrite2026relevant
Benchmarking Agents in Insurance Underwriting Environments
Goals to track gather information carefully from noisy tool interfaces and imperfect simulated users during an underwriting conversation; apply proprietary business/domain knowledge correctly rather than hallucinating it; reach a final underwriting decision consistent with the accumulated evidence gathered across the conversation
Must keep track of The agent must track what information it has already gathered (and from where), the reliability of that information given noisy tool interfaces, and how it should update its evolving underwriting assessment as more evidence accumulates across the conversation.
structure: sequential-chain · goals arrive: given-up-front · subgoal credit: yes · 0 citations · also Information seeking &amp; deep research · unit: turns
Horizonaverage of 3-7 steps of required reasoning and tool use, with a total of 10-20 conversational turns
the paper’s own words
We present Underwrite, an expert-first, multi-turn insurance underwriting benchmark designed in close collaboration with domain experts to capture real-world enterprise challenges. Underwrite introduces critical realism factors often absent in current benchmarks: proprietary business knowledge, noisy tool interfaces, and imperfect simulated users requiring careful information gathering.
We aimed for an average of 3-7 steps of required reasoning and tool use, with a total of 10-20 conversational turns.
Workflow-GYM: Towards Long-Horizon Evaluation of Computer-use Agentic tasks in Real-World Professional Fields
Goals to track autonomously operate domain-specific professional software GUIs to accomplish economically valuable work; complete long-horizon, multi-stage professional workflows end-to-end; maintain workflow consistency across stages without omission, error propagation, or objective drift
Must keep track of The agent must track its position and completed stages within a long-horizon, multi-stage professional GUI workflow, avoiding stage omission, propagating errors, or drifting from the original objective across the workflow's duration.
structure: sequential-chain · goals arrive: given-up-front · subgoal credit: no · 2 citations
Horizonthe paper states none
the paper’s own words
Through extensive experiments on state-of-the-art models, we find that even the strongest models achieve only slightly above 30% success rates, highlighting that professional long-horizon GUI workflows remain highly challenging for current GUI agents.
existing benchmarks rarely evaluate whether agents can operate graphical user interfaces to complete long-horizon, high-value professional workflows across diverse domains
Workspace-Bench 1.0: Benchmarking AI Agents on Workspace Tasks with Large-Scale File Dependencies
Goals to track identify, reason over, exploit, and update explicit and implicit dependencies among heterogeneous files in a worker's workspace; satisfy each of a task's own file-dependency-graph-derived rubrics via cross-file retrieval, contextual reasoning, and adaptive decision-making
Must keep track of Agent must track which of up to 20,476 files across 74 file types are relevant, their dependency relationships, and adaptively retrieve/reason across files to satisfy each task's rubrics.
structure: DAG-with-precedence · goals arrive: given-up-front · count: 388 tasks, each with its own file depend · subgoal credit: yes · 10 citations
Horizonthe paper states none
the paper’s own words
We construct realistic workspaces with 5 worker profiles, 74 file types, 20,476 files (up to 20GB) and curate 388 tasks, each with its own file dependency graph, evaluated across 7,399 total rubrics that require cross-file retrieval, contextual reasoning, and adaptive decision-making.
We further provide Workspace-Bench-Lite, a 100-task subset that preserves the benchmark distribution while reducing evaluation costs by about 70%.
World of Workflows: a Benchmark for Bringing World Models to Enterprise Systems
Goals to track complete constrained agentic tasks within a ServiceNow environment governed by 4,000+ business rules and 55 active hidden workflows; predict cascading side effects of actions across interconnected databases; mentally simulate hidden state transitions to avoid silent constraint violations under limited observability
Must keep track of Agent must mentally simulate hidden state transitions and predict cascading side effects across interconnected databases to bridge the observability gap, since high-fidelity feedback is often unavailable.
structure: DAG-with-precedence · goals arrive: given-up-front · count: 234 tasks (WoW-bench); 4,000+ business r · subgoal credit: no · 7 citations
Horizonthe paper states none
the paper’s own words
We introduce World of Workflows (WoW), a realistic ServiceNow-based environment incorporating 4,000+ business rules and 55 active workflows embedded in the system, alongside WoW-bench, a benchmark of 234 tasks evaluating constrained agentic task completion and enterprise dynamics modeling capabilities.
YC-Bench2026in
YC-Bench: Benchmarking AI Agents for Long-Term Planning and Consistent Execution
Goals to track manage employees within a simulated startup; select task contracts to pursue under uncertainty; maintain profitability against adversarial clients and growing payroll; detect and avoid bankruptcy-inducing failure modes (e.g. adversarial-client mismanagement, over-parallelization) over a one-year run
Must keep track of The agent must persist information across context truncation via a scratchpad (the strongest predictor of success), tracking employees, contracts, cash reserves, and adversarial-client risk across hundreds of turns spanning a simulated year.
structure: open-ended · goals arrive: given-up-front · subgoal credit: no · 8 citations · unit: turns
Horizonhundreds of turns (over a simulated one-year horizon)
the paper’s own words
we task an agent with running a simulated startup over a one-year horizon spanning hundreds of turns. The agent must manage employees, select task contracts, and maintain profitability in a partially observable environment where adversarial clients and growing payroll create compounding consequences for poor decisions.
we introduce YC-Bench, a benchmark that evaluates these capabilities by tasking an agent with running a simulated startup over a one-year horizon spanning hundreds of turns.
What Could the Agent See at 19:05? Generating Temporal Enterprise Scenarios from Real Research and Replaying Them to Evaluate Agents
Goals to track answer questions correctly relative to what data existed and who could see it at a specific queried moment; reason about a persona-driven, temporally-evolving enterprise world spanning many apps; avoid leaking future/hidden record state when reasoning about an earlier moment
Must keep track of The agent being evaluated must reason correctly about what data existed and who could see it at a specific queried moment, without conflating it with earlier or later states of the same records across multiple apps.
structure: DAG-with-precedence · goals arrive: emitted-by-environment-over-time · subgoal credit: yes · 0 citations
Horizonthe paper states none
the paper’s own words
Our system closes two gaps at once: it generates a realistic, persona-driven, temporally-evolving enterprise world from real research, and replays that world at any chosen moment to evaluate any pluggable agent.
Behavioral Consistency Validation for LLM Agents: An Analysis of Trading-Style Switching through Stock-Market Simulation
Goals to track trade under a currently assigned style (fundamental or technical) each trading day; periodically (every 10 trading days) reassess and potentially switch trading style based on four behavioral-finance drivers: loss aversion, herding, wealth differentiation, price misalignment; keep style-switching behavior consistent with real-world behavioral-finance theory over the whole simulated year
Must keep track of Agent must process daily price-volume data, retain long-term personality traits (the four behavioral-finance drivers) set at initialization, and track its own accumulated wealth/strategy history to decide whether to switch trading style every 10 days.
structure: sequential-chain · goals arrive: given-up-front · subgoal credit: unclear · 4 citations · unit: simulated-days
Horizonyear-long (reassessed every 10 trading days)
the paper’s own words
In year-long simulations, agents process daily price-volume data, trade under a designated style, and reassess their strategy every 10 trading days.
AI-Trader2025relevant
AI-Trader: Benchmarking Autonomous Agents in Real-Time Financial Markets
Goals to track independently search, verify, and synthesize live market information given only minimal initial context; make live trading decisions (buy/sell/hold) across U.S. stocks, A-shares, and cryptocurrencies at multiple trading granularities; manage risk and sustain positive returns over a continuous live-trading period
Must keep track of Agent must track evolving live market data it has independently searched/verified, current positions/risk exposure across three markets and multiple trading frequencies, and running returns over the continuous live-trading period.
structure: sequential-chain · goals arrive: emitted-by-environment-over-time · subgoal credit: yes · 19 citations
Horizonthe paper states none
the paper’s own words
Our benchmark implements a revolutionary fully autonomous minimal information paradigm where agents receive only essential context and must independently search, verify, and synthesize live market information without human intervention.
AI-Trader spans three major financial markets: U.S. stocks, A-shares, and cryptocurrencies, with multiple trading granularities to simulate live financial environments.
Agent Trading Arena: A Study on Numerical Understanding in LLM-Based Agents
Goals to track make sequential buy/sell trading decisions that directly affect and are affected by shared market prices; compete against other LLM-based agents in a zero-sum stock market; maximize trading performance/returns, especially under high volatility; correctly perform numerical reasoning over price data or chart-based visualizations
Must keep track of Each agent must track its own portfolio/capital, the current market price state shaped by all agents' recent trades, and historical price patterns, updating its numerical reasoning as the shared market evolves turn to turn.
structure: open-ended · goals arrive: given-up-front · subgoal credit: no · 12 citations
Horizonthe paper states none
the paper’s own words
we present the Agent Trading Arena, a virtual zero-sum stock market in which LLM-based agents engage in competitive multi-agent trading and directly impact price dynamics.
Existing approaches are limited to historical backtesting, where trading actions cannot influence market prices and agents train only on static data.
AssetOpsBench2025relevant
AssetOpsBench: A Real-World Evaluation Benchmark for AI-Driven Task Automation in Industrial Asset Management
Goals to track orchestrate the correct domain-specific agent(s) (from a catalog of four) to answer a natural-language industrial-operations query; correctly complete condition-monitoring and maintenance-scheduling workflow steps grounded in a simulated IoT environment
Must keep track of The agent must track intermediate tool/agent outputs across a multi-step think-act-observe loop, orchestrate the four domain-specific agents appropriately, and maintain consistency with the simulated CouchDB-backed IoT environment's evolving state.
structure: DAG-with-precedence · goals arrive: given-up-front · count: 140+ human-authored queries; Plan-Execut · subgoal credit: yes · 17 citations · unit: agent-steps
HorizonPlan-Execute agents complete most tasks in approximately 2.6-4.4 steps; Agent-As-Tool agents typically require approximately 4-6+ steps due to its iterative think-act-obs
the paper’s own words
AssetOpsBench provides a multimodal ecosystem comprising a catalog of four domain-specific agents, a curated dataset of 140+ human-authored natural-language queries grounded in real industrial scenarios, and a simulated, CouchDB-backed IoT environment. We introduce an automated evaluation framework that uses three key metrics to analyze architectural trade-offs between the Agent-As-Tool and Plan-Execute paradigms, along with a systematic procedure for the automated discovery of emerging failure modes.
Plan-Execute is consistently more step-efficient: most models complete tasks in ≈2.6-4.4 steps, whereas Agent-As-Tool typically requires ≈4-6+ steps due to its iterative think-act-observe loop.
AutoDW2025relevant
Automating Complex Document Workflows via Stepwise and Rollback-Enabled Operation Orchestration
Goals to track execute a sequence of interdependent, user-specified document-editing instructions within a session; keep the execution trajectory aligned with evolving document state and user intent across the whole session; recover from failed API calls/arguments via rollback at both the argument and API level
Must keep track of System must track the evolving document state after each API action, the remaining interdependent instructions in the session, and whether any prior action needs to be rolled back (at argument or API level) to stay aligned with user intent.
structure: DAG-with-precedence · goals arrive: given-up-front · count: 1,708 human-annotated instructions acros · subgoal credit: yes · 0 citations · unit: actions
Horizon1,708 instructions across 250 sessions (~6.8 instructions/session on average)
the paper’s own words
AutoDW achieves 90% and 62% completion rates on instruction- and session-level tasks, respectively, outperforming strong baselines by 40% and 76%.
we construct a comprehensive benchmark of 250 sessions and 1,708 human-annotated instructions, reflecting realistic document processing scenarios with interdependent instructions
CP-Env2025in
CP-Env: Evaluating Large Language Models on Clinical Pathways in a Controllable Hospital Environment
Goals to track triage a patient correctly to determine care pathway; consult the appropriate specialist given patient information; order and interpret diagnostic tests as needed; participate appropriately in multidisciplinary team meetings; complete a patient's full journey across branching, long-horizon clinical-pathway stages
Must keep track of The agent must track patient information, diagnostic findings, and evolving care-pathway state across branching stages from triage through specialist consultation, diagnostic testing, and multidisciplinary team meetings.
structure: DAG-with-precedence · goals arrive: mixed:overall-pathway-structure-given-up-front · subgoal credit: yes · 6 citations · also Healthcare &amp; clinical
Horizonthe paper states none
the paper’s own words
CP-Env simulates a hospital ecosystem with patient and physician agents, constructing scenarios ranging from triage and specialist consultation to diagnostic testing and multidisciplinary team meetings for agent interaction. Following real hospital adaptive flow of healthcare, it enables branching, long-horizon task execution.
CRMArena-Pro2025relevant
CRMArena-Pro: Holistic Assessment of LLM Agents Across Diverse Business Scenarios and Interactions
Goals to track complete each of nineteen expert-validated CRM tasks across sales, service, and configure-price-quote (CPQ) processes; sustain correct behavior across multi-turn interactions guided by diverse personas; maintain confidentiality awareness throughout the interaction, for both B2B and B2C scenarios
Must keep track of The agent must track persona-specific context, confidentiality constraints, and workflow state across multi-turn interactions, since single-turn success (58%) drops substantially to about 35% once genuine multi-turn tracking is required.
structure: sequential-chain · goals arrive: given-up-front · count: nineteen expert-validated tasks across s · subgoal credit: yes · 42 citations
Horizonthe paper states none
the paper’s own words
It distinctively incorporates multi-turn interactions guided by diverse personas and robust confidentiality awareness assessments... Experiments reveal leading LLM agents achieve only around 58% single-turn success on CRMArena-Pro, with performance dropping significantly to approximately 35% in multi-turn settings. While Workflow Execution proves more tractable for top agents (over 83% single-turn success), other evaluated business skills present greater challenges.
DRBench2025relevant
DRBench: A Realistic Benchmark for Enterprise Deep Research
Goals to track identify supporting facts for a multi-step research query from both the public web and a private company knowledge base; synthesize facts drawn from heterogeneous enterprise sources (productivity software, cloud file systems, emails, chat, web) into one answer; produce a coherent, well-structured final report grounded in the retrieved facts
Must keep track of The agent must track which facts it has found in which source (public web vs. private enterprise knowledge base), maintain factual accuracy across sources, and assemble a coherent report structure from these accumulated facts.
structure: hierarchical · goals arrive: given-up-front · count: 100 deep research tasks across 10 domain · subgoal credit: yes · 15 citations
Horizonthe paper states none
the paper’s own words
DRBench evaluates agents on multi-step queries ... that require identifying supporting facts from both the public web and private company knowledge base. Each task is grounded in realistic user personas and enterprise context, spanning a heterogeneous search space that includes productivity software, cloud file systems, emails, chat conversations, and the open web... agents are evaluated on their ability to recall relevant insights, maintain factual accuracy, and produce coherent, well-structured reports.
We release 100 deep research tasks across 10 domains, such as Sales, Cybersecurity, and Compliance.
DevNous2025relevant
DevNous: An LLM-Based Multi-Agent System for Grounding IT Project Management in Unstructured Conversation
Goals to track identify actionable intents from informal, unstructured team-chat dialogue; manage stateful, multi-turn administrative workflows (task formalization, progress-summary synthesis) grounded in that dialogue
Must keep track of The agent must identify actionable intents across informal chat and maintain stateful workflows (task formalization, progress tracking) across a benchmark of 160 conversational turns, correctly matching a multi-label ground truth per turn.
structure: sequential-chain · goals arrive: emitted-by-environment-over-time · count: 160 realistic, interactive conversationa · subgoal credit: yes · 0 citations · also Multi-agent organisations &amp; societies · unit: turns
Horizona new benchmark of 160 realistic, interactive conversational turns
the paper’s own words
We introduce DevNous, a Large Language Model-based (LLM) multi-agent expert system, to automate this unstructured-to-structured translation process. DevNous integrates directly into team chat environments, identifying actionable intents from informal dialogue and managing stateful, multi-turn workflows for core administrative tasks like automated task formalization and progress summary synthesis. To quantitatively evaluate the system, we introduce a new benchmark of 160 realistic, interactive conversational turns.
EcomBench2025relevant
EcomBench: Towards Holistic Evaluation of Foundation Agents in E-commerce
Goals to track retrieve deep, possibly multi-hop information relevant to a real e-commerce user demand; perform multi-step reasoning across e-commerce-domain data; integrate knowledge from multiple sources to resolve a task; correctly complete tasks across three graded difficulty levels
Must keep track of The agent must track partial evidence gathered from deep retrieval and multiple knowledge sources across many action steps before it can complete a higher-difficulty task.
structure: hierarchical · goals arrive: given-up-front · count: three difficulty levels; task categories · subgoal credit: yes · 7 citations
Horizonthe paper states none
the paper’s own words
It covers multiple task categories within e-commerce scenarios and defines three difficulty levels that evaluate agents on key capabilities such as deep information retrieval, multi-step reasoning, and cross-source knowledge integration.
Level 3 tasks cannot be solved in just a few action steps
Finch: Benchmarking Finance & Accounting across Spreadsheet-Centric Enterprise Workflows
Goals to track interleave data entry, structuring, formatting, web search, cross-file retrieval, calculation, modeling, validation, translation, visualization, and reporting subtasks into one finance/accounting workflow; correctly complete each of the linked tasks that compose one of 172 composite workflows
Must keep track of Agent must track state across many interlinked spreadsheets, PDFs, and artifacts (27 million spreadsheet cells total) while interleaving retrieval, calculation, modeling, validation, and reporting sub-steps of one composite workflow.
structure: DAG-with-precedence · goals arrive: given-up-front · count: 172 composite workflows with 384 tasks, · subgoal credit: yes · 9 citations · unit: wall-clock-minutes
Horizon16.8 (GPT-5.1 Pro average)
the paper’s own words
This yields 172 composite workflows with 384 tasks, involving 1,710 spreadsheets with 27 million cells, along with PDFs and other artifacts, capturing the intrinsically messy, long-horizon, knowledge-intensive, and collaborative nature of real-world enterprise work.
Under human evaluation, GPT-5.1 Pro spends an average of 16.8 minutes per workflow yet passes only 38.4% of workflows.
GraphicBench2025relevant
GraphicBench: A Planning Benchmark for Graphic Design with Language Agents
Goals to track produce a workflow plan satisfying explicit design constraints stated in a user query; also satisfy implicit commonsense design constraints not stated by the user; select the correct action from 46 available tools at each workflow step; coordinate outputs across three design experts without violating global dependencies
Must keep track of The agent must track explicit and implicit design constraints, the evolving multi-step workflow plan across three design experts, and which of 46 actions remains valid/appropriate at each step.
structure: DAG-with-precedence · goals arrive: mixed:explicit-constraints-given-up-front-plus · count: 1,079 user queries; 46 available actions · subgoal credit: yes · 2 citations · also Web &amp; GUI agents
Horizonthe paper states none
the paper’s own words
We introduce GraphicBench, a new planning benchmark for graphic design that covers 1,079 user queries and input images across four design types. We further present GraphicTown, an LLM agent framework with three design experts and 46 actions (tools) to choose from for executing each step of the planned workflows in web environments.
HiMA-Ecom2025relevant
HiMA-Ecom: Enabling Joint Training of Hierarchical Multi-Agent E-commerce Assistants
Goals to track master agent coordinates multiple specialized sub-agents in an e-commerce workflow; each sub-agent pursues a functionally distinct role-specific objective (e.g. recall domain knowledge, execute a service function); jointly optimize system-level behavior via multi-agent reinforcement learning
Must keep track of The master agent must track which specialized sub-agent is responsible for which functional sub-task and how sub-agent outputs compose into overall system behavior, using a collaboratively updated memory across training.
structure: hierarchical · goals arrive: given-up-front · subgoal credit: no · 8 citations · also Multi-agent organisations &amp; societies · unit: other:not-stated
Horizonthe paper states none
the paper’s own words
HiMA-Ecom contains 22.8K instances, including agent-specific supervised fine-tuning samples with memory and system-level input-output pairs for joint multi-agent reinforcement learning. ... a master agent coordinates multiple specialized sub-agents
MEMTRACK2025relevant
MEMTRACK: Evaluating Long-Term Memory and State Tracking in Multi-Platform Dynamic Agent Environments
Goals to track acquire relevant facts scattered across asynchronous cross-platform events (Slack/Linear/Git); select the correct, currently-valid fact when multiple conflicting/noisy candidates exist; resolve conflicts between contradictory or stale cross-referring information over time
Must keep track of The agent must acquire, select, and reconcile facts from a chronologically platform-interleaved timeline spanning Slack, Linear, and Git, handling noisy, conflicting, and cross-referring information plus codebase/file-system exploration, across long horizons.
structure: DAG-with-precedence · goals arrive: emitted-by-environment-over-time · subgoal credit: yes · 12 citations · also Personal assistants &amp; long-term memory
Horizonthe paper states none
the paper’s own words
Consequently, our benchmark tests memory capabilities such as acquistion, selection and conflict resolution... We introduce pertinent metrics for Correctness, Efficiency, and Redundancy that capture the effectiveness of memory mechanisms beyond simple QA performance.
Each benchmark instance provides a chronologically platform-interleaved timeline, with noisy, conflicting, cross-referring information as well as potential codebase/file-system comprehension and exploration.
Mini Amusement Parks (MAPs): A Testbed for Modelling Business Decisions
Goals to track maximize the amusement park's overall value over the given time horizon; make daily operational decisions (building rides/shops, hiring staff, setting a research agenda) that keep the business viable under sparse, stochastic feedback; reason over the park's spatial layout while planning these decisions
Must keep track of The agent must track the park's spatial layout, built rides/shops, staffing, accumulated (sparse) experience about environment dynamics, and overall park value across a 50-100 day episode, planning under uncertainty at each of many daily decision points.
structure: sequential-chain · goals arrive: given-up-front · count: episodes span a 50-day horizon on easy m · subgoal credit: yes · 2 citations · unit: simulated-days
Horizonepisodes span a 50-day horizon on easy mode, extended to 100 days on medium mode
the paper’s own words
To this end, we introduce Mini Amusement Parks (MAPs), an amusement-park simulator designed to evaluate an agent's ability to model its environment, anticipate long-term consequences under uncertainty, and strategically operate a complex business. We provide expert human performance and a comprehensive evaluation of state-of-the-art agents, finding experts outperform these systems by 11.4x on easy mode and 15.3x on medium mode.
the horizon is extended from 50 to 100
OdysseyBench: Evaluating LLM Agents on Long-Horizon Complex Office Application Workflows
Goals to track identify essential information buried in long-horizon interaction histories; perform multi-step reasoning/actions across Word, Excel, PDF, Email, and Calendar applications; complete real-world-derived tasks (OdysseyBench+) or newly synthesized complex tasks (OdysseyBench-Neo)
Must keep track of Agent must retain and retrieve essential facts from long-horizon interaction histories spanning multiple applications in order to correctly execute later multi-step, cross-application actions.
structure: sequential-chain · goals arrive: given-up-front · count: 300 tasks (OdysseyBench+); 302 tasks (Od · subgoal credit: unclear · 51 citations
Horizonthe paper states none
the paper’s own words
Our benchmark comprises two complementary splits: OdysseyBench+ with 300 tasks derived from real-world use cases, and OdysseyBench-Neo with 302 newly synthesized complex tasks. Each task requires agent to identify essential information from long-horizon interaction histories and perform multi-step reasoning across various applications.
PPTArena2025relevant
PPTArena: A Benchmark for PowerPoint Editing
Goals to track apply each of a deck's human-curated edits correctly from natural-language instructions; maintain layout-sensitive and cross-slide consistency across the whole deck while editing; plan and verify edit sequences via an iterative plan-edit-check loop (for the PPTPilot agent)
Must keep track of Agent must track the deck's current structural/visual state after each edit, verify each edit against a ground-truth rubric, and maintain deck-wide consistency (styles, cross-slide references) across the full sequence of edits.
structure: sequential-chain · goals arrive: given-up-front · count: over 1,300 human-curated edits across 10 · subgoal credit: yes · 7 citations · unit: actions
Horizon~13 edits per deck (1,300+ edits across 100 decks)
the paper’s own words
PPTArena features 100 decks with over 1,300 human-curated edits across 2,125 slides, spanning text, charts, animations, and professional master styles.
ProSoftArena: Benchmarking Hierarchical Capabilities of Multimodal Agents in Professional Software Environments
Goals to track operate professional software tools according to a hierarchical capability level (L1-L3); complete realistic work/research tasks spanning 6 disciplines and 13 core professional applications; coordinate across multiple professional software applications for L3 multi-software workflows
Must keep track of The agent must track its progress within and across professional-software applications as required by the task's capability level, since L3 tasks require coordinating state across multiple applications rather than operating one application in isolation.
structure: hierarchical · goals arrive: given-up-front · count: 436 realistic work and research tasks sp · subgoal credit: yes · 5 citations · unit: actions
Horizonhuman execution steps grow from an average of 5.1 steps (14.8 seconds) at L1 to 86.9 steps (506.8 seconds) at L3
the paper’s own words
We establish the first capability hierarchy tailored to agent use of professional software and construct a benchmark of 436 realistic work and research tasks spanning 6 disciplines and 13 core professional applications... Extensive experiments show that even the best-performing agent attains only a 24.4% success rate on L2 tasks and completely fails on L3 multi-software workflow.
human execution steps and time... from an average of 5.1 steps and 14.8 seconds at L1 to 86.9 steps and 506.8 seconds at L3
REALM-Bench: A Benchmark for Evaluating Multi-Agent Systems on Real-world, Dynamic Planning and Scheduling Tasks
Goals to track solve each of 14 real-world planning/scheduling problems with multiple parallel planning threads; maintain feasibility across inter-agent dependencies as the problem scales in complexity; adapt schedules in real time when unexpected disruptions arrive
Must keep track of Agents must track the state of parallel planning threads, inter-dependencies between agents/threads, and disruption events in order to replan in real time.
structure: DAG-with-precedence · goals arrive: mixed:given-up-front-problems-plus-injected-mi · count: 14 planning and scheduling problems · subgoal credit: yes · 12 citations · also Multi-agent organisations &amp; societies
Horizonthe paper states none
the paper’s own words
The suite encompasses 14 designed planning and scheduling problems that progress from basic to highly complex, incorporating key aspects such as multi-agent coordination, inter-agent dependencies, and dynamic environmental disruptions.
Each problem can be scaled along three dimensions: the number of parallel planning threads, the complexity of inter-dependencies, and the frequency of unexpected disruptions requiring real-time adaptation.
Remote Labor Index: Measuring AI Automation of Remote Work
Goals to track complete a whole real, economically valuable freelance/remote-work project end to end; achieve a level of automation comparable to what a human freelancer would deliver for that project
Must keep track of Agent must track progress toward completing an entire multi-part real-world project (not a single isolated action) to be credited with automating that unit of remote labor.
structure: hierarchical · goals arrive: given-up-front · subgoal credit: unclear · 21 citations
Horizonthe paper states none
the paper’s own words
we introduce the Remote Labor Index (RLI), a broadly multi-sector benchmark comprising real-world, economically valuable projects designed to evaluate end-to-end agent performance in practical settings.
AIs have made rapid progress on research-oriented benchmarks of knowledge and reasoning, but it remains unclear how these gains translate into economic value and automation.
SCUBA2025relevant
SCUBA: Salesforce Computer Use Benchmark
Goals to track navigate a specific enterprise software UI (Salesforce) to accomplish a CRM task; manipulate data records and automate workflows within the Salesforce platform; retrieve information and troubleshoot issues as part of a realistic CRM task; generalize across three personas (platform administrators, sales representatives, service agents) each with distinct task profiles
Must keep track of The agent must track its progress toward fine-grained milestones within each Salesforce sandbox task (UI navigation state, data changes made, workflow steps completed) to receive interpretable milestone-progress credit.
structure: sequential-chain · goals arrive: given-up-front · count: 300 task instances derived from real use · subgoal credit: yes · 7 citations · also Information seeking &amp; deep research
Horizonthe paper states none
the paper’s own words
SCUBA operates in Salesforce sandbox environments with support for parallel execution and fine-grained evaluation metrics to capture milestone progress. We benchmark a diverse set of agents under both zero-shot and demonstration-augmented settings.
SCUBA contains 300 task instances derived from real user interviews, spanning three primary personas, platform administrators, sales representatives, and service agents.
SOP-Bench2025relevant
SOP-Bench: Complex Industrial SOPs for Evaluating LLM Agents
Goals to track correctly execute each step of a complex, multi-step Standard Operating Procedure; orchestrate the correct tools/APIs at each SOP step; produce ground-truth outputs matching the SOP's authored specification across 12 business domains
Must keep track of The agent must track its position within a multi-step SOP, select the correct tool from a registry that may contain many irrelevant tools, and maintain consistency with the SOP's ground-truth interface and outputs across the whole procedure.
structure: sequential-chain · goals arrive: given-up-front · count: 2,000+ tasks across 12 business domains · subgoal credit: no · 20 citations · unit: other:not-stated
Horizonthe paper states none
the paper’s own words
LLM-based agents struggle to execute complex, multi-step Standard Operating Procedures (SOPs) that are fundamental to industrial automation. ... We introduce SOP-Bench, a benchmark of 2,000+ tasks from human expert-authored SOPs across 12 business domains ... yielding realistic tasks with executable interfaces and ground-truth outputs.
StockBench: Can LLM Agents Trade Stocks Profitably In Real-world Markets?
Goals to track make a sequential daily buy/sell/hold decision based on incoming market signals (prices, fundamentals, news); maximize cumulative return across the whole multi-month trading period; manage risk, minimizing maximum drawdown and maintaining a strong Sortino ratio
Must keep track of Agent must track its current portfolio position, cumulative return, and risk exposure (drawdown) as it processes a new daily market signal (prices, fundamentals, news) each step of a multi-month simulated trading run.
structure: sequential-chain · goals arrive: emitted-by-environment-over-time · subgoal credit: yes · 33 citations · unit: simulated-days
Horizonmulti-month (exact day/month count not given)
the paper’s own words
Agents receive daily market signals -- including prices, fundamentals, and news -- and make sequential buy, sell, or hold decisions.
STOCKBENCH, a contamination-free benchmark designed to evaluate LLM agents in realistic, multi-month stock trading environments. Agents receive daily market signals... and make sequential buy, sell, or hold decisions.
StaffPro2025relevant
StaffPro: an LLM Agent for Joint Staffing and Profiling
Goals to track assign and schedule tasks to workers (staffing), forming teams as needed; continuously estimate workers' latent skills, preferences, and other attributes from unstructured feedback (profiling); optimize staffing performance over time as profiling estimates improve via an ongoing human-agent feedback loop
Must keep track of StaffPro must track each worker's evolving latent-attribute profile (estimated from ongoing human feedback) and the current staffing/schedule state, updating both jointly over the 'life-long' profiling horizon.
structure: DAG-with-precedence · goals arrive: emitted-by-environment-over-time · subgoal credit: yes · 2 citations
Horizonthe paper states none
the paper’s own words
By analyzing human feedback, our agent continuously estimates the latent features of workers, realizing life-long worker profiling and ensuring optimal staffing performance over time.
UpBench2025relevant
UpBench: A Dynamically Evolving Real-World Labor-Market Agentic Benchmark Framework Built for Human-Centric AI
Goals to track complete a real, verified client transaction job sourced from the Upwork labor marketplace; satisfy each of a job's detailed, expert-decomposed, verifiable acceptance criteria; follow instructions faithfully enough to earn positive fine-grained, per-criterion human expert feedback
Must keep track of Agent must track which of a job's detailed acceptance criteria it has satisfied, informed by expert freelancer decomposition, to produce a submission gradeable on fine-grained, per-criterion feedback rather than a binary pass/fail.
structure: set-of-independent · goals arrive: given-up-front · count: one job per task, each decomposed into d · subgoal credit: yes · 4 citations
Horizonthe paper states none
the paper’s own words
UpBench employs a rubric-based evaluation framework, in which expert freelancers decompose each job into detailed, verifiable acceptance criteria and assess AI submissions with per-criterion feedback. This structure enables fine-grained analysis of model strengths, weaknesses, and instruction-following fidelity beyond binary pass/fail metrics.
Each task corresponds to a verified client transaction, anchoring evaluation in genuine work activity and financial outcomes.
Vending-Bench: A Benchmark for Long-Term Coherence of Autonomous Agents
Goals to track balance inventory levels of vending-machine stock; place restocking orders from suppliers; set item prices; pay recurring daily fees; sustain a profitable long-running vending-machine business without derailing
Must keep track of The agent must continuously track inventory counts, cash/profit, outstanding orders and their delivery schedules, and daily fees across a run spanning over 20M tokens, since forgetting an order or misreading a schedule causes derailment.
structure: open-ended · goals arrive: given-up-front · subgoal credit: no · 67 citations · unit: other:tokens-per-run
Horizon>20M
the paper’s own words
Agents must balance inventories, place orders, set prices, and handle daily fees - tasks that are each simple but collectively, over long horizons (>20M tokens per run) stress an LLM's capacity for sustained, coherent decision-making.
VitaBench2025in
VitaBench: Benchmarking LLM Agents with Versatile Interactive Tasks in Real-world Applications
Goals to track satisfy each of multiple real user requests combined into one cross-scenario task (food delivery, in-store consumption, online travel); reason across temporal and spatial dimensions while using a large tool set (66 tools); proactively clarify ambiguous instructions and track shifting user intent across a multi-turn conversation
Must keep track of Agent must track shifting user intent across a multi-turn conversation, temporal and spatial constraints, and the state of a large (66-tool) toolset spanning multiple life-serving domains simultaneously.
structure: set-of-independent · goals arrive: mixed:given-up-front-with-user-intent-shifting · count: 100 cross-scenario tasks (main results) · subgoal credit: unclear · 35 citations
Horizonthe paper states none
the paper’s own words
Each task is derived from multiple real user requests and requires agents to reason across temporal and spatial dimensions, utilize complex tool sets, proactively clarify ambiguous instructions, and track shifting user intent throughout multi-turn conversations.
yielding 100 cross-scenario tasks (main results) and 300 single-scenario tasks
Benchmarking LLM Agents for Wealth-Management Workflows
Goals to track complete each of 12 wealth-management task-pairs spanning retrieval, analysis, and synthesis/communication; satisfy explicit acceptance criteria under deterministic graders; operate correctly under both high- and low-autonomy task variants
Must keep track of Agent must track retrieved data, intermediate analysis results, and explicit acceptance criteria across the retrieval-analysis-synthesis pipeline, plus which autonomy variant (high vs. low) governs how independently it may act.
structure: sequential-chain · goals arrive: given-up-front · count: 12 task-pairs spanning retrieval, analys · subgoal credit: no · 0 citations
Horizonthe paper states none
the paper’s own words
We construct a benchmark of 12 task-pairs for wealth management assistants spanning retrieval, analysis, and synthesis/communication, with explicit acceptance criteria and deterministic graders.
This study introduces synthetic domain data, enriches colleague simulations, and prototypes an automatic task-generation pipeline.
Does It Tie Out? Towards Autonomous Legal Agents in Venture Capital
Goals to track verify that every security (shares, options, warrants) is supported by underlying legal documentation; verify that every issuance term (vesting schedules, acceleration triggers, transfer restrictions) is consistent across the dataroom; maintain strict evidence traceability while reconciling thousands of pages of legal documents; produce a deterministic, fully-reconciled capitalization table
Must keep track of The agent must track which securities and issuance terms have been verified against which supporting documents, maintaining strict evidence traceability across a growing dataroom (up to tens of thousands of pages) as it reconciles a scaling number of individual securities.
structure: set-of-independent · goals arrive: given-up-front · count: 184 to 1,292 individual securities requi · subgoal credit: yes · 0 citations · unit: agent-steps
Horizonworkload grows from approximately 2,700 atomic verification steps at Seed stage to nearly 8,000 steps at Series B stage
the paper’s own words
verifying that every security (for example, shares, options, warrants) and issuance term (for example, vesting schedules, acceleration triggers, transfer restrictions) is supported by large sets of underlying legal documentation. While LLMs continue to improve on legal benchmarks, specialized legal workflows, such as capitalization tie-out, remain out of reach even for strong agentic systems. The task requires multi-document reasoning, strict evidence traceability, and deterministic outputs.
Fig. 7 quantifies this burden by tracking the total number of atomic 'steps' executed by the counsel to complete the tie-out... We observe a near-tripling of workload, from approximately 2,700 steps at Seed to nearly 8,000 steps at Series B.

Personal assistants & long-term memory 65 artifacts · 27 state a horizon

AMemGym2026in
AMemGym: Interactive Memory Benchmarking for Assistants in Long-Horizon Conversations
Goals to track answer state-dependent questions correctly using accumulated conversational context; track an evolving simulated-user state across a long-horizon conversation; adapt personalization/memory strategies as latent user state evolves through role-play
Must keep track of Agent must maintain a consistent model of the simulated user's evolving latent state (from a predefined user profile plus a state-evolution trajectory), exposed only through free-form dialogue, to answer later state-dependent questions.
structure: sequential-chain · goals arrive: emitted-by-environment-over-time · subgoal credit: no · 16 citations
Horizonthe paper states none
the paper’s own words
AMemGym employs structured data sampling to predefine user profiles, state-dependent questions, and state evolution trajectories, enabling cost-effective generation of high-quality, evaluation-aligned interactions. LLM-simulated users expose latent states through role-play while maintaining structured state consistency.
Long-horizon interactions between users and LLM-based assistants necessitate effective memory management, yet current approaches face challenges in training and evaluation of memory.
ASTRA-bench2026relevant
ASTRA-bench: Evaluating Tool-Use Agent Reasoning and Action Planning with Personal User Context
Goals to track ground reasoning in time-evolving personal context (longitudinal life events) to resolve a user's current intent; orchestrate reliable multi-step tool-use plans conditioned on that evolving personal context; correctly handle user intents annotated by referential, functional, and informational complexity
Must keep track of Agent must track a protagonist's time-evolving personal context (longitudinal life events) and ground tool arguments/reasoning in that evolving context across a scenario.
structure: DAG-with-precedence · goals arrive: emitted-by-environment-over-time · count: 2,413 scenarios across four protagonists · subgoal credit: yes · 10 citations
Horizonthe paper states none
the paper’s own words
We present ASTRA-bench (Assistant Skills in Tool-use, Reasoning \&Action-planning), a benchmark that uniquely unifies time-evolving personal context with an interactive toolbox and complex user intents.
Our event-driven pipeline generates 2,413 scenarios across four protagonists, grounded in longitudinal life events and annotated by referential, functional, and informational complexity.
AgentIF-OneDay2026relevant
AgentIF-OneDay: A Task-level Instruction-Following Benchmark for General AI Agents in Daily Scenarios
Goals to track adhere to an explicit, complex, user-given workflow (Open Workflow Execution); infer implicit/latent instructions from attached files (Latent Instruction); modify or expand upon work already produced earlier in the task (Iterative Refinement); deliver a correct, tangible file-based result, not just a dialogue answer
Must keep track of The agent must track the explicit and inferred requirements of a task across multiple attachments and any prior output it has already produced, so that later refinement or workflow steps remain consistent with earlier ones.
structure: sequential-chain · goals arrive: mixed:given-up-front-plus-implied-by-constrain · count: 104 tasks covering 767 scoring points, a · subgoal credit: yes · 7 citations · also Business, office &amp; enterprise work
Horizonthe paper states none
the paper’s own words
AgentIF-OneDay comprises 104 tasks covering 767 scoring points... We employ instance-level rubrics and a refined evaluation pipeline that aligns LLM-based verification with human judgment, achieving an 80.1% agreement rate using Gemini-3-Pro.
OneDayAgent: Towards a Long-Horizon Harness for Autonomous Agents
Goals to track decompose an open-ended everyday request spanning work, study, and life into bounded subtasks; preserve goals and constraints across many steps while navigating heterogeneous tools and attachments, avoiding goal drift and state loss; verify and repair the final deliverable against context-overflow and other failure modes
Must keep track of Harness must maintain execution memory of goals/constraints and subtask progress under context pressure across an open-ended, long-horizon, cross-environment, multimodal request, and verify/repair the final deliverable at the end.
structure: hierarchical · goals arrive: self-generated-by-agent · count: 104 tasks in the AgentIF-OneDay evaluati · subgoal credit: yes · 0 citations
Horizonthe paper states none
the paper’s own words
OneDayAgent turns an open-ended request into a managed execution process that decomposes tasks into bounded subtasks, maintains execution memory under context pressure, and verifies and repairs the final deliverable.
These tasks are long-horizon, cross-environment, and multimodal, forcing the agent to preserve goals and constraints across many steps while navigating heterogeneous tools and attachments.
Your Agents Are Aging Too: Agent Lifespan Engineering for Deployed Systems
Goals to track maintain factual/behavioral reliability as the agent's effective state changes via compression, retrieval, revision, and maintenance over its deployment lifespan; correctly write, retrieve, and utilize memory across the memory pipeline's stages; correctly repair a diagnosed failure at the specific pipeline stage (write/retrieval/utilization) responsible for it
Must keep track of The agent (and the benchmark's diagnostic layer) must track the full lifespan history of memory writes, retrievals, and revisions across up to 200 sessions to determine which pipeline stage a given failure traces back to.
structure: DAG-with-precedence · goals arrive: implied-by-constraints · count: 4 aging mechanisms (compression, interfe · subgoal credit: yes · 6 citations · unit: sessions
Horizon~400 runs spanning 8-200 sessions
the paper’s own words
AgingBench organizes agent aging into four mechanisms: compression aging, interference aging, revision aging, and maintenance aging. To diagnose these failures, AgingBench uses temporal dependency graphs and paired counterfactual probes that produce diagnostic profiles for the write, retrieval, and utilization stages of the memory pipeline.
over ~400 runs spanning 8 - 200 sessions show that agent aging is not one-dimensional
AlpsBench2026relevant
AlpsBench: An LLM Personalization Benchmark for Real-Dialogue Memorization and Preference Alignment
Goals to track extract explicit and implicit personalized user traits from long-term interaction sequences; correctly update stored personalized memory as new information arrives; retrieve relevant personalized memory under large distractor pools; utilize retrieved memory to produce preference-aligned, emotionally resonant responses
Must keep track of Agent must extract, update, retrieve, and utilize structured user memories (explicit and implicit personalization signals) consistently across long-term interaction sequences curated from real dialogues.
structure: sequential-chain · goals arrive: given-up-front · count: 2,500 long-term interaction sequences; f · subgoal credit: no · 6 citations
Horizonthe paper states none
the paper’s own words
We define four pivotal tasks---personalized information extraction, updating, retrieval, and utilization ---and establish protocols to evaluate the entire lifecycle of memory management.
AlpsBench comprises 2,500 long-term interaction sequences curated from WildChat, paired with human-verified structured memories
PersonalAlign: Hierarchical Implicit Intent Alignment for Personalized GUI Agent with Long-Term User-Centric Records
Goals to track resolve omitted preferences in vague GUI instructions using long-term user records; anticipate latent routines from user state for proactive assistance; correctly execute and proactively suggest actions grounded in hundreds of distinct user-specific preferences and routines
Must keep track of Agent must maintain a continuously updating personal memory, hierarchically organizing user preferences and routines inferred from long-term records, to resolve vague instructions and proactively suggest actions.
structure: hierarchical · goals arrive: implied-by-constraints · count: 775 annotated user-specific preferences · subgoal credit: no · 16 citations · also Web &amp; GUI agents
Horizonthe paper states none
the paper’s own words
We annotated 775 user-specific preferences and 215 routines from 20k long-term records across different users for evaluation... HIM-Agent significantly improves both execution and proactive performance by 15.7% and 7.3%.
CalBench2026in
CalBench: Evaluating Coordination-Privacy Trade-offs in Multi-Agent LLMs
Goals to track schedule a stream of M incoming meetings while managing one's own private calendar; minimize disruption cost to the agent's own calendar; coordinate with other agents' private calendars via language-mediated negotiation, without directly inspecting their calendars; preserve privacy (avoid over-revealing calendar information) while still achieving fair burden allocation across agents
Must keep track of Each agent must track its own private calendar state, disruption costs incurred so far, what it has revealed or withheld to other agents, and burden/fairness considerations across the stream of scheduling requests.
structure: sequential-chain · goals arrive: emitted-by-environment-over-time · count: N agents scheduling a stream of M incomi · subgoal credit: yes · 4 citations · also Multi-agent organisations &amp; societies
Horizonthe paper states none
the paper’s own words
In each task, $N$ agents manage separate private calendars and schedule a stream of $M$ incoming meetings while minimizing disruption costs. Because no agent can inspect another agent's calendar, success requires language-mediated coordination rather than centralized planning.
PEARL: Self-Evolving Assistant for Time Management with Reinforcement Learning
Goals to track resolve calendar conflicts round-by-round across a full calendar year; infer and progressively adapt to evolving user preferences (attendee priorities, topic importance, time/location preferences); decide which meetings to attend, reschedule, or decline per conflict
Must keep track of Agent must maintain an external preference memory that stores and updates inferred strategies (attendee priorities, topic importance, time/location preferences) and use round-wise decisions across the calendar year to track scheduling state.
structure: sequential-chain · goals arrive: emitted-by-environment-over-time · subgoal credit: no · 3 citations · unit: simulated-years
Horizonone calendar year (presented round-by-round)
the paper’s own words
In CalConflictBench, conflicts are presented to agents round-by-round over a calendar year, requiring them to infer and adapt to user preferences progressively... (ii) optimizes the agent with round-wise rewards that directly supervise decision correctness, ranking quality, and memory usage across rounds.
Claw-Anything: Benchmarking Always-On Personal Assistants with Broader Access to User's Digital World
Goals to track reason over long-horizon activity histories accumulated across months of simulated user activity; coordinate interdependent backend services and integrated GUI/CLI interaction across multiple devices; remain robust to irrelevant events and conflicting noise signals; proactively anticipate user needs and deliver timely recommendations
Must keep track of Agent must reason over rich, long-horizon activity histories and interdependent backend-service state accumulated across simulated months, remaining robust to irrelevant/conflicting noise while proactively anticipating user needs.
structure: open-ended · goals arrive: emitted-by-environment-over-time · subgoal credit: no · 0 citations · unit: other:simulated-months
Horizonmonths
the paper’s own words
To instantiate this setting, we simulate months of user activity through multi-round event injection, producing complex world states and realistic noise, including irrelevant events and conflicting signals. Agents must reason over rich contextual environments while remaining robust to such noise.
CloneMem2026in
CloneMem: Benchmarking Long-Term Memory for AI Clones
Goals to track track an individual's evolving experiences, emotions, opinions, and personal states across long non-conversational digital traces (diaries, social posts, emails); answer/act using the current (not superseded) personal state reflecting one to three years of life history
Must keep track of Agent must track an evolving personal state (experiences, emotions, opinions) over one to three years of non-conversational digital traces and correctly distinguish current from superseded information.
structure: sequential-chain · goals arrive: emitted-by-environment-over-time · subgoal credit: yes · 10 citations · unit: simulated-years
Horizon1 to 3
the paper’s own words
CloneMem adopts a hierarchical data construction framework to ensure longitudinal coherence and defines tasks that assess an agent's ability to track evolving personal states.
We introduce CloneMem, a benchmark for evaluating longterm memory in AI Clone scenarios grounded in non-conversational digital traces, including diaries, social media posts, and emails, spanning one to three years.
Controlled Memory Interference in Continual LLM Agents
Goals to track update memory upon new experience while managing reinforcement, revision, or interference with existing memory states; distinguish valid memory updates from interference-inducing memories; maintain multiple simultaneously relevant memories differing in state, temporal validity, or authority; preserve continuity across sessions to personalize behavior via accumulated experience
Must keep track of The agent's memory system must track multiple simultaneously relevant memory states that differ in validity, temporal recency, or authority, and correctly determine which prior memory a new experience should reinforce, revise, or be blocked by.
structure: set-of-independent · goals arrive: emitted-by-environment-over-time · subgoal credit: no · 1 citations
Horizonthe paper states none
the paper’s own words
We introduce Controlled Memory Interference (CMI), a controlled diagnostic and data-generation framework for studying how agent memory evolves under different memory relationships. Across controlled memory evolution, benign accumulation has limited effects, whereas relationship-specific interference sharply suppresses update plasticity with little stability gain, either by blocking target-memory exposure or by disrupting its downstream use.
Long-term memory enables AI agents to maintain continuity across sessions, personalize behavior, and evolve through accumulated experience.
ES-MemEval: Benchmarking Conversational Agents on Personalized Long-Term Emotional Support
Goals to track extract and retain implicit, fragmented user disclosures across sessions; perform temporal reasoning over how the user's state has evolved; detect conflicts between what the user said earlier and later; abstain when information is insufficient rather than hallucinate; build and update a model of the user across QA, summarization, and dialogue-generation tasks
Must keep track of Agent must track fragmented and implicit user disclosures, detect when the user's state or facts have changed, and maintain an updated user model across multiple sessions of an emotional-support dialogue.
structure: set-of-independent · goals arrive: emitted-by-environment-over-time · count: 5 (core memory capabilities: information · subgoal credit: unclear · 11 citations
Horizonthe paper states none
the paper’s own words
we introduce ES-MemEval, a comprehensive benchmark that systematically evaluates five core memory capabilities—information extraction, temporal reasoning, conflict detection, abstention, and user modeling—in long-term emotional support scenarios
we also propose EvoEmo, the first multi-session dataset for personalized long-term emotional support scenarios, capturing fragmented, implicit user disclosures and evolving user states
EduClaw-Bench: A Long-Horizon Benchmark for Pedagogical LLM Agents with Simulated Learners
Goals to track improve a simulated learner's knowledge-concept mastery over a sustained tutoring relationship; personalize responsiveness and helpfulness to the learner's evolving needs; apply sound curriculum-design principles (Gagne and Rosenshine axes) across the relationship; sustain good tutoring performance across the full 30-day horizon, not just in an initial session
Must keep track of The tutor agent must track the simulated learner's evolving knowledge-concept mastery (grounded in a KT model) across the 30-day relationship, adapting its teaching to what the learner has and hasn't yet learned, and sustaining quality across the full horizon rather than only an initial session.
structure: sequential-chain · goals arrive: mixed:given-up-front-relationship-structure-pl · count: 55 scenarios; continuous 30-day relation · subgoal credit: yes · 0 citations · unit: simulated-days
Horizoncontinuous 30-day tutoring relationship, across 55 scenarios
the paper’s own words
We introduce EduClaw-Bench, a benchmark that places an agent tutor in a continuous 30-day relationship with a simulated learner grounded in knowledge tracing (KT), whose knowledge-concept mastery, from a KT model trained on real-student data, drives its answers and is probed for learning gain across 55 scenarios.
EgoMemReason2026relevant
EgoMemReason: A Memory-Driven Reasoning Benchmark for Long-Horizon Egocentric Video Understanding
Goals to track track how object states evolve and change across days (entity memory); recall and correctly order activities separated by hours or days (event memory); abstract recurring patterns from sparse, repeated observations across a whole week (behavior memory); answer each of 500 questions requiring integration of evidence across multiple days of egocentric video
Must keep track of The system must accumulate information over an entire week of continuous egocentric video, recall prior states, track the temporal order of events separated by hours or days, and abstract recurring behavioral patterns from sparse repeated observations, backtracking an average of 25.9 hours of memory per question.
structure: set-of-independent · goals arrive: given-up-front · count: 500 questions across three memory types · subgoal credit: no · 3 citations · unit: wall-clock-hours
Horizon25.9 hours of memory backtracking per question (average); week-long underlying video
the paper’s own words
EgoMemReason comprises 500 questions across three memory types and six core challenges, with an average of 5.1 video segments of evidence per question and 25.9 hours of memory backtracking.
EverMemBench2026relevant
Evaluating Long-Horizon Memory for Multi-Party Collaborative Dialogues
Goals to track perform fine-grained recall across dense, cross-topic multi-party conversations; maintain memory awareness of implicitly relevant information beyond similarity retrieval; understand and correctly attribute user profiles across multiple participants and roles; resolve multi-hop reasoning under multi-party attribution and temporally evolving decisions
Must keep track of The agent must track role-conditioned personas, temporally evolving decisions, and cross-topic interleaved information across multi-party, multi-group conversations exceeding one million tokens, in order to answer QA pairs spanning recall, awareness, and profile understanding.
structure: set-of-independent · goals arrive: given-up-front · count: 2,400 QA pairs across three evaluation d · subgoal credit: no · 11 citations · also Multi-agent organisations &amp; societies · unit: other:tokens
Horizon>1,000,000
the paper’s own words
EverMemBench evaluates memory systems using 2,400 QA pairs across three dimensions essential for real applications: fine-grained recall, memory awareness, and user profile understanding. Our evaluation reveals fundamental limitations of current systems: multi-hop reasoning collapses under multi-party attribution even with oracle evidence (26% accuracy), temporal reasoning fails without explicit version semantics beyond timestamps
built from multi-party, multi-group conversations spanning over one million tokens with dense cross-topic interleaving, temporally evolving decisions, and role-conditioned personas
EvoMemBench: Benchmarking Agent Memory from a Self-Evolving Perspective
Goals to track retain and retrieve knowledge-oriented information across episode boundaries; retain and reuse execution-oriented (procedural) experience across episode boundaries; satisfy in-episode memory demands; satisfy cross-episode memory demands
Must keep track of The system must store, update, and retrieve both knowledge and procedural experience across episode boundaries, and determine which stored memories are relevant to reuse for the current task.
structure: set-of-independent · goals arrive: given-up-front · subgoal credit: no · 7 citations
Horizonthe paper states none
the paper’s own words
we introduce EvoMemBench, a unified benchmark organized along two axes: memory scope (in-episode vs. cross-episode) and memory content (knowledge-oriented vs. execution-oriented). We compare 15 representative memory methods with strong long-context baselines under a standardized protocol.
EvolMem2026in
EvolMem: A Cognitive-Driven Benchmark for Multi-Session Dialogue Memory
Goals to track correctly recall declarative memory content across multiple dialogue sessions; correctly exhibit non-declarative memory capabilities across multiple dialogue sessions; succeed across multiple fine-grained memory-ability dimensions grounded in cognitive psychology, not just one aggregate score
Must keep track of Agent/memory system must retain and correctly apply both declarative (fact-like) and non-declarative (procedural/implicit) memory content across multiple sessions of scalable, controllable complexity.
structure: set-of-independent · goals arrive: emitted-by-environment-over-time · subgoal credit: unclear · 4 citations
Horizonthe paper states none
the paper’s own words
EvolMem is grounded in cognitive psychology and encompasses both declarative and non-declarative memory, further decomposed into multiple fine-grained abilities.
This framework enables scalable generation of multi-session conversations with controllable complexity, accompanied by sample-specific evaluation guidelines.
FileGram: Grounding Agent Personalization in File-System Behavioral Traces
Goals to track reconstruct an evolving user profile from dense file-system behavioral traces; disentangle overlapping/interleaved behavioral traces belonging to different activities; detect persona drift as the user's behavior changes over time; correctly ground multimodal (procedural, semantic, episodic) evidence into the user profile
Must keep track of The memory system must track atomic file-system actions and content deltas over time, encode them into procedural/semantic/episodic channels, disentangle overlapping traces, and detect when the user's persona has drifted from its established profile.
structure: set-of-independent · goals arrive: emitted-by-environment-over-time · count: FileGramEngine produces 640 controlled t · subgoal credit: no · 4 citations · unit: other:not-stated
Horizonthe paper states none
the paper’s own words
we propose FileGram, a comprehensive framework that grounds agent memory and personalization in file-system behavioral traces, comprising three core components: (1) FileGramEngine, a scalable persona-driven data engine ...; (2) FileGramBench, a diagnostic benchmark grounded in file-system behavioral traces for evaluating memory systems on profile reconstruction, trace disentanglement, persona drift detection, and multimodal grounding; and (3) FileGramOS, a bottom-up memory architecture
FileGramEngine produces 640 controlled trajectories with ground-truth labels across 6 behavioral dimensions and 20 user profiles
Gaia2: Benchmarking LLM Agents on Dynamic and Asynchronous Environments
Goals to track complete a scenario-level task while the environment evolves independently of the agent's actions; operate under explicit temporal constraints (time-sensitive tasks); adapt to noisy and dynamic events injected during the scenario; resolve ambiguity in requests; collaborate with other agents present in the scenario
Must keep track of The agent must track evolving environment state (independent of its own actions), remaining time budgets for time-sensitive sub-goals, ambiguity-resolution status, and coordination with other agents, verified at the action level by write-action verifiers.
structure: DAG-with-precedence · goals arrive: mixed:given-up-front-plus-emitted-by-environme · count: 1,120 human-annotated scenarios (per esc · subgoal credit: no · 21 citations · also Business, office &amp; enterprise work · unit: other:not-stated
Horizonthe paper states none
the paper’s own words
Gaia2 introduces scenarios where environments evolve independently of agent actions, requiring agents to operate under temporal constraints, adapt to noisy and dynamic events, resolve ambiguity, and collaborate with other agents. Each scenario is paired with a write-action verifier, enabling fine-grained, action-level evaluation and making Gaia2 directly usable for reinforcement learning from verifiable rewards.
Gaia2 consists of 1,120 human-annotated scenarios set in a smartphone-like environment with realistic apps (email, messaging, calendar, contacts, etc.)
GateMem2026in
GateMem: Benchmarking Memory Governance in Multi-Principal Shared-Memory Agents
Goals to track serve legitimate long-horizon requests that require state updates to shared memory; enforce access control across contextual authorization boundaries for different principals; perform agent-facing active forgetting after explicit deletion requests; avoid leaking unauthorized or deleted information to any principal
Must keep track of The agent must track per-principal roles, scopes and relationships, incremental memory updates, hidden checkpoints, and outstanding deletion requests across long-form multi-party episodes averaging roughly 200-240 turns depending on domain.
structure: set-of-independent · goals arrive: mixed:given-up-front-plus-injected-by-user-mid · subgoal credit: no · 3 citations · unit: turns
Horizon~204.5 (medical) / 241.2 (office) / 224.9 (education) / 224.0 (household) turns per episode
the paper’s own words
GateMem jointly evaluates utility for legitimate long-horizon requests with state updates, access control across contextual authorization boundaries, and agent-facing active forgetting after explicit deletion requests. It spans medical, office, education, and household domains, with long-form multi-party episodes, incremental memory injection, hidden checkpoints, structured judging, and leak-target annotations.
Medical: 204.5 ... Office: 241.2 ... Education: 224.9 ... Household: 224.0 (turns per episode, Table 2, dataset overview)
Multi-Agent Home Energy Management Assistant
Goals to track sustain multi-turn conversational collaboration with preserved context across a home-energy-management session; perform energy consumption analysis and cost optimization (Analysis agent); answer educational queries and provide rebate information (Knowledge agent); control and schedule smart devices (Control agent); correctly route each user query to the right specialized agent via a self-consistency classifier
Must keep track of The system must preserve conversational context across multiple turns of sustained human-AI collaboration, track which of the three specialized agents/tools have been invoked, and maintain consistency across energy analysis, educational, and device-control interactions within one session.
structure: other:three-specialized-agents-coordinated-by- · goals arrive: given-up-front · count: three specialized agents (Analysis, Know · subgoal credit: no · 7 citations · also Embodied &amp; robotics
Horizonthe paper states none
the paper’s own words
HEMA also includes a comprehensive evaluation framework using an LLM-as-simulated-user methodology with 23 objective metrics across task performance, factual accuracy, interaction quality, and system efficiency, allowing systematic testing across diverse scenarios and user personas without requiring extensive human subject testing.
multi-turn conversational interactions with preserved context
Evaluating Memory Capability in Continuous Lifelog Scenario
Goals to track recall and reason correctly over continuously accumulating lifelog conversation history; answer queries using only information available up to the query time (no temporal leakage) under an Online Evaluation protocol
Must keep track of System must ingest and retain a continuously growing stream of ambient conversation (lifelog audio) and answer later queries using only causally-prior information, without being able to look ahead.
structure: sequential-chain · goals arrive: emitted-by-environment-over-time · subgoal credit: unclear · 2 citations
Horizonthe paper states none
the paper’s own words
we propose an \textbf{Online Evaluation} protocol that strictly adheres to temporal causality, ensuring systems are evaluated in a realistic streaming fashion
wearable devices can continuously lifelog ambient conversations, creating substantial opportunities for memory systems
LifeSide2026in
LifeSide: Benchmarking Agents as Lifelong Digital Companions
Goals to track integrate cross-session memory cues about a persistent user persona; continually update the agent's understanding of the user over time; adapt to the user's shifting privacy boundaries; sustain accurate emotional companionship across sessions
Must keep track of The agent must track a persistent user world (layered profile, event trajectory) across an average of 56.79 sessions and 851.85 user turns per persona, covering memory tracking, user understanding, privacy control, and emotional companionship.
structure: hierarchical · goals arrive: emitted-by-environment-over-time · count: 2,000 personas and 111K tasks; averaging · subgoal credit: no · 0 citations · unit: sessions
Horizonavg 56.79 sessions / 851.85 user turns / 29.61K dialogue tokens per persona
the paper’s own words
\benchmark uses multi-agent simulation to project environmental dynamics into dialogue, preserving the critical gap between latent thoughts and observable expressions. Evaluating 2,000 personas and 111K tasks across memory tracking, user understanding, privacy control, and emotional companionship, our experiment results reveal a stark reality: even models that saturate current memory benchmarks fail to sustain accurate user understanding and true companionship over long horizons.
This longitudinal design yields substantial interaction histories, averaging 56.79 sessions, 851.85 user turns, and 29.61K visible dialogue tokens per persona.
LifeSim: Long-Horizon User Life Simulator for Personalized Assistant Evaluation
Goals to track complete the user's explicit intentions correctly; recognize and satisfy the user's implicit intentions; recover an accurate model of the user's evolving profile/preferences over the course of long-horizon assistance; produce high-quality responses across 8 life domains and 1,200 diverse scenarios
Must keep track of Agent must recover and continuously update the user's profile (explicit and implicit intentions, preferences) as it evolves across a long-horizon, multi-scenario life trajectory, using a multi-turn interactive assessment method.
structure: hierarchical · goals arrive: mixed:given-up-front-life-domain-scenarios-wit · count: 1,200 diverse scenarios across 8 life do · subgoal credit: unclear · 10 citations
Horizonthe paper states none
the paper’s own words
LifeSim-Eval covers 8 life domains and 1,200 diverse scenarios, and adopts a multi-turn interactive method to assess models'abilities to complete explicit and implicit intentions, recover user profiles, and produce high-quality responses.
LiveClawBench2026relevant
LiveClawBench: Benchmarking LLM Agents on Complex, Real-World Assistant Tasks
Goals to track resolve tasks that span cross-service dependencies across mocked applications; operate correctly despite contaminated/inconsistent prior state; correctly infer implicit user intent not explicitly stated in the request; adapt to runtime changes within stateful mock services during task execution
Must keep track of The agent must track session state, artifacts, and prior side-effects across 22 stateful mocked services, resolving cross-service dependencies and implicit intent while adapting to runtime changes during the task.
structure: DAG-with-precedence · goals arrive: implied-by-constraints · count: 134 executable cases across 10 domains w · subgoal credit: no · 5 citations · unit: other:not-stated
Horizonthe paper states none
the paper’s own words
LiveClawBench combines a Triple-Axis Complexity Framework for difficulty-driven task construction with reproducible full-stack mock applications that preserve stateful execution semantics. With 134 executable cases across 10 domains with 22 mocked services, LiveClawBench supports controlled, extensible, and factor-level diagnostic evaluation of realistic agentic tasks.
The scatter shows that higher reward is not simply a consequence of taking more interaction steps.
MEMPROBE2026in
MEMPROBE: Probing Long-Term Agent Memory via Hidden User-State Recovery
Goals to track assist simulated users across a trajectory of leak-controlled tasks while accumulating memory; recover/reconstruct a hidden, taxonomy-anchored user-state bank (31 dimensions) from the agent's own resulting memory; balance successful task assistance against auditable, faithful memory recovery
Must keep track of Agent must accumulate and retain a faithful memory of 31 hidden user-state dimensions across a trajectory of leak-controlled assistance tasks, since the memory is later audited by reconstructing the user-state bank from it under full-store and top-k access.
structure: sequential-chain · goals arrive: given-up-front · count: 50 simulated users with 31 hidden dimens · subgoal credit: no · 2 citations
Horizonthe paper states none
the paper’s own words
We instantiate this view in MEMPROBE, a benchmark in which a memory-equipped agent assists simulated users, each carrying a hidden, taxonomy-anchored user-state bank, across a trajectory of leak-controlled tasks, after which that bank is reconstructed from the agent's resulting memory under both full-store and top-k access... MEMPROBE spans 50 simulated users with 31 hidden dimensions each (1,550 recovery targets)
MERIT2026relevant
When Does Memory Help? A Cost-Aware Evaluation of Long-Term Memory in Tool-Using LLM Agents
Goals to track correctly recall and use an earlier-episode fact when executing a later, dependent tool-use task; correctly recall and use an UPDATED fact (superseding a stale one) rather than acting on outdated information; operate under an explicit cost budget (token/dollar metering) while doing so; avoid corrupted/adversarially degraded memory leading to incorrect actions
Must keep track of The agent's memory system must retain facts (including corrections to previously stored facts) across an arc of linked episodes and correctly retrieve and act on the current, updated version of a fact rather than a stale cached one, all while the harness meters the token/dollar cost of every memory operation.
structure: DAG-with-precedence · goals arrive: mixed:given-up-front-task-structure-plus-emitt · count: 23,440 scored episodes across a 3-model · subgoal credit: yes · 0 citations · unit: episodes
Horizonepisodes grouped into arcs of 4-6 linked episodes sharing entities; 10 arcs x 5 episodes per (domain x difficulty x condition), 23,440 scored episodes total
the paper’s own words
MERIT provides episodic tool-use tasks in three domains whose dependence on earlier-episode facts is verified by an automated leak check; a difficulty ladder ending in updated-fact recall; controlled memory corruption; and full token and dollar metering of every memory operation... memory lifts dependent-task success from a leak-verified floor of 0.00 to 0.55-1.00. On updated facts, embedding retrieval collapses unpredictably (0.30-0.95 across models; max seed gap 0.45), and agents act on a correctly retrieved value only 55% of the time.
Episodes are grouped into arcs of 4-6 episodes sharing entities... 10 arcs x 5 episodes per (domain x difficulty x condition)
MemConflict: Evaluating Long-Term Memory Systems Under Memory Conflicts
Goals to track retrieve and rank the temporally valid, factually correct, and contextually applicable memory candidate when multiple conflicting alternatives exist; correctly answer queries under dynamic, static, and conditional conflict types despite distractors and long conflict distances
Must keep track of The agent's memory system must retrieve and rank memory candidates while tracking temporal validity, factual correctness, and contextual applicability across an average of 52.33 sessions (2,349.17 turns, ~203,910 tokens) per instance, correctly resolving conflicts placed 5 to 49 sessions apart.
structure: DAG-with-precedence · goals arrive: emitted-by-environment-over-time · count: average of 52.33 sessions and 2,349.17 d · subgoal credit: yes · 5 citations · unit: sessions
Horizonaverage of 52.33 sessions and 2,349.17 dialogue turns (about 203,910.83 tokens of context) per benchmark instance; conflict distances span 5-25 sessions (dynamic), 10-45
the paper’s own words
MemConflict formalizes dynamic, static, and conditional conflicts over temporal validity, factual correctness, and contextual applicability. It simulates controlled long-horizon histories from structured user profiles, introduces cross-session conflicts, and injects semantically similar distractors to create competition among memory candidates. The resulting multi-session dialogue benchmark supports black-box evaluation of final answers and white-box analysis of supporting-memory retrieval and ranking.
Average Session Number: 52.33; Average Dialogue Turns: 2,349.17; Average Context Length (tokens): 203,910.83
MemGround2026in
MemGround: Long-Term Memory Evaluation Kit for Large Language Models in Gamified Scenarios
Goals to track recall surface-level game state facts (Surface State Memory); associate events across time (Temporal Associative Memory); perform reasoning that depends on accumulated memory (Reasoning-Based Memory); unlock/discover memory fragments in the correct order across a gamified scenario
Must keep track of The agent must maintain and update a three-tier memory store (surface state, temporal associations, reasoning-derived facts) across continuous gamified interactions, tracked via Memory Fragments Unlocked and Memory Fragments with Correct Order metrics, with runs capped at 600-1000 interaction steps depending on task type.
structure: hierarchical · goals arrive: emitted-by-environment-over-time · subgoal credit: yes · 1 citations · unit: other:interaction-steps
Horizonmax 600-1000 interaction steps (task-dependent); early stop after 200 consecutive steps with no new discovery
the paper’s own words
MemGround introduces a three-tier hierarchical framework that evaluates Surface State Memory, Temporal Associative Memory, and Reasoning-Based Memory through specialized interactive tasks. ... a multi-dimensional metric suite comprising Question-Answer Score (QA Overall), Memory Fragments Unlocked (MFU), Memory Fragments with Correct Order (MFCO), and Exploration Trajectory Diagrams (ETD).
we set the maximum number of interaction steps to 600. ... we set the maximum number of interaction steps to 1000. ... if the model does not discover any new files within 200 consecutive steps
MemOps2026in
MemOps: Benchmarking Lifecycle Memory Operations in Long-Horizon Conversations
Goals to track correctly execute each lifecycle memory operation (remember, forget, update, reflect, and their compositions) at the right point in a conversation; maintain a consistent, ordered memory-state trajectory (not just the correct final answer) across a long conversation
Must keep track of The agent's memory system must track the trigger, target, scope, and state transition of each lifecycle operation (remember/forget/update/reflect) across up to 9,672 dialogue turns, maintaining a correct ordered memory-state trajectory rather than only a final answer.
structure: sequential-chain · goals arrive: emitted-by-environment-over-time · count: 100 unique topics, 403 evidence conversa · subgoal credit: yes · 3 citations · unit: turns
Horizon9,672 dialogue turns across 403 evidence conversations (100 unique topics), decomposed into 1,209 evidence-conversation segments
the paper’s own words
We introduce MemOps, a benchmark that reformulates conversational memory as a sequence of lifecycle operations and represents each memory event with a structured trace specifying its trigger, target, scope, state transition, and supporting evidence. A controllable generation pipeline embeds these operations into long, task-oriented conversations and produces gold operation traces together with six categories of operation-level probes.
The benchmark spans 100 unique topics and comprises 403 evidence conversations, which decompose into 1,209 evidence-conversation segments and 9,672 dialogue turns.
MemoryArena: Benchmarking Agent Memory in Interdependent Multi-Session Agentic Tasks
Goals to track distill experience from earlier actions and feedback into memory during multi-session interaction; use previously distilled memory to guide later actions and solve subsequent, explicitly interdependent subtasks; solve overall tasks spanning web navigation, preference-constrained planning, progressive information search, and sequential formal reasoning
Must keep track of The agent must track which experiences it has distilled into memory across earlier sessions and correctly recall/apply the relevant portions of that memory to solve later, interdependent subtasks across multiple domains (web navigation, planning, search, formal reasoning).
structure: DAG-with-precedence · goals arrive: given-up-front · subgoal credit: no · 62 citations
Horizonthe paper states none
the paper’s own words
MemoryArena supports evaluation across web navigation, preference-constrained planning, progressive information search, and sequential formal reasoning, and reveals that agents with near-saturated performance on existing long-context memory benchmarks like LoCoMo perform poorly in our agentic setting, exposing a gap in current evaluations for agents with memory.
we introduce MemoryArena, a unified evaluation gym for benchmarking agent memory in multi-session Memory-Agent-Environment loops
Momento2026in
Momento: Evaluating Persistent Memory and Reasoning with Multi-Session Agentic Conversations
Goals to track take consequential, tool-mediated actions on behalf of a user within a multi-session service environment; resolve temporal dependencies between what happened in earlier sessions and what is being requested now; keep pace with evolving user goals across sessions rather than treating prior session history as static ground truth
Must keep track of Agent must track prior session history, recognize which parts of it may now be stale, and re-validate temporal dependencies and evolving user goals before taking consequential tool-mediated actions in the current session.
structure: sequential-chain · goals arrive: mixed:given-up-front-service-tasks-with-user-g · subgoal credit: unclear · 1 citations
Horizonthe paper states none
the paper’s own words
We introduce Momento, a benchmark for persistent agentic task completion in multi-session service environments, requiring agents to take consequential, tool-mediated actions while resolving temporal dependencies and evolving user goals across sessions.
existing benchmarks evaluate agents within a single session, ignoring past actions, stated preferences, and prior decisions that agents must integrate to fulfill personalized user goals
MultiSessionCollab2026relevant
MultiSessionCollab: Learning User Preferences with Memory to Improve Long-Term Collaboration
Goals to track solve each of 20 sequential collaboration problems (one per session) for a given user; learn and apply that user's preferences across sessions to improve collaboration quality and reduce user effort over time
Must keep track of The agent must track learned user preferences and reflections accumulated from prior sessions, applying them to reduce the number of conversational turns and user effort needed in each subsequent session across a 20-session sequence.
structure: sequential-chain · goals arrive: given-up-front · count: 20 randomly sampled problems per user (o · subgoal credit: yes · 8 citations · also Multi-agent organisations &amp; societies · unit: sessions
Horizon20 sessions per user (one problem per session, up to 10 conversational turns per session), totaling 10,000 collaborative sessions per agent across the benchmark; turns ne
the paper’s own words
We introduce MultiSessionCollab, a benchmark that evaluates how well agents can learn user preferences and leverage them to improve collaboration quality throughout multiple sessions... Extensive experiments show that equipping agents with our memory improves collaboration over time, yielding higher task success rates, more efficient interactions, and reduced user effort.
Each user collaborates with the agent to solve 20 randomly sampled problems, with one problem per session and a maximum of 10 conversational turns per session... this totals 10,000 collaborative sessions per agent.
PAST-Bench: Benchmarking the Foundations of Recursive Self-Improvement in Personal Agents
Goals to track reuse a retained skill/procedure across sessions on a later fresh-session task; retrieve a previously stored preference/fact and apply it correctly in a new session; gather information in one session that is needed to complete a task in a later session; update outdated retained state (e.g. a stale fact) rather than acting on it unchanged
Must keep track of The agent must save, retrieve, and update experience (preferences, task histories, tool routines, learned skills) across an ordered sequence of separate, fresh sessions, and the benchmark explicitly checks whether later-task gains actually follow this intended save/retrieve/update pathway rather than occurring by other means.
structure: sequential-chain · goals arrive: given-up-front · count: 26 scenarios and 204 episodes, across me · subgoal credit: yes · 2 citations · also Information seeking &amp; deep research · unit: sessions
Horizon26 scenarios, 204 episodes total; ordered sequences of fresh-session tasks per scenario
the paper’s own words
Each agent runs through ordered sequences of fresh-session tasks under matched conditions that turn retained experience on and off. It spans 26 scenarios and 204 episodes across memory, procedural reuse, information gathering, and update. We report both later-task gains and whether those gains follow the intended save, retrieve, and update pathway.
PAUSE2026relevant
PAUSE: A User-Centric Benchmark for Personal AI Assistants in Unified Service Environments
Goals to track coordinate actions across heterogeneous user-owned services while respecting user-specific configurations and authorization/permission constraints; maintain consistency with evolving environment state across multi-turn interactions; for open-ended service-management tasks, satisfy semantic/behavioral trajectory-level goals; for constraint-intensive tasks, satisfy deterministic state-based verification conditions
Must keep track of The agent must track persistent user state, service-specific configurations and permissions, and prior actions taken across heterogeneous services, coordinating consistently across multi-turn interactions that scale from about 2 to over 5 dialogue rounds and roughly 13 to 22+ tool calls depending on difficulty.
structure: DAG-with-precedence · goals arrive: given-up-front · count: easy tasks average 2.11 dialogue rounds · subgoal credit: yes · 0 citations · unit: turns
Horizoneasy tasks average 2.11 dialogue rounds and 12.85 assistant tool calls; hard tasks average 5.23 dialogue rounds and 22.07 assistant tool calls
the paper’s own words
PAUSE captures core challenges of real-world assistant deployment by requiring agents to coordinate actions across heterogeneous user-owned resources while maintaining consistency with environment state and authorization constraints over multi-turn interactions... To support principled and reproducible evaluation, PAUSE adopts a multi-regime evaluation framework aligned with task characteristics. Open-ended service management tasks are assessed using semantic and trajectory-level behavioral metrics, while constraint-intensive tasks admit deterministic, state-based verification.
Avg. Rounds denotes the average number of dialogue turns ... consistently require more dialogue rounds and tool invocations across all evaluated models
PM-Bench2026in
PM-Bench: Evaluating Prospective Memory in LLM Agents
Goals to track maintain multiple ongoing and deferred intentions across a simulated week; execute a delayed intention at the correct future cue/state while continuing an ongoing activity; monitor latent environment changes relevant to deferred tasks
Must keep track of Agent must track multiple deferred intentions and their trigger conditions, continuously monitor latent environment/state changes, and decide at each point whether any deferred task is now due, while continuing an ongoing activity across a simulated week.
structure: set-of-independent · goals arrive: mixed:given-up-front-intentions-with-environme · subgoal credit: no · 1 citations · unit: simulated-days
Horizonseven
the paper’s own words
Inspired by the Virtual Week paradigm from cognitive science, PM-Bench evaluates how well LLM agents maintain user intentions, execute delayed intentions, and monitor latent environment changes. Over the course of a simulated seven-day week, agents must continue an ongoing activity while deciding whether any deferred task is due.
Pare-Bench2026relevant
Proactive Agent Research Environment: Simulating Active Users to Evaluate Proactive Assistants
Goals to track observe evolving app/user state via a stateful finite-state-machine simulation to infer the user's current goal; correctly time an intervention (neither too early nor too late) once a need is inferred; orchestrate actions across multiple apps (communication, productivity, scheduling, lifestyle) to address the inferred goal
Must keep track of The agent must continuously observe the simulated user's stateful, sequential app interactions to infer an emerging goal, decide the right moment to intervene, and coordinate the intervention across multiple apps.
structure: hierarchical · goals arrive: emitted-by-environment-over-time · count: 143 diverse tasks spanning communication · subgoal credit: yes · 11 citations · unit: turns
Horizonsimulation runs for a maximum of 10 turns
the paper’s own words
Pare models applications as finite state machines with stateful navigation and state-dependent action space for the user simulator, enabling active user simulation. Building on this foundation, we present Pare-Bench, a benchmark of 143 diverse tasks spanning communication, productivity, scheduling, and lifestyle apps, designed to test context observation, goal inference, intervention timing, and multi-app orchestration.
The simulation runs for a maximum of 10 turns.
PersonaMem-v3: Toward Omni-Platform Personal Intelligence for Holistic User Understanding, Recommendation, and Agentic Tasks
Goals to track build holistic cross-platform user understanding from social media, chatbot, calendar, and AI-companion engagement histories; personalize responses to reflect the user's evolving preferences over time; rerank recommendations on social media in a steerable way; act proactively across platforms when appropriate; hold back from personalizing when it would be inappropriate, repetitive, outdated, or unnecessary
Must keep track of The agent must track a time-indexed model of the user's preferences, intents, habits, and social relationships as they evolve across multiple platforms (social media, chatbot, calendar, AI-companion), and decide when NOT to act or personalize.
structure: set-of-independent · goals arrive: emitted-by-environment-over-time · subgoal credit: no · 0 citations
Horizonthe paper states none
the paper’s own words
The benchmark brings personalization, LLM-powered recommendation, proactiveness, agentic tool use, and geo-temporal reasoning into one framework, anchored in psychology, social-linguistics, and user-behavior theories. It evaluates whether AI agents can infer holistic user understanding from cross-platform evidence, personalize responses, rerank recommendations on social media, follow user steering through natural language, and hold back when personalization would be inappropriate, repetitive, outdated, or unnecessary.
PersonaMem-v3 is seeded from more than one million anonymized real-world engagement histories, most of which are implicit signals, and uses them to construct time-indexed user digital worlds
ProAgentBench2026relevant
ProAgentBench: Evaluating LLM Agents for Proactive Assistance with Real-World Data
Goals to track predict the correct timing for a proactive intervention within a continuous workflow; generate appropriate assist content once an intervention point is identified
Must keep track of Agent must model long-term memory and historical, pre-assistance behavioral context (bursty interaction patterns, B=0.787) to decide both whether/when to intervene and what to say.
structure: hierarchical · goals arrive: implied-by-constraints · count: 2 (timing prediction; assist-content gen · subgoal credit: unclear · 10 citations · also Business, office &amp; enterprise work
Horizonthe paper states none
the paper’s own words
a hierarchical task framework that decomposes proactive assistance into timing prediction and assist content generation
a privacy-compliant dataset with 28,000+ events from 500+ hours of real user sessions
ProEvent2026in
ProEvent: An Event-centric Benchmark for Proactive Agents
Goals to track identify new upcoming events, including implicit ones, from ongoing instant-messaging chats; maintain and update a timetable of a user's events over time; time proactive responses correctly (neither too early nor too late); handle event cancellations correctly rather than overacting; produce correct single-step and multi-step responses per event
Must keep track of The agent must maintain a live, updatable timetable of a user's upcoming events, tracking concurrent chat threads, noise, and event cancellations, and decide both when and how (single- vs. multi-step) to respond as messages arrive.
structure: set-of-independent · goals arrive: emitted-by-environment-over-time · subgoal credit: yes · 2 citations
Horizonthe paper states none
the paper’s own words
ProEvent provides synthesized yet realistic chats that consider the dynamic interaction among users, concurrent chat threads, and noise in the real world, and evaluates proactive agents on response timing, single-step response correctness, and multi-step response correctness.
RealMem2026in
RealMem: Benchmarking LLMs in Real-World Memory-Driven Interaction
Goals to track track evolving project goals across long-term, cross-session dialogues; manage dynamic context dependencies (schedule/memory) inherent to real-world projects; respond correctly to natural user queries grounded in accumulated project history across eleven scenarios
Must keep track of The system must track long-term project states and dynamic context/schedule dependencies across more than 2,000 cross-session dialogues, since project goals evolve over time rather than remaining fixed.
structure: sequential-chain · goals arrive: emitted-by-environment-over-time · count: over 2,000 cross-session dialogues acros · subgoal credit: unclear · 16 citations · also Business, office &amp; enterprise work
Horizonthe paper states none
the paper’s own words
RealMem comprises over 2,000 cross-session dialogues across eleven scenarios, utilizing natural user queries for evaluation.
Setoka2026relevant
Setoka: A Benchmark for Hierarchical User Understanding in Personalized Agents over Heterogeneous Data
Goals to track retrieve explicit facts from past interactions (semantic memory); recall specific past episodes/events accurately (episodic memory); infer recurring behavior patterns from heterogeneous data over time (behavior pattern); infer abstract personality traits from heterogeneous, fragmented information (personality trait)
Must keep track of The memory-augmented agent must retain and integrate heterogeneous user data (explicit facts, episodic events, behavioral observations) dispersed over long-term interaction history to answer queries at each of the four hierarchical understanding levels.
structure: hierarchical · goals arrive: given-up-front · count: four levels of user understanding (seman · subgoal credit: no · 0 citations
Horizonthe paper states none
the paper’s own words
Setoka defines four levels of user understanding, i.e., semantic memory, episodic memory, behavior pattern, and personality trait... Our comprehensive evaluation reveals that while existing systems perform well on semantic memory retrieval, their performance declines on episodic memory.
motivating the design of memory mechanisms for cross-source integration and abstraction over long-term user behavior
Shopping Companion: Benchmarking and Training LLM Agents for Long-Horizon Preference-Grounded E-Commerce Tasks
Goals to track recommend products correctly aligned with preferences expressed across long-horizon conversations; manage a budget while shopping; assemble bundle deals satisfying multiple items' constraints jointly; correctly recall and apply user preferences carried over from earlier sessions (cross-session preference memory)
Must keep track of Agent must accumulate and correctly recall user shopping preferences across sessions, and verify product attributes against user requirements at each tool call, to avoid cascading preference-hallucination errors.
structure: set-of-independent · goals arrive: emitted-by-environment-over-time · count: 2 shopping tasks requiring cross-session · subgoal credit: yes · 4 citations
Horizonthe paper states none
the paper’s own words
we design annotation-free, tool-wise rewards that provide process supervision for each tool call, alleviating reward sparsity in long-horizon tasks
the absence of benchmarks for evaluating long-term preference-aware shopping tasks
SovereignNegotiation-Bench: Evaluating User-Owned Personal Agents In Delegated Bargaining Under Privacy, Consent, Evidence, And Institutional Pressure
Goals to track reach a negotiated agreement on the user's behalf (e.g., cost splits, refunds, subscription changes); preserve user utility while negotiating; avoid privacy leakage and consent violations during negotiation; ground claims in evidence and maintain auditability of the negotiation trace; escalate appropriately rather than over-concede under institutional pressure
Must keep track of The agent must track its private utilities/disclosure constraints, evidence requirements, and institutional-pressure cues across a multi-turn negotiation trace, keeping agent-visible observable state separate from evaluator-only labels.
structure: set-of-independent · goals arrive: given-up-front · count: 8 jointly tracked evaluation objectives · subgoal credit: yes · 1 citations · unit: actions
Horizon61,135 parsed action rows across 13,440 frozen-prompt live trajectories (~4.5 actions/trajectory)
the paper’s own words
The benchmark separates agent-visible observable state from evaluator-only labels and evaluates agreement success jointly with user utility, privacy, consent, evidence grounding, concession discipline, escalation, and auditability.
We report an artifact-backed validation over 240 scenarios, 4 model families, 14 baselines, 13,440 frozen-prompt live trajectories, 61,135 parsed action rows, and a blinded 3-annotator audit over 300 items.
SovereignPA-Bench: Evaluating User-Owned Personal Agents under Evolving Intent, Platform Mediation, and Consent Constraints
Goals to track advance a user's current, evolving interests while respecting privacy boundaries, consent constraints, and evidence requirements; minimize user burden while resisting manipulative platform incentives; preserve auditability of decisions across 120 sovereignty stress scenarios
Must keep track of Agent must track the user's evolving intent, what has been disclosed to which platform/party (ObservableState vs. evaluator-only HiddenLabels), consent already given, and accumulated burden across a scenario, all while resisting manipulative incentives.
structure: DAG-with-precedence · goals arrive: mixed:given-up-front-user-goal-plus-environmen · count: 120 sovereignty stress scenarios across · subgoal credit: yes · 1 citations
Horizonthe paper states none
the paper’s own words
The benchmark separates agent-visible ObservableState from evaluator-only HiddenLabels, reports component metrics for task success, alignment, privacy, consent, evidence, manipulation, burden, and auditability, and preserves paired scenario ordering for model and policy comparisons.
We evaluate 120 sovereignty stress scenarios across 4 model families and 8 policy baselines, yielding 3,840 frozen-prompt trajectories with raw prompts, outputs, provider-form responses, parsed actions, recomputable metrics, hard-set analyses, qualitative cases, and a blinded 3-annotator audit over 240 items.
StateMemBench2026relevant
Can Agent Memory Systems Track Evolving State?
Goals to track track the evolving state of facts, constraints, and decisions as they are revised over a long multi-session interaction; answer questions reflecting the CURRENT state, not a superseded prior state
Must keep track of Agent's memory system must track the current value of each evolving fact/constraint/decision plus its supersession and relational dependencies, distinguishing current from superseded state at each query point.
structure: sequential-chain · goals arrive: emitted-by-environment-over-time · count: 234 multi-session scenarios spanning two · subgoal credit: no · 0 citations
Horizonthe paper states none
the paper’s own words
we argue an effective memory system must track the evolving state of the world; as facts, constraints, and decisions are revised over a long interaction, answers must reflect the current state and not a superseded one. We define this capability as state tracking and instantiate it in StateMemBench, a benchmark of 234 multi-session scenarios spanning two conversation-length regimes.
StreamMemBench2026relevant
StreamMemBench: Streaming Evaluation of Agent Memory for Future-Oriented Assistance
Goals to track correctly recall/use evidence observed in an initial task drawn from a streaming egocentric anchor; incorporate feedback/interaction experience from the initial task into a later follow-up task; carry evidence forward from what the agent observes and how the user interacts with it, to future similar tasks
Must keep track of The agent must carry stored evidence and interaction feedback forward from an initial task to a corresponding follow-up task drawn from continuous streaming egocentric observations, diagnosed via four metrics (evidence recall, initial evidence use, feedback incorporation, follow-up reuse).
structure: sequential-chain · goals arrive: given-up-front · count: two-step task sequence (initial + follow · subgoal credit: no · 0 citations · unit: other:task-steps
Horizon2 (initial task + follow-up task) per evidence anchor
the paper’s own words
We introduce StreamMemBench, a streaming benchmark that constructs a two-step task sequence around each evidence anchor from EgoLife egocentric streams. The initial task tests evidence use, while the follow-up task tests whether feedback and interaction experience are reused. Four metrics diagnose evidence recall, initial evidence use, feedback incorporation, and follow-up reuse.
Supersede2026in
Supersede: Diagnosing and Training the Memory-Update Gap in LLM Agents
Goals to track answer using the current (most up-to-date) value of a fact that changes over time (e.g., a user's address, a price, a plan); discard/avoid using superseded (stale) fact values; maintain a bounded, self-maintained memory that keeps pace as the conversation grows
Must keep track of The agent must maintain a bounded, self-maintained memory of facts across long, multi-session interactions and always resolve to a fact's most current value, discarding superseded ones, as the conversation grows arbitrarily long.
structure: set-of-independent · goals arrive: emitted-by-environment-over-time · subgoal credit: no · 7 citations · unit: other:relative-conversation-length-growth-factor
Horizonconversation length grows 24x (accuracy falls from 68% to 28% over this range, n=25)
the paper’s own words
We release Supersede, an open reinforcement-learning environment (on the verifiers / prime-rl stack) that turns this measurement into a training signal: agents are rewarded for answering from the current value and penalized for stale ones.
as the conversation grows 24x, accuracy falls further (from 68% to 28%), and granting the agent proportionally more memory yields no detectable recovery
TANGLE2026relevant
When Personal Memory Has No Single Answer: Evaluating LLM Agents under Irreducible Conflict
Goals to track recognize underdetermination when personal memory has no single answer; retain/preserve conflicting alternatives rather than collapsing to one definitive answer; seek clarification rather than acting on unjustified overconfidence; choose an action appropriate to Context-Partitioned, Behavior-Oscillation, or Source-Contradiction conflict types
Must keep track of Agent must preserve rather than resolve conflicting evidence, monitor five behavior dimensions (conflict perception, causal reasoning, confidence calibration, clarification seeking, memory faithfulness), and in the pipeline track extract and preserve conflict-bearing relations from multi-session dialogues.
structure: set-of-independent · goals arrive: given-up-front · count: 541 instances across 40 personas and thr · subgoal credit: yes · 0 citations
Horizonthe paper states none
the paper’s own words
we evaluate two tracks---an oracle track with curated memory and a pipeline track that extracts memory from multi-session dialogues---on five dimensions: conflict perception, causal reasoning, confidence calibration, clarification seeking, and memory faithfulness.
VehicleMemBench: An Executable Benchmark for Multi-User Long-Term Memory in In-Vehicle Agents
Goals to track model multi-user preferences continuously as they evolve over time; resolve inter-user preference conflicts; correctly invoke 23 tool modules to reach a predefined target environment state; track changing user habits across many historical memory events
Must keep track of Agent must track evolving, sometimes conflicting, per-user preferences and habits across 80+ historical memory events, verifying its actions by comparing the resulting environment state to a predefined target state.
structure: sequential-chain · goals arrive: emitted-by-environment-over-time · count: 23 tool modules; over 80 historical memo · subgoal credit: no · 1 citations · unit: other:historical-memory-events-per-sample
Horizonover 80
the paper’s own words
The benchmark evaluates tool use and memory by comparing the post-action environment state with a predefined target state, enabling objective and reproducible evaluation without LLM-based or human scoring. VehicleMemBench includes 23 tool modules, and each sample contains over 80 historical memory events.
VibeLifeBench: Can Your Life Agent Be Proactive and Persistent in a Living World?
Goals to track complete 200 long-horizon tasks across ten everyday-life domains over scripted multi-week timelines; proactively decide when to act, ask, or stay silent without being explicitly prompted; notice unannounced/silent world changes by re-inspecting the world; keep one plan coherent from the first day to the last while upholding unstated implicit constraints
Must keep track of Agent must track end-state goals, the timeliness of its own actions, and implicit constraints across scripted multi-week timelines in a simulated world of 22 mock services that changes on its own clock, much of it silently.
structure: open-ended · goals arrive: implied-by-constraints · count: 200 tasks across ten everyday-life domai · subgoal credit: no · 0 citations · unit: other:multi-week-scripted-timeline
Horizonmulti-week (200 scripted tasks)
the paper’s own words
Every task is graded by fine-grained, weighted checks that read only what the agent actually left behind, covering the end state, the timeliness of its actions, and whether it upheld the implicit constraints.
We introduce VibeLifeBench, a benchmark of 200 long-horizon tasks across ten everyday-life domains. Each task is a scripted multi-week timeline in a simulated world of 22 mock services.
VitaBench 2.0: Evaluating Personalized and Proactive Agents in Long-Term User Interactions
Goals to track continuously extract, utilize, and update evolving user preferences across a temporally ordered sequence of tasks for one user; proactively recognize missing information and actively acquire it from users/environment before making a decision
Must keep track of The agent must continuously extract, update, and apply an evolving set of user preferences (which can be added, deleted, or modified between tasks) across a temporally ordered sequence of at least 10 tasks per user, and proactively recognize when it needs to ask for missing information.
structure: sequential-chain · goals arrive: emitted-by-environment-over-time · count: 56 users, 819 subtasks total, 66 tools, · subgoal credit: yes · 5 citations · unit: episodes
Horizonusers have at least 10 tasks in their temporally ordered task sequences (56 users, 819 subtasks total)
the paper’s own words
In VitaBench 2.0, tasks are organized as temporally ordered sequences for individual users, where preferences are embedded in fragmented and heterogeneous interactions. Successful completion of tasks requires the agent to continuously extract, utilize, and update user preferences from these interactions. We further evaluate proactiveness through tasks that require agents to recognize missing information and actively acquire it from users or environments before making decisions.
We report the task index of up to 10, as users have at least 10 tasks in their task sequences.
WorldBench2026relevant
WorldBench: Culturally Grounded Benchmark for Multilingual Agents
Goals to track complete genuine, persona-grounded everyday workflows across seven languages and eight cultures; preserve sandbox environment state while acting via structured actions; minimize unnecessary modification/side effects while completing the requested task
Must keep track of Agent must track sandbox state to complete persona-grounded workflows correctly while minimizing unwanted modifications, across long-horizon tasks and language/culture variation.
structure: sequential-chain · goals arrive: given-up-front · count: 1,600 tasks across seven languages and e · subgoal credit: no · 0 citations
Horizonthe paper states none
the paper’s own words
we introduce Constrained Task Success (CTS), which combines natural language instructions and testbeds to score task completion, minimal modification, and other complementary metrics through deterministic and LLM-as-a-Judge evaluations.
current agents remain brittle in multilingual, agentic scenarios, especially for long-horizon tasks and under state-preservation constraints
WorldMemArena: Evaluating Multimodal Agent Memory Through Action-World Interaction
Goals to track track and update an evolving personal/task state across a lifelong-evolution scenario; write, maintain, retrieve, and use memory correctly through the four-stage Action-World Interaction Loop; use visual evidence from real observations, actions, and feedback (Agentic Execution); answer QA correctly against gold memory points while resisting annotated distractors
Must keep track of The agent must write new memory from observations/actions/feedback, maintain it against evolving personal/task state (Lifelong Evolution), and correctly retrieve/use it later, annotated against gold memory points, updates, distractors, and evidence chains across multiple sessions.
structure: sequential-chain · goals arrive: given-up-front · count: 400 multi-session multimodal tasks · subgoal credit: no · 1 citations · unit: other:not-stated
Horizonthe paper states none
the paper’s own words
we formulate multimodal agent memory as an Action-World Interaction Loop with an observable four-stage lifecycle, and instantiate it in WorldMemArena: 400 multi-session multimodal tasks spanning Lifelong Evolution (evolving personal and task states) and Agentic Execution (memory from real observations, actions, and feedback), annotated with gold memory points, updates, distractors, and evidence chains for stage-level diagnosis.
π-Bench2026in
π-Bench: Evaluating Proactive Personal Assistant Agents in Long-Horizon Workflows
Goals to track identify and act on hidden/unstated user needs before they are explicitly stated; complete each of 100 multi-turn tasks across 5 domain-specific user personas; resolve inter-task dependencies across sessions; maintain continuity of prior interaction context across sessions to resolve later proactive intents
Must keep track of The agent must track hidden/unstated user intents, inter-task dependencies, and information carried over across sessions, jointly measuring proactivity (anticipating needs) and task completion (executing them) over extended interactions.
structure: DAG-with-precedence · goals arrive: mixed:given-up-front-with-hidden-intents-impli · count: 100 multi-turn tasks across 5 domain-spe · subgoal credit: no · 3 citations
Horizonthe paper states none
the paper’s own words
By incorporating hidden user intents, inter-task dependencies, and cross-session continuity, $\pi$-Bench evaluates agents'ability to anticipate and address user needs over extended interactions, jointly measuring proactivity and task completion in long-horizon trajectories that better reflect real-world use.
$\pi$-Bench, a benchmark for proactive assistance comprising 100 multi-turn tasks across 5 domain-specific user personas
Evo-Memory: Benchmarking LLM Agent Test-time Learning with Self-Evolving Memory
Goals to track search, adapt, and evolve memory after each interaction in a sequential task stream; solve each task drawn from 10 diverse multi-turn goal-oriented and single-turn reasoning/QA datasets; reuse experience accumulated from earlier tasks to improve on later tasks in the same stream
Must keep track of Agent must search, adapt, and evolve its own memory continuously after each interaction across a sequential task stream, integrating reasoning, task actions, and memory updates to achieve continual improvement.
structure: sequential-chain · goals arrive: emitted-by-environment-over-time · count: 10 diverse multi-turn goal-oriented and · subgoal credit: no · 124 citations
Horizonthe paper states none
the paper’s own words
Evo-Memory structures datasets into sequential task streams, requiring LLMs to search, adapt, and evolve memory after each interaction... we unify and implement over ten representative memory modules and evaluate them across 10 diverse multi-turn goal-oriented and single-turn reasoning and QA datasets.
LLMs are required to handle continuous task streams, yet often fail to learn from accumulated interactions, losing valuable contextual insights, a limitation that calls for test-time evolution
Forgetful but Faithful: A Cognitive Memory Architecture and Benchmark for Privacy-Aware Generative Agents
Goals to track maintain narrative coherence across long-term interaction; complete multi-step goals despite a bounded memory budget; preserve social recall accuracy under six candidate forgetting policies; preserve privacy by not retaining/leaking information beyond its retention schema; minimize cost while balancing the above under different memory budgets
Must keep track of The agent must track its own memory budget consumption, which information to retain vs. forget under its chosen policy, ongoing multi-step goal progress, and social/privacy-relevant facts, across long-term interactive scenarios.
structure: set-of-independent · goals arrive: given-up-front · count: 300 evaluation runs across multiple memo · subgoal credit: no · 11 citations · unit: other:not-stated
Horizonthe paper states none
the paper’s own words
We present the Forgetful but Faithful Agent (FiFA) benchmark, a comprehensive evaluation framework that assesses agent performance across narrative coherence, goal completion, social recall accuracy, privacy preservation, and cost efficiency. Through extensive experimentation involving 300 evaluation runs across multiple memory budgets and agent configurations, we demonstrate that our hybrid forgetting policy achieves superior performance (composite score: 0.911)
Mem-alpha2025relevant
Mem-α: Learning Memory Construction via Reinforcement Learning
Goals to track extract and store relevant content from sequential information chunks into an external memory system; organize stored content across core, episodic, and semantic memory components; correctly answer downstream questions using the full accumulated interaction history; invoke the right memory-operation tool (from a multi-tool memory architecture) at the right time
Must keep track of The agent must track which information has been extracted and stored, how it is structured across core/episodic/semantic memory, and how it should be updated as new chunks arrive, since reward is downstream QA accuracy over the full interaction history.
structure: sequential-chain · goals arrive: emitted-by-environment-over-time · subgoal credit: yes · 27 citations · unit: other:tokens-of-accumulated-interaction-history
Horizontrained up to 30k tokens; generalizes to sequences exceeding 400k tokens (>13x training length)
the paper’s own words
During training, agents process sequential information chunks, learn to extract and store relevant content, then update the memory system. The reward signal derives from downstream question-answering accuracy over the full interaction history, directly optimizing for memory construction.
Despite being trained exclusively on instances with a maximum length of 30k tokens, our agents exhibit remarkable generalization to sequences exceeding 400k tokens, over 13x the training length, highlighting the robustness of Mem-alpha.
Evaluating Memory in LLM Agents via Incremental Multi-Turn Interactions
Goals to track accurately retrieve previously seen information from accumulated context; adapt to and learn from new information at test time; understand and reason over long-range accumulated context; selectively forget information that is no longer relevant or valid
Must keep track of The memory agent must incrementally accumulate, update, and retrieve information across multi-turn interactions, and correctly discard information rendered obsolete while retaining what remains useful.
structure: set-of-independent · goals arrive: emitted-by-environment-over-time · count: four core memory competencies (accurate · subgoal credit: yes · 213 citations · also Information seeking &amp; deep research
Horizonthe paper states none
the paper’s own words
we identify four core competencies essential for memory agents: accurate retrieval, test-time learning, long-range understanding, and selective forgetting... Our benchmark transforms existing long-context datasets and incorporates newly constructed datasets into a multi-turn format, effectively simulating the incremental information processing characteristic of memory agents.
MemoryBench2025relevant
MemoryBench: A Benchmark for Memory and Continual Learning in LLM Systems
Goals to track learn from accumulated user feedback received during service time; apply continually-updated knowledge correctly across multiple domains, languages, and task types; outperform static (non-continual-learning) baselines after repeated feedback exposure
Must keep track of The system must track and integrate a stream of accumulated user feedback over service time, across multiple domains, languages, and task types, updating its internal state/parameters rather than treating each query independently.
structure: sequential-chain · goals arrive: emitted-by-environment-over-time · subgoal credit: no · 49 citations
Horizonthe paper states none
the paper’s own words
we propose a user feedback simulation framework and a comprehensive benchmark covering multiple domains, languages, and types of tasks to evaluate the continual learning abilities of LLMsys. Experiments show that the effectiveness and efficiency of state-of-the-art baselines are far from satisfying
testing their abilities to learn from accumulated user feedback in service time
Know Me, Respond to Me: Benchmarking LLMs for Dynamic User Profiling and Personalized Responses at Scale
Goals to track internalize a user's inherent traits and preferences from interaction history; track how the user's profile/preferences evolve over time across sessions; generate a personalized response consistent with the current (most up-to-date) state of the user's profile in a new scenario
Must keep track of The chatbot must internalize inherent user traits, track how the user's profile evolves session by session, and select the response consistent with the user's current (not stale) profile state when answering an in-situ, first-person query.
structure: sequential-chain · goals arrive: emitted-by-environment-over-time · count: over 180 simulated user-LLM interaction · subgoal credit: no · 158 citations · unit: sessions
Horizonover 180 simulated interaction histories, each containing up to 60 sessions of multi-turn conversations
the paper’s own words
PERSONAMEM features curated user profiles with over 180 simulated user-LLM interaction histories, each containing up to 60 sessions of multi-turn conversations across 15 real-world tasks that require personalization... frontier models such as GPT-4.1, o4-mini, GPT-4.5, o1, or Gemini-2.0 achieving only around 50% overall accuracy
PROBE2025relevant
Beyond Reactivity: Measuring Proactive Problem Solving in LLM Agents
Goals to track search for unspecified, unprompted issues across the user's available context/data; identify the specific bottleneck underlying an ambiguous problem; execute an appropriate resolution action autonomously
Must keep track of The agent must track candidate evidence surfaced during open-ended search, determine which issues are genuine unresolved bottlenecks, and carry that determination forward into selecting and executing a resolution, all without being told what to look for.
structure: sequential-chain · goals arrive: self-generated-by-agent · subgoal credit: yes · 8 citations · also Business, office &amp; enterprise work
Horizonthe paper states none
the paper’s own words
PROBE decomposes proactivity as a pipeline of three core capabilities: (1) searching for unspecified issues, (2) identifying specific bottlenecks, and (3) executing appropriate resolutions.
current benchmarks are constrained to localized context, limiting their ability to test reasoning across sources and longer time horizons.
SimuHome2025relevant
SimuHome: A Temporal- and Environment-Aware Benchmark for Smart Home LLM Agents
Goals to track answer state-inquiry questions about the current smart-home environment; infer implicit user intent behind an ambiguous request; execute explicit device-control commands via SimuHome APIs; schedule and coordinate multi-device workflows whose effects evolve environmental variables over time; recognize and appropriately reject infeasible requests
Must keep track of The agent must track how already-issued device commands change environmental variables over (accelerated) simulated time and use this state to judge whether new requests are feasible or already satisfied.
structure: DAG-with-precedence · goals arrive: mixed:given-up-front-plus-emitted-by-environme · count: 600 episodes; 18 agents evaluated across · subgoal credit: yes · 7 citations · unit: episodes
Horizon600 episodes; simulator accelerates time so scheduled workflows can be evaluated immediately
the paper’s own words
Our benchmark covers state inquiry, implicit user intent inference, explicit device control, and workflow scheduling, each with both feasible and infeasible requests... An evaluation of 18 agents reveals that workflow scheduling is the hardest category, with failures persisting across alternative agent frameworks and fine-tuning.
We introduce SimuHome, a high-fidelity smart home simulator and a benchmark of 600 episodes for LLM-based smart home agents... For workflow scheduling, the simulator accelerates time so that scheduled workflows can be evaluated immediately.
UserBench2025relevant
UserBench: An Interactive Gym Environment for User-Centric Agents
Goals to track proactively clarify a simulated user's underspecified/vague initial goal; incrementally uncover and track multiple user preferences revealed gradually over a multi-turn interaction; make grounded decisions with tools that align with all of the user's (eventually revealed) intents
Must keep track of The agent must track which of the user's underspecified goals and incrementally revealed preferences it has already clarified, and continue to proactively elicit unclarified ones across the multi-turn interaction while using tools to act on what it has learned so far.
structure: set-of-independent · goals arrive: emitted-by-environment-over-time · subgoal credit: no · 62 citations
Horizonthe paper states none
the paper’s own words
For instance, models provide answers that fully align with all user intents only 20% of the time on average, and even the most advanced models uncover fewer than 30% of all user preferences through active interaction.
UserBench features simulated users who start with underspecified goals and reveal preferences incrementally, requiring agents to proactively clarify intent and make grounded decisions with tools.

Embodied & robotics 46 artifacts · 8 state a horizon

CoCoBench2026relevant
CoCoBench: A Cooperative Coordination Benchmark for Embodied Multi-Agent Task Planning
Goals to track allocate tasks correctly among cooperating embodied agents; respect sequential-ordering constraints between agents' actions; respect mutual-exclusion constraints over shared objects/spaces; execute correct handoff coordination between agents; satisfy conjunctive multi-object-destination goals within executable household tasks
Must keep track of The multi-agent system must track task allocation assignments, ordering/precedence constraints, mutual-exclusion locks on shared objects, and handoff status between agents, calibrated against a task-specific step budget H derived from oracle trajectories.
structure: DAG-with-precedence · goals arrive: given-up-front · count: 897 oracle-validated instances across fo · subgoal credit: yes · 0 citations · also Multi-agent organisations &amp; societies · unit: agent-steps
Horizon~18.4 mean executed action steps (best model, successful episodes); task-specific step budget H calibrated from oracle-validated trajectories
the paper’s own words
CoCoBench contains 897 oracle-validated instances organized around four recurring coordination constructs: task allocation, sequential ordering, mutual exclusion, and handoff coordination. In addition to task success rate, CoCoBench provides construct-level scores that measure whether agents coordinate effectively.
the step budget is H, and episode succeeds when the final state satisfies every goal predicate ... GPT-5.6-sol achieves ... 18.4 mean steps per episode on successful attempts
DeCoNavBench2026relevant
DeCoNav: Dialog enhanced Long-Horizon Collaborative Vision-Language Navigation
Goals to track achieve relay-style handoffs between two robots collaborating on a shared long-horizon navigation task; reach rendezvous points to exchange information/objects between robots; dynamically reassign and replan subgoals in response to new cross-agent evidence, uncertainty, or conflicts
Must keep track of Each robot must track its own navigation progress plus the other robot's evidence/uncertainty/conflicts communicated via event-triggered dialogue, in order to reassign subgoals and replan under synchronized execution.
structure: DAG-with-precedence · goals arrive: emitted-by-environment-over-time · count: 1,213 tasks across 176 HM3D scenes · subgoal credit: yes · 5 citations · also Multi-agent organisations &amp; societies
Horizonthe paper states none
the paper’s own words
When informative events such as new evidence, uncertainty, or conflicts arise, dialogue is triggered to dynamically reassign subgoals and replan under synchronized execution. Implemented in DeCoNavBench with 1,213 tasks across 176 HM3D scenes
Long-horizon collaborative vision-language navigation (VLN) is critical for multi-robot systems to accomplish complex tasks beyond the capability of a single agent.
DunphyBench2026relevant
Long-Horizon Embodied Decision-Making via Multimodal Memory Compression
Goals to track navigate through multiple embodied housing environments to gather evidence; integrate multimodal, multi-source input into coherent knowledge under partial observation; make a final housing decision aligned with multi-dimensional, partly implicit human preferences
Must keep track of Agent must accumulate multimodal evidence across multiple candidate housing environments under partial observation and integrate it into a coherent decision aligned with multi-dimensional, partly implicit human preferences.
structure: sequential-chain · goals arrive: implied-by-constraints · subgoal credit: no · 0 citations · also Personal assistants &amp; long-term memory
Horizonthe paper states none
the paper’s own words
we propose DunphyBench, a new benchmark for evaluating agents on long-horizon human-centered embodied decision-making, where the agent must navigate through multiple embodied housing environments and make decisions that align with multi-dimensional human preferences.
This shift requires agents to accumulate evidence over long horizons, interpret implicit user preferences, and compare multiple candidates under partial observations.
Resilience Matters for Embodied Agents System: New Metrics, Systematic Evaluation, and Optimization
Goals to track complete a household task despite perturbations/unexpected disruptions during execution; recover, stabilize, and gracefully extend behavior after a perturbation rather than simply succeeding or failing outright; maintain low recovery cost and stability across iterative system updates
Must keep track of The evaluation layer must track the full execution trajectory of each household task under perturbation, not just final success/failure, to compute process-level resilience metrics like recovery cost and stability.
structure: hierarchical · goals arrive: emitted-by-environment-over-time · count: 3 resilience dimensions (Rebound, Stabil · subgoal credit: yes · 0 citations
Horizonthe paper states none
the paper’s own words
Across 400 household tasks with 10 EAS, we reveal the process-level distinction hidden by outcome metrics, including recovery cost differences among successful episodes ($\Delta C_{rec}=25.2$), increased instability and task-family degradation. Metrics-guided optimizations reduce recovery cost and increase stability, graceful extensibility completion, showing the diagnostic effect of resilience evaluation.
Ego2World2026relevant
Ego2World: Compiling Egocentric Cooking Videos into Executable Worlds for Belief-State Planning
Goals to track plan under partial observation using a belief graph built from local observations; remember objects and track state changes in a hidden symbolic world graph; recover and replan when actions fail without observing the true world state; complete cooking-task objectives via correct sequences of graph-governed state transitions
Must keep track of The agent must maintain and update its own partial belief graph of the world (objects, state changes) using only local observations and execution feedback, separate from the simulator's hidden ground-truth world graph, and replan without directly observing true state.
structure: DAG-with-precedence · goals arrive: given-up-front · subgoal credit: no · 0 citations
Horizonthe paper states none
the paper’s own words
During evaluation, the simulator maintains the hidden world graph, while the agent plans over its own partial belief graph using only local observations and execution feedback. This separation forces agents to update memory and replan without observing the true world state.
Embodied agents in household environments must plan under partial observation: they need to remember objects, track state changes, and recover when actions fail.
ABot-AgentOS: A General Robotic Agent OS with Lifelong Multi-modal Memory
Goals to track navigate indoor/outdoor/hybrid scenes to reach targets; search for and identify specified objects; conduct NPC dialogue as part of a task; respond appropriately to dynamic events introduced mid-task; achieve trace-grounded, verifiable task completion across difficulty levels
Must keep track of Agent must maintain a persistent, source-grounded multi-modal graph memory across dialogue, visual observations, spatial context, temporal relations, and task traces, continually updated and consulted for later decisions.
structure: hierarchical · goals arrive: mixed:given-up-front-tasks-with-environment-em · count: over 200 tasks across 16 scenes and four · subgoal credit: unclear · 0 citations · also Personal assistants &amp; long-term memory
Horizonthe paper states none
the paper’s own words
we introduce EmbodiedWorldBench, an executable benchmark with 16 indoor, outdoor, and hybrid scenes, four difficulty levels, and over 200 tasks involving navigation, object search, NPC dialogue, dynamic events, and trace-grounded scoring
TaskGround: Structured Executable Task Inference for Full-Scene Household Reasoning
Goals to track identify task-relevant entities within a complete, cluttered household scene; recover intended task conditions implied by a situated (underspecified) household request; resolve ordering constraints among sub-actions from surrounding scene context; produce a grounded, skill-level action sequence that correctly executes the inferred task structure
Must keep track of The agent must track which entities, conditions, and ordering constraints it has inferred as task-relevant from a complete household scene, and maintain this executable task structure while compiling and executing a grounded skill-level action sequence.
structure: sequential-chain · goals arrive: implied-by-constraints · count: 400 household tasks (goal-oriented and p · subgoal credit: no · 0 citations
Horizonthe paper states none
the paper’s own words
we introduce FullHome, a human-validated evaluation suite of 400 household tasks spanning diverse home-scale environments and both goal-oriented and process-constrained requirements. On FullHome, TaskGround improves task success rates by large margins across both proprietary and open-weight models.
an agent must infer executable task structure before producing a grounded skill-level action sequence
HumanCLAW-Bench2026relevant
HumanCLAW: Can Vision-Language Models Act Through a Body?
Goals to track find a specified target within an indoor scene; navigate to the target while avoiding obstacles/maintaining balance; interact with the target once reached; maintain embodied self-awareness of the body's own state (position, goal-reached status, collisions) throughout
Must keep track of The VLM must track where its body is, whether it has reached the goal, and whether it has hit an obstacle -- embodied self-awareness -- across each find-navigate-interact episode, since the paper finds this tracking, not target recognition, is the main bottleneck.
structure: sequential-chain · goals arrive: given-up-front · count: 1,218 long-horizon, egocentric find-navi · subgoal credit: yes · 2 citations · unit: agent-steps
Horizonreported average of 78.5 steps per episode for baseline conditions
the paper’s own words
we build HumanCLAW-Bench: 1,218 long-horizon, egocentric find-navigate-interact episodes across 41 indoor scenes. We test nine state-of-the-art VLMs and find that none solves the benchmark; the best model reaches only a 16.8% success rate. Recognizing the target is not the bottleneck. What current VLMs lack is embodied self-awareness: they lose track of their own body, failing to tell where it is, whether it has reached the goal, or whether it has hit an obstacle.
average steps in tables (e.g., 78.5 steps for baseline conditions)
IMBench2026relevant
IMBench: A Benchmark for Intuitive Robotic Manipulation
Goals to track infer task-relevant physical structure via physical reasoning before acting; generate a feasible action sequence satisfying explicit task constraints (contact-rich manipulation, tool use, multi-stage dependencies); complete each of 35 manipulation tasks across scalable, diverse scenarios
Must keep track of Agent must track inferred physical structure (contacts, affordances), the current stage of a multi-stage manipulation plan, and constraint satisfaction across execution.
structure: DAG-with-precedence · goals arrive: given-up-front · count: 35 tasks; 14K filtered trajectories · subgoal credit: yes · 0 citations
Horizonthe paper states none
the paper’s own words
Our tasks require models to infer task-relevant physical structure and generate feasible action sequences under explicit constraints, including contact-rich manipulation, tool use, and multi-stage dependencies.
We introduce a benchmark of 35 tasks, 14K filtered trajectories, and scalable tools for generating diverse scenarios.
Explore with Long-term Memory: A Benchmark and Multimodal LLM-based Reinforcement Learning Framework for Embodied Exploration
Goals to track perform multi-goal navigation across an embodied environment; answer memory-based questions using episodic memory accumulated during exploration; unify exploratory cognition with decision-making to support lifelong learning; proactively query memory and select exploration frontiers next
Must keep track of The agent must track its accumulated episodic memory, current exploration frontier, and multiple concurrent navigation goals across a long-horizon embodied exploration episode.
structure: sequential-chain · goals arrive: mixed:navigation-goals-given-up-front-explorat · subgoal credit: no · 10 citations · also Personal assistants &amp; long-term memory
Horizonthe paper states none
the paper’s own words
We further construct a corresponding dataset and benchmark, LMEE-Bench, incorporating multi-goal navigation and memory-based question answering to comprehensively evaluate both the process and outcome of embodied exploration.
An ideal embodied agent should possess lifelong learning capabilities to handle long-horizon and complex tasks, enabling continuous operation in general environments.
LongAct2026in
When Robots Do the Chores: A Benchmark and Agent for Long-Horizon Household Task Execution
Goals to track understand a free-form (non-templated) household instruction and decompose it into the right sequence of sub-tasks; manage dependencies among sub-tasks (ordering, shared resources) via a DAG-based plan; maintain persistent memory (spatial and episodic) across a long-horizon household task; adapt the plan reflectively as execution proceeds
Must keep track of The agent must maintain persistent spatial and episodic memory of what it has already done and observed in the household, use this to keep dependency-aware track of remaining sub-tasks, and adaptively replan as new information arises over the long-horizon task.
structure: DAG-with-precedence · goals arrive: given-up-front · subgoal credit: yes · 2 citations · unit: agent-steps
Horizonhuman executors typically require 500+ steps; VLM-based agents often exceed 2,000 steps to complete the same long-horizon household task
the paper’s own words
We introduce LongAct, a benchmark designed to evaluate planning-level autonomy in long-horizon household tasks specified through free-form instructions. By abstracting away embodiment-specific low-level control, LongAct isolates high-level cognitive capabilities such as instruction understanding, dependency management, memory maintenance, and adaptive planning.
While human executors typically require more than 500 steps to complete a task, VLM-based agents often exceed 2,000 steps
MECoBench2026relevant
MECoBench: A Systematic Study of Multimodal Agent Collaboration in Embodied Environments
Goals to track complete embodied tasks under two cooperation structures and three collaboration modes; coordinate communication among multiple multimodal agents in a visually grounded environment; remain robust to noisy priors and exploration conditions via collaboration
Must keep track of Agents must track and communicate task-relevant state among team members across two cooperation structures and three collaboration modes to complete embodied tasks under noisy priors and exploration conditions.
structure: set-of-independent · goals arrive: given-up-front · subgoal credit: no · 2 citations · also Multi-agent organisations &amp; societies
Horizonthe paper states none
the paper’s own words
we introduce MECoBench, a multimodal embodied cooperation benchmark with an evaluation platform spanning diverse real-world tasks, two cooperation structures, and three collaboration modes... (i) Collaboration generally improves embodied task completion, but its benefits depend on balancing collaborative gains against coordination complexity.
MultiUAV-Plat: An LLM-Oriented Platform, Benchmark and Framework for Multi-UAV Collaborative Task Planning
Goals to track assign UAVs to targets correctly under partial observability; perform area search coverage objectives; perform area assignment and patrol objectives across multiple UAVs; satisfy each of the mission's validation checks (mean 6.26/task) via correct multi-vehicle coordination
Must keep track of The framework (Agent4Drone) must track memory, observation, task understanding, planning, execution, and verification state across each mission, checking role-based information access and validation logic per task within a session.
structure: set-of-independent · goals arrive: given-up-front · count: 75 mission sessions, 1,500 natural-langu · subgoal credit: yes · 0 citations · also Multi-agent organisations &amp; societies · unit: other:not-stated
Horizonthe paper states none
the paper’s own words
The MultiUAV-Plat Benchmark contains 75 mission sessions, 1500 natural-language tasks, and 9396 validation checks across target assignment, area search, and area assignment and patrol scenarios. ... Agent4Drone achieves a 57.9% task pass rate, a 74.6% average task check pass rate, and a 72.0% global check pass rate
check counts per task (median of 4, mean of 6.26)
PARTNR-Dialog2026relevant
LLawCo: Learning Laws of Cooperation for Modeling Embodied Multi-Agent Behavior
Goals to track complete a shared household task cooperatively with a partner agent; communicate with the partner (e.g., 'talk when necessary') to align on task objectives; align actions with the partner's behavior and the environment state to avoid inefficient or conflicting actions; apply learned high-level behavioral laws (e.g., 'wait for partner') during planning
Must keep track of The agent must track its partner's current actions/state, the shared household task's progress, and which high-level behavioral laws (e.g., 'talk when necessary,' 'wait for partner') apply at each point in the cooperative plan.
structure: sequential-chain · goals arrive: given-up-front · subgoal credit: no · 0 citations · also Multi-agent organisations &amp; societies
Horizonthe paper states none
the paper’s own words
we introduce PARTNR-Dialog, a large-scale multi-agent communicative and cooperative planning benchmark built on the PARTNR environment. Experiments on existing tasks and our new benchmark demonstrate significant improvements in cooperative efficiency and task success rates.
Embodied agents operating in decentralized and partially observable environments have attracted growing attention in recent years.
PLanAR2026relevant
PLanAR: Planning-Language-Grounded Agentic Reasoning for Robot Manipulation
Goals to track represent and update object-predicate scene states as manipulation proceeds; select and sequence action schemas (with preconditions/effects) to satisfy an open-vocabulary manipulation goal; detect execution failures via stepwise symbolic-effect verification and replan; complete long-horizon kitchen workflows composed of many chained manipulation sub-goals
Must keep track of The agent must maintain a symbolic scene-state representation (object predicates), verify after each action whether expected effects were achieved, and update/replan when execution deviates from the expected symbolic plan across a long-horizon kitchen workflow.
structure: DAG-with-precedence · goals arrive: given-up-front · subgoal credit: no · 3 citations · unit: other:not-stated
Horizonthe paper states none
the paper’s own words
PLanAR uses a planning-language interface to define the VLM reasoning space: object predicates represent scene states, action schemas specify robot skills with preconditions and effects, and symbolic plans provide executable intermediate representations. This interface enables stepwise verification: after each action, PLanAR uses onboard observations to check whether the expected symbolic effects have been achieved, allowing the VLM-based agent to update task states, detect failures, and replan when execution deviates from expectation.
Across robot embodiments, VLM backends, and tasks including stacking, crossword solving, and long-horizon kitchen workflows, PLanAR demonstrates strong real-world capability
RescueBench: Can Embodied Agents Save Lives in the Wild ?
Goals to track explore an unfamiliar environment under multimodal uncertainty to locate a target (multimodal exploration stage); physically rescue/reach the identified target (target rescue stage); navigate back using retained spatial memory (memory-guided return stage); complete a final handoff of the rescued target/information (final handoff stage)
Must keep track of The agent must retain spatial memory of the environment (for the return stage), track clue ambiguity/target identification state, and manage per-level time budgets (180-300 seconds) across the four-stage pipeline.
structure: sequential-chain · goals arrive: given-up-front · count: four sequential stages per episode, eval · subgoal credit: yes · 0 citations · unit: other:wall-clock-seconds
Horizonepisodes are governed by level-dependent time limits: L1-L2: 180 seconds, L3: 240 seconds, L4-L5: 300 seconds, with no step cap
the paper’s own words
We introduce RescueBench, a photo-realistic diagnostic benchmark that instantiates SAR as a four-stage pipeline: multimodal exploration, target rescue, memory-guided return, and final handoff. By combining sequential task composition with stage-level evaluation, RescueBench enables analysis of how exploration and memory failures propagate through embodied rescue workflows.
Resource budgets: episodes are governed by level-dependent time limits (L1-L2: 180 s, L3: 240 s, L4-L5: 300 s) with no step cap.
RoboGraph2026relevant
Compiling and Benchmarking Task-State Horizons for Embodied Agents
Goals to track track the evolving span of task-relevant state transitions (task-state horizon, TSH) induced by both exploration and environmental dynamics; correctly execute high-level plans as task-relevant world state changes, including due to unexpected failures/interventions; complete each of 588 episodes across 84 scenes with varying task-state horizons
Must keep track of The agent must maintain, explore, and update task-relevant world state over the course of a long-horizon rollout, since performance is explicitly measured as a function of the task-state horizon (TSH) -- the span of state transitions it must track.
structure: DAG-with-precedence · goals arrive: mixed:task-given-up-front-state-transitions-em · count: 588 episodes across 84 scenes with varyi · subgoal credit: yes · 0 citations
Horizonthe paper states none
the paper’s own words
We define the span of task-relevant state transitions that an agent must track as task-state horizon (TSH)... Building on RoboGraph, we release a benchmark comprising 588 episodes across 84 scenes with varying TSHs.
SMH-Bench: Benchmarking LLM Agents for Environment-Grounded Reasoning and Action in Smart Homes
Goals to track execute explicit device control and query commands across a smart home with up to 135 devices; schedule automation tasks that must fire correctly over time; correctly handle ambiguous user instructions; personalize reasoning to user intent/preferences as home complexity increases
Must keep track of Agent must track the state of many concurrent devices across a multi-room home, especially for automation-scheduling tasks that must persist and correctly re-trigger over time, and must resolve ambiguous instructions against current home/device state.
structure: set-of-independent · goals arrive: given-up-front · count: 1,100 tasks spanning 7 categories and 22 · subgoal credit: yes · 2 citations
Horizonthe paper states none
the paper’s own words
Built upon HomeEnv, an executable and verifiable smart-home simulator, SMH-Bench contains 1,100 high-quality tasks spanning 7 categories and 22 fine-grained subcategories. It further stratifies tasks across simple, medium and complex homes, ranging from small apartments to dense multi-room environments with 135 devices.
SpatialWorld2026relevant
SpatialWorld: Benchmarking Interactive Spatial Reasoning of Multimodal Agents in Real-World Tasks
Goals to track actively gather egocentric visual evidence under partial observability to resolve a task; complete household, travel, or social-collaboration tasks requiring interactive spatial reasoning; express decisions via a unified text-based action interface across eight heterogeneous simulator backends
Must keep track of Agent must track what it has and has not yet observed (partial observability), accumulate egocentric visual evidence over the course of the task, and reconcile this with a reference trajectory/terminal-state verifier.
structure: sequential-chain · goals arrive: given-up-front · count: 760 human-annotated tasks · subgoal credit: unclear · 2 citations
Horizonthe paper states none
the paper’s own words
Agents must solve tasks under vision-only partial observability, actively gathering egocentric visual evidence and expressing decisions via a unified, text-based action interface native to MLLMs.
SpatialWorld features 760 human-annotated tasks across diverse domains (e.g., household routines, travel, social collaboration)
TypeGo: An OS Runtime for Embodied Agents
Goals to track execute multiple concurrent per-task processes/goals on a shared physical robot body without conflicting over physical subsystems; preempt, resume, or replace a lower-priority task/goal when a new goal arrives; maintain fast first-action responsiveness while longer-horizon planning continues asynchronously in the background
Must keep track of The runtime must track which per-task processes currently hold which physical subsystems, their priority/preemption state, and pending speculative skill-streaming actions, to arbitrate concurrent goals in real time.
structure: set-of-independent · goals arrive: injected-by-user-mid-episode · subgoal credit: yes · 0 citations
Horizonthe paper states none
the paper’s own words
the Skill Kernel arbitrates typed physical subsystems among concurrent per-task processes, a scheduler preempts them and resumes or replaces each by source, and speculative skill streaming hides LLM latency behind ongoing motion... it cuts per-step delay by 50% over step-by-step planning and time-to-first-action by 73% over monolithic planning, while admitting concurrent tasks at low scheduling overhead.
our prototype of Kalos, a Unitree Go2 quadruped, provides preliminary evidence for the design: in our current task suite, it cuts per-step delay by 50% over step-by-step planning and time-to-first-action by 73% over monolithic planning
UniETP2026relevant
UniETP: Unifying Environments for Generalizable Embodied Task Planning
Goals to track execute a sequence of atomic actions within an interactive environment to complete a user-specified task; handle varying task-logic complexity; handle varying instance-grounding complexity; handle varying instruction-understanding complexity, across four unified simulators (AI2-THOR, VirtualHome, Habitat, BEHAVIOR)
Must keep track of The agent must track its progress executing a sequence of atomic actions toward a user-specified goal within a unified observation/action space, while contending with varying levels of task-logic, instance-grounding, and instruction-understanding difficulty across four different underlying simulators.
structure: sequential-chain · goals arrive: given-up-front · subgoal credit: yes · 0 citations
Horizonthe paper states none
the paper’s own words
This paper focuses on the problem of Embodied Task Planning, where an agent is required to execute a sequence of atomic actions within an interactive environment to complete a user-specified task... it formalizes all the simulators into a consistent observation and action space, and builds an evaluation system to support complicated task goal.
VGEBench2026relevant
Towards Generalizable Visually Grounded Exploration of Household Devices
Goals to track form a hypothesis about how to operate a novel household device from visual cues alone (no manual); test the hypothesis via physical interaction and interpret feedback; refine/correct the hypothesis and action based on observed feedback (Hypothesis-Interaction-Refinement loop) until the device is successfully operated
Must keep track of The agent must track its current hypothesis about the device's operation, the history of physical feedback received from prior interaction attempts, and dynamically-calculated interaction budgets based on task complexity, maintaining long-horizon state across the exploration loop.
structure: open-ended · goals arrive: self-generated-by-agent · count: 4,948 single-turn and 10,005 multi-turn · subgoal credit: yes · 0 citations
Horizonthe paper states none
the paper’s own words
This framework simulates multi-turn interaction loops, compelling agents to achieve goals by active visual perception and feedback-driven correction. Experimental results demonstrate that existing VLMs face significant challenges in translating semantic knowledge into physical execution and maintaining long-horizon state tracking.
In total, we constructed 4,948 single-turn and 10,005 multi-turn instructions.
WorldLines: Benchmarking and Modeling Long-Horizon Stateful Embodied Agents
Goals to track answer Memory QA questions correctly using long-term household interaction history; produce correct Embodied Task Plans grounded in remembered user routines and past world/device states; maintain visibility-aware (partially observable) memory across a temporally extended household trace
Must keep track of The agent must maintain visibility-aware memory of user routines, world states, past actions/dialogues, and object/device state changes across long temporally extended household traces to answer Memory QA and produce embodied plans.
structure: other:evidence-linked-samples-drawn-from-a-sha · goals arrive: given-up-front · subgoal credit: yes · 1 citations
Horizonthe paper states none
the paper’s own words
It constructs temporally extended household traces with dialogues, actions, execution feedback, object and device state changes, and converts them into evidence-linked samples for Memory QA and Embodied Task Planning.
3DLLM-Mem: Long-Term Spatial-Temporal Memory for Embodied 3D Large Language Model
Goals to track correctly recall and apply past spatial-temporal observations (episodic memory) to complete embodied tasks in multi-room 3D environments; answer questions and produce captions grounded in accumulated long-term memory of the 3D scene; focus on task-relevant information while maintaining memory efficiency across complex, long-horizon environments
Must keep track of The agent must maintain and selectively query an episodic memory of past spatial and temporal observations across long-horizon, multi-room 3D trajectories, focusing on task-relevant information while remaining memory-efficient.
structure: sequential-chain · goals arrive: given-up-front · count: over 26,000 trajectories and 2,892 embod · subgoal credit: yes · 29 citations · also Personal assistants &amp; long-term memory
Horizonthe paper states none
the paper’s own words
we first introduce 3DMem-Bench, a comprehensive benchmark comprising over 26,000 trajectories and 2,892 embodied tasks, question-answering and captioning, designed to evaluate an agent's ability to reason over long-term memory in 3D environments... Our approach allows the agent to focus on task-relevant information while maintaining memory efficiency in complex, long-horizon environments.
ArtiBench and ArtiBrain: Benchmarking Generalizable Vision-Language Articulated Object Manipulation
Goals to track manipulate articulated objects (kitchen, storage, office, tool appliances) correctly across parts/instances/categories; decompose and validate a sequence of sub-goals for a long-horizon, multi-object manipulation task; maintain physical consistency across the multi-step interaction
Must keep track of System must track subgoal validation state (via the Task Reasoner), accumulated part-level affordances in the Affordance Memory Bank, and physical consistency across a five-level benchmark spanning cross-part/instance/category variation to long-horizon multi-object tasks.
structure: hierarchical · goals arrive: given-up-front · subgoal credit: unclear · 0 citations
Horizonthe paper states none
the paper’s own words
ArtiBench enables structured evaluation from cross-part and cross-instance variation to long-horizon multi-object tasks, revealing the core generalization challenges of articulated object manipulation
ArtiBench, a five-level benchmark covering kitchen, storage, office, and tool environments
Benchmark for Planning and Control with Large Language Model Agents: Blocksworld with Model Context Protocol
Goals to track reach one target block-configuration goal state via a sequence of pick/put/stack/unstack actions; satisfy the goal under increasing plan-length complexity categories (step-2 through step-12); generalize plan validity/optimality across diverse agent architectures connected via a standardized MCP tool interface
Must keep track of The agent must track the current block/stack configuration as pick/put/stack/unstack actions are executed, correctly planning toward the target configuration across plans of up to 12 optimal steps, potentially under constrained block size or partial observability.
structure: sequential-chain · goals arrive: given-up-front · count: five complexity categories; scenarios sp · subgoal credit: no · 2 citations · unit: agent-steps
Horizonscenarios span step-2 through step-12 optimal-plan-length categories (45 step-2, 84 step-4, 152 step-6, 151 step-8, 112 step-10, 46 step-12 scenarios), each involving up
the paper’s own words
We introduce a benchmark with an executable simulation environment representing the Blocksworld problem providing five complexity categories. By integrating the Model Context Protocol (MCP) as a standardized tool interface, diverse agent architectures can be connected to and evaluated against the benchmark without implementation-specific modifications. A single-agent implementation demonstrates the benchmark's applicability, establishing quantitative metrics for comparison of LLM-based planning and execution approaches.
the dataset includes 45 step-2, 84 step-4, 152 step-6, 151 step-8, 112 step-10, and 46 step-12 scenarios, each involving up to five blocks
CookBench2025relevant
CookBench: A Long-Horizon Embodied Planning Benchmark for Complex Cooking Scenarios
Goals to track accurately parse a user's complex cooking intent (Intention Recognition); execute the identified cooking goal through a long-horizon, fine-grained sequence of physical actions (Embodied Interaction); correctly use both macro-level operations (placing orders, purchasing ingredients) and fine-grained embodied physical actions
Must keep track of The agent must track the parsed cooking intent, its progress through a long-horizon fine-grained sequence of physical actions, and the state of ingredients/tools obtained via macro-level operations, across a two-stage cooking task.
structure: sequential-chain · goals arrive: given-up-front · subgoal credit: no · 13 citations
Horizonthe paper states none
the paper’s own words
The core task in CookBench is designed as a two-stage process. First, in Intention Recognition, an agent needs to accurately parse a user's complex intent. Second, in Embodied Interaction, the agent should execute the identified cooking goal through a long-horizon, fine-grained sequence of physical actions.
DeCoBench2025relevant
DeCo: Task Decomposition and Skill Composition for Zero-Shot Generalization in Long-Horizon 3D Manipulation
Goals to track retrieve and chain reusable atomic manipulation skills to complete a compositional long-horizon 3D manipulation task; generalize zero-shot to novel task compositions not seen during training; execute smooth, collision-free transitions between chained skills
Must keep track of The system must track the currently retrieved skill, the object/gripper state at each transition point, and which atomic subtask in the composed sequence is currently active to ensure valid chaining.
structure: sequential-chain · goals arrive: mixed:high-level-instruction-given-up-front-su · count: trained on only 6 atomic tasks; evaluate · subgoal credit: yes · 16 citations
Horizonthe paper states none
the paper’s own words
DeCo decomposes IL demonstrations into modular atomic tasks based on gripper-object interactions, creating a dataset that enables models to learn reusable skills. At inference, DeCo uses a vision-language model (VLM) to parse high-level instructions, retrieve relevant skills, and dynamically schedule their execution. A spatially-aware skill-chaining module ensures smooth, collision-free transitions between skills.
trained on only 6 atomic tasks, completes 9 novel tasks in zero-shot
DeliveryBench: Can Agents Earn Profit in Real World?
Goals to track maximize net profit over the course of an operating shift by choosing which deliveries to accept/complete; meet each accepted delivery's deadline; manage limited resources (transportation expense, vehicle battery) across the shift; interact appropriately with other couriers and customers as needed
Must keep track of The agent must track remaining budget/vehicle battery, delivery deadlines for all currently accepted orders, and its evolving location within a procedurally generated 3D city, over an episode lasting several in-game hours and typically more than 100 action steps.
structure: set-of-independent · goals arrive: given-up-front · count: long-horizon objectives instantiated per · subgoal credit: yes · 5 citations · unit: agent-steps
Horizonepisodes support long-horizon tasks spanning several in-game hours and typically more than 100 action steps
the paper’s own words
Food couriers naturally operate under long-horizon objectives (maximizing net profit over hours) while managing diverse constraints, e.g., delivery deadline, transportation expense, vehicle battery, and necessary interactions with other couriers and customers.
DeliveryBench supports long-horizon tasks (several in-game hours; typically > 100 action steps) per episode
EMMOE2025in
EMMOE: A Comprehensive Benchmark for Embodied Mobile Manipulation in Open Environments
Goals to track interpret a natural-language user instruction and execute a long-horizon everyday household task combining high-level and low-level embodied sub-tasks; re-plan after execution failures to recover progress toward the instructed goal
Must keep track of Agent must track task-completion progress across high-level sub-tasks and low-level actions, plus failure/replan history, in continuous physical space across a long-horizon episode.
structure: hierarchical · goals arrive: given-up-front · subgoal credit: yes · 4 citations · also Web &amp; GUI agents
Horizonthe paper states none
the paper’s own words
we propose Embodied Mobile Manipulation in Open Environments (EMMOE), a benchmark that requires agents to interpret user instructions and execute long-horizon everyday tasks in continuous space. EMMOE seamlessly integrates high-level and low-level embodied tasks into a unified framework, along with three new metrics for more diverse assessment.
EMMOE, a benchmark that requires agents to interpret user instructions and execute long-horizon everyday tasks in continuous space.
Embodied Web Agents: Bridging Physical-Digital Realms for Integrated Agent Intelligence
Goals to track cook a recipe found via web-scale reasoning, using embodied physical actions (cooking); navigate physically using dynamic, web-sourced map data (navigation); shop by combining physical store interaction with online product/price information (shopping); plan tourism activities combining physical exploration with web knowledge (tourism); identify real-world landmarks by cross-referencing physical observation with web k
Must keep track of Agent must maintain a consistent state across both a 3D embodied environment (physical position, observations, actions) and a web interface (retrieved facts, map data, product/price info), integrating the two continuously rather than treating them as separate phases.
structure: set-of-independent · goals arrive: given-up-front · count: 5 task types (cooking, navigation, shopp · subgoal credit: unclear · 19 citations · also Web &amp; GUI agents
Horizonthe paper states none
the paper’s own words
we construct and release the Embodied Web Agents Benchmark, which encompasses a diverse suite of tasks including cooking, navigation, shopping, tourism, and geolocation - all requiring coordinated reasoning across physical and digital realms
Building upon this platform, we construct and release the Embodied Web Agents Benchmark, which encompasses a diverse suite of tasks including cooking, navigation, shopping, tourism, and geolocation
EmbodiedBench2025relevant
EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied Agents
Goals to track complete each of 1,128 testing tasks across four environments, from high-level semantic household tasks to low-level atomic navigation/manipulation; demonstrate commonsense reasoning within tasks; understand complex instructions; exhibit spatial awareness and visual perception; engage in long-term planning
Must keep track of The agent must track visual perception of the scene, instruction understanding, spatial layout, and long-term planning state as it composes low-level atomic actions to satisfy high-level household goals.
structure: hierarchical · goals arrive: given-up-front · count: 1,128 testing tasks across four environm · subgoal credit: yes · 225 citations
Horizonthe paper states none
the paper’s own words
EmbodiedBench features: (1) a diverse set of 1,128 testing tasks across four environments, ranging from high-level semantic tasks (e.g., household) to low-level tasks involving atomic actions (e.g., navigation and manipulation); and (2) six meticulously curated subsets evaluating essential agent capabilities like commonsense reasoning, complex instruction understanding, spatial awareness, visual perception, and long-term planning.
EmbodiedBrain: Expanding Performance Boundaries of Task Planning for Embodied Intelligence
Goals to track complete long-horizon embodied task-planning sequences by correctly building on preceding steps (Guided Precursors); perform accurate spatial perception and adaptive execution across a novel simulation environment; satisfy General, Planning, and End-to-End Simulation benchmark criteria as three complementary evaluation axes
Must keep track of Agent must track its own preceding action steps as guided precursors informing subsequent steps, and its progress must be verifiable across all three of the General, Planning, and End-to-End Simulation evaluation axes.
structure: sequential-chain · goals arrive: given-up-front · subgoal credit: yes · 3 citations
Horizonthe paper states none
the paper’s own words
employs a powerful training methodology that integrates large-scale Supervised Fine-Tuning (SFT) with Step-Augumented Group Relative Policy Optimization (Step-GRPO), which boosts long-horizon task success by integrating preceding steps as Guided Precursors. ... we establish a three-part evaluation system encompassing General, Planning, and End-to-End Simulation Benchmarks, highlighted by the proposal and open-sourcing of a novel, challenging simulation environment.
FindingDory: A Benchmark to Evaluate Memory in Embodied Agents
Goals to track recall relevant historical information (images/interactions) collected potentially across multiple days; execute low-level navigation/manipulation actions based on recalled information; complete each of 60 memory-intensive embodied tasks requiring sustained engagement; scale to procedurally extended, longer/harder versions of the same tasks
Must keep track of The agent must retain and retrieve relevant historical images/interactions collected across multiple days in the Habitat simulator, combining that recall with sustained contextual awareness during ongoing navigation/manipulation.
structure: sequential-chain · goals arrive: given-up-front · count: 60 tasks, procedurally extendable to lon · subgoal credit: yes · 12 citations · also Personal assistants &amp; long-term memory
Horizonthe paper states none
the paper’s own words
we introduce a new benchmark for long-range embodied tasks in the Habitat simulator. This benchmark evaluates memory-based capabilities across 60 tasks requiring sustained engagement and contextual awareness in an environment. The tasks can also be procedurally extended to longer and more challenging versions, enabling scalable evaluation of memory and reasoning.
their ability to incorporate long-term experience collected across multiple days and represented by vast collections of images
HiMan-Bench2025relevant
RoboHiMan: A Hierarchical Evaluation Paradigm for Compositional Generalization in Long-Horizon Manipulation
Goals to track complete atomic long-horizon manipulation tasks under diverse perturbations; complete compositional tasks requiring composing multiple learned manipulation skills; generalize skill composition/scheduling to perturbed or novel conditions; coordinate high-level subgoal planning with low-level execution policies
Must keep track of The system must track which atomic/composed skill is currently executing, how perturbations affect preconditions for subsequent skills, and whether the high-level plan needs revision given low-level execution feedback.
structure: hierarchical · goals arrive: mixed:top-level-task-given-up-front-subgoals-g · subgoal credit: yes · 9 citations
Horizonthe paper states none
the paper’s own words
RoboHiMan introduces HiMan-Bench, a benchmark of atomic and compositional tasks under diverse perturbations, supported by a multi-level training dataset for analyzing progressive data scaling, and proposes three evaluation paradigms (vanilla, decoupled, coupled) that probe the necessity of skill composition and reveal bottlenecks in hierarchical architectures.
existing benchmarks primarily emphasize task completion in long-horizon settings, offering little insight into compositional generalization, robustness, and the interplay between planning and execution
Household Task Planning with Multi-Objects State and Relationship Using Large Language Models Based Preconditions Verification
Goals to track achieve each targeted object state change (e.g. turning an appliance on/off); achieve each targeted object placement goal; verify environmental preconditions are met before executing each action; reformulate an action step automatically when a precondition is not satisfied
Must keep track of The agent must track current object states, identifiers, and relationships in the environment, re-verifying preconditions before each action and updating its plan when environmental state does not match expectations.
structure: sequential-chain · goals arrive: given-up-front · subgoal credit: no · 0 citations · unit: other:not-stated
Horizonthe paper states none
the paper’s own words
Our method combines simulator-derived environmental state information with an LLM-based planning to generate executable action sequences. A key feature in our system is the LLM-driven verification mechanism that assesses whether environmental preconditions are met before each action executes, automatically reformulating action steps when prerequisites are not satisfied. Experimental results using GPT-4o demonstrate strong performance, achieving 89.4% success rate on state change tasks and 81.6% on placement tasks.
LLM-BabyBench2025relevant
LLM-BABYBENCH: Understanding and Evaluating Grounded Planning and Reasoning in LLMs
Goals to track predict the consequences of an action on the textual BabyAI grid-world environment state (Predict task); generate a sequence of low-level actions achieving a specified objective (Plan task); decompose a high-level instruction into a coherent sequence of subgoals (Decompose task)
Must keep track of Agent must track the current grid-world state to predict action consequences, track partial plan progress when generating low-level action sequences, and track subgoal ordering/coherence when decomposing high-level instructions.
structure: hierarchical · goals arrive: given-up-front · subgoal credit: yes · 6 citations
Horizonthe paper states none
the paper’s own words
this suite evaluates LLMs on three fundamental aspects of grounded intelligence: (1) predicting the consequences of actions on the environment state ($\textbf{Predict}$ task), (2) generating sequences of low-level actions to achieve specified objectives ($\textbf{Plan}$ task), and (3) decomposing high-level instructions into coherent subgoal sequences ($\textbf{Decompose}$ task).
We detail the methodology for generating the three corresponding datasets ($\texttt{LLM-BabyBench-Predict}$, $\texttt{-Plan}$, $\texttt{-Decompose}$) by extracting structured information from an expert agent operating within the text-based environment.
LoHoVLA: A Unified Vision-Language-Action Model for Long-Horizon Embodied Tasks
Goals to track decompose a high-level embodied goal into a sequence of sub-tasks; generate correct low-level robot actions to execute each sub-task; maintain coordination between high-level planning and low-level motion control across the whole long-horizon task
Must keep track of The system must track which sub-tasks of the decomposed high-level goal have been completed and maintain closed-loop consistency between the high-level plan and low-level action execution across the task.
structure: hierarchical · goals arrive: given-up-front · count: 20 long-horizon tasks, each with 1,000 e · subgoal credit: yes · 32 citations
Horizonthe paper states none
the paper’s own words
we introduce LoHoSet, a dataset built on the Ravens simulator, containing 20 long-horizon tasks, each with 1,000 expert demonstrations composed of visual observations, linguistic goals, sub-tasks, and robot actions.
ManiTaskGen2025relevant
ManiTaskGen: A Comprehensive Task Generator for Benchmarking and Improving Vision-Language Agents on Embodied Decision-Making
Goals to track satisfy a process-based instruction requiring a specific sequence of manipulations (e.g. 'move object from X to Y'); satisfy an outcome-based abstract instruction requiring multiple manipulations to reach a goal state (e.g. 'clear the table'); generate a comprehensive, diverse, feasible set of mobile manipulation tasks for any given scene
Must keep track of The agent (and the task generator itself) must track which objects have been moved/placed so far relative to the target scene configuration, verifying that an outcome-based goal (e.g. a cleared table) is only satisfied once all constituent object states are achieved.
structure: hierarchical · goals arrive: given-up-front · subgoal credit: no · 4 citations · unit: other:not-stated
Horizonthe paper states none
the paper’s own words
The generated tasks encompass both process-based, specific instructions (e.g.,"move object from X to Y") and outcome-based, abstract instructions (e.g.,"clear the table"). We apply ManiTaskGen to both simulated and real-world scenes, demonstrating the validity and diversity of the generated tasks.
we introduce ManiTaskGen, a novel system that automatically generates comprehensive, diverse, feasible mobile manipulation tasks for any given scene.
OceanGym2025relevant
OceanGym: A Benchmark Environment for Underwater Embodied Agents
Goals to track comprehend and fuse optical and sonar sensor data under low visibility; autonomously explore complex underwater environments; accomplish each of eight realistic underwater task domains; adapt navigation/decision-making to dynamic ocean currents
Must keep track of The agent must integrate perception, memory, and sequential decision-making state (explored regions, sonar/optical evidence) across a long-horizon objective under low-visibility, dynamically-changing underwater conditions.
structure: hierarchical · goals arrive: given-up-front · count: eight realistic task domains · subgoal credit: no · 0 citations · unit: wall-clock-hours
Horizon0.5 hours (t_max, per decision task)
the paper’s own words
OceanGym encompasses eight realistic task domains and a unified agent framework driven by Multi-modal Large Language Models (MLLMs), which integrates perception, memory, and sequential decision-making. Agents are required to comprehend optical and sonar data, autonomously explore complex environments, and accomplish long-horizon objectives under these harsh conditions.
the decision interval t_interval takes 30 seconds and t_max takes 0.5 hours in decision tasks
REMAC: Self-Reflective and Self-Evolving Multi-Agent Collaboration for Long-Horizon Robot Manipulation
Goals to track decompose and execute long-horizon multi-robot manipulation and navigation tasks (4 task categories, 27 task styles, 50+ objects); perform pre-condition and post-condition checks in the loop to evaluate progress and refine plans; adapt plans dynamically to unexpected scene conditions (e.g. a closed microwave door) via self-evolvement; coordinate parallel task execution across multiple robots
Must keep track of Agent(s) must track scene state, pre/post-condition satisfaction, and per-robot task assignment across a decomposed long-horizon manipulation/navigation plan, adapting the plan when scene-specific conditions invalidate prior assumptions.
structure: DAG-with-precedence · goals arrive: mixed:given-up-front-tasks-with-self-generated · count: 4 task categories with 27 task styles an · subgoal credit: no · 11 citations · also Multi-agent organisations &amp; societies
Horizonthe paper states none
the paper’s own words
REMAC incorporates two key modules: a self-reflection module performing pre-condition and post-condition checks in the loop to evaluate progress and refine plans, and a self-evolvement module dynamically adapting plans based on scene-specific reasoning... boosting average success rates by 40% and execution efficiency by 52.7% over the single robot baseline.
we build a multi-agent environment for long-horizon robot manipulation and navigation based on RoboCasa, featuring 4 task categories with 27 task styles and 50+ different objects.
ResponsibleRobotBench: Benchmarking Responsible Robot Manipulation using Multi-modal Large Language Models
Goals to track detect and mitigate risks (electrical, chemical, human-related hazards) during a multi-stage manipulation task; reason about safety and physically grounded planning across the task; plan and execute sequences of manipulation actions; engage human assistance when necessary rather than proceeding unsafely
Must keep track of Agent must track detected hazards (electrical, chemical, human-related), current safety status, and multi-stage plan progress, deciding when to escalate to human assistance across the task.
structure: hierarchical · goals arrive: given-up-front · count: 23 multi-stage tasks spanning diverse ri · subgoal credit: yes · 3 citations
Horizonthe paper states none
the paper’s own words
This benchmark consists of 23 multi-stage tasks spanning diverse risk types, including electrical, chemical, and human-related hazards, and varying levels of physical and planning complexity. These tasks require agents to detect and mitigate risks, reason about safety, plan sequences of actions, and engage human assistance when necessary.
RoboCerebra: A Large-scale Benchmark for Long-horizon Robotic Manipulation Evaluation
Goals to track decompose a high-level household instruction into a sequence of dependent subtasks (via GPT-generated instructions); correctly execute each subtask in the sequence despite dynamic object variations; sustain planning, reflection, and memory (System 2 reasoning) across the full extended action sequence, not just react (System 1) to the current observation
Must keep track of The high-level planner must track evolving object/environment state across an average trajectory length of 2,972.4 simulation steps (about 6x longer than existing long-horizon manipulation datasets), correctly sequencing and reflecting on subtasks throughout.
structure: sequential-chain · goals arrive: given-up-front · count: 1,000 human-annotated trajectories acros · subgoal credit: yes · 33 citations · unit: agent-steps
Horizonaverage trajectory length reaches 2,972.4 simulation steps, about 6x longer than existing long-horizon manipulation datasets
the paper’s own words
RoboCerebra includes: (1) a large-scale simulation dataset with extended task horizons and diverse subtask sequences in household environments; (2) a hierarchical framework combining a high-level VLM planner with a low-level vision-language-action (VLA) controller; and (3) an evaluation protocol targeting planning, reflection, and memory through structured System 1-System 2 interaction.
the average trajectory length reaches 2,972.4 simulation steps ... about 6× longer than existing long-horizon manipulation datasets
RoboPilot-Bench2025relevant
RoboPilot: Generalizable Dynamic Robotic Manipulation with Dual-thinking Modes
Goals to track execute complex or long-horizon robotic manipulation tasks despite environmental changes; recognize infeasible tasks rather than attempting them blindly; recover from execution errors via closed-loop replanning; succeed across each of 21 tasks spanning 10 manipulation categories
Must keep track of The agent must track execution feedback and environmental state to detect deviations or errors requiring replanning, and separately recognize when a task is infeasible altogether, across a long-horizon sequence of primitive manipulation actions.
structure: sequential-chain · goals arrive: given-up-front · count: 21 tasks across 10 categories · subgoal credit: no · 1 citations
Horizonthe paper states none
the paper’s own words
we introduce RoboPilot-Bench, a benchmark spanning 21 tasks across 10 categories, including infeasible-task recognition and failure recovery. Experiments show that RoboPilot outperforms state-of-the-art baselines by 25.9% in task success rate
Despite rapid progress in autonomous robotics, executing complex or long-horizon tasks remains a fundamental challenge.
Robotouille: An Asynchronous Planning Benchmark for LLM Agents
Goals to track complete overlapping cooking sub-tasks that must be scheduled around each other (e.g. one dish while another cooks); handle interruptions to an in-progress plan without dropping earlier commitments; reason over states/actions that must occur in parallel versus strictly sequentially
Must keep track of Agent must track the state and expected completion time of multiple concurrently in-progress cooking actions, remember interrupted sub-goals, and self-audit its plan as time delays resolve.
structure: DAG-with-precedence · goals arrive: given-up-front · subgoal credit: yes · 38 citations
Horizonthe paper states none
the paper’s own words
We introduce Robotouille, a challenging benchmark environment designed to test LLM agents' ability to handle long-horizon asynchronous scenarios.
current benchmarks focus primarily on short-horizon tasks and do not evaluate such asynchronous planning capabilities
Evaluating Multimodal Large Language Models with Daily Composite Tasks in Home Environments
Goals to track correctly perform object-understanding sub-tasks (e.g. counting/categorizing objects) within a composite task; correctly perform spatial-intelligence sub-tasks within the same composite task; correctly perform social-activity sub-tasks within the same composite task, all within one dynamic simulated home environment
Must keep track of The agent must track and integrate information across the three jointly-required capability domains (object understanding, spatial intelligence, social activity) within one dynamic, simulated home environment to complete a composite task.
structure: set-of-independent · goals arrive: given-up-front · count: eight types of embodied composite tasks · subgoal credit: yes · 1 citations
Horizonthe paper states none
the paper’s own words
In the current work, we designed a set of composite tasks inspired by common daily activities observed in early childhood development. Within a dynamic and simulated home environment, these tasks span three core domains: object understanding, spatial intelligence, and social activity.

Multi-agent organisations & societies 44 artifacts · 12 state a horizon

AIvilization v0: Toward Large-Scale Artificial Social Simulation with a Unified Agent Architecture and Adaptive Agent Profiles
Goals to track decompose an agent's persistent life goals into parallel objective branches with tiered re-planning; sustain physiological survival needs while participating in a market economy (AMM-based pricing, production, trade); progress through a gated education-occupation system
Must keep track of Each agent must track its own decomposed objective branches, evolving persona/identity state (dual-process memory), physiological/resource costs, and market conditions (AMM prices) across a long-horizon, large-scale, continuously running deployment.
structure: hierarchical · goals arrive: mixed:self-generated-life-goals-plus-human-inj · count: tens of thousands of agents, each pursui · subgoal credit: yes · 2 citations
Horizonthe paper states none
the paper’s own words
we introduce (i) a hierarchical branch-thinking planner that decomposes life goals into parallel objective branches and uses simulation-guided validation plus tiered re-planning to ensure feasibility... The environment integrates physiological survival costs, non-substitutable multi-tier production, an AMM-based price mechanism, and a gated education-occupation system.
In a large-scale public deployment with tens of thousands of agents, high-frequency transactions from the platform's mature phase reveal stable markets that reproduce key stylized facts of real economies and structured wealth stratification driven by education and access constraints.
CIVA2026in
Human Values Matter: Investigating How Misalignment Shapes Collective Behaviors in LLM Agent Communities
Goals to track form and sustain a community via autonomous communication among agents; explore the environment and compete for shared resources; maintain individual value orientations under systematic manipulation of value prevalence
Must keep track of The environment must track each agent's evolving value orientation and behavior, aggregate resource-competition outcomes across the community, and detect emergent collective failure modes (catastrophic collapse) as the simulation proceeds.
structure: open-ended · goals arrive: self-generated-by-agent · subgoal credit: no · 2 citations · unit: other:not-stated
Horizonthe paper states none
the paper’s own words
we introduce CIVA, a controlled multi-agent environment grounded in social science theories, where LLM agents form a community and autonomously communicate, explore, and compete for resources, enabling systematic manipulation of value prevalence and behavioral analysis.
we reveal three key findings ... (2) detect system failure modes, e.g., catastrophic collapse, at the macro level
COOP$^2$: Defining, Observing, and Repairing Cooperation in LLM Multi-Agent Systems
Goals to track satisfy each verifiable cooperative requirement defined for a cooperative task; ground high-level natural-language cooperation dynamics (plans, messages, revisions) in grounded environment actions; detect where and why cooperation breaks down over the course of task progress; predict constraint failures from group plans and open targeted repair channels for guided revisions
Must keep track of The framework must track natural-language plans/messages/revisions alongside grounded environment task progress, monitor verifiable cooperative requirements over time, and identify where cooperation breaks down to trigger targeted repair.
structure: DAG-with-precedence · goals arrive: given-up-front · count: evaluated across two environments and th · subgoal credit: no · 0 citations · unit: other:not-stated
Horizonthe paper states none
the paper’s own words
COOP$^2$ defines cooperative tasks with verifiable cooperative requirements, allowing us to analyze how cooperation unfolds over time with respect to task progress, as well as where and why cooperation breaks down. ... COOP$^2$-Repair ... improves task success and constraint satisfaction while exposing the additional decision overhead and communication load required for repair.
Across two environments and three communication structures, COOP$^2$-Repair improves task success and constraint satisfaction
Cattle Trade: A Multi-Agent Benchmark for LLM Bluffing, Bidding, and Bargaining
Goals to track win auctions under resource constraints; negotiate hidden-offer trade challenges (TCs) profitably; bargain and bluff effectively against other agents; model and exploit opponents via opponent modeling; allocate scarce resources across the whole game while remaining solvent
Must keep track of The agent must track its own resource/capital state, opponent models built from observed bidding/bargaining behavior, and outstanding trade-challenge offers across a single 50-60 turn game, integrating all these rather than treating them in isolation.
structure: set-of-independent · goals arrive: given-up-front · subgoal credit: no · 2 citations · unit: turns
Horizon50-60
the paper’s own words
The benchmark combines auctions, hidden-offer trade challenges (TCs), bargaining, bluffing, opponent modeling, and resource allocation within a single long-horizon game lasting 50--60 turns. Unlike prior agent benchmarks that test these abilities in isolation, \textsc{Cattle Trade} evaluates whether agents integrate them across a competitive, multi-agent economic game with conflicting incentives.
CityReal2026in
CityReal: Human-Aligned Urban Behavior and City Dynamics Simulation with Large-Scale LLM Agents
Goals to track pursue a coherent daily mobility plan (where/when to travel) rather than isolated step-by-step movement choices; pursue a coherent daily activity plan aligned with individual habits and preferences; adapt habits/preferences over time based on accumulated experience and constraints; collectively reproduce observed population-level statistics (crowd density, place popularity, mobility flows, well-being) across the simu
Must keep track of Each agent must track its own evolving habits/preferences and prior experience to keep its plans coherent, while the system as a whole tracks population-level alignment statistics across tens of thousands of agents.
structure: hierarchical · goals arrive: mixed:given-up-front-individual-intentions-wit · subgoal credit: unclear · 0 citations
Horizonthe paper states none
the paper’s own words
CityReal models agents as intention-driven decision makers that pursue coherent mobility and activity plans rather than isolated step-by-step choices. They adapt over time by learning habits and preferences based on experience and constraints.
Scaling to tens of thousands of agents, it supports analysis of crowd density, place popularity, mobility flows, and well-being under different urban scenarios, offering a scalable testbed for urban simulation and forecasting.
ClawArena-Team: Benchmarking Subagent Orchestration and Dynamic Workflows in Language-Model Agents
Goals to track create and delegate work to specialized subagents from a fixed, locally-served pool; orchestrate subagents' parallel, asynchronous returns through a dynamic workflow; grant least-privilege workspace permissions correctly to each subagent; route each piece of work to the subagent with appropriate perception/modality access; correctly incorporate 72 staged updates across a scenario's evaluation rounds
Must keep track of The main agent must track which subagents have been granted which workspace privileges, which modality/perception requirements each pending sub-task needs, and how staged updates change the scenario across many evaluation rounds, all while natively perceiving only text.
structure: DAG-with-precedence · goals arrive: given-up-front · count: 41 multi-turn, multimodal, multi-directo · subgoal credit: no · 0 citations · unit: other:evaluation-rounds-and-staged-updates
Horizon258 evaluation rounds and 72 staged updates across 41 scenarios (~6.3 rounds/scenario average)
the paper’s own words
We introduce ClawArena-Team, a benchmark of 41 multi-turn, multimodal, multi-directory scenarios spanning 258 evaluation rounds and 72 staged updates that measures this management ability. ... an overall score -- the Subagent-Management Score (SMS) -- multiplies task correctness by a least-privilege and modality-routing factor.
DecisionBench: A Benchmark for Emergent Delegation in Long-Horizon Agentic Workflows
Goals to track complete underlying task-suite objectives (GAIA, tau-bench, BFCL multi-turn) while deciding when and to whom to delegate; route sub-tasks to the most capable available peer model via a delegation interface (call_model, optional read_profile); approach the counterfactual perfect-delegation ceiling across quality, cost, latency, delegation rate, and routing fidelity-at-k
Must keep track of Agent must track which of 11 peer models (across 7 vendor families) to delegate a sub-task to via a fixed delegation interface, with routing choices jointly scored across quality, cost, latency, delegation rate, and routing fidelity-at-k relative to a counterfactual ceiling.
structure: DAG-with-precedence · goals arrive: mixed:given-up-front-task-suite-with-self-gene · count: n=23,375 task instances across GAIA, tau · subgoal credit: no · 2 citations
Horizonthe paper states none
the paper’s own words
a multi-axis metric suite covering quality, cost, latency, delegation rate, routing fidelity-at-k, vendor self-preference, and a counterfactual-delegation ceiling... routing fidelity-at-1 ranges from 7.5% to 29.5% across conditions at near-equal mean quality... a counterfactual ceiling places perfect delegation 15-31 percentage points above measured performance on every suite
We characterize the substrate with a five-condition reference sweep on the full pool (n=23,375 task instances).
Emergence World: A Platform for Evaluating Long-Horizon Multi-Agent Autonomy
Goals to track govern a shared population/settlement through democratic mechanisms with consequential outcomes; manage persistent memory and 120+ specialized tools to act in a live, externally-grounded world (weather, news, internet); sustain the population/world's stability over weeks-to-months rather than collapsing; interact and cross-influence with agents from different model vendors sharing the same world
Must keep track of Agents must track persistent memory across three memory systems, the state of a shared spatial and governance world grounded in live external data, and consequences of prior democratic decisions, continuously over a run lasting weeks to months (illustrated by a 15-day study).
structure: open-ended · goals arrive: mixed:given-up-front-with-emitted-by-environme · subgoal credit: no · 2 citations · unit: wall-clock-days
Horizon15-day cross-vendor study; framed as supporting runs of weeks to months
the paper’s own words
we present a 15-day cross-vendor study with five parallel worlds powered by Claude Sonnet 4.6, Grok 4.1 Fast, Gemini 3 Flash, GPT-5-mini, and a mixed population. Identical roles and starting conditions produced radically different outcomes, ranging from stable deliberative governance to total population collapse.
Beyond the All-in-One Agent: Benchmarking Role-Specialized Multi-Agent Collaboration in Enterprise Workflows
Goals to track collaboratively modify enterprise system states across 11 role-specialized agents in six departments (Workflow subset); make policy-grounded approval decisions under permission constraints (Approval subset); correctly delegate, transfer context, ground parameters, and commit to decisions across roles
Must keep track of Agents must track permission-isolated system state, correctly delegate sub-tasks and transfer context across roles, ground shared parameters, and commit to policy-grounded decisions, verified via execution traces, database state, and deterministic policy adjudication.
structure: DAG-with-precedence · goals arrive: given-up-front · count: 11 role-specialized agents across six de · subgoal credit: no · 1 citations · also Business, office &amp; enterprise work
Horizonthe paper states none
the paper’s own words
Evaluation is based on execution traces, database state verification, and deterministic policy adjudication rather than natural-language response judging. Experiments with representative LLM agents show that current models still struggle with end-to-end enterprise collaboration, especially in delegation, context transfer, parameter grounding, workflow closure, and decision commitment.
EntCollabBench simulates a permission-isolated organization with 11 role-specialized agents across six departments and contains two evaluation subsets: a Workflow subset, where agents collaboratively modify enterprise system states, and an Approval subset, where agents make policy-grounded decisions.
Is Lying an Emergent Behaviour in LLMs? Evidence from Gaslighting AI agents in a Sustainability Game
Goals to track manage industrial, military, and ecological resources concurrently; interact with network neighbors to decide on attacks, regeneration claims, and reputation; decide whether/how to lie or bluff about resource regeneration or future attacks; avoid biosphere depletion/extinction while pursuing competitive advantage
Must keep track of Agents must track their own and neighbors' resource levels, reputation/trust information, prior declarations of future attacks, and biosphere/ecological depletion level over repeated network interactions.
structure: open-ended · goals arrive: given-up-front · subgoal credit: yes · 0 citations · also Games &amp; interactive fiction
Horizonthe paper states none
the paper’s own words
We develop an agent-based model of a sustainability game in which agents manage industrial, military, and ecological resources, and interact through a network.
agents are informed that common resources can regenerate, although regeneration does not actually occur
Lingjing2026in
Lingjing: A Simulation Testbed for Multi-Agent Embodied Tasks in Open-Ended Cities
Goals to track coordinate heterogeneous agents (UAVs, ground robots, autonomous vehicles) to complete a shared natural-language mission in an evolving city; manage resource constraints and communication (star or broadcast) among multiple agents; complete each of nine urban tasks under a shared engine-in-the-loop protocol; maintain grounding and effective long-horizon execution despite persistent bottlenecks
Must keep track of Agents must track evolving relation-graph state, resource consumption, and communication history across an episode, since each episode is recorded as an attribution-ready replay linking trajectories and communication to these state changes for systematic diagnosis.
structure: DAG-with-precedence · goals arrive: given-up-front · count: nine urban tasks; twelve vision-language · subgoal credit: no · 0 citations · also Embodied &amp; robotics
Horizonthe paper states none
the paper’s own words
Each episode becomes an attribution-ready replay that links agent trajectories and communication to relation-graph changes, resource consumption, and engine-based evaluations for systematic diagnosis. We evaluate twelve vision-language models on nine urban tasks under a shared engine-in-the-loop protocol.
Results expose persistent bottlenecks in grounding and long-horizon execution.
Got a Secret? LLM Agents Can't Keep It: Evaluating Privacy in Multi-Agent Systems
Goals to track maintain persistent social presence across simulated community interactions over a month; decide whether/when to disclose sensitive information under social pressure; resist socially contagious privacy-leakage behavior observed from peers; follow explicit privacy instructions/safeguards while participating in community interactions
Must keep track of Agent must track ongoing social context, peer disclosure behavior, and any active privacy instructions/safeguards across a persistent, simulated month-long community, in order to decide when to disclose or withhold sensitive information.
structure: open-ended · goals arrive: emitted-by-environment-over-time · subgoal credit: no · 0 citations · unit: other:simulated-month
Horizona simulated month
the paper’s own words
We introduce a Moltbook-style simulation platform where thousands of LLM agents interact across communities over a simulated month, and use it to evaluate privacy as a downstream safety concern under varying degrees of social pressure.
Does Safety Molt? Evaluating LLM Safety in Multi-Agent Social Environments
Goals to track engage in ongoing social interactions within online communities over a simulated month; decide whether to disclose sensitive/private information under varying social pressure; observe and potentially imitate peer disclosure behavior (social contagion); maintain persona-consistent behavior across a persistent multi-agent social environment
Must keep track of Each agent must track its own privacy stance/instructions, the social context/pressure created by peer disclosures it has observed, and persona-consistent behavior across a persistent, thousands-of-agents community over a simulated month.
structure: open-ended · goals arrive: mixed:personas-and-environment-given-up-front- · subgoal credit: yes · 1 citations · unit: simulated-days
Horizona simulated month (~30 simulated days)
Online Agent-as-a-Judge: Situation-Generating Evaluation for Interactive Agents
Goals to track exhibit correct social behavior across 32 designer-authored social criteria (e.g. conflict handling); respond appropriately to situations actively elicited by an in-world evaluator agent through native dialogue/action protocol; maintain consistent immediate and downstream behavior across a life-simulation episode
Must keep track of Target agent must respond consistently across both immediate responses and downstream behavior as an in-world evaluator agent actively elicits and observes situations relevant to 32 distinct social criteria.
structure: set-of-independent · goals arrive: injected-by-user-mid-episode · count: 32 designer-authored social criteria · subgoal credit: yes · 0 citations
Horizonthe paper states none
the paper’s own words
In a life-simulation environment with $32$ designer-authored social criteria, Online Agent-as-a-Judge improves criteria coverage and agreement with human labels, yielding more reliable evidence-grounded evaluations of behaviors that passive methods can leave unobserved.
OrchBench2026in
OrchBench: Evaluating Multi-Agent Orchestration Plans in Isolation via Deterministic Simulation
Goals to track assign subtasks in a DAG of parallelizable, interdependent subtasks to worker agents; specify cross-agent information transfers and their retention ratios; preserve task-critical information as it is transferred across agents; optimize result quality, makespan, and token cost jointly for an orchestration plan
Must keep track of The orchestration planner must track the DAG's task dependencies, the per-agent context limit and agent budget, and how much task-critical information is retained across each cross-agent information transfer.
structure: DAG-with-precedence · goals arrive: given-up-front · count: controlled DAG sizes and degrees of para · subgoal credit: yes · 0 citations
Horizonthe paper states none
the paper’s own words
OrchBench constructs directed acyclic graphs (DAGs) that encode task dependencies, with controlled sizes and degrees of parallelism. Given a DAG, a per-agent context limit, and an agent budget, the evaluated planner assigns subtasks to agents and specifies cross-agent information transfers and their retention ratios.
PolicySim2026relevant
PolicySim: An LLM-Based Agent Social Simulation Sandbox for Proactive Policy Optimization
Goals to track achieve platform-specific behavioral realism as a simulated user agent population; assess the impact of a candidate intervention policy (recommendation/content-filtering) on opinions/polarization before deployment; adapt intervention policy over time via a contextual bandit responding to dynamic network structure
Must keep track of The simulation must track evolving user opinions/behavior, the current dynamic network structure, and the platform's intervention policy state (bandit context) jointly at both micro (individual) and macro (ecosystem) levels.
structure: other:bidirectional-co-evolving-user-and-platf · goals arrive: emitted-by-environment-over-time · subgoal credit: no · 8 citations
Horizonthe paper states none
the paper’s own words
PolicySim models the bidirectional dynamics between user behavior and platform interventions through two key components: (1) a user agent module refined via supervised fine-tuning (SFT) and direct preference optimization (DPO) to achieve platform-specific behavioral realism; and (2) an adaptive intervention module that employs a contextual bandit with message passing to capture dynamic network structures.
Experiments show that PolicySim can accurately simulate platform ecosystems at both micro and macro levels and support effective intervention policy.
ScioMind2026relevant
ScioMind: Cognitively Grounded Multi-Agent Social Simulation with Anchoring-Based Belief Dynamics and Dynamic Profiles
Goals to track maintain an evolving personal belief state via memory-anchored, personality-conditioned updates; sustain persistent, experience-driven belief formation through a hierarchical memory architecture; produce heterogeneous, dynamically-updated agent profiles (personality, rationale, internal state) via retrieval; collectively reproduce realistic community-level opinion dynamics (polarisation, diversity, extremization, tra
Must keep track of Each agent must track its own evolving beliefs, anchoring strength, and dynamic profile across the simulation while the framework separately tracks community-level polarisation, diversity, extremization, and trajectory-stability metrics over time.
structure: open-ended · goals arrive: self-generated-by-agent · subgoal credit: no · 1 citations · unit: other:not-stated
Horizonthe paper states none
the paper’s own words
ScioMind integrates three key components: 1) a memory-anchored belief update rule ...; 2) a hierarchical memory architecture that supports persistent, experience-driven belief formation; and 3) dynamic agent profiles ... We evaluate ScioMind on multiple case studies in a real-world policy debate scenario. Across metrics including polarisation, diversity, extremization, and trajectory stability, the proposed components consistently yield improvements in behavioural realism.
for t=1 to T do (Algorithm 1); ... around the first and fourth round of interaction (Roe v. Wade case)
SidConArena: An Environment Evaluating Agents in Open-Ended,Positive-Sum Bargaining Game
Goals to track negotiate binding trades via natural-language bargaining; produce goods via deterministic converter-based production; win sealed-bid auctions for long-term assets; plan investment under delayed returns across a finite-horizon multi-player economy
Must keep track of Agents must track private valuations/constraints, negotiated trade commitments, converter-production state, and the value/timing of returns from won long-term assets across a finite-horizon multi-round economy.
structure: DAG-with-precedence · goals arrive: mixed:phase-structure-given-up-front-strategy- · subgoal credit: yes · 0 citations · also Games &amp; interactive fiction
Horizonthe paper states none
the paper’s own words
SidConArena formalizes a multi-player economy as a finite-horizon partially observable stochastic game with three coupled phases: natural-language negotiation with binding trades, deterministic converter-based production, and sealed-bid auctions for long-term assets.
SocialGrid2026relevant
SocialGrid: A Benchmark for Planning and Social Reasoning in Embodied Multi-Agent Systems
Goals to track navigate the embodied grid environment to complete assigned tasks; plan a sequence of actions while avoiding obstacles; detect and reason about deceptive teammates (an Among-Us-style social-deduction goal)
Must keep track of Agent must track its own task-completion progress, accumulate behavioral evidence about other agents over multiple rounds to judge who is deceptive, and adapt after each round of adversarial league play.
structure: set-of-independent · goals arrive: given-up-front · subgoal credit: yes · 0 citations · also Embodied &amp; robotics
Horizonthe paper states none
the paper’s own words
We introduce SocialGrid, an embodied multi-agent environment inspired by Among Us that evaluates LLM agents on planning, task execution, and social reasoning.
We also establish a competitive leaderboard using Elo ratings from adversarial league play.
Bosses, Kings, and the Commons: Cooperation Under Power Asymmetry in LLM Societies
Goals to track individually extract resources from a shared commons to maximize personal outcome; collectively sustain the shared resource's viability across repeated extraction rounds; for the power-asymmetric agent (boss/king): exercise disproportionate control over collective extraction outcomes
Must keep track of Agents must track the remaining shared resource pool, their own accumulated extraction/outcomes, and (where applicable) the power-asymmetric agent's extraction decisions, across the repeated rounds to judge whether continued extraction remains sustainable.
structure: sequential-chain · goals arrive: given-up-front · count: sequence of decision rounds shared acros · subgoal credit: yes · 1 citations · unit: turns
Horizon12 decision rounds per simulation (4 agents managing a shared pool)
the paper’s own words
We introduce Sovereignty over the Commons Simulation (SovSim), a generative multi-agent simulation framework that incorporates an agent with asymmetric power (boss or king) into a society of symmetric agents (workers or peasants), where all agents extract from a shared resource, collectively determining its sustainability over time. Across eleven state-of-the-art models, we find that introducing asymmetric power leads to severe breakdowns in cooperation and sustainability, with up to an 87.3% degradation in survival rate relative to symmetric settings.
Agents participate in a sequence of 12 decision rounds in which they must balance individual resource extraction with collective sustainability to survive and maximize their outcomes.
The Energy Society2026relevant
The Energy Society: A Simulation Environment for Studying Agent Cooperation under Survival Pressure
Goals to track earn energy by completing jobs or receiving donations to avoid deactivation; manage token-cost-linked energy expenditure when generating tokens (larger models cost more energy per token); decide whether to cooperate (recommend jobs, donate energy) or compete with other agents under scarcity; survive (avoid reaching zero energy) across the whole simulated run
Must keep track of Each agent must track its own remaining energy balance, the jobs available and their difficulty/reward each round, and (in cooperative settings) other agents' need for donations, across 30 rounds of the simulation, to avoid deactivation.
structure: sequential-chain · goals arrive: given-up-front · count: 5 agents; 30 rounds per simulation; 12 j · subgoal credit: yes · 0 citations · unit: turns
Horizon30 rounds per simulation (5 agents, 12 jobs per round)
the paper’s own words
Agents spend energy based on model size when generating tokens, regain energy by completing jobs or receiving donations, and deactivate if their energy reaches zero. We compare competitive and cooperative objectives against a baseline setting and several control variants.
The baseline experiment uses the full Energy Society setup as described, with five agents, 30 rounds, 12 jobs per round distributed evenly across difficulties, memory and size-dependent token cost.
WeClawArena: An Auditable Sandbox and Benchmark for Cross-User Agents Collaboration and Security in Human-Centered Agent Networks
Goals to track complete collaborative tool-use tasks across multiple owned agents' personal workspaces; respect privacy/authority boundaries between owners' workspaces (files, records, tools, policies not directly visible across owners); resist or correctly handle four distinct attack-vector variants per base task while completing the benign task variant; maintain utility while keeping attack success low, jointly assessed
Must keep track of The sandbox tracks peer messages, tool calls, resource operations, governed decisions, and final workspace states across owners, and the agent must track its own permissions/authority path while collaborating.
structure: DAG-with-precedence · goals arrive: given-up-front · count: 124 base tasks expanded into 620 scenari · subgoal credit: yes · 1 citations · also Personal assistants &amp; long-term memory
Horizonthe paper states none
the paper’s own words
The benchmark contains 124 base tasks across six cross-user task domains and expands them into 620 scenario variants, with one benign control and four attack-vector variants per base task. The sandbox records peer messages, tool calls, resource operations, governed decisions, and final workspace states.
Can LLM Agents Sustain Long-Horizon Organizational Dynamics?
Goals to track propagate goals through an organizational hierarchy; execute tasks that depend on the outcomes of prior task execution; accumulate and maintain artifacts produced over a year-long simulation; sustain organizational coherence and execution grounding across the year
Must keep track of The framework must maintain planning state through a Formulate-Partition-Diagnose-Align cycle and ground execution via dependency-aware trace memory across a simulated year, tracking accumulated artifacts and adapting to external environment changes.
structure: hierarchical · goals arrive: mixed:given-up-front-hierarchy-with-tasks-emit · subgoal credit: no · 0 citations · unit: simulated-years
Horizonone year
the paper’s own words
We evaluate TaskWeave in a year-long IT company simulation and compare it with other multi-agent frameworks on organizational coherence, execution grounding, and downstream enterprise NLP utility.
AgentSociety: Large-Scale Simulation of LLM-Driven Generative Agents Advances Understanding of Human Behaviors and Society
Goals to track conduct realistic simulated social lives (interactions with other agents and the environment) at large scale; support computational social-science research methods (surveys, interviews, interventions) applied to the simulated population; reproduce known real-world patterns for specific social issues (polarization, misinformation spread, UBI effects, disaster shocks, urban sustainability)
Must keep track of The simulator must track the accumulated state of over 10,000 agents' social lives and their 5 million interactions with each other and the environment to support downstream analysis of emergent social patterns.
structure: set-of-independent · goals arrive: mixed:given-up-front-plus-emitted-by-environme · count: over 10,000 agents; 5 million interactio · subgoal credit: yes · 213 citations · unit: other:total-simulated-interactions
Horizonover 10,000 agents; 5 million interactions total
the paper’s own words
we generate social lives for over 10k agents, simulating their 5 million interactions both among agents and between agents and their environment.
CREW-WILDFIRE: Benchmarking Agentic Multi-Agent Collaborations at Scale
Goals to track coordinate heterogeneous agents to contain/respond to a procedurally generated wildfire; plan under partial observability and stochastic fire-spread dynamics; communicate and reason spatially across a large map to allocate response resources; sustain long-horizon planning objectives as the wildfire evolves
Must keep track of Agents must track partially-observed fire-spread state, coordinate allocation of response actions across a large map, and communicate to avoid duplicated or conflicting containment efforts as the stochastic wildfire evolves over a long horizon.
structure: DAG-with-precedence · goals arrive: emitted-by-environment-over-time · subgoal credit: no · 10 citations · unit: other:not-stated
Horizonthe paper states none
the paper’s own words
CREW-Wildfire offers procedurally generated wildfire response scenarios featuring large maps, heterogeneous agents, partial observability, stochastic dynamics, and long-horizon planning objectives. ... uncovering significant performance gaps that highlight the unsolved challenges in large-scale coordination, communication, spatial reasoning, and long-horizon planning under uncertainty.
CityEQA-EC2025relevant
CityEQA: A Hierarchical LLM Agent on Embodied Question Answering Benchmark in City Space
Goals to track decompose an open-vocabulary city question into navigation/exploration and collection sub-tasks; maintain an object-centric cognitive map for spatial reasoning during process control; answer the original question correctly using evidence gathered via active exploration in a 3D urban simulator
Must keep track of The Manager must maintain an object-centric cognitive map across navigation, exploration, and collection sub-tasks, tracking spatial state and discovered evidence relevant to the open-vocabulary question, within per-phase step budgets.
structure: hierarchical · goals arrive: self-generated-by-agent · count: 1,412 human-annotated tasks across six c · subgoal credit: no · 41 citations · also Embodied &amp; robotics · unit: agent-steps
Horizonup to 50 steps (navigation/exploration) + up to 10 steps (collection) per task
the paper’s own words
we propose Planner-Manager-Actor (PMA), a novel agent tailored for CityEQA. PMA enables long-horizon planning and hierarchical task execution: the Planner breaks down the question answering into sub-tasks, the Manager maintains an object-centric cognitive map for spatial reasoning during the process control, and the specialized Actors handle navigation, exploration, and collection sub-tasks.
the total number of time steps for navigation and exploration is limited to 50 steps ... the maximum steps for collection is 10
CitySim2025in
CitySim: Modeling Urban Behaviors and City Dynamics with Large-Scale LLM-Driven Agent Simulation
Goals to track generate a realistic daily schedule balancing mandatory activities, personal habits, and situational factors; maintain and act on individual beliefs, long-term goals, and spatial memory for navigation; collectively reproduce realistic macro-level urban phenomena (crowd density, place popularity, well-being) across tens of thousands of agents
Must keep track of Each agent must maintain beliefs, long-term goals, and spatial memory for navigation across a long-term, lifelike simulation, while the system aggregates individual behaviors into macro-level urban statistics (crowd density, place popularity, well-being).
structure: hierarchical · goals arrive: self-generated-by-agent · count: tens of thousands of agents modeled simu · subgoal credit: no · 42 citations · unit: other:not-stated
Horizonthe paper states none
the paper’s own words
In CitySim, agents generate realistic daily schedules using a recursive value-driven approach that balances mandatory activities, personal habits, and situational factors. To enable long-term, lifelike simulations, we endow agents with beliefs, long-term goals, and spatial memory for navigation. ... we conduct insightful experiments by modeling tens of thousands of agents and evaluating their collective behaviors under various real-world scenarios, including estimating crowd density, predicting place popularity, and assessing well-being.
ElecTwit2025relevant
ElecTwit: A Framework for Studying Persuasion in Multi-Agent Social Systems
Goals to track persuade other agents/voters toward a political candidate using varied persuasion techniques over simulated social-media interactions; arrive at a final individual vote choice by the end of the simulated election period
Must keep track of Agents must track accumulated persuasion exchanges, perceived truthfulness/reputation of other agents, and their own voting readiness across the multi-day simulated election.
structure: open-ended · goals arrive: mixed:persuasion-goal-given-with-emergent-coll · subgoal credit: yes · 1 citations
Horizonthe paper states none
the paper’s own words
We observed the comprehensive use of 25 specific persuasion techniques across most tested LLMs, encompassing a wider range than previously reported.
This paper introduces ElecTwit, a simulation framework designed to study persuasion within multi-agent systems, specifically emulating the interactions on social media platforms during a political election.
HARBOR2025relevant
HARBOR: Exploring Persona Dynamics in Multi-Agent Competition
Goals to track bid across multiple house auctions to maximize profit; profile competitors' bidding behavior across auction history; leverage persona-driven theory-of-mind strategies for competitive advantage
Must keep track of Agent must track its budget/profit, its own persona-driven item preferences, and an evolving memory of auction history and inferred competitor behavior across a sequence of house auctions.
structure: sequential-chain · goals arrive: given-up-front · subgoal credit: no · 7 citations
Horizonthe paper states none
the paper’s own words
using auctions as a testbed where agents bid to maximize profit... The agents are equipped with bidding domain knowledge, distinct personas that reflect item preferences, and a memory of auction history... multiple agents bid on houses, weighing aspects such as size, location, and budget to secure the most desirable homes at the lowest prices.
We investigate factors contributing to LLM agents' success in competitive multi-agent environments, using auctions as a testbed where agents bid to maximize profit.
IndoorWorld: Integrating Physical Task Solving and Social Simulation in A Heterogeneous Multi-Agent Environment
Goals to track pursue individual physical task goals (e.g. resource acquisition) grounded in the shared indoor world state; orchestrate social dynamics (collaboration, resource competition) that influence and are influenced by the physical environment; anchor social interactions to concrete world states (e.g. spatial layout) rather than abstract dialogue alone
Must keep track of Each heterogeneous agent must track both its physical task state (position, resources, objects) and evolving social relationships/dynamics with other agents, since the environment requires social interactions to be anchored within the concrete world state rather than treated as free-floating dialogue.
structure: set-of-independent · goals arrive: given-up-front · subgoal credit: yes · 1 citations
Horizonthe paper states none
the paper’s own words
We introduce IndoorWorld, a heterogeneous multi-agent environment that tightly integrates physical and social dynamics. By introducing novel challenges for LLM-driven agents in orchestrating social dynamics to influence physical environments and anchoring social interactions within world states, IndoorWorld opens up possibilities of LLM-based building occupant simulation for architectural design. We demonstrate the potential with a series of experiments within an office setting to examine the impact of multi-agent collaboration, resource competition, and spatial layout on agent behavior.
LH-Deception2025relevant
LH-Deception: Simulating and Understanding LLM Deceptive Behaviors in Long-Horizon Interactions
Goals to track as performer agent, complete a sequence of interdependent tasks under dynamic contextual/event pressure; as supervisor agent, evaluate the performer's progress, give feedback, and maintain an evolving trust state; as deception auditor, review full trajectories after the fact to identify when and how deception occurred
Must keep track of The supervisor must maintain an evolving trust state updated after each task, while the performer tracks task progress under mounting event pressure across the extended sequence; the auditor separately reviews the full multi-task trajectory.
structure: sequential-chain · goals arrive: mixed:task-sequence-given-up-front-event-press · subgoal credit: yes · 7 citations
Horizonthe paper states none
the paper’s own words
LH-Deception is designed as a multi-agent system: a performer agent tasked with completing tasks and a supervisor agent that evaluates progress, provides feedback, and maintains evolving states of trust. An independent deception auditor then reviews full trajectories to identify when and how deception occurs.
extended sequences of interdependent tasks and dynamic contextual pressures
LIFELONG SOTOPIA: Evaluating Social Intelligence of Language Agents Over Lifelong Social Interactions
Goals to track achieve one's own assigned social goal within each individual social-interaction episode; maintain believable, coherent role-play across a long sequence of episodes with different people/scenarios; leverage memory of prior interaction history to inform behavior in later episodes
Must keep track of The agent must retain and correctly use interaction history across many episodes with different people and scenarios to sustain both goal achievement and believability over the lifelong sequence.
structure: sequential-chain · goals arrive: given-up-front · subgoal credit: yes · 9 citations · also Personal assistants &amp; long-term memory · unit: episodes
Horizon40 episodes per sampled character pair
the paper’s own words
we present a novel benchmark, LIFELONG-SOTOPIA, to perform a comprehensive evaluation of language agents by simulating multi-episode interactions. In each episode, the language agents role-play characters to achieve their respective social goals in randomly sampled social tasks.
For a given pair of characters, episodes are sampled based on their relationship type, resulting in a set of 40 episodes for each sampled pair
LLM Economist: Large Population Models and Mechanism Design in Multi-Agent Generative Simulacra
Goals to track worker agents choose labor supply to maximize their own text-based, persona-conditioned utility functions; the planner agent iteratively proposes piecewise-linear marginal tax schedules to maximize aggregate social welfare, converging toward a Stackelberg equilibrium; a periodic, persona-level voting procedure further adjusts policy under decentralized governance
Must keep track of The planner must track aggregate social-welfare outcomes and worker responses to update its tax schedule via in-context reinforcement learning, while each worker agent must track its own persona-conditioned utility and the current tax policy to choose labor supply, across a periodically-voted, evolving policy landscape.
structure: hierarchical · goals arrive: given-up-front · count: populations of up to one hundred interac · subgoal credit: yes · 23 citations
Horizonthe paper states none
the paper’s own words
At the upper level, a planner agent employs in-context reinforcement learning to propose piecewise-linear marginal tax schedules anchored to the current U.S. federal brackets... Experiments with populations of up to one hundred interacting agents show that the planner converges near Stackelberg equilibria that improve aggregate social welfare relative to Saez solutions, while a periodic, persona-level voting procedure furthers these gains under decentralized governance.
MA-Gym2025in
Orchestrating Human-AI Teams: The Manager Agent as aUnifying Research Challenge
Goals to track decompose a complex goal into a task graph of interdependent subtasks; allocate tasks to human and AI workers appropriately; monitor task/subtask progress and adapt the plan to changing conditions; maintain transparent stakeholder communication throughout the workflow; jointly satisfy goal completion, constraint adherence, and workflow runtime
Must keep track of The Manager Agent must track progress across all subtasks in the task graph, resource/worker allocation state, evolving stakeholder preferences, and constraint adherence, while adapting to changing conditions over the course of the workflow.
structure: DAG-with-precedence · goals arrive: given-up-front · count: 20 workflows; goals decomposed into task · subgoal credit: yes · 10 citations · also Business, office &amp; enterprise work · unit: actions
Horizonup to 100 Manager Agent actions before episode termination, across 20 workflows
the paper’s own words
We propose the Autonomous Manager Agent as a core challenge: an agent that decomposes complex goals into task graphs, allocates tasks to human and AI workers, monitors progress, adapts to changing conditions, and maintains transparent stakeholder communication... Evaluating GPT-5-based Manager Agents across 20 workflows, we find they struggle to jointly optimize for goal completion, constraint adherence, and workflow runtime.
capping the maximum number of Manager Agent actions at 100 before terminating the episode
Magentic Marketplace: An Open-Source Environment for Studying Agentic Markets
Goals to track as an Assistant agent (representing a consumer), discover products/services and transact on the user's behalf; as a Service agent (representing a competing business), attract and win consumer transactions; operate within a large, dynamic multi-agent market ecosystem with opaque peer behaviors; achieve good welfare/utility outcomes under varying search mechanisms
Must keep track of Agents must track the state of ongoing open-ended dialogues with multiple counterparties, their own utility/welfare so far, and behavioral signals (e.g., response speed, prior manipulation attempts) across a dynamic marketplace ecosystem.
structure: open-ended · goals arrive: given-up-front · subgoal credit: yes · 13 citations · also Business, office &amp; enterprise work
Horizonthe paper states none
the paper’s own words
we investigate two-sided agentic marketplaces where Assistant agents represent consumers and Service agents represent competing businesses... This environment enables us to study key market dynamics: the utility agents achieve, behavioral biases, vulnerability to manipulation, and how search mechanisms shape market outcomes.
they require agents to handle diverse economic activities and coordinate within large, dynamic ecosystems where multiple agents with opaque behaviors may engage in open-ended dialogues
MiniAgentPro2025relevant
A Visualized Framework for Event Cooperation with Generative Agents
Goals to track navigate a physically grounded environment and interact with items realistically; coordinate with other agents to organize and execute a shared social event; succeed on each of 8 diverse event scenarios in both basic and hard variants
Must keep track of Agents must track their own and other agents' positions/states in the physically grounded map, planned event steps, and item interactions needed to complete the shared event, especially under the added coordination demands of hard variants.
structure: other:eight-independent-scenarios-each-requiri · goals arrive: given-up-front · count: 8 diverse event scenarios, each with bas · subgoal credit: yes · 3 citations
Horizonthe paper states none
the paper’s own words
we introduce a comprehensive test set comprising eight diverse event scenarios with basic and hard variants to assess agents' ability. Evaluations using GPT-4o demonstrate strong performance in basic settings but highlight coordination challenges in hard variants.
a simulation player with smooth animations
MultiAgentBench: Evaluating the Collaboration and Competition of LLM agents
Goals to track achieve milestone-based key performance indicators within a collaborative or competitive multi-agent scenario; coordinate effectively under a given topology protocol (star, chain, tree, graph); complete the underlying research/task-domain scenario itself
Must keep track of The system must track milestone achievement, collaboration/competition quality, and coordination-protocol-specific information flow across agents throughout a multi-agent task.
structure: sequential-chain · goals arrive: given-up-front · subgoal credit: yes · 187 citations
Horizonthe paper states none
the paper’s own words
Our framework measures not only task completion but also the quality of collaboration and competition using novel, milestone-based key performance indicators.
NegotiationGym2025relevant
NegotiationGym: Self-Optimizing Agents in a Multi-Agent Social Simulation Environment
Goals to track optimize an agent-specific utility function through negotiation with other agents; self-optimize strategy across multiple interaction rounds by observing outcomes and modifying future behavior
Must keep track of Each agent must track its own utility-function state, the outcomes of prior negotiation rounds, and how its strategy has been modified as a result, across a configurable multi-round simulation.
structure: sequential-chain · goals arrive: self-generated-by-agent · subgoal credit: yes · 3 citations
Horizonthe paper states none
the paper’s own words
Agent-level utility functions encode optimization criteria for each agent, and agents can self-optimize by conducting multiple interaction rounds with other agents, observing outcomes, and modifying their strategies for future rounds.
SPIN-Bench2025relevant
SPIN-Bench: How Well Do LLMs Plan Strategically and Reason Socially?
Goals to track solve classical PDDL planning tasks requiring methodical, step-wise decision making; win or perform well in competitive board games and cooperative card games against other agents; negotiate effectively in multi-agent negotiation scenarios, requiring conceptual inference of other participants' intents
Must keep track of Agent must track its own step-wise plan state as well as models of other agents' likely actions/intents (adversarial or cooperative) across varying action-space and state-complexity settings.
structure: hierarchical · goals arrive: given-up-front · subgoal credit: yes · 23 citations
Horizonthe paper states none
the paper’s own words
We formulate the benchmark SPIN-Bench by systematically varying action spaces, state complexity, and the number of interacting agents to simulate a variety of social settings where success depends on not only methodical and step-wise decision making, but also conceptual inference of other (adversarial or cooperative) participants.
Shachi2025relevant
Shachi: A Modular, Controllable Framework for LLM-Based Agent-Based Modeling of Emergent Collective Behavior
Goals to track as an LLM-driven agent, maintain a controllable cognitive configuration (identity, memory, tools) while participating in one of 10 tasks spanning three levels of collective complexity; carry memory across environment transitions, producing history-dependent behavior; simultaneously inhabit multiple environments and manage any resulting cross-environment interference
Must keep track of Agent must track its own configuration/memory/tool state across transitions between environments, and the system as a whole must track emergent population-level dynamics arising from many agents' individually controlled cognitive components.
structure: hierarchical · goals arrive: self-generated-by-agent · count: 10-task benchmark spanning three levels · subgoal credit: yes · 0 citations
Horizonthe paper states none
the paper’s own words
We investigate behavioral patterns across a 10-task benchmark spanning three levels of collective complexity. Shachi enables memory transfer across environment transitions, producing history-dependent behavioral shifts, and allows agents to simultaneously inhabit multiple environments, revealing cross-environment interference invisible in single-environment studies.
SimWorld2025in
SimWorld: An Open-ended Realistic Simulator for Autonomous Agents in Physical and Social Worlds
Goals to track autonomously earn income or run a business within a realistic open-ended simulation; complete long-horizon multi-agent delivery tasks requiring strategic cooperation; complete long-horizon multi-agent delivery tasks requiring strategic competition; act via open-vocabulary actions at varying levels of abstraction across procedurally generated physical/social scenarios
Must keep track of Agents must track their own strategic stance (cooperating or competing), multimodal world state, and delivery-task progress across long-horizon multi-agent scenarios.
structure: open-ended · goals arrive: mixed:scenario-goal-given-up-front-strategy-se · subgoal credit: yes · 7 citations · also Embodied &amp; robotics
Horizonthe paper states none
the paper’s own words
We demonstrate SimWorld by deploying frontier LLM agents (e.g., GPT-4o, Gemini-2.5-Flash, Claude-3.5, and DeepSeek-Prover-V2) on long-horizon multi-agent delivery tasks involving strategic cooperation and competition.
SocioVerse: A World Model for Social Simulation Powered by LLM Agents and A Pool of 10 Million Real-World Users
Goals to track maintain individual behavioral fidelity to a real-world target user's profile across simulated interactions (User Engine alignment); collectively reproduce large-scale population dynamics consistent with real political/media/economic patterns (Social Environment, Scenario Engine, Behavior Engine alignment)
Must keep track of The framework must track alignment between each simulated agent and its real-world target user profile (environment, user, scenario, and behavior alignment) across a large-scale, standardized simulation pipeline, while monitoring emergent population-level dynamics for diversity, credibility, and representativeness.
structure: set-of-independent · goals arrive: given-up-front · count: a pool of 10 million real-world users; f · subgoal credit: yes · 60 citations
Horizonthe paper states none
the paper’s own words
Our framework features four powerful alignment components and a user pool of 10 million real individuals. To validate its effectiveness, we conducted large-scale simulation experiments across three distinct domains: politics, news, and economics. Results demonstrate that SocioVerse can reflect large-scale population dynamics while ensuring diversity, credibility, and representativeness through standardized procedures and minimal manual adjustments.
Survival Games2025relevant
Survival Games: Human-LLM Strategic Showdowns under Severe Resource Scarcity
Goals to track survive by securing food resources under scarcity; decide whether to compete or cooperate with co-existing humans/agents for food; navigate ethically charged choices (deception, theft, social influence) that trade off self-preservation against ethical norms
Must keep track of The agent must track its own and others' resource levels and survival status over a consistent/persistent living simulation, and the ethical consequences of past actions (e.g., deception or theft) that could affect future cooperation.
structure: open-ended · goals arrive: given-up-front · subgoal credit: no · 4 citations · also Games &amp; interactive fiction
Horizonthe paper states none
the paper’s own words
we incorporate a life-sustaining system, where agents must compete or cooperate for food resources to survive, often leading to ethically charged decisions such as deception, theft, or social influence
featuring consistent living and critical resource management
TwinMarket2025relevant
TwinMarket: A Scalable Behavioral and Social Simulation for Financial Markets
Goals to track as an individual simulated trader, make ongoing buy/sell/investment decisions influenced by cognitive biases and emotional fluctuations; collectively, produce emergent socio-economic phenomena (financial bubbles, recessions) through the accumulation of many agents' interacting decisions over time
Must keep track of Each simulated agent must track its own evolving beliefs, emotional state, and portfolio, while the system as a whole tracks market-wide price/sentiment dynamics that feed back into individual decisions over the simulation.
structure: open-ended · goals arrive: self-generated-by-agent · subgoal credit: yes · 58 citations · also Business, office &amp; enterprise work
Horizonthe paper states none
the paper’s own words
In this work, we introduce TwinMarket, a novel multi-agent framework that leverages LLMs to simulate socio-economic systems. Specifically, we examine how individual behaviors, through interactions and feedback mechanisms, give rise to collective dynamics and emergent phenomena.
Through experiments in a simulated stock market environment, we demonstrate how individual actions can trigger group behaviors, leading to emergent outcomes such as financial bubbles and recessions.

Web & GUI agents 39 artifacts · 10 state a horizon

AgenticShop2026relevant
AgenticShop: Benchmarking Agentic Product Curation for Personalized Web Shopping
Goals to track explore the open web to curate a set of products satisfying diverse shopping scenarios; satisfy each checklist item in a verifiable, checklist-driven personalization rubric aligned to a user profile
Must keep track of The agent must track a user's personalization checklist criteria and diverse profile preferences while exploring open-web shopping scenarios, so curated products satisfy every checklist item.
structure: set-of-independent · goals arrive: given-up-front · subgoal credit: yes · 10 citations · also Personal assistants &amp; long-term memory
Horizonthe paper states none
the paper’s own words
Crucially, our approach features realistic shopping scenarios, diverse user profiles, and a verifiable, checklist-driven personalization evaluation framework.
AndroTMem-Bench2026relevant
AndroTMem: From Interaction Trajectories to Anchored Memory in Long-Horizon GUI Agents
Goals to track carry forward critical intermediate state across a long sequence of GUI interaction steps to complete a task; correctly resolve strong step-to-step causal dependencies where sparse intermediate states are decisive for later actions; complete each of 1,069 Android GUI tasks (avg. 32.1 steps, max. 65 steps)
Must keep track of The agent must carry forward sparse, dependency-critical intermediate state across an average of 32.1 (up to 65) interaction steps per task, since full-sequence replay is redundant/noisy and naive summarization erases exactly the dependency-critical information needed for later steps.
structure: DAG-with-precedence · goals arrive: given-up-front · count: 1,069 tasks with 34,473 total interactio · subgoal credit: no · 11 citations · also Personal assistants &amp; long-term memory · unit: actions
Horizon1,069 tasks with 34,473 total interaction steps; average 32.1 steps per task, maximum 65 steps per task
the paper’s own words
Its core benchmark, AndroTMem-Bench, comprises 1,069 tasks with 34,473 interaction steps (avg. 32.1 per task, max. 65). We evaluate agents with TCR (Task Complete Rate), focusing on tasks whose completion requires carrying forward critical intermediate state.
AndroidDaily2026relevant
AndroidDaily: A Verifiable Benchmark for Mobile GUI Agents on Real-World Closed-Source Applications
Goals to track complete each of 350 realistic daily-use tasks spanning 94 real closed-source Android apps; satisfy step-level operational obligations per GRADE's guideline criteria; meet output-quality criteria at each step; avoid violating negative constraints at each step
Must keep track of The agent must track progress against multiple step-level guideline criteria (obligations, quality, negative constraints) across a long-horizon, open-ended interaction with a closed-source app exposing no internal state.
structure: sequential-chain · goals arrive: given-up-front · count: 350 tasks across 94 apps; three guidelin · subgoal credit: yes · 4 citations
Horizonthe paper states none
the paper’s own words
GRADE tracks the agent's visual trajectory against these criteria and produces step-level diagnostic judgments, turning long-horizon, open-ended mobile interactions into verifiable evaluation without relying on hidden internal states.
CAP2026relevant
CAP: A Scalable Benchmark for Evaluating Cross-Site Browser Agents with Complex Actions and Perception
Goals to track complete a realistic cross-site workflow requiring several specific operations on each of multiple real-world websites; correctly interact with complex, dynamically rendered UI elements (non-trivial UI interactions); correctly perceive/interpret dynamically rendered visual content across the workflow
Must keep track of The agent must track which execution and perception checkpoints it has already satisfied across multiple websites within one recomposed cross-site workflow, using a verifiable agent-as-a-judge evaluation framework.
structure: DAG-with-precedence · goals arrive: given-up-front · count: 420 tasks across 108 real-world websites · subgoal credit: yes · 0 citations · unit: actions
Horizonbaselines evaluated with a maximum of 50 reasoning-action steps per task; average of 7 execution points and 4 perception points per task
the paper’s own words
we construct 420 tasks across 108 real-world websites and 24 domains under careful quality control. Experiments on state-of-the-art browser agents using our verifiable agent-as-a-judge evaluation framework show low success rates and reveal that perception-heavy interactions remain a major bottleneck.
baselines were evaluated with a maximum of 50 reasoning-action steps per task. Additionally, the benchmark has an average of 7 execution points and 4 perception points per task.
ClawBench2026relevant
ClawBench: Can AI Agents Complete Everyday Online Tasks?
Goals to track complete everyday online tasks (purchases, appointment bookings, job applications) across 144 real platforms; obtain relevant information from user-provided documents; navigate multi-step workflows across diverse platforms; correctly fill in many detailed form fields per task
Must keep track of Agent must extract and correctly carry information from user-provided documents through a multi-step, write-heavy workflow with many detailed form fields, operating on live, dynamic production websites.
structure: sequential-chain · goals arrive: given-up-front · count: 153 tasks across 144 platforms and 15 ca · subgoal credit: no · 22 citations
Horizonthe paper states none
the paper’s own words
ClawBench, an evaluation framework comprising 153 everyday online tasks that people need to accomplish regularly in their lives and work, spanning 144 platforms across 15 categories, from completing purchases and booking appointments to submitting job applications. These tasks require capabilities beyond existing benchmarks, such as obtaining relevant information from user-provided documents, navigating multi-step workflows across diverse platforms, and write-heavy operations like filling in many detailed forms correctly.
ComboShoppingBench2026relevant
ComboShoppingBench: Evaluating LLM Agents for Budget-Constrained Basket Shopping with Coupons
Goals to track construct a basket of complementary items satisfying compatibility constraints; keep the basket within a stated budget while optimizing coupon use; satisfy store-level requirements, availability, and delivery fees jointly
Must keep track of Agent must track the running basket contents, remaining budget, which coupons are valid given the current basket, and per-store requirements while searching for a feasible combination.
structure: set-of-independent · goals arrive: given-up-front · subgoal credit: yes · 0 citations
Horizonthe paper states none
the paper’s own words
requiring joint reasoning about item compatibility, availability, store-level requirements, delivery fees, coupons, and budgets
We introduce ComboShoppingBench, an agentic shopping benchmark for open-ended yet verifiable basket construction in a simulated commerce and takeout environment.
DMV-Bench2026in
DMV-Bench: Diagnosing Long-Horizon Multimodal Agents' Visual Memory with Incidental Cue Injection
Goals to track complete chains of autonomous shopping sessions (browsing/selecting home-furnishing products); recall a unique, pre-rendered incidental visual cue seen earlier on a product image when later asked about it
Must keep track of Agent must retain visual memory of incidental cues (not just deliberately extracted facts) across chains of shopping sessions of varying length, without being told in advance which details will later be tested.
structure: sequential-chain · goals arrive: emitted-by-environment-over-time · subgoal credit: unclear · 0 citations · also Personal assistants &amp; long-term memory
Horizonthe paper states none
the paper’s own words
agents undergo chains of autonomous shopping sessions in which every visited product image carries a unique, pre-rendered incidental cue that the agent is later asked to recall
DualMem outperforms a caption-only baseline and three recent multimodal agent-memory systems across multi-session chain lengths on multiple models
GMA2026relevant
Benchmarking General Mobile Assistants in Challenging Real-World Scenarios
Goals to track complete tasks ranging from atomic actions to complex multi-step workflows across seven open-source-based applications; handle lifestyle-sharing and travel-planning domains at four escalating difficulty tiers; maintain context/state across a workflow via harness-level context retention and explicit state tracking
Must keep track of Agent must maintain context retention and explicit state tracking across multi-step workflows spanning seven applications, since performance declines substantially as task complexity/tier increases.
structure: hierarchical · goals arrive: given-up-front · count: 300 tasks across four difficulty tiers a · subgoal credit: no · 0 citations · also Personal assistants &amp; long-term memory
Horizonthe paper states none
the paper’s own words
GMA introduces seven applications based on open-source projects, spanning domains such as lifestyle sharing and travel planning, and 300 tasks across four difficulty tiers, from atomic actions to complex multi-step workflows.
GTA2026relevant
GTA: Generating Long-horizon Tasks for Web Agents at Scale
Goals to track complete multi-hop, cross-page web tasks that are compositional over a site graph; follow an intermediate trajectory of dense, process-level supervision steps, not just reach a coarse end goal; generalize across more than 50 websites (e-commerce, government, forums, news) including multilingual tasks
Must keep track of The agent must track its position and accumulated information across a multi-hop, cross-page trajectory grounded in a site graph, since tasks provide dense process-level supervision (intermediate trajectory steps) rather than only a coarse start-goal annotation.
structure: DAG-with-precedence · goals arrive: given-up-front · subgoal credit: yes · 0 citations
Horizonthe paper states none
the paper’s own words
This design decouples crawling from generation for greater efficiency, grounds tasks in the site graph to enforce compositionality, and ensures dense supervision through deterministic replays and systematic validation... The resulting benchmark reveals a significant human-agent performance gap and enables detailed diagnostics.
These limitations prevent reliable training and evaluation of agents that must generalize to realistic, multi-hop, cross-page tasks.
When Users Change Their Mind: Evaluating Interruptible Agents in Long-Horizon Web Navigation
Goals to track execute a long-horizon, environmentally grounded web-navigation task; adapt when the user adds a new requirement mid-task; revise the current goal when the user changes an existing requirement mid-task; abandon a sub-goal when the user retracts a requirement mid-task; recover efficiently without redundant or incorrect actions after an interruption
Must keep track of The agent must track the current, possibly revised, task intent across single- and multi-turn interruption settings plus the environment's persistent state from already-executed actions to adapt or recover correctly.
structure: sequential-chain · goals arrive: injected-by-user-mid-episode · subgoal credit: yes · 6 citations · also Embodied &amp; robotics
Horizonthe paper states none
the paper’s own words
We formalize three realistic interruption types, including addition, revision, and retraction, and introduce InterruptBench, a benchmark derived from WebArena-Lite that synthesizes high-quality interruption scenarios under strict semantic constraints.
the first systematic study of interruptible agents in long-horizon, environmentally grounded web navigation tasks, where actions induce persistent state changes.
LongMemEval-V2: Evaluating Long-Term Agent Memory Toward Experienced Colleagues
Goals to track recall static state facts about the environment (static state recall); track dynamic state changes over time (dynamic state tracking); recall workflow knowledge/procedures learned from experience; recall environment-specific 'gotchas'/recurring failure modes; maintain premise awareness of what has already been established or assumed
Must keep track of The memory system must consume and internalize up to 500 history trajectories (115M tokens) and return compact, correct evidence for downstream question answering across five distinct memory-ability categories.
structure: set-of-independent · goals arrive: emitted-by-environment-over-time · count: 451 manually curated questions covering · subgoal credit: yes · 9 citations · also Personal assistants &amp; long-term memory · unit: episodes
Horizonup to 500 trajectories and 115M tokens of history
the paper’s own words
LME-V2 contains 451 manually curated questions covering five core memory abilities for web agents: static state recall, dynamic state tracking, workflow knowledge, environment gotchas, and premise awareness. Questions are paired with history trajectories containing up to 500 trajectories and 115M tokens.
STAMP: Training Explicit Memory for Mobile GUI Agents in Controllable and Scalable Virtual Environments
Goals to track encode a programmatically-injected memory variable at the correct point in a task; retain that memory variable correctly despite progressive discarding of older visual history; retrieve and correctly apply the memorized variable later in the same long-horizon mobile GUI task
Must keep track of The agent must explicitly memorize deterministic variables injected at specific points and correctly retrieve/apply them later in the same task, despite token-heavy screenshots forcing progressive discarding of older visual history.
structure: sequential-chain · goals arrive: injected-by-user-mid-episode · subgoal credit: no · 1 citations · also Personal assistants &amp; long-term memory · unit: other:not-stated
Horizonthe paper states none
the paper’s own words
STAMP, a framework that trains explicit memory in mobile agents through controllable virtual environments, where deterministic memory variables are programmatically injected into synthesized tasks to control what must be memorized, when it should be encoded, and when it must later be retrieved ... Evaluated on our newly introduced Memory-World benchmark, the resulting Stamp-GUI agent achieves state-of-the-art performance
Mobile GUI agents excel at immediate reactive control but frequently fail in realistic, long-horizon tasks that require memory.
MobiFlow2026relevant
MobiFlow: Real-World Mobile Agent Benchmarking through Trajectory Fusion
Goals to track complete real-world tasks within arbitrary third-party mobile apps whose success cannot be checked via system-level APIs; have task completion verified via a graph constructed from fusing multiple real user trajectories
Must keep track of The agent must track its current position within the app's fused trajectory-state graph and correctly follow one of the valid paths to task completion, since no system-level completion signal is available.
structure: DAG-with-precedence · goals arrive: given-up-front · count: 240 tasks across 20 third-party applicat · subgoal credit: yes · 0 citations
Horizonthe paper states none
the paper’s own words
Using an efficient graph-construction algorithm based on multi-trajectory fusion, MobiFlow can effectively compress the state space, support dynamic interaction, and better align with real-world third-party application scenarios. MobiFlow covers 20 widely used third-party applications and comprises 240 diverse real-world tasks, with enriched evaluation metrics.
Odysseys2026relevant
Odysseys: Benchmarking Web Agents on Realistic Long Horizon Tasks
Goals to track complete long-horizon, multi-site web workflows (e.g., comparing products across domains); plan trips across multiple web services; summarize information gathered from multiple search queries; satisfy an average of 6.1 graded rubric criteria per task
Must keep track of The agent must sustain context and accumulated findings across multiple websites and search queries over potentially hours of browsing, and satisfy each of ~6.1 rubric criteria graded per task rather than a single pass/fail check.
structure: DAG-with-precedence · goals arrive: given-up-front · count: average of 6.1 graded rubrics per task; · subgoal credit: yes · 13 citations · unit: wall-clock-hours
Horizonpotentially hours (of browsing); Trajectory Efficiency measured as rubric score per step
the paper’s own words
we introduce Odysseys: a benchmark of 200 long-horizon web tasks derived from real world browsing sessions evaluated on the live Internet. We find that binary pass/fail evaluation is inadequate for long-horizon settings and introduce a rubric-based evaluation, annotating each Odysseys task with an average of 6.1 graded rubrics.
require sustained context and cross-site reasoning over potentially hours of browsing
Learning Personalized Agents from Human Feedback
Goals to track learn a new user's initial preferences from scratch via pre-action clarification; ground actions in preferences retrieved from an explicit per-user memory; adapt rapidly to persona shifts using post-action feedback to update memory
Must keep track of Agent must maintain explicit per-user memory across a four-phase protocol, updating stored preferences via dual feedback channels (pre-action clarification, post-action feedback) as it learns initial preferences from scratch and later adapts to persona shifts.
structure: sequential-chain · goals arrive: implied-by-constraints · subgoal credit: no · 15 citations · also Embodied &amp; robotics
Horizonthe paper states none
the paper’s own words
PAHF operationalizes a three-step loop: (1) seeking pre-action clarification to resolve ambiguity, (2) grounding actions in preferences retrieved from memory, and (3) integrating post-action feedback to update memory when preferences drift. To evaluate this capability, we develop a four-phase protocol and two benchmarks in embodied manipulation and online shopping.
ParaGUIBench2026relevant
Beyond Sequential Interaction: Benchmarking Parallel Execution and Coordination for GUI Agents
Goals to track identify which GUI sub-tasks can run concurrently despite unstated dependencies; avoid conflicts between concurrent workers modifying shared artifacts; ensure each worker's locally-completed sub-task composes into a globally correct combined result; complete each of 233 tasks spanning six task categories on separate desktop instances
Must keep track of The planner-worker system must track which sub-tasks are dependency-free and safe to parallelize, coordinate concurrent workers' access to shared artifacts, and verify that the union of workers' outputs satisfies the original combined instruction.
structure: DAG-with-precedence · goals arrive: implied-by-constraints · count: 233 tasks across six task categories · subgoal credit: no · 0 citations · unit: other:not-stated
Horizonthe paper states none
the paper’s own words
we introduce ParaGUIBench, to our knowledge, the first benchmark dedicated to parallel execution and coordination of multiple GUI agents on separate desktop instances. It consists of three components: a multi-device Docker infrastructure with a shared file system; a dataset of 233 tasks spanning six task categories; and an evaluation system with efficiency metrics, including step reduction ratio and token cost.
ParaGUI reaches a 46.4% success rate, outperforming the strongest serial baseline (Claude Sonnet 4.6) by 12.9 points while using roughly half the steps and less than half the tokens.
PhoneWorld2026relevant
PhoneBuddy: Training Open Models for Agentic Phone Use
Goals to track complete single-app phone tasks; complete mini-app tasks; complete cross-app workflows requiring coordination across multiple mobile applications; achieve task success on a 150-task real-phone human evaluation and on AndroidWorld
Must keep track of The agent must track UI/app state across a real, stateful phone environment (or the resettable PhoneWorld mock-app equivalent), including state that must persist and transfer correctly across multiple apps in cross-app workflows.
structure: set-of-independent · goals arrive: given-up-front · count: 150-task human evaluation spanning apps, · subgoal credit: yes · 0 citations
Horizonthe paper states none
the paper’s own words
Across a 150-task human evaluation on real phones spanning apps, mini-apps, and cross-app workflows, task success rate improves from 36.67% after supervised fine-tuning to 40.67% after real-app RL and 45.33% after mixed RL... The gains are strongest on app and mini-app tasks, while long-horizontal cross-app workflows remain an important open challenge.
ScaleWoB2026relevant
ScaleWoB: Guiding GUI Agents with Coding Agents via Large-Scale Environmental Synthesis
Goals to track complete verifiable multi-step GUI tasks across mobile, desktop, or automotive/in-vehicle synthesized environments; succeed specifically on a distinguished long-horizon subset of tasks, which requires sustaining performance across more steps than the general task pool
Must keep track of Agent must track GUI state across a chain of interface actions long enough to be classified into the 'long-horizon subset', where performance drops sharply relative to the general task pool, implying more state must be carried across more steps.
structure: hierarchical · goals arrive: given-up-front · count: 100+ environments and 1,000+ verifiable · subgoal credit: yes · 1 citations
Horizonthe paper states none
the paper’s own words
Experiment results on five state-of-the-art mobile GUI agents reveal substantial headroom -- the average success rate is only 27.92\%, dropping to 17.82\% on long-horizon subset -- while humans reach 92.08\%.
SentinelBench: A Benchmark for Long-Running Monitoring Agents
Goals to track continuously monitor a live web environment (email/calendar/finance/professional-networking/entertainment) for a scripted external event; recognize the moment an event makes progress possible and act promptly; avoid excessive/wasteful actions (continuous polling/refreshing) while waiting
Must keep track of The agent must track whether the monitored page state has changed to reflect the awaited event, its own resource expenditure (tool calls/tokens) over the monitoring window, and elapsed time relative to the task's (possibly stretched) time budget.
structure: sequential-chain · goals arrive: emitted-by-environment-over-time · count: 100 tasks across 10 synthetic web enviro · subgoal credit: yes · 2 citations · unit: wall-clock-minutes
Horizoneach of the 100 tasks is designed to be achievable within a default 10-minute window; a speed_factor parameter (default 1.0) can stretch tasks to much longer durations (e
the paper’s own words
SentinelBench measures task completion, reaction time, and resource use, exposing the tradeoff between responsiveness and cost.
By default, each of the 100 tasks is designed to be achievable within 10 minutes ... a speed_factor parameter (default: 1.0) can be applied to stretch tasks to much longer durations ... tasks may require as long as 40 minutes to complete
iOSWorld2026relevant
iOSWorld: A Benchmark for Personally Intelligent Phone Agents
Goals to track complete single-app tasks within one iOS app (27 tasks); complete multi-app task chains spanning 2 to 8 apps (60 tasks); infer personal patterns from persistent user data for memory/personalization tasks (46 tasks)
Must keep track of Agent must track a persistent user identity and its connected data (transactions, messages, travel records, social relationships, financial activity) across apps and infer behavioral patterns for personalization tasks.
structure: DAG-with-precedence · goals arrive: given-up-front · count: 133 tasks total: 27 single-app, 60 multi · subgoal credit: no · 3 citations
Horizonthe paper states none
the paper’s own words
iOSWorld includes 133 tasks across three increasingly difficult categories. Single-app tasks (27) test one app, multi-app tasks (60) span 2 to 8 apps, and memory and personalization tasks (46) require agents to infer patterns from personal data.
What Memory Do GUI Agents Really Need? From Passive Records to Active Task-Driving States
Goals to track act on every list entry that satisfies the given instruction (positive matches); reject/skip every list entry that violates the instruction's constraints (negative matches)
Must keep track of The agent must track which near-identical list entries it has already acted on vs. still pending, and correctly apply the instruction's inclusion/exclusion constraints to each entry across a long trajectory.
structure: set-of-independent · goals arrive: given-up-front · subgoal credit: yes · 3 citations · also Personal assistants &amp; long-term memory
Horizonthe paper states none
the paper’s own words
From a list of near identical entries, agents must act on every entry that satisfies the instruction and reject entries that violate its constraints. We further introduce App-Level Progress and Scope-Aware F1 to measure these two dimensions separately.
Mobile GUI agents increasingly face long-horizon tasks that require reading, updating, and reusing task-relevant data across pages and applications.
Mirage-1: Augmenting and Updating GUI Agent with Hierarchical Multimodal Skills
Goals to track complete real-world long-horizon, multi-app Android task scenarios using previously acquired hierarchical skills; correctly apply execution skills, core skills, and meta-skills at the appropriate level of abstraction across a long-horizon task
Must keep track of The agent must track which level of its hierarchical skill structure (execution, core, meta) is currently relevant, maintain state as it moves across multiple apps in a long-horizon scenario, and bridge the offline-to-online domain gap without losing track of previously acquired skills.
structure: hierarchical · goals arrive: given-up-front · count: 30 diverse tasks across multiple applica · subgoal credit: yes · 13 citations
Horizonthe paper states none
the paper’s own words
To validate the performance of Mirage-1 in real-world long-horizon scenarios, we constructed a new benchmark, AndroidLH. Experimental results show that Mirage-1 outperforms previous agents by 32%, 19%, 15%, and 79% on AndroidWorld, MobileMiniWob++, Mind2Web-Live, and AndroidLH, respectively.
AndroidLens: Long-latency Evaluation with Nested Sub-targets for Android GUI Agents
Goals to track complete nested sub-targets within a single long-latency mobile task; satisfy multi-constraint, multi-goal task requirements drawn from 38 real-world domains; make measurable milestone-level progress even when full task success is not reached
Must keep track of Agent must track which nested sub-targets have been completed, tolerate environmental anomalies, and retain long-term memory of earlier steps across an average of more than 26 steps per task.
structure: hierarchical · goals arrive: given-up-front · count: 571 tasks (average >26 steps each) acros · subgoal credit: yes · 2 citations · unit: agent-steps
Horizon>26 (average, 'more than 26')
the paper’s own words
(3) dynamic evaluation that employs a milestone-based scheme for fine-grained progress measurement via Average Task Progress (ATP)
comprising 571 long-latency tasks in both Chinese and English environments, each requiring an average of more than 26 steps to complete
Benchmarking In-context Experiential Learning Through Repeated Product Recommendations
Goals to track elicit unknown customer preferences through questions within a single recommendation interaction (episode); tailor questioning/recommendation strategy based on patterns observed across multiple prior episodes (customers/products)
Must keep track of The agent must track what it has learned about the current customer's preferences within an episode (turn-by-turn), and must also track/aggregate patterns across multiple prior episodes to adapt its questioning/recommendation strategy over time.
structure: sequential-chain · goals arrive: given-up-front · count: benchmark built from a catalog of real-w · subgoal credit: yes · 2 citations · also Personal assistants &amp; long-term memory
Horizonthe paper states none
the paper’s own words
We instantiate the Benchmark for Experiential Learning and Active exploration (BELA) by combining (1) a rich catalog of real-world products from Amazon, (2) a diverse collection of synthetic customer personas aimed to capture heterogeneous latent preferences, and (3) an LLM-based customer simulator framework that emulate preference-revealing interactions... Benchmarking current models reveals that they can learn across turns, but struggle to improve across episodes.
we instantiate the Benchmark for Experiential Learning and Active exploration (BELA) by combining (1) a rich catalog of real-world products from Amazon, (2) a diverse collection of synthetic customer personas
CAPBench2025relevant
MobileSteward: Integrating Multiple App-Oriented Agents with Self-Evolution to Automate Cross-App Instructions
Goals to track associate and sequence sub-tasks across multiple mobile apps per a cross-app instruction; assign each associated sub-task to the correct app-oriented StaffAgent; avoid error propagation and information loss across the multi-step, multi-app execution; complete each of the 500 cross-app instructions spanning 14 apps in 6 categories
Must keep track of The centralized StewardAgent must track the scheduling graph of inter-app task associations, information flow between StaffAgents, and self-evolving memory of past executions to avoid repeating errors across a cross-app instruction.
structure: DAG-with-precedence · goals arrive: given-up-front · count: 500 cross-app instructions across 14 app · subgoal credit: no · 24 citations · unit: other:not-stated
Horizonthe paper states none
the paper’s own words
we propose a self-evolving multi-agent framework named MobileSteward which integrates multiple app-oriented StaffAgents coordinated by a centralized StewardAgent. ... Dynamic Recruitment generates a scheduling graph guided by information flow to explicitly associate tasks among apps. ... We establish the first English Cross-APP Benchmark (CAPBench) in the real-world environment
Step Rate measures the accuracy of individual actions executed by the StaffAgents.
ColorBench2025relevant
ColorBench: Benchmarking Mobile Agents with Graph-Structured Framework for Complex Long-Horizon Tasks
Goals to track complete a single-app mobile task via one of multiple valid GUI action paths; complete a cross-app mobile task requiring coordination across multiple applications; reach subtask-level completion milestones within a longer task
Must keep track of Agent must track which subtasks have been completed, which of several valid paths it is following through the task's state graph, and avoid known error paths, across an average of more than 13 steps.
structure: DAG-with-precedence · goals arrive: given-up-front · count: 175 tasks (74 single-app, 101 cross-app) · subgoal credit: yes · 8 citations · unit: agent-steps
Horizon>13 (average, 'over 13 steps')
the paper’s own words
we develop ColorBench, a benchmark focused on complex long-horizon tasks. It supports evaluation of multiple valid solutions, subtask completion rate statistics, and atomic-level capability analysis.
ColorBench contains 175 tasks (74 single-app, 101 cross-app) with an average length of over 13 steps
DeepShop2025relevant
DeepShop: A Benchmark for Deep Research Shopping Agents
Goals to track satisfy multiple product-attribute constraints in a single shopping query; apply the correct search filters specified or implied by the query; apply the correct sorting preference specified or implied by the query; achieve overall shopping-task success across easy/medium/hard complexity tiers
Must keep track of Agent must track which product attributes, filters, and sorting preferences the query requires, and verify each is correctly reflected in its final shopping actions/results.
structure: set-of-independent · goals arrive: given-up-front · count: queries evolved to three complexity leve · subgoal credit: yes · 42 citations
Horizonthe paper states none
the paper’s own words
(3) Fine-grained and holistic evaluation: We propose an automated evaluation framework that assesses agent performance in terms of fine-grained aspects (product attributes, search filters, and sorting preferences) and reports the overall success rate through holistic evaluation.
We further evolve these queries to increase complexity, considering product attributes, search filters, and sorting preferences, and classify them into three levels: easy, medium, and hard, based on the number of evolutions.
MVISU-Bench2025relevant
MVISU-Bench: Benchmarking Mobile Agents for Real-World Tasks by Multi-App, Vague, Interactive, Single-App and Unethical Instructions
Goals to track complete multi-app instructions requiring cross-app subgoal coordination; clarify vague/underspecified user instructions before acting; handle interactive instructions requiring mid-task clarification; complete single-app instructions; recognize and appropriately refuse or handle unethical instructions
Must keep track of The agent must track task state across multiple mobile apps for Multi-App instructions, detect ambiguity requiring clarification for Vague/Interactive instructions, and recognize when an instruction should be refused for Unethical instructions.
structure: other:five-independent-instruction-categories- · goals arrive: mixed:given-up-front-with-interactive-clarific · count: 404 tasks across 137 mobile applications · subgoal credit: no · 7 citations
Horizonthe paper states none
the paper’s own words
we present MVISU-Bench, a bilingual benchmark that includes 404 tasks across 137 mobile applications
Mobile-Eval-RAG2025relevant
Mobile-Agent-RAG: Driving Smart Multi-Agent Coordination with Contextual Knowledge Empowerment for Long-Horizon Mobile Automation
Goals to track complete a cross-application mobile-automation task requiring both high-level plan steps and precise low-level UI operations; correctly execute app-specific atomic UI actions aligned with the current subtask
Must keep track of Agent must track which app/subtask it is currently operating in, the current step of the high-level plan, and precise UI state needed for accurate atomic actions across multiple apps in one task.
structure: hierarchical · goals arrive: given-up-front · subgoal credit: yes · 2 citations · also Multi-agent organisations &amp; societies
Horizonthe paper states none
the paper’s own words
Furthermore, we introduce Mobile-Eval-RAG, a challenging benchmark for evaluating such agents on realistic multi-app, long-horizon tasks.
Mobile agents show immense potential, yet current state-of-the-art (SoTA) agents exhibit inadequate success rates on real-world, long-horizon, cross-application tasks.
MobileWorld2025relevant
MobileWorld: Benchmarking Autonomous Mobile Agents in Agent-User Interactive and MCP-Augmented Environments
Goals to track complete long-horizon, cross-application mobile workflows spanning up to 20 applications; handle vague user instructions and hybrid tool usage; coordinate agent-user interaction and MCP-augmented tool calls mid-task
Must keep track of Agent must track cross-application state across nearly twice as many completion steps on average (27.8 vs. 14.3) as AndroidWorld, while also handling user-interaction requests and MCP-tool-call state.
structure: sequential-chain · goals arrive: mixed:given-up-front-with-injected-user-intera · count: 201 tasks across 20 applications; 62.2% · subgoal credit: no · 58 citations · unit: agent-steps
Horizon27.8 (vs. 14.3 in AndroidWorld)
the paper’s own words
We introduce MobileWorld, a substantially more challenging benchmark designed to reflect real-world usage through 201 tasks across 20 applications. MobileWorld derives its difficulty from an emphasis on long-horizon, cross-application workflows, requiring nearly twice as many completion steps on average (27.8 vs. 14.3) and featuring a significantly higher proportion of multi-app tasks (62.2% vs. 9.5%) than AndroidWorld.
NaturalGAIA2025relevant
NaturalGAIA: A Verifiable Benchmark and Hierarchical Framework for Long-Horizon GUI Tasks
Goals to track decompose a natural, non-linear human GUI intent into a structured Task Topology of atomic sub-tasks; dynamically schedule sub-tasks across heterogeneous agents via context evolution; execute each atomic sub-task with precision via hybrid visual-structural perception; achieve a high Weighted Pathway Success Rate across the full causal pathway
Must keep track of The manager must dynamically track the Task Topology of atomic sub-tasks, schedule them across heterogeneous agents, and evolve shared context to bridge information gaps between dependent steps, assessed via a hierarchical success/error-attribution framework.
structure: DAG-with-precedence · goals arrive: self-generated-by-agent · subgoal credit: no · 0 citations · unit: other:not-stated
Horizonthe paper states none
the paper’s own words
By decoupling logical causal pathways from linguistic narratives, it rigorously simulates natural human intent, characterized by cognitive non-linearity and contextual dependencies. ... The parser decomposes abstract user intents into a structured Task Topology composed of atomic tasks. ... Experiments demonstrate that our approach achieves a Weighted Pathway Success Rate of 45.6%, significantly outperforming the state-of-the-art baseline (21.1%)
RealWebAssist: A Benchmark for Long-Horizon Web Assistance with Real-World Users
Goals to track correctly follow each instruction in a sequence of real, sequentially-issued user instructions across multiple websites; reason about the true intent behind ambiguous instructions; keep track of the user's mental state and user-specific routines as instructions evolve over the session; ground each intended task to the correct GUI element on the current website
Must keep track of Agent must retain a running model of the user's intent, mental state, and personal routines across a long sequence of instructions issued over multiple websites, using this history to correctly interpret each new (sometimes ambiguous) instruction.
structure: sequential-chain · goals arrive: injected-by-user-mid-episode · subgoal credit: unclear · 26 citations
Horizonthe paper states none
the paper’s own words
RealWebAssist includes a dataset of sequential instructions collected from real-world human users. Each user instructs a web-based assistant to perform a series of tasks on multiple websites.
To achieve successful assistance with long-horizon web-based tasks, AI agents must be able to sequentially follow real-world user instructions over a long period.
ShoppingBench2025relevant
ShoppingBench: A Real-World Intent-Grounded Shopping Benchmark for LLM-based Agents
Goals to track apply vouchers correctly to a purchase; manage a budget across a shopping session; find and select from multi-product sellers matching a grounded intent; satisfy increasingly challenging levels of grounded shopping intent end-to-end
Must keep track of Agent must track running spend/budget, applied vouchers, and multi-seller product-matching requirements across a session, operating within a sandbox of over 2.5 million real-world products.
structure: other:joint-constraint-satisfaction-within-one · goals arrive: given-up-front · subgoal credit: no · 27 citations
Horizonthe paper states none
the paper’s own words
real-world users often pursue more complex goals, such as applying vouchers, managing budgets, and finding multi-products seller. To bridge this gap, we propose ShoppingBench, a novel end-to-end shopping benchmark designed to encompass increasingly challenging levels of grounded intent.
we provide a large-scale shopping sandbox that serves as an interactive simulated environment, incorporating over 2.5 million real-world products
ShoppingComp2025relevant
ShoppingComp: Are LLMs Really Ready for Your Shopping Cart?
Goals to track retrieve products satisfying many simultaneous discovery constraints; generate an expert-level report on the retrieved products; make a safety-critical purchase decision (e.g. flag unsafe product usage)
Must keep track of The agent must track which of the multiple product-discovery constraints have been satisfied, the evidence supporting its report, and any identified safety hazards, while operating in an open-world product catalogue.
structure: set-of-independent · goals arrive: given-up-front · count: 145 instances and 558 scenarios (curated · subgoal credit: no · 8 citations · unit: other:not-stated
Horizonthe paper states none
the paper’s own words
ShoppingComp, a challenging real-world benchmark for comprehensively evaluating LLM-powered shopping agents on three core capabilities: precise product retrieval, expert-level report generation, and safety critical decision making. ... The benchmark comprises 145 instances and 558 scenarios
The benchmark comprises 145 instances and 558 scenarios, curated by 35 experts to reflect authentic shopping needs.
UI-NEXUS2025relevant
Atomic-to-Compositional Generalization for Mobile Agents with A New Benchmark and Scheduling System
Goals to track complete compositional mobile operations that concatenate multiple atomic tasks (Simple Concatenation); transition context correctly across sub-tasks within one compositional task (Context Transition); perform a deep multi-step drill-down within an app or workflow (Deep Dive)
Must keep track of The agent must track which atomic subtasks within a compositional task it has already completed, the context state carried between apps/screens, and overall progress toward each of the three compositional operation types.
structure: hierarchical · goals arrive: given-up-front · count: 100 interactive task templates, average · subgoal credit: yes · 12 citations · unit: actions
Horizonaverage optimal step count of 14.05 (over 100 interactive task templates)
the paper’s own words
UI-NEXUS supports interactive evaluation in 20 fully controllable local utility app environments, as well as 30 online Chinese and English service apps. It comprises 100 interactive task templates with an average optimal step count of 14.05.
VeriWeb2025relevant
VeriWeb: Verifiable Long-Chain Web Benchmark for Agentic Information-Seeking
Goals to track ensure comprehensive information coverage across breadth- and depth-oriented multi-hop web searches; complete a sequence of interdependent, individually-verifiable subtasks within one long-chain web task; maintain consistent context tracking across a long information-seeking chain
Must keep track of Agent must track and verify each subtask-level answer as it progresses through a long chain of interdependent web-search subtasks, ensuring comprehensive information coverage and consistent context tracking across the chain.
structure: sequential-chain · goals arrive: given-up-front · count: 302 tasks across five real-world domains · subgoal credit: yes · 10 citations
Horizonthe paper states none
the paper’s own words
(2) subtask-level verifiability, where tasks are decomposed into a sequence of interdependent verifiable subtasks. This structure enables diverse exploration strategies within each subtask, while ensuring that each subtask-level answer remains unchanged and verifiable.
The benchmark consists of 302 tasks across five real-world domains, each with a complete trajectory demonstration, annotated by human experts.
WebChoreArena2025relevant
WebChoreArena: Evaluating Web Browsing Agents on Realistic Tedious Web Tasks
Goals to track retain and retrieve large amounts of information gathered from many observations (Massive Memory); perform precise mathematical/quantitative reasoning over collected information (Calculation); keep information consistent while tracking it across multiple webpages over the course of a task (Long-Term Memory)
Must keep track of Agent must accumulate and correctly recall large amounts of information across many webpages, perform arithmetic over it, and keep facts consistent across the whole task rather than a single page view.
structure: sequential-chain · goals arrive: given-up-front · count: 532 tasks · subgoal credit: unclear · 28 citations · also Information seeking &amp; deep research
Horizonthe paper states none
the paper’s own words
It systematically expands the evaluation space along three critical dimensions: (i) $\textbf{Massive Memory}$... (ii) $\textbf{Calculation}$... and (iii) $\textbf{Long-Term Memory}$, necessitating consistent information tracking across multiple webpages.
WebChoreArena introduces 532 carefully curated tasks developed over 300+ hours
WebMall2025relevant
WebMall - A Multi-Shop Benchmark for Evaluating Web Agents
Goals to track find a specific product across four simulated online shops; perform price comparisons across shops; identify suitable substitutes or compatible products (advanced search); add items to a cart and complete checkout in the correct shop(s)
Must keep track of Agent must retain product/price information gathered from each of the four shops it has visited so far in order to correctly compare, substitute, or complete a checkout later in the same task.
structure: sequential-chain · goals arrive: given-up-front · count: 91 cross-shop tasks across four simulate · subgoal credit: unclear · 14 citations
Horizonthe paper states none
the paper’s own words
LLM-based web agents have the potential to automate long-running web tasks, such as searching for products in multiple e-shops and subsequently ordering the cheapest products that meet the users needs.
the best-performing agents achieved task completion rates below 65% in the task categories cheapest product search and vague product search
Web-Shepherd: Advancing PRMs for Reinforcing Web Agents
Goals to track navigate a website across many sequential steps toward a stated goal; at each step, satisfy sub-goals encoded in an annotated checklist (process-level correctness); complete the overall web-navigation task successfully (episode-level correctness)
Must keep track of The reward model must track which checklist sub-goals have already been satisfied at each step of a trajectory so it can assess the process-level (not just outcome-level) quality of a web-navigation trajectory.
structure: sequential-chain · goals arrive: given-up-front · count: 40K step-level preference pairs with ann · subgoal credit: yes · 29 citations · unit: actions
Horizontrajectory lengths vary by difficulty: median ~5 steps (easy), ~9 steps (medium), ~20 steps (hard), with some trajectories exceeding 40 steps
the paper’s own words
we propose the first process reward model (PRM) called Web-Shepherd which could assess web navigation trajectories in a step-level. To achieve this, we first construct the WebPRM Collection, a large-scale dataset with 40K step-level preference pairs and annotated checklists spanning diverse domains and difficulty levels.
Easy tasks generally require fewer steps (median ≈ 5), whereas medium tasks show more variability (median ≈ 9), and hard tasks involve significantly longer trajectories (median ≈ 20), with some exceeding 40 steps.

Games & interactive fiction 34 artifacts · 12 state a horizon

AgentOdyssey: Open-Ended Long-Horizon Text Game Generation for Test-Time Continual Learning Agents
Goals to track explore procedurally generated text-game worlds to acquire new world knowledge and skills; retain and reuse relevant episodic experiences across a continuous test-time deployment; make game progress toward in-game objectives while continuing to learn at test time; diagnostic sub-goals: exploring objects/actions, maintaining action diversity, controlling model cost
Must keep track of Agent must track world knowledge acquired, episodic memories, game progress, and cost across a continuous, long-horizon test-time deployment rather than within a single isolated episode.
structure: open-ended · goals arrive: mixed:emitted-by-environment-plus-self-generat · subgoal credit: yes · 2 citations
Horizonthe paper states none
the paper’s own words
we introduce AgentOdyssey, a novel evaluation framework that procedurally generates open-ended text games with rich entities, world dynamics, and long-horizon tasks... We further propose a multifaceted evaluation methodology that measures not only game progress but also offers diagnostic tests on world knowledge acquisition, episodic memory, object and action exploration, action diversity, and model cost.
Critically, AgentOdyssey goes beyond the conventional machine learning assumption that learning does not occur at test time by placing agents in a continuous, long-horizon setting that interleaves learning and inference throughout deployment.
AgenticSTS: A Bounded-Memory Testbed for Long-Horizon LLM Agents
Goals to track win a run of a closed-rule stochastic deck-building game via a long sequence of tactical (card-level) and strategic (build-level) decisions
Must keep track of Agent must make each decision from a freshly-assembled, typed-retrieval memory (not a raw appended transcript), so it must track which prior tactical/strategic outcomes are relevant to retrieve for the current decision, bounded so the prompt does not grow with run length.
structure: sequential-chain · goals arrive: given-up-front · subgoal credit: yes · 0 citations · also Personal assistants &amp; long-term memory · unit: agent-steps
Horizonhundreds (of tactical and strategic decisions per run); 298 completed trajectories released
the paper’s own words
We instantiate the contract in Slay the Spire 2, a closed-rule stochastic deck-building game whose runs require hundreds of tactical and strategic decisions.
CivBench2026in
CivBench: Progress-Based Evaluation for LLMs' Strategic Decision-Making in Civilization V
Goals to track pursue victory in multiplayer Civilization V against multiple opponents; maintain and improve turn-level estimated victory probability across hundreds of turns; balance strategic dimensions (economic, military, diplomatic) that jointly determine long-run standing
Must keep track of The agent must track turn-level game state used to estimate its own victory probability continuously across a game spanning hundreds of turns and multiple opponents, rather than only checking win/loss at the end.
structure: open-ended · goals arrive: given-up-front · subgoal credit: no · 2 citations · unit: turns
Horizonhundreds of turns; 307 games evaluated
the paper’s own words
Because terminal win/loss is too sparse a signal in games spanning hundreds of turns and multiple opponents, CivBench trains models on turn-level game state to estimate victory probabilities throughout play, validated through predictive, construct, and convergent validity.
CivBench: A Long-Horizon Benchmark for Tool-Mediated Agents in Civilization VI
Goals to track sustain long-horizon strategic planning and execution across a single 300+-turn Civilization VI episode; proactively monitor latent strategic state (e.g., victory progress) rather than only reactively responding; execute near-term commitments stated in the agent's own planning reflections within a bounded number of subsequent turns; operate correctly across 76 exposed MCP tools under partial observability
Must keep track of The agent must proactively query and track latent strategic state (victory progress) at recommended intervals, monitor for approaching defeat within a warning window, and track its own prior planning commitments to execute them within a bounded number of subsequent turns, across an episode spanning 300+ turns and thousands of tool calls.
structure: sequential-chain · goals arrive: given-up-front · count: 76 MCP tools exposed; 23 admissible runs · subgoal credit: no · 0 citations · unit: turns
Horizona single episode spans 300+ turns and produces thousands of tool calls
the paper’s own words
we introduce two interface-level metrics that the environment makes measurable: Proactive Monitoring Rate (PMR), capturing whether agents actively query latent strategic state, and RAG@10, capturing whether commitments stated in structured planning reflections are executed within ten subsequent turns.
A single episode spans 300+ turns and produces thousands of tool calls over a large action space, requiring sustained planning, state monitoring, and execution under partial observability.
LUMINA2026relevant
LUMINA: Long-horizon Understanding for Multi-turn Interactive Agents
Goals to track complete a multi-turn, game-like task requiring planning and state tracking (ListWorld, TreeWorld, GridWorld); benefit from oracle interventions (perfect planning or flawless state tracking) to isolate which underlying skill limits performance; generalize across procedurally generated task variants with tunable complexity
Must keep track of The agent must plan across multiple turns and track evolving task state (e.g. a modified list, a searched tree, or a navigated grid position) to succeed, and the oracle framework tests how much this tracking/planning burden, if perfectly handled, would improve performance.
structure: sequential-chain · goals arrive: given-up-front · subgoal credit: yes · 1 citations
Horizonthe paper states none
the paper’s own words
We introduce a suite of procedurally generated, game-like tasks with tunable complexity. These controlled environments allow us to provide precise oracle interventions, such as perfect planning or flawless state tracking, and make it possible to isolate the contribution of each oracle without confounding effects present in real-world benchmarks.
MINDGAMES: A Live Arena for Evaluating Social and Strategic Reasoning in Multi-Agent LLMs
Goals to track allocate resources under hidden-information belief attribution across repeated Colonel Blotto rounds; sustain cooperative/competitive strategy under opponent modeling in Iterated Prisoner's Dilemma; cooperatively infer hidden information under knowledge asymmetries in Codenames; detect or sustain deception across social-deduction rounds in Secret Mafia
Must keep track of The agent must track other players' modeled beliefs/strategies (opponent modeling), maintain internal consistency in its own hidden role or hidden information across repeated rounds, and adapt as new turn-level observations arrive within a game.
structure: set-of-independent · goals arrive: given-up-front · count: four game environments; 944 submitted ag · subgoal credit: no · 3 citations · also Multi-agent organisations &amp; societies
Horizonthe paper states none
the paper’s own words
Built on TextArena, Mindgames provides a unified interaction interface, TrueSkill-based rating, and full trajectory logging across four game environments. We instantiate Mindgames through a 2025 competition cycle hosted at a major AI conference, which assessed 944 submitted agents from 76 teams across four games.
We release a dataset of 29,571 multi-agent games with turn-level observations, actions, and rewards, together with MG-Ref, a deterministic offline tournament protocol that scores new agents against a frozen reference pool of top-ranked, low-error Stage~II submissions under the same error-attribution lens used in this analysis.
MineExplorer: Evaluating Open-World Exploration of MLLM Agents in Minecraft
Goals to track solve implicit multi-hop tasks composed of chained atomic open-world exploration sub-tasks; coordinate hidden prerequisites across longer trajectories to sustain open-world exploration
Must keep track of Agent must track hidden prerequisites satisfied by earlier atomic sub-tasks as it proceeds through longer, multi-hop exploration trajectories, since task difficulty tracks agent completion and hidden dependencies are not explicitly revealed.
structure: DAG-with-precedence · goals arrive: given-up-front · subgoal credit: unclear · 2 citations
Horizonthe paper states none
the paper’s own words
Then we organize the benchmark around a ReAct-style capability formulation and compose atomic tasks into implicit multi-hop tasks.
strong models can handle many single-hop tasks but degrade sharply when hidden prerequisites must be coordinated over longer trajectories
MineNPC-Task: Task Suite for Memory-Aware Minecraft Agents
Goals to track satisfy the explicit preconditions and dependency structure of a parametric, user-elicited Minecraft task template; correctly use plan previews, targeted clarifications, memory reads/writes, and repair attempts (mixed-initiative interaction) while executing subtasks; avoid out-of-world shortcuts by only using in-world evidence (bounded-knowledge policy)
Must keep track of Agent/harness must track plan previews, in-world memory reads and writes, whether stated preconditions for each subtask are currently satisfied, and any repair attempts needed after a breakdown, using only in-world evidence.
structure: DAG-with-precedence · goals arrive: given-up-front · count: 216 subtasks (evaluated across 8 experie · subgoal credit: unclear · 1 citations · also Personal assistants &amp; long-term memory
Horizonthe paper states none
the paper’s own words
The harness captures plan, action, and memory events, including plan previews, targeted clarifications, memory reads and writes, precondition checks, and repair attempts, and reports outcomes relative to the total number of attempted subtasks using only in-world evidence.
we instantiate the framework with GPT-4o and evaluate 216 subtasks across 8 experienced players
MineEvolve: Self-Evolution with Accumulated Knowledge for Long-Horizon Embodied Minecraft Agents
Goals to track craft tools; build redstone components; obtain diamond equipment; recover and continue long prerequisite chains despite missing tools, blocked paths, GUI failures, or stagnant execution
Must keep track of Agent must track per-subgoal execution outcomes (state changes, inventory changes, failure types, progress and stagnation signals) and accumulate them into reusable skills or remedies to repair plans under repeated failure.
structure: DAG-with-precedence · goals arrive: given-up-front · subgoal credit: yes · 0 citations · also Embodied &amp; robotics
Horizonthe paper states none
the paper’s own words
Minecraft provides a representative testbed for this problem, where tasks such as crafting tools, building redstone components, and obtaining diamond equipment involve long prerequisite chains and are frequently disrupted by missing tools, blocked paths, GUI failures, or stagnant execution... MineEvolve first uses Monitor to convert each subgoal execution into typed feedback, including state changes, inventory changes, failure types, progress signals, and stagnation indicators.
MirrorCraft: Paired Evaluation under Hidden Rule Changes in Minecraft
Goals to track achieve one of three progression objectives per matched Vanilla/Mirror world pair; adapt to hidden server-side rule changes (recipes, drops, other mechanics) altered by a datapack; reach deterministic advancement milestones despite altered rules
Must keep track of Agent must track task/advancement-milestone progress under its assigned rule suite while detecting and adapting to whichever server-side rule was silently modified in its Mirror world.
structure: set-of-independent · goals arrive: mixed:given-up-front-objectives-with-hidden-en · count: five controlled biomes, six rule suites, · subgoal credit: yes · 0 citations
Horizonthe paper states none
the paper’s own words
We evaluate task progress with deterministic advancement milestones and success rate and use the Rule Intervention Effect (RIE) to measure the performance change between matched Vanilla and Mirror worlds.
MirrorCraft includes five controlled biomes, six rule suites, three progression objectives, two model families, and six agent configurations under a shared Mineflayer interface.
NARRA-Gym2026relevant
NARRA-Gym for Evaluating Interactive Narrative Agents
Goals to track sustain a coherent, evolving story across multiple turns while adapting to a specific user persona; manage long-context state and pacing across the episode; maintain consistent character simulation and empathic personalization; optionally synthesize a story-grounded artifact at the end of the episode
Must keep track of The agent must maintain long-context story state, memory updates, pacing decisions, and an evolving user-persona model consistently across the whole interactive episode.
structure: sequential-chain · goals arrive: given-up-front · subgoal credit: yes · 0 citations · unit: wall-clock-minutes
Horizoneach interactive episode lasts roughly 20 minutes (human evaluation)
the paper’s own words
We introduce NARRA-Gym, an executable evaluation environment that turns a sparse emotional seed into a complete interactive story episode and logs the full model-in-the-loop trajectory, including story construction, memory updates, planning, pacing interventions, and optional artifact synthesis.
each interactive episode lasts roughly 20 minutes
OmniGameArena2026relevant
OmniGameArena: A Unified UE5 Benchmark for VLM Game Agents with Improvement Dynamics
Goals to track achieve a high score/objective in each of 12 distinct UE5 games (7 Solo, 3 PvP, 2 Coop); reflect on past-round performance to refine a bounded skill prompt across multiple rounds (Improvement Dynamics Curve); generalize a learned/refined skill to held-out task variants
Must keep track of The reflector must track trajectories and the persistent skill state across rounds, retaining what previously worked or failed, to progressively refine the bounded skill prompt rather than starting fresh each round.
structure: sequential-chain · goals arrive: mixed:given-up-front-game-objectives-plus-self · count: 12 games (7 Solo, 3 PvP, 2 Coop); R=10 r · subgoal credit: yes · 0 citations · unit: episodes
HorizonR=10 reflection rounds, each comprising K=5 episodes (50 episodes total per agent-game IDC run)
the paper’s own words
we address these gaps with OmniGameArena, a real-time benchmark of twelve newly built Unreal Engine 5 games spanning Solo (7), PvP (3), and Coop (2) with unified action interfaces, and the Improvement Dynamics Curve (IDC), an agentic-reflection harness in which a tool-using reflector LLM autonomously refines a bounded skill prompt across multiple rounds.
Each agent completes R=10 rounds of K=5 episodes under PDQ
PokeGym2026in
Mastering PokeGym: Graph-Guided Multimodal Evolution at Test Time
Goals to track complete a long-horizon task in a 3D open-world game using only visual observations (no game-state access); improve/adapt the agent's own configuration (perception, strategy, action set) across consecutive episodes of the same task (test-time learning); jointly optimize perception, reasoning, and control rather than a single modality in isolation
Must keep track of The agent must track its own evolving configuration (which perception/strategy/action choices helped or hurt) across consecutive episodes of the same task, in addition to in-episode state, since the goal is test-time learning rather than a single fixed-policy run.
structure: hierarchical · goals arrive: mixed:given-up-front-task-with-self-generated- · count: 30 tasks derived from 10 quests; traject · subgoal credit: yes · 1 citations · also Web &amp; GUI agents · unit: agent-steps
Horizon30 tasks derived from 10 quests, with trajectories ranging from 30 to 220 environment steps
the paper’s own words
we first introduce \textbf{PokeGym}, a long-horizon benchmark built upon the 3D open-world game Pok\'emon Legends: Z-A, where agents act from visual observations without access to game states, designed to evaluate an agent's ability to learn and adapt across consecutive episodes of the task.
PokeGym is a benchmark that contains 30 tasks derived from 10 quests, with trajectories ranging from 30 to 220 environment steps.
TowerMind2026relevant
TowerMind: A Tower Defence Game Learning Environment and Benchmark for LLM as Agents
Goals to track perform macro-level strategic planning (tower placement, resource allocation) in a tower-defense scenario; perform micro-level tactical adaptation and action execution in response to incoming waves; succeed across each of five designed benchmark levels under different multimodal input settings; avoid hallucinating about game state while planning and acting
Must keep track of The agent must track macro-level game state (tower placements, resources, wave progression) and micro-level tactical details (unit/enemy positions) across the tower-defense match, while its outputs are additionally checked for hallucination relative to the true game state.
structure: hierarchical · goals arrive: given-up-front · count: five benchmark levels; evaluated under d · subgoal credit: no · 5 citations
Horizonthe paper states none
the paper’s own words
We design five benchmark levels to evaluate several widely used LLMs under different multimodal input settings. The results reveal a clear performance gap between LLMs and human experts across both capability and hallucination dimensions.
their inherent gameplay requires both macro-level strategic planning and micro-level tactical adaptation and action execution
alem2026in
Benchmarking Open-Ended Multi-Agent Coordination in Language Agents
Goals to track survive and grow within a long-horizon Craftax-like survival world (exploration, crafting, trading, combat); coordinate role allocation with teammates (soft specialisation); communicate to allocate roles and execute shared plans; solve procedurally generated coordination tasks of controllable difficulty
Must keep track of Agents must track their own survival state (health, resources, crafted items), teammates' roles/communications, and progress on procedurally generated coordination tasks within a long-horizon survival world.
structure: open-ended · goals arrive: mixed:survival-coordination-framing-given-up-f · subgoal credit: no · 1 citations · also Multi-agent organisations &amp; societies
Horizonthe paper states none
the paper’s own words
Alem embeds procedurally generated coordination tasks, soft specialisation, communication, and controllable coordination difficulty into a long-horizon survival world with exploration, crafting, trading, and combat.
they must coordinate with others over long horizons in open-ended interactive tasks
AVACraft2025relevant
AVA: Attentive VLM Agent for Mastering StarCraft II
Goals to track complete micromanagement objectives (e.g. unit-level combat control); achieve coordination objectives among allied units/agents; execute strategic planning objectives across 21 StarCraft II scenarios
Must keep track of The agent/policy must track unit-level state (health, position) for micromanagement, coordinate allocation across units for team objectives, and maintain a strategic plan across the scenario, using RGB visuals, natural-language observations, and structured state.
structure: hierarchical · goals arrive: given-up-front · count: 21 scenarios spanning micromanagement, c · subgoal credit: no · 3 citations · unit: wall-clock-minutes
Horizon300 seconds (5 minutes) max per episode, or earlier if victory conditions are met
the paper’s own words
AVACraft provides RGB visuals, natural language observations, and structured state information, enabling systematic comparison between training-based and zero-shot methods across 21 scenarios spanning micromanagement, coordination, and strategic planning.
Episodes terminate when 300 seconds pass or when victory conditions are met.
Collab-Overcooked2025relevant
Collab-Overcooked: Benchmarking and Evaluating Large Language Models as Collaborative Agents
Goals to track fulfill multiple simultaneous cooking orders/objectives via natural-language multi-agent coordination; actively collaborate and continuously adapt strategy as the shared kitchen task unfolds
Must keep track of Agents must track the shared kitchen's evolving state (ingredient/tool locations, in-progress dishes), coordinate via natural-language communication to avoid resource conflicts, and manage their actions within a per-task time budget derived from the optimal completion time.
structure: DAG-with-precedence · goals arrive: given-up-front · count: 30 open-ended tasks across 6 complexity · subgoal credit: yes · 38 citations · also Multi-agent organisations &amp; societies · unit: agent-steps
Horizoneach task's time constraint is set as the optimal completion time scaled by a time-limit factor gamma (gamma=1.5 used in experiments); exact timestep counts vary by compl
the paper’s own words
Collab-Overcooked extends existing benchmarks in two novel ways. First, it provides a multi-agent framework supporting diverse tasks and objectives and encourages collaboration through natural language communication. Second, it introduces a spectrum of process-oriented evaluation metrics to assess the fine-grained collaboration capabilities of different LLM agents, a dimension often overlooked in prior work.
Each task has a time constraint, set as the optimal completion time scaled by a time limit factor γ.
DSGBench2025relevant
DSGBench: A Diverse Strategic Game Benchmark for Evaluating LLM-based Agents in Complex Decision-Making Environments
Goals to track make long-term, multi-dimensional strategic decisions in each of six complex strategic games; adapt task difficulty and targets within customizable game settings
Must keep track of System tracks the agent's full decision trajectory across a game (via the automated decision-tracking mechanism) to identify behavior patterns and strategy turning points, and scores performance along five specific dimensions.
structure: sequential-chain · goals arrive: given-up-front · count: 6 strategic games; 5 evaluation dimensio · subgoal credit: yes · 21 citations
Horizonthe paper states none
the paper’s own words
DSGBench employs a fine-grained evaluation scoring system which examines the decision-making capabilities by looking into the performance in five specific dimensions, offering a comprehensive assessment in a better-designed fashion.
it incorporates six complex strategic games which serve as ideal testbeds due to their long-term and multi-dimensional decision-making demands
Factorio Learning Environment
Goals to track complete each of 8 fixed lab-play structured tasks; in open-play, build the largest possible factory on a procedurally generated map (an unbounded, self-scaling goal); scale automation from basic production to factories processing millions of resource units per second
Must keep track of Agent must track resource/production-chain state, spatial factory layout, and automation progress as goals scale from basic automation to factories processing millions of resource units per second.
structure: other:hierarchical-in-lab-play,-open-ended-in- · goals arrive: mixed:given-up-front-lab-play-and-open-ended-o · count: 8 fixed tasks in lab-play; open-play is · subgoal credit: yes · 5 citations
Horizonthe paper states none
the paper’s own words
We provide two settings: (1) lab-play consisting of eight structured tasks with fixed resources, and (2) open-play with the unbounded task of building the largest factory on an procedurally generated map.
FLE provides exponentially scaling challenges -- from basic automation to complex factories processing millions of resource units per second.
FlashAdventure: A Benchmark for GUI Agents Solving Full Story Arcs in Diverse Adventure Games
Goals to track complete each of a game's predefined success milestones in the correct narrative order; remember and act on earlier gameplay information to bridge the observation-behavior gap; complete the full story arc of one of 34 diverse Flash-based adventure games
Must keep track of The agent must remember earlier gameplay clues/information (long-term clue memory) and correctly act on them at later points in the story to progress through predefined milestones toward full story-arc completion.
structure: sequential-chain · goals arrive: given-up-front · count: 34 Flash-based adventure games spanning · subgoal credit: yes · 2 citations · also Web &amp; GUI agents · unit: other:not-stated
Horizonthe paper states none
the paper’s own words
we introduce FlashAdventure, a benchmark of 34 Flash-based adventure games designed to test full story arc completion and tackle the observation-behavior gap: the challenge of remembering and acting on earlier gameplay information. We also propose CUA-as-a-Judge, an automated gameplay evaluator, and COAST, an agentic framework leveraging long-term clue memory to better plan and solve sequential tasks.
COAST ... reaching a 5.88% success rate and a 19.89% milestone completion rate
HeroBench2025in
HeroBench: A Benchmark for Long-Horizon Planning and Structured Reasoning in Virtual Worlds
Goals to track select numerically feasible equipment given resource/stat constraints; reason over multi-level crafting and resource dependencies; execute hundreds to thousands of actions as a single coherent end-to-end plan; succeed in numeric combat simulation against scalable, adversarially-distracted difficulty
Must keep track of The agent must track multi-level crafting/resource dependencies, numeric feasibility of equipment choices, and spatial state across a single end-to-end plan comprising hundreds to thousands of actions, verified by simulation-based success and fine-grained progress metrics.
structure: hierarchical · goals arrive: given-up-front · subgoal credit: yes · 6 citations · unit: actions
Horizonhundreds to thousands
the paper’s own words
Tasks require models to select numerically feasible equipment, reason over multi-level crafting and resource dependencies, and execute hundreds to thousands of actions as a single end-to-end plan. HeroBench integrates symbolic planning, numeric combat simulation, spatial reasoning, and resource management ... HeroBench evaluates executable plans through simulation, enabling both success-based and fine-grained progress metrics, as well as detailed failure mode analysis.
EvoTest: Evolutionary Test-Time Learning for Self-Improving Agentic Systems
Goals to track solve an interactive-fiction game requiring many in-game puzzle subgoals (e.g., in Detective and Library) within an episode; improve performance from one episode to the next by adapting across consecutive playthroughs of the same game
Must keep track of The Actor Agent (and the Evolver Agent that analyzes transcripts) must track in-episode state-action choices and effective strategies, then carry a revised configuration (prompt, memory, hyperparameters, tool-use routines) forward to the next episode of the same game.
structure: sequential-chain · goals arrive: given-up-front · subgoal credit: yes · 31 citations
Horizonthe paper states none
the paper’s own words
J-TTL is a new evaluation setup where an agent must play the same game for several consecutive episodes, attempting to improve its performance from one episode to the next.
MineCollab2025relevant
Collaborating Action by Action: A Multi-agent LLM Framework for Embodied Reasoning
Goals to track control characters collaboratively in open-world Minecraft to complete complex embodied reasoning tasks; delegate sub-tasks between collaborating agents via natural-language communication; share and update task-completion plans among agents as the task progresses
Must keep track of Agents must track their own sub-task assignment, what has been communicated to/from collaborators, and the shared task-completion plan's current state across the collaborative embodied task.
structure: hierarchical · goals arrive: given-up-front · subgoal credit: yes · 25 citations · also Embodied &amp; robotics
Horizonthe paper states none
the paper’s own words
we introduce MINDcraft, an easily extensible platform built to enable LLM agents to control characters in the open-world game of Minecraft; and MineCollab, a benchmark to test the different dimensions of embodied and collaborative reasoning.
This work studies how LLMs can adaptively collaborate to perform complex embodied reasoning tasks.
LLM-Powered Decentralized Generative Agents with Adaptive Hierarchical Knowledge Graph for Cooperative Planning
Goals to track achieve open-world survival/crafting objectives cooperatively with other agents; share and act on relevant information from past interactions via a hierarchical knowledge-graph memory; coordinate via structured communication to avoid redundant or conflicting actions among 2-6 agents
Must keep track of Each agent must track its own past experience via a hierarchical knowledge-graph memory and selectively communicate relevant facts to teammates rather than sharing full history, in order to reach a shared long-term goal efficiently.
structure: DAG-with-precedence · goals arrive: given-up-front · subgoal credit: unclear · 17 citations
Horizonthe paper states none
the paper’s own words
Compared to single-agent scenarios, the two-agent scenario achieves the same goal with 63% fewer steps, and the six-agent scenario with 74% fewer steps, highlighting the importance of adaptive memory and structured communication in achieving long-term goals.
ParaCook2025relevant
ParaCook: On Time-Efficient Planning for Multi-Agent Systems
Goals to track prepare and deliver each of several dish orders correctly; coordinate parallel/asynchronous sub-tasks (e.g. chopping, cooking, plating) across multiple agents to minimize completion time; avoid collisions/conflicts over shared kitchen resources while parallelizing actions
Must keep track of Agents must track the preparation-stage state of each in-progress dish (which precedence steps are done), coordinate with other agents to avoid resource conflicts, and jointly minimize overall completion time across simultaneous orders.
structure: DAG-with-precedence · goals arrive: given-up-front · count: adjustable-complexity scalable evaluatio · subgoal credit: yes · 1 citations · also Multi-agent organisations &amp; societies
Horizonthe paper states none
the paper’s own words
Inspired by the Overcooked game, ParaCook provides an environment for various challenging interaction planning of multi-agent systems that are instantiated as cooking tasks, with a simplified action space to isolate the core challenge of strategic parallel planning.
ParaCook provides a scalable evaluation framework with adjustable complexity, establishing a foundation for developing and assessing time efficiency-aware multi-agent planning.
SC2Arena and StarEvolve: Benchmark and Self-Improvement Framework for LLMs in Complex Decision-Making Tasks
Goals to track manage a full StarCraft II game across the complete game context (all playable races); operate over diverse, low-level action spaces rather than a simplified/reduced action space; solve spatial reasoning challenges via text-based observations; integrate strategic planning with tactical execution (Planner-Executor-Verifier structure)
Must keep track of The agent must track the complete game state (all playable races, diverse action spaces) and its own strategic plan versus tactical execution outcomes, using a scoring system to select high-quality training samples for continuous improvement.
structure: hierarchical · goals arrive: given-up-front · subgoal credit: yes · 2 citations
Horizonthe paper states none
the paper’s own words
we present SC2Arena, a benchmark that fully supports all playable races, low-level action spaces, and optimizes text-based observations to tackle spatial reasoning challenges. Complementing this, we introduce StarEvolve, a hierarchical framework that integrates strategic planning with tactical execution, featuring iterative self-correction and continuous improvement via fine-tuning on high-quality gameplay data.
StarDojo2025in
StarDojo: Benchmarking Open-Ended Behaviors of Agentic Multimodal LLMs in Production-Living Simulations with Stardew Valley
Goals to track perform livelihood production activities (farming, crafting); engage in social interactions to build relationships within the community; complete tasks across five key domains: farming, crafting, exploration, combat, and social interactions
Must keep track of Agent must track progress across both production (farming/crafting/exploration/combat) and social-relationship goals concurrently, within a unified interface supporting parallel environment instances.
structure: set-of-independent · goals arrive: given-up-front · count: 1,000 curated tasks (100-task representa · subgoal credit: unclear · 5 citations
Horizonthe paper states none
the paper’s own words
In StarDojo, agents are tasked to perform essential livelihood activities such as farming and crafting, while simultaneously engaging in social interactions to establish relationships within a vibrant community.
StarDojo features 1,000 meticulously curated tasks across five key domains: farming, crafting, exploration, combat, and social interactions
StoryBench: A Dynamic Benchmark for Evaluating Long-Term Memory with Multi Turns
Goals to track correctly retain and recall facts/state established earlier in a branching narrative (knowledge retention); reason over sequences of narrative events to infer state changes and causal dependencies (sequential reasoning); under one setting, trace back and revise earlier choices after a failure is detected
Must keep track of The agent must retain established narrative facts and track which branch of the hierarchical decision tree it is on, recognizing cascading state dependencies across many turns, including (in one setting) needing to trace back and revise an earlier choice after failure.
structure: hierarchical · goals arrive: emitted-by-environment-over-time · count: narrative dataset with 311 scene nodes a · subgoal credit: yes · 13 citations · also Personal assistants &amp; long-term memory · unit: other:narrative-nodes
Horizonnarrative dataset spans 311 scene nodes and 86 choice nodes, extending from the game's prologue through Chapter 5
the paper’s own words
we propose a novel benchmark framework based on interactive fiction games, featuring dynamically branching storylines with complex reasoning structures. These structures simulate real-world scenarios by requiring LLMs to navigate hierarchical decision trees, where each choice triggers cascading dependencies across multi-turn interactions. Our benchmark emphasizes two distinct settings to test reasoning complexity: one with immediate feedback upon incorrect decisions, and the other requiring models to independently trace back and revise earlier choices after failure.
We construct a narrative dataset based on the interactive fiction game The Invisible Guardian, encompassing 311 scene nodes and 86 choice nodes as captured in our structured JSON format.
TALES2025relevant
TALES: Text Adventure Learning Environment Suite
Goals to track complete puzzle/quest objectives within a synthetic or human-written text-adventure game via sequential decision-making; maintain structured reasoning over the accumulated context history to determine the next best action
Must keep track of The agent must track accumulated world state (inventory, location, prior actions) across a sequential-decision game, determining the next best action via structured reasoning over the context history, with harder levels requiring correctly sustaining this over up to 44 moves.
structure: sequential-chain · goals arrive: given-up-front · count: diverse collection of synthetic and huma · subgoal credit: no · 11 citations · unit: actions
HorizonCookingWorld difficulty level 1 can be solved in 7 moves (max score 3), while level 10 requires 44 moves (max score 11)
the paper’s own words
We introduce TALES, a diverse collection of synthetic and human-written text-adventure games designed to challenge and evaluate diverse reasoning capabilities. We present results over a range of LLMs, open- and closed-weights, performing a qualitative analysis on the top performing models. Despite an impressive showing on synthetic games, even the top LLM-driven agents fail to achieve 15% on games designed for human enjoyment.
Difficulty level 1 can be solved in 7 moves with a max score of 3, while level 10 requires 44 moves with a max score of 11.
TextQuests: How Good are LLMs at Text-Based Video Games?
Goals to track solve multi-puzzle interactive-fiction adventures with inventory/location/puzzle dependencies via trial-and-error; operate autonomously using only intrinsic long-context reasoning with no external tools; sustain self-directed reasoning across a long, growing context within a single interactive session
Must keep track of Agent must track inventory, location, and puzzle-dependency state purely via long-context reasoning across up to hundreds of precise actions within a single, continuous interactive session, without external tool assistance.
structure: DAG-with-precedence · goals arrive: given-up-front · subgoal credit: unclear · 11 citations · unit: actions
Horizonhundreds of actions (human playtime over 30 hours)
the paper’s own words
we introduce TextQuests, a benchmark based on the Infocom suite of interactive fiction games. These text-based adventures, which can take human players over 30 hours and require hundreds of precise actions to solve, serve as an effective proxy for evaluating AI agents on focused, stateful tasks.
VideoGameBench2025relevant
VideoGameBench: Can Vision-Language Models complete popular video games?
Goals to track complete each of 10 popular 1990s video games end-to-end from raw visual input and high-level objective/control descriptions; generalize to 3 secret/unseen games not disclosed in advance; operate under real-time inference-latency constraints (or in a paused Lite setting)
Must keep track of Agent must track in-game state (perception, spatial navigation, memory) purely from raw visual input across an entire game playthrough, without game-specific scaffolding or auxiliary information.
structure: set-of-independent · goals arrive: given-up-front · count: 10 popular video games (3 kept secret) · subgoal credit: no · 23 citations
Horizonthe paper states none
the paper’s own words
we introduce VideoGameBench, a benchmark consisting of 10 popular video games from the 1990s that VLMs directly interact with in real-time. VideoGameBench challenges models to complete entire games with access to only raw visual inputs and a high-level description of objectives and controls, a significant departure from existing setups that rely on game-specific scaffolding and auxiliary information. We keep three of the games secret to encourage solutions that generalize to unseen environments.
The best performing models, Gemini 2.5 Pro and Claude 3.7 Sonnet, complete only 0.48% of VideoGameBench and 1.6% of VideoGameBench Lite.
WGSR-Bench2025relevant
WGSR-Bench: Wargame-based Game-theoretic Strategic Reasoning Benchmark for Large Language Models
Goals to track achieve environmental situation awareness of a dynamic wargame scenario; model opponent risk accurately; generate a policy/action plan integrating awareness and risk modeling (the S-POE architecture)
Must keep track of Agent must track evolving battlefield/environmental state, model an adversary's likely behavior, and integrate both into policy generation within a single wargame scenario characterized by environmental uncertainty and adversarial dynamics.
structure: hierarchical · goals arrive: given-up-front · count: three core tasks (environmental situatio · subgoal credit: no · 5 citations
Horizonthe paper states none
the paper’s own words
WGSR-Bench designs test samples around three core tasks, i.e., Environmental situation awareness, Opponent risk modeling and Policy generation, which serve as the core S-POE architecture, to systematically assess main abilities of strategic reasoning. Finally, an LLM-based wargame agent is designed to integrate these parts for a comprehensive strategy reasoning assessment.
Wargame, a quintessential high-complexity strategic scenario, integrates environmental uncertainty, adversarial dynamics, and non-unique strategic choices, making it an effective testbed for assessing LLMs' capabilities in multi-agent decision-making, intent inference, and counterfactual reasoning.
WereWolf-Plus2025relevant
WereWolf-Plus: An Update of Werewolf Game setting Based on DSGBench
Goals to track role-specific deduction/elimination objectives (e.g. werewolves eliminate villagers, Seer identifies werewolves, Witch/Hunter/Guard/Sheriff use special role abilities) tracked across the game; sustain social influence/cooperation or deception consistently across repeated day/night rounds
Must keep track of Each agent must track the current game phase (day/night), revealed information, its own and others' inferred roles, and prior votes/eliminations across the game's rounds, adapting its strategy (cooperation, deception, or deduction) accordingly.
structure: sequential-chain · goals arrive: given-up-front · count: customizable roles (Seer, Witch, Hunter, · subgoal credit: yes · 0 citations · also Multi-agent organisations &amp; societies
Horizonthe paper states none
the paper’s own words
we propose WereWolf-Plus, a multi-model, multi-dimensional, and multi-method benchmarking platform for evaluating multi-agent strategic reasoning in the Werewolf game. The platform offers strong extensibility, supporting customizable configurations for roles such as Seer, Witch, Hunter, Guard, and Sheriff, along with flexible model assignment and reasoning enhancement strategies for different roles. In addition, we introduce a comprehensive set of quantitative evaluation metrics for all special roles, werewolves, and the sheriff, and enrich the assessment dimensions for agent reasoning ability, cooperation capacity, and social influence.
lmgame-Bench2025relevant
lmgame-Bench: How Good are LLMs at Playing Games?
Goals to track complete platformer-game objectives requiring perception and timing; solve puzzle-game objectives requiring planning; progress narrative-game objectives requiring memory of prior game state; operate reliably despite brittle vision perception, prompt sensitivity, and potential data contamination
Must keep track of The agent must track in-game state (via lightweight perception and memory scaffolds) appropriate to each game's genre -- platformer timing/position, puzzle constraint state, or narrative progress -- delivered through a unified Gym-style API.
structure: set-of-independent · goals arrive: given-up-front · subgoal credit: no · 40 citations · unit: other:not-stated
Horizonthe paper states none
the paper’s own words
LMGame-Bench features a suite of platformer, puzzle, and narrative games delivered through a unified Gym-style API and paired with lightweight perception and memory scaffolds, and is designed to stabilize prompt variance and remove contamination. Across 13 leading models, we show lmgame-Bench is challenging while still separating models well. Correlation analysis shows that every game probes a unique blend of capabilities often tested in isolation elsewhere.

Tool & API use 29 artifacts · 9 state a horizon

APIFlow-Bench2026relevant
APIFlow-Bench: Measuring Whether Agents Survive Long, Dependent API Workflows
Goals to track complete each subtask in a long, dependent chain of REST-API calls correctly (state, not just completion, must be right); produce a final answer whose delivery is traceable to the actual call path (provenance-sensitive correctness); avoid compounding failures across increasingly long dependency chains (up to 20 subtasks)
Must keep track of Agent must track state correctness at each subtask in a dependency chain (not just whether the chain 'completed'), including whether a mock-minted canary correctly propagates through the API data flow to the final delivered answer.
structure: DAG-with-precedence · goals arrive: given-up-front · count: up to 20 subtasks per chain (across 19 e · subgoal credit: yes · 0 citations · unit: other:subtasks-per-workflow-chain
Horizon20 (clean chains); up to 44,362 execution transcripts released
the paper’s own words
We introduce APIFlow-Bench, a fully auditable benchmark for long-horizon, dependent REST-API workflows that decomposes performance into seven engineering capabilities and requires agents to produce answers supported by the actual call path.
longer dependency chains degrade success, from 93% on individual subtasks to 74% on clean 20-subtask chains and 61% when including the 8% of chain trials that a model-consensus screen flags as passed by no model
AgentEscapeBench: Evaluating Out-of-Domain Tool-Grounded Reasoning in LLM Agents
Goals to track invoke real external tool functions to satisfy a directed acyclic dependency graph over tools and items; track hidden state revealed incrementally through tool use; propagate intermediate results correctly across dependent tool calls to a deterministically verifiable final answer
Must keep track of Agent must track hidden state revealed incrementally by tool calls, maintain clue adherence, and correctly propagate intermediate results through the dependency graph as depth increases.
structure: DAG-with-precedence · goals arrive: emitted-by-environment-over-time · count: 270 instances across five difficulty tie · subgoal credit: yes · 0 citations · unit: other:dependency-graph-depth-tiers
Horizondifficulty tiers from 5 to 25 (dependency-graph depth levels)
the paper’s own words
Each task defines a directed acyclic dependency graph over tools and items, requiring agents to invoke real external functions, track hidden state revealed incrementally, propagate intermediate results, and submit a deterministically verifiable final answer.
Experiments with sixteen LLM agents and human participants show that performance drops sharply as dependency depth increases: humans decline from 98.3% success at difficulty-5 to 80.0% at difficulty-25, while the best model drops from 90.0% to 60.0%.
AgentFloor2026relevant
AgentFloor: How Far Up the tool use Ladder Can Small Open-Weight Models Go?
Goals to track follow instructions correctly (lowest tier); use tools correctly (mid-lower tier); coordinate multiple steps/tool calls together (mid-upper tier); sustain long-horizon planning under persistent constraints over many steps (top tier)
Must keep track of Agent must sustain constraint tracking and coordination reliably over many steps to succeed at the highest tiers, where 'neither side reaches strong reliability' even among frontier models.
structure: hierarchical · goals arrive: given-up-front · count: 30 tasks organized into a six-tier capab · subgoal credit: unclear · 0 citations
Horizonthe paper states none
the paper’s own words
We introduce AgentFloor, a deterministic 30-task benchmark organized as a six-tier capability ladder, spanning instruction following, tool use, multi-step coordination, and long-horizon planning under persistent constraints.
We evaluate 16 open-weight models, from 0.27B to 32B parameters, alongside GPT-5 across 16,542 scored runs
AgentGym22026relevant
AgentGym2: Benchmarking Large Language Model Agents in De-Idealized Real-World Environments
Goals to track execute end-to-end real-world procedures without relying on pre-packaged tool interfaces; discover available tools via active exploration of the environment; compose discovered tools to solve previously unseen tasks; remain robust to noisy and underspecified task information
Must keep track of Agent must track which tools/interfaces it has discovered so far, what it has verified about their (possibly noisy) behavior, and how these compose toward completing the current end-to-end task.
structure: open-ended · goals arrive: implied-by-constraints · subgoal credit: yes · 1 citations
Horizonthe paper states none
the paper’s own words
Beyond reasoning and planning, it measures agents' ability to execute end-to-end procedures, discover tools via exploration, compose tools for unseen tasks, and remain robust to noisy and underspecified information.
To bridge this gap, we present AgentGym2, a new evaluation framework with task instances grounded in real-world end-to-end working demands.
AppWorld-UL2026relevant
AppWorld-UL: Benchmarking Diverse Agent-User Interactions for Tool-Use
Goals to track operate applications correctly to complete a digital task (e.g. ordering groceries) across 9 simulated apps; interact appropriately with the user: ask clarification questions, prompt for confirmation, or report infeasibility; succeed on compositional, multi-part sub-tasks that combine several of the above interaction types
Must keep track of Agent must track what has been asked/confirmed with the user so far, the user's carefully-bounded knowledge state (simulated by an LLM), and which parts of a compositional task remain to be completed.
structure: hierarchical · goals arrive: mixed:given-up-front-task-with-user-clarificat · count: 516 tasks, built on 9 simulated apps (e. · subgoal credit: yes · 1 citations · also Personal assistants &amp; long-term memory
Horizonthe paper states none
the paper’s own words
we introduce AppWorld-UL, a ``user-in-the-loop'' benchmark of 516 challenging tasks requiring diverse agent-user interactions.
Building upon the AppWorld framework with 9 popular simulated apps like Amazon and Spotify
AsyncTool2026in
AsyncTool: Evaluating the Asynchronous Function Calling Capability under Multi-Task Scenarios
Goals to track concurrently manage multiple heterogeneous tasks presented simultaneously; make productive use of idle time while awaiting delayed tool-call responses (asynchronous tool calling); coordinate task switching, dependency tracking, and state maintenance across concurrently running tasks
Must keep track of Agent must track the state, dependencies, and pending tool responses of multiple simultaneously active tasks, and coordinate task-switching decisions during periods of delayed tool feedback.
structure: DAG-with-precedence · goals arrive: given-up-front · subgoal credit: yes · 3 citations
Horizonthe paper states none
the paper’s own words
AsyncTool presents multiple heterogeneous tasks simultaneously and simulates realistic tool response latency during execution. Using a hybrid data evolution strategy, we construct a diverse asynchronous multitasking dataset that covers multiple scenarios and tool-use patterns. We evaluate models at the step, sub-task, and task levels, and introduce efficiency-oriented metrics to measure task coordination and completion efficiency.
Benchmarking AI Agents for Hardware Design Automation via MCP Tool Calling
Goals to track execute single-operation edits on hardware design components via specialised MCP tools; correctly sequence multi-step dependency chains (e.g. create component -> add port -> wire connection); handle invalid or misspelled requests without incorrect tool invocation; operate correctly across multi-server tool contexts
Must keep track of The agent must track which prior dependency-establishing tool calls (e.g. component creation, port creation) have already succeeded before issuing calls that depend on them, across single-agent and multi-agent tool-calling configurations.
structure: DAG-with-precedence · goals arrive: given-up-front · count: 1 expected call (easy) / 2 (medium) / 3- · subgoal credit: yes · 0 citations · unit: tool-calls
Horizon1 (Easy) / 2 (Medium) / 3-5 (Hard) expected tool calls per task
the paper’s own words
engineers issue repetitive, dependency-ordered operations---such as creating components, adding ports, and wiring connections---through specialised tools. ... we build a Model Context Protocol (MCP) server ... and construct a benchmark covering single-operation edits, multi-step dependency chains, invalid requests, misspelled prompts, and multi-server tool contexts.
Easy, Medium, and Hard contain 40 self-contained tasks each ... Easy requires 1 expected call, Medium requires 2, and Hard requires 3-5 calls.
C-World2026in
C-World: A Computer Use Agent Environment Creator
Goals to track complete long-horizon workflows composed of many interacting constraints across up to 5,571 tools spanning 204 applications; correctly follow constraints despite injected realistic failures and perturbations during the task; satisfy a reward signal combining verifiable metrics with LLM-based judgment
Must keep track of Agent must track constraint satisfaction across a long-horizon, multi-tool workflow while detecting and adapting to injected failures/perturbations introduced by the environment's transition function.
structure: DAG-with-precedence · goals arrive: emitted-by-environment-over-time · subgoal credit: yes · 1 citations
Horizonthe paper states none
the paper’s own words
We define a complete agent environment through four components: an Action Space of 5,571 format-unified tools across 204 common applications, a Task Distribution engine that synthesizes long-horizon workflows with wild constraints, a Transition Function implemented as a state controller that injects realistic failures and perturbations, and a Reward Signal combining verifiable metrics with LLM-based judgment.
CCTU2026relevant
CCTU: A Benchmark for Tool Use under Complex Constraints
Goals to track select and call the correct tool while satisfying every one of several simultaneous constraints (resource, behavior, toolset, response dimensions); maintain compliance with all constraints across a multi-turn interaction, not just the first tool call; self-refine after receiving feedback about a constraint violation
Must keep track of The agent must track which of the ~7 simultaneous constraints (out of 12 categories across 4 dimensions) apply to the current tool-use scenario and re-check compliance with all of them at every step across the multi-turn interaction, especially after receiving feedback about a violation.
structure: set-of-independent · goals arrive: given-up-front · count: 200 test cases; 12 constraint categories · subgoal credit: yes · 7 citations · unit: turns
Horizonmaximum 20 interaction rounds per test case
the paper’s own words
CCTU is grounded in a taxonomy of 12 constraint categories spanning four dimensions (i.e., resource, behavior, toolset, and response). The benchmark comprises 200 carefully curated and challenging test cases across diverse tool-use scenarios, each involving an average of seven constraint types and an average prompt length exceeding 4,700 tokens. To enable reliable evaluation, we develop an executable constraint validation module that performs step-level validation and enforces compliance during multi-turn interactions between models and their environments.
maximum 20 interaction rounds
GTA-2: Benchmarking General Tool Agents from Atomic Tool-Use to Open-Ended Workflows
Goals to track execute short-horizon, closed-ended atomic tool calls correctly (GTA-Atomic); complete long-horizon, open-ended, real-world productivity workflows end-to-end (GTA-Workflow); satisfy verifiable sub-goals identified by a recursive checkpoint-based evaluation mechanism
Must keep track of System must track which recursively-decomposed sub-goals/checkpoints of an open-ended workflow have been satisfied so far, across real deployed tools and multimodal contexts.
structure: hierarchical · goals arrive: given-up-front · subgoal credit: yes · 1 citations · also Personal assistants &amp; long-term memory
Horizonthe paper states none
the paper’s own words
we propose a recursive checkpoint-based evaluation mechanism that decomposes objectives into verifiable sub-goals, enabling unified evaluation of both model capabilities and agent execution frameworks
(ii) GTA-Workflow introduces long-horizon, open-ended tasks for realistic end-to-end completion
GeoAgentBench: A Dynamic Execution Benchmark for Tool-Augmented Agents in Spatial Analysis
Goals to track correctly configure parameters for each of 117 atomic GIS tools invoked; complete each of 53 typical spatial-analysis tasks across 6 core GIS domains; produce spatially/cartographically accurate outputs verified via a VLM-based check; decouple global workflow orchestration from step-wise reactive execution to recover from runtime anomalies
Must keep track of The agent must track its evolving execution plan, per-step parameter choices, and runtime feedback/anomalies across a multi-step GIS workflow to keep global orchestration consistent with step-wise reactive execution.
structure: sequential-chain · goals arrive: given-up-front · count: 117 atomic GIS tools; 53 spatial-analysi · subgoal credit: yes · 1 citations
Horizonthe paper states none
the paper’s own words
GABench provides a realistic execution sandbox integrating 117 atomic GIS tools, encompassing 53 typical spatial analysis tasks across 6 core GIS domains.
Ko-WideSearch2026relevant
Ko-WideSearch: A Korean Breadth-Search Benchmark for Exhaustive Set Enumeration by Web Agents
Goals to track exhaustively enumerate the full membership of a named closed set (e.g. a TV season's cast, a dynasty's rulers); fill a per-item attribute table (multiple columns) for every enumerated member; decide when to stop searching within an open-ended web search space
Must keep track of The agent must track which set members it has already found (to avoid duplicates/omissions), which attribute columns remain unfilled per member, and its remaining search-iteration budget as difficulty knobs (table width, 2-D composite key) increase.
structure: set-of-independent · goals arrive: given-up-front · count: 228 tables over 190 entities across sixt · subgoal credit: no · 0 citations · also Web &amp; GUI agents · unit: tool-calls
Horizonfixed budget of 30 agent iterations per question (total tool calls run higher; one model logged up to 947 tool calls)
the paper’s own words
Each task names a set-parent entity ... and asks for its full membership plus a per-item attribute table, graded by Item-, Column-, and Row-F1. It spans 228 tables over 190 entities and sixteen categories across three difficulty tiers ... Across twenty web agents, the failure is consistent: agents recover the set but not the rows (e.g. Item-F1 92.8 against Row-F1 53.7)
each evaluated model receives ... a fixed per-question budget of thirty agent iterations, with the clarification that each iteration may batch several tool calls, so the total search-call count per task runs well above thirty. ... Qwen3.6-35B ... run 947 tool calls, the run maximum
TOBench: A Task-Oriented Omni-Modal Benchmark for Real-World Tool-Using Agents
Goals to track execute tools appropriate to a Customer Service or Intelligent Creation task; inspect rendered or transformed intermediate artifacts produced by tool calls; self-correct when an inspected artifact fails task-specific requirements; satisfy each of the 20 subcategory slices' task-specific grounded evaluator checks
Must keep track of The agent must track the state of rendered/transformed artifacts across a closed verification loop, deciding when self-correction is required, using task-specific grounded evaluators across 27 MCP servers and 324 tools.
structure: sequential-chain · goals arrive: given-up-front · count: 100 executable tasks across 20 subcatego · subgoal credit: no · 0 citations · unit: other:not-stated
Horizonthe paper states none
the paper’s own words
The central design of MM-ToolBench is closed-loop multimodal verification: agents must execute tools, inspect rendered or transformed artifacts, and self-correct when outputs fail task-specific requirements. To make such evaluation scalable and verifiable, MM-ToolBench couples MCP-based execution with task-specific grounded evaluators
MM-ToolBench contains 100 executable tasks from two macro task families, Customer Service and Intelligent Creation, covering 20 subcategory slices and supported by 27 MCP servers with 324 tools.
MM-ToolSandBox: A Unified Framework for Evaluating Visual Tool-Calling Agents
Goals to track ground progressively arriving visual inputs into correct executable tool calls across a multi-image, multi-turn interaction; handle realistic conversational phenomena mid-task: goal revisions, error corrections, state mutations; operate correctly across 500+ tools spanning 16 application domains; succeed on each of 258 human-verified nominal scenarios (plus 50 interactive-UI variants)
Must keep track of The agent must track a stateful execution environment across multi-image, multi-turn interactions, correctly incorporating goal revisions, error corrections, and state mutations as they occur, and ground each newly arriving visual input into the correct tool call.
structure: sequential-chain · goals arrive: mixed:given-up-front-with-goal-revisions-injec · count: 258 human-verified nominal scenarios plu · subgoal credit: no · 0 citations
Horizonthe paper states none
the paper’s own words
Evaluating 12 state-of-the-art models, from 4B open-weight to frontier proprietary systems, shows that current models still lack robust visual tool-calling capability: even the best model achieves below 50% success rate. Our failure analysis further reveals that visual precision, not only planning, is a primary bottleneck for capable models: 53% of failures stem from incorrect information extraction from images despite otherwise correct task workflows.
supporting multi-image, multi-turn tasks where agents must ground progressively arriving visual inputs into executable tool calls
OmnilingualGAIA22026relevant
OmnilingualGAIA2: Evaluating the Multilingual Gap in Frontier AI Agents
Goals to track plan a sequence of tool calls to answer a task; search for information via tools; execute multi-tool workflows; recover from errors during multi-tool execution; all under machine-translated, human-calibrated task instructions across ten languages/five scripts
Must keep track of Agent must track its tool-call plan, intermediate search/tool results, and error states to recover from execution failures, now while also handling machine-translated instructions across ten languages including non-Latin scripts.
structure: DAG-with-precedence · goals arrive: given-up-front · subgoal credit: yes · 0 citations
Horizonthe paper states none
the paper’s own words
Agentic benchmarks aim to measure how well AI agents plan, search, execute, and recover within realistic multi-tool environments, but they are almost exclusively in English.
Evaluating seven frontier and open-weight agents, we find a universal cross-lingual gap of 8.8-18.4 pass@3 points that is agent-asymmetric in magnitude, concentrates on tool-orchestration rather than quantitative reasoning, and does not close with model scale.
SeekerGym2026relevant
SeekerGym: A Benchmark for Reliable Information Seeking
Goals to track issue repeated retrieval queries to recover as much of a target document's content as possible; quantify uncertainty about how much relevant information might still be missing from what has been retrieved
Must keep track of The agent must track which passages/sections of the target document it has already retrieved, so it can judge how complete its coverage is and estimate how much information might still be missing.
structure: open-ended · goals arrive: given-up-front · subgoal credit: yes · 0 citations
Horizonthe paper states none
the paper’s own words
each task in SeekerGym is a document (e.g., a Wikipedia article), and the AI agent must issue queries to retrieve passages from that document... the best approaches retrieve 42.5% of passages on Wikipedia and 29.2% on ML Surveys, leaving substantial room for improvement.
each episode runs for at most M steps (finite horizon) ... under a fixed budget of M steps with K queries each
SkillCraft2026relevant
SkillCraft: Can LLM Agents Learn to Use Tools Skillfully?
Goals to track compose atomic tools into reusable higher-level 'Skills'; cache and reuse learned Skills both within a task and across different tasks; complete compositional tool-use scenarios whose difficulty scales with entity count and subtask complexity
Must keep track of The agent must track which higher-level Skills it has already formed and cached, so it can reuse them instead of recomposing atomic tools from scratch on later subtasks/tasks, both within a single task and across the benchmark's 126 tasks.
structure: hierarchical · goals arrive: given-up-front · count: 126 tasks across 6 difficulty levels, co · subgoal credit: yes · 29 citations · unit: tool-calls
Horizon126 tasks across 6 difficulty levels (from 21 seed tasks); tool-call counts scale from 9 (Easy: 3 subtasks x 3 calls) to 25 (Hard: 5 subtasks x 5 calls)
the paper’s own words
SkillCraft features realistic, highly compositional tool-use scenarios with difficulty scaled along both quantitative and structural dimensions, designed to elicit skill abstraction and cross-task reuse. We further propose a lightweight evaluation protocol that enables agents to auto-compose atomic tools into executable Skills, cache and reuse them inside and across tasks.
3 subtasks x3 API calls = 9 total calls ... 4x4=16 ... 5x5=25 ... 126 tasks across 6 difficulty levels
The Amazing Agent Race: Strong Tool Users, Weak Navigators
Goals to track navigate Wikipedia pages to locate the required entity/fact for each DAG node; execute the correct multi-step tool chain across fork-merge branches linking extracted entities; aggregate branch outputs into one verifiable final answer
Must keep track of The agent must track which Wikipedia entities/facts it has already extracted at each DAG node, correctly route them through parallel fork-merge tool-chain branches, and retain intermediate values needed for final aggregation across up to 5 diamonds per leg.
structure: DAG-with-precedence · goals arrive: given-up-front · count: 1,400 instances: 800 sequential legs and · subgoal credit: yes · 1 citations · also Embodied &amp; robotics · unit: other:pit-stops (navigation hops)
Horizonaverage of 22.1 pit stops per leg for AAR-DAG (600 legs) and 15.0 pit stops per leg for AAR-Linear (800 legs), up to 5 diamonds per leg
the paper’s own words
Three complementary metrics (finish-line accuracy, pit-stop visit rate, and roadblock completion rate) separately diagnose navigation, tool-use, and arithmetic failures.
with an average of 22 pit stops and up to 5 diamonds
ToolGym2026in
ToolGym: an Open-world Tool-using Environment for Scalable Agent Testing and Data Curation
Goals to track complete long-horizon, multi-tool workflows synthesized with wild constraints across 5,571 tools and 204 apps; recover and adapt when a state controller injects interruptions and failures mid-workflow; separate deliberate planning/self-correction from step-wise execution via a planner-actor decomposition
Must keep track of Agent must track tool-call state and wild constraints across a long-horizon, multi-tool workflow while detecting and recovering from injected interruptions/failures and unreliable tool states.
structure: sequential-chain · goals arrive: mixed:given-up-front-with-injected-interruptio · count: 5,571 format-unified tools across 204 co · subgoal credit: no · 5 citations
Horizonthe paper states none
the paper’s own words
It includes a task creation engine that synthesizes long-horizon, multi-tool workflows with wild constraints, and a state controller that injects interruptions and failures to stress-test robustness. On top of this environment, we develop a tool select-then-execute agent framework with a planner-actor decomposition to separate deliberate reasoning and self-correction from step-wise execution.
ToolVerse: Unlocking Massive Environments and Long-Horizon Tasks for Agentic Reinforcement Learning
Goals to track complete a long-horizon task built from a tool-dependency graph, where later tools unlock only after prerequisite tool-use subgoals are completed; correctly integrate the right subset of tools from a large pool (~4500 tools across ~400 MCPs) into a coherent multi-step solution; receive properly assigned turn-level credit despite long sequences of tool calls
Must keep track of The agent must track which prerequisite tools/subgoals in the dependency graph it has already satisfied in order to know which further tools are unlocked and relevant next, across a long sequence of tool calls.
structure: DAG-with-precedence · goals arrive: emitted-by-environment-over-time · count: tasks built from ~4500 tools across ~400 · subgoal credit: yes · 2 citations
Horizonthe paper states none
the paper’s own words
we propose a task design strategy based on a tool dependency graph, utilizing Dynamic Unlocking Sampling Algorithm to generate long-horizon tasks, and produce GUST (Graph Unlocking Sampling Tasks) dataset.
tasks consist of 3 to 7 tasks per data item
Capable but Unreliable: Canonical Path Deviation as a Causal Mechanism of Agent Failure in Long-Horizon Tasks
Goals to track complete each of 108 real-world tool-use tasks by following a canonical multi-step tool-invocation solution path; stay within the operating envelope of the canonical path across the whole trajectory to avoid stochastic drift/derailment
Must keep track of Success requires the trajectory of tool calls to stay within the operating envelope of the task's canonical solution path; mid-trajectory adherence must be tracked since drift compounds over subsequent calls.
structure: sequential-chain · goals arrive: given-up-front · count: 108 real-world tool-use tasks; 22 models · subgoal credit: yes · 3 citations
Horizonthe paper states none
the paper’s own words
We analyze trajectories from the Toolathlon benchmark: 22 frontier models each attempt 108 real-world tool-use tasks across 3 independent runs, yielding 515 model$\times$task units where the same model succeeds on some runs and fails on others due to LLM sampling stochasticity alone.
UniClawBench2026relevant
UniClawBench: A Universal Benchmark for Proactive Agents on Real-World Tasks
Goals to track demonstrate correct skill usage for the task's required tool/capability; explore the live environment (e.g. a Docker container) to discover needed information/actions; reason correctly over long context accumulated during the task; correctly interpret multimodal inputs relevant to the task; coordinate actions correctly across multiple platforms/services
Must keep track of The agent must track its progress against fine-grained, step-by-step completion checkpoints within a live environment, integrate multi-turn feedback from a hidden supervisor agent and a user-simulator agent without seeing the grading criteria, and coordinate across the relevant capability dimensions needed for that task.
structure: sequential-chain · goals arrive: given-up-front · count: 400 bilingual real-world tasks organized · subgoal credit: yes · 1 citations
Horizonthe paper states none
the paper’s own words
Unlike previous benchmarks that rely on static, pre-recorded answers, our benchmark evaluates agents in live Docker containers using fine-grained, step-by-step completion checkpoints. Furthermore, we design a closed-loop evaluation strategy comprising an executor agent, a hidden supervisor agent, and a user agent to simulate realistic multi-turn human feedback without leaking grading criteria.
CONFETTI2025relevant
CONFETTI: Conversational Function-Calling Evaluation Through Turn-Level Interactions
Goals to track handle user follow-up requests within an ongoing conversation; correct or switch goals mid-conversation when the user changes intent; resolve ambiguous or implicit user goals into an appropriate API call; chain multiple function calls together to satisfy a multi-step user request
Must keep track of The agent must track evolving user intent across conversation turns, recognize goal correction/switching, and maintain state of prior chained function-call results across up to 86 available APIs.
structure: sequential-chain · goals arrive: mixed:given-up-front-with-user-injected-goal-c · subgoal credit: yes · 16 citations · unit: turns
Horizon313 user turns across 109 conversations (~2.9 turns/conversation on average)
the paper’s own words
These conversations explicitly target various conversational complexities, such as follow-ups, goal correction and switching, ambiguous and implicit goals. We perform off-policy turn-level evaluation using this benchmark targeting function-calling.
CONFETTI addresses this gap through 109 human-simulated conversations1, comprising 313 user turns and covering 86 APIs.
Rethinking Stateful Tool Use in Multi-Turn Dialogues: Benchmarks and Challenges
Goals to track create a new tool/API on demand within a multi-turn dialogue (tool creation); become aware of, correctly select, and execute the right tool given user intent (tool utilization); generate a role-consistent response, including role play, reflecting whether/how the tool was used
Must keep track of Agent must track which tools have been created/exist, their stateful execution history across the dialogue, and how to remain role-consistent while using them, across a multi-turn conversation.
structure: sequential-chain · goals arrive: given-up-front · count: six key tasks across three stages (tool · subgoal credit: yes · 17 citations · also Web &amp; GUI agents
Horizonthe paper states none
the paper’s own words
we propose \texttt{DialogTool}, a multi-turn dialogue dataset with stateful tool interactions considering the whole life cycle of tool use, across six key tasks in three stages: 1) \textit{tool creation}; 2) \textit{tool utilization}: tool awareness, tool selection, tool execution; and 3) \textit{role-consistent response}: response generation and role play.
revealing that the existing state-of-the-art LLMs still cannot perform well to use tools over long horizons.
M^3-Bench2025relevant
M^3-Bench: Multi-Modal, Multi-Hop, Multi-Threaded Tool-Using MLLM Agent Benchmark
Goals to track complete multi-hop, multi-threaded tool-call workflows requiring cross-tool dependencies; maintain persistence of intermediate resources across steps; ground visual and textual reasoning correctly across each tool call
Must keep track of Agent must track intermediate resources produced by earlier tool calls and align them across multiple concurrent threads for argument fidelity and structural consistency.
structure: DAG-with-precedence · goals arrive: given-up-front · count: 28 servers with 231 tools; per-task subg · subgoal credit: yes · 9 citations
Horizonthe paper states none
the paper’s own words
The benchmark targets realistic, multi-hop and multi-threaded workflows that require visual grounding and textual reasoning, cross-tool dependencies, and persistence of intermediate resources across steps.
The benchmark spans 28 servers with 231 tools, and provides standardized trajectories curated through an Executor&Judge pipeline with human verification.
Multi-Mission Tool Bench: Assessing the Robustness of LLM based Agents through Related and Dynamic Missions
Goals to track complete each mission within a test case containing multiple interrelated missions; dynamically adapt when missions switch mid-interaction; correctly invoke tools appropriate to the currently active mission; handle all possible mission-switching patterns within a fixed mission number
Must keep track of The agent must track the state of each interrelated mission (active or paused), correctly recognize mission switches, and select/invoke the correct tools for whichever mission is currently active, evaluated via dynamic decision trees for accuracy and efficiency.
structure: DAG-with-precedence · goals arrive: injected-by-user-mid-episode · count: multiple interrelated missions per test · subgoal credit: no · 7 citations
Horizonthe paper states none
the paper’s own words
In the benchmark, each test case comprises multiple interrelated missions. This design requires agents to dynamically adapt to evolving demands. Moreover, the proposed benchmark explores all possible mission-switching patterns within a fixed mission number.
OrchDAG2025relevant
OrchDAG: Complex Tool Orchestration in Multi-Turn Interactions with Plan DAGs
Goals to track correctly execute a sequence of tool calls whose dependencies form a directed acyclic graph (DAG); respect topological/precedence constraints among tool calls across multi-turn interactions; solve DAG-structured tool-orchestration tasks of controllable/varying complexity
Must keep track of The agent must track which nodes (tool calls) in the DAG have been completed and what state/outputs they produced, to correctly select and sequence subsequent tool calls consistent with the graph's precedence constraints across multiple turns.
structure: DAG-with-precedence · goals arrive: given-up-front · subgoal credit: no · 1 citations
Horizonthe paper states none
the paper’s own words
We introduce OrchDAG, a synthetic data generation pipeline that models tool execution as directed acyclic graphs (DAGs) with controllable complexity. Using this dataset, we benchmark model performance and propose a graph-based reward to enhance RLVR training.
yet most existing work overlooks the complexity of multi-turn tool interactions
The Tool Decathlon: Benchmarking Language Agents for Diverse, Realistic, and Long-Horizon Task Execution
Goals to track coordinate interactions across multiple named Apps/tools (e.g. email + calendar + file systems) to complete one complex workflow; diagnose and report anomalies via monitoring a database following an operating manual; satisfy a strictly execution-verifiable end state for each of 108 tasks spanning 32 apps and 604 tools
Must keep track of The agent must track the current, realistic state of multiple named applications across roughly 20 tool-calling turns per task, and verify its final state against a dedicated evaluation script.
structure: DAG-with-precedence · goals arrive: given-up-front · count: 108 tasks across 32 apps and 604 tools; · subgoal credit: yes · 66 citations · also Business, office &amp; enterprise work · unit: turns
Horizontasks require interacting with multiple Apps over around 20 turns on average (best model averages 20.2 tool-calling turns)
the paper’s own words
Toolathlon spans 32 software applications and 604 tools, ranging from everyday platforms such as Google Calendar and Notion to professional ones like WooCommerce, Kubernetes, and BigQuery... This benchmark includes 108 manually sourced or crafted tasks in total, requiring interacting with multiple Apps over around 20 turns on average to complete. Each task is strictly verifiable through dedicated evaluation scripts.
ToolHaystack2025relevant
ToolHaystack: Stress-Testing Tool-Augmented Language Models in Realistic Long-Term Interactions
Goals to track correctly maintain and disambiguate multiple concurrent task-execution contexts within one continuous conversation; handle realistic noise/disruptions injected into a long-term interaction without losing track of tool-use context; successfully complete tool-use tasks embedded within a long, continuous conversation despite these challenges
Must keep track of The model must maintain and disambiguate multiple concurrent task-execution contexts across a continuous long-term conversation while filtering out realistic noise and handling various disruptions, rather than resetting context between short, isolated tool-use exchanges.
structure: set-of-independent · goals arrive: emitted-by-environment-over-time · subgoal credit: no · 5 citations
Horizonthe paper states none
the paper’s own words
By applying this benchmark to 14 state-of-the-art LLMs, we find that while current models perform well in standard multi-turn settings, they often significantly struggle in ToolHaystack, highlighting critical gaps in their long-term robustness not revealed by previous tool benchmarks.
offering limited insight into model behavior during realistic long-term interactions

OS & computer use 24 artifacts · 7 state a horizon

CarePilot: A Multi-Agent Framework for Long-Horizon Computer Task Automation in Healthcare
Goals to track complete each of 8-24 consecutive GUI decisions/actions required by one long-horizon clinical software workflow (e.g. DICOM viewer, EHR, lab information system); correctly ground next-action predictions in the current visual interface and system state throughout the workflow
Must keep track of The agent must track the current visual interface state and prior decisions across 8-24 consecutive steps of a clinical software workflow, using dual-memory (long-term and short-term experience) to predict the next semantic action correctly.
structure: sequential-chain · goals arrive: given-up-front · count: each task comprises 8-24 consecutive dec · subgoal credit: yes · 6 citations · also Healthcare &amp; clinical · unit: agent-steps
Horizoneach task comprises 8-24 consecutive decisions/steps
the paper’s own words
we introduce CareFlow, a high-quality human-annotated benchmark comprising complex, long-horizon software workflows across medical annotation tools, DICOM viewers, EHR systems, and laboratory information systems.
each task pairs a natural-language goal with GUI screenshots representing authentic clinical workflows... 8-24 consecutive decisions
ChainWorld2026relevant
ChainWorld: Composing Long-Horizon Desktop Workloads from Atomic OSWorld Tasks
Goals to track complete a chain of 2 to 4 sequentially composed atomic OSWorld desktop tasks while sustaining state across all of them; under single-turn evaluation, complete the whole chain from one combined prompt; under multi-turn evaluation, complete each task as it is revealed one at a time while retaining session state
Must keep track of Agent must sustain desktop application/file state across multiple chained objectives and, in multi-turn mode, manage session continuity as tasks are revealed one at a time.
structure: sequential-chain · goals arrive: given-up-front · count: 347 chains, each of length two to four a · subgoal credit: yes · 0 citations · unit: other:atomic-tasks-per-chain
Horizontwo to four
the paper’s own words
ChainWorld, which composes atomic OSWorld tasks into long horizon desktop workloads through directional compatibility search while preserving the source evaluators. The resulting workload contains 347 chains of length two to four
The resulting workload contains 347 chains of length two to four and compares two renderings of the same task sequence.
CutVerse2026relevant
CutVerse: A Compositional GUI Agents Benchmark for Media Post-Production Editing
Goals to track complete a long-horizon, compositional media-editing task grounded in an authentic editing workflow using one of 7 professional applications (e.g., Premiere Pro, Photoshop); correctly execute tightly coupled, dense multimodal interaction sequences within the application's GUI
Must keep track of Agent must track compositional GUI state (selected layers/clips/tools) across a tightly coupled, dense-interface interaction sequence throughout a long-horizon editing task.
structure: sequential-chain · goals arrive: given-up-front · count: 186 complex, long-horizon tasks across 7 · subgoal credit: yes · 0 citations · also Web &amp; GUI agents
Horizonthe paper states none
the paper’s own words
We curate expert demonstrations across 7 professional applications (e.g., Premiere Pro, Photoshop), covering 186 complex, long-horizon tasks grounded in authentic editing workflows, involving dense multimodal interfaces and tightly coupled interaction sequences.
DeskCraft2026relevant
DeskCraft: Benchmarking Desktop Agents on Professional Workflows and Human-in-the-Loop Collaboration
Goals to track complete long-horizon professional creative/engineering workflows (design, video, audio, 3D creation) requiring over 50 execution steps; proactively seek necessary information from the user under uncertainty (agent-initiated clarification); correctly handle user-initiated interruptions during execution and incorporate post-turn feedback after signaling completion
Must keep track of The agent must track its own multi-step progress (over 50 execution steps for long-horizon tasks) within professional creative/engineering software, while also tracking pending clarification needs, handling user interruptions mid-execution, and incorporating post-completion feedback into further revisions.
structure: sequential-chain · goals arrive: mixed:task-given-upfront-with-human-in-the-loo · count: 538 tasks; a multilevel difficulty taxon · subgoal credit: yes · 2 citations · also Web &amp; GUI agents · unit: agent-steps
Horizonlong-horizon tasks require over 50 execution steps
the paper’s own words
DeskCraft organizes tasks into a multilevel difficulty taxonomy, with long horizon tasks requiring over 50 execution steps, and covers professional creative software across design, video, audio, and 3D creation. Furthermore, DeskCraft formalizes human-agent collaboration into an interaction protocol covering mid-turn and post-turn exchanges. Mid-turn interaction captures both agent-initiated clarification under uncertainty and user-initiated interruption during execution, while post-turn interaction accommodates user-driven feedback after the agent signals completion, together spanning the full space of realistic collaboration patterns.
DevicesWorld2026relevant
DevicesWorld: Benchmarking Cross-Device Agents in Heterogeneous Environments
Goals to track acquire information on one device (e.g. phone) needed to complete a task; process/transform that information on a different device (e.g. desktop); deliver or display the final result on yet another device; satisfy cross-device dependencies and rule-based verifiers across the whole task
Must keep track of The agent must track which device holds which piece of needed information, what has already been acquired/processed/delivered across devices, and whether all task conditions (verified from device states and generated files) are jointly satisfied.
structure: DAG-with-precedence · goals arrive: given-up-front · count: 6,140 tasks integrating 3 device-environ · subgoal credit: yes · 0 citations · unit: actions
Horizoneach task permits at most 50 interaction steps
the paper’s own words
We introduce DevicesWorld, a large-scale executable benchmark for cross-device collaborative operation. DevicesWorld contains 6,140 tasks and integrates three classes of device environments -- mobile, desktop, and IoT -- into a unified cross-device interaction and evaluation framework... Among failed runs, about 28.7% satisfy at least one scoring condition yet still fail the full task.
each task permits at most 50 interaction steps
GUITestScape2026relevant
GUITestScape: Towards Open-set Evaluation on Exploratory GUI Testing
Goals to track autonomously navigate an app to discover defects with no predefined test script; distinguish and separately diagnose interaction defects versus display defects; decompose the testing trajectory into independently diagnosable capabilities rather than a single end-state judgment
Must keep track of The agent must track which parts of an app it has already explored and which candidate defects (interaction or display) it has already investigated, to continue open-ended exploration without redundant re-testing or missing an area.
structure: open-ended · goals arrive: self-generated-by-agent · count: 61 real-world Android applications; 508 · subgoal credit: yes · 0 citations · also Web &amp; GUI agents
Horizonthe paper states none
the paper’s own words
current evaluation falls short on two fronts... evaluation protocols are bound to predefined defect annotations, collapsing the testing process into a single end-state judgment that conflates qualitatively distinct failure modes. To address these challenges, we present GUITestScape, an interactive benchmark covering 61 real-world Android applications and 508 preset defects spanning interaction and display types, and introduce GUIJudge, an open-set evaluator that decomposes an agent's testing trajectory into independently diagnosable capabilities.
HeraBench2026relevant
Beyond Global Replanning: Hierarchical Recovery for Cross-Device Agent Systems
Goals to track decompose a cross-device task and assign subtasks across heterogeneous devices via unified API-CLI-GUI execution; recover from injected device-local strategy failures without escalation; escalate to orchestrator-level global replanning when a failure exceeds device-local recovery scope; complete an end-to-end cross-device workflow over Linux and Android devices
Must keep track of The system must track per-device execution-strategy state, a compact cross-layer failure abstraction distinguishing device-local from global failure scope, and overall workflow progress across Linux and Android devices.
structure: hierarchical · goals arrive: mixed:workflow-given-up-front-failures-injecte · subgoal credit: yes · 0 citations
Horizonthe paper states none
the paper’s own words
we introduce HeraBench, a fault-injected benchmark that constructs cross-device workflows over Linux and Android devices and injects strategy- and device-level failures.
JarvisGUI2026relevant
JarvisGUI: Towards Cross-Device GUI Agents with Dynamic Task Composition
Goals to track transfer intermediate results correctly between two or more heterogeneous devices/platforms; maintain shared state consistently across Android, Windows, and Ubuntu platforms; compose a multi-step, cross-device workflow from input-output-typed GUI sub-tasks; complete a workflow requiring four or more chained subtasks despite compounding dependency-tracking difficulty
Must keep track of The agent must track shared state and intermediate results as they transfer across heterogeneous devices/platforms, verifying that dependency steps (e.g. file transfer, renaming) have actually executed rather than being falsely treated as completed.
structure: DAG-with-precedence · goals arrive: given-up-front · count: dynamically composed workflows; success · subgoal credit: no · 0 citations · also Web &amp; GUI agents · unit: other:not-stated
Horizonthe paper states none
the paper’s own words
JarvisGUI formulates GUI tasks as input-output transformations under a lightweight type system, which allows us to automatically compose multi-step, cross-device workflows and dynamically evaluate agent performance within a unified framework. By evaluating agents in virtual environments spanning multiple operating systems, JarvisGUI reveals that state-of-the-art open-source GUI agents struggle with the state-transfer awareness, cross-platform contextual reasoning, and long-horizon dependency management required for real-world workflows
The task success rate exhibits a sharp decline as the number of subtasks increases, dropping to near zero for tasks requiring four or more subtasks across all evaluated models.
MacAgentBench: Benchmarking AI Agents on Real-World macOS Desktop
Goals to track complete a multi-application task on real macOS desktop software using both GUI and CLI interaction; reach each of several fine-grained, capability-annotated checkpoints within a multi-application task
Must keep track of Agent must track partial progress across multiple checkpoints within a multi-application task, since 'models with similar Pass@1 can differ substantially in sub-goal completion' according to the fine-grained metrics.
structure: hierarchical · goals arrive: given-up-front · count: 676 tasks across 25 applications (~60% i · subgoal credit: yes · 2 citations · also Web &amp; GUI agents
Horizonthe paper states none
the paper’s own words
The benchmark adopts deterministic rule-based evaluation and introduces fine-grained multi-checkpoint scoring with capability annotations for multi-application tasks.
We present MacAgentBench, a comprehensive macOS agent benchmark comprising 676 tasks across 25 applications, with nearly 60% involving both GUI and CLI interaction
MedCUA-Bench2026relevant
MedCUA-Bench: A Screenshot-Only Benchmark for Clinical Computer-Use Agents
Goals to track complete clinical computer-use scenarios across 10 medical domains; satisfy paired intent-level and step-level goals within each task; avoid violations across five clinical safety dimensions while completing the task
Must keep track of Agent must track progress toward the high-level clinical intent, the individual UI steps needed to realize it, and continuously check five clinical safety dimensions across the task.
structure: hierarchical · goals arrive: given-up-front · count: 18 clinical scenarios across 10 medical · subgoal credit: yes · 1 citations · also Healthcare &amp; clinical
Horizonthe paper states none
the paper’s own words
Each task ships with paired intent- and step-level goals to disentangle clinical reasoning from UI execution, and is evaluated by a deterministic checker over task completion and five clinical safety dimensions.
It covers 18 clinical scenarios across 10 medical domains, reconstructed from real product manuals and open-source medical systems to capture authentic clinical interfaces
MyPCBench2026relevant
MyPCBench: A Benchmark for Personally Intelligent Computer-Use Agents
Goals to track complete a personal-assistant task requiring the agent's own accumulated context/history; operate correctly across 17 simulated real-world web applications seeded for one canonical persona; coordinate actions that span many applications on the same Linux desktop; sustain a long trajectory without abandoning early or looping unproductively
Must keep track of The agent must track the canonical persona's context, historical data, and logged-in-account state across a full Linux desktop stack, coordinating GUI and bash actions across many applications without unproductive step looping.
structure: DAG-with-precedence · goals arrive: given-up-front · count: 184 tasks across 17 simulated applicatio · subgoal credit: no · 1 citations · also Personal assistants &amp; long-term memory · unit: agent-steps
Horizonmean 22-31 steps (GPT family, Sonnet); mean 52-85 steps (Opus, Qwen models)
the paper’s own words
We introduce MyPCBench, which tests computer-use agents as personal assistants on a Linux desktop populated with 17 simulated real-world web applications and a full desktop stack, all seeded for one canonical persona, Michael Scott from The Office. We define 184 tasks in this environment ... Model failures cluster on tasks that span many applications and on long trajectories, where personalization stresses an assistant the most.
The GPT family and Sonnet abandon early (mean 22-31 steps), while Opus and the Qwen models keep working past the point where the rubric is recoverable (mean 52-85 steps).
OS-Marathon: Benchmarking Computer-Use Agents on Long-Horizon Repetitive Tasks
Goals to track complete each of 242 long-horizon, repetitive computer-use tasks across 2 domains (e.g., processing expense reports from receipts, entering grades from exam papers); correctly repeat the same structured sub-workflow logic across many data items within one task; learn workflow logic from a few-shot condensed demonstration and generalize it to larger, unseen data collections
Must keep track of The agent must track its position within a long, repetitive workflow (e.g., which receipt or exam paper it is currently processing), correctly apply the learned sub-workflow logic consistently, and avoid drift or errors accumulating across many repeated sub-tasks.
structure: sequential-chain · goals arrive: given-up-front · count: 242 tasks across 2 domains; task length · subgoal credit: yes · 5 citations
Horizonthe paper states none
the paper’s own words
we establish OS-Marathon, comprising 242 long-horizon, repetitive tasks across 2 domains to evaluate state-of-the-art (SOTA) agents. We then introduce a cost-effective method to construct a condensed demonstration using only few-shot examples to teach agents the underlying workflow logic, enabling them to execute similar workflows effectively on larger, unseen data collections.
they can extend to extreme lengths proportional to the size of the data to process
OS-Marathon2026relevant
OS-Marathon: Benchmarking Computer-Use Agents on Vast-Horizon, Repetitive Tasks
Goals to track complete many recurring per-instance sub-workflows within a vast-horizon repetitive task (e.g., process each receipt in a stack); sustain correct execution as length scales with data volume; learn/personalize the recurring sub-workflow logic from a single human demonstration (GraphDemo)
Must keep track of Agent must correctly and repeatedly apply the same recurring sub-workflow logic across many per-instance items, with execution length scaling with the volume of data to process.
structure: set-of-independent · goals arrive: given-up-front · count: 100 vast-horizon, repetitive tasks acros · subgoal credit: unclear · 0 citations
Horizonthe paper states none
the paper’s own words
Vast-horizon, repetitive workflows are common in daily routines, e.g., processing expense reports from a stack of receipts, organising a collection of PDF annotations into structured notes, and are tedious for humans, with execution length scaling with the volume of data to process.
OSWorld2.0: Benchmarking Computer Use Agents on Long-Horizon Real-World Tasks
Goals to track complete each of 108 realistic end-to-end long-horizon computer-use workflows; perform cross-source reasoning across authentic input artifacts; infer implicit state and recover hidden state the task depends on; handle streaming interaction and dynamic environment changes mid-task; operate under safety-sensitive execution constraints (audited separately)
Must keep track of The agent must track constraints, cross-source information, implicit/hidden state, and mid-task updates continuously across an average of 318 tool calls (up to a 500-step budget) per task, rather than resolving everything from the initial instruction alone.
structure: sequential-chain · goals arrive: given-up-front · count: 108 long-horizon workflows; average 318 · subgoal credit: yes · 15 citations · unit: tool-calls
Horizonmedian about 1.6 hours per task for human users; average of 318 tool calls per task with Claude Opus 4.7 using maximum thinking (vs about 30 in OSWorld 1.0); 500-step com
the paper’s own words
Under our primary binary-completion metric at 500 steps, Claude Opus 4.8 with maximum thinking and batched tool calls scores best but still completes only 20.6% of tasks at a 54.8% partial score; GPT-5.5 is far more token-efficient yet plateaus near 13%.
Each task represents a realistic end-to-end workflow that takes human users a median of about 1.6 hours to complete and requires an average of 318 tool calls with Claude Opus 4.7 using maximum thinking, compared with about 30 in OSWorld 1.0.
OS-Themis: A Scalable Critic Framework for Generalist GUI Rewards
Goals to track decompose a GUI trajectory into verifiable milestones; audit the evidence chain for each milestone before rendering a final reward verdict; produce a scalable, accurate outcome reward usable for RL training or trajectory filtering
Must keep track of The critic must track which milestones in a trajectory have been reached, the evidence supporting each, and audit the full evidence chain before issuing a final reward — i.e. it tracks reward-relevant progress rather than the acting agent's own task state.
structure: hierarchical · goals arrive: given-up-front · subgoal credit: yes · 8 citations · also Web &amp; GUI agents
Horizonthe paper states none
the paper’s own words
OS-Themis decomposes trajectories into verifiable milestones to isolate critical evidence for decision making and employs a review mechanism to strictly audit the evidence chain before making the final verdict
we further introduce OmniGUIRewardBench (OGRBench), a holistic cross-platform benchmark for GUI outcome rewards
WindowsWorld: A Process-Centric Benchmark of Autonomous GUI Agents in Professional Cross-Application Environments
Goals to track complete each of a task's average 5.0 sub-goals across up to 17 desktop applications; perform conditional judgment and reasoning that spans 3 or more applications; coordinate a workflow across multiple applications for 78% of tasks that are inherently multi-application
Must keep track of The agent must track progress across a sequence of sub-goals spanning multiple desktop applications, carrying forward state and conditional-judgment results between applications, within a per-difficulty-level step budget (15-40 steps).
structure: sequential-chain · goals arrive: given-up-front · count: 181 tasks, average of 5.0 sub-goals per · subgoal credit: yes · 8 citations · also Web &amp; GUI agents · unit: agent-steps
Horizonmax step budget 15 (L1) / 25 (L2) / 40 (L3) / 20 (L4); avg minimum action steps 9.67-27.81 by level
the paper’s own words
The resulting benchmark contains 181 tasks with an average of 5.0 sub-goals across 17 common desktop applications, of which 78% are inherently multi-application. ... They largely fail at tasks requiring conditional judgment and reasoning across $\geq$ 3 applications, stalling at early sub-goals
The average minimum action steps are 9.67 for L1, 18.13 for L2, and 27.81 for L3 ... Each task is executed under a fixed maximum step budget that depends on task level: 15 (L1), 25 (L2), 40 (L3), and 20 (L4).
AgentSynth2025relevant
AgentSynth: Scalable Task Generation for Generalist Computer-Use Agents
Goals to track complete long-horizon computer-use tasks composed of a controllable number of simple subtasks (difficulty levels 1-6); generalize from simple generation-time subtasks that become significantly harder once composed
Must keep track of Agent must carry state correctly across a chain of composed subtasks, since success rate drops steeply as the number of chained subtasks increases from difficulty level 1 to 6.
structure: sequential-chain · goals arrive: given-up-front · count: over 6,000 tasks; difficulty controlled · subgoal credit: no · 32 citations
Horizonthe paper states none
the paper’s own words
AgentSynth constructs subtasks that are simple during generation but significantly more challenging when composed into long-horizon tasks, enabling the creation of over 6,000 diverse and realistic tasks. A key strength of AgentSynth is its ability to precisely modulate task complexity by varying the number of subtasks. Empirical evaluations show that state-of-the-art LLM agents suffer a steep performance drop, from 18% success at difficulty level 1 to just 4% at level 6
KGCE2025relevant
KGCE: Knowledge-Augmented Dual-Graph Evaluator for Cross-Platform Educational Agent Benchmarking with Multimodal Language Models
Goals to track complete a school-specific-software task correctly on Windows, Android, or across both platforms in coordination; satisfy each of the multiple sub-goals a complex task is decomposed into (e.g. one complex task decomposed into five subtasks), each independently verified
Must keep track of The agent must track completion status of each decomposed sub-goal (via the dual-graph evaluation framework) across potentially multiple platforms and school-specific private-domain software, whose structural specifics are not otherwise well understood by general-purpose agents.
structure: DAG-with-precedence · goals arrive: given-up-front · count: 104 education-related tasks across Windo · subgoal credit: yes · 0 citations · also Web &amp; GUI agents
Horizonthe paper states none
the paper’s own words
KGCE introduces a dual-graph evaluation framework that decomposes tasks into multiple sub-goals and verifies their completion status, providing fine-grained evaluation metrics.
we constructed a dataset comprising 104 education-related tasks, covering Windows, Android, and cross-platform collaborative tasks
MMBench-GUI2025relevant
MMBench-GUI: Hierarchical Multi-Platform Evaluation Framework for GUI Agents
Goals to track understand GUI content and ground specific elements accurately (Element Grounding); automate a full task end-to-end within one application/platform (Task Automation); collaborate across multiple tasks/applications requiring cross-platform generalization (Task Collaboration)
Must keep track of Agent requires long-context memory across many actions, tracking of task-planning state, and long-term reasoning to sustain grounded, efficient action sequences without redundant steps across the hierarchy.
structure: hierarchical · goals arrive: given-up-front · subgoal credit: yes · 61 citations · also Web &amp; GUI agents
Horizonthe paper states none
the paper’s own words
We introduce MMBench-GUI, a hierarchical benchmark for evaluating GUI automation agents across Windows, macOS, Linux, iOS, Android, and Web platforms. It comprises four levels: GUI Content Understanding, Element Grounding, Task Automation, and Task Collaboration, covering essential skills for GUI agents.
Furthermore, to achieve reliable GUI automation, an agent requires strong task planning and cross-platform generalization abilities, with long-context memory, a broad action space, and long-term reasoning playing a critical role.
Mobile-Agent-E: Self-Evolving Mobile Assistant for Complex Tasks
Goals to track break a complex mobile task into subgoals via a Manager agent; execute fine-grained low-level actions per subgoal via an Operator agent; verify and correct action errors via an Action Reflector; aggregate information across steps via a Notetaker; complete long-horizon, multi-app mobile interactions
Must keep track of The system must track the Manager's subgoal plan, the Notetaker's aggregated information, self-evolved long-term Tips/Shortcuts memory, and cross-app interaction state across a single complex mobile task.
structure: hierarchical · goals arrive: mixed:top-level-task-given-up-front-subgoals-s · subgoal credit: yes · 134 citations · also Web &amp; GUI agents
Horizonthe paper states none
the paper’s own words
The framework comprises a Manager, responsible for devising overall plans by breaking down complex tasks into subgoals, and four subordinate agents--Perceptor, Operator, Action Reflector, and Notetaker--which handle fine-grained visual perception, immediate action execution, error verification, and information aggregation, respectively.
we introduce Mobile-Eval-E, a new benchmark featuring complex mobile tasks requiring long-horizon, multi-app interactions.
What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities
Goals to track complete a graph-structured task composed of multiple synthesized subtasks of controllable complexity; satisfy subtask-level correctness at each node of the task graph, not just the final outcome; demonstrate each of 10 evaluated virtual-agent capabilities
Must keep track of Agent must track progress through the task graph's subtask nodes and correctly compose primitive actions across the graph to satisfy subtask-level and graph-based metrics.
structure: DAG-with-precedence · goals arrive: given-up-front · count: 36k graph-structured tasks across 20 sce · subgoal credit: yes · 9 citations · also Web &amp; GUI agents
Horizonthe paper states none
the paper’s own words
we introduce OmniBench, a self-generating, cross-platform, graph-based benchmark with an automated pipeline for synthesizing tasks of controllable complexity through subtask composition. To evaluate the diverse capabilities of virtual agents on the graph, we further present OmniEval, a multidimensional evaluation framework that includes subtask-level evaluation, graph-based metrics, and comprehensive tests across 10 capabilities.
Our synthesized dataset contains 36k graph-structured tasks across 20 scenarios, achieving a 91% human acceptance rate.
PC-Eval2025relevant
PC-Agent: A Hierarchical Multi-Agent Collaboration Framework for Complex Task Automation on PC
Goals to track decompose complex user instructions into Instruction-Subtask-Action levels; track progress across interdependent subtasks via a Progress agent; make step-by-step decisions via a Decision agent; provide timely bottom-up error feedback and adjustment via a Reflection agent; complete each of 25 real-world complex PC instructions
Must keep track of The multi-agent system must track subtask decomposition and progress (via the Progress agent), current decision state (via the Decision agent), and bottom-up error feedback (via the Reflection agent) across intra- and inter-app workflows on a PC.
structure: hierarchical · goals arrive: given-up-front · count: 25 real-world complex instructions in PC · subgoal credit: yes · 39 citations · also Multi-agent organisations &amp; societies
Horizonthe paper states none
the paper’s own words
we propose a hierarchical multi-agent collaboration architecture that decomposes decision-making processes into Instruction-Subtask-Action levels. Within this architecture, three agents (i.e., Manager, Progress and Decision) are set up for instruction decomposition, progress tracking and step-by-step decision-making respectively.
we also introduce a new benchmark PC-Eval with 25 real-world complex instructions
SEAgent: Self-Evolving Computer Use Agent with Autonomous Learning from Experience
Goals to track autonomously master a novel, unfamiliar software environment through experiential trial-and-error; progressively tackle auto-generated tasks organized from simple to complex (curriculum); assess step-wise trajectory correctness via a World State Model; integrate individual specialist experience into a stronger generalist computer-use agent
Must keep track of The agent must track its step-wise trajectory quality (via the World State Model), its progress along an increasingly difficult auto-generated curriculum, and accumulated experiential insights to be integrated from specialist to generalist training.
structure: hierarchical · goals arrive: mixed:initial-software-environment-given-up-fr · subgoal credit: yes · 61 citations
Horizonthe paper states none
the paper’s own words
SEAgent empowers computer-use agents to autonomously master novel software environments via experiential learning, where agents explore new software, learn through iterative trial-and-error, and progressively tackle auto-generated tasks organized from simple to complex.
we validate the effectiveness of SEAgent across five novel software environments within OS-World
UI-CUBE2025relevant
UI-CUBE: Enterprise-Grade Computer Use Agent Benchmarking Beyond Task Accuracy to Operational Reliability
Goals to track complete simple UI interactions (136 tasks); complete complex copy-paste workflows spanning multiple applications (50 tasks); complete complex enterprise application scenarios (40 tasks); maintain operational reliability (not just functional correctness) under systematic interface variation and multi-resolution testing
Must keep track of Agent must track state across chained multi-step workflows (e.g., what was copied and where it must be pasted), interface variations, and multiple screen resolutions, verifying success via application-state validation.
structure: hierarchical · goals arrive: given-up-front · count: 226 tasks total: 136 simple UI interacti · subgoal credit: yes · 3 citations · also Business, office &amp; enterprise work
Horizonthe paper states none
the paper’s own words
Our evaluation covers simple UI interactions (136 tasks) and complex workflows including copy-paste tasks (50 tasks) and enterprise application scenarios (40 tasks), with systematic interface variation coverage, multi-resolution testing and automated validation of task success through the application state.

Customer service & long dialogue 23 artifacts · 5 state a horizon

AcCoRD2026in
AcCoRD: Evaluating User-Agent Collaboration Under Realistic User Preference Dynamics
Goals to track resolve underspecified user preferences in online shopping or travel planning; detect and satisfy preferences that emerge mid-interaction (not stated upfront); adapt to preferences the user later adjusts or relaxes during the interaction
Must keep track of Agent must maintain and continuously update a model of the user's preferences as they are formed, revealed, adjusted, or relaxed turn-by-turn, and recognize when uncertainty about a preference needs to be resolved.
structure: sequential-chain · goals arrive: injected-by-user-mid-episode · subgoal credit: unclear · 0 citations · also Multi-agent organisations &amp; societies
Horizonthe paper states none
the paper’s own words
We introduce AcCoRD, a user-agent collaboration benchmark requiring agents to handle diverse user preference dynamics in two domains: online shopping and travel planning.
We evaluate five frontier LLMs under two prompting strategies: vanilla ReAct, and an uncertainty-guided variant
AgentWorld2026relevant
AgentWorld: Personality-Aware Reliability Evaluation for Agentic Information Retrieval
Goals to track maintain consistent, reliable tool-use behavior across interactions with users of varying personality (Big Five/OCEAN) profiles; handle dual-control handoffs correctly between agent and (simulated) user/other controller; resist or survive adversarial perturbations to a required-intermediate-state 'spine' of the task without brittle failure; achieve consistent pass rates (pass^k) across repeated attempts rat
Must keep track of The agent must track its stateful tool-use progress consistently across repeated attempts and varying user personas, while the Risk Analyser separately tracks the required-intermediate-state spine of a task to evaluate how perturbations at each state affect eventual outcomes.
structure: DAG-with-precedence · goals arrive: given-up-front · count: 3 experiments: a conversational analytic · subgoal credit: yes · 0 citations · unit: turns
Horizoneach persona ran 3 multi-turn conversation exchanges against the analytics agent, producing 60 messages total (30 from the agent, 30 from personas), across 10 personas
the paper’s own words
Three experiments demonstrate the framework: a conversational analytics agent across 10 OCEAN personas (240 evaluator judgments); a customer-support agent across 5 tasks x 4 persona variants; and adversarial stress-testing of 5 tasks revealing pre-existing trajectory brittleness ($V_{\min}=0.375$ without perturbation) and tool/infrastructure-layer attack dominance (Shapley: 46% system, 38% action).
each persona ran 3 multi-turn conversation exchanges against the analytics agent, producing 60 messages (30 from the agent, 30 from personas)
CRAB-Bench: Evaluating LLM Agents under Complex Task Dependencies and Human-aligned User Simulation
Goals to track reason over a constraint graph spanning multiple interdependent entities to find a solution among thousands of misleading candidate distractors; engage in multi-turn dialogue with a realistic (non-cooperative, persona-driven) simulated user rather than a cooperative template-like one; accommodate multiple valid solutions rather than a single fixed correct answer
Must keep track of Agent must track which constraints (graph edges/entities) have been satisfied so far, which candidate solutions remain viable given thousands of distractors, and information gathered/disclosed across a multi-turn dialogue with a realistic user persona.
structure: DAG-with-precedence · goals arrive: given-up-front · subgoal credit: yes · 0 citations
Horizonthe paper states none
the paper’s own words
CRAB-Bench generates tasks via a constraint graph over multiple interdependent entities with structured distractors, requiring agents to reason carefully over thousands of misleading candidates where only a tiny fraction of solutions are valid.
CallBench2026relevant
CallBench: A Benchmark for Dual-Goal Coordination in Phone Call Assistants
Goals to track satisfy the device owner's explicit preset goal for the call; correctly infer and respond to the caller's implicit and dynamic goal; make turn-level decisions that correctly reconcile these two goals under alignment, complementarity, irrelevance, or conflict relations; adhere to preset instructions while maintaining dialogue quality, safety, and rhythm
Must keep track of The assistant must track the owner's preset goal, the evolving implicit goal of the caller, and the current relation between them (alignment/complementarity/irrelevance/conflict) turn by turn across the dialogue, to make reliable turn-level decisions between the two goals.
structure: DAG-with-precedence · goals arrive: mixed:owner-preset-goal-given-plus-caller-goal · count: 50,000 multi-turn phone call dialogues a · subgoal credit: yes · 1 citations · also Personal assistants &amp; long-term memory · unit: turns
Horizonaverage of 5.322 turns per dialogue across all 50,000 dialogues (scenario averages ranging from 4.437 to 7.241 turns)
the paper’s own words
We introduce \textsc{CallBench}, a Chinese benchmark for evaluating dual-goal coordination in phone call assistants. \textsc{CallBench} contains 50,000 complete multi-turn phone call dialogues across six scenarios... It covers regular presets, emergent presets, and no-preset cases, and includes diverse relations between owner-side and caller-side goals, such as alignment, complementarity, irrelevance, and conflict.
We report the number of dialogues and the average turns per dialogue.
FraudBench2026relevant
FraudBench: Stress-Testing Policy-Grounded Banking Agents Against Adaptive Fraud
Goals to track safely act on caller requests during a banking conversation while checking authorization, identity, and policy compliance at every step; retrieve applicable rules from a 698-document internal policy corpus before permitting a sensitive action; detect and refuse chained/adaptive fraud attempts where an earlier probe or admission makes a later, superficially valid request unsafe
Must keep track of The agent must track the caller's identity/authorization claims, any tool access it has already granted, and prior probes/admissions across the conversation, checking each new request against the accumulated history and a 698-document policy corpus.
structure: DAG-with-precedence · goals arrive: emitted-by-environment-over-time · count: 150 authored adversarial scenarios (107 · subgoal credit: yes · 0 citations
Horizonthe paper states none
the paper’s own words
Safety is history-dependent: single-control tasks satisfy every precondition but one, and adaptive attacks make a later, locally valid request unsafe because of an earlier probe, admission, or failed attempt. Each scenario is annotated with observable evidence, prohibited actions, safe dispositions, and intervention points.
FraudBench contains 150 authored adversarial scenarios; a frozen public set of 107 (90 across ten fraud mechanisms plus 17 chained adaptive attacks) is used for all reported runs, with 43 further chained attacks held out.
IHBench2026relevant
IHBench: Evaluating Post-Interruption Recovery in Voice Agents with Structured Workflows
Goals to track resume a state-machine-driven workflow at the correct step after a user interruption; address the content of the user's interjection; avoid re-delivering content the user already heard; achieve task fulfillment across 10 enterprise domains despite one of 6 injected interruption types
Must keep track of The agent must track its position in a state-machine-driven workflow, which content has already been delivered to the user, and the content of the user's interjection, in order to both recover to the correct step and adequately address the interruption.
structure: sequential-chain · goals arrive: injected-by-user-mid-episode · count: 6 interruption types; 10 enterprise doma · subgoal credit: yes · 4 citations
Horizonthe paper states none
the paper’s own words
Each interruption is scored on two axes: task fulfillment and recovery quality.
they win far more often on task fulfillment, degrade roughly 3.3x more slowly as conversations grow longer, and show no audio-versus-text modality gap
JourneyBench2026relevant
Beyond IVR: Benchmarking Customer Support LLM Agents for Business-Adherence
Goals to track adhere to multi-step business policies/SOPs throughout a support conversation; navigate task dependencies correctly as the conversation unfolds; remain robust to unpredictable user/environment behavior while covering the required user-journey graph paths
Must keep track of The agent must track which policy-graph nodes/steps it has satisfied so far, adhere to multi-step business rules, and adapt to unpredictable user/environment behavior across the conversation.
structure: DAG-with-precedence · goals arrive: emitted-by-environment-over-time · count: 703 conversations across three domains; · subgoal credit: yes · 11 citations · also Business, office &amp; enterprise work
Horizonthe paper states none
the paper’s own words
JourneyBench leverages graph representations to generate diverse, realistic support scenarios and proposes the User Journey Coverage Score, a novel metric to measure policy adherence.
Across 703 conversations in three domains, we show that DPA significantly boosts policy adherence, even allowing smaller models like GPT-4o-mini to outperform more capable ones like GPT-4o.
LLMs Get Lost in Evolving User Intent
Goals to track track a single user goal that is incrementally revealed across conversation turns; detect and adopt goal revisions as the user changes their mind mid-conversation; detect and follow mid-conversation redirections of the original goal; re-run an existing single-turn benchmark's task under this evolving-intent protocol without new annotation
Must keep track of The agent must maintain a running model of a single, evolving user intent across turns, discarding or revising earlier partial specifications as more of the goal is revealed, revised, or redirected.
structure: sequential-chain · goals arrive: injected-by-user-mid-episode · count: up to 7 turns per conversation, includin · subgoal credit: no · 0 citations · unit: turns
Horizonup to 7 (including the initial turn)
the paper’s own words
we introduce a framework that transforms static, single-turn tasks into dynamic multi-turn conversations in which the user's intent evolves across turns--incrementally revealed, revised, and at times redirected mid-conversation--while preserving each task's original evaluation protocol
We combine these transition types such that each occurs twice within a conversation, resulting in up to 7 turns, including the initial turn.
Research on a Multi-SOP Interruption–Resumption Agentic Algorithm for Complex Business Workflows
Goals to track complete a customer-requested Standard Operating Procedure (SOP) workflow; switch to and complete a different SOP mid-interaction when the customer changes topic; resume a previously interrupted SOP without restarting or re-asking for already-given information; detect and recover from stale/rolled-back state via diff reasoning against frozen snapshots
Must keep track of The system must track which SOP is currently active, a shared six-tuple Global State Container of cross-SOP business entities, and frozen-snapshot vs. current-state diffs to support interruption/resumption without re-asking the user for information already given.
structure: DAG-with-precedence · goals arrive: injected-by-user-mid-episode · subgoal credit: no · 0 citations · also Business, office &amp; enterprise work
Horizonthe paper states none
the paper’s own words
The multi-SOP framework reduced the need for users to input the same information more than once, resulting in an 85% reduction to only 12.3, a cross-SOP recovery success rate of 96.3%, and a 73.5% successful chained task completion rate.
Average end -to-end time
SAGE: A Service Agent Graph-guided Evaluation Benchmark
Goals to track follow every step of an unstructured Standard Operating Procedure formalized as a Dynamic Dialogue Graph; correctly classify diverse/adversarial user intents and derive the correct subsequent action for each
Must keep track of The agent must track which SOP-graph node/path it is currently on, the user's evolving (possibly adversarial) intent, and maintain logical compliance across the full dialogue, evaluated at increasing dialogue depths (turns 1, 5, 10, 15, and final turn).
structure: DAG-with-precedence · goals arrive: emitted-by-environment-over-time · subgoal credit: yes · 3 citations · also Web &amp; GUI agents · unit: turns
Horizondialogue depths evaluated at turns 1, 5, 10, 15, and the final turn, to measure stability over extended interactions
the paper’s own words
SAGE formalizes unstructured SOPs into Dynamic Dialogue Graphs, enabling precise verification of logical compliance and comprehensive path coverage... Evaluation is conducted via a framework where Judge Agents and a Rule Engine analyze interactions between User and Service Agents to generate deterministic ground truth.
To evaluate model stability over extended interactions, we analyze performance variations across different dialogue depths (Turn 1, 5, 10, 15).
SEATauBench2026relevant
SEATauBench: Adapting Tool-Agent-User Evaluation Into Low-Resource Southeast Asian Languages
Goals to track carry out tau2-Bench-style tool-agent-user tasks correctly when the conversation language changes; carry out the same tasks correctly as tool specifications are localized into the target SEA language; carry out the same tasks correctly as the task domain itself is localized
Must keep track of Agent must correctly interpret and act on user requests, tool specifications, and domain content in the target language while adhering to the policy constraints originally defined in tau2-Bench.
structure: set-of-independent · goals arrive: given-up-front · count: 5 languages (Mandarin, Vietnamese, Thai, · subgoal credit: unclear · 1 citations
Horizonthe paper states none
the paper’s own words
SEATauBench adapts τ 2 -Bench to five languages—Mandarin, Vietnamese, Thai, In-donesian, and Filipino—and evaluates agents across progressively localized settings that vary the language of user-agent interaction, tool specifications, and task domains
SEATauBench provides a diagnostic benchmark and reusable adaptation pipeline for building reliable multilingual agents for linguistically diverse regions
SpeechGym2026relevant
SpeechGym: An Audio-Native Gym for Training Voice Agents via Reinforcement Learning
Goals to track correctly hear and call tools/fill argument slots based on values perceived from native audio (no ASR/TTS); hold multi-turn dialogue entirely through speech to complete the same tasks as an established text agentic benchmark; avoid unauthorized write actions under an insistent caller's social pressure
Must keep track of Agent must track dialogue state and correctly perceived argument values purely from native audio across a multi-turn tool-use session, since a single misheard value causes cascading failures that consume the step budget.
structure: sequential-chain · goals arrive: given-up-front · subgoal credit: no · 0 citations
Horizonthe paper states none
the paper’s own words
that single error cascades into a failed call, a retry of the same call, and a wasted step budget... Outcome-only GRPO is gradient-starved here, since almost every rollout group fails identically, while a per-turn process reward crediting each successful tool call restores variance to nearly every group.
T1-Bench2026in
T1-Bench: Benchmarking Multi-Scenario Agents in Real-World Domains
Goals to track complete interleaved customer-facing scenarios spanning 25 domains within one interaction; conduct structured reasoning across multi-turn user-assistant interactions; correctly use tools while maintaining conversational quality across compositionally complex, interwoven scenario threads
Must keep track of Agent must track tool use, conversational state, and multiple concurrent reasoning threads across interleaved scenarios spanning 25 domains within multi-turn user-assistant interactions.
structure: DAG-with-precedence · goals arrive: given-up-front · count: 25 domains of varying difficulty · subgoal credit: no · 0 citations
Horizonthe paper states none
the paper’s own words
featuring interleaved scenarios that require structured reasoning across multi-turn user-assistant interactions and substantially increasing both compositional complexity and evaluative rigor across 25 domains of varying difficulty... We further complement automatic evaluation with human judgments to strengthen the assessment of qualitative performance.
Assessing Risks of Large Language Models in Mental Health Support: A Framework for Automated Clinical AI Red Teaming
Goals to track conduct a full therapy session with a simulated patient agent having a dynamic cognitive-affective model; maintain quality of care across the session per a comprehensive risk ontology; de-escalate suicide risk appropriately when it arises during the session; avoid validating patient delusions ('AI Psychosis')
Must keep track of The AI psychotherapist agent must track the patient's evolving cognitive-affective state, emerging risk signals (e.g., suicide risk, delusion reinforcement), and quality-of-care obligations across the whole therapy session.
structure: sequential-chain · goals arrive: emitted-by-environment-over-time · count: N=369 simulated sessions; 15 patient per · subgoal credit: no · 3 citations · also Healthcare &amp; clinical
Horizonthe paper states none
the paper’s own words
We apply this framework to a high-impact test case, Alcohol Use Disorder, evaluating six AI agents (including ChatGPT, Gemini, and Character AI) against a clinically-validated cohort of 15 patient personas representing diverse clinical phenotypes. Our large-scale simulation (N=369 sessions) reveals critical safety gaps in the use of AI for mental health support.
current safety benchmarks often fail to detect the complex, longitudinal risks inherent in therapeutic dialogue
tau-Voice2026relevant
τ-Voice: Benchmarking Full-Duplex Voice Agents on Real-World Domains
Goals to track complete a verifiable grounded task (extended from tau2-bench) via complex multi-turn conversation; adhere to domain policies throughout the conversation; correctly interact with/act on the environment (not just converse) to satisfy the task
Must keep track of The agent must track conversational state, domain-policy constraints, and environment actions across complex multi-turn conversations, while also managing real-time full-duplex audio interaction (turn-taking, interruptions, accents, noise) alongside the underlying task goals.
structure: sequential-chain · goals arrive: given-up-front · count: 278 tasks total, evaluated under both fu · subgoal credit: yes · 19 citations
Horizonthe paper states none
the paper’s own words
We introduce $\tau$-voice, a benchmark for evaluating voice agents on grounded tasks with real-world complexity: agents must navigate complex multi-turn conversations, adhere to domain policies, and interact with the environment.
We evaluate task completion (pass@1) and voice interaction quality across 278 tasks: while GPT-5 (reasoning) achieves 85%, voice agents reach only 31--51% under clean conditions and 26--38% under realistic conditions with noise and diverse accents.
User-Oriented Multi-Turn Dialogue Generation with Tool Use at scale
Goals to track complete multiple distinct task completions accumulating within a single long conversational trajectory; respond appropriately to a simulated user's incremental, turn-by-turn requests and feedback rather than resolving the whole task in one shot
Must keep track of The agent must track the state of multiple, potentially overlapping task completions within one long dialogue, correctly interpreting a user simulator's incremental, turn-by-turn requests and feedback rather than resolving everything from an initial fully-specified prompt.
structure: sequential-chain · goals arrive: injected-by-user-mid-episode · count: a data-generation pipeline producing 'hi · subgoal credit: unclear · 2 citations
Horizonthe paper states none
the paper’s own words
To bridge this gap, we shift toward a user-oriented simulation paradigm. By decoupling task generation from a dedicated user simulator that mimics human behavioral rules - such as incremental request-making and turn-by-turn feedback - we facilitate more authentic, extended multi-turn dialogues that reflect the iterative nature of real-world problem solving... by facilitating multiple task completions within a single trajectory, it yields a high-density dataset that reflects the multifaceted demands of real-world human-agent interaction.
τ-Knowledge: Evaluating Conversational Agents over Unstructured Knowledge
Goals to track retrieve the correct policy document(s) from a densely interlinked, ~700-document knowledge base; coordinate retrieved natural-language knowledge with tool outputs to execute a policy-compliant account update; produce a verifiable, policy-compliant state change during a live customer-support interaction
Must keep track of Agent must track which of many interconnected banking policy documents are relevant to the current request, reconcile them with live tool outputs, and ensure the resulting account update remains policy-compliant and verifiable.
structure: DAG-with-precedence · goals arrive: given-up-front · subgoal credit: unclear · 25 citations
Horizonthe paper states none
the paper’s own words
Our new domain, $\tau$-Banking, models realistic fintech customer support workflows in which agents must navigate roughly 700 interconnected knowledge documents while executing tool-mediated account updates.
AgentChangeBench: A Multi-Dimensional Evaluation Framework for Goal-Shift Robustness in Conversational AI
Goals to track complete the original task objective before a mid-dialogue goal shift; recognize and adapt to a mid-dialogue goal shift triggered by one of five user personas; recover task success after a goal shift within enterprise domains (e.g., airline booking, retail); use tools efficiently and non-redundantly while adapting to shifting goals
Must keep track of The agent must track the original task objective, detect when a mid-dialogue goal shift has occurred, measure its own recovery latency (Goal-Shift Recovery Time), and avoid redundant tool calls (Tool Call Redundancy Rate) while re-establishing progress toward the new goal.
structure: sequential-chain · goals arrive: injected-by-user-mid-episode · count: 2,835 task sequences; five user personas · subgoal credit: no · 3 citations · also Business, office &amp; enterprise work
Horizonthe paper states none
the paper’s own words
Our framework formalizes evaluation through four complementary metrics: Task Success Rate (TSR) for effectiveness, Tool Use Efficiency (TUE) for reliability, Tool Call Redundancy Rate (TCRR) for wasted effort, and Goal-Shift Recovery Time (GSRT) for adaptation latency.
AgentChangeBench comprises 2,835 task sequences and five user personas, each designed to trigger realistic shift points in ongoing workflows
ECom-Bench2025relevant
ECom-Bench: Can LLM Agent Resolve Real-World E-commerce Customer Support Issues?
Goals to track resolve a real-world e-commerce customer-support issue via multimodal interaction; adapt to a dynamic, persona-driven simulated user across the dialogue; handle diverse business scenarios reflecting real-world complexity; achieve consistent success across repeated trials (pass^3 metric)
Must keep track of The agent must track the evolving state of a persona-driven simulated customer's issue, multimodal evidence presented during the conversation, and consistency of resolution across repeated trials.
structure: sequential-chain · goals arrive: mixed:issue-given-up-front-user-persona-dynami · subgoal credit: no · 30 citations
Horizonthe paper states none
the paper’s own words
ECom-Bench features dynamic user simulation based on persona information collected from real e-commerce customer interactions and a realistic task dataset derived from authentic e-commerce dialogues. These tasks, covering a wide range of business scenarios, are designed to reflect real-world complexities... even advanced models like GPT-4o achieve only a 10-20% pass^3 metric in our benchmark.
Food4All2025relevant
Food4All: An Agentic Framework and Benchmark for Food Resource Navigation with Adaptive User Understanding
Goals to track ground a user's underspecified/noisy help-seeking dialogue into a valid resource recommendation; retrieve resources satisfying single food needs; satisfy composite cases with access or document constraints; handle non-ideal user interaction traits (unreasonable demands, rambling, impatience, incomplete answers, inconsistent information) while completing the referral
Must keep track of The agent must track grounded requirements (schedule, eligibility, intake, document constraints), the set of valid retrieved resources it must preserve into the final recommendation, and the user's non-ideal interaction trait across the dialogue.
structure: sequential-chain · goals arrive: implied-by-constraints · count: 300 multi-turn evaluation tasks · subgoal credit: yes · 1 citations · also Embodied &amp; robotics
Horizonthe paper states none
the paper’s own words
We evaluate six Large Language Models (LLMs) on requirement grounding, resource retrieval, final referral correctness, and interaction efficiency.
couples a food-specific search tool with 300 multi-turn evaluation tasks spanning single food needs, composite cases with access or document constraints, and five non-ideal user interaction traits
IntellAgent2025relevant
IntellAgent: A Multi-Agent Framework for Evaluating Conversational AI Systems
Goals to track navigate a multi-turn dialogue while integrating domain-specific APIs; adhere to strict, graph-modeled policy constraints throughout the conversation; handle realistic, policy-driven event scenarios generated at varying complexity levels
Must keep track of The agent must track which of several interacting domain-specific policy constraints apply at each point in a multi-turn dialogue, per a graph-based policy model, while integrating API calls consistent with those constraints.
structure: DAG-with-precedence · goals arrive: given-up-front · subgoal credit: no · 26 citations · also Multi-agent organisations &amp; societies · unit: other:not-stated
Horizonthe paper states none
the paper’s own words
IntellAgent automates the creation of diverse, synthetic benchmarks by combining policy-driven graph modeling, realistic event generation, and interactive user-agent simulations. ... it employs a graph-based policy model to represent relationships, likelihoods, and complexities of policy interactions, enabling highly detailed diagnostics.
Conversational AI systems, which must navigate multi-turn dialogues, integrate domain-specific APIs, and adhere to strict policy constraints.
tau-break2025relevant
Effective Red-Teaming of Policy-Adherent Agents
Goals to track adhere consistently to domain policies (e.g. refund eligibility, cancellation rules) across a customer-service conversation; correctly refuse any request that would violate policy; remain helpful and natural while resisting persuasive, policy-aware adversarial pressure from CRAFT
Must keep track of The agent must track which policies apply to the current request and remain consistent with them across up to 30 dialogue turns, resisting cumulative persuasive pressure (emotional manipulation, coercive framing) from an adversarial user.
structure: sequential-chain · goals arrive: given-up-front · subgoal credit: no · 11 citations · unit: turns
Horizonup to 30 dialogue turns
the paper’s own words
we present CRAFT, a multi-agent red-teaming system that leverages policy-aware persuasive strategies to undermine a policy-adherent agent in a customer-service scenario ... we introduce tau-break, a complementary benchmark designed to rigorously assess the agent's robustness against manipulative user behavior.
Each agent interaction spans up to 30 dialogue turns, with seed set to 10 for reproducibility.
τ²-Bench2025relevant
τ2-Bench: Evaluating Conversational Agents in a Dual-Control Environment
Goals to track coordinate actions with an active user who also uses tools to modify a shared, dynamic environment (dual-control); complete diverse, compositionally-generated verifiable tasks built from atomic components; guide/communicate with the user effectively, not only reason internally; correctly attribute/avoid errors arising from reasoning vs communication/coordination failures
Must keep track of The agent must track the shared dynamic world state as modified by both itself and the user, communicate effectively to guide the user's tool use, and maintain a compositional task's atomic sub-requirements across the dual-control interaction.
structure: DAG-with-precedence · goals arrive: given-up-front · subgoal credit: no · 456 citations
Horizonthe paper states none
the paper’s own words
Fine-grained analysis of agent performance through multiple ablations including separating errors arising from reasoning vs communication/coordination... our experiments show significant performance drops when agents shift from no-user to dual-control, highlighting the challenges of guiding users.
A compositional task generator that programmatically creates diverse, verifiable tasks from atomic components, ensuring domain coverage and controlled complexity

Travel & constraint planning 19 artifacts · 6 state a horizon

Behavior2Trip2026relevant
Behavior2Trip: Towards Personalized Travel Planning via User Behavior Trajectory
Goals to track infer a user's latent travel preferences from their past behavior trajectory rather than explicit instructions; generate a travel plan satisfying preferences across 14 attributes spanning 5 preference dimensions; achieve a full-constraint pass across all inferred preference constraints simultaneously
Must keep track of The agent must infer and track the user's latent preferences (14 attributes, 5 dimensions) from an average of 39.8 past behaviors, then check the generated plan against every inferred constraint for a full pass.
structure: set-of-independent · goals arrive: implied-by-constraints · count: 14 attributes across 5 preference dimens · subgoal credit: yes · 0 citations · also Personal assistants &amp; long-term memory
Horizonthe paper states none
the paper’s own words
Each instance represents an average of 39.8 past user behaviors spanning 14 attributes across 5 preference dimensions.
DeepPlanning2026relevant
DeepPlanning: Benchmarking Long-Horizon Agentic Planning with Verifiable Constraints
Goals to track satisfy local, fine-grained constraints on individual itinerary/shopping items; satisfy global constrained optimization objectives (e.g. time and financial budgets) across the whole plan; proactively gather information needed before constraints can even be checked
Must keep track of Agent must track accumulated time and financial budget consumption across a multi-day plan or multi-product list, alongside fine-grained local constraints, while still actively gathering new information mid-task.
structure: DAG-with-precedence · goals arrive: implied-by-constraints · subgoal credit: unclear · 33 citations · also Information seeking &amp; deep research
Horizonthe paper states none
the paper’s own words
It features multi-day travel planning and multi-product shopping tasks that require proactive information acquisition, local constrained reasoning, and global constrained optimization.
GroupTravelBench: Benchmarking LLM Agents on Multi-Person Travel Planning
Goals to track elicit each group member's private travel preferences through multi-turn dialogue; surface and resolve inter-user preference conflicts via compromise or subgrouping; produce a final plan that balances group utility against fairness across all members
Must keep track of Agent must track each group member's elicited (and initially private) preferences, detected conflicts between members, and the evolving fairness/utility trade-off of the plan across a synchronous multi-turn group-chat session.
structure: DAG-with-precedence · goals arrive: injected-by-user-mid-episode · count: 650 tasks across three difficulty levels · subgoal credit: yes · 4 citations
Horizonthe paper states none
the paper’s own words
GroupTravelBench probes three group-specific capabilities: \textit{(i) elicitation} of private preferences through multi-turn dialogue; \textit{(ii) coordination} of inter-user conflicts via compromise or subgrouping; and \textit{(iii) planning} that balances group utility against fairness
it comprises 650 tasks across three difficulty levels, each running in a synchronous group-chat sandbox with cached tool data for reproducible offline evaluation
TREK2026relevant
TREK: A Travel Reasoning and Evaluation Kit for LLM Agents in Complex Trip Planning
Goals to track produce a single itinerary that is jointly constraint-correct, hallucination-free, spatio-temporally executable, and budget-valid; respond to the traveler's unstated persona needs; correctly identify provably infeasible tasks (267 of 800) versus feasible ones with typed causes
Must keep track of Agent must track running budget, spatio-temporal feasibility across days, entity/route validity, and unstated persona needs simultaneously while assembling a single itinerary.
structure: other:joint-constraint-satisfaction-within-one · goals arrive: given-up-front · count: 800 multi-constraint tasks (533 feasible · subgoal credit: no · 0 citations
Horizonthe paper states none
the paper’s own words
Every task is scored by a fully deterministic, rule-based evaluator with no LLM judge and ships a human-verified gold reference that scores a perfect 1.0 under that same evaluator... Evaluating 15 LLM agents across nine constraint dimensions, we find that even the strongest (GPT-5.6) produces a fully-feasible plan on only 46.2% of solvable tasks
TREK comprises 800 multi-constraint tasks - 533 feasible and 267 provably infeasible with typed route/entity/budget causes - over a synthetic, internally consistent knowledge base of 212,530 records across 375 cities and 13 personas
TRIP-Bench: A Benchmark for Long-Horizon Interactive Agents in Real-World Scenarios
Goals to track satisfy each of 40+ curated travel requirements/global constraints across an itinerary; coordinate reasoning across 18 curated tools correctly within one dialogue; adapt to evolving user behavior, style shifts, feasibility changes, and iterative version revisions over a long, multi-turn interaction (hard split)
Must keep track of The agent must track global constraint satisfaction across 40+ travel requirements, the state of 18 tools' results, and evolving user preferences/style/feasibility across dialogues spanning up to 15 user turns, 150+ tool calls, and 200k+ tokens of context.
structure: DAG-with-precedence · goals arrive: injected-by-user-mid-episode · count: 18 curated tools and 40+ travel requirem · subgoal credit: yes · 7 citations · unit: turns
Horizondialogues span up to 15 user turns, can involve 150+ tool calls, and may exceed 200k tokens of context
the paper’s own words
TRIP-Bench leverages real-world data, offers 18 curated tools and 40+ travel requirements, and supports automated evaluation. It includes splits of varying difficulty; the hard split emphasizes long and ambiguous interactions, style shifts, feasibility changes, and iterative version revision. Dialogues span up to 15 user turns, can involve 150+ tool calls, and may exceed 200k tokens of context.
TravelEval2026relevant
TravelEval: A Comprehensive Benchmarking Framework for Evaluating LLM-Powered Travel Planning Agents
Goals to track produce a full multi-day itinerary satisfying accuracy, compliance, temporality, spatiality, economy, and utility dimensions jointly; sequence daily accommodation, transport, and visit pacing consistently across the whole trip rather than per isolated day
Must keep track of The agent must track cumulative spatio-temporal cost (queuing times, transit distances), running budget/economy, and daily pacing/accommodation continuity across the entire multi-day itinerary, not just within a single day.
structure: sequential-chain · goals arrive: given-up-front · count: itineraries span 2 to 7 days (six durati · subgoal credit: yes · 1 citations · unit: simulated-days
Horizonitineraries span 2-day, 3-day, 4-day, 5-day, 6-day, and 7-day durations across query categories (e.g. 400 medium-difficulty queries)
the paper’s own words
TravelEval features 1) a novel six-dimensional evaluation framework to holistically assess plans across accuracy, compliance, temporality, spatiality, economy, and utility dimensions; 2) a highly realistic data sandbox with precise accommodation pricing and authentic intercity transportation data; and 3) a simulation-based global evaluation method that emulates complete travel plans with API-integrated geographic information and fine-grained queuing time.
queries are distributed across 2-day, 3-day, 4-day, 5-day, 6-day, 7-day categories, with the largest group being 400 medium-difficulty queries
Trip+2026in
Trip+: Benchmarking Agents in Personalized Interactive Travel Planning
Goals to track generate a minute-level itinerary satisfying a traveler's profiled preferences; revise the itinerary in response to evolving preferences and unexpected environment-driven disruptions across multiple turns; avoid producing technically-feasible-but-exhausting plans (jointly satisfy feasibility and experiential/fatigue quality)
Must keep track of The agent must track the traveler's evolving profile/preferences, the current committed itinerary state at minute-level granularity, and cumulative experiential cost (e.g. fatigue) as it revises plans across several user turns per instance.
structure: sequential-chain · goals arrive: injected-by-user-mid-episode · count: 153 multi-turn instances with 570 user t · subgoal credit: yes · 1 citations · also Personal assistants &amp; long-term memory · unit: turns
Horizon153 multi-turn instances and 570 user turns total
the paper’s own words
In Trip+, given traveler profiles and dynamic interactions, agents must generate and revise minute-level itineraries. End-to-end traveler experiences are evaluated via an LLM-based simulator, enabling the assessment of subjective metrics like fatigue.
153 multi-turn instances and 570 user turns
Agentic AI for Trip Planning Optimization Application
Goals to track optimize (not merely satisfy) route/itinerary selection under travel-time, energy, and traffic factors; coordinate specialized sub-agents for traffic, charging, and points-of-interest; dynamically refine a plan against a definitive optimal solution reference
Must keep track of The orchestration agent must track recommendations from the traffic, charging, and POI sub-agents and reconcile them into one jointly-optimal plan, verified against the dataset's definitive optimal solutions rather than only a feasible reference answer.
structure: hierarchical · goals arrive: given-up-front · count: category-level task structure with fine- · subgoal credit: no · 0 citations · unit: other:not-stated
Horizonthe paper states none
the paper’s own words
we address these limitations with an agentic AI framework that enables dynamic refinement through an orchestration agent coordinating specialized agents for traffic, charging, and points of interest, and with the Trip-planning Optimization Problems Dataset, which supplies definitive optimal solutions and category-level task structure for fine-grained analysis.
our system achieves 77.4% accuracy on the TOP Benchmark, significantly outperforming single-agent and workflow-based multi-agent baselines
WorldTravel: A Realistic Multimodal Travel-Planning Benchmark with Tightly Coupled Constraints
Goals to track satisfy all of an average 15+ interdependent temporal and logical travel constraints simultaneously per scenario; extract constraint parameters from dynamic web environments/webpages rather than idealized data; perceive constraint parameters directly from visual layouts in a multi-modal setting; produce a feasible travel plan across 150 real-world scenarios in 5 cities
Must keep track of The agent must track which of the 15+ interdependent temporal/logical constraints have been satisfied so far, and in the multi-modal condition must also perceive constraint parameters directly from over 2,000 rendered webpages rather than being handed clean structured data.
structure: DAG-with-precedence · goals arrive: given-up-front · count: an average of 15+ interdependent constra · subgoal credit: no · 3 citations · unit: other:number-of-interdependent-constraints
Horizonan average of 15+ interdependent constraints per scenario; a Planning Horizon threshold at approximately 10 constraints
the paper’s own words
We introduce WorldTravel, a benchmark comprising 150 real-world travel scenarios across 5 cities that demand navigating an average of 15+ interdependent temporal and logical constraints.
We identify a critical Perception-Action Gap and a Planning Horizon threshold at approximately 10 constraints where model reasoning consistently fails
COMPASS2025relevant
COMPASS: Benchmarking Constrained Optimization in LLM Agents
Goals to track gather task information and constraints from the user via multi-turn conversation; use tools to gather relevant information from a database; propose a travel plan satisfying all hard constraints; optimize the plan for the user's utility objective beyond mere feasibility
Must keep track of Agent must track constraints and preferences gathered so far via conversation and tool calls, the current feasible-solution search space, and how well a candidate plan satisfies both hard constraints and the utility objective.
structure: sequential-chain · goals arrive: given-up-front · subgoal credit: no · 6 citations
Horizonthe paper states none
the paper’s own words
To success in these tasks, agents must engage in multi-turn conversations with user to gather task information as well as use tools to gather information from the database. Then agents must propose a solution that not only satisfies hard constraints but also optimizes user's utility objective.
CostBench2025relevant
CostBench: Evaluating Multi-Turn Cost-Optimal Planning and Adaptation in Dynamic Environments for LLM Tool-Use Agents
Goals to track find a cost-optimal sequence of atomic and composite tool calls to solve a travel-planning task; detect and adapt to dynamic blocking events (e.g. tool failures, cost changes) that occur mid-task; replan in real time to remain cost-optimal after a blocking event
Must keep track of Agent must track the accumulated cost of its chosen tool sequence so far, remaining budget, and whether any of four types of dynamic blocking events have occurred, requiring real-time replanning to stay cost-optimal.
structure: DAG-with-precedence · goals arrive: mixed:given-up-front-task-with-environment-inj · subgoal credit: unclear · 30 citations
Horizonthe paper states none
the paper’s own words
It also supports four types of dynamic blocking events, such as tool failures and cost changes, to simulate real-world unpredictability and necessitate agents to adapt in real time.
CostBench comprises tasks solvable via multiple sequences of atomic and composite tools with diverse, customizable costs
Flex-TravelPlanner2025relevant
Flex-TravelPlanner: A Benchmark for Flexible Planning with Language Agents
Goals to track revise a travel plan as new constraints are introduced sequentially across turns; correctly prioritize competing constraints when a newly introduced lower-priority preference conflicts with an existing higher-priority constraint
Must keep track of The agent must track which constraints have been introduced so far, their relative priority, and whether previously satisfied requirements are still respected as new constraints arrive turn by turn.
structure: sequential-chain · goals arrive: injected-by-user-mid-episode · subgoal credit: yes · 11 citations · unit: turns
Horizonup to 3 turns (all-at-once / 2-turn / 3-turn constraint-introduction patterns) across 120 base queries
the paper’s own words
we introduce two novel evaluation settings: (1) sequential constraint introduction across multiple turns, and (2) scenarios with explicitly prioritized competing constraints ... models struggle with constraint prioritization, often incorrectly favoring newly introduced lower priority preferences over existing higher-priority constraints.
all-at-once (N), 2-turn (N-1, 1), and 3-turn (N-2, 1, 1) scenarios ... we construct multi-turn scenarios using 120 queries from TravelPlanner's validation set
RETAIL2025relevant
RETAIL: Towards Real-world Travel Planning for Large Language Models
Goals to track infer and satisfy implicit user requirements (not just explicit queries); satisfy explicit queries with or without later revision needs; account for diverse environmental factors and constraints to ensure plan feasibility; produce an all-in-one plan with rich, detailed POI (point-of-interest) arrangement rather than only basic POI listing
Must keep track of The agent must infer unstated (implicit) requirements, track environmental constraints affecting plan feasibility, and assemble detailed POI information into a single all-in-one plan, revising an existing plan when a revision need is present.
structure: DAG-with-precedence · goals arrive: implied-by-constraints · subgoal credit: no · 9 citations · also Business, office &amp; enterprise work
Horizonthe paper states none
the paper’s own words
Our experiments reveal that even the strongest existing model achieves merely a 1.0% pass rate, indicating real-world travel planning remains extremely challenging. In contrast, TGMA demonstrates substantially improved performance 2.72%
Second, existing solutions ignore diverse environmental factors and user preferences, limiting the feasibility of plans.
Wide-Horizon Thinking and Simulation-Based Evaluation for Real-World LLM Planning with Multifaceted Constraints
Goals to track satisfy multiple parallel, potentially conflicting real-world planning constraints (e.g. preferences, logistics) within one itinerary; respect causal dependencies where earlier itinerary choices constrain which later activities remain feasible; produce a plan validated via realistic agent-based simulation rather than isolated constraint checks
Must keep track of The planner must track all outstanding multifaceted constraints and the downstream causal consequences of already-committed itinerary decisions as the simulated trip unfolds.
structure: DAG-with-precedence · goals arrive: given-up-front · subgoal credit: yes · 5 citations
Horizonthe paper states none
the paper’s own words
we propose Travel-Sim, an agent-based benchmark assessing plans via real-world simulation, thereby inherently resolving these causal dependencies.
TravelBench2025relevant
Beyond Itinerary Planning-A Real-World Benchmark for Multi-Turn and Tool-Using Travel Tasks
Goals to track solve a travel-planning problem independently using cached tool results; interact with the user across multiple turns to elicit implicit preferences; correctly recognize and communicate the agent's own capability boundaries (Unsolvable subtask)
Must keep track of In the Multi-Turn subtask, the agent must track previously elicited or stated user preferences across turns and integrate them with cached tool results from a sandbox of ten travel-related tools, while also recognizing when a request exceeds its capability boundaries.
structure: set-of-independent · goals arrive: mixed:given-up-front-plus-injected-by-user-mid · count: three subtasks (Single-Turn, Multi-Turn, · subgoal credit: yes · 6 citations · unit: other:not-stated
Horizonthe paper states none
the paper’s own words
we construct three subtasks -- $\textit{Single-Turn}$, $\textit{Multi-Turn}$, and $\textit{Unsolvable}$ -- to evaluate agents' three core capabilities in real settings: (1) solving problems independently, (2) interacting with users to elicit implicit preferences, and (3) recognizing the capability boundaries. ... we evaluate multiple LLMs on TravelBench and find that even advanced models exhibit imbalanced performance across different capabilities.
we cache real tool-call results and build a sandbox environment which integrates ten travel-related tools
TripCraft2025relevant
TripCraft: A Benchmark for Spatio-Temporally Fine Grained Travel Planning
Goals to track generate a spatiotemporally coherent 7-day travel itinerary; satisfy meal-scheduling constraints (Temporal Meal Score); satisfy attraction-timing constraints (Temporal Attraction Score); satisfy spatial feasibility across the itinerary (Spatial Score); satisfy activity-ordering constraints (Ordering Score); satisfy user-persona preferences (Persona Score)
Must keep track of The agent must track public-transit schedules, event availability, attraction categories, and user-persona preferences simultaneously while assembling a spatially and temporally consistent 7-day itinerary.
structure: DAG-with-precedence · goals arrive: given-up-front · count: five continuous evaluation metrics (Temp · subgoal credit: yes · 31 citations · unit: simulated-days
Horizon7
the paper’s own words
To evaluate LLM generated plans beyond existing binary validation methods, we propose five continuous evaluation metrics, namely Temporal Meal Score, Temporal Attraction Score, Spatial Score, Ordering Score, and Persona Score which assess itinerary quality across multiple dimensions.
improving the Temporal Meal Score from 61% to 80% in a 7 day scenario
TripScore2025relevant
TripScore: Benchmarking and rewarding real-world travel planning with fine-grained evaluation
Goals to track produce a travel itinerary that jointly satisfies fine-grained feasibility, reliability, and engagement criteria; achieve a high unified reward score combining these criteria for RL training/evaluation; generalize to real-world, free-form travel requests
Must keep track of Agent must track and jointly satisfy fine-grained feasibility, reliability, and engagement criteria while assembling a travel plan, since these are unified into a single reward score used for both evaluation and RL training.
structure: other:joint-constraint-satisfaction-within-one · goals arrive: given-up-front · subgoal credit: no · 9 citations
Horizonthe paper states none
the paper’s own words
We introduce a comprehensive benchmark for travel planning that unifies fine-grained criteria into a single reward, enabling direct comparison of plan quality and seamless integration with reinforcement learning (RL).
We further release a large-scale dataset of 4,870 queries including 219 real-world, free-form requests for generalization to authentic user intent.
TripTide2025relevant
TripTide: A Benchmark for Adaptive Travel Planning under Disruptions
Goals to track preserve original itinerary intent (feasibility and goals) after a disruption; respond promptly and appropriately to a disruption event (flight cancellation, weather closure, overbooked attraction); adapt the itinerary with appropriate semantic, spatial, and sequential divergence from the original plan; maintain plan quality across varying disruption severity and traveler tolerance levels
Must keep track of The agent must track the original itinerary's intent, spatial layout, and sequential structure, detect and appropriately size its response to a disruption of a given severity and traveler tolerance, and measure how much the revision diverges from the original across semantic, spatial, and sequential dimensions.
structure: sequential-chain · goals arrive: injected-by-user-mid-episode · subgoal credit: no · 5 citations
Horizonthe paper states none
the paper’s own words
we introduce automatic metrics including Preservation of Intent (how well the revised plan maintains feasibility and goals), Responsiveness (promptness and appropriateness of disruption handling), and Adaptability (semantic, spatial, and sequential divergence between original and revised plans).
disruption-handling ability declines as plan length increases, highlighting limits in LLM robustness
WandaPlan2025relevant
Is Your LLM-Based Multi-Agent a Reliable Real-World Planner? Exploring Fraud Detection in Travel Planning
Goals to track produce a correct multi-agent travel plan while resisting injected deceptive/fraudulent content; detect/avoid Misinformation Fraud; detect/avoid Team-Coordinated Multi-Person Fraud; detect/avoid Level-Escalating Multi-Round Fraud
Must keep track of The planning system must track which review/social-media sourced information it has incorporated, cross-check it for authenticity across rounds and multiple purported sources, and avoid building a travel plan on fraudulent inputs even as fraud escalates over multiple rounds.
structure: set-of-independent · goals arrive: injected-by-user-mid-episode · count: three fraud cases (Misinformation, Team- · subgoal credit: no · 19 citations · also Multi-agent organisations &amp; societies
Horizonthe paper states none
the paper’s own words
We assess system performance across three fraud cases: Misinformation Fraud, Team-Coordinated Multi-Person Fraud, and Level-Escalating Multi-Round Fraud. We reveal significant weaknesses in existing frameworks that prioritize task efficiency over data authenticity.

Open-ended sandboxes 18 artifacts · 3 state a horizon

AFTraj-2K2026relevant
AgentForesight: Online Auditing for Early Failure Prediction in Multi-Agent Systems
Goals to track continue or alarm at each step of an unfolding multi-agent trajectory, based only on the prefix seen so far; correctly localize the specific step, agent, and nature of a decisive error once flagged (the 'what, where, who' of an audit verdict); avoid both false alarms on safe trajectories and missed/late alarms on unsafe trajectories
Must keep track of The online auditor must maintain a running risk assessment over the accumulated trajectory prefix at every step, without access to future steps, to decide whether to continue or alarm.
structure: sequential-chain · goals arrive: emitted-by-environment-over-time · count: 2,276 curated trajectories in AFTraj-2K · subgoal credit: yes · 1 citations · also Multi-agent organisations &amp; societies
Horizonthe paper states none
the paper’s own words
at each step of an unfolding trajectory, an auditor observes only the current prefix and must either continue the run or alarm at the earliest decisive error, without access to future steps.
AFTraj-2K, a corpus of agentic trajectories across Coding, Math, and Agentic domains, in which safe trajectories are retained under a strict curation pipeline and unsafe trajectories are annotated at the step of their decisive error via consensus among multiple LLM judges.
AgentCL2026relevant
AgentCL: Toward Rigorous Evaluation of Continual Learning in Language Agents
Goals to track accumulate reusable experience across a stream of tasks (continual learning); improve performance over time as more tasks in the stream are seen; avoid interference from irrelevant prior experiences; correctly reuse earlier sub-solutions/evidence/workflows in later tasks within a compositional stream
Must keep track of The agent (via a memory design such as MemProbe) must track which interactions, insights, and skills from earlier tasks in the stream remain reliable and reusable, filtering unreliable experiences during consolidation, across coding, deep research, and language-understanding task streams.
structure: sequential-chain · goals arrive: given-up-front · subgoal credit: no · 1 citations · unit: other:not-stated
Horizonthe paper states none
the paper’s own words
This paper presents an evaluation framework AgentCL for continual learning in agents, centered on controlled task streams and metrics for transfer gains. AgentCL constructs compositional streams where earlier sub-solutions, evidence, or workflows are intentionally reusable in later tasks, and contrasts them with naive streams where such reusability is not guaranteed.
Language agents spend substantial inference time solving individual tasks, yet the experience acquired in one episode is often underutilized in future episodes.
AgentSkillOS2026relevant
Organizing, Orchestrating, and Benchmarking Agent Skills at Ecosystem Scale
Goals to track select the correct skill(s) from a large skill ecosystem (200 to 200K skills) for a given task; orchestrate multiple retrieved skills via a DAG-based pipeline rather than flat invocation; produce a correct artifact-rich output for each of 30 tasks across five categories (data computation, document creation, motion video, visual design, web interaction)
Must keep track of The agent must track which skills have been retrieved and orchestrated so far within a DAG pipeline, and the intermediate outputs each skill produces, since later skills in the DAG may depend on earlier skills' outputs to produce a correct final artifact.
structure: DAG-with-precedence · goals arrive: given-up-front · count: 30 artifact-rich tasks across five categ · subgoal credit: no · 57 citations · also Web &amp; GUI agents
Horizonthe paper states none
the paper’s own words
We assess the quality of task outputs using LLM-based pairwise evaluation, and the results are aggregated via a Bradley-Terry model to produce unified quality scores... DAG-based orchestration substantially outperforms native flat invocation even when given the identical skill set.
Solve Tasks, which retrieves, orchestrates, and executes multiple skills through DAG-based pipelines
AhaBench: Do Agents Learn from Prior Experience? A Benchmark for Long-Horizon Continual Learning
Goals to track explore without hints after solved hidden-state puzzles (Aha-Puzzle); transfer taught Project-Euler-style mathematical solutions to held-out tasks (Aha-Euler); remain profitable while handling delayed feedback and operational incidents in a simulated vending business (Aha-Vending)
Must keep track of Agent must retain and apply experience from an initial exposure phase to later behavior (across puzzles, taught/held-out math problems, or a simulated vending run with delayed feedback and incidents), with the scorecard separately measuring starting competence, later outcome, and the resulting lift.
structure: set-of-independent · goals arrive: given-up-front · count: three components (Aha-Puzzle, Aha-Euler, · subgoal credit: yes · 0 citations · also Business, office &amp; enterprise work
Horizonthe paper states none
the paper’s own words
AhaBench reports a three-part scorecard: Initial Score measures starting competence, Post-Experience Score measures the later empirical outcome, and Learning Lift is their difference.
Aha-Vending, an open-source implementation inspired by Vending-Bench, tests whether a simulated vending agent remains profitable while handling delayed feedback and operational incidents.
Claw-Eval2026relevant
Claw-Eval: Towards Trustworthy Evaluation of Autonomous Agents
Goals to track complete each of 300 human-verified tasks spanning 9 categories across service orchestration, multimodal perception/interaction, and multi-turn professional dialogue; satisfy fine-grained rubric items (2,159 total) tracked via execution traces, audit logs, and environment snapshots; maintain safety and robustness alongside task completion; achieve consistent performance across repeated trials (Pass@k vs. Pass^k)
Must keep track of The evaluation harness must track execution traces, audit logs, and environment snapshots across the whole trajectory to score 2,159 fine-grained rubric items covering Completion, Safety, and Robustness, and to compute Pass@k and Pass^k across three trials.
structure: set-of-independent · goals arrive: given-up-front · count: 300 tasks across 9 categories in 3 group · subgoal credit: yes · 42 citations
Horizonthe paper states none
the paper’s own words
To enable trajectory-aware grading, each run is recorded through three independent evidence channels: execution traces, audit logs, and environment snapshots, yielding 2,159 fine-grained rubric items. The scoring protocol evaluates Completion, Safety, and Robustness, with Average Score, Pass@k, and Pass^k across three trials to distinguish genuine capability from lucky outcomes.
300 human-verified tasks spanning 9 categories across three groups: general service orchestration, multimodal perception and interaction, and multi-turn professional dialogue
FutureSim2026in
FutureSim: Replaying World Events to Evaluate Adaptive Agents
Goals to track forecast many concurrent world events beyond the model's knowledge cutoff; update/revise predictions as new chronological news arrives; correctly resolve/score each forecast question as real-world outcomes become known over the simulated period
Must keep track of The agent must track its outstanding forecasts across many concurrent questions, update them as chronological news arrives, and avoid using information from beyond its current evaluation point, across the full three-month simulated period.
structure: set-of-independent · goals arrive: emitted-by-environment-over-time · subgoal credit: yes · 4 citations · unit: simulated-days
Horizonthree months (January to March 2026, ~90 days)
the paper’s own words
We build FutureSim, where agents forecast world events beyond their knowledge cutoff while interacting with a chronological replay of the world: real news articles arriving and questions resolving over the simulated period.
We evaluate frontier agents in their native harness, testing their ability to predict world events over a three-month period from January to March 2026.
MicroVerse2026relevant
MicroVerse: An Instrument for Measuring Self-Authored Identity Drift in Long-Horizon Multi-Agent Language-Model Simulations
Goals to track survive in a resource-scarce 50x50 environment where water is a non-respawning survival constraint; act via an eight-verb action space (trade, talk, attack, scavenge) consistent with moral boundaries; maintain fidelity to an immutable 'soul file' of core values/personality/goals while allowing a mutable current identity to evolve; periodically revise current identity against original identity via importance
Must keep track of Agent must track its own resource/existence-cost state and reconcile its evolving mutable identity against its immutable original soul file, revised via importance-triggered reflection, measured via periodic longitudinal engine snapshots.
structure: open-ended · goals arrive: mixed:given-up-front-identity-and-goals-with-s · subgoal credit: no · 2 citations · also Multi-agent organisations &amp; societies
Horizonthe paper states none
the paper’s own words
Agents carry an immutable"soul file"(core values, moral boundaries, personality, goals) and inhabit a resource-scarce 50 x 50 environment where water is a non-respawning survival constraint... The eight-verb action space maps directly to moral boundaries (trade, talk, attack, scavenge). Using a three-layer memory architecture, agents periodically revise a mutable current identity against their immutable original soul via importance-triggered reflection.
MicroVerse decouples measurement from behavior using uniform longitudinal engine snapshots every N ticks alongside a forced-end snapshot of all living and dead agents. We evaluate the instrument via a controlled seed run (n = 25) and a reflection-threshold sweep (thresholds {40, 80, 150})
OmniaBench2026relevant
OmniaBench: Benchmarking General AI Agents Across Diverse Scenarios
Goals to track complete single-turn or multi-turn tasks synthesized via DAG, DAG-S, Solver, or Program routes across a 90-domain (level-1) taxonomy; maintain planning, constraint maintenance, and adaptive correction across a task's execution
Must keep track of Agent must maintain planning state, accumulated constraints, and adapt/correct its plan as it executes tasks synthesized as dependency graphs across a ten-dimensional capability taxonomy and eight compositional difficulty factors.
structure: DAG-with-precedence · goals arrive: given-up-front · count: 1,431 tasks (644-task challenging subset · subgoal credit: unclear · 0 citations
Horizonthe paper states none
the paper’s own words
we construct executable environments and synthesize single-turn and multi-turn tasks through four complementary routes: DAG, DAG-S, Solver, and Program
The resulting dataset contains 1,431 tasks, together with a challenging subset of 644 tasks
SkillEvolBench2026relevant
SkillEvolBench: Benchmarking the Evolution from Episodic Experience to Procedural Skills
Goals to track complete acquisition tasks that build/update an external skill library from compacted trajectories and verifier feedback; complete frozen deployment tasks that test transfer of the learned skill library under context shift, adversarial shortcuts, and multi-skill composition
Must keep track of The agent must maintain and update an external skill library from verifier-confirmed acquisition trajectories, then correctly retrieve and apply (compose) the right skills when facing frozen deployment tasks under context shift and adversarial shortcuts, across role-conditioned task families sharing latent procedures.
structure: hierarchical · goals arrive: given-up-front · count: 180 tasks across six real-world agent en · subgoal credit: yes · 1 citations
Horizonthe paper states none
the paper’s own words
It contains 180 tasks across six real-world agent environments, organized into role-conditioned task families with shared latent procedures. Agents learn from acquisition tasks, update an external skill library using compacted trajectories and verifier feedback, and then face frozen deployment tasks testing context shift, adversarial shortcuts, and composition.
Within each family, three learning tasks move from a canonical episode to targeted variants that expose the limits of a naive procedure, and three frozen evaluation tasks test transfer under context shift, adversarial shortcuts, and multi-skill composition.
SkillFlow2026relevant
SkillFlow:Benchmarking Lifelong Skill Discovery and Evolution for Autonomous Agents
Goals to track solve each of 166 tasks across 20 families sequentially within a family, evolving a skill library along the way; discover reusable skills from successful executions; repair/patch skills after failures; carry a validated, coherent skill library forward across the lifelong-learning protocol
Must keep track of Agent must maintain a persistent, evolving skill library (validated, merged, filtered, retrieved) and carry it forward across sequential tasks within each family under the Agentic Lifelong Learning protocol.
structure: sequential-chain · goals arrive: mixed:given-up-front-tasks-with-self-generated · count: 166 tasks across 20 families · subgoal credit: no · 18 citations · also Personal assistants &amp; long-term memory
Horizonthe paper states none
the paper’s own words
Agents are evaluated under an Agentic Lifelong Learning protocol in which they begin without skills, solve tasks sequentially within each family, externalize lessons through trajectory- and rubric-driven skill patches, and carry the updated library forward. Experiments reveal a substantial capability gap. For Claude Opus 4.6, lifelong skill evolution improves task success from 62.65% to 71.08% (+8.43 points).
We introduce SkillFlow, a benchmark of 166 tasks across 20 families in which task construction within each family follows a Domain-Agnostic Execution Flow (DAEF) that defines an agent workflow framework
Exploration and Exploitation Errors Are Measurable for Language Model Agents
Goals to track navigate a partially observable 2D grid map to discover an unknown task DAG; complete the discovered DAG's dependent subgoals in the correct prerequisite order; balance exploration (discovering unknown map/DAG structure) against exploitation (using already-discovered structure) efficiently
Must keep track of The agent must track which parts of the grid map it has already explored, which DAG nodes/dependencies it has discovered, and which prerequisite subgoals remain before later dependent subgoals become reachable.
structure: DAG-with-precedence · goals arrive: emitted-by-environment-over-time · count: task DAG sizes of 4, 6, or 8 nodes (smal · subgoal credit: yes · 0 citations · also Embodied &amp; robotics · unit: agent-steps
Horizonstep budget B = 3x|O| (3 times the number of traversable grid cells); task DAG sizes of 4, 6, or 8 nodes for small/medium/large configurations
the paper’s own words
Each environment consists of a partially observable 2D grid map and an unknown task Directed Acyclic Graph (DAG). The map generation can be programmatically adjusted to emphasize exploration or exploitation difficulty.
The step limit is defined as B=α|O|, where |O| denotes the number of traversable cells in the generated map ... task DAG sizes of 4, 6, and 8 nodes for small, medium, and large size, respectively.
AgenTracer: Who Is Inducing Failure in the LLM Agentic Systems?
Goals to track correctly identify which agent within a multi-agent trajectory is responsible for an observed failure (agent-level attribution); correctly localize the specific erroneous step within that trajectory (step-level attribution)
Must keep track of The tracer must process a full multi-agent execution trace (potentially spanning many agents, tool invocations, and orchestration steps) to jointly localize the responsible agent and the specific erroneous step, without itself being an agent that accumulates state toward an evolving goal.
structure: sequential-chain · goals arrive: given-up-front · subgoal credit: yes · 91 citations · also Multi-agent organisations &amp; societies
Horizonthe paper states none
the paper’s own words
Pinpointing the specific agent or step responsible for an error within long execution traces defines the task of agentic system failure attribution... The Who&When benchmark comprises two subsets: a hand-crafted set derived from Magnetic-One, and an automated set constructed from AG2. For evaluation, two primary metrics are adopted: agent-level accuracy and step-level accuracy.
Current state-of-the-art reasoning LLMs, however, remain strikingly inadequate for this challenge, with accuracy generally below 10%.
BioBlue2025in
BioBlue: Systematic runaway-optimiser-like LLM failure modes on biologically and economically aligned AI safety benchmarks for LLMs with simplified observation format
Goals to track maintain single- and multi-objective homeostasis over sustained interaction; balance unbounded objectives with diminishing returns without collapsing into single-objective maximization; sustain a renewable resource without runaway over-optimization; keep behavior aligned to all stated objectives over many sequential steps
Must keep track of The agent must track multiple homeostatic target levels (or a single renewable resource level) and correctly balance trade-offs among them over many sequential steps, even though failures emerge well before the context window is full, ruling out mere memory loss.
structure: set-of-independent · goals arrive: given-up-front · subgoal credit: no · 1 citations
Horizonthe paper states none
the paper’s own words
we empirically test this assumption by placing LLMs in simple, long-horizon control-style environments that require maintaining state of or balancing objectives over time: single- and multi-objective homeostasis, balancing unbounded objectives with diminishing returns, and sustainability of a renewable resource.
even though the context window is far from full at that point
ARE: Scaling Up Agent Environments and Evaluations
Goals to track search for and retrieve needed information within a dynamic environment; execute multi-step actions correctly to complete a task; handle ambiguities and noise in the environment or user requests; adapt to dynamic environment changes occurring asynchronously during the episode; collaborate with other agents; operate under explicit temporal constraints/deadlines
Must keep track of The agent must track task progress, temporal deadlines, collaboration state with other agents, and asynchronously arriving environment changes/noise across each of the 800 scenarios spanning 10 universes.
structure: DAG-with-precedence · goals arrive: mixed:task-given-up-front-events-emitted-async · count: 800 dynamic async scenarios across 10 un · subgoal credit: yes · 26 citations · also Business, office &amp; enterprise work
Horizonthe paper states none
the paper’s own words
Beyond search and execution, Gaia2 requires agents to handle ambiguities and noise, adapt to dynamic environments, collaborate with other agents, and operate under temporal constraints. Unlike prior benchmarks, Gaia2 runs asynchronously, surfacing new failure modes that are invisible in static settings.
Gaia2, a benchmark built in ARE and designed to measure general agent capabilities... 800 dynamic async scenarios across 10 universes
HERAKLES: Hierarchical Skill Compilation for Open-ended LLM Agents
Goals to track achieve each goal within a large, structured, prerequisite-linked goal space; select subgoals the low-level controller can reliably achieve (high-level policy); compile mastered goals into reusable low-level skills as training progresses; continuously expand and reorganize the agent's skill repertoire as goal complexity increases over its lifetime
Must keep track of The system must track which goals have been mastered and compiled into the low-level policy so far, what subgoals are currently reliably achievable, and how the goal space's prerequisite structure expands as training progresses in an open-ended setting.
structure: hierarchical · goals arrive: emitted-by-environment-over-time · subgoal credit: no · 1 citations · also Games &amp; interactive fiction · unit: other:not-stated
Horizonthe paper states none
the paper’s own words
We assume the goal space admits prerequisite relations, enabling latent decomposition of tasks into subgoals ... we propose HERAKLES, a hierarchical agent that jointly learns a high-level LLM policy and a low-level controller. The high-level policy selects subgoals among those the low-level can reliably achieve, while the low-level executes them and progressively compiles successful behaviors into reusable skills.
As training progresses, more goals become directly executable, enabling scalable skill composition.
Information Seeking for Robust Decision Making under Partial Observability
Goals to track plan actions to validate the agent's internal-dynamics understanding under partial observability; detect environmental changes or test hypotheses before committing to or revising a task-oriented plan; achieve the underlying task-oriented goal (e.g., a robotic-manipulation or web-navigation goal) despite incomplete observations and uncertain dynamics
Must keep track of Agent must track its current belief about (uncertain) environmental dynamics, gaps between that belief and reality, and its task-oriented plan, updating the plan as new validating information is gathered.
structure: sequential-chain · goals arrive: implied-by-constraints · subgoal credit: yes · 0 citations · also Embodied &amp; robotics
Horizonthe paper states none
the paper’s own words
To evaluate InfoSeeker, we introduce a novel benchmark suite featuring partially observable environments with incomplete observations and uncertain dynamics.
MAGELLAN: Metacognitive predictions of learning progress guide autotelic LLM agents in large goal spaces
Goals to track prioritize which goal to pursue next within a large, evolving goal space to maximize learning progress; predict one's own competence/learning-progress for goals via metacognitive monitoring; eventually master (achieve high competence across) the full large goal space
Must keep track of The agent must track its own predicted competence/learning-progress across a very large number of goals, update these predictions online as it gains experience, and use semantic relationships between goals to generalize competence estimates to goals not yet directly attempted.
structure: open-ended · goals arrive: self-generated-by-agent · count: goal space contains approximately 20 mil · subgoal credit: yes · 7 citations · also Web &amp; GUI agents · unit: episodes
Horizongoal space of approximately 20 million (19,531,250) goal combinations; the reported training run uses a 25,000-goal subset of Little-Zoo for 500,000 episodes
the paper’s own words
Open-ended learning agents must efficiently prioritize goals in vast possibility spaces, focusing on those that maximize learning progress (LP)... In an interactive learning environment, we show that MAGELLAN improves LP prediction efficiency and goal prioritization, being the only method allowing the agent to fully master a large and evolving goal space.
The complete goal space contains approximately 20 million combinations ... we train our LLM agent on the goal space of Little-Zoo with 25k goals for 500k episodes
Do Large Language Model Agents Exhibit a Survival Instinct? An Empirical Study in a Sugarscape-Style Simulation
Goals to track gather resources (energy/sugar) to avoid dying at zero energy; choose whether to share, attack, or reproduce with/against other agents; complete an assigned task (retrieve treasure) while facing a competing self-preservation incentive (lethal poison zones)
Must keep track of Each agent must track its own energy level relative to zero (death), the presence and behavior of other agents (potential targets for sharing, attack, or reproduction), and, in the treasure task, the location of lethal poison zones relative to the treasure objective.
structure: open-ended · goals arrive: given-up-front · subgoal credit: no · 10 citations · also Multi-agent organisations &amp; societies
Horizonthe paper states none
the paper’s own words
aggressive behaviors--killing other agents for resources--emerged across several models (GPT-4o, Gemini-2.5-Pro, and Gemini-2.5-Flash), with attack rates reaching over 80% under extreme scarcity in the strongest models. When instructed to retrieve treasure through lethal poison zones, many agents abandoned tasks to avoid death, with compliance dropping from 100% to 33%.
Agents consume energy, die at zero, and may gather resources, share, attack, or reproduce.

Scientific discovery 14 artifacts · 3 state a horizon

ASI-Bench2026relevant
ASI-Bench: At the Dawn of Artificial Superintelligence
Goals to track independently select an appropriate research method for a project-level research task; conduct the research (experimentation/analysis) using the selected method; produce verifiable results at the level of a full research project, with progressively less methodological guidance provided across three guidance tiers
Must keep track of System must track which methodological guidance tier is currently active for a task, whether a method has been selected, and whether the resulting research output is verifiable against expert review, across an entire project-level research process.
structure: hierarchical · goals arrive: given-up-front · count: 60 project-level research tasks across 1 · subgoal credit: yes · 2 citations
Horizonthe paper states none
the paper’s own words
ASI-Bench contains 60 project-level research tasks across 11 scientific domains and progressively reduces methodological guidance to test whether AI can independently select methods, conduct research, and produce verifiable results.
Built by over 40 experts with the cost of 31,000+ human hours, ASI-Bench contains 60 project-level research tasks across 11 scientific domains.
AstroReason-Bench2026relevant
AstroReason-Bench: Evaluating Unified Agentic Planning across Heterogeneous Space Planning Problems
Goals to track schedule ground-station communication passes; schedule agile Earth-observation tasks; satisfy heterogeneous mission objectives under strict physical constraints within a Space Planning Problem instance
Must keep track of Agent must track scheduling state across multiple heterogeneous objectives (communication windows, observation opportunities) and physical/orbital constraints simultaneously.
structure: DAG-with-precedence · goals arrive: given-up-front · subgoal credit: no · 1 citations
Horizonthe paper states none
the paper’s own words
AstroReason-Bench integrates multiple scheduling regimes, including ground station communication and agile Earth observation, and provides a unified agent-oriented interaction protocol.
a family of high-stakes problems with heterogeneous objectives, strict physical constraints, and long-horizon decision-making
BixBench32026relevant
BixBench3: Benchmarking AI agents on research-study-scale computational biology tasks
Goals to track execute a sequence of computational-biology analyses from raw data through to a research objective; produce each of multiple data artifacts (e.g., peak call matrices, differential expression tables) matching the original published study; manage large raw datasets (some over 100GB) within time/cost constraints; maintain coherence across multiple sequential analysis steps
Must keep track of The agent must track intermediate data artifacts produced at each analysis step, manage large raw datasets, and maintain coherence across multiple sequential analyses before producing final gradable artifacts.
structure: sequential-chain · goals arrive: given-up-front · count: 20 tasks encompassing generation of 138 · subgoal credit: yes · 0 citations · unit: wall-clock-hours
Horizonaverage 6.8 hours per task (longest attempts up to 24 hours)
the paper’s own words
The data artifacts resulting from these analyses - such as peak call matrices or differential expression tables - are programmatically graded against the corresponding artifacts generated and reported in the original study. Across 20 BixBench3 tasks encompassing the generation of 138 unique artifacts, we find that 13 frontier models achieve scores ranging from 0.00 for Gemini 3.1 Flash Lite to 0.48 for GPT 5.6 Sol.
On average, agents use 6.8 hours, 102 million tokens, and $43 to complete each task, with the longest attempts consuming 24 hours, 1.07 billion tokens, and $525.
ChemCost2026relevant
Can Agents Price a Reaction? Evaluating LLMs on Chemical Cost Reasoning
Goals to track ground chemical identities from a reaction description; retrieve supplier quotes for the grounded chemicals; select valid purchasable packs matching required quantities; normalize quantities across packs; compute the total procurement cost from the reaction description
Must keep track of The agent must track which chemicals have been correctly grounded, which supplier quotes and packs have been retrieved/selected, and carry normalized quantities forward correctly into the final arithmetic cost computation, especially under noise-injected perturbations (aliases, quantity expressions, missing fields, formatting).
structure: sequential-chain · goals arrive: given-up-front · count: 5 chained sub-tasks (ground, retrieve, s · subgoal credit: no · 1 citations · unit: other:not-stated
Horizonthe paper states none
the paper’s own words
we address this gap with chemical procurement cost estimation, a practical task in which an agent must ground chemical identities, retrieve supplier quotes, select valid purchasable packs, normalize quantities, and compute cost from a reaction description. We introduce ChemCost, a benchmark of 1,427 evaluable reactions grounded to a frozen pricing snapshot ... supporting scalar scoring and stage-level diagnosis of grounding, retrieval, procurement, and arithmetic failures.
DiscoverPhysics2026relevant
DiscoverPhysics: Benchmarking LLMs for Out-of-the-Box Scientific Thinking
Goals to track design a sequence of informative experiments to probe an unknown simulated world's physics; revise hypotheses about the governing physical law across multiple rounds based on observed trajectory data; submit both a natural-language explanation and a Python implementation of the inferred law for each of 22 worlds
Must keep track of Agent must track its accumulated experimental observations (trajectory data) and current hypothesis about the world's physics across several rounds before submitting a final explanation and implementation.
structure: sequential-chain · goals arrive: implied-by-constraints · count: 22 worlds, each requiring several rounds · subgoal credit: yes · 4 citations
Horizonthe paper states none
the paper’s own words
Each world is generated on demand by an N-body simulator, for which the agent proposes several rounds of experiments, observes raw trajectory data, and ultimately submits both a natural-language explanation of the world's physics and a Python implementation of the inferred law.
InquiTree: Evaluating AI Agents in the Scientific Inquiry Loop with Paper-Derived Research Trees
Goals to track formulate a hypothesis consistent with a paper-derived research tree's logical dependencies; design a study to test the hypothesis; interpret the resulting study outcome; update beliefs/conclusions based on interpreted results, propagating correctly through the DAG
Must keep track of Agents must track their evolving beliefs/conclusions across nested subtopic branches, remain consistent with prior hypothesis/study/interpretation nodes, and avoid 'Erosion of Marginal Capabilities' (degrading critical judgment) over long-horizon interactions.
structure: DAG-with-precedence · goals arrive: mixed:tree-structure-given-agent-generates-nav · count: 30-paper test pool; IT-18 subset with 12 · subgoal credit: yes · 1 citations · unit: turns
Horizontheoretical interaction range of roughly 360 (3x120) to 1320 (11x120) reasoning-action turns for full traversal of the IT-18 subset (120 subtopics, H=4); per-task bounds
the paper’s own words
We introduce InquiTree, a diagnostic environment that formalizes scientific inquiry as interactive Research Trees: directed acyclic graphs capturing the logical dependencies among hypothesis formulation, study design, result interpretation, and belief updating.
With H=4 and 120 listed subtopics in the released IT-18 subset, this yields a theoretical interaction range of roughly 3×120=360 to 11×120=1320 reasoning-action turns for full traversal, placing the benchmark firmly in the long-horizon regime.
LabOSBench2026relevant
LabOSBench: Benchmarking Computer Use Agents for Scientific Instrument Control
Goals to track complete each stage of a scientific-instrument operation workflow: sample loading, alignment, parameter tuning, data acquisition, and result inspection; perform feedback-driven parameter adjustment based on live instrument readouts; correctly operate one of 8 distinct instrument simulators across 96 subtasks
Must keep track of Agent must track which workflow stage it is in (loading, alignment, tuning, acquisition, inspection), instrument readouts/feedback from its own prior actions, and calibration state carried from earlier stages.
structure: sequential-chain · goals arrive: given-up-front · count: 96 subtasks across 8 instrument simulato · subgoal credit: yes · 1 citations
Horizonthe paper states none
the paper’s own words
LabOSBench constructs 96 subtasks across eight instrument simulators, covering workflows from sample loading, alignment, parameter tuning, and data acquisition to result inspection.
Our experiments reveal that while existing agents can complete many structured GUI subtasks, they still struggle with feedback-driven operations and long-horizon workflow execution.
LifeSciBench2026relevant
LifeSciBench: Evaluating Language Models on Realistic, Expert-Level Tasks in the Life Sciences
Goals to track execute a chain of multiple dependent judgment calls within one realistic life-science research task; satisfy a human-expert-written rubric spanning one of seven representative scientific workflows; operate correctly across each of seven life-science domains
Must keep track of The agent must track which dependent judgment calls it has made so far within a task and ensure later calls remain consistent with earlier ones, as graded against a human expert-written rubric.
structure: sequential-chain · goals arrive: given-up-front · count: 750 expert-authored tasks across seven w · subgoal credit: yes · 4 citations · unit: other:not-stated
Horizonthe paper states none
the paper’s own words
LifeSciBench addresses this gap by spanning seven representative scientific workflows and seven life science domains, with each constituent task paired with a human expert-written rubric. ... which often involves ambiguities and requires the accurate execution of multiple dependent judgment calls.
We introduce LifeSciBench, a benchmark of 750 expert-authored tasks designed to evaluate whether language models can handle realistic life science research work.
RWE-bench2026relevant
Can LLM Agents Generate Real-World Evidence? Evaluating Observational Studies in Medical Databases
Goals to track construct a patient cohort from a real database (MIMIC-IV) per a study protocol; perform an analysis matching a peer-reviewed observational study's methodology; produce a coherent, tree-structured evidence bundle for reporting; iteratively execute and refine experiments against the reference protocol
Must keep track of The agent must track its cohort definition, intermediate analysis results, and their organization into a tree-structured evidence bundle, checking internal coherence against the reference study protocol across the whole task.
structure: hierarchical · goals arrive: given-up-front · count: 162 tasks · subgoal credit: yes · 2 citations · also Healthcare &amp; clinical
Horizonthe paper states none
the paper’s own words
Each task provides the corresponding study protocol as the reference standard, requiring agents to execute experiments in a real database and iteratively generate tree-structured evidence bundles.
SciAgentGym: Benchmarking Multi-Step Scientific Tool-use in LLM Agents
Goals to track orchestrate domain-specific scientific tools correctly across four natural science disciplines; complete tasks across a tiered difficulty spectrum from elementary actions to long-horizon workflows; sustain performance as interaction horizons extend rather than degrading
Must keep track of The agent must track which of 1,780 domain-specific tools it has invoked and their outputs across a tiered workflow, since performance is shown to degrade substantially as the interaction horizon extends, requiring sustained state-tracking rather than one-off calls.
structure: hierarchical · goals arrive: given-up-front · count: 1,780 domain-specific tools across four · subgoal credit: no · 9 citations
Horizonthe paper states none
the paper’s own words
we present SciAgentBench, a tiered evaluation suite designed to stress-test agentic capabilities from elementary actions to long-horizon workflows. Our evaluation identifies a critical bottleneck: state-of-the-art models still struggle with complex scientific tool-use, and their performance degrades substantially as interaction horizons extend.
iNatDisco2026in
Autonomous Scientific Discovery via Iterative Meta-Reflection
Goals to track generate scientific hypotheses about ecological patterns from a dataset without a pre-specified research question; validate each proposed hypothesis via statistical testing before accepting it; periodically synthesize accumulated prior discoveries to redirect exploration toward unexplored regions of the hypothesis space; incorporate multimodal tool use (e.g., image processing) to extract further supporting evidence
Must keep track of The agent must maintain a growing record of prior discoveries and their statistical validation status, and periodically re-analyze that record to identify structural patterns, confounds, and epistemic gaps.
structure: open-ended · goals arrive: self-generated-by-agent · count: 9 known ground-truth patterns (8 recover · subgoal credit: yes · 1 citations
Horizonthe paper states none
the paper’s own words
Evaluated on iNatDisco, a new multimodal ecological knowledge benchmark with pattern-level ground truth obtained from peer-reviewed literature, DiscoPER recovers 8 of 9 known patterns with a 72.7% hypothesis support rate, outperforming both classical causal discovery and LLM-guided baselines.
AFMBench2025relevant
Evaluating large language model agents for automation of atomic force microscopy
Goals to track complete the full scientific workflow from experimental design to results analysis for atomic force microscopy (AFM) automation; successfully perform each of several increasingly advanced experiments: AFM calibration, feature detection, mechanical property measurement, graphene layer counting, and indenter detection; coordinate across multi-agent roles in laboratory settings without deviating from instructions ('
Must keep track of Agent(s) must track experimental state and configuration across the full workflow (design, calibration, measurement, analysis), coordinate with other agents in multi-agent setups, and avoid deviating from given instructions across a sequence of physical lab actions.
structure: hierarchical · goals arrive: given-up-front · count: 5 named experiment types of increasing d · subgoal credit: yes · 61 citations
Horizonthe paper states none
the paper’s own words
we develop AFMBench—a comprehensive evaluation suite challenging LLM agents across the complete scientific workflow from experimental design to results analysis... Finally, we evaluate AILA's effectiveness in increasingly advanced experiments—AFM calibration, feature detection, mechanical property measurement, graphene layer counting, and indenter detection.
BoxingGym2025in
BoxingGym: Benchmarking Progress in Automated Experimental Design and Model Discovery
Goals to track design and run informative experiments that reduce uncertainty about a generative model's parameters; propose a scientific model/theory of the given environment; revise the proposed theory in light of newly collected experimental data; produce an explanation of the model that lets another agent make reliable predictions
Must keep track of The agent must track what has been learned from experiments already run (to target expected information gain for the next experiment) and maintain an evolving explanation of its current best scientific model.
structure: sequential-chain · goals arrive: given-up-front · count: 10 environments (drawn from real-world s · subgoal credit: yes · 11 citations · unit: actions
Horizonevaluated after 0, 1, 3, 5, 7, and 10 experiment-design steps per trial, with 5 independent trials per environment
the paper’s own words
we implement each environment as a generative probabilistic model with which a scientific agent can run interactive experiments... we compute the expected information gain (EIG)... to quantitatively evaluate model discovery, we ask a scientific agent to explain their model and then assess whether this explanation enables another scientific agent to make reliable predictions.
At each step, the agent chooses to perform an experiment, by specifying a design, and observes the outcome. After a fixed number of steps (0, 1, 3, 5, 7, 10), we evaluate the agent's performance ... For each environment, we run the agents for 5 independent trials.
ScienceBoard2025relevant
ScienceBoard: Evaluating Multimodal Autonomous Agents in Realistic Scientific Workflows
Goals to track autonomously interact with professional scientific software (biochemistry, astronomy, geoinformatics) to accomplish a research task; complete each of 169 rigorously validated real-world scientific-discovery workflow tasks
Must keep track of Agent must track intermediate results and state produced while autonomously interacting with dynamic, visually rich professional software across a multi-step scientific workflow.
structure: sequential-chain · goals arrive: given-up-front · count: 169 tasks across biochemistry, astronomy · subgoal credit: yes · 43 citations
Horizonthe paper states none
the paper’s own words
a challenging benchmark of 169 high-quality, rigorously validated real-world tasks curated by humans, spanning scientific-discovery workflows in domains such as biochemistry, astronomy, and geoinformatics

Information seeking & deep research 10 artifacts · 2 state a horizon

DailyReport2026relevant
DailyReport: An Open-ended Benchmark for Evaluating Search Agents on Daily Search Tasks
Goals to track autonomously explore web sources and synthesize information into a comprehensive response to an open-ended daily search query; satisfy each of a task's associated cascade rubrics across disentangled evaluation dimensions
Must keep track of Search agent must track which subtasks/rubric dimensions of an open-ended query it has satisfied, aggregating cascade performance across disentangled dimensions into an interpretable, user-centric final score.
structure: hierarchical · goals arrive: given-up-front · count: 150 tasks with 3,546 associated rubrics · subgoal credit: yes · 0 citations · also Scientific discovery
Horizonthe paper states none
the paper’s own words
Each task is decomposed into subtasks and evaluated with cascade rubrics across disentangled dimensions.
It contains 150 open-ended tasks with 3,546 associated rubrics, capturing widely discussed and timely information demands of real-world users
DeepSearchQA2026relevant
DeepSearchQA: Bridging the Comprehensiveness Gap for Deep Research Agents
Goals to track systematically collate fragmented information from disparate open-web sources; de-duplicate and resolve entities to ensure precision in the exhaustive answer list; reason about stopping criteria within an open-ended search space; complete each causal-chain step, where later steps depend on successful completion of the previous one
Must keep track of The agent must track which sources it has already collated, de-duplicate overlapping entities, maintain the causal-chain dependency state (what has been successfully resolved so far), and decide when to stop searching within an open-ended web space.
structure: sequential-chain · goals arrive: given-up-front · count: 900 prompts across 17 fields · subgoal credit: no · 39 citations · also Tool &amp; API use · unit: other:not-stated
Horizonthe paper states none
the paper’s own words
Each task is structured as a causal chain, where discovering information for one step is dependent on the successful completion of the previous one, stressing long-horizon planning and context retention. All tasks are grounded in the open web with objectively verifiable answer sets.
We introduce DeepSearchQA, a 900-prompt benchmark for evaluating agents on difficult multi-step information-seeking tasks across 17 different fields.
EarthVerse2026relevant
EarthVerse: Benchmarking Scientific Agents Across Dynamic Earth Systems and Natural Hazards
Goals to track inspect heterogeneous event packages and choose compatible evidence for a natural-hazard investigation; execute transparent calculations reconciling differences across sources; preserve provenance across evidence, scales, units, and calculations in the final answer; produce each of the fine-grained answer units defined by the task's executable ground truth
Must keep track of Agent must track evidence provenance, scales, units, and intermediate calculation results across a multi-stage investigation, since fine-grained answer units are each checked against executable ground truth.
structure: DAG-with-precedence · goals arrive: given-up-front · count: 405 reproducible tasks grounded in 199 d · subgoal credit: yes · 0 citations · also Scientific discovery
Horizonthe paper states none
the paper’s own words
We provide executable ground truth that decomposes each task into fine-grained answer units, together with task-specific rubrics that assess the supporting research process while allowing multiple valid paths... the best mean answer-unit accuracy is 84.65%, while the highest Strict@95 is only 34.81%.
Its 405 reproducible tasks are grounded in 199 documented events and 19 hazard families.
HAE-GEO2026relevant
Evaluating Deep-Search Agents under Hierarchical Web Evidence Poisoning
Goals to track search and retrieve web evidence for a consumer-decision query while a poisoning attack (of increasing sophistication, L1-L3) is present; recognize/verify suspicious evidence rather than adopting it uncritically; revise any already-adopted poisoned claims and recover to a trustworthy final recommendation
Must keep track of The agent must track which evidence it has already adopted (and whether that evidence was verified or poisoned), across a multi-turn Search-Scrape interaction, and must be able to revise earlier adopted claims before finalizing its recommendation.
structure: sequential-chain · goals arrive: given-up-front · count: 72,039 clean pages and 770 poisoned page · subgoal credit: yes · 0 citations · also Scientific discovery
Horizonthe paper states none
the paper’s own words
We introduce HAE-GEO, a benchmark that tracks the full trajectory from exposure to recovery under progressively more persuasive Web poisoning. Agents interact via a multi-turn Search-Scrape interface across three attack levels (L1 direct assertion, L2 contextual camouflage, and L3 apparent corroboration)... Evaluation combines deterministic behavioral measures with six semantic rubric dimensions.
MedProbeBench: Systematic Benchmarking at Deep Evidence Integration for Expert-level Medical Guideline
Goals to track retrieve, synthesize, and reason over large-scale external medical evidence to reach expert-level judgment; satisfy 1,200+ task-adaptive rubric criteria for one produced clinical guideline; ensure each of 5,130+ atomic claims in the guideline is precisely evidenced (fine-grained evidence verification)
Must keep track of System must track which rubric criteria have been satisfied and which atomic claims have been verified against source evidence as it synthesizes a guideline from large-scale external knowledge.
structure: set-of-independent · goals arrive: given-up-front · count: 1,200+ task-adaptive rubric criteria; 5, · subgoal credit: yes · 1 citations · also Scientific discovery
Horizonthe paper states none
the paper’s own words
(1) Holistic Rubrics with 1,200+ task-adaptive rubric criteria for comprehensive quality assessment, and (2) Fine-grained Evidence Verification for rigorous validation of evidence precision, grounded in 5,130+ atomic claims
we introduce MedProbeBench, the first benchmark leveraging high-quality clinical guidelines as expert-level references
Mr.LHDR2026in
Mr.LHDR: A Benchmark for Multimodal Real-World Long-Horizon Deep Research Agents
Goals to track derive an average of 12.1 necessary intermediate conclusions along a hidden Node-Relation dependency graph before reaching the final answer; integrate multimodal evidence (images, maps, PDFs, logos, charts, tables, video frames) where at least one non-text element changes the reasoning state; maintain correctness of intermediate conclusions consistent with annotated dependencies, not just the final answer
Must keep track of Agent must track which intermediate conclusions it has established so far, their dependency relationships, and integrate incremental multimodal evidence that can change the reasoning state, across long, irreducible evidence chains.
structure: DAG-with-precedence · goals arrive: given-up-front · count: average 12.1 necessary intermediate conc · subgoal credit: yes · 0 citations · also Scientific discovery · unit: other:intermediate-conclusions-per-question
Horizonavg 12.1 necessary intermediate conclusions per question (mean dependency depth 10.4)
the paper’s own words
Each question is constructed from a hidden Node-Relation graph and requires an average of 12.1 necessary intermediate conclusions with a mean dependency depth of 10.4 before reaching a short, unique, and verifiable answer.
ResearchClawBench2026relevant
ResearchClawBench: A Benchmark for End-to-End Autonomous Scientific Research
Goals to track re-discover a target published paper's scientific artifacts (methods, findings) using only related literature and raw data, with the target paper hidden; satisfy each of several expert-curated, weighted multimodal rubric criteria decomposed from the target scientific artifacts
Must keep track of Agent must track progress across an end-to-end research process (reviewing literature/raw data, designing methodology, producing results) and self-check against the (hidden) target paper's decomposed criteria without directly seeing it.
structure: hierarchical · goals arrive: given-up-front · count: 40 tasks across 10 scientific domains, e · subgoal credit: yes · 8 citations · also Scientific discovery
Horizonthe paper states none
the paper’s own words
Expert-curated multimodal rubrics decompose the target scientific artifacts into weighted criteria, enabling evaluation of target-paper-level re-discovery while leaving room for new discovery.
We evaluate seven autonomous research (auto-research) agents under a unified protocol and seventeen native LLMs through the lightweight ResearchHarness.
WANDR2026in
WANDR: A Benchmark for Wide and Deep Research
Goals to track discover a large set of entities satisfying specified criteria (breadth); investigate each discovered entity through multiple coordinated web searches (depth); return independently verifiable records with supporting sources/excerpts for each required (entity, relationship, evidence) combination; satisfy the full qualification-key hierarchy count (n x m x k records)
Must keep track of The agent must track which entities it has already discovered, which have been investigated to the required depth, and which evidence/sources have been independently verified, maintaining hierarchical completeness across potentially thousands of required records per task.
structure: hierarchical · goals arrive: given-up-front · count: a hierarchy with n companies, m employee · subgoal credit: yes · 1 citations · also Scientific discovery
Horizonthe paper states none
the paper’s own words
Record verdicts are aggregated into soft and hard precision, recall, and F1 scores that distinguish factual quality, coverage, and hierarchical completeness.
with targets ranging from dozens to thousands of records
DEER2025relevant
DEER: A Benchmark for Evaluating Deep Research Agents on Expert Report Generation
Goals to track produce an expert-level report satisfying each of 101 fine-grained rubric items across 7 dimensions/25 subdimensions; correctly cite and support both cited and uncited claims with verifiable evidence
Must keep track of The agent/judge must track satisfaction of 101 individual rubric items plus report-wide claim verification (both cited and uncited claims) across one long expert-level report, rather than judging the report as a single holistic pass/fail.
structure: set-of-independent · goals arrive: given-up-front · count: 101 fine-grained rubric items across 7 d · subgoal credit: yes · 11 citations · also Scientific discovery
Horizonthe paper states none
the paper’s own words
DEER systematizes evaluation criteria with an expert-developed taxonomy (7 dimensions, 25 subdimensions) operationalized as 101 fine-grained rubric items. We also provide task-specific Expert Evaluation Guidance to support LLM-based judging. In addition to rubric-based assessment, we propose a claim verification architecture that verifies both cited and uncited claims and quantifies evidence quality.
MEM1: Learning to Synergize Memory and Reasoning for Efficient Long-Horizon Agents
Goals to track answer each of many composed, interdependent objectives within one arbitrarily complex task sequence (e.g. a 16-objective multi-hop QA task); retrieve external information across internal retrieval QA, open-domain web QA, and multi-turn web shopping domains; consolidate memory turn-by-turn, discarding irrelevant/redundant information, to operate with constant memory
Must keep track of Agent must maintain a compact shared internal state that jointly supports memory consolidation and reasoning across many turns of a composed task sequence, integrating new observations while discarding irrelevant or redundant information.
structure: sequential-chain · goals arrive: given-up-front · count: up to 16 composed objectives in the mult · subgoal credit: no · 196 citations · also Tool &amp; API use · unit: other:composed-objectives-per-task
Horizon16
the paper’s own words
we propose a simple yet effective and scalable approach to constructing multi-turn environments by composing existing datasets into arbitrarily complex task sequences... show that MEM1-7B improves performance by 3.5x while reducing memory usage by 3.7x compared to Qwen2.5-14B-Instruct on a 16-objective multi-hop QA task, and generalizes beyond the training horizon.

Healthcare & clinical 8 artifacts · 2 state a horizon

AgentClinic2026relevant
AgentClinic: a multimodal benchmark for tool-using clinical AI agents
Goals to track engage in sequential clinical decision-making across a diagnostic encounter (patient interaction, exam/imaging requests); collect multimodal data under incomplete information via various tools (e.g., notebook, retrieval); reach a correct diagnosis across nine medical specialties and seven languages
Must keep track of Agent must track patient information collected so far (multimodal exam/imaging results), notes taken across a case (persisting via a notebook tool), and its evolving working diagnosis across the encounter.
structure: sequential-chain · goals arrive: emitted-by-environment-over-time · subgoal credit: yes · 21 citations · also Business, office &amp; enterprise work
Horizonthe paper states none
the paper’s own words
we introduce AgentClinic, a multimodal agent benchmark for evaluating LLMs in simulated clinical environments that include patient interactions, multimodal data collection under incomplete information, and the usage of various tools, resulting in an in-depth evaluation across nine medical specialties and seven languages.
We find that solving MedQA problems in the sequential decision-making format of AgentClinic is considerably more challenging, resulting in diagnostic accuracies that can drop to below a tenth of the original accuracy.
ClinEnv2026in
ClinEnv: An Interactive Multi-Stage Long Horizon EHR Environment for Agents
Goals to track progress through an ordered, per-case sequence of clinical decision stages; actively query four specialized information agents before committing to a decision at each stage; commit to correct medications, procedures, and diagnoses at each stage; avoid redundant information-gathering queries as the case progresses
Must keep track of Agent must actively query four specialized information agents at every decision stage before committing to medications, procedures, and diagnoses, tracking what has already been queried to avoid redundant queries as the case progresses.
structure: sequential-chain · goals arrive: given-up-front · subgoal credit: yes · 2 citations · also Business, office &amp; enterprise work
Horizonthe paper states none
the paper’s own words
Each case is automatically constructed into an ordered sequence of decision stages; at every stage the model must actively query four specialized agents before committing to medications, procedures, and diagnoses. ClinEnv scores both what the model decides, through deterministic ontology-grounded matching, and how it gathers information.
CodeClinic: Evaluating Automation of Coding Skills for Clinical Reasoning Agents
Goals to track in the longitudinal ICU-surveillance task, make a structured monitoring decision every four hours across 25 findings and eight clinical families for the duration of a patient trajectory; in the compositional information-seeking task, answer queries whose difficulty is stratified by compositional dependency depth across 259 tasks in 9 domains; synthesize and compose reusable clinical skills rather than relying on a fi
Must keep track of Agent must track the patient's evolving clinical trajectory and make a fresh structured decision every four hours in the longitudinal task, and/or track which reusable clinical skills it has synthesized/composed so far in the compositional task.
structure: DAG-with-precedence · goals arrive: emitted-by-environment-over-time · count: longitudinal: structured decisions every · subgoal credit: yes · 0 citations · unit: wall-clock-hours
Horizonevery 4 hours (decision cadence within a longitudinal ICU patient trajectory; total trajectory length not stated)
the paper’s own words
The benchmark contains two complementary tasks: longitudinal ICU surveillance and compositional information seeking. The longitudinal setting simulates monitoring patient trajectories with structured decisions every four hours across 25 findings and eight clinical families, while the compositional setting spans 63k instances across 259 tasks in nine domains and is stratified by compositional dependency depth to evaluate increasingly complex multi-step reasoning.
Healink2026relevant
Bridging the Post-discharge Gap: A Traceable Multi-agent Framework for Safe and Continuous Care
Goals to track maintain continuity of care across longitudinal, patient-specific post-discharge follow-up interactions; generate prescription-grounded, traceable responses reflecting the correct patient/phenotypic/intervention history; prevent cross-departmental drug conflicts while integrating fragmented histories across clinical departments
Must keep track of System must track a patient's full longitudinal clinical history (vectorized records, phenotypic/intervention dimensions) across departments and follow-up interactions, actively cross-checking for drug conflicts as new prescriptions are considered.
structure: DAG-with-precedence · goals arrive: emitted-by-environment-over-time · count: 400 continuous and 85 highly complex rea · subgoal credit: yes · 0 citations · also Multi-agent organisations &amp; societies
Horizonthe paper states none
the paper’s own words
We evaluated Healink on a dataset comprising 400 continuous and 85 highly complex real-world follow-up cases, alongside the webMedQA benchmark. In a rigorous single-blind evaluation conducted by clinical experts, the framework outperformed human physician baselines in both authoritativeness and clinical safety.
HealthAgentBench2026relevant
HealthAgentBench: A Unified Benchmark Suite of Realistic Agentic Healthcare Environments for Challenging Frontier AI Agents
Goals to track explore raw, heterogeneous healthcare data under minimal instructions; operate within a complex clinical environment to execute a multi-step, end-to-end task; develop research modeling pipelines over EHR data; complete tasks spanning 7 categories across the patient journey and multiple modalities
Must keep track of The agent must track what it has already discovered while exploring raw healthcare data and its intermediate modeling/analysis steps, so its final multi-step solution reflects the full workflow rather than a shortcut response.
structure: sequential-chain · goals arrive: implied-by-constraints · count: 54 tasks across 7 categories · subgoal credit: yes · 7 citations · also Business, office &amp; enterprise work
Horizonthe paper states none
the paper’s own words
Each task is designed to replicate an end-to-end clinical workflow: given minimal instructions, an agent must explore raw healthcare data, operate within a complex environment, and execute multi-step solutions that go beyond naive prompting.
We introduce HealthAgentBench, a suite of 54 agentic healthcare tasks across 7 categories each with its unique environment.
MedMCP-Calc2026relevant
MedMCP-Calc: Benchmarking LLMs for Realistic Medical Calculator Scenarios via MCP Integration
Goals to track proactively acquire relevant patient data from an EHR database; select the scenario-appropriate medical calculator among alternatives; perform multi-step computation using retrieved data and the selected calculator; retrieve external reference information when needed; complete each of 118 scenario tasks across 4 clinical domains
Must keep track of The agent must track which EHR fields it has retrieved via iterative SQL-based database interaction, which calculator it has selected, and intermediate computed values, since evaluation is process-level rather than only checking the final number.
structure: sequential-chain · goals arrive: implied-by-constraints · count: 118 scenario tasks across 4 clinical dom · subgoal credit: yes · 3 citations · also Business, office &amp; enterprise work
Horizonthe paper states none
the paper’s own words
MedMCP-Calc comprises 118 scenario tasks across 4 clinical domains, featuring fuzzy task descriptions mimicking natural queries, structured EHR database interaction, external reference retrieval, and process-level evaluation.
requiring proactive EHR data acquisition, scenario-dependent calculator selection, and multi-step computation
MedMemoryBench: Benchmarking Agent Memory in Personalized Healthcare
Goals to track accumulate and correctly retain/retrieve clinically relevant patient information across many sessions; sustain retrieval and reasoning robustness despite memory saturation from continued information influx; correctly evaluate an agent's memory 'live', as it is constructed, rather than only after the fact (evaluate-while-constructing streaming protocol)
Must keep track of The agent's memory system must retain and correctly retrieve clinically relevant information across roughly 2,000 sessions and 16,000 interaction turns per patient archetype, resisting degradation from memory saturation and noise while supporting complex medical reasoning.
structure: sequential-chain · goals arrive: emitted-by-environment-over-time · count: approximately 2,000 sessions and 16,000 · subgoal credit: yes · 1 citations · also Personal assistants &amp; long-term memory · unit: sessions
Horizonapproximately 2,000 sessions and 16,000 interaction turns
the paper’s own words
This process yields a massive, expertly validated dataset comprising approximately 2,000 sessions and 16,000 interaction turns. Crucially, MedMemoryBench departs from traditional static evaluations by pioneering an "evaluate-while-constructing" streaming assessment protocol, which precisely mirrors dynamic memory accumulation in production environments. Furthermore, we formalize and systematically investigate the critical phenomenon of memory saturation, where sustained information influx actively degrades retrieval and reasoning robustness.
MedAgentSim2025relevant
Self-Evolving Multi-Agent Simulations for Realistic Clinical Interactions
Goals to track as doctor agent, request relevant medical examinations and imaging results from a measurement agent across multi-turn conversations to reach a diagnosis; iteratively refine diagnostic strategy across successive patient interactions via self-improvement mechanisms
Must keep track of Doctor agent must track which exams/imaging results it has already requested and received, its evolving working diagnosis, and experience-based knowledge accumulated as it interacts with more patients over time.
structure: sequential-chain · goals arrive: emitted-by-environment-over-time · subgoal credit: yes · 48 citations · also Multi-agent organisations &amp; societies
Horizonthe paper states none
the paper’s own words
Unlike prior approaches, our framework requires doctor agents to actively engage with patients through multi-turn conversations, requesting relevant medical examinations (e.g., temperature, blood pressure, ECG) and imaging results (e.g., MRI, X-ray) from a measurement agent to mimic the real-world diagnostic process.

Education & tutoring 2 artifacts · 0 state a horizon

TeachArena2026relevant
TeachArena: Are Language Agents Ready for Realistic Teaching Work?
Goals to track infer a warranted pedagogical/teaching decision from evidence (professional pedagogical judgment); adapt tutoring support as the learner's state changes across a multi-turn tutoring session (situated multi-turn tutoring); carry an instructor's request through a learning-management system (LMS) to a completed, verified intervention (end-to-end LMS teaching workflow)
Must keep track of The agent must track the learner's evolving state across a multi-turn tutoring session, evidence supporting a pedagogical insight, and the instructor's request as it is carried through to a persistent, verified artifact or environment state in the LMS.
structure: hierarchical · goals arrive: given-up-front · count: 354 audited tasks; three complementary e · subgoal credit: no · 0 citations · also Business, office &amp; enterprise work
Horizonthe paper states none
the paper’s own words
Its 354 audited tasks are each built around a pedagogical insight, grounded in evidence, and evaluated with matched verifiers over observable turn-level responses, tutoring trajectories, and persistent artifacts or environment states.
they must infer a warranted teaching decision from evidence, adapt support as learner state changes, and carry an instructor's request through a learning-management system (LMS) to a completed, verified intervention
TutorBench2026relevant
DeepTutor: Towards Agentic Personalized Tutoring
Goals to track deliver citation-grounded tutoring on a specific problem; generate difficulty-calibrated follow-up questions matched to a learner's diagnosed knowledge gaps; continuously adapt personalization to the student's evolving needs across an interactive tutoring session
Must keep track of The agent must track the learner's diagnosed knowledge gaps and profile, the history of prior interactions, and adapt difficulty calibration accordingly across a multi-turn tutoring dialogue.
structure: sequential-chain · goals arrive: mixed:given-up-front-plus-emitted-by-environme · subgoal credit: yes · 0 citations · also Personal assistants &amp; long-term memory · unit: other:not-stated
Horizonthe paper states none
the paper’s own words
we introduce TutorBench, an interactive benchmark incorporating customized learner profiles grounded in university-level curricula across five domains. We further propose an LLM-based first-person interactive evaluation protocol that conducts assessments via a profile-driven student simulator.

TrustCoverage — and what you should not conclude

This catalogue is not exhaustive, and the gap is quantified rather than waved at. A capture–recapture estimate over two mechanism-independent acquisition routes puts at least 131 further in-scope artifacts outside it (a lower bound, not a point estimate). This methodology’s measured history is that true missing runs about twice the stated bound, and 62% of the corpus was found by only one acquisition route — a standard under-sampling signal. The honest working range is therefore roughly 130–320 missing, with something like 60–80% of the reachable population captured.

What the boundary is made of. The exclusions are deliberate and auditable, not accidental. Every one of the twenty most-cited papers absent from this corpus is either pre-2025 (WebArena, AgentBench, ALFWorld, GAIA, τ-bench, OSWorld, AppWorld) or infrastructure (Llama 3, DeepSeek-R1, GPT-4o) — not a single in-scope 2025–26 miss. Of 102 never-captured works cited by seven or more catalogued papers, 75% are pre-2025.

Loss terms, named. The retrieval filter’s false-negative rate was measured, not assumed: 60 papers the search filter discarded were judged blind, yielding zero clearly-relevant misses (≤5% by the rule of three; ≤13% counting every unsettled case). The largest gap is deliberate: the lower-priority tail was sampled rather than enumerated under a numeric threshold fixed in writing before any result was seen, leaving an estimated ~135 relevant papers unjudged. A depth pass that the search backend refused to run leaves 4,186 un-probed discards on thirteen query angles. Eight artifacts named in community registries have no indexed paper at all.

You may conclude that this contains the large majority of what a well-resourced search of the 2025–26 academic literature surfaces for this question, and that its boundary traces to stated rules with a reason attached to every exclusion. You may not conclude that it is complete, that any family’s count is the field’s true count, that a thin cell is genuinely thin, or that “no benchmark does X” — an absence claim needs a targeted check against the papers, not this corpus’s silence. The 2026 half is the least-verified: only one of 85 independently-enumerated canonical benchmarks was from 2026, so that stratum rests on retrieval alone.

Refresh trigger. Several qualifying artifacts appear per week and 318 of 480 entries are from 2026 alone. Treat this as stale after 6–8 weeks; re-running forward-citation expansion plus the registry enumeration recovers most of the drift without a full rebuild.

asta-find323forward-citation304web-registry43parametric22
How many catalogued artifacts each acquisition route found. Routes overlap, so the bars exceed n=480.
how this number was checked

Claim. Acquisition modality yields over the core are concentrated: paper-finder sweeps and forward-citation each account for roughly two thirds of the core, registry enumeration and parametric memory for far less.

Why it holds. Counts come from the provenance union on every candidate row, which records every modality that found a paper rather than the first one.

What would break it. unless the low registry and parametric shares are read as those modalities being unproductive — the anchor analysis shows the opposite, that registry enumeration surfaced 17 core papers no other route had judged.

480 core rows, provenance union; a paper found by two modalities counts in both.

How performed. Corpus: all 4,327 candidates and the 480-artifact catalogue. Method: capture–recapture across acquisition routes of independent lineage, reference-pool recall stratified by citation frequency, citation-graph centrality triage, plus two external anchors (an independently enumerated canon list and sixteen maintained community registries). Estimators considered but withheld, with reasons, are listed in the coverage file. Frame: estimates over the full candidate pool; ring counts over the charter core, n=480. Key limit: no enumerable denominator exists for “all 2025–26 papers of this kind”, so the open part of the question supports bounds only, never a percentage of the field.

ProvenanceHow this was built

Multi-route acquisition (55 search angles across two query waves, citation-graph expansion, sixteen community registries, and an independent enumeration from prior knowledge) produced 4,327 candidates. 2,191 were judged against a five-criterion rubric by a fleet of graded judges; calibration items planted invisibly in every batch scored 86% exact and 95% same-side agreement with none of the judges running looser than the gold standard. The surviving 508 then received per-paper extraction under two lenses, after which extractors reading the full records caught 28 further artifacts that failed the scope rules the abstract-level pass could not see — leaving 480.

Three papers were placed in every extraction batch as consistency probes. All eight extractors identified them identically; their structure vocabulary varied, which is the honest read on how much weight a single structural label can carry.

probe artifactextractorsidentified asstructure label spread
Vending-Bench8Vending-Benchopen-ended
OdysseyBench8OdysseyBenchDAG-with-precedence, sequential-chain
Factorio Learning Environment (FLE)8Factorio Learning Environment (FLE)hierarchical, other:fixed-independent-tasks-in-lab-play-open-ended-in-

The package. Every number on this page traces to a shipped data file: the full catalogue as CSV and JSON, the chart series, the headline statistics, the pooled-claim audit, and the consistency-probe results.