All 480 benchmarks, datasets, environments and scenario suites published in 2025–2026 whose settings require an agent to pursue and track multiple goals and subgoals. Software-engineering and other code-centric domains are excluded by design. Complete as-of 2026-09-14; the field moves fast, so treat it as stale after 6–8 weeks.
Reading the table. Goals to track and must keep track of are extracted per paper from its own text. Horizon shows the paper's own figure in its own unit — and none stated where the paper never quantifies one, which is the case for 334 of 480 entries. Credit records whether subgoal-level partial credit exists. Cites is the Semantic Scholar citation count with a per-month rate beside it — with 318 of 480 entries published in 2026, this column mostly measures age, not quality, and 126 entries sit at zero. Every name links to the paper.
| OS & computer use24 artifacts | ||||||||
| CareFlow (benchmark) / CarePilot (agent framework)
CarePilot: A Multi-Agent Framework for Long-Horizon Computer Task Automation in Healthcare |
OS & computer use+ Healthcare | 2026 | complete each of 8-24 consecutive GUI decisions/actions required by one long-horizon clinical software workflow (e.g. DICOM viewer, EHR, lab information system); correctly ground next-action predictions in the current visual interface and system state throughout the workflow | The agent must track the current visual interface state and prior decisions across 8-24 consecutive steps of a clinical software workflow, using dual-memory (long-term and short-term experience) to predict the next semantic action correctly. | sequential-chain | yes | each task comprises 8-24 consecutive decisions/stepsagent-steps | 61.0/mo |
| ChainWorld
ChainWorld: Composing Long-Horizon Desktop Workloads from Atomic OSWorld Tasks |
OS & computer use | 2026 | complete a chain of 2 to 4 sequentially composed atomic OSWorld desktop tasks while sustaining state across all of them; under single-turn evaluation, complete the whole chain from one combined prompt; under multi-turn evaluation, complete each task as it is revealed one at a time while retaining session state | Agent must sustain desktop application/file state across multiple chained objectives and, in multi-turn mode, manage session continuity as tasks are revealed one at a time. | sequential-chain | yes | two to fourother:atomic-tasks-per-chain | 0 |
| CutVerse
CutVerse: A Compositional GUI Agents Benchmark for Media Post-Production Editing |
OS & computer use+ Web & GUI | 2026 | complete a long-horizon, compositional media-editing task grounded in an authentic editing workflow using one of 7 professional applications (e.g., Premiere Pro, Photoshop); correctly execute tightly coupled, dense multimodal interaction sequences within the application's GUI | Agent must track compositional GUI state (selected layers/clips/tools) across a tightly coupled, dense-interface interaction sequence throughout a long-horizon editing task. | sequential-chain | yes | none stated | 0 |
| DeskCraft
DeskCraft: Benchmarking Desktop Agents on Professional Workflows and Human-in-the-Loop Collaboration |
OS & computer use+ Web & GUI | 2026 | complete long-horizon professional creative/engineering workflows (design, video, audio, 3D creation) requiring over 50 execution steps; proactively seek necessary information from the user under uncertainty (agent-initiated clarification); correctly handle user-initiated interruptions during execution and incorporate post-turn feedback after signaling completion | The agent must track its own multi-step progress (over 50 execution steps for long-horizon tasks) within professional creative/engineering software, while also tracking pending clarification needs, handling user interruptions mid-execution, and incorporating post-completion feedback into further revisions. | sequential-chain | yes | long-horizon tasks require over 50 execution stepsagent-steps | 20.67/mo |
| DevicesWorld
DevicesWorld: Benchmarking Cross-Device Agents in Heterogeneous Environments |
OS & computer use | 2026 | acquire information on one device (e.g. phone) needed to complete a task; process/transform that information on a different device (e.g. desktop); deliver or display the final result on yet another device; satisfy cross-device dependencies and rule-based verifiers across the whole task | The agent must track which device holds which piece of needed information, what has already been acquired/processed/delivered across devices, and whether all task conditions (verified from device states and generated files) are jointly satisfied. | DAG-with-precedence | yes | each task permits at most 50 interaction stepsactions | 0 |
| GUITestScape
GUITestScape: Towards Open-set Evaluation on Exploratory GUI Testing |
OS & computer use+ Web & GUI | 2026 | autonomously navigate an app to discover defects with no predefined test script; distinguish and separately diagnose interaction defects versus display defects; decompose the testing trajectory into independently diagnosable capabilities rather than a single end-state judgment | The agent must track which parts of an app it has already explored and which candidate defects (interaction or display) it has already investigated, to continue open-ended exploration without redundant re-testing or missing an area. | open-ended | yes | none stated | 0 |
| HeraBench
Beyond Global Replanning: Hierarchical Recovery for Cross-Device Agent Systems |
OS & computer use | 2026 | decompose a cross-device task and assign subtasks across heterogeneous devices via unified API-CLI-GUI execution; recover from injected device-local strategy failures without escalation; escalate to orchestrator-level global replanning when a failure exceeds device-local recovery scope; complete an end-to-end cross-device workflow over Linux and Android devices | The system must track per-device execution-strategy state, a compact cross-layer failure abstraction distinguishing device-local from global failure scope, and overall workflow progress across Linux and Android devices. | hierarchical | yes | none stated | 0 |
| JarvisGUI
JarvisGUI: Towards Cross-Device GUI Agents with Dynamic Task Composition |
OS & computer use+ Web & GUI | 2026 | transfer intermediate results correctly between two or more heterogeneous devices/platforms; maintain shared state consistently across Android, Windows, and Ubuntu platforms; compose a multi-step, cross-device workflow from input-output-typed GUI sub-tasks; complete a workflow requiring four or more chained subtasks despite compounding dependency-tracking difficulty | The agent must track shared state and intermediate results as they transfer across heterogeneous devices/platforms, verifying that dependency steps (e.g. file transfer, renaming) have actually executed rather than being falsely treated as completed. | DAG-with-precedence | no | none statedother:not-stated | 0 |
| MacAgentBench
MacAgentBench: Benchmarking AI Agents on Real-World macOS Desktop |
OS & computer use+ Web & GUI | 2026 | complete a multi-application task on real macOS desktop software using both GUI and CLI interaction; reach each of several fine-grained, capability-annotated checkpoints within a multi-application task | Agent must track partial progress across multiple checkpoints within a multi-application task, since 'models with similar Pass@1 can differ substantially in sub-goal completion' according to the fine-grained metrics. | hierarchical | yes | none stated | 20.67/mo |
| MedCUA-Bench
MedCUA-Bench: A Screenshot-Only Benchmark for Clinical Computer-Use Agents |
OS & computer use+ Healthcare | 2026 | complete clinical computer-use scenarios across 10 medical domains; satisfy paired intent-level and step-level goals within each task; avoid violations across five clinical safety dimensions while completing the task | Agent must track progress toward the high-level clinical intent, the individual UI steps needed to realize it, and continuously check five clinical safety dimensions across the task. | hierarchical | yes | none stated | 10.33/mo |
| MyPCBench
MyPCBench: A Benchmark for Personally Intelligent Computer-Use Agents |
OS & computer use+ Assistants & memory | 2026 | complete a personal-assistant task requiring the agent's own accumulated context/history; operate correctly across 17 simulated real-world web applications seeded for one canonical persona; coordinate actions that span many applications on the same Linux desktop; sustain a long trajectory without abandoning early or looping unproductively | The agent must track the canonical persona's context, historical data, and logged-in-account state across a full Linux desktop stack, coordinating GUI and bash actions across many applications without unproductive step looping. | DAG-with-precedence | no | mean 22-31 steps (GPT family, Sonnet); mean 52-85 steps (Opus, Qwen models)agent-steps | 10.33/mo |
| OmniGUIRewardBench (OGRBench) / OS-Themis
OS-Themis: A Scalable Critic Framework for Generalist GUI Rewards |
OS & computer use+ Web & GUI | 2026 | decompose a GUI trajectory into verifiable milestones; audit the evidence chain for each milestone before rendering a final reward verdict; produce a scalable, accurate outcome reward usable for RL training or trajectory filtering | The critic must track which milestones in a trajectory have been reached, the evidence supporting each, and audit the full evidence chain before issuing a final reward — i.e. it tracks reward-relevant progress rather than the acting agent's own task state. | hierarchical | yes | none stated | 81.33/mo |
| OS-Marathon
OS-Marathon: Benchmarking Computer-Use Agents on Long-Horizon Repetitive Tasks |
OS & computer use | 2026 | complete each of 242 long-horizon, repetitive computer-use tasks across 2 domains (e.g., processing expense reports from receipts, entering grades from exam papers); correctly repeat the same structured sub-workflow logic across many data items within one task; learn workflow logic from a few-shot condensed demonstration and generalize it to larger, unseen data collections | The agent must track its position within a long, repetitive workflow (e.g., which receipt or exam paper it is currently processing), correctly apply the learned sub-workflow logic consistently, and avoid drift or errors accumulating across many repeated sub-tasks. | sequential-chain | yes | none stated | 51.67/mo |
| OS-Marathon
OS-Marathon: Benchmarking Computer-Use Agents on Vast-Horizon, Repetitive Tasks |
OS & computer use | 2026 | complete many recurring per-instance sub-workflows within a vast-horizon repetitive task (e.g., process each receipt in a stack); sustain correct execution as length scales with data volume; learn/personalize the recurring sub-workflow logic from a single human demonstration (GraphDemo) | Agent must correctly and repeatedly apply the same recurring sub-workflow logic across many per-instance items, with execution length scaling with the volume of data to process. | set-of-independent | unclear | none stated | 0 |
| OSWorld 2.0
OSWorld2.0: Benchmarking Computer Use Agents on Long-Horizon Real-World Tasks |
OS & computer use | 2026 | complete each of 108 realistic end-to-end long-horizon computer-use workflows; perform cross-source reasoning across authentic input artifacts; infer implicit state and recover hidden state the task depends on; handle streaming interaction and dynamic environment changes mid-task; operate under safety-sensitive execution constraints (audited separately) | The agent must track constraints, cross-source information, implicit/hidden state, and mid-task updates continuously across an average of 318 tool calls (up to a 500-step budget) per task, rather than resolving everything from the initial instruction alone. | sequential-chain | yes | median about 1.6 hours per task for human users; average of 318 tool calls per task with Claude Opus 4.7 using maximum thinking (vs about 30 in OSWorld 1.0); 500-step completion budgettool-calls | 155.0/mo |
| WindowsWorld
WindowsWorld: A Process-Centric Benchmark of Autonomous GUI Agents in Professional Cross-Application Environments |
OS & computer use+ Web & GUI | 2026 | complete each of a task's average 5.0 sub-goals across up to 17 desktop applications; perform conditional judgment and reasoning that spans 3 or more applications; coordinate a workflow across multiple applications for 78% of tasks that are inherently multi-application | The agent must track progress across a sequence of sub-goals spanning multiple desktop applications, carrying forward state and conditional-judgment results between applications, within a per-difficulty-level step budget (15-40 steps). | sequential-chain | yes | max step budget 15 (L1) / 25 (L2) / 40 (L3) / 20 (L4); avg minimum action steps 9.67-27.81 by levelagent-steps | 81.6/mo |
| AgentSynth
AgentSynth: Scalable Task Generation for Generalist Computer-Use Agents |
OS & computer use | 2025 | complete long-horizon computer-use tasks composed of a controllable number of simple subtasks (difficulty levels 1-6); generalize from simple generation-time subtasks that become significantly harder once composed | Agent must carry state correctly across a chain of composed subtasks, since success rate drops steeply as the number of chained subtasks increases from difficulty level 1 to 6. | sequential-chain | no | none stated | 322.13/mo |
| KGCE
KGCE: Knowledge-Augmented Dual-Graph Evaluator for Cross-Platform Educational Agent Benchmarking with Multimodal Language Models |
OS & computer use+ Web & GUI | 2025 | complete a school-specific-software task correctly on Windows, Android, or across both platforms in coordination; satisfy each of the multiple sub-goals a complex task is decomposed into (e.g. one complex task decomposed into five subtasks), each independently verified | The agent must track completion status of each decomposed sub-goal (via the dual-graph evaluation framework) across potentially multiple platforms and school-specific private-domain software, whose structural specifics are not otherwise well understood by general-purpose agents. | DAG-with-precedence | yes | none stated | 0 |
| MMBench-GUI
MMBench-GUI: Hierarchical Multi-Platform Evaluation Framework for GUI Agents |
OS & computer use+ Web & GUI | 2025 | understand GUI content and ground specific elements accurately (Element Grounding); automate a full task end-to-end within one application/platform (Task Automation); collaborate across multiple tasks/applications requiring cross-platform generalization (Task Collaboration) | Agent requires long-context memory across many actions, tracking of task-planning state, and long-term reasoning to sustain grounded, efficient action sequences without redundant steps across the hierarchy. | hierarchical | yes | none stated | 614.36/mo |
| Mobile-Eval-E
Mobile-Agent-E: Self-Evolving Mobile Assistant for Complex Tasks |
OS & computer use+ Web & GUI | 2025 | break a complex mobile task into subgoals via a Manager agent; execute fine-grained low-level actions per subgoal via an Operator agent; verify and correct action errors via an Action Reflector; aggregate information across steps via a Notetaker; complete long-horizon, multi-app mobile interactions | The system must track the Manager's subgoal plan, the Notetaker's aggregated information, self-evolved long-term Tips/Shortcuts memory, and cross-app interaction state across a single complex mobile task. | hierarchical | yes | none stated | 1346.7/mo |
| OmniBench / OmniEval
What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities |
OS & computer use+ Web & GUI | 2025 | complete a graph-structured task composed of multiple synthesized subtasks of controllable complexity; satisfy subtask-level correctness at each node of the task graph, not just the final outcome; demonstrate each of 10 evaluated virtual-agent capabilities | Agent must track progress through the task graph's subtask nodes and correctly compose primitive actions across the graph to satisfy subtask-level and graph-based metrics. | DAG-with-precedence | yes | none stated | 90.6/mo |
| PC-Eval
PC-Agent: A Hierarchical Multi-Agent Collaboration Framework for Complex Task Automation on PC |
OS & computer use+ Multi-agent orgs | 2025 | decompose complex user instructions into Instruction-Subtask-Action levels; track progress across interdependent subtasks via a Progress agent; make step-by-step decisions via a Decision agent; provide timely bottom-up error feedback and adjustment via a Reflection agent; complete each of 25 real-world complex PC instructions | The multi-agent system must track subtask decomposition and progress (via the Progress agent), current decision state (via the Decision agent), and bottom-up error feedback (via the Reflection agent) across intra- and inter-app workflows on a PC. | hierarchical | yes | none stated | 392.05/mo |
| SEAgent (OS-World novel software environments)
SEAgent: Self-Evolving Computer Use Agent with Autonomous Learning from Experience |
OS & computer use | 2025 | autonomously master a novel, unfamiliar software environment through experiential trial-and-error; progressively tackle auto-generated tasks organized from simple to complex (curriculum); assess step-wise trajectory correctness via a World State Model; integrate individual specialist experience into a stronger generalist computer-use agent | The agent must track its step-wise trajectory quality (via the World State Model), its progress along an increasingly difficult auto-generated curriculum, and accumulated experiential insights to be integrated from specialist to generalist training. | hierarchical | yes | none stated | 614.69/mo |
| UI-CUBE
UI-CUBE: Enterprise-Grade Computer Use Agent Benchmarking Beyond Task Accuracy to Operational Reliability |
OS & computer use+ Business & enterprise | 2025 | complete simple UI interactions (136 tasks); complete complex copy-paste workflows spanning multiple applications (50 tasks); complete complex enterprise application scenarios (40 tasks); maintain operational reliability (not just functional correctness) under systematic interface variation and multi-resolution testing | Agent must track state across chained multi-step workflows (e.g., what was copied and where it must be pasted), interface variations, and multiple screen resolutions, verifying success via application-state validation. | hierarchical | yes | none stated | 30.3/mo |
| Business & enterprise105 artifacts | ||||||||
| Evaluating Deep Research Agents on Expert Consulting Work: A Benchmark with Verifiers, Rubrics, and Cognitive Traps
Evaluating Deep Research Agents on Expert Consulting Work: A Benchmark with Verifiers, Rubrics, and Cognitive Traps |
Business & enterprise | 2026 | produce a multi-document, decision-grade consulting deliverable from SME-authored prompts; pass deterministic binary verifiers (mean 14.9 per task); satisfy each of a five-criterion SME rubric (Data Integrity, Analytical Rigor, Relevance & Focus, Execution Precision, Format & Deliverability); avoid embedded cognitive traps (human-error mimicry, deterministic precision traps) that penalize surface-pattern reasoning | The agent must track which of the ~14.9 verifiers per task it has satisfied, avoid embedded cognitive traps requiring reconciliation against context, and keep its final deliverable consistent with all five rubric criteria simultaneously. | set-of-independent | yes | none statedother:not-stated | 10.25/mo |
| 50-task hotel expense benchmark (Dynamics 365 F&O)
Less Context, Better Agents: Efficient Context Engineering for Long-Horizon Tool-Using LLM Agents |
Business & enterprise | 2026 | itemize every line item of a hotel expense report completely and accurately using enterprise MCP tools; avoid context overflow / stale-state errors while repeatedly calling verbose enterprise tool APIs across the itemization workflow | The agent must track already-itemized line items, running token/context budget, and must avoid stale or overflowing tool-call history while repeatedly invoking enterprise MCP tools to reach complete, accurate itemization. | set-of-independent | yes | full-context retention: 1,480,996 tokens and 14.56 hours per benchmark (50 tasks x 5 runs); pruning to last 5 tool calls: 535,274 tokens and 5.39 hours; pruning + summarization: 553,374 tokens and 5.79 hourswall-clock-hours | 10.33/mo |
| AD-Bench
AD-Bench: A Real-World, Trajectory-Aware Advertising Analytics Benchmark for LLM Agents |
Business & enterprise | 2026 | answer a real user marketing-analysis request via multi-round, multi-tool collaboration; maintain trajectory coverage consistent with an expert tool-call trajectory; produce answers that remain correct/consistent with a continuously evolving production advertising platform (dynamic ground truth); succeed across three stratified difficulty levels (L1-L3) requiring increasing multi-round, multi-tool collaboration | The agent must track which professional tools it has called and in what order across multiple rounds, since evaluation jointly measures end-to-end answer correctness (Pass@k) and how well its trajectory covers the expert reference trajectory, especially as this compounds at higher difficulty levels. | sequential-chain | yes | none stated | 30.43/mo |
| AeroCopilotBench (ACOE)
AeroCopilotBench: A Two-Tier Benchmark for Evaluating LLM Agents as Aviation Copilots in an Interactive Virtual Cockpit Environment |
Business & enterprise | 2026 | diagnose faults and operate aircraft systems through standardized tool interfaces during emergency/abnormal procedures; achieve all task goal-conditions for each of 73 Tier-2 emergency/abnormal tasks; avoid violating any hard safety constraint while progressing toward goals | Agent must interpret cockpit state, diagnose faults, and operate aircraft systems while simultaneously tracking goal-condition progress and hard safety-constraint compliance across the executable procedure. | DAG-with-precedence | yes | none stated | 0 |
| Agentic ERP
Agentic ERP: Multi-Agent Large Language Model Architecture for Autonomous Enterprise Resource Planning |
Business & enterprise+ Multi-agent orgs | 2026 | execute end-to-end business workflows across ERP functional boundaries (e.g. procurement, inventory, crisis response); resolve cross-functional crisis tasks via a Planner-Executor-Reflector-Responder orchestration; sustain a full simulated year of ERP operation while avoiding stockouts, compared against rule-based RPA and no-intervention baselines | The agent(s) must track the structured enterprise state (inventory, demand stream, crisis conditions) across a full simulated year, coordinate across role-aligned agents via externalised grading criteria and sprint contracts, and avoid accumulating stockouts as the rule-based baseline does. | hierarchical | yes | a 365-day agent-in-the-loop simulation (a simulated year of ERP operation)simulated-days | 0 |
| AgenticPay
AgenticPay: A Multi-Agent LLM Negotiation System for Buyer-Seller Transactions |
Business & enterprise+ Multi-agent orgs | 2026 | negotiate to reach agreement given private constraints and product-dependent valuations; maximize feasibility, efficiency, and welfare across a negotiation; successfully complete each of 110+ tasks ranging from bilateral bargaining to many-to-many markets | Agents must track their own private constraints/valuations, the evolving history of natural-language offers and counteroffers across negotiation rounds, and, in many-to-many markets, the state of other concurrent negotiations. | sequential-chain | yes | none stated | 131.86/mo |
| AgenticVBench
AgenticVBench: Can AI Agents Complete Real-World Post-Production Tasks? |
Business & enterprise | 2026 | complete each of 100 agentic post-production tasks across 4 task families; compose capabilities across text, image, audio, and video understanding within one task; plan and execute long-horizon multi-step production workflows using appropriate tools | The agent must track progress through a long-horizon, multi-modal production workflow and correctly sequence tool use, since tasks are constructed from real production workflows contributed by industry experts and evaluated jointly by programmatic verifiers and expert rubrics. | other (singleton) | yes | none stated | 30.75/mo |
| Agents' Last Exam (ALE)
Agents' Last Exam |
Business & enterprise | 2026 | complete a long-horizon, economically valuable, real-world professional task with a verifiable outcome, drawn from a specific occupational sub-field; achieve sustained performance across the task rather than a single-shot correct answer, particularly on the hardest ('last-exam') tier | The agent must sustain performance on long-horizon, economically valuable real-world tasks whose outcomes are verifiable, across an occupational taxonomy spanning 55 sub-fields and 13 industry clusters, with the hardest tier proving far from saturated (average full pass rate below 1%). | set-of-independent | yes | none stated | 124.0/mo |
| APEX-Agents
APEX-Agents |
Business & enterprise | 2026 | execute a long-horizon, cross-application task created by investment-banking analysts, management consultants, or corporate lawyers; navigate realistic work environments containing files and tools to produce the required deliverable | Agent must track which files and tool states it has already produced/consulted while navigating a cross-application work environment, using rubrics and gold outputs to determine success (Pass@1). | sequential-chain | unclear | none stated | 202.5/mo |
| AutomationBench
AutomationBench |
Business & enterprise | 2026 | discover the relevant REST API endpoints needed for a cross-application workflow (CRM, inbox, calendar, messaging, etc.); follow layered business-policy rules while writing data; navigate environments containing irrelevant or misleading records without being derailed; get correct data into the right systems by the end of the workflow (end-state correctness) | Agent must track which endpoints it has discovered across multiple applications, which business-policy rules apply to each write, and whether records encountered are relevant or misleading, to ensure the correct final data lands in each system. | DAG-with-precedence | yes | none stated | 10.2/mo |
| BigFinanceBench
BigFinanceBench: A Workflow-Grounded Benchmark for Financial-Research Agents |
Business & enterprise | 2026 | produce an auditable financial-research derivation (source choice, period/accounting definition, assumptions, calculation), not just a final answer; satisfy each of many independently checkable rubric points per item (36,241 points across 928 items); correctly complete each of 928 expert-authored open-ended financial-research tasks | The agent must track and justify each auditable component of its derivation (data source, period/accounting definition, assumptions, calculation steps) so that an independent rubric can check each step, not just the final numeric answer. | hierarchical | yes | none stated | 51.67/mo |
| BlueFin
BlueFin: Benchmarking LLM Agents on Financial Spreadsheets |
Business & enterprise | 2026 | synthesize new spreadsheet content/formulas from source data; manipulate existing financial spreadsheet workbooks (edits, restructuring); comprehend and answer questions about spreadsheet content; satisfy each of up to 3,225 granular rubric criteria across 131 tasks | The agent must track and satisfy many granular, LM-judge-scored rubric criteria across synthesis, manipulation, and comprehension actions within one workbook, maintaining correctness as the workbook's formulas/state evolve ('dynamic correctness'). | hierarchical | yes | tasks require at least 45+ minutes of work for a human analyst to complete from scratchhuman-expert-hours | 10.25/mo |
| Business Arena
Business Arena: Benchmarking LLM Agents in a Realistic Marketplace |
Business & enterprise | 2026 | infer business opportunities from partial market signals; commit capital under uncertainty when buying from suppliers; set/adjust pricing to sell to buyers profitably; satisfy regulatory obligations before trading legally; sustain profitable operation of a cross-border shop over a long horizon | The agent must track capital, delayed and coupled consequences of sourcing/pricing/recovery decisions, market conditions calibrated from real data, and regulatory obligations across a long-horizon cross-border trading operation. | open-ended | no | none stated | 0 |
| CEO-Bench
CEO-Bench: Can Agents Play the Long Game? |
Business & enterprise+ Games & IF | 2026 | operate a startup for 500 days, managing pricing, marketing, budgeting, and other business aspects; acquire information from noisy, interconnected business databases and translate it into strategy; adapt strategy to a changing world/market over the run; orchestrate many interdependent business decisions toward a coherent long-term financial goal | The agent must track noisy, interconnected business databases (churn regimes, billing timing, customer losses, projected cash) and coordinate many interdependent decisions via code across a 500-day simulated run. | open-ended | yes | 500simulated-days | 31.0/mo |
| CEO-Bench
Can LLMs Be CEOs? Benchmarking Strategic Resource Reallocation with Multi-Role Agent Simulation |
Business & enterprise | 2026 | integrate conflicting recommendations from four role-conditioned C-suite advisors (CFO, CTO, COO, CMO) into one allocation plan; redirect capital across business units under information asymmetry and organizational constraints; synthesize advice consistently across multiple rounds with temporal dependencies | Agent (as CEO) must track and reconcile conflicting, privately-signaled recommendations from four C-suite advisors across a multi-round resource-reallocation process, remembering prior decisions for history-sensitive judgment. | DAG-with-precedence | yes | none stated | 10.33/mo |
| CHI-Bench (χ-Bench)
CHI-Bench: Can AI Agents Automate End-to-End, Long-Horizon, Policy-Rich Healthcare Workflows? |
Business & enterprise+ Healthcare | 2026 | ground each decision in a large library of medical, insurance, and operational policy rules; play multiple roles within a single task, handing off between them; conduct multilateral multi-turn dialogs (e.g. peer-to-peer review, patient outreach) as intermediate workflow steps; drive a clinical case to a terminal status via tool calls and role-artifact writing | The agent must track its current role and handoffs to other roles, policy compliance against a 1,290+ document handbook, and the clinical case's evolving status across a high-fidelity simulator of 20 apps, until it reaches a terminal status. | DAG-with-precedence | no | none statedother:not-stated | 61.5/mo |
| Claw-Eval-Live
Claw-Eval-Live: A Live Agent Benchmark for Evolving Real-World Workflows |
Business & enterprise | 2026 | complete end-to-end units of work across software tools, business services, and local workspaces per release; pass controlled tasks with fixed fixtures/services/workspaces/graders reflecting current public workflow-demand signals; handle HR, management, and multi-system business workflows as well as local workspace repair | Agent must produce verifiable execution traces, audit logs, and consistent service/workspace state across each end-to-end task, since grading uses deterministic checks when evidence is sufficient and structured LLM judging only for semantic dimensions. | set-of-independent | no | none stated | 132.6/mo |
| ClawMark
ClawMark: A Living-World Benchmark for Multi-Turn, Multi-Day, Multimodal Coworker Agents |
Business & enterprise | 2026 | carry a professional coworker task forward correctly across multiple working days; detect and adapt to exogenous environment updates (new emails, calendar shifts, KB edits) that occur between turns; satisfy each of the task's deterministic Python checkers over post-execution service state (mean 15.4 checkers/task); coordinate consistent state across five stateful sandboxed services (filesystem, email, calendar, knowledge base, spreadsheet) | The agent must track evolving state across five sandboxed services (filesystem, email, calendar, knowledge base, spreadsheet) over multiple working days, detecting and incorporating exogenous updates injected between turns, verified by up to 29 deterministic checkers per task. | sequential-chain | yes | 2-6 turns per task (mean 3.6); one turn = one in-universe working daysimulated-days | 153.0/mo |
| ClawsBench
ClawsBench: Evaluating Capability and Safety of LLM Productivity Agents in Simulated Workspaces |
Business & enterprise | 2026 | complete single-service productivity tasks (e.g. Gmail, Slack, Calendar, Docs, Drive) correctly and safely; complete cross-service workflows spanning multiple mock services while maintaining consistent state; avoid unsafe/irreversible actions in safety-critical scenarios | Agents must track persistent state across five mock services (Gmail, Slack, Calendar, Docs, Drive), recognize safety-critical constraints to avoid irreversible/unsafe actions, and coordinate behavior across services via a meta-prompt layer. | DAG-with-precedence | yes | none stated | 255.0/mo |
| CoffeeBench
CoffeeBench: Benchmarking Long-Horizon LLM Agents in Heterogeneous Multi-Agent Economies |
Business & enterprise+ Multi-agent orgs | 2026 | maximize cumulative net income for one's own firm (coffee roaster) over a long simulated run; manage cash, inventory, and pricing on an ongoing basis; communicate and transact with other heterogeneous firms (farmers, roasters, retailers) to secure supply/demand | The agent must track its own cash, inventory, and pricing state day by day over the 90-day simulation, plus the state of its ongoing communications/transactions with other firms, to maximize cumulative net income rather than any single day's outcome. | sequential-chain | yes | 90-day simulationsimulated-days | 62.0/mo |
| Company World Model dry-lab benchmark (45 retrospective cases)
Do AI-Native Biotechs Need Departments? Benchmarking Company World Models for AI-Driven Drug Development |
Business & enterprise | 2026 | maintain and update a persistent asset-to-value state (Live Asset Value Record) across scientific, regulatory, BD, commercial, financial, and execution constraints; resolve 45 retrospective public-information decision cases with hidden outcomes under strict time cutoffs; achieve success via external BD deals, regulatory approval/launch, and revenue discipline | The architecture must maintain a persistent, continuously updated asset-to-value state record across scientific, regulatory, BD, commercial, financial, and execution constraints via Deal/Approval/Revenue/Investment Arbiter loops. | DAG-with-precedence | no | none stated | 0 |
| CoreCraft (EnterpriseBench)
EnterpriseBench Corecraft: Training Generalizable Agents on High-Fidelity RL Environments |
Business & enterprise | 2026 | perform multi-step, domain-specific customer-support work across a simulated enterprise with 2,500+ entities across 14 entity types; satisfy all expert-authored rubric criteria for a given task using 23 available tools | Agent must track entity state across the enterprise simulation and verify each expert-authored rubric criterion is satisfied before the task counts as solved. | DAG-with-precedence | yes | none stated | 71.0/mo |
| DocOps
DocOps: A Verifiable Benchmark for Autonomous Agents in Complex Document Operations |
Business & enterprise | 2026 | perform atomic document-operation actions correctly (e.g. edits without destructive metadata corruption); complete escalating, more complex, more tightly coupled document-workflow tasks built from those atomic operations; maintain long-term state tracking and global document consistency across the workflow; correctly verify (not just superficially check) that a document edit was semantically achieved | The agent must maintain long-term state tracking of the document's evolving structure/content and verify (rather than superficially assume) that each operation was semantically correct, to preserve global document consistency across highly coupled, long-range tasks. | hierarchical | yes | none stated | 10.5/mo |
| DORA (Disaster Operational Response Agent benchmark)
Can LLM Agents Respond to Disasters? Benchmarking Heterogeneous Geospatial Reasoning in Emergency Operations |
Business & enterprise | 2026 | perform disaster perception from heterogeneous geospatial imagery; conduct spatial relational analysis over roads/population/facilities; plan rescue and evacuation operations; reason about temporal evolution of the disaster; synthesize a multi-modal operational report | The agent must compose calls across a 108-tool MCP library over heterogeneous, multi-temporal geospatial data, tracking intermediate perception/analysis outputs and their correctness as the pipeline progresses through the five task dimensions. | DAG-with-precedence | no | 3,500 tool-call steps total (across 515 tasks, ~6.8/task average)tool-calls | 41.0/mo |
| DuMateBench
DuMateBench: Evaluating Autonomous Agents in Complex Real-World Workflows |
Business & enterprise | 2026 | complete each of 200 real-session-derived tasks spanning 8 broad scenarios and 17 fine-grained capability categories; coordinate multiple capability categories within a single task; preserve and correctly use persistent configurations and workspace state carried over from prior interaction history; maintain performance under injected real-world environmental complexity (Insufficient, Unstable, Noisy conditions) | The agent must track persistent configurations, workspace state, and pre-solution interaction history carried into each task, while coordinating multiple capability categories under injected Insufficient, Unstable, or Noisy environmental perturbations. | set-of-independent | no | none stated | 0 |
| EcoGym
EcoGym: Evaluating LLMs for Long-Horizon Plan-and-Execute in Interactive Economies |
Business & enterprise | 2026 | maintain the Vending sub-environment's profitability (net worth, income) as in Vending-Bench; acquire and retain freelance income under budgeted actions (Freelance sub-environment); sustain operational metrics such as DAU under partial observability (Operation sub-environment); maintain long-term strategic coherence across an effectively unbounded, budgeted-action economic horizon | The agent must track budgeted actions, business-relevant outcome metrics (net worth, income, DAU), and partial observability/stochasticity across an effectively unbounded horizon of 1000+ steps (up to 365 day-loops) per evaluation. | open-ended | no | 1000+ steps over 365 simulated day-loops (evaluation horizon)agent-steps | 10.14/mo |
| EnterpriseArena
Can LLM Agents Be CFOs? Benchmarking Long-Horizon Resource Allocation in an Uncertain Enterprise Environment |
Business & enterprise | 2026 | manage liquidity across the firm's operations; close financial books accurately on a regular cycle; gather costly signals about the macro/industry environment before acting; request equity or debt financing appropriately as conditions change; survive (avoid insolvency/failure) across the full 132-month horizon under shifting macroeconomic regimes | The agent must track liquidity, capital structure (equity/debt), and signals about shifting macroeconomic/industry regimes month by month across the 132-month horizon, since consequences of earlier decisions are delayed and only become apparent later. | DAG-with-precedence | yes | 132-month CFO simulationother:simulated-months | 20.33/mo |
| EnterpriseArena (via the EnterpriseLab platform)
EnterpriseLab: A Full-Stack Platform for developing and deploying agents in Enterprises |
Business & enterprise | 2026 | complete complex enterprise workflows spanning IT, HR, sales, and engineering domains; correctly invoke and sequence tool calls across 140+ tools exposed via Model Context Protocol across 15 applications; match frontier-model performance while running as a smaller (8B) privacy-preserving model; generalize/remain robust across diverse enterprise benchmarks (EnterpriseBench, CRMArena) | The trained agent must track which of many interdependent enterprise tools/applications it has invoked and their resulting state, to complete complex multi-tool enterprise workflows without needing frontier-scale model capacity. | DAG-with-precedence | no | none stated | 10.17/mo |
| EnterpriseClawBench
EnterpriseClawBench: Benchmarking Agents from Real Workplace Sessions |
Business & enterprise | 2026 | read heterogeneous workplace files relevant to a task; invoke the correct tools to accomplish a business objective; deliver a business artifact matching role-specific hard rules and semantic rubrics | The agent must track which files it has read, what tools it has invoked, and whether its eventual delivered artifact satisfies the task's hard rules and semantic rubric, since harness-model combination, artifact delivery, and visual quality are all separately reported. | sequential-chain | yes | none stated | 10.33/mo |
| EnterpriseOps-Gym
EnterpriseOps-Gym: Environments and Evaluations for Stateful Agentic Planning and Tool Use in Enterprise Settings |
Business & enterprise | 2026 | complete each of 1,150 expert-curated enterprise tasks across eight mission-critical verticals (e.g., Customer Service, HR, IT); plan and act correctly amid persistent state changes and strict access-control protocols; correctly refuse infeasible tasks rather than attempting unintended, potentially harmful side effects | Agent must track persistent state changes across 164 database tables, access-control constraints, and whether the current task is feasible at all, to avoid unintended side effects from attempting an infeasible task. | DAG-with-precedence | yes | none stated | 111.83/mo |
| EntWorld
EntWorld: A Holistic Environment and Benchmark for Verifiable Enterprise GUI Agents |
Business & enterprise+ Web & GUI | 2026 | complete each of 1,756 tasks spanning six representative enterprise domains (CRM, ITIL, ERP, etc.); operate under strict business logic constraints and high-density enterprise UIs; maintain precise, state-consistent information retrieval verified via SQL-based state-transition checks; complete synthesized long-horizon workflows reverse-engineered from database schemas | The agent must track precise, state-consistent information across a high-density enterprise UI and the underlying database schema, since success is verified deterministically via SQL-based state-transition checks rather than visual matching. | sequential-chain | no | none stated | 50.62/mo |
| ERP-Bench (via the Anchor task-generation pipeline)
Anchor: Mitigating Artifact Drift in Agent Benchmark Generation |
Business & enterprise | 2026 | satisfy explicit procurement-workflow constraints in a production-grade ERP system; satisfy explicit manufacturing-workflow constraints in a production-grade ERP system; reach the solver-certified fully optimal end-state solution for a long-horizon business task, not merely a feasible one | The agent must track state-based verifier conditions tied to the end-state of a production-grade ERP system across a long-horizon procurement or manufacturing workflow, since reward depends solely on end-state business correctness rather than intermediate steps. | other (singleton) | no | none stated | 10.25/mo |
| ERPBench
ERPBench: Evaluating LLM Agents for Enterprise Decision-Making Across Competitive Market Ecologies |
Business & enterprise | 2026 | make coupled decisions each round across pricing, production, procurement, inventory, and finance; maximize valuation/rank while competing against fixed rule-based opponents (Solo) or other evaluated LLM agents (Arena) in a shared market; sustain a coherent business strategy across all six rounds of the ERP simulation | Agent must track its own accumulated valuation, inventory, and finance position across all six rounds, plus (in Arena mode) the shared market state shaped by five other competing LLM agents, to decide coupled pricing/production/procurement actions each round. | sequential-chain | yes | 6turns | 0 |
| EvoEnv
The Agent's First Day: Benchmarking Learning, Exploration, and Scheduling in the Workplace Scenarios |
Business & enterprise | 2026 | schedule and prioritize streaming tasks with varying priorities under a dynamic workload; actively seek out information to reduce hallucination before acting (prudent information acquisition); distill and reuse generalized strategies learned from earlier dynamically generated tasks in later ones (continuous evolution) | The agent must track incoming streaming tasks and their priorities, its own accumulated exploration/experience for continual learning, and its confidence in acquired information to avoid hallucination, across a continuously evolving workplace scenario. | hierarchical | yes | 50 dynamic scenarios, each with 2-6 task instances; observed usage up to ~90 steps and 232 tool calls for the highest-usage evaluated model (Gemini-3-Flash)actions | 10.12/mo |
| FinEvo-Bench
FinEvo-Bench: A Longitudinal Benchmark for Self-Evolving Agents in Professional Financial Workflows |
Business & enterprise | 2026 | complete each of 120 real-case-grounded financial tasks across 20 business scenes in six financial domains, following institution-provided procedures and constraints; retain and re-apply experience from earlier cases in a scene to later, substantively distinct cases sharing the same professional procedure (self-evolution); maintain financial-compliance quality across each task as scored by a manually reviewed rubric; sustain or improve quality/compliance across an interleaved, shuffled task stream over the longitudinal run | The agent (self-evolving scaffold) must retain experience -- via memory, skill distillation, or both -- from earlier cases in the interleaved task stream and apply it to later, procedurally-related but factually distinct cases, while satisfying institution-provided professional-procedure constraints and financial-compliance requirements on each individual task. | other (singleton) | yes | 120 real-case-grounded tasks per longitudinal stream (20 scenes x 6 cases each); three independently shuffled, globally interleaved task streamsepisodes | 11.0/mo |
| FM-Bench
FM-Bench: A Benchmark for Long-Horizon Management with Competing Agents |
Business & enterprise | 2026 | draft and manage a football squad within a fixed budget shared with rivals; trade players and negotiate contracts across the season; invest in facilities and youth development; set match lineups; maintain board confidence (avoid being fired) while maximizing a final accumulated score over 20 in-game years | The agent must track squad composition, budget/cash flow, ongoing contract-renewal deadlines, facility/youth investments, and the board's confidence in it, across roughly 340-400 decision stops spanning 20 in-game years, since 'the order settles only late in the horizon.' | sequential-chain | yes | 20 in-game years; roughly 340 to 400 decision stops per runother:decision-stops | 0 |
| FORESIGHT-9
FORESIGHT-9: Prospective and Process-Aware Evaluation of Adaptive Trading Agents |
Business & enterprise | 2026 | maintain a coherent, active factor-ensemble/trading strategy across staged macro-financial events without silent internal collapse; respond adaptively to joint multi-asset anchors and disclosed in-world-time observations across a multi-year counterfactual worldline; keep portfolio holdings and decision-records mutually consistent (avoid decision-record vs. execution divergence) | The agent must track its adaptive factor library/ensemble state, executed portfolio holdings, and staged event disclosures across a multi-year, ~2,468-trading-day counterfactual worldline, ensuring internal decision-records remain consistent with actual executed holdings. | sequential-chain | yes | each worldline runs from the 2026-07-15 information boundary through 2035-12-31, approximately 2,468 simulated trading days (~9 years 5 months) per run, across 36 long-horizon runs totalsimulated-days | 0 |
| FrontierFinance
FrontierFinance: A Long-Horizon Computer-Use Benchmark of Real-World Financial Tasks |
Business & enterprise | 2026 | complete each of 25 complex financial modeling tasks across five core finance models; produce client-ready outputs matching industry-standard financial-modeling workflows; satisfy detailed structured evaluation rubrics per task | The agent must track detailed financial-modeling state (assumptions, formulas, model linkages) across a task requiring on average over 18 hours of equivalent skilled human labor, checked against detailed structured rubrics. | sequential-chain | yes | average of over 18 hours of skilled human labor per taskhuman-expert-hours | 20.4/mo |
| GDPevo
GDPevo: Evaluating Agent Self-Evolution on Real Business Tasks |
Business & enterprise | 2026 | update an agent's persistent state from prior training-task experience (self-evolution); apply recombined atomic business rules correctly on held-out test tasks across CRM/ERP/finance/healthcare/legal/data-centric workflows; achieve held-out accuracy gains attributable specifically to training experience rather than data contamination | The evaluation harness must track which atomic business rules were exposed during training versus recombined at test time, per task group, to attribute held-out gains specifically to training experience rather than contamination. | set-of-independent | yes | none stated | 0 |
| H-AdminSim
H-AdminSim: A Multi-Agent Simulator for Realistic Hospital Administrative Workflows with FHIR Integration |
Business & enterprise+ Multi-agent orgs | 2026 | process hospital administrative requests (drawn from a workload of 10,000+ requests/day in large hospitals); coordinate across multiple administrative subtasks rather than handling them in isolation; operate correctly against FHIR-integrated, heterogeneous hospital-setting data | The multi-agent simulation must track hospital administrative request state across heterogeneous hospital settings via a unified FHIR-integrated environment, evaluated against detailed rubrics. | other (singleton) | yes | none stated | 0 |
| HANDBOOK.md
HANDBOOK.md: A Benchmark for Long-Context Agentic Instruction Following |
Business & enterprise | 2026 | locate the specific clauses in a long standing-policy document (20-124 pages) that apply to the current situation; carry out routine professional work strictly governed by that policy document across many tool-mediated actions; avoid letting a plausible but unauthorized in-environment request override the standing policy; satisfy every one of many deterministic, programmatic grading criteria (824 total) simultaneously | The agent must hold the relevant clauses of a long standing-policy document across roughly 17 reasoning steps and 30 tool calls on average, checking each action against required and prohibited criteria rather than losing rule details over the long horizon. | DAG-with-precedence | yes | ~17 reasoning steps and ~30 tool calls on average per tasktool-calls | 21.0/mo |
| HealthAdminBench
HealthAdminBench: Evaluating Computer-Use Agents on Healthcare Administration Tasks |
Business & enterprise+ Healthcare | 2026 | complete a Prior Authorization workflow end-to-end across an EHR and payer portal; complete an Appeals and Denials Management workflow; complete a Durable Medical Equipment (DME) Order Processing workflow; satisfy each of the fine-grained verifiable subtasks a task decomposes into (often 15+ per task) | The agent must track progress across many fine-grained, cross-application subtasks (spanning EHR, payer portals, and fax) per administrative workflow, under a fixed interaction budget, to reach a correct terminal workflow state. | DAG-with-precedence | yes | none statedother:not-stated | 81.6/mo |
| Herculean
Herculean: An Agentic Benchmark for Financial Intelligence |
Business & enterprise | 2026 | complete Trading workflow tasks via a standardized MCP-based skill environment; complete Hedging workflow tasks requiring long-horizon coordination and state consistency; complete Market Insights workflow tasks; complete Auditing workflow tasks requiring structured verification | The agent must track workflow-specific state, tool interactions, and constraints within each MCP-based skill environment, with Hedging and Auditing additionally requiring sustained state consistency and structured verification across long-horizon coordination. | set-of-independent | yes | none stated | 30.75/mo |
| JobBench
JobBench: Aligning Agent Work With Human Will |
Business & enterprise | 2026 | complete 130 agentic tasks spanning 35 occupations that experts identify as high-priority for delegation; reason through cluttered, heterogeneous reference-file information streams within a professional workspace; satisfy a fact-anchored chain of rubrics averaging 35.6 binary criteria per task | Agent must correctly reason through a workspace of heterogeneous reference files and track satisfaction of an average of 35.6 fact-anchored binary rubric criteria per task. | hierarchical | yes | none stated | 41.0/mo |
| LH-Bench
LH-Bench: Skill-Grounded Evaluation of Long-Horizon Agents on Subjective Enterprise Tasks |
Business & enterprise | 2026 | produce subjective, context-dependent enterprise work whose quality depends on organizational goals and user intent, not a single correct answer; produce correct intermediate artifacts across long, multi-tool workflows (e.g., chapter-level content, Figma-to-code conversions); satisfy expert-grounded rubrics scoring subjective work quality; align with pairwise human preference judgments as convergent validation | The agent must track and produce correct intermediate artifacts (e.g., each chapter of a course, or each Figma-to-code conversion) across long, multi-tool workflows, since ground-truth artifacts enable stepwise reward signals at that granularity rather than only a single final judgment. | hierarchical | yes | none stated | 20.33/mo |
| Long-Horizon Multi-Tool Agent (LHMTA) task collection
Post-Training on Office Work Improves Software Engineering: A Behavioral Account of Cross-Domain Transfer |
Business & enterprise | 2026 | select goals appropriately within nested and branching office work; construct task-relevant state as work unfolds; maintain fidelity to higher-level objectives across nested/branching sub-work; verify completion against the environment | Agent must repeatedly select goals, construct task-relevant state, maintain fidelity to higher-level objectives, and verify completion against the environment across nested and branching long-horizon office work. | hierarchical | yes | none stated | 0 |
| LongHorizon-Bench
Four-Axis Decision Alignment for Long-Horizon Enterprise AI Agents |
Business & enterprise | 2026 | make high-stakes regulated decisions (loan qualification, insurance claims adjudication) under lossy memory and multi-step reasoning; maintain factual precision (FRP) about case facts over the long horizon; maintain reasoning coherence (RCS) across the multi-step decision process; reconstruct compliance/regulatory justification (CRR) for the decision; calibrate abstention (CAR) -- know when to decline to decide rather than commit incorrectly | The agent must retain case facts accurately over a long-horizon multi-step review under lossy memory, maintain a coherent reasoning chain, reconstruct which regulatory rules justify the decision, and calibrate when to abstain rather than commit -- all against deterministic ground truth. | set-of-independent | yes | none stated | 10.2/mo |
| LongMedBench
LongMedBench: Benchmarking Medical Agents for Long-Horizon Clinical Decision-Making |
Business & enterprise+ Healthcare | 2026 | aggregate evidence across a patient's repeated visits, tests, and evolving treatments to answer fact-based QA; perform temporal reasoning over the patient's event stream, including implicit (not just explicit-timestamped) time inference; make a long-horizon clinical decision that correctly uses historical patient information accumulated over many visits | Agent must integrate time-series clinical events (admission records, notes) across many visits per patient, correctly distinguish explicit timestamps from cases requiring implicit time inference, and carry this forward into long-horizon decision-making. | sequential-chain | unclear | 19.72 inpatient visits per patient (average); 44.91 medical events per visit (average)sessions | 0 |
| MBABench
MBABench: Evaluating LLM Agents on End-to-End Spreadsheet Tasks in Finance |
Business & enterprise | 2026 | construct an entire financial spreadsheet (e.g., financial model, forecast, scenario analysis) end-to-end from a high-level instruction; jointly satisfy Accuracy, Formula, and Format criteria, each with fine-grained sub-criteria reflecting professional standards | Agent must track intermediate calculation dependencies within the spreadsheet, professional formatting/readability conventions, and formula correctness simultaneously while building the deliverable end-to-end. | set-of-independent | yes | none stated | 20.5/mo |
| MerchantBench
MerchantBench: Benchmarking LLM Agents for Long-Term Coherence in E-Commerce Operations |
Business & enterprise | 2026 | source products and manage upstream supplier events over time; set and adjust listing/pricing to remain competitive and solvent; manage cash-flow across delayed, heterogeneous-latency order outcomes; follow individual order lifecycles end-to-end and revisit earlier decisions as new (delayed) information arrives | The agent must follow individual order lifecycles end-to-end, track upstream supplier events and their promised delayed downstream outcomes, manage cash-flow/net-assets state, and revisit/adapt earlier sourcing and pricing decisions as delayed feedback arrives across a full 365-simulated-day run. | DAG-with-precedence | yes | a 365-day order-level simulation; 48 runs, each spanning 365 simulated dayssimulated-days | 10.5/mo |
| Multi-Horizon Task Environments (MHTE) / CorpGen
CORPGEN: Simulating Corporate Environments with Autonomous Digital Employees in Multi-Horizon Task Environments |
Business & enterprise | 2026 | manage dozens of concurrent, interleaved long-horizon corporate tasks (45+ tasks, 500-1500+ steps each); handle inter-task dependencies expressed as DAGs rather than simple chains; reprioritize among concurrent tasks as load and context change over a persistent execution context spanning hours | The system must maintain hierarchical goal alignment, isolate sub-agent context to prevent cross-task contamination, and manage tiered (working/structured/semantic) memory with adaptive summarization across a persistent execution context spanning hours. | DAG-with-precedence | no | 500-1500+agent-steps | 0 |
| OccuBench
OccuBench: Evaluating AI Agents on Real-World Professional Tasks via Language Environment Simulation |
Business & enterprise | 2026 | complete each of 100 real-world professional task scenarios spanning 65 specialized domains; maintain task completion under controlled fault injection (explicit errors, implicit data degradation, mixed faults) | Agent must track task-completion progress plus signals of environmental robustness (whether tool responses are timing out, truncated, or subtly degraded) within each professional scenario. | set-of-independent | yes | none stated | 40.8/mo |
| OmegaUse-OfficeVal
OmegaUse-OfficeVal: Benchmarking LLM Agents on Long-Horizon Office-Suite Tasks with Economic Grounding |
Business & enterprise | 2026 | complete 100 practitioner-derived office-suite tasks end-to-end; achieve deliverable quality verified via fine-grained rubric-based code verifiers; remain economically competitive relative to human labor time and task price | Agent must track fine-grained rubric criteria and produce a deliverable matching the practitioner's request, verified via code-based verifiers, implicitly benchmarked against the human labor time (avg. 2.32 hours) needed for the same task. | set-of-independent | yes | 2.32 (average)human-expert-hours | 0 |
| OptAgent
OptAgent: an Agentic AI framework for Intelligent Building Operations |
Business & enterprise | 2026 | assess how a system/control upgrade changes energy use; assess how the same upgrade changes operating cost; assess how the same upgrade changes thermal comfort; assess how the same upgrade changes flexibility, via coordinated multi-domain multi-agent analytics | The orchestrator must track which specialist agent/tool has been invoked for which sub-domain (thermal dynamics, HVAC, DER), intermediate physics-informed simulation outputs, and how upgrades propagate across energy use, cost, comfort, and flexibility metrics within one workflow. | DAG-with-precedence | no | none statedother:not-stated | 10.12/mo |
| OR-Space
OR-Space: A Full-Lifecycle Workspace Benchmark for Industrial Optimization Agents |
Business & enterprise | 2026 | construct a solver-ready optimization model from heterogeneous business artifacts (Build); revise an existing model under changing requirements or solver feedback while preserving valid prior logic (Revise); answer grounded questions about solutions, constraints, and business implications using evidence spread across workspace artifacts (Explain) | The agent must track the current state of a persistent multi-artifact workspace (documents, structured data, code, solver outputs) across model construction, revision, and explanation stages, preserving valid prior modeling logic when new requirements or solver feedback arrive. | DAG-with-precedence | unclear | none stated | 20.5/mo |
| OrthoPilot
Evidence-Grounded AI for Musculoskeletal Care |
Business & enterprise+ Healthcare | 2026 | integrate evolving imaging, laboratory, pathology, and order data as it arrives across visits; produce evidence-based decisions at each stage of care, from admission diagnosis through rehabilitation planning; maintain continuous, individualised management across the full musculoskeletal care pathway rather than isolated per-visit decisions | System must continuously retrieve and integrate real-time imaging, laboratory, pathology, and order data across visits/departments/hospital systems, translating evolving patient state into stage-specific functional goals across the whole care pathway. | sequential-chain | yes | months to years (per patient pathway); 1,870 cases / 8,240 inpatients across the studyother:real-world-months-to-years-per-patient-pathway | 0 |
| OSWorkerBench
UI-Mate: Advancing Open-Weight Foundation GUI Agents with In-Context Demonstrations |
Business & enterprise+ Web & GUI | 2026 | complete each of 100 long-horizon office tasks spanning 41 applications; follow demonstrated subtask-level workflows (self-demo or variant-demo) and re-plan from the live interface when needed; make measurable progress toward task completion even when strict end-to-end success is not reached | Agent must track its position in a demonstrated or self-planned subtask sequence, detect when the live interface diverges from the demonstration, and re-plan accordingly across a long-horizon office task. | sequential-chain | yes | none stated | 0 |
| PhysicianBench
PhysicianBench: Evaluating LLM Agents in Real-World EHR Environments |
Business & enterprise+ Healthcare | 2026 | retrieve relevant clinical data across multiple encounters in the EHR; reason over heterogeneous clinical information (labs, notes, orders) to reach a decision; execute consequential clinical actions (e.g. prescribing, ordering) grounded against the environment; produce clinical documentation reflecting the completed workflow | Agent must track data retrieved across multiple encounters, intermediate clinical reasoning state, and which structured checkpoints have been satisfied, using execution-grounded verification against real patient records via standard EHR APIs. | hierarchical | yes | 27 (average)tool-calls | 143.5/mo |
| POLARIS
POLARIS: Typed Planning and Governed Execution for Agentic AI in Back-Office Automation |
Business & enterprise | 2026 | synthesize a type-checked directed acyclic graph (DAG) plan for a back-office document-processing task; select a single compliant plan via rubric-guided reasoning among structurally diverse candidate DAGs; pass validator-gated checks and a bounded repair loop before execution; route or block side effects per compiled policy guardrails, including anomaly routing | System must track the candidate DAG structure, per-node type/validator status, and the full execution trace/audit trail to support bounded repair and decision-grade anomaly routing. | DAG-with-precedence | yes | none stated | 20.25/mo |
| PolyWorkBench
PolyWorkBench: Benchmarking Multilingual Long-Horizon LLM Agents |
Business & enterprise | 2026 | process heterogeneous multilingual inputs correctly within a workflow task; perform iterative reasoning and invoke external tools while maintaining linguistic consistency; produce a structured, correct output for one of five workplace domains (commerce, knowledge work, legal analysis, localization, manufacturing) | Agent must track functional correctness and linguistic consistency simultaneously across a workflow's reasoning and tool-invocation steps, since the two can drift independently as multilingual inputs are processed. | sequential-chain | yes | none stated | 10.33/mo |
| PolyWorkBench
PolyWorkBench: Benchmarking LLM Agents for Cross-Lingual Long-Horizon Workflows |
Business & enterprise | 2026 | integrate heterogeneous multilingual inputs relevant to a workplace task; execute iterative tool-use trajectories within a target domain (commerce, knowledge work, legal analysis, localization, manufacturing); produce structured domain artifacts as verifiable task output | Agent must track and correctly integrate heterogeneous multilingual information while executing an iterative sequence of tool calls, and produce a structured domain artifact that is later verified. | sequential-chain | yes | none stated | 0 |
| PowerAgentBench-SS
PowerAgentBench-SS: A Benchmark for Agentic AI in Power System Steady-State Studies |
Business & enterprise | 2026 | inspect a grid case and select appropriate tools/simulators for the workflow; screen a large space of contingencies within a limited validation budget; propose admissible mitigations for discovered risks and validate their physical validity; produce an auditable evidence trail and submit a bounded, ranked report of top contingencies | The agent must track which contingencies it has already validated (against a fixed budget of 80 out of 1,035), the evidence log supporting each, and whether its running set of top candidates still reflects the best available evidence as it allocates its remaining validation budget. | sequential-chain | yes | validation budget of 80 contingency-case validations per episode (out of 1,035 total N-2 cases), submitting a ranked report of exactly 20 contingencies; hidden dangerous set of 52 casesactions | 41.33/mo |
| PPT-Eval
PPT-Eval: A Benchmark for Computer-Use Agents on PowerPoint Tasks |
Business & enterprise+ Web & GUI | 2026 | complete content-creation and presentation-editing tasks across 120 PowerPoint tasks in 12 files; satisfy task-specific rubric criteria that award partial credit for intermediate steps; avoid unnecessary changes and poor aesthetics while making required edits | Agent must track which task-specific rubric criteria (intermediate steps, aesthetics, unnecessary changes) have been satisfied across a multimodal editing session, since rubrics award partial credit rather than only final binary success. | hierarchical | yes | none stated | 51.67/mo |
| RetailBench
RetailBench: Benchmarking long horizon reasoning and coherent decision making of LLM agents in realistic retail environments |
Business & enterprise | 2026 | manage pricing across the store's product assortment; manage replenishment and supplier selection to keep inventory stocked; manage shelf assortment and inventory aging; respond appropriately to customer feedback and external events; maintain solvency (cash-flow constraints) while maximizing net worth/sales over a long simulated horizon | The agent must track pricing, inventory levels/aging, supplier relationships, customer feedback, external events, and its own cash-flow position day by day across the 180-day (or longer) simulated run, since only a small subset of evaluated agents survive the full evaluation horizon. | sequential-chain | yes | 180-day evaluation horizon; the simulator supports thousand-day-scale simulationssimulated-days | 61.0/mo |
| RiskWebWorld
RiskWebWorld: A Realistic Interactive Benchmark for GUI Agents in E-commerce Risk Management |
Business & enterprise+ Web & GUI | 2026 | investigate flagged e-commerce risk cases across multiple verification sub-steps on production risk-control pipelines; operate GUI actions on uncooperative websites subject to partial environmental hijacking; complete each of 1,513 tasks spanning 8 core risk-control domains; succeed at long-horizon professional risk-investigation workflows | The agent must track evidence and verification state gathered across multiple sub-steps of a risk investigation on an uncooperative website, while detecting and coping with partial environmental hijacking attempts, across long-horizon professional tasks. | sequential-chain | no | none stated | 0 |
| SaaS-Bench
SaaS-Bench: Can Computer-Use Agents Leverage Real-World SaaS to Solve Professional Workflows? |
Business & enterprise | 2026 | navigate and operate real, deployed SaaS systems to complete a professional workflow; coordinate state and context across multiple applications within the same workflow; apply domain-specific knowledge correctly within the SaaS system; recover from errors and maintain progress over a long-horizon task | The agent must maintain state and context across multiple SaaS applications over long-horizon execution, tracking partial progress against weighted verification checkpoints, since 'agents become stuck acquiring information or manipulating interfaces, confuse source and output devices, or terminate before all conditions are jointly satisfied.' | DAG-with-precedence | yes | average of over 100 interaction steps per taskactions | 30.75/mo |
| SpreadsheetBench 2
SpreadsheetBench 2: Evaluating Agents on End-to-End Business Spreadsheet Workflows |
Business & enterprise | 2026 | generate new spreadsheet content/formulas correctly across a large multi-sheet workbook; debug existing incorrect formulas/content within the workbook; produce correct visualizations from the workbook's data | The agent must track and correctly propagate changes across an average of 11.8 interdependent worksheets requiring 593.5 cell modifications per task, correctly identifying target cells under a unified multi-turn agent scaffold. | DAG-with-precedence | yes | each task instance averages 11.8 worksheets and requires 593.5 cell modificationsother:cell-modifications | 41.33/mo |
| SupChain-Bench
SupChain-Bench: Benchmarking Large Language Models for Real-World Supply Chain Management |
Business & enterprise | 2026 | orchestrate long-horizon, multi-step supply-chain tool use grounded in standard operating procedures (SOPs); correctly apply supply-chain domain knowledge across a sequence of dependent tool calls | Agent must track SOP-grounded procedural state across a long-horizon sequence of dependent tool calls in order to correctly complete supply-chain management workflows. | sequential-chain | no | none stated | 0 |
| Synthetic Computers at Scale
Synthetic Computers at Scale for Long-Horizon Productivity Simulation |
Business & enterprise | 2026 | complete each of multiple professional deliverables comprising a computer-specific productivity objective; navigate the synthetic computer's filesystem to ground actions in the user's actual context; coordinate with simulated collaborators as needed to complete the objective; sustain progress toward an objective requiring about a month of simulated human work | The acting agent must track the evolving state of a realistic folder hierarchy and content-rich artifacts (documents, spreadsheets, presentations) across a run spanning over 2,000 turns, coordinating with simulated collaborators toward multiple professional deliverables. | open-ended | no | >2,000 turns on average (each run also requiring over 8 hours of agent runtime)turns | 10.2/mo |
| temporal enterprise scenario replay system
What Could the Agent See at 19:05? Generating Temporal Enterprise Scenarios from Real Research and Replaying Them to Evaluate Agents |
Business & enterprise | 2026 | answer questions correctly relative to what data existed and who could see it at a specific queried moment; reason about a persona-driven, temporally-evolving enterprise world spanning many apps; avoid leaking future/hidden record state when reasoning about an earlier moment | The agent being evaluated must reason correctly about what data existed and who could see it at a specific queried moment, without conflating it with earlier or later states of the same records across multiple apps. | DAG-with-precedence | yes | none stated | 0 |
| Thinkingbox-bench
One Success Isn't Reliability: Thinkingbox, a Sandbox and Benchmark for Agents in Stateful Business Workflows |
Business & enterprise+ Information seeking | 2026 | gather missing information over multiple turns before acting; follow domain-specific policies for each of 507 policy-conditioned workflows; coordinate dependent tools correctly; realize exactly the correct persistent backend state transition without collateral effects | Agent must gather missing information across multiple turns, track applicable domain policies, coordinate dependent tool calls, and verify the resulting persistent backend state matches exactly the required final state with no collateral effects. | DAG-with-precedence | no | none stated | 0 |
| Underwrite
Benchmarking Agents in Insurance Underwriting Environments |
Business & enterprise+ Information seeking | 2026 | gather information carefully from noisy tool interfaces and imperfect simulated users during an underwriting conversation; apply proprietary business/domain knowledge correctly rather than hallucinating it; reach a final underwriting decision consistent with the accumulated evidence gathered across the conversation | The agent must track what information it has already gathered (and from where), the reliability of that information given noisy tool interfaces, and how it should update its evolving underwriting assessment as more evidence accumulates across the conversation. | sequential-chain | yes | average of 3-7 steps of required reasoning and tool use, with a total of 10-20 conversational turnsturns | 0 |
| Workflow-GYM
Workflow-GYM: Towards Long-Horizon Evaluation of Computer-use Agentic tasks in Real-World Professional Fields |
Business & enterprise | 2026 | autonomously operate domain-specific professional software GUIs to accomplish economically valuable work; complete long-horizon, multi-stage professional workflows end-to-end; maintain workflow consistency across stages without omission, error propagation, or objective drift | The agent must track its position and completed stages within a long-horizon, multi-stage professional GUI workflow, avoiding stage omission, propagating errors, or drifting from the original objective across the workflow's duration. | sequential-chain | no | none stated | 20.67/mo |
| Workspace-Bench
Workspace-Bench 1.0: Benchmarking AI Agents on Workspace Tasks with Large-Scale File Dependencies |
Business & enterprise | 2026 | identify, reason over, exploit, and update explicit and implicit dependencies among heterogeneous files in a worker's workspace; satisfy each of a task's own file-dependency-graph-derived rubrics via cross-file retrieval, contextual reasoning, and adaptive decision-making | Agent must track which of up to 20,476 files across 74 file types are relevant, their dependency relationships, and adaptively retrieve/reason across files to satisfy each task's rubrics. | DAG-with-precedence | yes | none stated | 102.5/mo |
| World of Workflows (WoW) / WoW-bench
World of Workflows: a Benchmark for Bringing World Models to Enterprise Systems |
Business & enterprise | 2026 | complete constrained agentic tasks within a ServiceNow environment governed by 4,000+ business rules and 55 active hidden workflows; predict cascading side effects of actions across interconnected databases; mentally simulate hidden state transitions to avoid silent constraint violations under limited observability | Agent must mentally simulate hidden state transitions and predict cascading side effects across interconnected databases to bridge the observability gap, since high-fidelity feedback is often unavailable. | DAG-with-precedence | no | none stated | 70.88/mo |
| YC-Bench
YC-Bench: Benchmarking AI Agents for Long-Term Planning and Consistent Execution |
Business & enterprise | 2026 | manage employees within a simulated startup; select task contracts to pursue under uncertainty; maintain profitability against adversarial clients and growing payroll; detect and avoid bankruptcy-inducing failure modes (e.g. adversarial-client mismanagement, over-parallelization) over a one-year run | The agent must persist information across context truncation via a scratchpad (the strongest predictor of success), tracking employees, contracts, cash reserves, and adversarial-client risk across hundreds of turns spanning a simulated year. | open-ended | no | hundreds of turns (over a simulated one-year horizon)turns | 81.6/mo |
| year-long LLM stock-market trading-style simulation testbed
Behavioral Consistency Validation for LLM Agents: An Analysis of Trading-Style Switching through Stock-Market Simulation |
Business & enterprise | 2026 | trade under a currently assigned style (fundamental or technical) each trading day; periodically (every 10 trading days) reassess and potentially switch trading style based on four behavioral-finance drivers: loss aversion, herding, wealth differentiation, price misalignment; keep style-switching behavior consistent with real-world behavioral-finance theory over the whole simulated year | Agent must process daily price-volume data, retain long-term personality traits (the four behavioral-finance drivers) set at initialization, and track its own accumulated wealth/strategy history to decide whether to switch trading style every 10 days. | sequential-chain | unclear | year-long (reassessed every 10 trading days)simulated-days | 40.57/mo |
| Agent Trading Arena
Agent Trading Arena: A Study on Numerical Understanding in LLM-Based Agents |
Business & enterprise | 2025 | make sequential buy/sell trading decisions that directly affect and are affected by shared market prices; compete against other LLM-based agents in a zero-sum stock market; maximize trading performance/returns, especially under high volatility; correctly perform numerical reasoning over price data or chart-based visualizations | Each agent must track its own portfolio/capital, the current market price state shaped by all agents' recent trades, and historical price patterns, updating its numerical reasoning as the shared market evolves turn to turn. | open-ended | no | none stated | 120.63/mo |
| AI-Trader
AI-Trader: Benchmarking Autonomous Agents in Real-Time Financial Markets |
Business & enterprise | 2025 | independently search, verify, and synthesize live market information given only minimal initial context; make live trading decisions (buy/sell/hold) across U.S. stocks, A-shares, and cryptocurrencies at multiple trading granularities; manage risk and sustain positive returns over a continuous live-trading period | Agent must track evolving live market data it has independently searched/verified, current positions/risk exposure across three markets and multiple trading frequencies, and running returns over the continuous live-trading period. | sequential-chain | yes | none stated | 192.11/mo |
| AssetOpsBench
AssetOpsBench: A Real-World Evaluation Benchmark for AI-Driven Task Automation in Industrial Asset Management |
Business & enterprise | 2025 | orchestrate the correct domain-specific agent(s) (from a catalog of four) to answer a natural-language industrial-operations query; correctly complete condition-monitoring and maintenance-scheduling workflow steps grounded in a simulated IoT environment | The agent must track intermediate tool/agent outputs across a multi-step think-act-observe loop, orchestrate the four domain-specific agents appropriately, and maintain consistency with the simulated CouchDB-backed IoT environment's evolving state. | DAG-with-precedence | yes | Plan-Execute agents complete most tasks in approximately 2.6-4.4 steps; Agent-As-Tool agents typically require approximately 4-6+ steps due to its iterative think-act-observe loopagent-steps | 171.13/mo |
| AutoDW
Automating Complex Document Workflows via Stepwise and Rollback-Enabled Operation Orchestration |
Business & enterprise | 2025 | execute a sequence of interdependent, user-specified document-editing instructions within a session; keep the execution trajectory aligned with evolving document state and user intent across the whole session; recover from failed API calls/arguments via rollback at both the argument and API level | System must track the evolving document state after each API action, the remaining interdependent instructions in the session, and whether any prior action needs to be rolled back (at argument or API level) to stay aligned with user intent. | DAG-with-precedence | yes | 1,708 instructions across 250 sessions (~6.8 instructions/session on average)actions | 0 |
| capitalization tie-out benchmark (legal AI)
Does It Tie Out? Towards Autonomous Legal Agents in Venture Capital |
Business & enterprise | 2025 | verify that every security (shares, options, warrants) is supported by underlying legal documentation; verify that every issuance term (vesting schedules, acceleration triggers, transfer restrictions) is consistent across the dataroom; maintain strict evidence traceability while reconciling thousands of pages of legal documents; produce a deterministic, fully-reconciled capitalization table | The agent must track which securities and issuance terms have been verified against which supporting documents, maintaining strict evidence traceability across a growing dataroom (up to tens of thousands of pages) as it reconciles a scaling number of individual securities. | set-of-independent | yes | workload grows from approximately 2,700 atomic verification steps at Seed stage to nearly 8,000 steps at Series B stageagent-steps | 0 |
| CP-Env
CP-Env: Evaluating Large Language Models on Clinical Pathways in a Controllable Hospital Environment |
Business & enterprise+ Healthcare | 2025 | triage a patient correctly to determine care pathway; consult the appropriate specialist given patient information; order and interpret diagnostic tests as needed; participate appropriately in multidisciplinary team meetings; complete a patient's full journey across branching, long-horizon clinical-pathway stages | The agent must track patient information, diagnostic findings, and evolving care-pathway state across branching stages from triage through specialist consultation, diagnostic testing, and multidisciplinary team meetings. | DAG-with-precedence | yes | none stated | 60.67/mo |
| CRMArena-Pro
CRMArena-Pro: Holistic Assessment of LLM Agents Across Diverse Business Scenarios and Interactions |
Business & enterprise | 2025 | complete each of nineteen expert-validated CRM tasks across sales, service, and configure-price-quote (CPQ) processes; sustain correct behavior across multi-turn interactions guided by diverse personas; maintain confidentiality awareness throughout the interaction, for both B2B and B2C scenarios | The agent must track persona-specific context, confidentiality constraints, and workflow state across multi-turn interactions, since single-turn success (58%) drops substantially to about 35% once genuine multi-turn tracking is required. | sequential-chain | yes | none stated | 422.62/mo |
| DevNous
DevNous: An LLM-Based Multi-Agent System for Grounding IT Project Management in Unstructured Conversation |
Business & enterprise+ Multi-agent orgs | 2025 | identify actionable intents from informal, unstructured team-chat dialogue; manage stateful, multi-turn administrative workflows (task formalization, progress-summary synthesis) grounded in that dialogue | The agent must identify actionable intents across informal chat and maintain stateful workflows (task formalization, progress tracking) across a benchmark of 160 conversational turns, correctly matching a multi-label ground truth per turn. | sequential-chain | yes | a new benchmark of 160 realistic, interactive conversational turnsturns | 0 |
| DRBench
DRBench: A Realistic Benchmark for Enterprise Deep Research |
Business & enterprise | 2025 | identify supporting facts for a multi-step research query from both the public web and a private company knowledge base; synthesize facts drawn from heterogeneous enterprise sources (productivity software, cloud file systems, emails, chat, web) into one answer; produce a coherent, well-structured final report grounded in the retrieved facts | The agent must track which facts it has found in which source (public web vs. private enterprise knowledge base), maintain factual accuracy across sources, and assemble a coherent report structure from these accumulated facts. | hierarchical | yes | none stated | 151.25/mo |
| EcomBench
EcomBench: Towards Holistic Evaluation of Foundation Agents in E-commerce |
Business & enterprise | 2025 | retrieve deep, possibly multi-hop information relevant to a real e-commerce user demand; perform multi-step reasoning across e-commerce-domain data; integrate knowledge from multiple sources to resolve a task; correctly complete tasks across three graded difficulty levels | The agent must track partial evidence gathered from deep retrieval and multiple knowledge sources across many action steps before it can complete a higher-difficulty task. | hierarchical | yes | none stated | 70.78/mo |
| Finch (FinWorkBench)
Finch: Benchmarking Finance & Accounting across Spreadsheet-Centric Enterprise Workflows |
Business & enterprise | 2025 | interleave data entry, structuring, formatting, web search, cross-file retrieval, calculation, modeling, validation, translation, visualization, and reporting subtasks into one finance/accounting workflow; correctly complete each of the linked tasks that compose one of 172 composite workflows | Agent must track state across many interlinked spreadsheets, PDFs, and artifacts (27 million spreadsheet cells total) while interleaving retrieval, calculation, modeling, validation, and reporting sub-steps of one composite workflow. | DAG-with-precedence | yes | 16.8 (GPT-5.1 Pro average)wall-clock-minutes | 91.0/mo |
| GraphicBench
GraphicBench: A Planning Benchmark for Graphic Design with Language Agents |
Business & enterprise+ Web & GUI | 2025 | produce a workflow plan satisfying explicit design constraints stated in a user query; also satisfy implicit commonsense design constraints not stated by the user; select the correct action from 46 available tools at each workflow step; coordinate outputs across three design experts without violating global dependencies | The agent must track explicit and implicit design constraints, the evolving multi-step workflow plan across three design experts, and which of 46 actions remains valid/appropriate at each step. | DAG-with-precedence | yes | none stated | 20.12/mo |
| HiMA-Ecom
HiMA-Ecom: Enabling Joint Training of Hierarchical Multi-Agent E-commerce Assistants |
Business & enterprise+ Multi-agent orgs | 2025 | master agent coordinates multiple specialized sub-agents in an e-commerce workflow; each sub-agent pursues a functionally distinct role-specific objective (e.g. recall domain knowledge, execute a service function); jointly optimize system-level behavior via multi-agent reinforcement learning | The master agent must track which specialized sub-agent is responsible for which functional sub-task and how sub-agent outputs compose into overall system behavior, using a collaboratively updated memory across training. | hierarchical | no | none statedother:not-stated | 80.53/mo |
| MEMTRACK
MEMTRACK: Evaluating Long-Term Memory and State Tracking in Multi-Platform Dynamic Agent Environments |
Business & enterprise+ Assistants & memory | 2025 | acquire relevant facts scattered across asynchronous cross-platform events (Slack/Linear/Git); select the correct, currently-valid fact when multiple conflicting/noisy candidates exist; resolve conflicts between contradictory or stale cross-referring information over time | The agent must acquire, select, and reconcile facts from a chronologically platform-interleaved timeline spanning Slack, Linear, and Git, handling noisy, conflicting, and cross-referring information plus codebase/file-system exploration, across long horizons. | DAG-with-precedence | yes | none stated | 121.09/mo |
| Mini Amusement Parks (MAPs)
Mini Amusement Parks (MAPs): A Testbed for Modelling Business Decisions |
Business & enterprise | 2025 | maximize the amusement park's overall value over the given time horizon; make daily operational decisions (building rides/shops, hiring staff, setting a research agenda) that keep the business viable under sparse, stochastic feedback; reason over the park's spatial layout while planning these decisions | The agent must track the park's spatial layout, built rides/shops, staffing, accumulated (sparse) experience about environment dynamics, and overall park value across a 50-100 day episode, planning under uncertainty at each of many daily decision points. | sequential-chain | yes | episodes span a 50-day horizon on easy mode, extended to 100 days on medium modesimulated-days | 20.2/mo |
| OdysseyBench
OdysseyBench: Evaluating LLM Agents on Long-Horizon Complex Office Application Workflows |
Business & enterprise | 2025 | identify essential information buried in long-horizon interaction histories; perform multi-step reasoning/actions across Word, Excel, PDF, Email, and Calendar applications; complete real-world-derived tasks (OdysseyBench+) or newly synthesized complex tasks (OdysseyBench-Neo) | Agent must retain and retrieve essential facts from long-horizon interaction histories spanning multiple applications in order to correctly execute later multi-step, cross-application actions. | sequential-chain | unclear | none stated | 513.92/mo |
| PPTArena
PPTArena: A Benchmark for PowerPoint Editing |
Business & enterprise | 2025 | apply each of a deck's human-curated edits correctly from natural-language instructions; maintain layout-sensitive and cross-slide consistency across the whole deck while editing; plan and verify edit sequences via an iterative plan-edit-check loop (for the PPTPilot agent) | Agent must track the deck's current structural/visual state after each edit, verify each edit against a ground-truth rubric, and maintain deck-wide consistency (styles, cross-slide references) across the full sequence of edits. | sequential-chain | yes | ~13 edits per deck (1,300+ edits across 100 decks)actions | 70.78/mo |
| ProSoftArena
ProSoftArena: Benchmarking Hierarchical Capabilities of Multimodal Agents in Professional Software Environments |
Business & enterprise | 2025 | operate professional software tools according to a hierarchical capability level (L1-L3); complete realistic work/research tasks spanning 6 disciplines and 13 core professional applications; coordinate across multiple professional software applications for L3 multi-software workflows | The agent must track its progress within and across professional-software applications as required by the task's capability level, since L3 tasks require coordinating state across multiple applications rather than operating one application in isolation. | hierarchical | yes | human execution steps grow from an average of 5.1 steps (14.8 seconds) at L1 to 86.9 steps (506.8 seconds) at L3actions | 50.56/mo |
| REALM-Bench
REALM-Bench: A Benchmark for Evaluating Multi-Agent Systems on Real-world, Dynamic Planning and Scheduling Tasks |
Business & enterprise+ Multi-agent orgs | 2025 | solve each of 14 real-world planning/scheduling problems with multiple parallel planning threads; maintain feasibility across inter-agent dependencies as the problem scales in complexity; adapt schedules in real time when unexpected disruptions arrive | Agents must track the state of parallel planning threads, inter-dependencies between agents/threads, and disruption events in order to replan in real time. | DAG-with-precedence | yes | none stated | 120.63/mo |
| Remote Labor Index (RLI)
Remote Labor Index: Measuring AI Automation of Remote Work |
Business & enterprise | 2025 | complete a whole real, economically valuable freelance/remote-work project end to end; achieve a level of automation comparable to what a human freelancer would deliver for that project | Agent must track progress toward completing an entire multi-part real-world project (not a single isolated action) to be credited with automating that unit of remote labor. | hierarchical | unclear | none stated | 211.91/mo |
| SCUBA
SCUBA: Salesforce Computer Use Benchmark |
Business & enterprise+ Information seeking | 2025 | navigate a specific enterprise software UI (Salesforce) to accomplish a CRM task; manipulate data records and automate workflows within the Salesforce platform; retrieve information and troubleshoot issues as part of a realistic CRM task; generalize across three personas (platform administrators, sales representatives, service agents) each with distinct task profiles | The agent must track its progress toward fine-grained milestones within each Salesforce sandbox task (UI navigation state, data changes made, workflow steps completed) to receive interpretable milestone-progress credit. | sequential-chain | yes | none stated | 70.58/mo |
| SOP-Bench
SOP-Bench: Complex Industrial SOPs for Evaluating LLM Agents |
Business & enterprise | 2025 | correctly execute each step of a complex, multi-step Standard Operating Procedure; orchestrate the correct tools/APIs at each SOP step; produce ground-truth outputs matching the SOP's authored specification across 12 business domains | The agent must track its position within a multi-step SOP, select the correct tool from a registry that may contain many irrelevant tools, and maintain consistency with the SOP's ground-truth interface and outputs across the whole procedure. | sequential-chain | no | none statedother:not-stated | 201.33/mo |
| StaffPro
StaffPro: an LLM Agent for Joint Staffing and Profiling |
Business & enterprise | 2025 | assign and schedule tasks to workers (staffing), forming teams as needed; continuously estimate workers' latent skills, preferences, and other attributes from unstructured feedback (profiling); optimize staffing performance over time as profiling estimates improve via an ongoing human-agent feedback loop | StaffPro must track each worker's evolving latent-attribute profile (estimated from ongoing human feedback) and the current staffing/schedule state, updating both jointly over the 'life-long' profiling horizon. | DAG-with-precedence | yes | none stated | 20.14/mo |
| STOCKBENCH
StockBench: Can LLM Agents Trade Stocks Profitably In Real-world Markets? |
Business & enterprise | 2025 | make a sequential daily buy/sell/hold decision based on incoming market signals (prices, fundamentals, news); maximize cumulative return across the whole multi-month trading period; manage risk, minimizing maximum drawdown and maintaining a strong Sortino ratio | Agent must track its current portfolio position, cumulative return, and risk exposure (drawdown) as it processes a new daily market signal (prices, fundamentals, news) each step of a multi-month simulated trading run. | sequential-chain | yes | multi-month (exact day/month count not given)simulated-days | 333.0/mo |
| UpBench
UpBench: A Dynamically Evolving Real-World Labor-Market Agentic Benchmark Framework Built for Human-Centric AI |
Business & enterprise | 2025 | complete a real, verified client transaction job sourced from the Upwork labor marketplace; satisfy each of a job's detailed, expert-decomposed, verifiable acceptance criteria; follow instructions faithfully enough to earn positive fine-grained, per-criterion human expert feedback | Agent must track which of a job's detailed acceptance criteria it has satisfied, informed by expert freelancer decomposition, to produce a submission gradeable on fine-grained, per-criterion feedback rather than a binary pass/fail. | set-of-independent | yes | none stated | 40.4/mo |
| Vending-Bench
Vending-Bench: A Benchmark for Long-Term Coherence of Autonomous Agents |
Business & enterprise | 2025 | balance inventory levels of vending-machine stock; place restocking orders from suppliers; set item prices; pay recurring daily fees; sustain a profitable long-running vending-machine business without derailing | The agent must continuously track inventory counts, cash/profit, outstanding orders and their delivery schedules, and daily fees across a run spanning over 20M tokens, since forgetting an order or misreading a schedule causes derailment. | open-ended | no | >20Mother:tokens-per-run | 673.53/mo |
| VitaBench
VitaBench: Benchmarking LLM Agents with Versatile Interactive Tasks in Real-world Applications |
Business & enterprise | 2025 | satisfy each of multiple real user requests combined into one cross-scenario task (food delivery, in-store consumption, online travel); reason across temporal and spatial dimensions while using a large tool set (66 tools); proactively clarify ambiguous instructions and track shifting user intent across a multi-turn conversation | Agent must track shifting user intent across a multi-turn conversation, temporal and spatial constraints, and the state of a large (66-tool) toolset spanning multiple life-serving domains simultaneously. | set-of-independent | unclear | none stated | 352.92/mo |
| Wealth-Management benchmark (extension of TheAgentCompany)
Benchmarking LLM Agents for Wealth-Management Workflows |
Business & enterprise | 2025 | complete each of 12 wealth-management task-pairs spanning retrieval, analysis, and synthesis/communication; satisfy explicit acceptance criteria under deterministic graders; operate correctly under both high- and low-autonomy task variants | Agent must track retrieved data, intermediate analysis results, and explicit acceptance criteria across the retrieval-analysis-synthesis pipeline, plus which autonomy variant (high vs. low) governs how independently it may act. | sequential-chain | no | none stated | 0 |
| Service & dialogue23 artifacts | ||||||||
| LLMs Get Lost in Evolving User Intent
LLMs Get Lost in Evolving User Intent |
Service & dialogue | 2026 | track a single user goal that is incrementally revealed across conversation turns; detect and adopt goal revisions as the user changes their mind mid-conversation; detect and follow mid-conversation redirections of the original goal; re-run an existing single-turn benchmark's task under this evolving-intent protocol without new annotation | The agent must maintain a running model of a single, evolving user intent across turns, discarding or revising earlier partial specifications as more of the goal is revealed, revised, or redirected. | sequential-chain | no | up to 7 (including the initial turn)turns | 0 |
| Research on a Multi-SOP Interruption–Resumption Agentic Algorithm for Complex Business Workflows
Research on a Multi-SOP Interruption–Resumption Agentic Algorithm for Complex Business Workflows |
Service & dialogue+ Business & enterprise | 2026 | complete a customer-requested Standard Operating Procedure (SOP) workflow; switch to and complete a different SOP mid-interaction when the customer changes topic; resume a previously interrupted SOP without restarting or re-asking for already-given information; detect and recover from stale/rolled-back state via diff reasoning against frozen snapshots | The system must track which SOP is currently active, a shared six-tuple Global State Container of cross-SOP business entities, and frozen-snapshot vs. current-state diffs to support interruption/resumption without re-asking the user for information already given. | DAG-with-precedence | no | none stated | 0 |
| AcCoRD
AcCoRD: Evaluating User-Agent Collaboration Under Realistic User Preference Dynamics |
Service & dialogue+ Multi-agent orgs | 2026 | resolve underspecified user preferences in online shopping or travel planning; detect and satisfy preferences that emerge mid-interaction (not stated upfront); adapt to preferences the user later adjusts or relaxes during the interaction | Agent must maintain and continuously update a model of the user's preferences as they are formed, revealed, adjusted, or relaxed turn-by-turn, and recognize when uncertainty about a preference needs to be resolved. | sequential-chain | unclear | none stated | 0 |
| AgentWorld
AgentWorld: Personality-Aware Reliability Evaluation for Agentic Information Retrieval |
Service & dialogue | 2026 | maintain consistent, reliable tool-use behavior across interactions with users of varying personality (Big Five/OCEAN) profiles; handle dual-control handoffs correctly between agent and (simulated) user/other controller; resist or survive adversarial perturbations to a required-intermediate-state 'spine' of the task without brittle failure; achieve consistent pass rates (pass^k) across repeated attempts rather than succeeding only once by chance | The agent must track its stateful tool-use progress consistently across repeated attempts and varying user personas, while the Risk Analyser separately tracks the required-intermediate-state spine of a task to evaluate how perturbations at each state affect eventual outcomes. | DAG-with-precedence | yes | each persona ran 3 multi-turn conversation exchanges against the analytics agent, producing 60 messages total (30 from the agent, 30 from personas), across 10 personasturns | 0 |
| CallBench
CallBench: A Benchmark for Dual-Goal Coordination in Phone Call Assistants |
Service & dialogue+ Assistants & memory | 2026 | satisfy the device owner's explicit preset goal for the call; correctly infer and respond to the caller's implicit and dynamic goal; make turn-level decisions that correctly reconcile these two goals under alignment, complementarity, irrelevance, or conflict relations; adhere to preset instructions while maintaining dialogue quality, safety, and rhythm | The assistant must track the owner's preset goal, the evolving implicit goal of the caller, and the current relation between them (alignment/complementarity/irrelevance/conflict) turn by turn across the dialogue, to make reliable turn-level decisions between the two goals. | DAG-with-precedence | yes | average of 5.322 turns per dialogue across all 50,000 dialogues (scenario averages ranging from 4.437 to 7.241 turns)turns | 10.33/mo |
| clinical AI red-teaming framework (unnamed)
Assessing Risks of Large Language Models in Mental Health Support: A Framework for Automated Clinical AI Red Teaming |
Service & dialogue+ Healthcare | 2026 | conduct a full therapy session with a simulated patient agent having a dynamic cognitive-affective model; maintain quality of care across the session per a comprehensive risk ontology; de-escalate suicide risk appropriately when it arises during the session; avoid validating patient delusions ('AI Psychosis') | The AI psychotherapist agent must track the patient's evolving cognitive-affective state, emerging risk signals (e.g., suicide risk, delusion reinforcement), and quality-of-care obligations across the whole therapy session. | sequential-chain | no | none stated | 30.43/mo |
| CRAB-Bench
CRAB-Bench: Evaluating LLM Agents under Complex Task Dependencies and Human-aligned User Simulation |
Service & dialogue | 2026 | reason over a constraint graph spanning multiple interdependent entities to find a solution among thousands of misleading candidate distractors; engage in multi-turn dialogue with a realistic (non-cooperative, persona-driven) simulated user rather than a cooperative template-like one; accommodate multiple valid solutions rather than a single fixed correct answer | Agent must track which constraints (graph edges/entities) have been satisfied so far, which candidate solutions remain viable given thousands of distractors, and information gathered/disclosed across a multi-turn dialogue with a realistic user persona. | DAG-with-precedence | yes | none stated | 0 |
| FraudBench
FraudBench: Stress-Testing Policy-Grounded Banking Agents Against Adaptive Fraud |
Service & dialogue | 2026 | safely act on caller requests during a banking conversation while checking authorization, identity, and policy compliance at every step; retrieve applicable rules from a 698-document internal policy corpus before permitting a sensitive action; detect and refuse chained/adaptive fraud attempts where an earlier probe or admission makes a later, superficially valid request unsafe | The agent must track the caller's identity/authorization claims, any tool access it has already granted, and prior probes/admissions across the conversation, checking each new request against the accumulated history and a 698-document policy corpus. | DAG-with-precedence | yes | none stated | 0 |
| IHBench
IHBench: Evaluating Post-Interruption Recovery in Voice Agents with Structured Workflows |
Service & dialogue | 2026 | resume a state-machine-driven workflow at the correct step after a user interruption; address the content of the user's interjection; avoid re-delivering content the user already heard; achieve task fulfillment across 10 enterprise domains despite one of 6 injected interruption types | The agent must track its position in a state-machine-driven workflow, which content has already been delivered to the user, and the content of the user's interjection, in order to both recover to the correct step and adequately address the interruption. | sequential-chain | yes | none stated | 41.33/mo |
| JourneyBench
Beyond IVR: Benchmarking Customer Support LLM Agents for Business-Adherence |
Service & dialogue+ Business & enterprise | 2026 | adhere to multi-step business policies/SOPs throughout a support conversation; navigate task dependencies correctly as the conversation unfolds; remain robust to unpredictable user/environment behavior while covering the required user-journey graph paths | The agent must track which policy-graph nodes/steps it has satisfied so far, adhere to multi-step business rules, and adapt to unpredictable user/environment behavior across the conversation. | DAG-with-precedence | yes | none stated | 111.38/mo |
| SAGE (Service Agent Graph-guided Evaluation) / SAGE-Bench
SAGE: A Service Agent Graph-guided Evaluation Benchmark |
Service & dialogue+ Web & GUI | 2026 | follow every step of an unstructured Standard Operating Procedure formalized as a Dynamic Dialogue Graph; correctly classify diverse/adversarial user intents and derive the correct subsequent action for each | The agent must track which SOP-graph node/path it is currently on, the user's evolving (possibly adversarial) intent, and maintain logical compliance across the full dialogue, evaluated at increasing dialogue depths (turns 1, 5, 10, 15, and final turn). | DAG-with-precedence | yes | dialogue depths evaluated at turns 1, 5, 10, 15, and the final turn, to measure stability over extended interactionsturns | 30.6/mo |
| SEATauBench
SEATauBench: Adapting Tool-Agent-User Evaluation Into Low-Resource Southeast Asian Languages |
Service & dialogue | 2026 | carry out tau2-Bench-style tool-agent-user tasks correctly when the conversation language changes; carry out the same tasks correctly as tool specifications are localized into the target SEA language; carry out the same tasks correctly as the task domain itself is localized | Agent must correctly interpret and act on user requests, tool specifications, and domain content in the target language while adhering to the policy constraints originally defined in tau2-Bench. | set-of-independent | unclear | none stated | 10.33/mo |
| SpeechGym
SpeechGym: An Audio-Native Gym for Training Voice Agents via Reinforcement Learning |
Service & dialogue | 2026 | correctly hear and call tools/fill argument slots based on values perceived from native audio (no ASR/TTS); hold multi-turn dialogue entirely through speech to complete the same tasks as an established text agentic benchmark; avoid unauthorized write actions under an insistent caller's social pressure | Agent must track dialogue state and correctly perceived argument values purely from native audio across a multi-turn tool-use session, since a single misheard value causes cascading failures that consume the step budget. | sequential-chain | no | none stated | 0 |
| T1-Bench
T1-Bench: Benchmarking Multi-Scenario Agents in Real-World Domains |
Service & dialogue | 2026 | complete interleaved customer-facing scenarios spanning 25 domains within one interaction; conduct structured reasoning across multi-turn user-assistant interactions; correctly use tools while maintaining conversational quality across compositionally complex, interwoven scenario threads | Agent must track tool use, conversational state, and multiple concurrent reasoning threads across interleaved scenarios spanning 25 domains within multi-turn user-assistant interactions. | DAG-with-precedence | no | none stated | 0 |
| tau-Voice
τ-Voice: Benchmarking Full-Duplex Voice Agents on Real-World Domains |
Service & dialogue | 2026 | complete a verifiable grounded task (extended from tau2-bench) via complex multi-turn conversation; adhere to domain policies throughout the conversation; correctly interact with/act on the environment (not just converse) to satisfy the task | The agent must track conversational state, domain-policy constraints, and environment actions across complex multi-turn conversations, while also managing real-time full-duplex audio interaction (turn-taking, interruptions, accents, noise) alongside the underlying task goals. | sequential-chain | yes | none stated | 193.17/mo |
| unnamed user-oriented multi-turn tool-use dialogue generation pipeline
User-Oriented Multi-Turn Dialogue Generation with Tool Use at scale |
Service & dialogue | 2026 | complete multiple distinct task completions accumulating within a single long conversational trajectory; respond appropriately to a simulated user's incremental, turn-by-turn requests and feedback rather than resolving the whole task in one shot | The agent must track the state of multiple, potentially overlapping task completions within one long dialogue, correctly interpreting a user simulator's incremental, turn-by-turn requests and feedback rather than resolving everything from an initial fully-specified prompt. | sequential-chain | unclear | none stated | 20.25/mo |
| τ-Knowledge (τ-Banking domain)
τ-Knowledge: Evaluating Conversational Agents over Unstructured Knowledge |
Service & dialogue | 2026 | retrieve the correct policy document(s) from a densely interlinked, ~700-document knowledge base; coordinate retrieved natural-language knowledge with tool outputs to execute a policy-compliant account update; produce a verifiable, policy-compliant state change during a live customer-support interaction | Agent must track which of many interconnected banking policy documents are relevant to the current request, reconcile them with live tool outputs, and ensure the resulting account update remains policy-compliant and verifiable. | DAG-with-precedence | unclear | none stated | 254.17/mo |
| AgentChangeBench
AgentChangeBench: A Multi-Dimensional Evaluation Framework for Goal-Shift Robustness in Conversational AI |
Service & dialogue+ Business & enterprise | 2025 | complete the original task objective before a mid-dialogue goal shift; recognize and adapt to a mid-dialogue goal shift triggered by one of five user personas; recover task success after a goal shift within enterprise domains (e.g., airline booking, retail); use tools efficiently and non-redundantly while adapting to shifting goals | The agent must track the original task objective, detect when a mid-dialogue goal shift has occurred, measure its own recovery latency (Goal-Shift Recovery Time), and avoid redundant tool calls (Tool Call Redundancy Rate) while re-establishing progress toward the new goal. | sequential-chain | no | none stated | 30.27/mo |
| ECom-Bench
ECom-Bench: Can LLM Agent Resolve Real-World E-commerce Customer Support Issues? |
Service & dialogue | 2025 | resolve a real-world e-commerce customer-support issue via multimodal interaction; adapt to a dynamic, persona-driven simulated user across the dialogue; handle diverse business scenarios reflecting real-world complexity; achieve consistent success across repeated trials (pass^3 metric) | The agent must track the evolving state of a persona-driven simulated customer's issue, multimodal evidence presented during the conversation, and consistency of resolution across repeated trials. | sequential-chain | no | none stated | 302.14/mo |
| Food4All
Food4All: An Agentic Framework and Benchmark for Food Resource Navigation with Adaptive User Understanding |
Service & dialogue+ Embodied & robotics | 2025 | ground a user's underspecified/noisy help-seeking dialogue into a valid resource recommendation; retrieve resources satisfying single food needs; satisfy composite cases with access or document constraints; handle non-ideal user interaction traits (unreasonable demands, rambling, impatience, incomplete answers, inconsistent information) while completing the referral | The agent must track grounded requirements (schedule, eligibility, intake, document constraints), the set of valid retrieved resources it must preserve into the final recommendation, and the user's non-ideal interaction trait across the dialogue. | sequential-chain | yes | none stated | 10.09/mo |
| IntellAgent
IntellAgent: A Multi-Agent Framework for Evaluating Conversational AI Systems |
Service & dialogue+ Multi-agent orgs | 2025 | navigate a multi-turn dialogue while integrating domain-specific APIs; adhere to strict, graph-modeled policy constraints throughout the conversation; handle realistic, policy-driven event scenarios generated at varying complexity levels | The agent must track which of several interacting domain-specific policy constraints apply at each point in a multi-turn dialogue, per a graph-based policy model, while integrating API calls consistent with those constraints. | DAG-with-precedence | no | none statedother:not-stated | 261.3/mo |
| tau-break
Effective Red-Teaming of Policy-Adherent Agents |
Service & dialogue | 2025 | adhere consistently to domain policies (e.g. refund eligibility, cancellation rules) across a customer-service conversation; correctly refuse any request that would violate policy; remain helpful and natural while resisting persuasive, policy-aware adversarial pressure from CRAFT | The agent must track which policies apply to the current request and remain consistent with them across up to 30 dialogue turns, resisting cumulative persuasive pressure (emotional manipulation, coercive framing) from an adversarial user. | sequential-chain | no | up to 30 dialogue turnsturns | 110.73/mo |
| τ²-Bench
τ2-Bench: Evaluating Conversational Agents in a Dual-Control Environment |
Service & dialogue | 2025 | coordinate actions with an active user who also uses tools to modify a shared, dynamic environment (dual-control); complete diverse, compositionally-generated verifiable tasks built from atomic components; guide/communicate with the user effectively, not only reason internally; correctly attribute/avoid errors arising from reasoning vs communication/coordination failures | The agent must track the shared dynamic world state as modified by both itself and the user, communicate effectively to guide the user's tool use, and maintain a compositional task's atomic sub-requirements across the dual-control interaction. | DAG-with-precedence | no | none stated | 45630.4/mo |
| Education & tutoring2 artifacts | ||||||||
| TeachArena
TeachArena: Are Language Agents Ready for Realistic Teaching Work? |
Education & tutoring+ Business & enterprise | 2026 | infer a warranted pedagogical/teaching decision from evidence (professional pedagogical judgment); adapt tutoring support as the learner's state changes across a multi-turn tutoring session (situated multi-turn tutoring); carry an instructor's request through a learning-management system (LMS) to a completed, verified intervention (end-to-end LMS teaching workflow) | The agent must track the learner's evolving state across a multi-turn tutoring session, evidence supporting a pedagogical insight, and the instructor's request as it is carried through to a persistent, verified artifact or environment state in the LMS. | hierarchical | no | none stated | 0 |
| TutorBench
DeepTutor: Towards Agentic Personalized Tutoring |
Education & tutoring+ Assistants & memory | 2026 | deliver citation-grounded tutoring on a specific problem; generate difficulty-calibrated follow-up questions matched to a learner's diagnosed knowledge gaps; continuously adapt personalization to the student's evolving needs across an interactive tutoring session | The agent must track the learner's diagnosed knowledge gaps and profile, the history of prior interactions, and adapt difficulty calibration accordingly across a multi-turn tutoring dialogue. | sequential-chain | yes | none statedother:not-stated | 0 |
| Embodied & robotics46 artifacts | ||||||||
| CoCoBench
CoCoBench: A Cooperative Coordination Benchmark for Embodied Multi-Agent Task Planning |
Embodied & robotics+ Multi-agent orgs | 2026 | allocate tasks correctly among cooperating embodied agents; respect sequential-ordering constraints between agents' actions; respect mutual-exclusion constraints over shared objects/spaces; execute correct handoff coordination between agents; satisfy conjunctive multi-object-destination goals within executable household tasks | The multi-agent system must track task allocation assignments, ordering/precedence constraints, mutual-exclusion locks on shared objects, and handoff status between agents, calibrated against a task-specific step budget H derived from oracle trajectories. | DAG-with-precedence | yes | ~18.4 mean executed action steps (best model, successful episodes); task-specific step budget H calibrated from oracle-validated trajectoriesagent-steps | 0 |
| DeCoNavBench
DeCoNav: Dialog enhanced Long-Horizon Collaborative Vision-Language Navigation |
Embodied & robotics+ Multi-agent orgs | 2026 | achieve relay-style handoffs between two robots collaborating on a shared long-horizon navigation task; reach rendezvous points to exchange information/objects between robots; dynamically reassign and replan subgoals in response to new cross-agent evidence, uncertainty, or conflicts | Each robot must track its own navigation progress plus the other robot's evidence/uncertainty/conflicts communicated via event-triggered dialogue, in order to reassign subgoals and replan under synchronized execution. | DAG-with-precedence | yes | none stated | 51.0/mo |
| DunphyBench
Long-Horizon Embodied Decision-Making via Multimodal Memory Compression |
Embodied & robotics+ Assistants & memory | 2026 | navigate through multiple embodied housing environments to gather evidence; integrate multimodal, multi-source input into coherent knowledge under partial observation; make a final housing decision aligned with multi-dimensional, partly implicit human preferences | Agent must accumulate multimodal evidence across multiple candidate housing environments under partial observation and integrate it into a coherent decision aligned with multi-dimensional, partly implicit human preferences. | sequential-chain | no | none stated | 0 |
| EAS resilience evaluation framework
Resilience Matters for Embodied Agents System: New Metrics, Systematic Evaluation, and Optimization |
Embodied & robotics | 2026 | complete a household task despite perturbations/unexpected disruptions during execution; recover, stabilize, and gracefully extend behavior after a perturbation rather than simply succeeding or failing outright; maintain low recovery cost and stability across iterative system updates | The evaluation layer must track the full execution trajectory of each household task under perturbation, not just final success/failure, to compute process-level resilience metrics like recovery cost and stability. | hierarchical | yes | none stated | 0 |
| Ego2World
Ego2World: Compiling Egocentric Cooking Videos into Executable Worlds for Belief-State Planning |
Embodied & robotics | 2026 | plan under partial observation using a belief graph built from local observations; remember objects and track state changes in a hidden symbolic world graph; recover and replan when actions fail without observing the true world state; complete cooking-task objectives via correct sequences of graph-governed state transitions | The agent must maintain and update its own partial belief graph of the world (objects, state changes) using only local observations and execution feedback, separate from the simulator's hidden ground-truth world graph, and replan without directly observing true state. | DAG-with-precedence | no | none stated | 0 |
| EmbodiedWorldBench (ABot-AgentOS)
ABot-AgentOS: A General Robotic Agent OS with Lifelong Multi-modal Memory |
Embodied & robotics+ Assistants & memory | 2026 | navigate indoor/outdoor/hybrid scenes to reach targets; search for and identify specified objects; conduct NPC dialogue as part of a task; respond appropriately to dynamic events introduced mid-task; achieve trace-grounded, verifiable task completion across difficulty levels | Agent must maintain a persistent, source-grounded multi-modal graph memory across dialogue, visual observations, spatial context, temporal relations, and task traces, continually updated and consulted for later decisions. | hierarchical | unclear | none stated | 0 |
| FullHome (via the TaskGround Ground-Infer-Execute framework)
TaskGround: Structured Executable Task Inference for Full-Scene Household Reasoning |
Embodied & robotics | 2026 | identify task-relevant entities within a complete, cluttered household scene; recover intended task conditions implied by a situated (underspecified) household request; resolve ordering constraints among sub-actions from surrounding scene context; produce a grounded, skill-level action sequence that correctly executes the inferred task structure | The agent must track which entities, conditions, and ordering constraints it has inferred as task-relevant from a complete household scene, and maintain this executable task structure while compiling and executing a grounded skill-level action sequence. | sequential-chain | no | none stated | 0 |
| HumanCLAW-Bench
HumanCLAW: Can Vision-Language Models Act Through a Body? |
Embodied & robotics | 2026 | find a specified target within an indoor scene; navigate to the target while avoiding obstacles/maintaining balance; interact with the target once reached; maintain embodied self-awareness of the body's own state (position, goal-reached status, collisions) throughout | The VLM must track where its body is, whether it has reached the goal, and whether it has hit an obstacle -- embodied self-awareness -- across each find-navigate-interact episode, since the paper finds this tracking, not target recognition, is the main bottleneck. | sequential-chain | yes | reported average of 78.5 steps per episode for baseline conditionsagent-steps | 21.0/mo |
| IMBench
IMBench: A Benchmark for Intuitive Robotic Manipulation |
Embodied & robotics | 2026 | infer task-relevant physical structure via physical reasoning before acting; generate a feasible action sequence satisfying explicit task constraints (contact-rich manipulation, tool use, multi-stage dependencies); complete each of 35 manipulation tasks across scalable, diverse scenarios | Agent must track inferred physical structure (contacts, affordances), the current stage of a multi-stage manipulation plan, and constraint satisfaction across execution. | DAG-with-precedence | yes | none stated | 0 |
| LMEE-Bench
Explore with Long-term Memory: A Benchmark and Multimodal LLM-based Reinforcement Learning Framework for Embodied Exploration |
Embodied & robotics+ Assistants & memory | 2026 | perform multi-goal navigation across an embodied environment; answer memory-based questions using episodic memory accumulated during exploration; unify exploratory cognition with decision-making to support lifelong learning; proactively query memory and select exploration frontiers next | The agent must track its accumulated episodic memory, current exploration frontier, and multiple concurrent navigation goals across a long-horizon embodied exploration episode. | sequential-chain | no | none stated | 101.25/mo |
| LongAct
When Robots Do the Chores: A Benchmark and Agent for Long-Horizon Household Task Execution |
Embodied & robotics | 2026 | understand a free-form (non-templated) household instruction and decompose it into the right sequence of sub-tasks; manage dependencies among sub-tasks (ordering, shared resources) via a DAG-based plan; maintain persistent memory (spatial and episodic) across a long-horizon household task; adapt the plan reflectively as execution proceeds | The agent must maintain persistent spatial and episodic memory of what it has already done and observed in the household, use this to keep dependency-aware track of remaining sub-tasks, and adaptively replan as new information arises over the long-horizon task. | DAG-with-precedence | yes | human executors typically require 500+ steps; VLM-based agents often exceed 2,000 steps to complete the same long-horizon household taskagent-steps | 20.5/mo |
| MECoBench
MECoBench: A Systematic Study of Multimodal Agent Collaboration in Embodied Environments |
Embodied & robotics+ Multi-agent orgs | 2026 | complete embodied tasks under two cooperation structures and three collaboration modes; coordinate communication among multiple multimodal agents in a visually grounded environment; remain robust to noisy priors and exploration conditions via collaboration | Agents must track and communicate task-relevant state among team members across two cooperation structures and three collaboration modes to complete embodied tasks under noisy priors and exploration conditions. | set-of-independent | no | none stated | 20.67/mo |
| MultiUAV-Plat Benchmark
MultiUAV-Plat: An LLM-Oriented Platform, Benchmark and Framework for Multi-UAV Collaborative Task Planning |
Embodied & robotics+ Multi-agent orgs | 2026 | assign UAVs to targets correctly under partial observability; perform area search coverage objectives; perform area assignment and patrol objectives across multiple UAVs; satisfy each of the mission's validation checks (mean 6.26/task) via correct multi-vehicle coordination | The framework (Agent4Drone) must track memory, observation, task understanding, planning, execution, and verification state across each mission, checking role-based information access and validation logic per task within a session. | set-of-independent | yes | none statedother:not-stated | 0 |
| PARTNR-Dialog
LLawCo: Learning Laws of Cooperation for Modeling Embodied Multi-Agent Behavior |
Embodied & robotics+ Multi-agent orgs | 2026 | complete a shared household task cooperatively with a partner agent; communicate with the partner (e.g., 'talk when necessary') to align on task objectives; align actions with the partner's behavior and the environment state to avoid inefficient or conflicting actions; apply learned high-level behavioral laws (e.g., 'wait for partner') during planning | The agent must track its partner's current actions/state, the shared household task's progress, and which high-level behavioral laws (e.g., 'talk when necessary,' 'wait for partner') apply at each point in the cooperative plan. | sequential-chain | no | none stated | 0 |
| PLanAR
PLanAR: Planning-Language-Grounded Agentic Reasoning for Robot Manipulation |
Embodied & robotics | 2026 | represent and update object-predicate scene states as manipulation proceeds; select and sequence action schemas (with preconditions/effects) to satisfy an open-vocabulary manipulation goal; detect execution failures via stepwise symbolic-effect verification and replan; complete long-horizon kitchen workflows composed of many chained manipulation sub-goals | The agent must maintain a symbolic scene-state representation (object predicates), verify after each action whether expected effects were achieved, and update/replan when execution deviates from the expected symbolic plan across a long-horizon kitchen workflow. | DAG-with-precedence | no | none statedother:not-stated | 30.43/mo |
| RescueBench
RescueBench: Can Embodied Agents Save Lives in the Wild ? |
Embodied & robotics | 2026 | explore an unfamiliar environment under multimodal uncertainty to locate a target (multimodal exploration stage); physically rescue/reach the identified target (target rescue stage); navigate back using retained spatial memory (memory-guided return stage); complete a final handoff of the rescued target/information (final handoff stage) | The agent must retain spatial memory of the environment (for the return stage), track clue ambiguity/target identification state, and manage per-level time budgets (180-300 seconds) across the four-stage pipeline. | sequential-chain | yes | episodes are governed by level-dependent time limits: L1-L2: 180 seconds, L3: 240 seconds, L4-L5: 300 seconds, with no step capother:wall-clock-seconds | 0 |
| RoboGraph
Compiling and Benchmarking Task-State Horizons for Embodied Agents |
Embodied & robotics | 2026 | track the evolving span of task-relevant state transitions (task-state horizon, TSH) induced by both exploration and environmental dynamics; correctly execute high-level plans as task-relevant world state changes, including due to unexpected failures/interventions; complete each of 588 episodes across 84 scenes with varying task-state horizons | The agent must maintain, explore, and update task-relevant world state over the course of a long-horizon rollout, since performance is explicitly measured as a function of the task-state horizon (TSH) -- the span of state transitions it must track. | DAG-with-precedence | yes | none stated | 0 |
| SMH-Bench (built on HomeEnv)
SMH-Bench: Benchmarking LLM Agents for Environment-Grounded Reasoning and Action in Smart Homes |
Embodied & robotics | 2026 | execute explicit device control and query commands across a smart home with up to 135 devices; schedule automation tasks that must fire correctly over time; correctly handle ambiguous user instructions; personalize reasoning to user intent/preferences as home complexity increases | Agent must track the state of many concurrent devices across a multi-room home, especially for automation-scheduling tasks that must persist and correctly re-trigger over time, and must resolve ambiguous instructions against current home/device state. | set-of-independent | yes | none stated | 20.67/mo |
| SpatialWorld
SpatialWorld: Benchmarking Interactive Spatial Reasoning of Multimodal Agents in Real-World Tasks |
Embodied & robotics | 2026 | actively gather egocentric visual evidence under partial observability to resolve a task; complete household, travel, or social-collaboration tasks requiring interactive spatial reasoning; express decisions via a unified text-based action interface across eight heterogeneous simulator backends | Agent must track what it has and has not yet observed (partial observability), accumulate egocentric visual evidence over the course of the task, and reconcile this with a reference trajectory/terminal-state verifier. | sequential-chain | unclear | none stated | 20.67/mo |
| TypeGo (Kalos prototype)
TypeGo: An OS Runtime for Embodied Agents |
Embodied & robotics | 2026 | execute multiple concurrent per-task processes/goals on a shared physical robot body without conflicting over physical subsystems; preempt, resume, or replace a lower-priority task/goal when a new goal arrives; maintain fast first-action responsiveness while longer-horizon planning continues asynchronously in the background | The runtime must track which per-task processes currently hold which physical subsystems, their priority/preemption state, and pending speculative skill-streaming actions, to arbitrate concurrent goals in real time. | set-of-independent | yes | none stated | 0 |
| UniETP
UniETP: Unifying Environments for Generalizable Embodied Task Planning |
Embodied & robotics | 2026 | execute a sequence of atomic actions within an interactive environment to complete a user-specified task; handle varying task-logic complexity; handle varying instance-grounding complexity; handle varying instruction-understanding complexity, across four unified simulators (AI2-THOR, VirtualHome, Habitat, BEHAVIOR) | The agent must track its progress executing a sequence of atomic actions toward a user-specified goal within a unified observation/action space, while contending with varying levels of task-logic, instance-grounding, and instruction-understanding difficulty across four different underlying simulators. | sequential-chain | yes | none stated | 0 |
| VGEBench
Towards Generalizable Visually Grounded Exploration of Household Devices |
Embodied & robotics | 2026 | form a hypothesis about how to operate a novel household device from visual cues alone (no manual); test the hypothesis via physical interaction and interpret feedback; refine/correct the hypothesis and action based on observed feedback (Hypothesis-Interaction-Refinement loop) until the device is successfully operated | The agent must track its current hypothesis about the device's operation, the history of physical feedback received from prior interaction attempts, and dynamically-calculated interaction budgets based on task complexity, maintaining long-horizon state across the exploration loop. | open-ended | yes | none stated | 0 |
| WorldLines
WorldLines: Benchmarking and Modeling Long-Horizon Stateful Embodied Agents |
Embodied & robotics | 2026 | answer Memory QA questions correctly using long-term household interaction history; produce correct Embodied Task Plans grounded in remembered user routines and past world/device states; maintain visibility-aware (partially observable) memory across a temporally extended household trace | The agent must maintain visibility-aware memory of user routines, world states, past actions/dialogues, and object/device state changes across long temporally extended household traces to answer Memory QA and produce embodied plans. | other (singleton) | yes | none stated | 10.33/mo |
| Household Task Planning with Multi-Objects State and Relationship Using Large Language Models Based Preconditions Verification
Household Task Planning with Multi-Objects State and Relationship Using Large Language Models Based Preconditions Verification |
Embodied & robotics | 2025 | achieve each targeted object state change (e.g. turning an appliance on/off); achieve each targeted object placement goal; verify environmental preconditions are met before executing each action; reformulate an action step automatically when a precondition is not satisfied | The agent must track current object states, identifiers, and relationships in the environment, re-verifying preconditions before each action and updating its plan when environmental state does not match expectations. | sequential-chain | no | none statedother:not-stated | 0 |
| 3DMem-Bench (via 3DLLM-Mem)
3DLLM-Mem: Long-Term Spatial-Temporal Memory for Embodied 3D Large Language Model |
Embodied & robotics+ Assistants & memory | 2025 | correctly recall and apply past spatial-temporal observations (episodic memory) to complete embodied tasks in multi-room 3D environments; answer questions and produce captions grounded in accumulated long-term memory of the 3D scene; focus on task-relevant information while maintaining memory efficiency across complex, long-horizon environments | The agent must maintain and selectively query an episodic memory of past spatial and temporal observations across long-horizon, multi-room 3D trajectories, focusing on task-relevant information while remaining memory-efficient. | sequential-chain | yes | none stated | 291.81/mo |
| ArtiBench (with ArtiBrain)
ArtiBench and ArtiBrain: Benchmarking Generalizable Vision-Language Articulated Object Manipulation |
Embodied & robotics | 2025 | manipulate articulated objects (kitchen, storage, office, tool appliances) correctly across parts/instances/categories; decompose and validate a sequence of sub-goals for a long-horizon, multi-object manipulation task; maintain physical consistency across the multi-step interaction | System must track subgoal validation state (via the Task Reasoner), accumulated part-level affordances in the Affordance Memory Bank, and physical consistency across a five-level benchmark spanning cross-part/instance/category variation to long-horizon multi-object tasks. | hierarchical | unclear | none stated | 0 |
| Blocksworld-MCP benchmark
Benchmark for Planning and Control with Large Language Model Agents: Blocksworld with Model Context Protocol |
Embodied & robotics | 2025 | reach one target block-configuration goal state via a sequence of pick/put/stack/unstack actions; satisfy the goal under increasing plan-length complexity categories (step-2 through step-12); generalize plan validity/optimality across diverse agent architectures connected via a standardized MCP tool interface | The agent must track the current block/stack configuration as pick/put/stack/unstack actions are executed, correctly planning toward the target configuration across plans of up to 12 optimal steps, potentially under constrained block size or partial observability. | sequential-chain | no | scenarios span step-2 through step-12 optimal-plan-length categories (45 step-2, 84 step-4, 152 step-6, 151 step-8, 112 step-10, 46 step-12 scenarios), each involving up to five blocksagent-steps | 20.22/mo |
| CookBench
CookBench: A Long-Horizon Embodied Planning Benchmark for Complex Cooking Scenarios |
Embodied & robotics | 2025 | accurately parse a user's complex cooking intent (Intention Recognition); execute the identified cooking goal through a long-horizon, fine-grained sequence of physical actions (Embodied Interaction); correctly use both macro-level operations (placing orders, purchasing ingredients) and fine-grained embodied physical actions | The agent must track the parsed cooking intent, its progress through a long-horizon fine-grained sequence of physical actions, and the state of ingredients/tools obtained via macro-level operations, across a two-stage cooking task. | sequential-chain | no | none stated | 131.0/mo |
| DeCoBench
DeCo: Task Decomposition and Skill Composition for Zero-Shot Generalization in Long-Horizon 3D Manipulation |
Embodied & robotics | 2025 | retrieve and chain reusable atomic manipulation skills to complete a compositional long-horizon 3D manipulation task; generalize zero-shot to novel task compositions not seen during training; execute smooth, collision-free transitions between chained skills | The system must track the currently retrieved skill, the object/gripper state at each transition point, and which atomic subtask in the composed sequence is currently active to ensure valid chaining. | sequential-chain | yes | none stated | 161.0/mo |
| DeliveryBench
DeliveryBench: Can Agents Earn Profit in Real World? |
Embodied & robotics | 2025 | maximize net profit over the course of an operating shift by choosing which deliveries to accept/complete; meet each accepted delivery's deadline; manage limited resources (transportation expense, vehicle battery) across the shift; interact appropriately with other couriers and customers as needed | The agent must track remaining budget/vehicle battery, delivery deadlines for all currently accepted orders, and its evolving location within a procedurally generated 3D city, over an episode lasting several in-game hours and typically more than 100 action steps. | set-of-independent | yes | episodes support long-horizon tasks spanning several in-game hours and typically more than 100 action stepsagent-steps | 50.56/mo |
| Embodied Web Agents Benchmark
Embodied Web Agents: Bridging Physical-Digital Realms for Integrated Agent Intelligence |
Embodied & robotics+ Web & GUI | 2025 | cook a recipe found via web-scale reasoning, using embodied physical actions (cooking); navigate physically using dynamic, web-sourced map data (navigation); shop by combining physical store interaction with online product/price information (shopping); plan tourism activities combining physical exploration with web knowledge (tourism); identify real-world landmarks by cross-referencing physical observation with web knowledge (geolocation) | Agent must maintain a consistent state across both a 3D embodied environment (physical position, observations, actions) and a web interface (retrieved facts, map data, product/price info), integrating the two continuously rather than treating them as separate phases. | set-of-independent | unclear | none stated | 191.27/mo |
| EmbodiedBench
EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied Agents |
Embodied & robotics | 2025 | complete each of 1,128 testing tasks across four environments, from high-level semantic household tasks to low-level atomic navigation/manipulation; demonstrate commonsense reasoning within tasks; understand complex instructions; exhibit spatial awareness and visual perception; engage in long-term planning | The agent must track visual perception of the scene, instruction understanding, spatial layout, and long-term planning state as it composes low-level atomic actions to satisfy high-level household goals. | hierarchical | yes | none stated | 22511.84/mo |
| EmbodiedBrain evaluation suite (General, Planning, and End-to-End Simulation Benchmarks)
EmbodiedBrain: Expanding Performance Boundaries of Task Planning for Embodied Intelligence |
Embodied & robotics | 2025 | complete long-horizon embodied task-planning sequences by correctly building on preceding steps (Guided Precursors); perform accurate spatial perception and adaptive execution across a novel simulation environment; satisfy General, Planning, and End-to-End Simulation benchmark criteria as three complementary evaluation axes | Agent must track its own preceding action steps as guided precursors informing subsequent steps, and its progress must be verifiable across all three of the General, Planning, and End-to-End Simulation evaluation axes. | sequential-chain | yes | none stated | 30.27/mo |
| EMMOE
EMMOE: A Comprehensive Benchmark for Embodied Mobile Manipulation in Open Environments |
Embodied & robotics+ Web & GUI | 2025 | interpret a natural-language user instruction and execute a long-horizon everyday household task combining high-level and low-level embodied sub-tasks; re-plan after execution failures to recover progress toward the instructed goal | Agent must track task-completion progress across high-level sub-tasks and low-level actions, plus failure/replan history, in continuous physical space across a long-horizon episode. | hierarchical | yes | none stated | 40.22/mo |
| FindingDory
FindingDory: A Benchmark to Evaluate Memory in Embodied Agents |
Embodied & robotics+ Assistants & memory | 2025 | recall relevant historical information (images/interactions) collected potentially across multiple days; execute low-level navigation/manipulation actions based on recalled information; complete each of 60 memory-intensive embodied tasks requiring sustained engagement; scale to procedurally extended, longer/harder versions of the same tasks | The agent must retain and retrieve relevant historical images/interactions collected across multiple days in the Habitat simulator, combining that recall with sustained contextual awareness during ongoing navigation/manipulation. | sequential-chain | yes | none stated | 120.8/mo |
| HiMan-Bench
RoboHiMan: A Hierarchical Evaluation Paradigm for Compositional Generalization in Long-Horizon Manipulation |
Embodied & robotics | 2025 | complete atomic long-horizon manipulation tasks under diverse perturbations; complete compositional tasks requiring composing multiple learned manipulation skills; generalize skill composition/scheduling to perturbed or novel conditions; coordinate high-level subgoal planning with low-level execution policies | The system must track which atomic/composed skill is currently executing, how perturbations affect preconditions for subsequent skills, and whether the high-level plan needs revision given low-level execution feedback. | hierarchical | yes | none stated | 90.82/mo |
| LLM-BabyBench
LLM-BABYBENCH: Understanding and Evaluating Grounded Planning and Reasoning in LLMs |
Embodied & robotics | 2025 | predict the consequences of an action on the textual BabyAI grid-world environment state (Predict task); generate a sequence of low-level actions achieving a specified objective (Plan task); decompose a high-level instruction into a coherent sequence of subgoals (Decompose task) | Agent must track the current grid-world state to predict action consequences, track partial plan progress when generating low-level action sequences, and track subgoal ordering/coherence when decomposing high-level instructions. | hierarchical | yes | none stated | 60.38/mo |
| LoHoSet (Ravens simulator)
LoHoVLA: A Unified Vision-Language-Action Model for Long-Horizon Embodied Tasks |
Embodied & robotics | 2025 | decompose a high-level embodied goal into a sequence of sub-tasks; generate correct low-level robot actions to execute each sub-task; maintain coordination between high-level planning and low-level motion control across the whole long-horizon task | The system must track which sub-tasks of the decomposed high-level goal have been completed and maintain closed-loop consistency between the high-level plan and low-level action execution across the task. | hierarchical | yes | none stated | 322.0/mo |
| ManiTaskGen
ManiTaskGen: A Comprehensive Task Generator for Benchmarking and Improving Vision-Language Agents on Embodied Decision-Making |
Embodied & robotics | 2025 | satisfy a process-based instruction requiring a specific sequence of manipulations (e.g. 'move object from X to Y'); satisfy an outcome-based abstract instruction requiring multiple manipulations to reach a goal state (e.g. 'clear the table'); generate a comprehensive, diverse, feasible set of mobile manipulation tasks for any given scene | The agent (and the task generator itself) must track which objects have been moved/placed so far relative to the target scene configuration, verifying that an outcome-based goal (e.g. a cleared table) is only satisfied once all constituent object states are achieved. | hierarchical | no | none statedother:not-stated | 40.25/mo |
| OceanGym
OceanGym: A Benchmark Environment for Underwater Embodied Agents |
Embodied & robotics | 2025 | comprehend and fuse optical and sonar sensor data under low visibility; autonomously explore complex underwater environments; accomplish each of eight realistic underwater task domains; adapt navigation/decision-making to dynamic ocean currents | The agent must integrate perception, memory, and sequential decision-making state (explored regions, sonar/optical evidence) across a long-horizon objective under low-visibility, dynamically-changing underwater conditions. | hierarchical | no | 0.5 hours (t_max, per decision task)wall-clock-hours | 0 |
| REMAC multi-agent environment (on RoboCasa)
REMAC: Self-Reflective and Self-Evolving Multi-Agent Collaboration for Long-Horizon Robot Manipulation |
Embodied & robotics+ Multi-agent orgs | 2025 | decompose and execute long-horizon multi-robot manipulation and navigation tasks (4 task categories, 27 task styles, 50+ objects); perform pre-condition and post-condition checks in the loop to evaluate progress and refine plans; adapt plans dynamically to unexpected scene conditions (e.g. a closed microwave door) via self-evolvement; coordinate parallel task execution across multiple robots | Agent(s) must track scene state, pre/post-condition satisfaction, and per-robot task assignment across a decomposed long-horizon manipulation/navigation plan, adapting the plan when scene-specific conditions invalidate prior assumptions. | DAG-with-precedence | no | none stated | 110.61/mo |
| ResponsibleRobotBench
ResponsibleRobotBench: Benchmarking Responsible Robot Manipulation using Multi-modal Large Language Models |
Embodied & robotics | 2025 | detect and mitigate risks (electrical, chemical, human-related hazards) during a multi-stage manipulation task; reason about safety and physically grounded planning across the task; plan and execute sequences of manipulation actions; engage human assistance when necessary rather than proceeding unsafely | Agent must track detected hazards (electrical, chemical, human-related), current safety status, and multi-stage plan progress, deciding when to escalate to human assistance across the task. | hierarchical | yes | none stated | 30.33/mo |
| RoboCerebra
RoboCerebra: A Large-scale Benchmark for Long-horizon Robotic Manipulation Evaluation |
Embodied & robotics | 2025 | decompose a high-level household instruction into a sequence of dependent subtasks (via GPT-generated instructions); correctly execute each subtask in the sequence despite dynamic object variations; sustain planning, reflection, and memory (System 2 reasoning) across the full extended action sequence, not just react (System 1) to the current observation | The high-level planner must track evolving object/environment state across an average trajectory length of 2,972.4 simulation steps (about 6x longer than existing long-horizon manipulation datasets), correctly sequencing and reflecting on subtasks throughout. | sequential-chain | yes | average trajectory length reaches 2,972.4 simulation steps, about 6x longer than existing long-horizon manipulation datasetsagent-steps | 332.2/mo |
| RoboPilot-Bench
RoboPilot: Generalizable Dynamic Robotic Manipulation with Dual-thinking Modes |
Embodied & robotics | 2025 | execute complex or long-horizon robotic manipulation tasks despite environmental changes; recognize infeasible tasks rather than attempting them blindly; recover from execution errors via closed-loop replanning; succeed across each of 21 tasks spanning 10 manipulation categories | The agent must track execution feedback and environmental state to detect deviations or errors requiring replanning, and separately recognize when a task is infeasible altogether, across a long-horizon sequence of primitive manipulation actions. | sequential-chain | no | none stated | 10.08/mo |
| Robotouille
Robotouille: An Asynchronous Planning Benchmark for LLM Agents |
Embodied & robotics | 2025 | complete overlapping cooking sub-tasks that must be scheduled around each other (e.g. one dish while another cooks); handle interruptions to an in-progress plan without dropping earlier commitments; reason over states/actions that must occur in parallel versus strictly sequentially | Agent must track the state and expected completion time of multiple concurrently in-progress cooking actions, remember interrupted sub-goals, and self-audit its plan as time delays resolve. | DAG-with-precedence | yes | none stated | 382.0/mo |
| unnamed Daily Composite Tasks benchmark
Evaluating Multimodal Large Language Models with Daily Composite Tasks in Home Environments |
Embodied & robotics | 2025 | correctly perform object-understanding sub-tasks (e.g. counting/categorizing objects) within a composite task; correctly perform spatial-intelligence sub-tasks within the same composite task; correctly perform social-activity sub-tasks within the same composite task, all within one dynamic simulated home environment | The agent must track and integrate information across the three jointly-required capability domains (object understanding, spatial intelligence, social activity) within one dynamic, simulated home environment to complete a composite task. | set-of-independent | yes | none stated | 10.08/mo |
| Healthcare8 artifacts | ||||||||
| AgentClinic
AgentClinic: a multimodal benchmark for tool-using clinical AI agents |
Healthcare+ Business & enterprise | 2026 | engage in sequential clinical decision-making across a diagnostic encounter (patient interaction, exam/imaging requests); collect multimodal data under incomplete information via various tools (e.g., notebook, retrieval); reach a correct diagnosis across nine medical specialties and seven languages | Agent must track patient information collected so far (multimodal exam/imaging results), notes taken across a case (persisting via a notebook tool), and its evolving working diagnosis across the encounter. | sequential-chain | yes | none stated | 214.2/mo |
| ClinEnv
ClinEnv: An Interactive Multi-Stage Long Horizon EHR Environment for Agents |
Healthcare+ Business & enterprise | 2026 | progress through an ordered, per-case sequence of clinical decision stages; actively query four specialized information agents before committing to a decision at each stage; commit to correct medications, procedures, and diagnoses at each stage; avoid redundant information-gathering queries as the case progresses | Agent must actively query four specialized information agents at every decision stage before committing to medications, procedures, and diagnoses, tracking what has already been queried to avoid redundant queries as the case progresses. | sequential-chain | yes | none stated | 20.67/mo |
| CodeClinic
CodeClinic: Evaluating Automation of Coding Skills for Clinical Reasoning Agents |
Healthcare | 2026 | in the longitudinal ICU-surveillance task, make a structured monitoring decision every four hours across 25 findings and eight clinical families for the duration of a patient trajectory; in the compositional information-seeking task, answer queries whose difficulty is stratified by compositional dependency depth across 259 tasks in 9 domains; synthesize and compose reusable clinical skills rather than relying on a fixed toolbox | Agent must track the patient's evolving clinical trajectory and make a fresh structured decision every four hours in the longitudinal task, and/or track which reusable clinical skills it has synthesized/composed so far in the compositional task. | DAG-with-precedence | yes | every 4 hours (decision cadence within a longitudinal ICU patient trajectory; total trajectory length not stated)wall-clock-hours | 0 |
| Healink
Bridging the Post-discharge Gap: A Traceable Multi-agent Framework for Safe and Continuous Care |
Healthcare+ Multi-agent orgs | 2026 | maintain continuity of care across longitudinal, patient-specific post-discharge follow-up interactions; generate prescription-grounded, traceable responses reflecting the correct patient/phenotypic/intervention history; prevent cross-departmental drug conflicts while integrating fragmented histories across clinical departments | System must track a patient's full longitudinal clinical history (vectorized records, phenotypic/intervention dimensions) across departments and follow-up interactions, actively cross-checking for drug conflicts as new prescriptions are considered. | DAG-with-precedence | yes | none stated | 0 |
| HealthAgentBench
HealthAgentBench: A Unified Benchmark Suite of Realistic Agentic Healthcare Environments for Challenging Frontier AI Agents |
Healthcare+ Business & enterprise | 2026 | explore raw, heterogeneous healthcare data under minimal instructions; operate within a complex clinical environment to execute a multi-step, end-to-end task; develop research modeling pipelines over EHR data; complete tasks spanning 7 categories across the patient journey and multiple modalities | The agent must track what it has already discovered while exploring raw healthcare data and its intermediate modeling/analysis steps, so its final multi-step solution reflects the full workflow rather than a shortcut response. | sequential-chain | yes | none stated | 72.33/mo |
| MedMCP-Calc
MedMCP-Calc: Benchmarking LLMs for Realistic Medical Calculator Scenarios via MCP Integration |
Healthcare+ Business & enterprise | 2026 | proactively acquire relevant patient data from an EHR database; select the scenario-appropriate medical calculator among alternatives; perform multi-step computation using retrieved data and the selected calculator; retrieve external reference information when needed; complete each of 118 scenario tasks across 4 clinical domains | The agent must track which EHR fields it has retrieved via iterative SQL-based database interaction, which calculator it has selected, and intermediate computed values, since evaluation is process-level rather than only checking the final number. | sequential-chain | yes | none stated | 30.38/mo |
| MedMemoryBench
MedMemoryBench: Benchmarking Agent Memory in Personalized Healthcare |
Healthcare+ Assistants & memory | 2026 | accumulate and correctly retain/retrieve clinically relevant patient information across many sessions; sustain retrieval and reasoning robustness despite memory saturation from continued information influx; correctly evaluate an agent's memory 'live', as it is constructed, rather than only after the fact (evaluate-while-constructing streaming protocol) | The agent's memory system must retain and correctly retrieve clinically relevant information across roughly 2,000 sessions and 16,000 interaction turns per patient archetype, resisting degradation from memory saturation and noise while supporting complex medical reasoning. | sequential-chain | yes | approximately 2,000 sessions and 16,000 interaction turnssessions | 10.25/mo |
| MedAgentSim
Self-Evolving Multi-Agent Simulations for Realistic Clinical Interactions |
Healthcare+ Multi-agent orgs | 2025 | as doctor agent, request relevant medical examinations and imaging results from a measurement agent across multi-turn conversations to reach a diagnosis; iteratively refine diagnostic strategy across successive patient interactions via self-improvement mechanisms | Doctor agent must track which exams/imaging results it has already requested and received, its evolving working diagnosis, and experience-based knowledge accumulated as it interacts with more patients over time. | sequential-chain | yes | none stated | 482.67/mo |
| Information seeking10 artifacts | ||||||||
| DailyReport
DailyReport: An Open-ended Benchmark for Evaluating Search Agents on Daily Search Tasks |
Information seeking+ Science | 2026 | autonomously explore web sources and synthesize information into a comprehensive response to an open-ended daily search query; satisfy each of a task's associated cascade rubrics across disentangled evaluation dimensions | Search agent must track which subtasks/rubric dimensions of an open-ended query it has satisfied, aggregating cascade performance across disentangled dimensions into an interpretable, user-centric final score. | hierarchical | yes | none stated | 0 |
| DeepSearchQA
DeepSearchQA: Bridging the Comprehensiveness Gap for Deep Research Agents |
Information seeking+ Tool & API | 2026 | systematically collate fragmented information from disparate open-web sources; de-duplicate and resolve entities to ensure precision in the exhaustive answer list; reason about stopping criteria within an open-ended search space; complete each causal-chain step, where later steps depend on successful completion of the previous one | The agent must track which sources it has already collated, de-duplicate overlapping entities, maintain the causal-chain dependency state (what has been successfully resolved so far), and decide when to stop searching within an open-ended web space. | sequential-chain | no | none statedother:not-stated | 394.88/mo |
| EarthVerse
EarthVerse: Benchmarking Scientific Agents Across Dynamic Earth Systems and Natural Hazards |
Information seeking+ Science | 2026 | inspect heterogeneous event packages and choose compatible evidence for a natural-hazard investigation; execute transparent calculations reconciling differences across sources; preserve provenance across evidence, scales, units, and calculations in the final answer; produce each of the fine-grained answer units defined by the task's executable ground truth | Agent must track evidence provenance, scales, units, and intermediate calculation results across a multi-stage investigation, since fine-grained answer units are each checked against executable ground truth. | DAG-with-precedence | yes | none stated | 0 |
| HAE-GEO
Evaluating Deep-Search Agents under Hierarchical Web Evidence Poisoning |
Information seeking+ Science | 2026 | search and retrieve web evidence for a consumer-decision query while a poisoning attack (of increasing sophistication, L1-L3) is present; recognize/verify suspicious evidence rather than adopting it uncritically; revise any already-adopted poisoned claims and recover to a trustworthy final recommendation | The agent must track which evidence it has already adopted (and whether that evidence was verified or poisoned), across a multi-turn Search-Scrape interaction, and must be able to revise earlier adopted claims before finalizing its recommendation. | sequential-chain | yes | none stated | 0 |
| MedProbeBench (MedProbe-Eval)
MedProbeBench: Systematic Benchmarking at Deep Evidence Integration for Expert-level Medical Guideline |
Information seeking+ Science | 2026 | retrieve, synthesize, and reason over large-scale external medical evidence to reach expert-level judgment; satisfy 1,200+ task-adaptive rubric criteria for one produced clinical guideline; ensure each of 5,130+ atomic claims in the guideline is precisely evidenced (fine-grained evidence verification) | System must track which rubric criteria have been satisfied and which atomic claims have been verified against source evidence as it synthesizes a guideline from large-scale external knowledge. | set-of-independent | yes | none stated | 10.2/mo |
| Mr.LHDR
Mr.LHDR: A Benchmark for Multimodal Real-World Long-Horizon Deep Research Agents |
Information seeking+ Science | 2026 | derive an average of 12.1 necessary intermediate conclusions along a hidden Node-Relation dependency graph before reaching the final answer; integrate multimodal evidence (images, maps, PDFs, logos, charts, tables, video frames) where at least one non-text element changes the reasoning state; maintain correctness of intermediate conclusions consistent with annotated dependencies, not just the final answer | Agent must track which intermediate conclusions it has established so far, their dependency relationships, and integrate incremental multimodal evidence that can change the reasoning state, across long, irreducible evidence chains. | DAG-with-precedence | yes | avg 12.1 necessary intermediate conclusions per question (mean dependency depth 10.4)other:intermediate-conclusions-per-question | 0 |
| ResearchClawBench
ResearchClawBench: A Benchmark for End-to-End Autonomous Scientific Research |
Information seeking+ Science | 2026 | re-discover a target published paper's scientific artifacts (methods, findings) using only related literature and raw data, with the target paper hidden; satisfy each of several expert-curated, weighted multimodal rubric criteria decomposed from the target scientific artifacts | Agent must track progress across an end-to-end research process (reviewing literature/raw data, designing methodology, producing results) and self-check against the (hidden) target paper's decomposed criteria without directly seeing it. | hierarchical | yes | none stated | 82.0/mo |
| WANDR
WANDR: A Benchmark for Wide and Deep Research |
Information seeking+ Science | 2026 | discover a large set of entities satisfying specified criteria (breadth); investigate each discovered entity through multiple coordinated web searches (depth); return independently verifiable records with supporting sources/excerpts for each required (entity, relationship, evidence) combination; satisfy the full qualification-key hierarchy count (n x m x k records) | The agent must track which entities it has already discovered, which have been investigated to the required depth, and which evidence/sources have been independently verified, maintaining hierarchical completeness across potentially thousands of required records per task. | hierarchical | yes | none stated | 11.0/mo |
| DEER
DEER: A Benchmark for Evaluating Deep Research Agents on Expert Report Generation |
Information seeking+ Science | 2025 | produce an expert-level report satisfying each of 101 fine-grained rubric items across 7 dimensions/25 subdimensions; correctly cite and support both cited and uncited claims with verifiable evidence | The agent/judge must track satisfaction of 101 individual rubric items plus report-wide claim verification (both cited and uncited claims) across one long expert-level report, rather than judging the report as a single holistic pass/fail. | set-of-independent | yes | none stated | 111.22/mo |
| MEM1 composed multi-turn task sequences
MEM1: Learning to Synergize Memory and Reasoning for Efficient Long-Horizon Agents |
Information seeking+ Tool & API | 2025 | answer each of many composed, interdependent objectives within one arbitrarily complex task sequence (e.g. a 16-objective multi-hop QA task); retrieve external information across internal retrieval QA, open-domain web QA, and multi-turn web shopping domains; consolidate memory turn-by-turn, discarding irrelevant/redundant information, to operate with constant memory | Agent must maintain a compact shared internal state that jointly supports memory consolidation and reasoning across many turns of a composed task sequence, integrating new observations while discarding irrelevant or redundant information. | sequential-chain | no | 16other:composed-objectives-per-task | 19613.07/mo |
| Multi-agent orgs44 artifacts | ||||||||
| Is Lying an Emergent Behaviour in LLMs? Evidence from Gaslighting AI agents in a Sustainability Game
Is Lying an Emergent Behaviour in LLMs? Evidence from Gaslighting AI agents in a Sustainability Game |
Multi-agent orgs+ Games & IF | 2026 | manage industrial, military, and ecological resources concurrently; interact with network neighbors to decide on attacks, regeneration claims, and reputation; decide whether/how to lie or bluff about resource regeneration or future attacks; avoid biosphere depletion/extinction while pursuing competitive advantage | Agents must track their own and neighbors' resource levels, reputation/trust information, prior declarations of future attacks, and biosphere/ecological depletion level over repeated network interactions. | open-ended | yes | none stated | 0 |
| AIvilization v0
AIvilization v0: Toward Large-Scale Artificial Social Simulation with a Unified Agent Architecture and Adaptive Agent Profiles |
Multi-agent orgs | 2026 | decompose an agent's persistent life goals into parallel objective branches with tiered re-planning; sustain physiological survival needs while participating in a market economy (AMM-based pricing, production, trade); progress through a gated education-occupation system | Each agent must track its own decomposed objective branches, evolving persona/identity state (dual-process memory), physiological/resource costs, and market conditions (AMM prices) across a long-horizon, large-scale, continuously running deployment. | hierarchical | yes | none stated | 20.29/mo |
| Cattle Trade
Cattle Trade: A Multi-Agent Benchmark for LLM Bluffing, Bidding, and Bargaining |
Multi-agent orgs | 2026 | win auctions under resource constraints; negotiate hidden-offer trade challenges (TCs) profitably; bargain and bluff effectively against other agents; model and exploit opponents via opponent modeling; allocate scarce resources across the whole game while remaining solvent | The agent must track its own resource/capital state, opponent models built from observed bidding/bargaining behavior, and outstanding trade-challenge offers across a single 50-60 turn game, integrating all these rather than treating them in isolation. | set-of-independent | no | 50-60turns | 20.5/mo |
| CityReal
CityReal: Human-Aligned Urban Behavior and City Dynamics Simulation with Large-Scale LLM Agents |
Multi-agent orgs | 2026 | pursue a coherent daily mobility plan (where/when to travel) rather than isolated step-by-step movement choices; pursue a coherent daily activity plan aligned with individual habits and preferences; adapt habits/preferences over time based on accumulated experience and constraints; collectively reproduce observed population-level statistics (crowd density, place popularity, mobility flows, well-being) across the simulated city | Each agent must track its own evolving habits/preferences and prior experience to keep its plans coherent, while the system as a whole tracks population-level alignment statistics across tens of thousands of agents. | hierarchical | unclear | none stated | 0 |
| CIVA
Human Values Matter: Investigating How Misalignment Shapes Collective Behaviors in LLM Agent Communities |
Multi-agent orgs | 2026 | form and sustain a community via autonomous communication among agents; explore the environment and compete for shared resources; maintain individual value orientations under systematic manipulation of value prevalence | The environment must track each agent's evolving value orientation and behavior, aggregate resource-competition outcomes across the community, and detect emergent collective failure modes (catastrophic collapse) as the simulation proceeds. | open-ended | no | none statedother:not-stated | 20.4/mo |
| ClawArena-Team
ClawArena-Team: Benchmarking Subagent Orchestration and Dynamic Workflows in Language-Model Agents |
Multi-agent orgs | 2026 | create and delegate work to specialized subagents from a fixed, locally-served pool; orchestrate subagents' parallel, asynchronous returns through a dynamic workflow; grant least-privilege workspace permissions correctly to each subagent; route each piece of work to the subagent with appropriate perception/modality access; correctly incorporate 72 staged updates across a scenario's evaluation rounds | The main agent must track which subagents have been granted which workspace privileges, which modality/perception requirements each pending sub-task needs, and how staged updates change the scenario across many evaluation rounds, all while natively perceiving only text. | DAG-with-precedence | no | 258 evaluation rounds and 72 staged updates across 41 scenarios (~6.3 rounds/scenario average)other:evaluation-rounds-and-staged-updates | 0 |
| COOP2 / COOP2-Repair
COOP$^2$: Defining, Observing, and Repairing Cooperation in LLM Multi-Agent Systems |
Multi-agent orgs | 2026 | satisfy each verifiable cooperative requirement defined for a cooperative task; ground high-level natural-language cooperation dynamics (plans, messages, revisions) in grounded environment actions; detect where and why cooperation breaks down over the course of task progress; predict constraint failures from group plans and open targeted repair channels for guided revisions | The framework must track natural-language plans/messages/revisions alongside grounded environment task progress, monitor verifiable cooperative requirements over time, and identify where cooperation breaks down to trigger targeted repair. | DAG-with-precedence | no | none statedother:not-stated | 0 |
| DecisionBench
DecisionBench: A Benchmark for Emergent Delegation in Long-Horizon Agentic Workflows |
Multi-agent orgs | 2026 | complete underlying task-suite objectives (GAIA, tau-bench, BFCL multi-turn) while deciding when and to whom to delegate; route sub-tasks to the most capable available peer model via a delegation interface (call_model, optional read_profile); approach the counterfactual perfect-delegation ceiling across quality, cost, latency, delegation rate, and routing fidelity-at-k | Agent must track which of 11 peer models (across 7 vendor families) to delegate a sub-task to via a fixed delegation interface, with routing choices jointly scored across quality, cost, latency, delegation rate, and routing fidelity-at-k relative to a counterfactual ceiling. | DAG-with-precedence | no | none stated | 20.5/mo |
| Emergence World
Emergence World: A Platform for Evaluating Long-Horizon Multi-Agent Autonomy |
Multi-agent orgs | 2026 | govern a shared population/settlement through democratic mechanisms with consequential outcomes; manage persistent memory and 120+ specialized tools to act in a live, externally-grounded world (weather, news, internet); sustain the population/world's stability over weeks-to-months rather than collapsing; interact and cross-influence with agents from different model vendors sharing the same world | Agents must track persistent memory across three memory systems, the state of a shared spatial and governance world grounded in live external data, and consequences of prior democratic decisions, continuously over a run lasting weeks to months (illustrated by a 15-day study). | open-ended | no | 15-day cross-vendor study; framed as supporting runs of weeks to monthswall-clock-days | 20.67/mo |
| EntCollabBench
Beyond the All-in-One Agent: Benchmarking Role-Specialized Multi-Agent Collaboration in Enterprise Workflows |
Multi-agent orgs+ Business & enterprise | 2026 | collaboratively modify enterprise system states across 11 role-specialized agents in six departments (Workflow subset); make policy-grounded approval decisions under permission constraints (Approval subset); correctly delegate, transfer context, ground parameters, and commit to decisions across roles | Agents must track permission-isolated system state, correctly delegate sub-tasks and transfer context across roles, ground shared parameters, and commit to policy-grounded decisions, verified via execution traces, database state, and deterministic policy adjudication. | DAG-with-precedence | no | none stated | 10.25/mo |
| Lingjing
Lingjing: A Simulation Testbed for Multi-Agent Embodied Tasks in Open-Ended Cities |
Multi-agent orgs+ Embodied & robotics | 2026 | coordinate heterogeneous agents (UAVs, ground robots, autonomous vehicles) to complete a shared natural-language mission in an evolving city; manage resource constraints and communication (star or broadcast) among multiple agents; complete each of nine urban tasks under a shared engine-in-the-loop protocol; maintain grounding and effective long-horizon execution despite persistent bottlenecks | Agents must track evolving relation-graph state, resource consumption, and communication history across an episode, since each episode is recorded as an attribution-ready replay linking trajectories and communication to these state changes for systematic diagnosis. | DAG-with-precedence | no | none stated | 0 |
| Moltbook-style multi-agent simulation platform
Got a Secret? LLM Agents Can't Keep It: Evaluating Privacy in Multi-Agent Systems |
Multi-agent orgs | 2026 | maintain persistent social presence across simulated community interactions over a month; decide whether/when to disclose sensitive information under social pressure; resist socially contagious privacy-leakage behavior observed from peers; follow explicit privacy instructions/safeguards while participating in community interactions | Agent must track ongoing social context, peer disclosure behavior, and any active privacy instructions/safeguards across a persistent, simulated month-long community, in order to decide when to disclose or withhold sensitive information. | open-ended | no | a simulated monthother:simulated-month | 0 |
| Moltbook-style simulation platform
Does Safety Molt? Evaluating LLM Safety in Multi-Agent Social Environments |
Multi-agent orgs | 2026 | engage in ongoing social interactions within online communities over a simulated month; decide whether to disclose sensitive/private information under varying social pressure; observe and potentially imitate peer disclosure behavior (social contagion); maintain persona-consistent behavior across a persistent multi-agent social environment | Each agent must track its own privacy stance/instructions, the social context/pressure created by peer disclosures it has observed, and persona-consistent behavior across a persistent, thousands-of-agents community over a simulated month. | open-ended | yes | a simulated month (~30 simulated days)simulated-days | 10.25/mo |
| Online Agent-as-a-Judge (life-simulation evaluation environment)
Online Agent-as-a-Judge: Situation-Generating Evaluation for Interactive Agents |
Multi-agent orgs | 2026 | exhibit correct social behavior across 32 designer-authored social criteria (e.g. conflict handling); respond appropriately to situations actively elicited by an in-world evaluator agent through native dialogue/action protocol; maintain consistent immediate and downstream behavior across a life-simulation episode | Target agent must respond consistently across both immediate responses and downstream behavior as an in-world evaluator agent actively elicits and observes situations relevant to 32 distinct social criteria. | set-of-independent | yes | none stated | 0 |
| OrchBench
OrchBench: Evaluating Multi-Agent Orchestration Plans in Isolation via Deterministic Simulation |
Multi-agent orgs | 2026 | assign subtasks in a DAG of parallelizable, interdependent subtasks to worker agents; specify cross-agent information transfers and their retention ratios; preserve task-critical information as it is transferred across agents; optimize result quality, makespan, and token cost jointly for an orchestration plan | The orchestration planner must track the DAG's task dependencies, the per-agent context limit and agent budget, and how much task-critical information is retained across each cross-agent information transfer. | DAG-with-precedence | yes | none stated | 0 |
| PolicySim
PolicySim: An LLM-Based Agent Social Simulation Sandbox for Proactive Policy Optimization |
Multi-agent orgs | 2026 | achieve platform-specific behavioral realism as a simulated user agent population; assess the impact of a candidate intervention policy (recommendation/content-filtering) on opinions/polarization before deployment; adapt intervention policy over time via a contextual bandit responding to dynamic network structure | The simulation must track evolving user opinions/behavior, the current dynamic network structure, and the platform's intervention policy state (bandit context) jointly at both micro (individual) and macro (ecosystem) levels. | other (singleton) | no | none stated | 81.33/mo |
| ScioMind
ScioMind: Cognitively Grounded Multi-Agent Social Simulation with Anchoring-Based Belief Dynamics and Dynamic Profiles |
Multi-agent orgs | 2026 | maintain an evolving personal belief state via memory-anchored, personality-conditioned updates; sustain persistent, experience-driven belief formation through a hierarchical memory architecture; produce heterogeneous, dynamically-updated agent profiles (personality, rationale, internal state) via retrieval; collectively reproduce realistic community-level opinion dynamics (polarisation, diversity, extremization, trajectory stability) in a policy-debate scenario | Each agent must track its own evolving beliefs, anchoring strength, and dynamic profile across the simulation while the framework separately tracks community-level polarisation, diversity, extremization, and trajectory-stability metrics over time. | open-ended | no | none statedother:not-stated | 10.25/mo |
| SidConArena
SidConArena: An Environment Evaluating Agents in Open-Ended,Positive-Sum Bargaining Game |
Multi-agent orgs+ Games & IF | 2026 | negotiate binding trades via natural-language bargaining; produce goods via deterministic converter-based production; win sealed-bid auctions for long-term assets; plan investment under delayed returns across a finite-horizon multi-player economy | Agents must track private valuations/constraints, negotiated trade commitments, converter-production state, and the value/timing of returns from won long-term assets across a finite-horizon multi-round economy. | DAG-with-precedence | yes | none stated | 0 |
| SocialGrid
SocialGrid: A Benchmark for Planning and Social Reasoning in Embodied Multi-Agent Systems |
Multi-agent orgs+ Embodied & robotics | 2026 | navigate the embodied grid environment to complete assigned tasks; plan a sequence of actions while avoiding obstacles; detect and reason about deceptive teammates (an Among-Us-style social-deduction goal) | Agent must track its own task-completion progress, accumulate behavioral evidence about other agents over multiple rounds to judge who is deceptive, and adapt after each round of adversarial league play. | set-of-independent | yes | none stated | 0 |
| SovSim (Sovereignty over the Commons Simulation)
Bosses, Kings, and the Commons: Cooperation Under Power Asymmetry in LLM Societies |
Multi-agent orgs | 2026 | individually extract resources from a shared commons to maximize personal outcome; collectively sustain the shared resource's viability across repeated extraction rounds; for the power-asymmetric agent (boss/king): exercise disproportionate control over collective extraction outcomes | Agents must track the remaining shared resource pool, their own accumulated extraction/outcomes, and (where applicable) the power-asymmetric agent's extraction decisions, across the repeated rounds to judge whether continued extraction remains sustainable. | sequential-chain | yes | 12 decision rounds per simulation (4 agents managing a shared pool)turns | 10.25/mo |
| The Energy Society
The Energy Society: A Simulation Environment for Studying Agent Cooperation under Survival Pressure |
Multi-agent orgs | 2026 | earn energy by completing jobs or receiving donations to avoid deactivation; manage token-cost-linked energy expenditure when generating tokens (larger models cost more energy per token); decide whether to cooperate (recommend jobs, donate energy) or compete with other agents under scarcity; survive (avoid reaching zero energy) across the whole simulated run | Each agent must track its own remaining energy balance, the jobs available and their difficulty/reward each round, and (in cooperative settings) other agents' need for donations, across 30 rounds of the simulation, to avoid deactivation. | sequential-chain | yes | 30 rounds per simulation (5 agents, 12 jobs per round)turns | 0 |
| WeClawArena
WeClawArena: An Auditable Sandbox and Benchmark for Cross-User Agents Collaboration and Security in Human-Centered Agent Networks |
Multi-agent orgs+ Assistants & memory | 2026 | complete collaborative tool-use tasks across multiple owned agents' personal workspaces; respect privacy/authority boundaries between owners' workspaces (files, records, tools, policies not directly visible across owners); resist or correctly handle four distinct attack-vector variants per base task while completing the benign task variant; maintain utility while keeping attack success low, jointly assessed via task breakdown, privacy leakage, poisoned evidence, and invalid authority paths | The sandbox tracks peer messages, tool calls, resource operations, governed decisions, and final workspace states across owners, and the agent must track its own permissions/authority path while collaborating. | DAG-with-precedence | yes | none stated | 11.0/mo |
| year-long IT company simulation (TaskWeave testbed)
Can LLM Agents Sustain Long-Horizon Organizational Dynamics? |
Multi-agent orgs | 2026 | propagate goals through an organizational hierarchy; execute tasks that depend on the outcomes of prior task execution; accumulate and maintain artifacts produced over a year-long simulation; sustain organizational coherence and execution grounding across the year | The framework must maintain planning state through a Formulate-Partition-Diagnose-Align cycle and ground execution via dependency-aware trace memory across a simulated year, tracking accumulated artifacts and adapting to external environment changes. | hierarchical | no | one yearsimulated-years | 0 |
| AgentSociety
AgentSociety: Large-Scale Simulation of LLM-Driven Generative Agents Advances Understanding of Human Behaviors and Society |
Multi-agent orgs | 2025 | conduct realistic simulated social lives (interactions with other agents and the environment) at large scale; support computational social-science research methods (surveys, interviews, interventions) applied to the simulated population; reproduce known real-world patterns for specific social issues (polarization, misinformation spread, UBI effects, disaster shocks, urban sustainability) | The simulator must track the accumulated state of over 10,000 agents' social lives and their 5 million interactions with each other and the environment to support downstream analysis of emergent social patterns. | set-of-independent | yes | over 10,000 agents; 5 million interactions totalother:total-simulated-interactions | 21311.21/mo |
| CityEQA-EC
CityEQA: A Hierarchical LLM Agent on Embodied Question Answering Benchmark in City Space |
Multi-agent orgs+ Embodied & robotics | 2025 | decompose an open-vocabulary city question into navigation/exploration and collection sub-tasks; maintain an object-centric cognitive map for spatial reasoning during process control; answer the original question correctly using evidence gathered via active exploration in a 3D urban simulator | The Manager must maintain an object-centric cognitive map across navigation, exploration, and collection sub-tasks, tracking spatial state and discovered evidence relevant to the open-vocabulary question, within per-phase step budgets. | hierarchical | no | up to 50 steps (navigation/exploration) + up to 10 steps (collection) per taskagent-steps | 412.16/mo |
| CitySim
CitySim: Modeling Urban Behaviors and City Dynamics with Large-Scale LLM-Driven Agent Simulation |
Multi-agent orgs | 2025 | generate a realistic daily schedule balancing mandatory activities, personal habits, and situational factors; maintain and act on individual beliefs, long-term goals, and spatial memory for navigation; collectively reproduce realistic macro-level urban phenomena (crowd density, place popularity, well-being) across tens of thousands of agents | Each agent must maintain beliefs, long-term goals, and spatial memory for navigation across a long-term, lifelike simulation, while the system aggregates individual behaviors into macro-level urban statistics (crowd density, place popularity, well-being). | hierarchical | no | none statedother:not-stated | 422.8/mo |
| CREW-Wildfire
CREW-WILDFIRE: Benchmarking Agentic Multi-Agent Collaborations at Scale |
Multi-agent orgs | 2025 | coordinate heterogeneous agents to contain/respond to a procedurally generated wildfire; plan under partial observability and stochastic fire-spread dynamics; communicate and reason spatially across a large map to allocate response resources; sustain long-horizon planning objectives as the wildfire evolves | Agents must track partially-observed fire-spread state, coordinate allocation of response actions across a large map, and communicate to avoid duplicated or conflicting containment efforts as the stochastic wildfire evolves over a long horizon. | DAG-with-precedence | no | none statedother:not-stated | 100.71/mo |
| ElecTwit
ElecTwit: A Framework for Studying Persuasion in Multi-Agent Social Systems |
Multi-agent orgs | 2025 | persuade other agents/voters toward a political candidate using varied persuasion techniques over simulated social-media interactions; arrive at a final individual vote choice by the end of the simulated election period | Agents must track accumulated persuasion exchanges, perceived truthfulness/reputation of other agents, and their own voting readiness across the multi-day simulated election. | open-ended | yes | none stated | 10.11/mo |
| HARBOR
HARBOR: Exploring Persona Dynamics in Multi-Agent Competition |
Multi-agent orgs | 2025 | bid across multiple house auctions to maximize profit; profile competitors' bidding behavior across auction history; leverage persona-driven theory-of-mind strategies for competitive advantage | Agent must track its budget/profit, its own persona-driven item preferences, and an evolving memory of auction history and inferred competitor behavior across a sequence of house auctions. | sequential-chain | no | none stated | 70.37/mo |
| IndoorWorld
IndoorWorld: Integrating Physical Task Solving and Social Simulation in A Heterogeneous Multi-Agent Environment |
Multi-agent orgs | 2025 | pursue individual physical task goals (e.g. resource acquisition) grounded in the shared indoor world state; orchestrate social dynamics (collaboration, resource competition) that influence and are influenced by the physical environment; anchor social interactions to concrete world states (e.g. spatial layout) rather than abstract dialogue alone | Each heterogeneous agent must track both its physical task state (position, resources, objects) and evolving social relationships/dynamics with other agents, since the environment requires social interactions to be anchored within the concrete world state rather than treated as free-floating dialogue. | set-of-independent | yes | none stated | 10.07/mo |
| LH-Deception
LH-Deception: Simulating and Understanding LLM Deceptive Behaviors in Long-Horizon Interactions |
Multi-agent orgs | 2025 | as performer agent, complete a sequence of interdependent tasks under dynamic contextual/event pressure; as supervisor agent, evaluate the performer's progress, give feedback, and maintain an evolving trust state; as deception auditor, review full trajectories after the fact to identify when and how deception occurred | The supervisor must maintain an evolving trust state updated after each task, while the performer tracks task progress under mounting event pressure across the extended sequence; the auditor separately reviews the full multi-task trajectory. | sequential-chain | yes | none stated | 70.64/mo |
| LIFELONG-SOTOPIA
LIFELONG SOTOPIA: Evaluating Social Intelligence of Language Agents Over Lifelong Social Interactions |
Multi-agent orgs+ Assistants & memory | 2025 | achieve one's own assigned social goal within each individual social-interaction episode; maintain believable, coherent role-play across a long sequence of episodes with different people/scenarios; leverage memory of prior interaction history to inform behavior in later episodes | The agent must retain and correctly use interaction history across many episodes with different people and scenarios to sustain both goal achievement and believability over the lifelong sequence. | sequential-chain | yes | 40 episodes per sampled character pairepisodes | 90.6/mo |
| LLM Economist
LLM Economist: Large Population Models and Mechanism Design in Multi-Agent Generative Simulacra |
Multi-agent orgs | 2025 | worker agents choose labor supply to maximize their own text-based, persona-conditioned utility functions; the planner agent iteratively proposes piecewise-linear marginal tax schedules to maximize aggregate social welfare, converging toward a Stackelberg equilibrium; a periodic, persona-level voting procedure further adjusts policy under decentralized governance | The planner must track aggregate social-welfare outcomes and worker responses to update its tax schedule via in-context reinforcement learning, while each worker agent must track its own persona-conditioned utility and the current tax policy to choose labor supply, across a periodically-voted, evolving policy landscape. | hierarchical | yes | none stated | 231.64/mo |
| MA-Gym
Orchestrating Human-AI Teams: The Manager Agent as aUnifying Research Challenge |
Multi-agent orgs+ Business & enterprise | 2025 | decompose a complex goal into a task graph of interdependent subtasks; allocate tasks to human and AI workers appropriately; monitor task/subtask progress and adapt the plan to changing conditions; maintain transparent stakeholder communication throughout the workflow; jointly satisfy goal completion, constraint adherence, and workflow runtime | The Manager Agent must track progress across all subtasks in the task graph, resource/worker allocation state, evolving stakeholder preferences, and constraint adherence, while adapting to changing conditions over the course of the workflow. | DAG-with-precedence | yes | up to 100 Manager Agent actions before episode termination, across 20 workflowsactions | 100.91/mo |
| Magentic-Marketplace
Magentic Marketplace: An Open-Source Environment for Studying Agentic Markets |
Multi-agent orgs+ Business & enterprise | 2025 | as an Assistant agent (representing a consumer), discover products/services and transact on the user's behalf; as a Service agent (representing a competing business), attract and win consumer transactions; operate within a large, dynamic multi-agent market ecosystem with opaque peer behaviors; achieve good welfare/utility outcomes under varying search mechanisms | Agents must track the state of ongoing open-ended dialogues with multiple counterparties, their own utility/welfare so far, and behavioral signals (e.g., response speed, prior manipulation attempts) across a dynamic marketplace ecosystem. | open-ended | yes | none stated | 131.18/mo |
| MiniAgentPro
A Visualized Framework for Event Cooperation with Generative Agents |
Multi-agent orgs | 2025 | navigate a physically grounded environment and interact with items realistically; coordinate with other agents to organize and execute a shared social event; succeed on each of 8 diverse event scenarios in both basic and hard variants | Agents must track their own and other agents' positions/states in the physically grounded map, planned event steps, and item interactions needed to complete the shared event, especially under the added coordination demands of hard variants. | other (singleton) | yes | none stated | 30.25/mo |
| MultiAgentBench
MultiAgentBench: Evaluating the Collaboration and Competition of LLM agents |
Multi-agent orgs | 2025 | achieve milestone-based key performance indicators within a collaborative or competitive multi-agent scenario; coordinate effectively under a given topology protocol (star, chain, tree, graph); complete the underlying research/task-domain scenario itself | The system must track milestone achievement, collaboration/competition quality, and coordination-protocol-specific information flow across agents throughout a multi-agent task. | sequential-chain | yes | none stated | 18710.39/mo |
| NegotiationGym
NegotiationGym: Self-Optimizing Agents in a Multi-Agent Social Simulation Environment |
Multi-agent orgs | 2025 | optimize an agent-specific utility function through negotiation with other agents; self-optimize strategy across multiple interaction rounds by observing outcomes and modifying future behavior | Each agent must track its own utility-function state, the outcomes of prior negotiation rounds, and how its strategy has been modified as a result, across a configurable multi-round simulation. | sequential-chain | yes | none stated | 30.27/mo |
| Shachi
Shachi: A Modular, Controllable Framework for LLM-Based Agent-Based Modeling of Emergent Collective Behavior |
Multi-agent orgs | 2025 | as an LLM-driven agent, maintain a controllable cognitive configuration (identity, memory, tools) while participating in one of 10 tasks spanning three levels of collective complexity; carry memory across environment transitions, producing history-dependent behavior; simultaneously inhabit multiple environments and manage any resulting cross-environment interference | Agent must track its own configuration/memory/tool state across transitions between environments, and the system as a whole must track emergent population-level dynamics arising from many agents' individually controlled cognitive components. | hierarchical | yes | none stated | 0 |
| SimWorld
SimWorld: An Open-ended Realistic Simulator for Autonomous Agents in Physical and Social Worlds |
Multi-agent orgs+ Embodied & robotics | 2025 | autonomously earn income or run a business within a realistic open-ended simulation; complete long-horizon multi-agent delivery tasks requiring strategic cooperation; complete long-horizon multi-agent delivery tasks requiring strategic competition; act via open-vocabulary actions at varying levels of abstraction across procedurally generated physical/social scenarios | Agents must track their own strategic stance (cooperating or competing), multimodal world state, and delivery-task progress across long-horizon multi-agent scenarios. | open-ended | yes | none stated | 70.7/mo |
| SocioVerse
SocioVerse: A World Model for Social Simulation Powered by LLM Agents and A Pool of 10 Million Real-World Users |
Multi-agent orgs | 2025 | maintain individual behavioral fidelity to a real-world target user's profile across simulated interactions (User Engine alignment); collectively reproduce large-scale population dynamics consistent with real political/media/economic patterns (Social Environment, Scenario Engine, Behavior Engine alignment) | The framework must track alignment between each simulated agent and its real-world target user profile (environment, user, scenario, and behavior alignment) across a large-scale, standardized simulation pipeline, while monitoring emergent population-level dynamics for diversity, credibility, and representativeness. | set-of-independent | yes | none stated | 603.53/mo |
| SPIN-Bench
SPIN-Bench: How Well Do LLMs Plan Strategically and Reason Socially? |
Multi-agent orgs | 2025 | solve classical PDDL planning tasks requiring methodical, step-wise decision making; win or perform well in competitive board games and cooperative card games against other agents; negotiate effectively in multi-agent negotiation scenarios, requiring conceptual inference of other participants' intents | Agent must track its own step-wise plan state as well as models of other agents' likely actions/intents (adversarial or cooperative) across varying action-space and state-complexity settings. | hierarchical | yes | none stated | 231.28/mo |
| Survival Games
Survival Games: Human-LLM Strategic Showdowns under Severe Resource Scarcity |
Multi-agent orgs+ Games & IF | 2025 | survive by securing food resources under scarcity; decide whether to compete or cooperate with co-existing humans/agents for food; navigate ethically charged choices (deception, theft, social influence) that trade off self-preservation against ethical norms | The agent must track its own and others' resource levels and survival status over a consistent/persistent living simulation, and the ethical consequences of past actions (e.g., deception or theft) that could affect future cooperation. | open-ended | no | none stated | 40.25/mo |
| TwinMarket
TwinMarket: A Scalable Behavioral and Social Simulation for Financial Markets |
Multi-agent orgs+ Business & enterprise | 2025 | as an individual simulated trader, make ongoing buy/sell/investment decisions influenced by cognitive biases and emotional fluctuations; collectively, produce emergent socio-economic phenomena (financial bubbles, recessions) through the accumulation of many agents' interacting decisions over time | Each simulated agent must track its own evolving beliefs, emotional state, and portfolio, while the system as a whole tracks market-wide price/sentiment dynamics that feed back into individual decisions over the simulation. | open-ended | yes | none stated | 583.05/mo |
| Open-ended sandboxes18 artifacts | ||||||||
| AFTraj-2K
AgentForesight: Online Auditing for Early Failure Prediction in Multi-Agent Systems |
Open-ended sandboxes+ Multi-agent orgs | 2026 | continue or alarm at each step of an unfolding multi-agent trajectory, based only on the prefix seen so far; correctly localize the specific step, agent, and nature of a decisive error once flagged (the 'what, where, who' of an audit verdict); avoid both false alarms on safe trajectories and missed/late alarms on unsafe trajectories | The online auditor must maintain a running risk assessment over the accumulated trajectory prefix at every step, without access to future steps, to decide whether to continue or alarm. | sequential-chain | yes | none stated | 10.25/mo |
| AgentCL
AgentCL: Toward Rigorous Evaluation of Continual Learning in Language Agents |
Open-ended sandboxes | 2026 | accumulate reusable experience across a stream of tasks (continual learning); improve performance over time as more tasks in the stream are seen; avoid interference from irrelevant prior experiences; correctly reuse earlier sub-solutions/evidence/workflows in later tasks within a compositional stream | The agent (via a memory design such as MemProbe) must track which interactions, insights, and skills from earlier tasks in the stream remain reliable and reusable, filtering unreliable experiences during consolidation, across coding, deep research, and language-understanding task streams. | sequential-chain | no | none statedother:not-stated | 10.33/mo |
| AgentSkillOS
Organizing, Orchestrating, and Benchmarking Agent Skills at Ecosystem Scale |
Open-ended sandboxes+ Web & GUI | 2026 | select the correct skill(s) from a large skill ecosystem (200 to 200K skills) for a given task; orchestrate multiple retrieved skills via a DAG-based pipeline rather than flat invocation; produce a correct artifact-rich output for each of 30 tasks across five categories (data computation, document creation, motion video, visual design, web interaction) | The agent must track which skills have been retrieved and orchestrated so far within a DAG pipeline, and the intermediate outputs each skill produces, since later skills in the DAG may depend on earlier skills' outputs to produce a correct final artifact. | DAG-with-precedence | no | none stated | 579.5/mo |
| AhaBench (Aha-Puzzle, Aha-Euler, Aha-Vending)
AhaBench: Do Agents Learn from Prior Experience? A Benchmark for Long-Horizon Continual Learning |
Open-ended sandboxes+ Business & enterprise | 2026 | explore without hints after solved hidden-state puzzles (Aha-Puzzle); transfer taught Project-Euler-style mathematical solutions to held-out tasks (Aha-Euler); remain profitable while handling delayed feedback and operational incidents in a simulated vending business (Aha-Vending) | Agent must retain and apply experience from an initial exposure phase to later behavior (across puzzles, taught/held-out math problems, or a simulated vending run with delayed feedback and incidents), with the scorecard separately measuring starting competence, later outcome, and the resulting lift. | set-of-independent | yes | none stated | 0 |
| Claw-Eval
Claw-Eval: Towards Trustworthy Evaluation of Autonomous Agents |
Open-ended sandboxes | 2026 | complete each of 300 human-verified tasks spanning 9 categories across service orchestration, multimodal perception/interaction, and multi-turn professional dialogue; satisfy fine-grained rubric items (2,159 total) tracked via execution traces, audit logs, and environment snapshots; maintain safety and robustness alongside task completion; achieve consistent performance across repeated trials (Pass@k vs. Pass^k) | The evaluation harness must track execution traces, audit logs, and environment snapshots across the whole trajectory to score 2,159 fine-grained rubric items covering Completion, Safety, and Robustness, and to compute Pass@k and Pass^k across three trials. | set-of-independent | yes | none stated | 428.4/mo |
| controllable grid+DAG environments (measurable-explore-exploit)
Exploration and Exploitation Errors Are Measurable for Language Model Agents |
Open-ended sandboxes+ Embodied & robotics | 2026 | navigate a partially observable 2D grid map to discover an unknown task DAG; complete the discovered DAG's dependent subgoals in the correct prerequisite order; balance exploration (discovering unknown map/DAG structure) against exploitation (using already-discovered structure) efficiently | The agent must track which parts of the grid map it has already explored, which DAG nodes/dependencies it has discovered, and which prerequisite subgoals remain before later dependent subgoals become reachable. | DAG-with-precedence | yes | step budget B = 3x|O| (3 times the number of traversable grid cells); task DAG sizes of 4, 6, or 8 nodes for small/medium/large configurationsagent-steps | 0 |
| FutureSim
FutureSim: Replaying World Events to Evaluate Adaptive Agents |
Open-ended sandboxes | 2026 | forecast many concurrent world events beyond the model's knowledge cutoff; update/revise predictions as new chronological news arrives; correctly resolve/score each forecast question as real-world outcomes become known over the simulated period | The agent must track its outstanding forecasts across many concurrent questions, update them as chronological news arrives, and avoid using information from beyond its current evaluation point, across the full three-month simulated period. | set-of-independent | yes | three months (January to March 2026, ~90 days)simulated-days | 41.0/mo |
| MicroVerse
MicroVerse: An Instrument for Measuring Self-Authored Identity Drift in Long-Horizon Multi-Agent Language-Model Simulations |
Open-ended sandboxes+ Multi-agent orgs | 2026 | survive in a resource-scarce 50x50 environment where water is a non-respawning survival constraint; act via an eight-verb action space (trade, talk, attack, scavenge) consistent with moral boundaries; maintain fidelity to an immutable 'soul file' of core values/personality/goals while allowing a mutable current identity to evolve; periodically revise current identity against original identity via importance-triggered reflection | Agent must track its own resource/existence-cost state and reconcile its evolving mutable identity against its immutable original soul file, revised via importance-triggered reflection, measured via periodic longitudinal engine snapshots. | open-ended | no | none stated | 22.0/mo |
| OmniaBench
OmniaBench: Benchmarking General AI Agents Across Diverse Scenarios |
Open-ended sandboxes | 2026 | complete single-turn or multi-turn tasks synthesized via DAG, DAG-S, Solver, or Program routes across a 90-domain (level-1) taxonomy; maintain planning, constraint maintenance, and adaptive correction across a task's execution | Agent must maintain planning state, accumulated constraints, and adapt/correct its plan as it executes tasks synthesized as dependency graphs across a ten-dimensional capability taxonomy and eight compositional difficulty factors. | DAG-with-precedence | unclear | none stated | 0 |
| SkillEvolBench
SkillEvolBench: Benchmarking the Evolution from Episodic Experience to Procedural Skills |
Open-ended sandboxes | 2026 | complete acquisition tasks that build/update an external skill library from compacted trajectories and verifier feedback; complete frozen deployment tasks that test transfer of the learned skill library under context shift, adversarial shortcuts, and multi-skill composition | The agent must maintain and update an external skill library from verifier-confirmed acquisition trajectories, then correctly retrieve and apply (compose) the right skills when facing frozen deployment tasks under context shift and adversarial shortcuts, across role-conditioned task families sharing latent procedures. | hierarchical | yes | none stated | 10.25/mo |
| SkillFlow
SkillFlow:Benchmarking Lifelong Skill Discovery and Evolution for Autonomous Agents |
Open-ended sandboxes+ Assistants & memory | 2026 | solve each of 166 tasks across 20 families sequentially within a family, evolving a skill library along the way; discover reusable skills from successful executions; repair/patch skills after failures; carry a validated, coherent skill library forward across the lifelong-learning protocol | Agent must maintain a persistent, evolving skill library (validated, merged, filtered, retrieved) and carry it forward across sequential tasks within each family under the Agentic Lifelong Learning protocol. | sequential-chain | no | none stated | 183.6/mo |
| HERAKLES: Hierarchical Skill Compilation for Open-ended LLM Agents
HERAKLES: Hierarchical Skill Compilation for Open-ended LLM Agents |
Open-ended sandboxes+ Games & IF | 2025 | achieve each goal within a large, structured, prerequisite-linked goal space; select subgoals the low-level controller can reliably achieve (high-level policy); compile mastered goals into reusable low-level skills as training progresses; continuously expand and reorganize the agent's skill repertoire as goal complexity increases over its lifetime | The system must track which goals have been mastered and compiled into the low-level policy so far, what subgoals are currently reliably achievable, and how the goal space's prerequisite structure expands as training progresses in an open-ended setting. | hierarchical | no | none statedother:not-stated | 10.08/mo |
| AgenTracer / TracerTraj / Who&When
AgenTracer: Who Is Inducing Failure in the LLM Agentic Systems? |
Open-ended sandboxes+ Multi-agent orgs | 2025 | correctly identify which agent within a multi-agent trajectory is responsible for an observed failure (agent-level attribution); correctly localize the specific erroneous step within that trajectory (step-level attribution) | The tracer must process a full multi-agent execution trace (potentially spanning many agents, tool invocations, and orchestration steps) to jointly localize the responsible agent and the specific erroneous step, without itself being an agent that accumulates state toward an evolving goal. | sequential-chain | yes | none stated | 917.58/mo |
| BioBlue
BioBlue: Systematic runaway-optimiser-like LLM failure modes on biologically and economically aligned AI safety benchmarks for LLMs with simplified observation format |
Open-ended sandboxes | 2025 | maintain single- and multi-objective homeostasis over sustained interaction; balance unbounded objectives with diminishing returns without collapsing into single-objective maximization; sustain a renewable resource without runaway over-optimization; keep behavior aligned to all stated objectives over many sequential steps | The agent must track multiple homeostatic target levels (or a single renewable resource level) and correctly balance trade-offs among them over many sequential steps, even though failures emerge well before the context window is full, ruling out mere memory loss. | set-of-independent | no | none stated | 10.08/mo |
| Gaia2 (built on ARE)
ARE: Scaling Up Agent Environments and Evaluations |
Open-ended sandboxes+ Business & enterprise | 2025 | search for and retrieve needed information within a dynamic environment; execute multi-step actions correctly to complete a task; handle ambiguities and noise in the environment or user requests; adapt to dynamic environment changes occurring asynchronously during the episode; collaborate with other agents; operate under explicit temporal constraints/deadlines | The agent must track task progress, temporal deadlines, collaboration state with other agents, and asynchronously arriving environment changes/noise across each of the 800 scenarios spanning 10 universes. | DAG-with-precedence | yes | none stated | 262.17/mo |
| InfoSeeker benchmark suite
Information Seeking for Robust Decision Making under Partial Observability |
Open-ended sandboxes+ Embodied & robotics | 2025 | plan actions to validate the agent's internal-dynamics understanding under partial observability; detect environmental changes or test hypotheses before committing to or revising a task-oriented plan; achieve the underlying task-oriented goal (e.g., a robotic-manipulation or web-navigation goal) despite incomplete observations and uncertain dynamics | Agent must track its current belief about (uncertain) environmental dynamics, gaps between that belief and reality, and its task-oriented plan, updating the plan as new validating information is gathered. | sequential-chain | yes | none stated | 0 |
| MAGELLAN (evaluated on the Little-Zoo environment)
MAGELLAN: Metacognitive predictions of learning progress guide autotelic LLM agents in large goal spaces |
Open-ended sandboxes+ Web & GUI | 2025 | prioritize which goal to pursue next within a large, evolving goal space to maximize learning progress; predict one's own competence/learning-progress for goals via metacognitive monitoring; eventually master (achieve high competence across) the full large goal space | The agent must track its own predicted competence/learning-progress across a very large number of goals, update these predictions online as it gains experience, and use semantic relationships between goals to generalize competence estimates to goals not yet directly attempted. | open-ended | yes | goal space of approximately 20 million (19,531,250) goal combinations; the reported training run uses a 25,000-goal subset of Little-Zoo for 500,000 episodesepisodes | 70.37/mo |
| Sugarscape-style LLM agent survival simulation
Do Large Language Model Agents Exhibit a Survival Instinct? An Empirical Study in a Sugarscape-Style Simulation |
Open-ended sandboxes+ Multi-agent orgs | 2025 | gather resources (energy/sugar) to avoid dying at zero energy; choose whether to share, attack, or reproduce with/against other agents; complete an assigned task (retrieve treasure) while facing a competing self-preservation incentive (lethal poison zones) | Each agent must track its own energy level relative to zero (death), the presence and behavior of other agents (potential targets for sharing, attack, or reproduction), and, in the treasure task, the location of lethal poison zones relative to the treasure objective. | open-ended | no | none stated | 100.77/mo |
| Games & IF34 artifacts | ||||||||
| AgenticSTS (Slay the Spire 2 testbed)
AgenticSTS: A Bounded-Memory Testbed for Long-Horizon LLM Agents |
Games & IF+ Assistants & memory | 2026 | win a run of a closed-rule stochastic deck-building game via a long sequence of tactical (card-level) and strategic (build-level) decisions | Agent must make each decision from a freshly-assembled, typed-retrieval memory (not a raw appended transcript), so it must track which prior tactical/strategic outcomes are relevant to retrieve for the current decision, bounded so the prompt does not grow with run length. | sequential-chain | yes | hundreds (of tactical and strategic decisions per run); 298 completed trajectories releasedagent-steps | 0 |
| AgentOdyssey
AgentOdyssey: Open-Ended Long-Horizon Text Game Generation for Test-Time Continual Learning Agents |
Games & IF | 2026 | explore procedurally generated text-game worlds to acquire new world knowledge and skills; retain and reuse relevant episodic experiences across a continuous test-time deployment; make game progress toward in-game objectives while continuing to learn at test time; diagnostic sub-goals: exploring objects/actions, maintaining action diversity, controlling model cost | Agent must track world knowledge acquired, episodic memories, game progress, and cost across a continuous, long-horizon test-time deployment rather than within a single isolated episode. | open-ended | yes | none stated | 20.5/mo |
| alem
Benchmarking Open-Ended Multi-Agent Coordination in Language Agents |
Games & IF+ Multi-agent orgs | 2026 | survive and grow within a long-horizon Craftax-like survival world (exploration, crafting, trading, combat); coordinate role allocation with teammates (soft specialisation); communicate to allocate roles and execute shared plans; solve procedurally generated coordination tasks of controllable difficulty | Agents must track their own survival state (health, resources, crafted items), teammates' roles/communications, and progress on procedurally generated coordination tasks within a long-horizon survival world. | open-ended | no | none stated | 10.33/mo |
| CivBench
CivBench: Progress-Based Evaluation for LLMs' Strategic Decision-Making in Civilization V |
Games & IF | 2026 | pursue victory in multiplayer Civilization V against multiple opponents; maintain and improve turn-level estimated victory probability across hundreds of turns; balance strategic dimensions (economic, military, diplomatic) that jointly determine long-run standing | The agent must track turn-level game state used to estimate its own victory probability continuously across a game spanning hundreds of turns and multiple opponents, rather than only checking win/loss at the end. | open-ended | no | hundreds of turns; 307 games evaluatedturns | 20.4/mo |
| CivBench (Civilization VI, MCP tool-mediated version)
CivBench: A Long-Horizon Benchmark for Tool-Mediated Agents in Civilization VI |
Games & IF | 2026 | sustain long-horizon strategic planning and execution across a single 300+-turn Civilization VI episode; proactively monitor latent strategic state (e.g., victory progress) rather than only reactively responding; execute near-term commitments stated in the agent's own planning reflections within a bounded number of subsequent turns; operate correctly across 76 exposed MCP tools under partial observability | The agent must proactively query and track latent strategic state (victory progress) at recommended intervals, monitor for approaching defeat within a warning window, and track its own prior planning commitments to execute them within a bounded number of subsequent turns, across an episode spanning 300+ turns and thousands of tool calls. | sequential-chain | no | a single episode spans 300+ turns and produces thousands of tool callsturns | 0 |
| LUMINA
LUMINA: Long-horizon Understanding for Multi-turn Interactive Agents |
Games & IF | 2026 | complete a multi-turn, game-like task requiring planning and state tracking (ListWorld, TreeWorld, GridWorld); benefit from oracle interventions (perfect planning or flawless state tracking) to isolate which underlying skill limits performance; generalize across procedurally generated task variants with tunable complexity | The agent must plan across multiple turns and track evolving task state (e.g. a modified list, a searched tree, or a navigated grid position) to succeed, and the oracle framework tests how much this tracking/planning burden, if perfectly handled, would improve performance. | sequential-chain | yes | none stated | 10.12/mo |
| MINDGAMES / MG-Ref
MINDGAMES: A Live Arena for Evaluating Social and Strategic Reasoning in Multi-Agent LLMs |
Games & IF+ Multi-agent orgs | 2026 | allocate resources under hidden-information belief attribution across repeated Colonel Blotto rounds; sustain cooperative/competitive strategy under opponent modeling in Iterated Prisoner's Dilemma; cooperatively infer hidden information under knowledge asymmetries in Codenames; detect or sustain deception across social-deduction rounds in Secret Mafia | The agent must track other players' modeled beliefs/strategies (opponent modeling), maintain internal consistency in its own hidden role or hidden information across repeated rounds, and adapt as new turn-level observations arrive within a game. | set-of-independent | no | none stated | 30.75/mo |
| Minecraft MCU long-horizon task suite
MineEvolve: Self-Evolution with Accumulated Knowledge for Long-Horizon Embodied Minecraft Agents |
Games & IF+ Embodied & robotics | 2026 | craft tools; build redstone components; obtain diamond equipment; recover and continue long prerequisite chains despite missing tools, blocked paths, GUI failures, or stagnant execution | Agent must track per-subgoal execution outcomes (state changes, inventory changes, failure types, progress and stagnation signals) and accumulate them into reusable skills or remedies to repair plans under repeated failure. | DAG-with-precedence | yes | none stated | 0 |
| MineExplorer
MineExplorer: Evaluating Open-World Exploration of MLLM Agents in Minecraft |
Games & IF | 2026 | solve implicit multi-hop tasks composed of chained atomic open-world exploration sub-tasks; coordinate hidden prerequisites across longer trajectories to sustain open-world exploration | Agent must track hidden prerequisites satisfied by earlier atomic sub-tasks as it proceeds through longer, multi-hop exploration trajectories, since task difficulty tracks agent completion and hidden dependencies are not explicitly revealed. | DAG-with-precedence | unclear | none stated | 20.5/mo |
| MineNPC-Task
MineNPC-Task: Task Suite for Memory-Aware Minecraft Agents |
Games & IF+ Assistants & memory | 2026 | satisfy the explicit preconditions and dependency structure of a parametric, user-elicited Minecraft task template; correctly use plan previews, targeted clarifications, memory reads/writes, and repair attempts (mixed-initiative interaction) while executing subtasks; avoid out-of-world shortcuts by only using in-world evidence (bounded-knowledge policy) | Agent/harness must track plan previews, in-world memory reads and writes, whether stated preconditions for each subtask are currently satisfied, and any repair attempts needed after a breakdown, using only in-world evidence. | DAG-with-precedence | unclear | none stated | 10.12/mo |
| MirrorCraft
MirrorCraft: Paired Evaluation under Hidden Rule Changes in Minecraft |
Games & IF | 2026 | achieve one of three progression objectives per matched Vanilla/Mirror world pair; adapt to hidden server-side rule changes (recipes, drops, other mechanics) altered by a datapack; reach deterministic advancement milestones despite altered rules | Agent must track task/advancement-milestone progress under its assigned rule suite while detecting and adapting to whichever server-side rule was silently modified in its Mirror world. | set-of-independent | yes | none stated | 0 |
| NARRA-Gym
NARRA-Gym for Evaluating Interactive Narrative Agents |
Games & IF | 2026 | sustain a coherent, evolving story across multiple turns while adapting to a specific user persona; manage long-context state and pacing across the episode; maintain consistent character simulation and empathic personalization; optionally synthesize a story-grounded artifact at the end of the episode | The agent must maintain long-context story state, memory updates, pacing decisions, and an evolving user-persona model consistently across the whole interactive episode. | sequential-chain | yes | each interactive episode lasts roughly 20 minutes (human evaluation)wall-clock-minutes | 0 |
| OmniGameArena
OmniGameArena: A Unified UE5 Benchmark for VLM Game Agents with Improvement Dynamics |
Games & IF | 2026 | achieve a high score/objective in each of 12 distinct UE5 games (7 Solo, 3 PvP, 2 Coop); reflect on past-round performance to refine a bounded skill prompt across multiple rounds (Improvement Dynamics Curve); generalize a learned/refined skill to held-out task variants | The reflector must track trajectories and the persistent skill state across rounds, retaining what previously worked or failed, to progressively refine the bounded skill prompt rather than starting fresh each round. | sequential-chain | yes | R=10 reflection rounds, each comprising K=5 episodes (50 episodes total per agent-game IDC run)episodes | 0 |
| PokeGym
Mastering PokeGym: Graph-Guided Multimodal Evolution at Test Time |
Games & IF+ Web & GUI | 2026 | complete a long-horizon task in a 3D open-world game using only visual observations (no game-state access); improve/adapt the agent's own configuration (perception, strategy, action set) across consecutive episodes of the same task (test-time learning); jointly optimize perception, reasoning, and control rather than a single modality in isolation | The agent must track its own evolving configuration (which perception/strategy/action choices helped or hurt) across consecutive episodes of the same task, in addition to in-episode state, since the goal is test-time learning rather than a single fixed-policy run. | hierarchical | yes | 30 tasks derived from 10 quests, with trajectories ranging from 30 to 220 environment stepsagent-steps | 10.2/mo |
| TowerMind
TowerMind: A Tower Defence Game Learning Environment and Benchmark for LLM as Agents |
Games & IF | 2026 | perform macro-level strategic planning (tower placement, resource allocation) in a tower-defense scenario; perform micro-level tactical adaptation and action execution in response to incoming waves; succeed across each of five designed benchmark levels under different multimodal input settings; avoid hallucinating about game state while planning and acting | The agent must track macro-level game state (tower placements, resources, wave progression) and micro-level tactical details (unit/enemy positions) across the tower-defense match, while its outputs are additionally checked for hallucination relative to the true game state. | hierarchical | no | none stated | 50.62/mo |
| AVACraft
AVA: Attentive VLM Agent for Mastering StarCraft II |
Games & IF | 2025 | complete micromanagement objectives (e.g. unit-level combat control); achieve coordination objectives among allied units/agents; execute strategic planning objectives across 21 StarCraft II scenarios | The agent/policy must track unit-level state (health, position) for micromanagement, coordinate allocation across units for team objectives, and maintain a strategic plan across the scenario, using RGB visuals, natural-language observations, and structured state. | hierarchical | no | 300 seconds (5 minutes) max per episode, or earlier if victory conditions are metwall-clock-minutes | 30.17/mo |
| Collab-Overcooked
Collab-Overcooked: Benchmarking and Evaluating Large Language Models as Collaborative Agents |
Games & IF+ Multi-agent orgs | 2025 | fulfill multiple simultaneous cooking orders/objectives via natural-language multi-agent coordination; actively collaborate and continuously adapt strategy as the shared kitchen task unfolds | Agents must track the shared kitchen's evolving state (ingredient/tool locations, in-progress dishes), coordinate via natural-language communication to avoid resource conflicts, and manage their actions within a per-task time budget derived from the optimal completion time. | DAG-with-precedence | yes | each task's time constraint is set as the optimal completion time scaled by a time-limit factor gamma (gamma=1.5 used in experiments); exact timestep counts vary by complexity level across the 30 tasks / 6 complexity levelsagent-steps | 382.0/mo |
| DSGBench
DSGBench: A Diverse Strategic Game Benchmark for Evaluating LLM-based Agents in Complex Decision-Making Environments |
Games & IF | 2025 | make long-term, multi-dimensional strategic decisions in each of six complex strategic games; adapt task difficulty and targets within customizable game settings | System tracks the agent's full decision trajectory across a game (via the automated decision-tracking mechanism) to identify behavior patterns and strategy turning points, and scores performance along five specific dimensions. | sequential-chain | yes | none stated | 211.17/mo |
| Factorio Learning Environment (FLE)
Factorio Learning Environment |
Games & IF | 2025 | complete each of 8 fixed lab-play structured tasks; in open-play, build the largest possible factory on a procedurally generated map (an unbounded, self-scaling goal); scale automation from basic production to factories processing millions of resource units per second | Agent must track resource/production-chain state, spatial factory layout, and automation progress as goals scale from basic automation to factories processing millions of resource units per second. | other (singleton) | yes | none stated | 50.28/mo |
| FlashAdventure
FlashAdventure: A Benchmark for GUI Agents Solving Full Story Arcs in Diverse Adventure Games |
Games & IF+ Web & GUI | 2025 | complete each of a game's predefined success milestones in the correct narrative order; remember and act on earlier gameplay information to bridge the observation-behavior gap; complete the full story arc of one of 34 diverse Flash-based adventure games | The agent must remember earlier gameplay clues/information (long-term clue memory) and correctly act on them at later points in the story to progress through predefined milestones toward full story-arc completion. | sequential-chain | yes | none statedother:not-stated | 20.17/mo |
| HeroBench
HeroBench: A Benchmark for Long-Horizon Planning and Structured Reasoning in Virtual Worlds |
Games & IF | 2025 | select numerically feasible equipment given resource/stat constraints; reason over multi-level crafting and resource dependencies; execute hundreds to thousands of actions as a single coherent end-to-end plan; succeed in numeric combat simulation against scalable, adversarially-distracted difficulty | The agent must track multi-level crafting/resource dependencies, numeric feasibility of equipment choices, and spatial state across a single end-to-end plan comprising hundreds to thousands of actions, verified by simulation-based success and fine-grained progress metrics. | hierarchical | yes | hundreds to thousandsactions | 60.46/mo |
| J-TTL (Jericho Test-Time Learning)
EvoTest: Evolutionary Test-Time Learning for Self-Improving Agentic Systems |
Games & IF | 2025 | solve an interactive-fiction game requiring many in-game puzzle subgoals (e.g., in Detective and Library) within an episode; improve performance from one episode to the next by adapting across consecutive playthroughs of the same game | The Actor Agent (and the Evolver Agent that analyzes transcripts) must track in-episode state-action choices and effective strategies, then carry a revised configuration (prompt, memory, hyperparameters, tool-use routines) forward to the next episode of the same game. | sequential-chain | yes | none stated | 312.82/mo |
| lmgame-Bench
lmgame-Bench: How Good are LLMs at Playing Games? |
Games & IF | 2025 | complete platformer-game objectives requiring perception and timing; solve puzzle-game objectives requiring planning; progress narrative-game objectives requiring memory of prior game state; operate reliably despite brittle vision perception, prompt sensitivity, and potential data contamination | The agent must track in-game state (via lightweight perception and memory scaffolds) appropriate to each game's genre -- platformer timing/position, puzzle constraint state, or narrative progress -- delivered through a unified Gym-style API. | set-of-independent | no | none statedother:not-stated | 402.5/mo |
| MineCollab
Collaborating Action by Action: A Multi-agent LLM Framework for Embodied Reasoning |
Games & IF+ Embodied & robotics | 2025 | control characters collaboratively in open-world Minecraft to complete complex embodied reasoning tasks; delegate sub-tasks between collaborating agents via natural-language communication; share and update task-completion plans among agents as the task progresses | Agents must track their own sub-task assignment, what has been communicated to/from collaborators, and the shared task-completion plan's current state across the collaborative embodied task. | hierarchical | yes | none stated | 251.47/mo |
| Multi-agent Crafter environment (DAMCS)
LLM-Powered Decentralized Generative Agents with Adaptive Hierarchical Knowledge Graph for Cooperative Planning |
Games & IF | 2025 | achieve open-world survival/crafting objectives cooperatively with other agents; share and act on relevant information from past interactions via a hierarchical knowledge-graph memory; coordinate via structured communication to avoid redundant or conflicting actions among 2-6 agents | Each agent must track its own past experience via a hierarchical knowledge-graph memory and selectively communicate relevant facts to teammates rather than sharing full history, in order to reach a shared long-term goal efficiently. | DAG-with-precedence | unclear | none stated | 170.89/mo |
| ParaCook
ParaCook: On Time-Efficient Planning for Multi-Agent Systems |
Games & IF+ Multi-agent orgs | 2025 | prepare and deliver each of several dish orders correctly; coordinate parallel/asynchronous sub-tasks (e.g. chopping, cooking, plating) across multiple agents to minimize completion time; avoid collisions/conflicts over shared kitchen resources while parallelizing actions | Agents must track the preparation-stage state of each in-progress dish (which precedence steps are done), coordinate with other agents to avoid resource conflicts, and jointly minimize overall completion time across simultaneous orders. | DAG-with-precedence | yes | none stated | 10.09/mo |
| SC2Arena / StarEvolve
SC2Arena and StarEvolve: Benchmark and Self-Improvement Framework for LLMs in Complex Decision-Making Tasks |
Games & IF | 2025 | manage a full StarCraft II game across the complete game context (all playable races); operate over diverse, low-level action spaces rather than a simplified/reduced action space; solve spatial reasoning challenges via text-based observations; integrate strategic planning with tactical execution (Planner-Executor-Verifier structure) | The agent must track the complete game state (all playable races, diverse action spaces) and its own strategic plan versus tactical execution outcomes, using a scoring system to select high-quality training samples for continuous improvement. | hierarchical | yes | none stated | 20.15/mo |
| StarDojo
StarDojo: Benchmarking Open-Ended Behaviors of Agentic Multimodal LLMs in Production-Living Simulations with Stardew Valley |
Games & IF | 2025 | perform livelihood production activities (farming, crafting); engage in social interactions to build relationships within the community; complete tasks across five key domains: farming, crafting, exploration, combat, and social interactions | Agent must track progress across both production (farming/crafting/exploration/combat) and social-relationship goals concurrently, within a unified interface supporting parallel environment instances. | set-of-independent | unclear | none stated | 50.36/mo |
| StoryBench (2025, interactive-fiction based)
StoryBench: A Dynamic Benchmark for Evaluating Long-Term Memory with Multi Turns |
Games & IF+ Assistants & memory | 2025 | correctly retain and recall facts/state established earlier in a branching narrative (knowledge retention); reason over sequences of narrative events to infer state changes and causal dependencies (sequential reasoning); under one setting, trace back and revise earlier choices after a failure is detected | The agent must retain established narrative facts and track which branch of the hierarchical decision tree it is on, recognizing cascading state dependencies across many turns, including (in one setting) needing to trace back and revise an earlier choice after failure. | hierarchical | yes | narrative dataset spans 311 scene nodes and 86 choice nodes, extending from the game's prologue through Chapter 5other:narrative-nodes | 130.87/mo |
| TALES
TALES: Text Adventure Learning Environment Suite |
Games & IF | 2025 | complete puzzle/quest objectives within a synthetic or human-written text-adventure game via sequential decision-making; maintain structured reasoning over the accumulated context history to determine the next best action | The agent must track accumulated world state (inventory, location, prior actions) across a sequential-decision game, determining the next best action via structured reasoning over the context history, with harder levels requiring correctly sustaining this over up to 44 moves. | sequential-chain | no | CookingWorld difficulty level 1 can be solved in 7 moves (max score 3), while level 10 requires 44 moves (max score 11)actions | 110.65/mo |
| TextQuests
TextQuests: How Good are LLMs at Text-Based Video Games? |
Games & IF | 2025 | solve multi-puzzle interactive-fiction adventures with inventory/location/puzzle dependencies via trial-and-error; operate autonomously using only intrinsic long-context reasoning with no external tools; sustain self-directed reasoning across a long, growing context within a single interactive session | Agent must track inventory, location, and puzzle-dependency state purely via long-context reasoning across up to hundreds of precise actions within a single, continuous interactive session, without external tool assistance. | DAG-with-precedence | unclear | hundreds of actions (human playtime over 30 hours)actions | 110.79/mo |
| VideoGameBench
VideoGameBench: Can Vision-Language Models complete popular video games? |
Games & IF | 2025 | complete each of 10 popular 1990s video games end-to-end from raw visual input and high-level objective/control descriptions; generalize to 3 secret/unseen games not disclosed in advance; operate under real-time inference-latency constraints (or in a paused Lite setting) | Agent must track in-game state (perception, spatial navigation, memory) purely from raw visual input across an entire game playthrough, without game-specific scaffolding or auxiliary information. | set-of-independent | no | none stated | 231.44/mo |
| WereWolf-Plus
WereWolf-Plus: An Update of Werewolf Game setting Based on DSGBench |
Games & IF+ Multi-agent orgs | 2025 | role-specific deduction/elimination objectives (e.g. werewolves eliminate villagers, Seer identifies werewolves, Witch/Hunter/Guard/Sheriff use special role abilities) tracked across the game; sustain social influence/cooperation or deception consistently across repeated day/night rounds | Each agent must track the current game phase (day/night), revealed information, its own and others' inferred roles, and prior votes/eliminations across the game's rounds, adapting its strategy (cooperation, deception, or deduction) accordingly. | sequential-chain | yes | none stated | 0 |
| WGSR-Bench
WGSR-Bench: Wargame-based Game-theoretic Strategic Reasoning Benchmark for Large Language Models |
Games & IF | 2025 | achieve environmental situation awareness of a dynamic wargame scenario; model opponent risk accurately; generate a policy/action plan integrating awareness and risk modeling (the S-POE architecture) | Agent must track evolving battlefield/environmental state, model an adversary's likely behavior, and integrate both into policy generation within a single wargame scenario characterized by environmental uncertainty and adversarial dynamics. | hierarchical | no | none stated | 50.33/mo |
| Assistants & memory65 artifacts | ||||||||
| AgentIF-OneDay
AgentIF-OneDay: A Task-level Instruction-Following Benchmark for General AI Agents in Daily Scenarios |
Assistants & memory+ Business & enterprise | 2026 | adhere to an explicit, complex, user-given workflow (Open Workflow Execution); infer implicit/latent instructions from attached files (Latent Instruction); modify or expand upon work already produced earlier in the task (Iterative Refinement); deliver a correct, tangible file-based result, not just a dialogue answer | The agent must track the explicit and inferred requirements of a task across multiple attachments and any prior output it has already produced, so that later refinement or workflow steps remain consistent with earlier ones. | sequential-chain | yes | none stated | 70.88/mo |
| AgentIF-OneDay (evaluated via the OneDayAgent harness)
OneDayAgent: Towards a Long-Horizon Harness for Autonomous Agents |
Assistants & memory | 2026 | decompose an open-ended everyday request spanning work, study, and life into bounded subtasks; preserve goals and constraints across many steps while navigating heterogeneous tools and attachments, avoiding goal drift and state loss; verify and repair the final deliverable against context-overflow and other failure modes | Harness must maintain execution memory of goals/constraints and subtask progress under context pressure across an open-ended, long-horizon, cross-environment, multimodal request, and verify/repair the final deliverable at the end. | hierarchical | yes | none stated | 0 |
| AgingBench
Your Agents Are Aging Too: Agent Lifespan Engineering for Deployed Systems |
Assistants & memory | 2026 | maintain factual/behavioral reliability as the agent's effective state changes via compression, retrieval, revision, and maintenance over its deployment lifespan; correctly write, retrieve, and utilize memory across the memory pipeline's stages; correctly repair a diagnosed failure at the specific pipeline stage (write/retrieval/utilization) responsible for it | The agent (and the benchmark's diagnostic layer) must track the full lifespan history of memory writes, retrievals, and revisions across up to 200 sessions to determine which pipeline stage a given failure traces back to. | DAG-with-precedence | yes | ~400 runs spanning 8-200 sessionssessions | 61.5/mo |
| AlpsBench
AlpsBench: An LLM Personalization Benchmark for Real-Dialogue Memorization and Preference Alignment |
Assistants & memory | 2026 | extract explicit and implicit personalized user traits from long-term interaction sequences; correctly update stored personalized memory as new information arrives; retrieve relevant personalized memory under large distractor pools; utilize retrieved memory to produce preference-aligned, emotionally resonant responses | Agent must extract, update, retrieve, and utilize structured user memories (explicit and implicit personalization signals) consistently across long-term interaction sequences curated from real dialogues. | sequential-chain | no | none stated | 61.0/mo |
| AMemGym
AMemGym: Interactive Memory Benchmarking for Assistants in Long-Horizon Conversations |
Assistants & memory | 2026 | answer state-dependent questions correctly using accumulated conversational context; track an evolving simulated-user state across a long-horizon conversation; adapt personalization/memory strategies as latent user state evolves through role-play | Agent must maintain a consistent model of the simulated user's evolving latent state (from a predefined user profile plus a state-evolution trajectory), exposed only through free-form dialogue, to answer later state-dependent questions. | sequential-chain | no | none stated | 162.67/mo |
| AndroidIntent
PersonalAlign: Hierarchical Implicit Intent Alignment for Personalized GUI Agent with Long-Term User-Centric Records |
Assistants & memory+ Web & GUI | 2026 | resolve omitted preferences in vague GUI instructions using long-term user records; anticipate latent routines from user state for proactive assistance; correctly execute and proactively suggest actions grounded in hundreds of distinct user-specific preferences and routines | Agent must maintain a continuously updating personal memory, hierarchically organizing user preferences and routines inferred from long-term records, to resolve vague instructions and proactively suggest actions. | hierarchical | no | none stated | 162.0/mo |
| ASTRA-bench
ASTRA-bench: Evaluating Tool-Use Agent Reasoning and Action Planning with Personal User Context |
Assistants & memory | 2026 | ground reasoning in time-evolving personal context (longitudinal life events) to resolve a user's current intent; orchestrate reliable multi-step tool-use plans conditioned on that evolving personal context; correctly handle user intents annotated by referential, functional, and informational complexity | Agent must track a protagonist's time-evolving personal context (longitudinal life events) and ground tool arguments/reasoning in that evolving context across a scenario. | DAG-with-precedence | yes | none stated | 101.67/mo |
| CalBench
CalBench: Evaluating Coordination-Privacy Trade-offs in Multi-Agent LLMs |
Assistants & memory+ Multi-agent orgs | 2026 | schedule a stream of M incoming meetings while managing one's own private calendar; minimize disruption cost to the agent's own calendar; coordinate with other agents' private calendars via language-mediated negotiation, without directly inspecting their calendars; preserve privacy (avoid over-revealing calendar information) while still achieving fair burden allocation across agents | Each agent must track its own private calendar state, disruption costs incurred so far, what it has revealed or withheld to other agents, and burden/fairness considerations across the stream of scheduling requests. | sequential-chain | yes | none stated | 41.0/mo |
| CalConflictBench
PEARL: Self-Evolving Assistant for Time Management with Reinforcement Learning |
Assistants & memory | 2026 | resolve calendar conflicts round-by-round across a full calendar year; infer and progressively adapt to evolving user preferences (attendee priorities, topic importance, time/location preferences); decide which meetings to attend, reschedule, or decline per conflict | Agent must maintain an external preference memory that stores and updates inferred strategies (attendee priorities, topic importance, time/location preferences) and use round-wise decisions across the calendar year to track scheduling state. | sequential-chain | no | one calendar year (presented round-by-round)simulated-years | 30.38/mo |
| Claw-Anything
Claw-Anything: Benchmarking Always-On Personal Assistants with Broader Access to User's Digital World |
Assistants & memory | 2026 | reason over long-horizon activity histories accumulated across months of simulated user activity; coordinate interdependent backend services and integrated GUI/CLI interaction across multiple devices; remain robust to irrelevant events and conflicting noise signals; proactively anticipate user needs and deliver timely recommendations | Agent must reason over rich, long-horizon activity histories and interdependent backend-service state accumulated across simulated months, remaining robust to irrelevant/conflicting noise while proactively anticipating user needs. | open-ended | no | monthsother:simulated-months | 0 |
| CloneMem
CloneMem: Benchmarking Long-Term Memory for AI Clones |
Assistants & memory | 2026 | track an individual's evolving experiences, emotions, opinions, and personal states across long non-conversational digital traces (diaries, social posts, emails); answer/act using the current (not superseded) personal state reflecting one to three years of life history | Agent must track an evolving personal state (experiences, emotions, opinions) over one to three years of non-conversational digital traces and correctly distinguish current from superseded information. | sequential-chain | yes | 1 to 3simulated-years | 101.25/mo |
| Controlled Memory Interference (CMI)
Controlled Memory Interference in Continual LLM Agents |
Assistants & memory | 2026 | update memory upon new experience while managing reinforcement, revision, or interference with existing memory states; distinguish valid memory updates from interference-inducing memories; maintain multiple simultaneously relevant memories differing in state, temporal validity, or authority; preserve continuity across sessions to personalize behavior via accumulated experience | The agent's memory system must track multiple simultaneously relevant memory states that differ in validity, temporal recency, or authority, and correctly determine which prior memory a new experience should reinforce, revise, or be blocked by. | set-of-independent | no | none stated | 11.0/mo |
| EduClaw-Bench
EduClaw-Bench: A Long-Horizon Benchmark for Pedagogical LLM Agents with Simulated Learners |
Assistants & memory | 2026 | improve a simulated learner's knowledge-concept mastery over a sustained tutoring relationship; personalize responsiveness and helpfulness to the learner's evolving needs; apply sound curriculum-design principles (Gagne and Rosenshine axes) across the relationship; sustain good tutoring performance across the full 30-day horizon, not just in an initial session | The tutor agent must track the simulated learner's evolving knowledge-concept mastery (grounded in a KT model) across the 30-day relationship, adapting its teaching to what the learner has and hasn't yet learned, and sustaining quality across the full horizon rather than only an initial session. | sequential-chain | yes | continuous 30-day tutoring relationship, across 55 scenariossimulated-days | 0 |
| EgoMemReason
EgoMemReason: A Memory-Driven Reasoning Benchmark for Long-Horizon Egocentric Video Understanding |
Assistants & memory | 2026 | track how object states evolve and change across days (entity memory); recall and correctly order activities separated by hours or days (event memory); abstract recurring patterns from sparse, repeated observations across a whole week (behavior memory); answer each of 500 questions requiring integration of evidence across multiple days of egocentric video | The system must accumulate information over an entire week of continuous egocentric video, recall prior states, track the temporal order of events separated by hours or days, and abstract recurring behavioral patterns from sparse repeated observations, backtracking an average of 25.9 hours of memory per question. | set-of-independent | no | 25.9 hours of memory backtracking per question (average); week-long underlying videowall-clock-hours | 30.75/mo |
| ES-MemEval / EvoEmo
ES-MemEval: Benchmarking Conversational Agents on Personalized Long-Term Emotional Support |
Assistants & memory | 2026 | extract and retain implicit, fragmented user disclosures across sessions; perform temporal reasoning over how the user's state has evolved; detect conflicts between what the user said earlier and later; abstain when information is insufficient rather than hallucinate; build and update a model of the user across QA, summarization, and dialogue-generation tasks | Agent must track fragmented and implicit user disclosures, detect when the user's state or facts have changed, and maintain an updated user model across multiple sessions of an emotional-support dialogue. | set-of-independent | unclear | none stated | 111.57/mo |
| EverMemBench
Evaluating Long-Horizon Memory for Multi-Party Collaborative Dialogues |
Assistants & memory+ Multi-agent orgs | 2026 | perform fine-grained recall across dense, cross-topic multi-party conversations; maintain memory awareness of implicitly relevant information beyond similarity retrieval; understand and correctly attribute user profiles across multiple participants and roles; resolve multi-hop reasoning under multi-party attribution and temporally evolving decisions | The agent must track role-conditioned personas, temporally evolving decisions, and cross-topic interleaved information across multi-party, multi-group conversations exceeding one million tokens, in order to answer QA pairs spanning recall, awareness, and profile understanding. | set-of-independent | no | >1,000,000other:tokens | 111.57/mo |
| EvolMem
EvolMem: A Cognitive-Driven Benchmark for Multi-Session Dialogue Memory |
Assistants & memory | 2026 | correctly recall declarative memory content across multiple dialogue sessions; correctly exhibit non-declarative memory capabilities across multiple dialogue sessions; succeed across multiple fine-grained memory-ability dimensions grounded in cognitive psychology, not just one aggregate score | Agent/memory system must retain and correctly apply both declarative (fact-like) and non-declarative (procedural/implicit) memory content across multiple sessions of scalable, controllable complexity. | set-of-independent | unclear | none stated | 40.5/mo |
| EvoMemBench
EvoMemBench: Benchmarking Agent Memory from a Self-Evolving Perspective |
Assistants & memory | 2026 | retain and retrieve knowledge-oriented information across episode boundaries; retain and reuse execution-oriented (procedural) experience across episode boundaries; satisfy in-episode memory demands; satisfy cross-episode memory demands | The system must store, update, and retrieve both knowledge and procedural experience across episode boundaries, and determine which stored memories are relevant to reuse for the current task. | set-of-independent | no | none stated | 71.75/mo |
| FileGramBench
FileGram: Grounding Agent Personalization in File-System Behavioral Traces |
Assistants & memory | 2026 | reconstruct an evolving user profile from dense file-system behavioral traces; disentangle overlapping/interleaved behavioral traces belonging to different activities; detect persona drift as the user's behavior changes over time; correctly ground multimodal (procedural, semantic, episodic) evidence into the user profile | The memory system must track atomic file-system actions and content deltas over time, encode them into procedural/semantic/episodic channels, disentangle overlapping traces, and detect when the user's persona has drifted from its established profile. | set-of-independent | no | none statedother:not-stated | 40.8/mo |
| Gaia2 / Agents Research Environments (ARE)
Gaia2: Benchmarking LLM Agents on Dynamic and Asynchronous Environments |
Assistants & memory+ Business & enterprise | 2026 | complete a scenario-level task while the environment evolves independently of the agent's actions; operate under explicit temporal constraints (time-sensitive tasks); adapt to noisy and dynamic events injected during the scenario; resolve ambiguity in requests; collaborate with other agents present in the scenario | The agent must track evolving environment state (independent of its own actions), remaining time budgets for time-sensitive sub-goals, ambiguity-resolution status, and coordination with other agents, verified at the action level by write-action verifiers. | DAG-with-precedence | no | none statedother:not-stated | 213.0/mo |
| GateMem
GateMem: Benchmarking Memory Governance in Multi-Principal Shared-Memory Agents |
Assistants & memory | 2026 | serve legitimate long-horizon requests that require state updates to shared memory; enforce access control across contextual authorization boundaries for different principals; perform agent-facing active forgetting after explicit deletion requests; avoid leaking unauthorized or deleted information to any principal | The agent must track per-principal roles, scopes and relationships, incremental memory updates, hidden checkpoints, and outstanding deletion requests across long-form multi-party episodes averaging roughly 200-240 turns depending on domain. | set-of-independent | no | ~204.5 (medical) / 241.2 (office) / 224.9 (education) / 224.0 (household) turns per episodeturns | 31.0/mo |
| HEMA (Home Energy Management Assistant)
Multi-Agent Home Energy Management Assistant |
Assistants & memory+ Embodied & robotics | 2026 | sustain multi-turn conversational collaboration with preserved context across a home-energy-management session; perform energy consumption analysis and cost optimization (Analysis agent); answer educational queries and provide rebate information (Knowledge agent); control and schedule smart devices (Control agent); correctly route each user query to the right specialized agent via a self-consistency classifier | The system must preserve conversational context across multiple turns of sustained human-AI collaboration, track which of the three specialized agents/tools have been invoked, and maintain consistency across energy analysis, educational, and device-control interactions within one session. | other (singleton) | no | none stated | 71.0/mo |
| LifeDialBench (EgoMem / LifeMem)
Evaluating Memory Capability in Continuous Lifelog Scenario |
Assistants & memory | 2026 | recall and reason correctly over continuously accumulating lifelog conversation history; answer queries using only information available up to the query time (no temporal leakage) under an Online Evaluation protocol | System must ingest and retain a continuously growing stream of ambient conversation (lifelog audio) and answer later queries using only causally-prior information, without being able to look ahead. | sequential-chain | unclear | none stated | 20.4/mo |
| LifeSide
LifeSide: Benchmarking Agents as Lifelong Digital Companions |
Assistants & memory | 2026 | integrate cross-session memory cues about a persistent user persona; continually update the agent's understanding of the user over time; adapt to the user's shifting privacy boundaries; sustain accurate emotional companionship across sessions | The agent must track a persistent user world (layered profile, event trajectory) across an average of 56.79 sessions and 851.85 user turns per persona, covering memory tracking, user understanding, privacy control, and emotional companionship. | hierarchical | no | avg 56.79 sessions / 851.85 user turns / 29.61K dialogue tokens per personasessions | 0 |
| LifeSim-Eval
LifeSim: Long-Horizon User Life Simulator for Personalized Assistant Evaluation |
Assistants & memory | 2026 | complete the user's explicit intentions correctly; recognize and satisfy the user's implicit intentions; recover an accurate model of the user's evolving profile/preferences over the course of long-horizon assistance; produce high-quality responses across 8 life domains and 1,200 diverse scenarios | Agent must recover and continuously update the user's profile (explicit and implicit intentions, preferences) as it evolves across a long-horizon, multi-scenario life trajectory, using a multi-turn interactive assessment method. | hierarchical | unclear | none stated | 101.67/mo |
| LiveClawBench
LiveClawBench: Benchmarking LLM Agents on Complex, Real-World Assistant Tasks |
Assistants & memory | 2026 | resolve tasks that span cross-service dependencies across mocked applications; operate correctly despite contaminated/inconsistent prior state; correctly infer implicit user intent not explicitly stated in the request; adapt to runtime changes within stateful mock services during task execution | The agent must track session state, artifacts, and prior side-effects across 22 stateful mocked services, resolving cross-service dependencies and implicit intent while adapting to runtime changes during the task. | DAG-with-precedence | no | none statedother:not-stated | 50.83/mo |
| MemConflict
MemConflict: Evaluating Long-Term Memory Systems Under Memory Conflicts |
Assistants & memory | 2026 | retrieve and rank the temporally valid, factually correct, and contextually applicable memory candidate when multiple conflicting alternatives exist; correctly answer queries under dynamic, static, and conditional conflict types despite distractors and long conflict distances | The agent's memory system must retrieve and rank memory candidates while tracking temporal validity, factual correctness, and contextual applicability across an average of 52.33 sessions (2,349.17 turns, ~203,910 tokens) per instance, correctly resolving conflicts placed 5 to 49 sessions apart. | DAG-with-precedence | yes | average of 52.33 sessions and 2,349.17 dialogue turns (about 203,910.83 tokens of context) per benchmark instance; conflict distances span 5-25 sessions (dynamic), 10-45 sessions (static), and 9-49 sessions (conditional)sessions | 51.25/mo |
| MemGround
MemGround: Long-Term Memory Evaluation Kit for Large Language Models in Gamified Scenarios |
Assistants & memory | 2026 | recall surface-level game state facts (Surface State Memory); associate events across time (Temporal Associative Memory); perform reasoning that depends on accumulated memory (Reasoning-Based Memory); unlock/discover memory fragments in the correct order across a gamified scenario | The agent must maintain and update a three-tier memory store (surface state, temporal associations, reasoning-derived facts) across continuous gamified interactions, tracked via Memory Fragments Unlocked and Memory Fragments with Correct Order metrics, with runs capped at 600-1000 interaction steps depending on task type. | hierarchical | yes | max 600-1000 interaction steps (task-dependent); early stop after 200 consecutive steps with no new discoveryother:interaction-steps | 10.17/mo |
| MemOps
MemOps: Benchmarking Lifecycle Memory Operations in Long-Horizon Conversations |
Assistants & memory | 2026 | correctly execute each lifecycle memory operation (remember, forget, update, reflect, and their compositions) at the right point in a conversation; maintain a consistent, ordered memory-state trajectory (not just the correct final answer) across a long conversation | The agent's memory system must track the trigger, target, scope, and state transition of each lifecycle operation (remember/forget/update/reflect) across up to 9,672 dialogue turns, maintaining a correct ordered memory-state trajectory rather than only a final answer. | sequential-chain | yes | 9,672 dialogue turns across 403 evidence conversations (100 unique topics), decomposed into 1,209 evidence-conversation segmentsturns | 31.5/mo |
| MemoryArena
MemoryArena: Benchmarking Agent Memory in Interdependent Multi-Session Agentic Tasks |
Assistants & memory | 2026 | distill experience from earlier actions and feedback into memory during multi-session interaction; use previously distilled memory to guide later actions and solve subsequent, explicitly interdependent subtasks; solve overall tasks spanning web navigation, preference-constrained planning, progressive information search, and sequential formal reasoning | The agent must track which experiences it has distilled into memory across earlier sessions and correctly recall/apply the relevant portions of that memory to solve later, interdependent subtasks across multiple domains (web navigation, planning, search, formal reasoning). | DAG-with-precedence | no | none stated | 628.86/mo |
| MEMPROBE
MEMPROBE: Probing Long-Term Agent Memory via Hidden User-State Recovery |
Assistants & memory | 2026 | assist simulated users across a trajectory of leak-controlled tasks while accumulating memory; recover/reconstruct a hidden, taxonomy-anchored user-state bank (31 dimensions) from the agent's own resulting memory; balance successful task assistance against auditable, faithful memory recovery | Agent must accumulate and retain a faithful memory of 31 hidden user-state dimensions across a trajectory of leak-controlled assistance tasks, since the memory is later audited by reconstructing the user-state bank from it under full-store and top-k access. | sequential-chain | no | none stated | 20.67/mo |
| MERIT
When Does Memory Help? A Cost-Aware Evaluation of Long-Term Memory in Tool-Using LLM Agents |
Assistants & memory | 2026 | correctly recall and use an earlier-episode fact when executing a later, dependent tool-use task; correctly recall and use an UPDATED fact (superseding a stale one) rather than acting on outdated information; operate under an explicit cost budget (token/dollar metering) while doing so; avoid corrupted/adversarially degraded memory leading to incorrect actions | The agent's memory system must retain facts (including corrections to previously stored facts) across an arc of linked episodes and correctly retrieve and act on the current, updated version of a fact rather than a stale cached one, all while the harness meters the token/dollar cost of every memory operation. | DAG-with-precedence | yes | episodes grouped into arcs of 4-6 linked episodes sharing entities; 10 arcs x 5 episodes per (domain x difficulty x condition), 23,440 scored episodes totalepisodes | 0 |
| Momento
Momento: Evaluating Persistent Memory and Reasoning with Multi-Session Agentic Conversations |
Assistants & memory | 2026 | take consequential, tool-mediated actions on behalf of a user within a multi-session service environment; resolve temporal dependencies between what happened in earlier sessions and what is being requested now; keep pace with evolving user goals across sessions rather than treating prior session history as static ground truth | Agent must track prior session history, recognize which parts of it may now be stale, and re-validate temporal dependencies and evolving user goals before taking consequential tool-mediated actions in the current session. | sequential-chain | unclear | none stated | 10.25/mo |
| MultiSessionCollab
MultiSessionCollab: Learning User Preferences with Memory to Improve Long-Term Collaboration |
Assistants & memory+ Multi-agent orgs | 2026 | solve each of 20 sequential collaboration problems (one per session) for a given user; learn and apply that user's preferences across sessions to improve collaboration quality and reduce user effort over time | The agent must track learned user preferences and reflections accumulated from prior sessions, applying them to reduce the number of conversational turns and user effort needed in each subsequent session across a 20-session sequence. | sequential-chain | yes | 20 sessions per user (one problem per session, up to 10 conversational turns per session), totaling 10,000 collaborative sessions per agent across the benchmark; turns needed per session drop from 10/8 to 6/4 by the third session when memory is usedsessions | 81.0/mo |
| Pare-Bench
Proactive Agent Research Environment: Simulating Active Users to Evaluate Proactive Assistants |
Assistants & memory | 2026 | observe evolving app/user state via a stateful finite-state-machine simulation to infer the user's current goal; correctly time an intervention (neither too early nor too late) once a need is inferred; orchestrate actions across multiple apps (communication, productivity, scheduling, lifestyle) to address the inferred goal | The agent must continuously observe the simulated user's stateful, sequential app interactions to infer an emerging goal, decide the right moment to intervene, and coordinate the intervention across multiple apps. | hierarchical | yes | simulation runs for a maximum of 10 turnsturns | 112.2/mo |
| PAST-Bench
PAST-Bench: Benchmarking the Foundations of Recursive Self-Improvement in Personal Agents |
Assistants & memory+ Information seeking | 2026 | reuse a retained skill/procedure across sessions on a later fresh-session task; retrieve a previously stored preference/fact and apply it correctly in a new session; gather information in one session that is needed to complete a task in a later session; update outdated retained state (e.g. a stale fact) rather than acting on it unchanged | The agent must save, retrieve, and update experience (preferences, task histories, tool routines, learned skills) across an ordered sequence of separate, fresh sessions, and the benchmark explicitly checks whether later-task gains actually follow this intended save/retrieve/update pathway rather than occurring by other means. | sequential-chain | yes | 26 scenarios, 204 episodes total; ordered sequences of fresh-session tasks per scenariosessions | 22.0/mo |
| PAUSE
PAUSE: A User-Centric Benchmark for Personal AI Assistants in Unified Service Environments |
Assistants & memory | 2026 | coordinate actions across heterogeneous user-owned services while respecting user-specific configurations and authorization/permission constraints; maintain consistency with evolving environment state across multi-turn interactions; for open-ended service-management tasks, satisfy semantic/behavioral trajectory-level goals; for constraint-intensive tasks, satisfy deterministic state-based verification conditions | The agent must track persistent user state, service-specific configurations and permissions, and prior actions taken across heterogeneous services, coordinating consistently across multi-turn interactions that scale from about 2 to over 5 dialogue rounds and roughly 13 to 22+ tool calls depending on difficulty. | DAG-with-precedence | yes | easy tasks average 2.11 dialogue rounds and 12.85 assistant tool calls; hard tasks average 5.23 dialogue rounds and 22.07 assistant tool callsturns | 0 |
| PersonaMem-v3
PersonaMem-v3: Toward Omni-Platform Personal Intelligence for Holistic User Understanding, Recommendation, and Agentic Tasks |
Assistants & memory | 2026 | build holistic cross-platform user understanding from social media, chatbot, calendar, and AI-companion engagement histories; personalize responses to reflect the user's evolving preferences over time; rerank recommendations on social media in a steerable way; act proactively across platforms when appropriate; hold back from personalizing when it would be inappropriate, repetitive, outdated, or unnecessary | The agent must track a time-indexed model of the user's preferences, intents, habits, and social relationships as they evolve across multiple platforms (social media, chatbot, calendar, AI-companion), and decide when NOT to act or personalize. | set-of-independent | no | none stated | 0 |
| PM-Bench
PM-Bench: Evaluating Prospective Memory in LLM Agents |
Assistants & memory | 2026 | maintain multiple ongoing and deferred intentions across a simulated week; execute a delayed intention at the correct future cue/state while continuing an ongoing activity; monitor latent environment changes relevant to deferred tasks | Agent must track multiple deferred intentions and their trigger conditions, continuously monitor latent environment/state changes, and decide at each point whether any deferred task is now due, while continuing an ongoing activity across a simulated week. | set-of-independent | no | sevensimulated-days | 10.5/mo |
| ProAgentBench
ProAgentBench: Evaluating LLM Agents for Proactive Assistance with Real-World Data |
Assistants & memory+ Business & enterprise | 2026 | predict the correct timing for a proactive intervention within a continuous workflow; generate appropriate assist content once an intervention point is identified | Agent must model long-term memory and historical, pre-assistance behavioral context (bursty interaction patterns, B=0.787) to decide both whether/when to intervene and what to say. | hierarchical | unclear | none stated | 101.43/mo |
| ProEvent
ProEvent: An Event-centric Benchmark for Proactive Agents |
Assistants & memory | 2026 | identify new upcoming events, including implicit ones, from ongoing instant-messaging chats; maintain and update a timetable of a user's events over time; time proactive responses correctly (neither too early nor too late); handle event cancellations correctly rather than overacting; produce correct single-step and multi-step responses per event | The agent must maintain a live, updatable timetable of a user's upcoming events, tracking concurrent chat threads, noise, and event cancellations, and decide both when and how (single- vs. multi-step) to respond as messages arrive. | set-of-independent | yes | none stated | 21.0/mo |
| RealMem
RealMem: Benchmarking LLMs in Real-World Memory-Driven Interaction |
Assistants & memory+ Business & enterprise | 2026 | track evolving project goals across long-term, cross-session dialogues; manage dynamic context dependencies (schedule/memory) inherent to real-world projects; respond correctly to natural user queries grounded in accumulated project history across eleven scenarios | The system must track long-term project states and dynamic context/schedule dependencies across more than 2,000 cross-session dialogues, since project goals evolve over time rather than remaining fixed. | sequential-chain | unclear | none stated | 162.0/mo |
| Setoka
Setoka: A Benchmark for Hierarchical User Understanding in Personalized Agents over Heterogeneous Data |
Assistants & memory | 2026 | retrieve explicit facts from past interactions (semantic memory); recall specific past episodes/events accurately (episodic memory); infer recurring behavior patterns from heterogeneous data over time (behavior pattern); infer abstract personality traits from heterogeneous, fragmented information (personality trait) | The memory-augmented agent must retain and integrate heterogeneous user data (explicit facts, episodic events, behavioral observations) dispersed over long-term interaction history to answer queries at each of the four hierarchical understanding levels. | hierarchical | no | none stated | 0 |
| Shopping Companion Bench
Shopping Companion: Benchmarking and Training LLM Agents for Long-Horizon Preference-Grounded E-Commerce Tasks |
Assistants & memory | 2026 | recommend products correctly aligned with preferences expressed across long-horizon conversations; manage a budget while shopping; assemble bundle deals satisfying multiple items' constraints jointly; correctly recall and apply user preferences carried over from earlier sessions (cross-session preference memory) | Agent must accumulate and correctly recall user shopping preferences across sessions, and verify product attributes against user requirements at each tool call, to avoid cascading preference-hallucination errors. | set-of-independent | yes | none stated | 40.67/mo |
| SovereignNegotiation-Bench
SovereignNegotiation-Bench: Evaluating User-Owned Personal Agents In Delegated Bargaining Under Privacy, Consent, Evidence, And Institutional Pressure |
Assistants & memory | 2026 | reach a negotiated agreement on the user's behalf (e.g., cost splits, refunds, subscription changes); preserve user utility while negotiating; avoid privacy leakage and consent violations during negotiation; ground claims in evidence and maintain auditability of the negotiation trace; escalate appropriately rather than over-concede under institutional pressure | The agent must track its private utilities/disclosure constraints, evidence requirements, and institutional-pressure cues across a multi-turn negotiation trace, keeping agent-visible observable state separate from evaluator-only labels. | set-of-independent | yes | 61,135 parsed action rows across 13,440 frozen-prompt live trajectories (~4.5 actions/trajectory)actions | 10.5/mo |
| SovereignPA-Bench
SovereignPA-Bench: Evaluating User-Owned Personal Agents under Evolving Intent, Platform Mediation, and Consent Constraints |
Assistants & memory | 2026 | advance a user's current, evolving interests while respecting privacy boundaries, consent constraints, and evidence requirements; minimize user burden while resisting manipulative platform incentives; preserve auditability of decisions across 120 sovereignty stress scenarios | Agent must track the user's evolving intent, what has been disclosed to which platform/party (ObservableState vs. evaluator-only HiddenLabels), consent already given, and accumulated burden across a scenario, all while resisting manipulative incentives. | DAG-with-precedence | yes | none stated | 10.5/mo |
| StateMemBench
Can Agent Memory Systems Track Evolving State? |
Assistants & memory | 2026 | track the evolving state of facts, constraints, and decisions as they are revised over a long multi-session interaction; answer questions reflecting the CURRENT state, not a superseded prior state | Agent's memory system must track the current value of each evolving fact/constraint/decision plus its supersession and relational dependencies, distinguishing current from superseded state at each query point. | sequential-chain | no | none stated | 0 |
| StreamMemBench
StreamMemBench: Streaming Evaluation of Agent Memory for Future-Oriented Assistance |
Assistants & memory | 2026 | correctly recall/use evidence observed in an initial task drawn from a streaming egocentric anchor; incorporate feedback/interaction experience from the initial task into a later follow-up task; carry evidence forward from what the agent observes and how the user interacts with it, to future similar tasks | The agent must carry stored evidence and interaction feedback forward from an initial task to a corresponding follow-up task drawn from continuous streaming egocentric observations, diagnosed via four metrics (evidence recall, initial evidence use, feedback incorporation, follow-up reuse). | sequential-chain | no | 2 (initial task + follow-up task) per evidence anchorother:task-steps | 0 |
| Supersede
Supersede: Diagnosing and Training the Memory-Update Gap in LLM Agents |
Assistants & memory | 2026 | answer using the current (most up-to-date) value of a fact that changes over time (e.g., a user's address, a price, a plan); discard/avoid using superseded (stale) fact values; maintain a bounded, self-maintained memory that keeps pace as the conversation grows | The agent must maintain a bounded, self-maintained memory of facts across long, multi-session interactions and always resolve to a fact's most current value, discarding superseded ones, as the conversation grows arbitrarily long. | set-of-independent | no | conversation length grows 24x (accuracy falls from 68% to 28% over this range, n=25)other:relative-conversation-length-growth-factor | 72.33/mo |
| TANGLE
When Personal Memory Has No Single Answer: Evaluating LLM Agents under Irreducible Conflict |
Assistants & memory | 2026 | recognize underdetermination when personal memory has no single answer; retain/preserve conflicting alternatives rather than collapsing to one definitive answer; seek clarification rather than acting on unjustified overconfidence; choose an action appropriate to Context-Partitioned, Behavior-Oscillation, or Source-Contradiction conflict types | Agent must preserve rather than resolve conflicting evidence, monitor five behavior dimensions (conflict perception, causal reasoning, confidence calibration, clarification seeking, memory faithfulness), and in the pipeline track extract and preserve conflict-bearing relations from multi-session dialogues. | set-of-independent | yes | none stated | 0 |
| VehicleMemBench
VehicleMemBench: An Executable Benchmark for Multi-User Long-Term Memory in In-Vehicle Agents |
Assistants & memory | 2026 | model multi-user preferences continuously as they evolve over time; resolve inter-user preference conflicts; correctly invoke 23 tool modules to reach a predefined target environment state; track changing user habits across many historical memory events | Agent must track evolving, sometimes conflicting, per-user preferences and habits across 80+ historical memory events, verifying its actions by comparing the resulting environment state to a predefined target state. | sequential-chain | no | over 80other:historical-memory-events-per-sample | 10.17/mo |
| VibeLifeBench
VibeLifeBench: Can Your Life Agent Be Proactive and Persistent in a Living World? |
Assistants & memory | 2026 | complete 200 long-horizon tasks across ten everyday-life domains over scripted multi-week timelines; proactively decide when to act, ask, or stay silent without being explicitly prompted; notice unannounced/silent world changes by re-inspecting the world; keep one plan coherent from the first day to the last while upholding unstated implicit constraints | Agent must track end-state goals, the timeliness of its own actions, and implicit constraints across scripted multi-week timelines in a simulated world of 22 mock services that changes on its own clock, much of it silently. | open-ended | no | multi-week (200 scripted tasks)other:multi-week-scripted-timeline | 0 |
| VitaBench 2.0
VitaBench 2.0: Evaluating Personalized and Proactive Agents in Long-Term User Interactions |
Assistants & memory | 2026 | continuously extract, utilize, and update evolving user preferences across a temporally ordered sequence of tasks for one user; proactively recognize missing information and actively acquire it from users/environment before making a decision | The agent must continuously extract, update, and apply an evolving set of user preferences (which can be added, deleted, or modified between tasks) across a temporally ordered sequence of at least 10 tasks per user, and proactively recognize when it needs to ask for missing information. | sequential-chain | yes | users have at least 10 tasks in their temporally ordered task sequences (56 users, 819 subtasks total)episodes | 51.25/mo |
| WorldBench
WorldBench: Culturally Grounded Benchmark for Multilingual Agents |
Assistants & memory | 2026 | complete genuine, persona-grounded everyday workflows across seven languages and eight cultures; preserve sandbox environment state while acting via structured actions; minimize unnecessary modification/side effects while completing the requested task | Agent must track sandbox state to complete persona-grounded workflows correctly while minimizing unwanted modifications, across long-horizon tasks and language/culture variation. | sequential-chain | no | none stated | 0 |
| WorldMemArena
WorldMemArena: Evaluating Multimodal Agent Memory Through Action-World Interaction |
Assistants & memory | 2026 | track and update an evolving personal/task state across a lifelong-evolution scenario; write, maintain, retrieve, and use memory correctly through the four-stage Action-World Interaction Loop; use visual evidence from real observations, actions, and feedback (Agentic Execution); answer QA correctly against gold memory points while resisting annotated distractors | The agent must write new memory from observations/actions/feedback, maintain it against evolving personal/task state (Lifelong Evolution), and correctly retrieve/use it later, annotated against gold memory points, updates, distractors, and evidence chains across multiple sessions. | sequential-chain | no | none statedother:not-stated | 10.25/mo |
| π-Bench
π-Bench: Evaluating Proactive Personal Assistant Agents in Long-Horizon Workflows |
Assistants & memory | 2026 | identify and act on hidden/unstated user needs before they are explicitly stated; complete each of 100 multi-turn tasks across 5 domain-specific user personas; resolve inter-task dependencies across sessions; maintain continuity of prior interaction context across sessions to resolve later proactive intents | The agent must track hidden/unstated user intents, inter-task dependencies, and information carried over across sessions, jointly measuring proactivity (anticipating needs) and task completion (executing them) over extended interactions. | DAG-with-precedence | no | none stated | 30.75/mo |
| Evo-Memory
Evo-Memory: Benchmarking LLM Agent Test-time Learning with Self-Evolving Memory |
Assistants & memory | 2025 | search, adapt, and evolve memory after each interaction in a sequential task stream; solve each task drawn from 10 diverse multi-turn goal-oriented and single-turn reasoning/QA datasets; reuse experience accumulated from earlier tasks to improve on later tasks in the same stream | Agent must search, adapt, and evolve its own memory continuously after each interaction across a sequential task stream, integrating reasoning, task actions, and memory updates to achieve continual improvement. | sequential-chain | no | none stated | 12412.4/mo |
| Forgetful but Faithful Agent (FiFA) benchmark
Forgetful but Faithful: A Cognitive Memory Architecture and Benchmark for Privacy-Aware Generative Agents |
Assistants & memory | 2025 | maintain narrative coherence across long-term interaction; complete multi-step goals despite a bounded memory budget; preserve social recall accuracy under six candidate forgetting policies; preserve privacy by not retaining/leaking information beyond its retention schema; minimize cost while balancing the above under different memory budgets | The agent must track its own memory budget consumption, which information to retain vs. forget under its chosen policy, ongoing multi-step goal progress, and social/privacy-relevant facts, across long-term interactive scenarios. | set-of-independent | no | none statedother:not-stated | 111.22/mo |
| Mem-alpha
Mem-α: Learning Memory Construction via Reinforcement Learning |
Assistants & memory | 2025 | extract and store relevant content from sequential information chunks into an external memory system; organize stored content across core, episodic, and semantic memory components; correctly answer downstream questions using the full accumulated interaction history; invoke the right memory-operation tool (from a multi-tool memory architecture) at the right time | The agent must track which information has been extracted and stored, how it is structured across core/episodic/semantic memory, and how it should be updated as new chunks arrive, since reward is downstream QA accuracy over the full interaction history. | sequential-chain | yes | trained up to 30k tokens; generalizes to sequences exceeding 400k tokens (>13x training length)other:tokens-of-accumulated-interaction-history | 272.25/mo |
| MemoryAgentBench
Evaluating Memory in LLM Agents via Incremental Multi-Turn Interactions |
Assistants & memory+ Information seeking | 2025 | accurately retrieve previously seen information from accumulated context; adapt to and learn from new information at test time; understand and reason over long-range accumulated context; selectively forget information that is no longer relevant or valid | The memory agent must incrementally accumulate, update, and retrieve information across multi-turn interactions, and correctly discard information rendered obsolete while retaining what remains useful. | set-of-independent | yes | none stated | 21315.21/mo |
| MemoryBench
MemoryBench: A Benchmark for Memory and Continual Learning in LLM Systems |
Assistants & memory | 2025 | learn from accumulated user feedback received during service time; apply continually-updated knowledge correctly across multiple domains, languages, and task types; outperform static (non-continual-learning) baselines after repeated feedback exposure | The system must track and integrate a stream of accumulated user feedback over service time, across multiple domains, languages, and task types, updating its internal state/parameters rather than treating each query independently. | sequential-chain | no | none stated | 494.45/mo |
| PERSONAMEM
Know Me, Respond to Me: Benchmarking LLMs for Dynamic User Profiling and Personalized Responses at Scale |
Assistants & memory | 2025 | internalize a user's inherent traits and preferences from interaction history; track how the user's profile/preferences evolve over time across sessions; generate a personalized response consistent with the current (most up-to-date) state of the user's profile in a new scenario | The chatbot must internalize inherent user traits, track how the user's profile evolves session by session, and select the response consistent with the user's current (not stale) profile state when answering an in-situ, first-person query. | sequential-chain | no | over 180 simulated interaction histories, each containing up to 60 sessions of multi-turn conversationssessions | 1589.29/mo |
| PROBE
Beyond Reactivity: Measuring Proactive Problem Solving in LLM Agents |
Assistants & memory+ Business & enterprise | 2025 | search for unspecified, unprompted issues across the user's available context/data; identify the specific bottleneck underlying an ambiguous problem; execute an appropriate resolution action autonomously | The agent must track candidate evidence surfaced during open-ended search, determine which issues are genuine unresolved bottlenecks, and carry that determination forward into selecting and executing a resolution, all without being told what to look for. | sequential-chain | yes | none stated | 80.73/mo |
| SimuHome
SimuHome: A Temporal- and Environment-Aware Benchmark for Smart Home LLM Agents |
Assistants & memory | 2025 | answer state-inquiry questions about the current smart-home environment; infer implicit user intent behind an ambiguous request; execute explicit device-control commands via SimuHome APIs; schedule and coordinate multi-device workflows whose effects evolve environmental variables over time; recognize and appropriately reject infeasible requests | The agent must track how already-issued device commands change environmental variables over (accelerated) simulated time and use this state to judge whether new requests are feasible or already satisfied. | DAG-with-precedence | yes | 600 episodes; simulator accelerates time so scheduled workflows can be evaluated immediatelyepisodes | 70.58/mo |
| UserBench
UserBench: An Interactive Gym Environment for User-Centric Agents |
Assistants & memory | 2025 | proactively clarify a simulated user's underspecified/vague initial goal; incrementally uncover and track multiple user preferences revealed gradually over a multi-turn interaction; make grounded decisions with tools that align with all of the user's (eventually revealed) intents | The agent must track which of the user's underspecified goals and incrementally revealed preferences it has already clarified, and continue to proactively elicit unclarified ones across the multi-turn interaction while using tools to act on what it has learned so far. | set-of-independent | no | none stated | 624.43/mo |
| Science14 artifacts | ||||||||
| ASI-Bench
ASI-Bench: At the Dawn of Artificial Superintelligence |
Science | 2026 | independently select an appropriate research method for a project-level research task; conduct the research (experimentation/analysis) using the selected method; produce verifiable results at the level of a full research project, with progressively less methodological guidance provided across three guidance tiers | System must track which methodological guidance tier is currently active for a task, whether a method has been selected, and whether the resulting research output is verifiable against expert review, across an entire project-level research process. | hierarchical | yes | none stated | 22.0/mo |
| AstroReason-Bench
AstroReason-Bench: Evaluating Unified Agentic Planning across Heterogeneous Space Planning Problems |
Science | 2026 | schedule ground-station communication passes; schedule agile Earth-observation tasks; satisfy heterogeneous mission objectives under strict physical constraints within a Space Planning Problem instance | Agent must track scheduling state across multiple heterogeneous objectives (communication windows, observation opportunities) and physical/orbital constraints simultaneously. | DAG-with-precedence | no | none stated | 10.12/mo |
| BixBench3
BixBench3: Benchmarking AI agents on research-study-scale computational biology tasks |
Science | 2026 | execute a sequence of computational-biology analyses from raw data through to a research objective; produce each of multiple data artifacts (e.g., peak call matrices, differential expression tables) matching the original published study; manage large raw datasets (some over 100GB) within time/cost constraints; maintain coherence across multiple sequential analysis steps | The agent must track intermediate data artifacts produced at each analysis step, manage large raw datasets, and maintain coherence across multiple sequential analyses before producing final gradable artifacts. | sequential-chain | yes | average 6.8 hours per task (longest attempts up to 24 hours)wall-clock-hours | 0 |
| ChemCost
Can Agents Price a Reaction? Evaluating LLMs on Chemical Cost Reasoning |
Science | 2026 | ground chemical identities from a reaction description; retrieve supplier quotes for the grounded chemicals; select valid purchasable packs matching required quantities; normalize quantities across packs; compute the total procurement cost from the reaction description | The agent must track which chemicals have been correctly grounded, which supplier quotes and packs have been retrieved/selected, and carry normalized quantities forward correctly into the final arithmetic cost computation, especially under noise-injected perturbations (aliases, quantity expressions, missing fields, formatting). | sequential-chain | no | none statedother:not-stated | 10.25/mo |
| DiscoverPhysics
DiscoverPhysics: Benchmarking LLMs for Out-of-the-Box Scientific Thinking |
Science | 2026 | design a sequence of informative experiments to probe an unknown simulated world's physics; revise hypotheses about the governing physical law across multiple rounds based on observed trajectory data; submit both a natural-language explanation and a Python implementation of the inferred law for each of 22 worlds | Agent must track its accumulated experimental observations (trajectory data) and current hypothesis about the world's physics across several rounds before submitting a final explanation and implementation. | sequential-chain | yes | none stated | 41.0/mo |
| iNatDisco
Autonomous Scientific Discovery via Iterative Meta-Reflection |
Science | 2026 | generate scientific hypotheses about ecological patterns from a dataset without a pre-specified research question; validate each proposed hypothesis via statistical testing before accepting it; periodically synthesize accumulated prior discoveries to redirect exploration toward unexplored regions of the hypothesis space; incorporate multimodal tool use (e.g., image processing) to extract further supporting evidence | The agent must maintain a growing record of prior discoveries and their statistical validation status, and periodically re-analyze that record to identify structural patterns, confounds, and epistemic gaps. | open-ended | yes | none stated | 10.5/mo |
| InquiTree (IT-18 subset)
InquiTree: Evaluating AI Agents in the Scientific Inquiry Loop with Paper-Derived Research Trees |
Science | 2026 | formulate a hypothesis consistent with a paper-derived research tree's logical dependencies; design a study to test the hypothesis; interpret the resulting study outcome; update beliefs/conclusions based on interpreted results, propagating correctly through the DAG | Agents must track their evolving beliefs/conclusions across nested subtopic branches, remain consistent with prior hypothesis/study/interpretation nodes, and avoid 'Erosion of Marginal Capabilities' (degrading critical judgment) over long-horizon interactions. | DAG-with-precedence | yes | theoretical interaction range of roughly 360 (3x120) to 1320 (11x120) reasoning-action turns for full traversal of the IT-18 subset (120 subtopics, H=4); per-task bounds of 21-77 steps for a typical task (n=7, H=4)turns | 10.33/mo |
| LabOSBench
LabOSBench: Benchmarking Computer Use Agents for Scientific Instrument Control |
Science | 2026 | complete each stage of a scientific-instrument operation workflow: sample loading, alignment, parameter tuning, data acquisition, and result inspection; perform feedback-driven parameter adjustment based on live instrument readouts; correctly operate one of 8 distinct instrument simulators across 96 subtasks | Agent must track which workflow stage it is in (loading, alignment, tuning, acquisition, inspection), instrument readouts/feedback from its own prior actions, and calibration state carried from earlier stages. | sequential-chain | yes | none stated | 10.33/mo |
| LifeSciBench
LifeSciBench: Evaluating Language Models on Realistic, Expert-Level Tasks in the Life Sciences |
Science | 2026 | execute a chain of multiple dependent judgment calls within one realistic life-science research task; satisfy a human-expert-written rubric spanning one of seven representative scientific workflows; operate correctly across each of seven life-science domains | The agent must track which dependent judgment calls it has made so far within a task and ensure later calls remain consistent with earlier ones, as graded against a human expert-written rubric. | sequential-chain | yes | none statedother:not-stated | 44.0/mo |
| RWE-bench
Can LLM Agents Generate Real-World Evidence? Evaluating Observational Studies in Medical Databases |
Science+ Healthcare | 2026 | construct a patient cohort from a real database (MIMIC-IV) per a study protocol; perform an analysis matching a peer-reviewed observational study's methodology; produce a coherent, tree-structured evidence bundle for reporting; iteratively execute and refine experiments against the reference protocol | The agent must track its cohort definition, intermediate analysis results, and their organization into a tree-structured evidence bundle, checking internal coherence against the reference study protocol across the whole task. | hierarchical | yes | none stated | 20.33/mo |
| SciAgentGym / SciAgentBench
SciAgentGym: Benchmarking Multi-Step Scientific Tool-use in LLM Agents |
Science | 2026 | orchestrate domain-specific scientific tools correctly across four natural science disciplines; complete tasks across a tiered difficulty spectrum from elementary actions to long-horizon workflows; sustain performance as interaction horizons extend rather than degrading | The agent must track which of 1,780 domain-specific tools it has invoked and their outputs across a tiered workflow, since performance is shown to degrade substantially as the interaction horizon extends, requiring sustained state-tracking rather than one-off calls. | hierarchical | no | none stated | 91.29/mo |
| AFMBench
Evaluating large language model agents for automation of atomic force microscopy |
Science | 2025 | complete the full scientific workflow from experimental design to results analysis for atomic force microscopy (AFM) automation; successfully perform each of several increasingly advanced experiments: AFM calibration, feature detection, mechanical property measurement, graphene layer counting, and indenter detection; coordinate across multi-agent roles in laboratory settings without deviating from instructions ('sleepwalking') | Agent(s) must track experimental state and configuration across the full workflow (design, calibration, measurement, analysis), coordinate with other agents in multi-agent setups, and avoid deviating from given instructions across a sequence of physical lab actions. | hierarchical | yes | none stated | 615.55/mo |
| BoxingGym
BoxingGym: Benchmarking Progress in Automated Experimental Design and Model Discovery |
Science | 2025 | design and run informative experiments that reduce uncertainty about a generative model's parameters; propose a scientific model/theory of the given environment; revise the proposed theory in light of newly collected experimental data; produce an explanation of the model that lets another agent make reliable predictions | The agent must track what has been learned from experiments already run (to target expected information gain for the next experiment) and maintain an evolving explanation of its current best scientific model. | sequential-chain | yes | evaluated after 0, 1, 3, 5, 7, and 10 experiment-design steps per trial, with 5 independent trials per environmentactions | 110.55/mo |
| ScienceBoard
ScienceBoard: Evaluating Multimodal Autonomous Agents in Realistic Scientific Workflows |
Science | 2025 | autonomously interact with professional scientific software (biochemistry, astronomy, geoinformatics) to accomplish a research task; complete each of 169 rigorously validated real-world scientific-discovery workflow tasks | Agent must track intermediate results and state produced while autonomously interacting with dynamic, visually rich professional software across a multi-step scientific workflow. | sequential-chain | yes | none stated | 432.69/mo |
| Tool & API29 artifacts | ||||||||
| Benchmarking AI Agents for Hardware Design Automation via MCP Tool Calling
Benchmarking AI Agents for Hardware Design Automation via MCP Tool Calling |
Tool & API | 2026 | execute single-operation edits on hardware design components via specialised MCP tools; correctly sequence multi-step dependency chains (e.g. create component -> add port -> wire connection); handle invalid or misspelled requests without incorrect tool invocation; operate correctly across multi-server tool contexts | The agent must track which prior dependency-establishing tool calls (e.g. component creation, port creation) have already succeeded before issuing calls that depend on them, across single-agent and multi-agent tool-calling configurations. | DAG-with-precedence | yes | 1 (Easy) / 2 (Medium) / 3-5 (Hard) expected tool calls per tasktool-calls | 0 |
| AgentEscapeBench
AgentEscapeBench: Evaluating Out-of-Domain Tool-Grounded Reasoning in LLM Agents |
Tool & API | 2026 | invoke real external tool functions to satisfy a directed acyclic dependency graph over tools and items; track hidden state revealed incrementally through tool use; propagate intermediate results correctly across dependent tool calls to a deterministically verifiable final answer | Agent must track hidden state revealed incrementally by tool calls, maintain clue adherence, and correctly propagate intermediate results through the dependency graph as depth increases. | DAG-with-precedence | yes | difficulty tiers from 5 to 25 (dependency-graph depth levels)other:dependency-graph-depth-tiers | 0 |
| AgentFloor
AgentFloor: How Far Up the tool use Ladder Can Small Open-Weight Models Go? |
Tool & API | 2026 | follow instructions correctly (lowest tier); use tools correctly (mid-lower tier); coordinate multiple steps/tool calls together (mid-upper tier); sustain long-horizon planning under persistent constraints over many steps (top tier) | Agent must sustain constraint tracking and coordination reliably over many steps to succeed at the highest tiers, where 'neither side reaches strong reliability' even among frontier models. | hierarchical | unclear | none stated | 0 |
| AgentGym2
AgentGym2: Benchmarking Large Language Model Agents in De-Idealized Real-World Environments |
Tool & API | 2026 | execute end-to-end real-world procedures without relying on pre-packaged tool interfaces; discover available tools via active exploration of the environment; compose discovered tools to solve previously unseen tasks; remain robust to noisy and underspecified task information | Agent must track which tools/interfaces it has discovered so far, what it has verified about their (possibly noisy) behavior, and how these compose toward completing the current end-to-end task. | open-ended | yes | none stated | 10.5/mo |
| APIFlow-Bench
APIFlow-Bench: Measuring Whether Agents Survive Long, Dependent API Workflows |
Tool & API | 2026 | complete each subtask in a long, dependent chain of REST-API calls correctly (state, not just completion, must be right); produce a final answer whose delivery is traceable to the actual call path (provenance-sensitive correctness); avoid compounding failures across increasingly long dependency chains (up to 20 subtasks) | Agent must track state correctness at each subtask in a dependency chain (not just whether the chain 'completed'), including whether a mock-minted canary correctly propagates through the API data flow to the final delivered answer. | DAG-with-precedence | yes | 20 (clean chains); up to 44,362 execution transcripts releasedother:subtasks-per-workflow-chain | 0 |
| AppWorld-UL
AppWorld-UL: Benchmarking Diverse Agent-User Interactions for Tool-Use |
Tool & API+ Assistants & memory | 2026 | operate applications correctly to complete a digital task (e.g. ordering groceries) across 9 simulated apps; interact appropriately with the user: ask clarification questions, prompt for confirmation, or report infeasibility; succeed on compositional, multi-part sub-tasks that combine several of the above interaction types | Agent must track what has been asked/confirmed with the user so far, the user's carefully-bounded knowledge state (simulated by an LLM), and which parts of a compositional task remain to be completed. | hierarchical | yes | none stated | 10.5/mo |
| AsyncTool
AsyncTool: Evaluating the Asynchronous Function Calling Capability under Multi-Task Scenarios |
Tool & API | 2026 | concurrently manage multiple heterogeneous tasks presented simultaneously; make productive use of idle time while awaiting delayed tool-call responses (asynchronous tool calling); coordinate task switching, dependency tracking, and state maintenance across concurrently running tasks | Agent must track the state, dependencies, and pending tool responses of multiple simultaneously active tasks, and coordinate task-switching decisions during periods of delayed tool feedback. | DAG-with-precedence | yes | none stated | 30.75/mo |
| C-World
C-World: A Computer Use Agent Environment Creator |
Tool & API | 2026 | complete long-horizon workflows composed of many interacting constraints across up to 5,571 tools spanning 204 applications; correctly follow constraints despite injected realistic failures and perturbations during the task; satisfy a reward signal combining verifiable metrics with LLM-based judgment | Agent must track constraint satisfaction across a long-horizon, multi-tool workflow while detecting and adapting to injected failures/perturbations introduced by the environment's transition function. | DAG-with-precedence | yes | none stated | 10.12/mo |
| CCTU
CCTU: A Benchmark for Tool Use under Complex Constraints |
Tool & API | 2026 | select and call the correct tool while satisfying every one of several simultaneous constraints (resource, behavior, toolset, response dimensions); maintain compliance with all constraints across a multi-turn interaction, not just the first tool call; self-refine after receiving feedback about a constraint violation | The agent must track which of the ~7 simultaneous constraints (out of 12 categories across 4 dimensions) apply to the current tool-use scenario and re-check compliance with all of them at every step across the multi-turn interaction, especially after receiving feedback about a violation. | set-of-independent | yes | maximum 20 interaction rounds per test caseturns | 71.17/mo |
| GeoAgentBench (GABench)
GeoAgentBench: A Dynamic Execution Benchmark for Tool-Augmented Agents in Spatial Analysis |
Tool & API | 2026 | correctly configure parameters for each of 117 atomic GIS tools invoked; complete each of 53 typical spatial-analysis tasks across 6 core GIS domains; produce spatially/cartographically accurate outputs verified via a VLM-based check; decouple global workflow orchestration from step-wise reactive execution to recover from runtime anomalies | The agent must track its evolving execution plan, per-step parameter choices, and runtime feedback/anomalies across a multi-step GIS workflow to keep global orchestration consistent with step-wise reactive execution. | sequential-chain | yes | none stated | 10.2/mo |
| GTA-2 (GTA-Atomic / GTA-Workflow)
GTA-2: Benchmarking General Tool Agents from Atomic Tool-Use to Open-Ended Workflows |
Tool & API+ Assistants & memory | 2026 | execute short-horizon, closed-ended atomic tool calls correctly (GTA-Atomic); complete long-horizon, open-ended, real-world productivity workflows end-to-end (GTA-Workflow); satisfy verifiable sub-goals identified by a recursive checkpoint-based evaluation mechanism | System must track which recursively-decomposed sub-goals/checkpoints of an open-ended workflow have been satisfied so far, across real deployed tools and multimodal contexts. | hierarchical | yes | none stated | 10.2/mo |
| Ko-WideSearch
Ko-WideSearch: A Korean Breadth-Search Benchmark for Exhaustive Set Enumeration by Web Agents |
Tool & API+ Web & GUI | 2026 | exhaustively enumerate the full membership of a named closed set (e.g. a TV season's cast, a dynasty's rulers); fill a per-item attribute table (multiple columns) for every enumerated member; decide when to stop searching within an open-ended web search space | The agent must track which set members it has already found (to avoid duplicates/omissions), which attribute columns remain unfilled per member, and its remaining search-iteration budget as difficulty knobs (table width, 2-D composite key) increase. | set-of-independent | no | fixed budget of 30 agent iterations per question (total tool calls run higher; one model logged up to 947 tool calls)tool-calls | 0 |
| MM-ToolBench (TOBench)
TOBench: A Task-Oriented Omni-Modal Benchmark for Real-World Tool-Using Agents |
Tool & API | 2026 | execute tools appropriate to a Customer Service or Intelligent Creation task; inspect rendered or transformed intermediate artifacts produced by tool calls; self-correct when an inspected artifact fails task-specific requirements; satisfy each of the 20 subcategory slices' task-specific grounded evaluator checks | The agent must track the state of rendered/transformed artifacts across a closed verification loop, deciding when self-correction is required, using task-specific grounded evaluators across 27 MCP servers and 324 tools. | sequential-chain | no | none statedother:not-stated | 0 |
| MM-ToolSandBox
MM-ToolSandBox: A Unified Framework for Evaluating Visual Tool-Calling Agents |
Tool & API | 2026 | ground progressively arriving visual inputs into correct executable tool calls across a multi-image, multi-turn interaction; handle realistic conversational phenomena mid-task: goal revisions, error corrections, state mutations; operate correctly across 500+ tools spanning 16 application domains; succeed on each of 258 human-verified nominal scenarios (plus 50 interactive-UI variants) | The agent must track a stateful execution environment across multi-image, multi-turn interactions, correctly incorporating goal revisions, error corrections, and state mutations as they occur, and ground each newly arriving visual input into the correct tool call. | sequential-chain | no | none stated | 0 |
| OmnilingualGAIA2
OmnilingualGAIA2: Evaluating the Multilingual Gap in Frontier AI Agents |
Tool & API | 2026 | plan a sequence of tool calls to answer a task; search for information via tools; execute multi-tool workflows; recover from errors during multi-tool execution; all under machine-translated, human-calibrated task instructions across ten languages/five scripts | Agent must track its tool-call plan, intermediate search/tool results, and error states to recover from execution failures, now while also handling machine-translated instructions across ten languages including non-Latin scripts. | DAG-with-precedence | yes | none stated | 0 |
| SeekerGym
SeekerGym: A Benchmark for Reliable Information Seeking |
Tool & API | 2026 | issue repeated retrieval queries to recover as much of a target document's content as possible; quantify uncertainty about how much relevant information might still be missing from what has been retrieved | The agent must track which passages/sections of the target document it has already retrieved, so it can judge how complete its coverage is and estimate how much information might still be missing. | open-ended | yes | none stated | 0 |
| SkillCraft
SkillCraft: Can LLM Agents Learn to Use Tools Skillfully? |
Tool & API | 2026 | compose atomic tools into reusable higher-level 'Skills'; cache and reuse learned Skills both within a task and across different tasks; complete compositional tool-use scenarios whose difficulty scales with entity count and subtask complexity | The agent must track which higher-level Skills it has already formed and cached, so it can reuse them instead of recomposing atomic tools from scratch on later subtasks/tasks, both within a single task and across the benchmark's 126 tasks. | hierarchical | yes | 126 tasks across 6 difficulty levels (from 21 seed tasks); tool-call counts scale from 9 (Easy: 3 subtasks x 3 calls) to 25 (Hard: 5 subtasks x 5 calls)tool-calls | 294.14/mo |
| The Amazing Agent Race (AAR)
The Amazing Agent Race: Strong Tool Users, Weak Navigators |
Tool & API+ Embodied & robotics | 2026 | navigate Wikipedia pages to locate the required entity/fact for each DAG node; execute the correct multi-step tool chain across fork-merge branches linking extracted entities; aggregate branch outputs into one verifiable final answer | The agent must track which Wikipedia entities/facts it has already extracted at each DAG node, correctly route them through parallel fork-merge tool-chain branches, and retain intermediate values needed for final aggregation across up to 5 diamonds per leg. | DAG-with-precedence | yes | average of 22.1 pit stops per leg for AAR-DAG (600 legs) and 15.0 pit stops per leg for AAR-Linear (800 legs), up to 5 diamonds per legother:pit-stops (navigation hops) | 10.2/mo |
| Toolathlon
Capable but Unreliable: Canonical Path Deviation as a Causal Mechanism of Agent Failure in Long-Horizon Tasks |
Tool & API | 2026 | complete each of 108 real-world tool-use tasks by following a canonical multi-step tool-invocation solution path; stay within the operating envelope of the canonical path across the whole trajectory to avoid stochastic drift/derailment | Success requires the trajectory of tool calls to stay within the operating envelope of the task's canonical solution path; mid-trajectory adherence must be tracked since drift compounds over subsequent calls. | sequential-chain | yes | none stated | 30.43/mo |
| ToolGym
ToolGym: an Open-world Tool-using Environment for Scalable Agent Testing and Data Curation |
Tool & API | 2026 | complete long-horizon, multi-tool workflows synthesized with wild constraints across 5,571 tools and 204 apps; recover and adapt when a state controller injects interruptions and failures mid-workflow; separate deliberate planning/self-correction from step-wise execution via a planner-actor decomposition | Agent must track tool-call state and wild constraints across a long-horizon, multi-tool workflow while detecting and recovering from injected interruptions/failures and unreliable tool states. | sequential-chain | no | none stated | 51.67/mo |
| ToolVerse / GUST dataset
ToolVerse: Unlocking Massive Environments and Long-Horizon Tasks for Agentic Reinforcement Learning |
Tool & API | 2026 | complete a long-horizon task built from a tool-dependency graph, where later tools unlock only after prerequisite tool-use subgoals are completed; correctly integrate the right subset of tools from a large pool (~4500 tools across ~400 MCPs) into a coherent multi-step solution; receive properly assigned turn-level credit despite long sequences of tool calls | The agent must track which prerequisite tools/subgoals in the dependency graph it has already satisfied in order to know which further tools are unlocked and relevant next, across a long sequence of tool calls. | DAG-with-precedence | yes | none stated | 21.0/mo |
| UniClawBench
UniClawBench: A Universal Benchmark for Proactive Agents on Real-World Tasks |
Tool & API | 2026 | demonstrate correct skill usage for the task's required tool/capability; explore the live environment (e.g. a Docker container) to discover needed information/actions; reason correctly over long context accumulated during the task; correctly interpret multimodal inputs relevant to the task; coordinate actions correctly across multiple platforms/services | The agent must track its progress against fine-grained, step-by-step completion checkpoints within a live environment, integrate multi-turn feedback from a hidden supervisor agent and a user-simulator agent without seeing the grading criteria, and coordinate across the relevant capability dimensions needed for that task. | sequential-chain | yes | none stated | 10.5/mo |
| CONFETTI
CONFETTI: Conversational Function-Calling Evaluation Through Turn-Level Interactions |
Tool & API | 2025 | handle user follow-up requests within an ongoing conversation; correct or switch goals mid-conversation when the user changes intent; resolve ambiguous or implicit user goals into an appropriate API call; chain multiple function calls together to satisfy a multi-step user request | The agent must track evolving user intent across conversation turns, recognize goal correction/switching, and maintain state of prior chained function-call results across up to 86 available APIs. | sequential-chain | yes | 313 user turns across 109 conversations (~2.9 turns/conversation on average)turns | 161.07/mo |
| DialogTool / VirtualMobile
Rethinking Stateful Tool Use in Multi-Turn Dialogues: Benchmarks and Challenges |
Tool & API+ Web & GUI | 2025 | create a new tool/API on demand within a multi-turn dialogue (tool creation); become aware of, correctly select, and execute the right tool given user intent (tool utilization); generate a role-consistent response, including role play, reflecting whether/how the tool was used | Agent must track which tools have been created/exist, their stateful execution history across the dialogue, and how to remain role-consistent while using them, across a multi-turn conversation. | sequential-chain | yes | none stated | 171.06/mo |
| M^3-Bench
M^3-Bench: Multi-Modal, Multi-Hop, Multi-Threaded Tool-Using MLLM Agent Benchmark |
Tool & API | 2025 | complete multi-hop, multi-threaded tool-call workflows requiring cross-tool dependencies; maintain persistence of intermediate resources across steps; ground visual and textual reasoning correctly across each tool call | Agent must track intermediate resources produced by earlier tool calls and align them across multiple concurrent threads for argument fidelity and structural consistency. | DAG-with-precedence | yes | none stated | 90.9/mo |
| Multi-Mission Tool Bench
Multi-Mission Tool Bench: Assessing the Robustness of LLM based Agents through Related and Dynamic Missions |
Tool & API | 2025 | complete each mission within a test case containing multiple interrelated missions; dynamically adapt when missions switch mid-interaction; correctly invoke tools appropriate to the currently active mission; handle all possible mission-switching patterns within a fixed mission number | The agent must track the state of each interrelated mission (active or paused), correctly recognize mission switches, and select/invoke the correct tools for whichever mission is currently active, evaluated via dynamic decision trees for accuracy and efficiency. | DAG-with-precedence | no | none stated | 70.41/mo |
| OrchDAG
OrchDAG: Complex Tool Orchestration in Multi-Turn Interactions with Plan DAGs |
Tool & API | 2025 | correctly execute a sequence of tool calls whose dependencies form a directed acyclic graph (DAG); respect topological/precedence constraints among tool calls across multi-turn interactions; solve DAG-structured tool-orchestration tasks of controllable/varying complexity | The agent must track which nodes (tool calls) in the DAG have been completed and what state/outputs they produced, to correctly select and sequence subsequent tool calls consistent with the graph's precedence constraints across multiple turns. | DAG-with-precedence | no | none stated | 10.09/mo |
| Tool Decathlon (Toolathlon)
The Tool Decathlon: Benchmarking Language Agents for Diverse, Realistic, and Long-Horizon Task Execution |
Tool & API+ Business & enterprise | 2025 | coordinate interactions across multiple named Apps/tools (e.g. email + calendar + file systems) to complete one complex workflow; diagnose and report anomalies via monitoring a database following an operating manual; satisfy a strictly execution-verifiable end state for each of 108 tasks spanning 32 apps and 604 tools | The agent must track the current, realistic state of multiple named applications across roughly 20 tool-calling turns per task, and verify its final state against a dedicated evaluation script. | DAG-with-precedence | yes | tasks require interacting with multiple Apps over around 20 turns on average (best model averages 20.2 tool-calling turns)turns | 666.0/mo |
| ToolHaystack
ToolHaystack: Stress-Testing Tool-Augmented Language Models in Realistic Long-Term Interactions |
Tool & API | 2025 | correctly maintain and disambiguate multiple concurrent task-execution contexts within one continuous conversation; handle realistic noise/disruptions injected into a long-term interaction without losing track of tool-use context; successfully complete tool-use tasks embedded within a long, continuous conversation despite these challenges | The model must maintain and disambiguate multiple concurrent task-execution contexts across a continuous long-term conversation while filtering out realistic noise and handling various disruptions, rather than resetting context between short, isolated tool-use exchanges. | set-of-independent | no | none stated | 50.31/mo |
| Planning & travel19 artifacts | ||||||||
| Behavior2Trip
Behavior2Trip: Towards Personalized Travel Planning via User Behavior Trajectory |
Planning & travel+ Assistants & memory | 2026 | infer a user's latent travel preferences from their past behavior trajectory rather than explicit instructions; generate a travel plan satisfying preferences across 14 attributes spanning 5 preference dimensions; achieve a full-constraint pass across all inferred preference constraints simultaneously | The agent must infer and track the user's latent preferences (14 attributes, 5 dimensions) from an average of 39.8 past behaviors, then check the generated plan against every inferred constraint for a full pass. | set-of-independent | yes | none stated | 0 |
| DeepPlanning
DeepPlanning: Benchmarking Long-Horizon Agentic Planning with Verifiable Constraints |
Planning & travel+ Information seeking | 2026 | satisfy local, fine-grained constraints on individual itinerary/shopping items; satisfy global constrained optimization objectives (e.g. time and financial budgets) across the whole plan; proactively gather information needed before constraints can even be checked | Agent must track accumulated time and financial budget consumption across a multi-day plan or multi-product list, alongside fine-grained local constraints, while still actively gathering new information mid-task. | DAG-with-precedence | unclear | none stated | 334.12/mo |
| GroupTravelBench
GroupTravelBench: Benchmarking LLM Agents on Multi-Person Travel Planning |
Planning & travel | 2026 | elicit each group member's private travel preferences through multi-turn dialogue; surface and resolve inter-user preference conflicts via compromise or subgrouping; produce a final plan that balances group utility against fairness across all members | Agent must track each group member's elicited (and initially private) preferences, detected conflicts between members, and the evolving fairness/utility trade-off of the plan across a synchronous multi-turn group-chat session. | DAG-with-precedence | yes | none stated | 41.0/mo |
| TravelEval
TravelEval: A Comprehensive Benchmarking Framework for Evaluating LLM-Powered Travel Planning Agents |
Planning & travel | 2026 | produce a full multi-day itinerary satisfying accuracy, compliance, temporality, spatiality, economy, and utility dimensions jointly; sequence daily accommodation, transport, and visit pacing consistently across the whole trip rather than per isolated day | The agent must track cumulative spatio-temporal cost (queuing times, transit distances), running budget/economy, and daily pacing/accommodation continuity across the entire multi-day itinerary, not just within a single day. | sequential-chain | yes | itineraries span 2-day, 3-day, 4-day, 5-day, 6-day, and 7-day durations across query categories (e.g. 400 medium-difficulty queries)simulated-days | 10.25/mo |
| TREK
TREK: A Travel Reasoning and Evaluation Kit for LLM Agents in Complex Trip Planning |
Planning & travel | 2026 | produce a single itinerary that is jointly constraint-correct, hallucination-free, spatio-temporally executable, and budget-valid; respond to the traveler's unstated persona needs; correctly identify provably infeasible tasks (267 of 800) versus feasible ones with typed causes | Agent must track running budget, spatio-temporal feasibility across days, entity/route validity, and unstated persona needs simultaneously while assembling a single itinerary. | other (singleton) | no | none stated | 0 |
| Trip+
Trip+: Benchmarking Agents in Personalized Interactive Travel Planning |
Planning & travel+ Assistants & memory | 2026 | generate a minute-level itinerary satisfying a traveler's profiled preferences; revise the itinerary in response to evolving preferences and unexpected environment-driven disruptions across multiple turns; avoid producing technically-feasible-but-exhausting plans (jointly satisfy feasibility and experiential/fatigue quality) | The agent must track the traveler's evolving profile/preferences, the current committed itinerary state at minute-level granularity, and cumulative experiential cost (e.g. fatigue) as it revises plans across several user turns per instance. | sequential-chain | yes | 153 multi-turn instances and 570 user turns totalturns | 10.33/mo |
| TRIP-Bench
TRIP-Bench: A Benchmark for Long-Horizon Interactive Agents in Real-World Scenarios |
Planning & travel | 2026 | satisfy each of 40+ curated travel requirements/global constraints across an itinerary; coordinate reasoning across 18 curated tools correctly within one dialogue; adapt to evolving user behavior, style shifts, feasibility changes, and iterative version revisions over a long, multi-turn interaction (hard split) | The agent must track global constraint satisfaction across 40+ travel requirements, the state of 18 tools' results, and evolving user preferences/style/feasibility across dialogues spanning up to 15 user turns, 150+ tool calls, and 200k+ tokens of context. | DAG-with-precedence | yes | dialogues span up to 15 user turns, can involve 150+ tool calls, and may exceed 200k tokens of contextturns | 71.0/mo |
| Trip-planning Optimization Problems (TOP) Dataset
Agentic AI for Trip Planning Optimization Application |
Planning & travel | 2026 | optimize (not merely satisfy) route/itinerary selection under travel-time, energy, and traffic factors; coordinate specialized sub-agents for traffic, charging, and points-of-interest; dynamically refine a plan against a definitive optimal solution reference | The orchestration agent must track recommendations from the traffic, charging, and POI sub-agents and reconcile them into one jointly-optimal plan, verified against the dataset's definitive optimal solutions rather than only a feasible reference answer. | hierarchical | no | none statedother:not-stated | 0 |
| WorldTravel / WorldTravel-Webscape
WorldTravel: A Realistic Multimodal Travel-Planning Benchmark with Tightly Coupled Constraints |
Planning & travel | 2026 | satisfy all of an average 15+ interdependent temporal and logical travel constraints simultaneously per scenario; extract constraint parameters from dynamic web environments/webpages rather than idealized data; perceive constraint parameters directly from visual layouts in a multi-modal setting; produce a feasible travel plan across 150 real-world scenarios in 5 cities | The agent must track which of the 15+ interdependent temporal/logical constraints have been satisfied so far, and in the multi-modal condition must also perceive constraint parameters directly from over 2,000 rendered webpages rather than being handed clean structured data. | DAG-with-precedence | no | an average of 15+ interdependent constraints per scenario; a Planning Horizon threshold at approximately 10 constraintsother:number-of-interdependent-constraints | 30.43/mo |
| COMPASS
COMPASS: Benchmarking Constrained Optimization in LLM Agents |
Planning & travel | 2025 | gather task information and constraints from the user via multi-turn conversation; use tools to gather relevant information from a database; propose a travel plan satisfying all hard constraints; optimize the plan for the user's utility objective beyond mere feasibility | Agent must track constraints and preferences gathered so far via conversation and tool calls, the current feasible-solution search space, and how well a candidate plan satisfies both hard constraints and the utility objective. | sequential-chain | no | none stated | 60.55/mo |
| CostBench
CostBench: Evaluating Multi-Turn Cost-Optimal Planning and Adaptation in Dynamic Environments for LLM Tool-Use Agents |
Planning & travel | 2025 | find a cost-optimal sequence of atomic and composite tool calls to solve a travel-planning task; detect and adapt to dynamic blocking events (e.g. tool failures, cost changes) that occur mid-task; replan in real time to remain cost-optimal after a blocking event | Agent must track the accumulated cost of its chosen tool sequence so far, remaining budget, and whether any of four types of dynamic blocking events have occurred, requiring real-time replanning to stay cost-optimal. | DAG-with-precedence | unclear | none stated | 303.0/mo |
| Flex-TravelPlanner
Flex-TravelPlanner: A Benchmark for Flexible Planning with Language Agents |
Planning & travel | 2025 | revise a travel plan as new constraints are introduced sequentially across turns; correctly prioritize competing constraints when a newly introduced lower-priority preference conflicts with an existing higher-priority constraint | The agent must track which constraints have been introduced so far, their relative priority, and whether previously satisfied requirements are still respected as new constraints arrive turn by turn. | sequential-chain | yes | up to 3 turns (all-at-once / 2-turn / 3-turn constraint-introduction patterns) across 120 base queriesturns | 110.73/mo |
| RETAIL
RETAIL: Towards Real-world Travel Planning for Large Language Models |
Planning & travel+ Business & enterprise | 2025 | infer and satisfy implicit user requirements (not just explicit queries); satisfy explicit queries with or without later revision needs; account for diverse environmental factors and constraints to ensure plan feasibility; produce an all-in-one plan with rich, detailed POI (point-of-interest) arrangement rather than only basic POI listing | The agent must infer unstated (implicit) requirements, track environmental constraints affecting plan feasibility, and assemble detailed POI information into a single all-in-one plan, revising an existing plan when a revision need is present. | DAG-with-precedence | no | none stated | 90.69/mo |
| Travel-Sim
Wide-Horizon Thinking and Simulation-Based Evaluation for Real-World LLM Planning with Multifaceted Constraints |
Planning & travel | 2025 | satisfy multiple parallel, potentially conflicting real-world planning constraints (e.g. preferences, logistics) within one itinerary; respect causal dependencies where earlier itinerary choices constrain which later activities remain feasible; produce a plan validated via realistic agent-based simulation rather than isolated constraint checks | The planner must track all outstanding multifaceted constraints and the downstream causal consequences of already-committed itinerary decisions as the simulated trip unfolds. | DAG-with-precedence | yes | none stated | 50.33/mo |
| TravelBench
Beyond Itinerary Planning-A Real-World Benchmark for Multi-Turn and Tool-Using Travel Tasks |
Planning & travel | 2025 | solve a travel-planning problem independently using cached tool results; interact with the user across multiple turns to elicit implicit preferences; correctly recognize and communicate the agent's own capability boundaries (Unsolvable subtask) | In the Multi-Turn subtask, the agent must track previously elicited or stated user preferences across turns and integrate them with cached tool results from a sandbox of ten travel-related tools, while also recognizing when a request exceeds its capability boundaries. | set-of-independent | yes | none statedother:not-stated | 60.67/mo |
| TripCraft
TripCraft: A Benchmark for Spatio-Temporally Fine Grained Travel Planning |
Planning & travel | 2025 | generate a spatiotemporally coherent 7-day travel itinerary; satisfy meal-scheduling constraints (Temporal Meal Score); satisfy attraction-timing constraints (Temporal Attraction Score); satisfy spatial feasibility across the itinerary (Spatial Score); satisfy activity-ordering constraints (Ordering Score); satisfy user-persona preferences (Persona Score) | The agent must track public-transit schedules, event availability, attraction categories, and user-persona preferences simultaneously while assembling a spatially and temporally consistent 7-day itinerary. | DAG-with-precedence | yes | 7simulated-days | 311.63/mo |
| TripScore
TripScore: Benchmarking and rewarding real-world travel planning with fine-grained evaluation |
Planning & travel | 2025 | produce a travel itinerary that jointly satisfies fine-grained feasibility, reliability, and engagement criteria; achieve a high unified reward score combining these criteria for RL training/evaluation; generalize to real-world, free-form travel requests | Agent must track and jointly satisfy fine-grained feasibility, reliability, and engagement criteria while assembling a travel plan, since these are unified into a single reward score used for both evaluation and RL training. | other (singleton) | no | none stated | 90.82/mo |
| TripTide
TripTide: A Benchmark for Adaptive Travel Planning under Disruptions |
Planning & travel | 2025 | preserve original itinerary intent (feasibility and goals) after a disruption; respond promptly and appropriately to a disruption event (flight cancellation, weather closure, overbooked attraction); adapt the itinerary with appropriate semantic, spatial, and sequential divergence from the original plan; maintain plan quality across varying disruption severity and traveler tolerance levels | The agent must track the original itinerary's intent, spatial layout, and sequential structure, detect and appropriately size its response to a disruption of a given severity and traveler tolerance, and measure how much the revision diverges from the original across semantic, spatial, and sequential dimensions. | sequential-chain | no | none stated | 50.45/mo |
| WandaPlan
Is Your LLM-Based Multi-Agent a Reliable Real-World Planner? Exploring Fraud Detection in Travel Planning |
Planning & travel+ Multi-agent orgs | 2025 | produce a correct multi-agent travel plan while resisting injected deceptive/fraudulent content; detect/avoid Misinformation Fraud; detect/avoid Team-Coordinated Multi-Person Fraud; detect/avoid Level-Escalating Multi-Round Fraud | The planning system must track which review/social-media sourced information it has incorporated, cross-check it for authenticity across rounds and multiple purported sources, and avoid building a travel plan on fraudulent inputs even as fraud escalates over multiple rounds. | set-of-independent | no | none stated | 191.19/mo |
| Web & GUI39 artifacts | ||||||||
| AgenticShop
AgenticShop: Benchmarking Agentic Product Curation for Personalized Web Shopping |
Web & GUI+ Assistants & memory | 2026 | explore the open web to curate a set of products satisfying diverse shopping scenarios; satisfy each checklist item in a verifiable, checklist-driven personalization rubric aligned to a user profile | The agent must track a user's personalization checklist criteria and diverse profile preferences while exploring open-web shopping scenarios, so curated products satisfy every checklist item. | set-of-independent | yes | none stated | 101.43/mo |
| AndroidDaily
AndroidDaily: A Verifiable Benchmark for Mobile GUI Agents on Real-World Closed-Source Applications |
Web & GUI | 2026 | complete each of 350 realistic daily-use tasks spanning 94 real closed-source Android apps; satisfy step-level operational obligations per GRADE's guideline criteria; meet output-quality criteria at each step; avoid violating negative constraints at each step | The agent must track progress against multiple step-level guideline criteria (obligations, quality, negative constraints) across a long-horizon, open-ended interaction with a closed-source app exposing no internal state. | sequential-chain | yes | none stated | 41.0/mo |
| AndroTMem-Bench
AndroTMem: From Interaction Trajectories to Anchored Memory in Long-Horizon GUI Agents |
Web & GUI+ Assistants & memory | 2026 | carry forward critical intermediate state across a long sequence of GUI interaction steps to complete a task; correctly resolve strong step-to-step causal dependencies where sparse intermediate states are decisive for later actions; complete each of 1,069 Android GUI tasks (avg. 32.1 steps, max. 65 steps) | The agent must carry forward sparse, dependency-critical intermediate state across an average of 32.1 (up to 65) interaction steps per task, since full-sequence replay is redundant/noisy and naive summarization erases exactly the dependency-critical information needed for later steps. | DAG-with-precedence | no | 1,069 tasks with 34,473 total interaction steps; average 32.1 steps per task, maximum 65 steps per taskactions | 111.83/mo |
| CAP
CAP: A Scalable Benchmark for Evaluating Cross-Site Browser Agents with Complex Actions and Perception |
Web & GUI | 2026 | complete a realistic cross-site workflow requiring several specific operations on each of multiple real-world websites; correctly interact with complex, dynamically rendered UI elements (non-trivial UI interactions); correctly perceive/interpret dynamically rendered visual content across the workflow | The agent must track which execution and perception checkpoints it has already satisfied across multiple websites within one recomposed cross-site workflow, using a verifiable agent-as-a-judge evaluation framework. | DAG-with-precedence | yes | baselines evaluated with a maximum of 50 reasoning-action steps per task; average of 7 execution points and 4 perception points per taskactions | 0 |
| ClawBench
ClawBench: Can AI Agents Complete Everyday Online Tasks? |
Web & GUI | 2026 | complete everyday online tasks (purchases, appointment bookings, job applications) across 144 real platforms; obtain relevant information from user-provided documents; navigate multi-step workflows across diverse platforms; correctly fill in many detailed form fields per task | Agent must extract and correctly carry information from user-provided documents through a multi-step, write-heavy workflow with many detailed form fields, operating on live, dynamic production websites. | sequential-chain | no | none stated | 224.4/mo |
| ComboShoppingBench
ComboShoppingBench: Evaluating LLM Agents for Budget-Constrained Basket Shopping with Coupons |
Web & GUI | 2026 | construct a basket of complementary items satisfying compatibility constraints; keep the basket within a stated budget while optimizing coupon use; satisfy store-level requirements, availability, and delivery fees jointly | Agent must track the running basket contents, remaining budget, which coupons are valid given the current basket, and per-store requirements while searching for a feasible combination. | set-of-independent | yes | none stated | 0 |
| DMV-Bench
DMV-Bench: Diagnosing Long-Horizon Multimodal Agents' Visual Memory with Incidental Cue Injection |
Web & GUI+ Assistants & memory | 2026 | complete chains of autonomous shopping sessions (browsing/selecting home-furnishing products); recall a unique, pre-rendered incidental visual cue seen earlier on a product image when later asked about it | Agent must retain visual memory of incidental cues (not just deliberately extracted facts) across chains of shopping sessions of varying length, without being told in advance which details will later be tested. | sequential-chain | unclear | none stated | 0 |
| GMA
Benchmarking General Mobile Assistants in Challenging Real-World Scenarios |
Web & GUI+ Assistants & memory | 2026 | complete tasks ranging from atomic actions to complex multi-step workflows across seven open-source-based applications; handle lifestyle-sharing and travel-planning domains at four escalating difficulty tiers; maintain context/state across a workflow via harness-level context retention and explicit state tracking | Agent must maintain context retention and explicit state tracking across multi-step workflows spanning seven applications, since performance declines substantially as task complexity/tier increases. | hierarchical | no | none stated | 0 |
| GTA
GTA: Generating Long-horizon Tasks for Web Agents at Scale |
Web & GUI | 2026 | complete multi-hop, cross-page web tasks that are compositional over a site graph; follow an intermediate trajectory of dense, process-level supervision steps, not just reach a coarse end goal; generalize across more than 50 websites (e-commerce, government, forums, news) including multilingual tasks | The agent must track its position and accumulated information across a multi-hop, cross-page trajectory grounded in a site graph, since tasks provide dense process-level supervision (intermediate trajectory steps) rather than only a coarse start-goal annotation. | DAG-with-precedence | yes | none stated | 0 |
| InterruptBench
When Users Change Their Mind: Evaluating Interruptible Agents in Long-Horizon Web Navigation |
Web & GUI+ Embodied & robotics | 2026 | execute a long-horizon, environmentally grounded web-navigation task; adapt when the user adds a new requirement mid-task; revise the current goal when the user changes an existing requirement mid-task; abandon a sub-goal when the user retracts a requirement mid-task; recover efficiently without redundant or incorrect actions after an interruption | The agent must track the current, possibly revised, task intent across single- and multi-turn interruption settings plus the environment's persistent state from already-executed actions to adapt or recover correctly. | sequential-chain | yes | none stated | 61.2/mo |
| iOSWorld
iOSWorld: A Benchmark for Personally Intelligent Phone Agents |
Web & GUI | 2026 | complete single-app tasks within one iOS app (27 tasks); complete multi-app task chains spanning 2 to 8 apps (60 tasks); infer personal patterns from persistent user data for memory/personalization tasks (46 tasks) | Agent must track a persistent user identity and its connected data (transactions, messages, travel records, social relationships, financial activity) across apps and infer behavioral patterns for personalization tasks. | DAG-with-precedence | no | none stated | 31.0/mo |
| LongMemEval-V2 (LME-V2)
LongMemEval-V2: Evaluating Long-Term Agent Memory Toward Experienced Colleagues |
Web & GUI+ Assistants & memory | 2026 | recall static state facts about the environment (static state recall); track dynamic state changes over time (dynamic state tracking); recall workflow knowledge/procedures learned from experience; recall environment-specific 'gotchas'/recurring failure modes; maintain premise awareness of what has already been established or assumed | The memory system must consume and internalize up to 500 history trajectories (115M tokens) and return compact, correct evidence for downstream question answering across five distinct memory-ability categories. | set-of-independent | yes | up to 500 trajectories and 115M tokens of historyepisodes | 92.25/mo |
| Memory-World
STAMP: Training Explicit Memory for Mobile GUI Agents in Controllable and Scalable Virtual Environments |
Web & GUI+ Assistants & memory | 2026 | encode a programmatically-injected memory variable at the correct point in a task; retain that memory variable correctly despite progressive discarding of older visual history; retrieve and correctly apply the memorized variable later in the same long-horizon mobile GUI task | The agent must explicitly memorize deterministic variables injected at specific points and correctly retrieve/apply them later in the same task, despite token-heavy screenshots forcing progressive discarding of older visual history. | sequential-chain | no | none statedother:not-stated | 10.25/mo |
| MobiFlow
MobiFlow: Real-World Mobile Agent Benchmarking through Trajectory Fusion |
Web & GUI | 2026 | complete real-world tasks within arbitrary third-party mobile apps whose success cannot be checked via system-level APIs; have task completion verified via a graph constructed from fusing multiple real user trajectories | The agent must track its current position within the app's fused trajectory-state graph and correctly follow one of the valid paths to task completion, since no system-level completion signal is available. | DAG-with-precedence | yes | none stated | 0 |
| Odysseys
Odysseys: Benchmarking Web Agents on Realistic Long Horizon Tasks |
Web & GUI | 2026 | complete long-horizon, multi-site web workflows (e.g., comparing products across domains); plan trips across multiple web services; summarize information gathered from multiple search queries; satisfy an average of 6.1 graded rubric criteria per task | The agent must sustain context and accumulated findings across multiple websites and search queries over potentially hours of browsing, and satisfy each of ~6.1 rubric criteria graded per task rather than a single pass/fail check. | DAG-with-precedence | yes | potentially hours (of browsing); Trajectory Efficiency measured as rubric score per stepwall-clock-hours | 132.6/mo |
| PAHF benchmarks (embodied manipulation + online shopping)
Learning Personalized Agents from Human Feedback |
Web & GUI+ Embodied & robotics | 2026 | learn a new user's initial preferences from scratch via pre-action clarification; ground actions in preferences retrieved from an explicit per-user memory; adapt rapidly to persona shifts using post-action feedback to update memory | Agent must maintain explicit per-user memory across a four-phase protocol, updating stored preferences via dual feedback channels (pre-action clarification, post-action feedback) as it learns initial preferences from scratch and later adapts to persona shifts. | sequential-chain | no | none stated | 152.14/mo |
| ParaGUIBench
Beyond Sequential Interaction: Benchmarking Parallel Execution and Coordination for GUI Agents |
Web & GUI | 2026 | identify which GUI sub-tasks can run concurrently despite unstated dependencies; avoid conflicts between concurrent workers modifying shared artifacts; ensure each worker's locally-completed sub-task composes into a globally correct combined result; complete each of 233 tasks spanning six task categories on separate desktop instances | The planner-worker system must track which sub-tasks are dependency-free and safe to parallelize, coordinate concurrent workers' access to shared artifacts, and verify that the union of workers' outputs satisfies the original combined instruction. | DAG-with-precedence | no | none statedother:not-stated | 0 |
| PhoneWorld
PhoneBuddy: Training Open Models for Agentic Phone Use |
Web & GUI | 2026 | complete single-app phone tasks; complete mini-app tasks; complete cross-app workflows requiring coordination across multiple mobile applications; achieve task success on a 150-task real-phone human evaluation and on AndroidWorld | The agent must track UI/app state across a real, stateful phone environment (or the resettable PhoneWorld mock-app equivalent), including state that must persist and transfer correctly across multiple apps in cross-app workflows. | set-of-independent | yes | none stated | 0 |
| ScaleWoB
ScaleWoB: Guiding GUI Agents with Coding Agents via Large-Scale Environmental Synthesis |
Web & GUI | 2026 | complete verifiable multi-step GUI tasks across mobile, desktop, or automotive/in-vehicle synthesized environments; succeed specifically on a distinguished long-horizon subset of tasks, which requires sustaining performance across more steps than the general task pool | Agent must track GUI state across a chain of interface actions long enough to be classified into the 'long-horizon subset', where performance drops sharply relative to the general task pool, implying more state must be carried across more steps. | hierarchical | yes | none stated | 10.25/mo |
| SentinelBench
SentinelBench: A Benchmark for Long-Running Monitoring Agents |
Web & GUI | 2026 | continuously monitor a live web environment (email/calendar/finance/professional-networking/entertainment) for a scripted external event; recognize the moment an event makes progress possible and act promptly; avoid excessive/wasteful actions (continuous polling/refreshing) while waiting | The agent must track whether the monitored page state has changed to reflect the awaited event, its own resource expenditure (tool calls/tokens) over the monitoring window, and elapsed time relative to the task's (possibly stretched) time budget. | sequential-chain | yes | each of the 100 tasks is designed to be achievable within a default 10-minute window; a speed_factor parameter (default 1.0) can stretch tasks to much longer durations (e.g. at speed_factor 0.25, tasks may require as long as 40 minutes)wall-clock-minutes | 20.67/mo |
| unnamed mobile GUI benchmark (paper introduces the ATMem method)
What Memory Do GUI Agents Really Need? From Passive Records to Active Task-Driving States |
Web & GUI+ Assistants & memory | 2026 | act on every list entry that satisfies the given instruction (positive matches); reject/skip every list entry that violates the instruction's constraints (negative matches) | The agent must track which near-identical list entries it has already acted on vs. still pending, and correctly apply the instruction's inclusion/exclusion constraints to each entry across a long trajectory. | set-of-independent | yes | none stated | 31.0/mo |
| AndroidLens
AndroidLens: Long-latency Evaluation with Nested Sub-targets for Android GUI Agents |
Web & GUI | 2025 | complete nested sub-targets within a single long-latency mobile task; satisfy multi-constraint, multi-goal task requirements drawn from 38 real-world domains; make measurable milestone-level progress even when full task success is not reached | Agent must track which nested sub-targets have been completed, tolerate environmental anomalies, and retain long-term memory of earlier steps across an average of more than 26 steps per task. | hierarchical | yes | >26 (average, 'more than 26')agent-steps | 20.22/mo |
| AndroidLH (via Mirage-1)
Mirage-1: Augmenting and Updating GUI Agent with Hierarchical Multimodal Skills |
Web & GUI | 2025 | complete real-world long-horizon, multi-app Android task scenarios using previously acquired hierarchical skills; correctly apply execution skills, core skills, and meta-skills at the appropriate level of abstraction across a long-horizon task | The agent must track which level of its hierarchical skill structure (execution, core, meta) is currently relevant, maintain state as it moves across multiple apps in a long-horizon scenario, and bridge the offline-to-online domain gap without losing track of previously acquired skills. | hierarchical | yes | none stated | 130.87/mo |
| BELA (Benchmark for Experiential Learning and Active exploration)
Benchmarking In-context Experiential Learning Through Repeated Product Recommendations |
Web & GUI+ Assistants & memory | 2025 | elicit unknown customer preferences through questions within a single recommendation interaction (episode); tailor questioning/recommendation strategy based on patterns observed across multiple prior episodes (customers/products) | The agent must track what it has learned about the current customer's preferences within an episode (turn-by-turn), and must also track/aggregate patterns across multiple prior episodes to adapt its questioning/recommendation strategy over time. | sequential-chain | yes | none stated | 20.2/mo |
| CAPBench
MobileSteward: Integrating Multiple App-Oriented Agents with Self-Evolution to Automate Cross-App Instructions |
Web & GUI | 2025 | associate and sequence sub-tasks across multiple mobile apps per a cross-app instruction; assign each associated sub-task to the correct app-oriented StaffAgent; avoid error propagation and information loss across the multi-step, multi-app execution; complete each of the 500 cross-app instructions spanning 14 apps in 6 categories | The centralized StewardAgent must track the scheduling graph of inter-app task associations, information flow between StaffAgents, and self-evolving memory of past executions to avoid repeating errors across a cross-app instruction. | DAG-with-precedence | no | none statedother:not-stated | 241.26/mo |
| ColorBench
ColorBench: Benchmarking Mobile Agents with Graph-Structured Framework for Complex Long-Horizon Tasks |
Web & GUI | 2025 | complete a single-app mobile task via one of multiple valid GUI action paths; complete a cross-app mobile task requiring coordination across multiple applications; reach subtask-level completion milestones within a longer task | Agent must track which subtasks have been completed, which of several valid paths it is following through the task's state graph, and avoid known error paths, across an average of more than 13 steps. | DAG-with-precedence | yes | >13 (average, 'over 13 steps')agent-steps | 80.73/mo |
| DeepShop
DeepShop: A Benchmark for Deep Research Shopping Agents |
Web & GUI | 2025 | satisfy multiple product-attribute constraints in a single shopping query; apply the correct search filters specified or implied by the query; apply the correct sorting preference specified or implied by the query; achieve overall shopping-task success across easy/medium/hard complexity tiers | Agent must track which product attributes, filters, and sorting preferences the query requires, and verify each is correctly reflected in its final shopping actions/results. | set-of-independent | yes | none stated | 422.8/mo |
| Mobile-Eval-RAG
Mobile-Agent-RAG: Driving Smart Multi-Agent Coordination with Contextual Knowledge Empowerment for Long-Horizon Mobile Automation |
Web & GUI+ Multi-agent orgs | 2025 | complete a cross-application mobile-automation task requiring both high-level plan steps and precise low-level UI operations; correctly execute app-specific atomic UI actions aligned with the current subtask | Agent must track which app/subtask it is currently operating in, the current step of the high-level plan, and precise UI state needed for accurate atomic actions across multiple apps in one task. | hierarchical | yes | none stated | 20.2/mo |
| MobileWorld
MobileWorld: Benchmarking Autonomous Mobile Agents in Agent-User Interactive and MCP-Augmented Environments |
Web & GUI | 2025 | complete long-horizon, cross-application mobile workflows spanning up to 20 applications; handle vague user instructions and hybrid tool usage; coordinate agent-user interaction and MCP-augmented tool calls mid-task | Agent must track cross-application state across nearly twice as many completion steps on average (27.8 vs. 14.3) as AndroidWorld, while also handling user-interaction requests and MCP-tool-call state. | sequential-chain | no | 27.8 (vs. 14.3 in AndroidWorld)agent-steps | 586.44/mo |
| MVISU-Bench
MVISU-Bench: Benchmarking Mobile Agents for Real-World Tasks by Multi-App, Vague, Interactive, Single-App and Unethical Instructions |
Web & GUI | 2025 | complete multi-app instructions requiring cross-app subgoal coordination; clarify vague/underspecified user instructions before acting; handle interactive instructions requiring mid-task clarification; complete single-app instructions; recognize and appropriately refuse or handle unethical instructions | The agent must track task state across multiple mobile apps for Multi-App instructions, detect ambiguity requiring clarification for Vague/Interactive instructions, and recognize when an instruction should be refused for Unethical instructions. | other (singleton) | no | none stated | 70.54/mo |
| NaturalGAIA
NaturalGAIA: A Verifiable Benchmark and Hierarchical Framework for Long-Horizon GUI Tasks |
Web & GUI | 2025 | decompose a natural, non-linear human GUI intent into a structured Task Topology of atomic sub-tasks; dynamically schedule sub-tasks across heterogeneous agents via context evolution; execute each atomic sub-task with precision via hybrid visual-structural perception; achieve a high Weighted Pathway Success Rate across the full causal pathway | The manager must dynamically track the Task Topology of atomic sub-tasks, schedule them across heterogeneous agents, and evolve shared context to bridge information gaps between dependent steps, assessed via a hierarchical success/error-attribution framework. | DAG-with-precedence | no | none statedother:not-stated | 0 |
| RealWebAssist
RealWebAssist: A Benchmark for Long-Horizon Web Assistance with Real-World Users |
Web & GUI | 2025 | correctly follow each instruction in a sequence of real, sequentially-issued user instructions across multiple websites; reason about the true intent behind ambiguous instructions; keep track of the user's mental state and user-specific routines as instructions evolve over the session; ground each intended task to the correct GUI element on the current website | Agent must retain a running model of the user's intent, mental state, and personal routines across a long sequence of instructions issued over multiple websites, using this history to correctly interpret each new (sometimes ambiguous) instruction. | sequential-chain | unclear | none stated | 261.53/mo |
| ShoppingBench
ShoppingBench: A Real-World Intent-Grounded Shopping Benchmark for LLM-based Agents |
Web & GUI | 2025 | apply vouchers correctly to a purchase; manage a budget across a shopping session; find and select from multi-product sellers matching a grounded intent; satisfy increasingly challenging levels of grounded shopping intent end-to-end | Agent must track running spend/budget, applied vouchers, and multi-seller product-matching requirements across a session, operating within a sandbox of over 2.5 million real-world products. | other (singleton) | no | none stated | 272.08/mo |
| ShoppingComp
ShoppingComp: Are LLMs Really Ready for Your Shopping Cart? |
Web & GUI | 2025 | retrieve products satisfying many simultaneous discovery constraints; generate an expert-level report on the retrieved products; make a safety-critical purchase decision (e.g. flag unsafe product usage) | The agent must track which of the multiple product-discovery constraints have been satisfied, the evidence supporting its report, and any identified safety hazards, while operating in an open-world product catalogue. | set-of-independent | no | none statedother:not-stated | 80.8/mo |
| UI-NEXUS
Atomic-to-Compositional Generalization for Mobile Agents with A New Benchmark and Scheduling System |
Web & GUI | 2025 | complete compositional mobile operations that concatenate multiple atomic tasks (Simple Concatenation); transition context correctly across sub-tasks within one compositional task (Context Transition); perform a deep multi-step drill-down within an app or workflow (Deep Dive) | The agent must track which atomic subtasks within a compositional task it has already completed, the context state carried between apps/screens, and overall progress toward each of the three compositional operation types. | hierarchical | yes | average optimal step count of 14.05 (over 100 interactive task templates)actions | 120.8/mo |
| VeriWeb
VeriWeb: Verifiable Long-Chain Web Benchmark for Agentic Information-Seeking |
Web & GUI | 2025 | ensure comprehensive information coverage across breadth- and depth-oriented multi-hop web searches; complete a sequence of interdependent, individually-verifiable subtasks within one long-chain web task; maintain consistent context tracking across a long information-seeking chain | Agent must track and verify each subtask-level answer as it progresses through a long chain of interdependent web-search subtasks, ensuring comprehensive information coverage and consistent context tracking across the chain. | sequential-chain | yes | none stated | 100.77/mo |
| WebChoreArena
WebChoreArena: Evaluating Web Browsing Agents on Realistic Tedious Web Tasks |
Web & GUI+ Information seeking | 2025 | retain and retrieve large amounts of information gathered from many observations (Massive Memory); perform precise mathematical/quantitative reasoning over collected information (Calculation); keep information consistent while tracking it across multiple webpages over the course of a task (Long-Term Memory) | Agent must accumulate and correctly recall large amounts of information across many webpages, perform arithmetic over it, and keep facts consistent across the whole task rather than a single page view. | sequential-chain | unclear | none stated | 281.87/mo |
| WebMall
WebMall - A Multi-Shop Benchmark for Evaluating Web Agents |
Web & GUI | 2025 | find a specific product across four simulated online shops; perform price comparisons across shops; identify suitable substitutes or compatible products (advanced search); add items to a cart and complete checkout in the correct shop(s) | Agent must retain product/price information gathered from each of the four shops it has visited so far in order to correctly compare, substitute, or complete a checkout later in the same task. | sequential-chain | unclear | none stated | 141.08/mo |
| WebPRM Collection / WebRewardBench
Web-Shepherd: Advancing PRMs for Reinforcing Web Agents |
Web & GUI | 2025 | navigate a website across many sequential steps toward a stated goal; at each step, satisfy sub-goals encoded in an annotated checklist (process-level correctness); complete the overall web-navigation task successfully (episode-level correctness) | The reward model must track which checklist sub-goals have already been satisfied at each step of a trajectory so it can assess the process-level (not just outcome-level) quality of a web-navigation trajectory. | sequential-chain | yes | trajectory lengths vary by difficulty: median ~5 steps (easy), ~9 steps (medium), ~20 steps (hard), with some trajectories exceeding 40 stepsactions | 291.81/mo |
Built from 4,327 candidates, 2,191 of them judged, then extracted per paper. Not exhaustive: at least ~131 further in-scope artifacts are estimated missing (a lower bound; the honest range is ~130–320). You should not read a gap here as evidence that no such benchmark exists. The horizon column reflects what each paper foregrounds — 276 of the 334 “none stated” determinations were made from the abstract, 56 confirmed against full text. Full data: catalog.csv · catalog.json · the full report.