Benchmarks that make LLM agents juggle many goals — the catalogue

All 480 benchmarks, datasets, environments and scenario suites published in 2025–2026 whose settings require an agent to pursue and track multiple goals and subgoals. Software-engineering and other code-centric domains are excluded by design. Complete as-of 2026-09-14; the field moves fast, so treat it as stale after 6–8 weeks.

Reading the table. Goals to track and must keep track of are extracted per paper from its own text. Horizon shows the paper's own figure in its own unit — and none stated where the paper never quantifies one, which is the case for 334 of 480 entries. Credit records whether subgoal-level partial credit exists. Cites is the Semantic Scholar citation count with a per-month rate beside it — with 318 of 480 entries published in 2026, this column mostly measures age, not quality, and 126 entries sit at zero. Every name links to the paper.

OS & computer use24 artifacts
CareFlow (benchmark) / CarePilot (agent framework)
CarePilot: A Multi-Agent Framework for Long-Horizon Computer Task Automation in Healthcare
OS & computer use+ Healthcare 2026 complete each of 8-24 consecutive GUI decisions/actions required by one long-horizon clinical software workflow (e.g. DICOM viewer, EHR, lab information system); correctly ground next-action predictions in the current visual interface and system state throughout the workflow The agent must track the current visual interface state and prior decisions across 8-24 consecutive steps of a clinical software workflow, using dual-memory (long-term and short-term experience) to predict the next semantic action correctly. sequential-chain yes each task comprises 8-24 consecutive decisions/stepsagent-steps 61.0/mo
ChainWorld
ChainWorld: Composing Long-Horizon Desktop Workloads from Atomic OSWorld Tasks
OS & computer use 2026 complete a chain of 2 to 4 sequentially composed atomic OSWorld desktop tasks while sustaining state across all of them; under single-turn evaluation, complete the whole chain from one combined prompt; under multi-turn evaluation, complete each task as it is revealed one at a time while retaining session state Agent must sustain desktop application/file state across multiple chained objectives and, in multi-turn mode, manage session continuity as tasks are revealed one at a time. sequential-chain yes two to fourother:atomic-tasks-per-chain 0
CutVerse
CutVerse: A Compositional GUI Agents Benchmark for Media Post-Production Editing
OS & computer use+ Web & GUI 2026 complete a long-horizon, compositional media-editing task grounded in an authentic editing workflow using one of 7 professional applications (e.g., Premiere Pro, Photoshop); correctly execute tightly coupled, dense multimodal interaction sequences within the application's GUI Agent must track compositional GUI state (selected layers/clips/tools) across a tightly coupled, dense-interface interaction sequence throughout a long-horizon editing task. sequential-chain yes none stated 0
DeskCraft
DeskCraft: Benchmarking Desktop Agents on Professional Workflows and Human-in-the-Loop Collaboration
OS & computer use+ Web & GUI 2026 complete long-horizon professional creative/engineering workflows (design, video, audio, 3D creation) requiring over 50 execution steps; proactively seek necessary information from the user under uncertainty (agent-initiated clarification); correctly handle user-initiated interruptions during execution and incorporate post-turn feedback after signaling completion The agent must track its own multi-step progress (over 50 execution steps for long-horizon tasks) within professional creative/engineering software, while also tracking pending clarification needs, handling user interruptions mid-execution, and incorporating post-completion feedback into further revisions. sequential-chain yes long-horizon tasks require over 50 execution stepsagent-steps 20.67/mo
DevicesWorld
DevicesWorld: Benchmarking Cross-Device Agents in Heterogeneous Environments
OS & computer use 2026 acquire information on one device (e.g. phone) needed to complete a task; process/transform that information on a different device (e.g. desktop); deliver or display the final result on yet another device; satisfy cross-device dependencies and rule-based verifiers across the whole task The agent must track which device holds which piece of needed information, what has already been acquired/processed/delivered across devices, and whether all task conditions (verified from device states and generated files) are jointly satisfied. DAG-with-precedence yes each task permits at most 50 interaction stepsactions 0
GUITestScape
GUITestScape: Towards Open-set Evaluation on Exploratory GUI Testing
OS & computer use+ Web & GUI 2026 autonomously navigate an app to discover defects with no predefined test script; distinguish and separately diagnose interaction defects versus display defects; decompose the testing trajectory into independently diagnosable capabilities rather than a single end-state judgment The agent must track which parts of an app it has already explored and which candidate defects (interaction or display) it has already investigated, to continue open-ended exploration without redundant re-testing or missing an area. open-ended yes none stated 0
HeraBench
Beyond Global Replanning: Hierarchical Recovery for Cross-Device Agent Systems
OS & computer use 2026 decompose a cross-device task and assign subtasks across heterogeneous devices via unified API-CLI-GUI execution; recover from injected device-local strategy failures without escalation; escalate to orchestrator-level global replanning when a failure exceeds device-local recovery scope; complete an end-to-end cross-device workflow over Linux and Android devices The system must track per-device execution-strategy state, a compact cross-layer failure abstraction distinguishing device-local from global failure scope, and overall workflow progress across Linux and Android devices. hierarchical yes none stated 0
JarvisGUI
JarvisGUI: Towards Cross-Device GUI Agents with Dynamic Task Composition
OS & computer use+ Web & GUI 2026 transfer intermediate results correctly between two or more heterogeneous devices/platforms; maintain shared state consistently across Android, Windows, and Ubuntu platforms; compose a multi-step, cross-device workflow from input-output-typed GUI sub-tasks; complete a workflow requiring four or more chained subtasks despite compounding dependency-tracking difficulty The agent must track shared state and intermediate results as they transfer across heterogeneous devices/platforms, verifying that dependency steps (e.g. file transfer, renaming) have actually executed rather than being falsely treated as completed. DAG-with-precedence no none statedother:not-stated 0
MacAgentBench
MacAgentBench: Benchmarking AI Agents on Real-World macOS Desktop
OS & computer use+ Web & GUI 2026 complete a multi-application task on real macOS desktop software using both GUI and CLI interaction; reach each of several fine-grained, capability-annotated checkpoints within a multi-application task Agent must track partial progress across multiple checkpoints within a multi-application task, since 'models with similar Pass@1 can differ substantially in sub-goal completion' according to the fine-grained metrics. hierarchical yes none stated 20.67/mo
MedCUA-Bench
MedCUA-Bench: A Screenshot-Only Benchmark for Clinical Computer-Use Agents
OS & computer use+ Healthcare 2026 complete clinical computer-use scenarios across 10 medical domains; satisfy paired intent-level and step-level goals within each task; avoid violations across five clinical safety dimensions while completing the task Agent must track progress toward the high-level clinical intent, the individual UI steps needed to realize it, and continuously check five clinical safety dimensions across the task. hierarchical yes none stated 10.33/mo
MyPCBench
MyPCBench: A Benchmark for Personally Intelligent Computer-Use Agents
OS & computer use+ Assistants & memory 2026 complete a personal-assistant task requiring the agent's own accumulated context/history; operate correctly across 17 simulated real-world web applications seeded for one canonical persona; coordinate actions that span many applications on the same Linux desktop; sustain a long trajectory without abandoning early or looping unproductively The agent must track the canonical persona's context, historical data, and logged-in-account state across a full Linux desktop stack, coordinating GUI and bash actions across many applications without unproductive step looping. DAG-with-precedence no mean 22-31 steps (GPT family, Sonnet); mean 52-85 steps (Opus, Qwen models)agent-steps 10.33/mo
OmniGUIRewardBench (OGRBench) / OS-Themis
OS-Themis: A Scalable Critic Framework for Generalist GUI Rewards
OS & computer use+ Web & GUI 2026 decompose a GUI trajectory into verifiable milestones; audit the evidence chain for each milestone before rendering a final reward verdict; produce a scalable, accurate outcome reward usable for RL training or trajectory filtering The critic must track which milestones in a trajectory have been reached, the evidence supporting each, and audit the full evidence chain before issuing a final reward — i.e. it tracks reward-relevant progress rather than the acting agent's own task state. hierarchical yes none stated 81.33/mo
OS-Marathon
OS-Marathon: Benchmarking Computer-Use Agents on Long-Horizon Repetitive Tasks
OS & computer use 2026 complete each of 242 long-horizon, repetitive computer-use tasks across 2 domains (e.g., processing expense reports from receipts, entering grades from exam papers); correctly repeat the same structured sub-workflow logic across many data items within one task; learn workflow logic from a few-shot condensed demonstration and generalize it to larger, unseen data collections The agent must track its position within a long, repetitive workflow (e.g., which receipt or exam paper it is currently processing), correctly apply the learned sub-workflow logic consistently, and avoid drift or errors accumulating across many repeated sub-tasks. sequential-chain yes none stated 51.67/mo
OS-Marathon
OS-Marathon: Benchmarking Computer-Use Agents on Vast-Horizon, Repetitive Tasks
OS & computer use 2026 complete many recurring per-instance sub-workflows within a vast-horizon repetitive task (e.g., process each receipt in a stack); sustain correct execution as length scales with data volume; learn/personalize the recurring sub-workflow logic from a single human demonstration (GraphDemo) Agent must correctly and repeatedly apply the same recurring sub-workflow logic across many per-instance items, with execution length scaling with the volume of data to process. set-of-independent unclear none stated 0
OSWorld 2.0
OSWorld2.0: Benchmarking Computer Use Agents on Long-Horizon Real-World Tasks
OS & computer use 2026 complete each of 108 realistic end-to-end long-horizon computer-use workflows; perform cross-source reasoning across authentic input artifacts; infer implicit state and recover hidden state the task depends on; handle streaming interaction and dynamic environment changes mid-task; operate under safety-sensitive execution constraints (audited separately) The agent must track constraints, cross-source information, implicit/hidden state, and mid-task updates continuously across an average of 318 tool calls (up to a 500-step budget) per task, rather than resolving everything from the initial instruction alone. sequential-chain yes median about 1.6 hours per task for human users; average of 318 tool calls per task with Claude Opus 4.7 using maximum thinking (vs about 30 in OSWorld 1.0); 500-step completion budgettool-calls 155.0/mo
WindowsWorld
WindowsWorld: A Process-Centric Benchmark of Autonomous GUI Agents in Professional Cross-Application Environments
OS & computer use+ Web & GUI 2026 complete each of a task's average 5.0 sub-goals across up to 17 desktop applications; perform conditional judgment and reasoning that spans 3 or more applications; coordinate a workflow across multiple applications for 78% of tasks that are inherently multi-application The agent must track progress across a sequence of sub-goals spanning multiple desktop applications, carrying forward state and conditional-judgment results between applications, within a per-difficulty-level step budget (15-40 steps). sequential-chain yes max step budget 15 (L1) / 25 (L2) / 40 (L3) / 20 (L4); avg minimum action steps 9.67-27.81 by levelagent-steps 81.6/mo
AgentSynth
AgentSynth: Scalable Task Generation for Generalist Computer-Use Agents
OS & computer use 2025 complete long-horizon computer-use tasks composed of a controllable number of simple subtasks (difficulty levels 1-6); generalize from simple generation-time subtasks that become significantly harder once composed Agent must carry state correctly across a chain of composed subtasks, since success rate drops steeply as the number of chained subtasks increases from difficulty level 1 to 6. sequential-chain no none stated 322.13/mo
KGCE
KGCE: Knowledge-Augmented Dual-Graph Evaluator for Cross-Platform Educational Agent Benchmarking with Multimodal Language Models
OS & computer use+ Web & GUI 2025 complete a school-specific-software task correctly on Windows, Android, or across both platforms in coordination; satisfy each of the multiple sub-goals a complex task is decomposed into (e.g. one complex task decomposed into five subtasks), each independently verified The agent must track completion status of each decomposed sub-goal (via the dual-graph evaluation framework) across potentially multiple platforms and school-specific private-domain software, whose structural specifics are not otherwise well understood by general-purpose agents. DAG-with-precedence yes none stated 0
MMBench-GUI
MMBench-GUI: Hierarchical Multi-Platform Evaluation Framework for GUI Agents
OS & computer use+ Web & GUI 2025 understand GUI content and ground specific elements accurately (Element Grounding); automate a full task end-to-end within one application/platform (Task Automation); collaborate across multiple tasks/applications requiring cross-platform generalization (Task Collaboration) Agent requires long-context memory across many actions, tracking of task-planning state, and long-term reasoning to sustain grounded, efficient action sequences without redundant steps across the hierarchy. hierarchical yes none stated 614.36/mo
Mobile-Eval-E
Mobile-Agent-E: Self-Evolving Mobile Assistant for Complex Tasks
OS & computer use+ Web & GUI 2025 break a complex mobile task into subgoals via a Manager agent; execute fine-grained low-level actions per subgoal via an Operator agent; verify and correct action errors via an Action Reflector; aggregate information across steps via a Notetaker; complete long-horizon, multi-app mobile interactions The system must track the Manager's subgoal plan, the Notetaker's aggregated information, self-evolved long-term Tips/Shortcuts memory, and cross-app interaction state across a single complex mobile task. hierarchical yes none stated 1346.7/mo
OmniBench / OmniEval
What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities
OS & computer use+ Web & GUI 2025 complete a graph-structured task composed of multiple synthesized subtasks of controllable complexity; satisfy subtask-level correctness at each node of the task graph, not just the final outcome; demonstrate each of 10 evaluated virtual-agent capabilities Agent must track progress through the task graph's subtask nodes and correctly compose primitive actions across the graph to satisfy subtask-level and graph-based metrics. DAG-with-precedence yes none stated 90.6/mo
PC-Eval
PC-Agent: A Hierarchical Multi-Agent Collaboration Framework for Complex Task Automation on PC
OS & computer use+ Multi-agent orgs 2025 decompose complex user instructions into Instruction-Subtask-Action levels; track progress across interdependent subtasks via a Progress agent; make step-by-step decisions via a Decision agent; provide timely bottom-up error feedback and adjustment via a Reflection agent; complete each of 25 real-world complex PC instructions The multi-agent system must track subtask decomposition and progress (via the Progress agent), current decision state (via the Decision agent), and bottom-up error feedback (via the Reflection agent) across intra- and inter-app workflows on a PC. hierarchical yes none stated 392.05/mo
SEAgent (OS-World novel software environments)
SEAgent: Self-Evolving Computer Use Agent with Autonomous Learning from Experience
OS & computer use 2025 autonomously master a novel, unfamiliar software environment through experiential trial-and-error; progressively tackle auto-generated tasks organized from simple to complex (curriculum); assess step-wise trajectory correctness via a World State Model; integrate individual specialist experience into a stronger generalist computer-use agent The agent must track its step-wise trajectory quality (via the World State Model), its progress along an increasingly difficult auto-generated curriculum, and accumulated experiential insights to be integrated from specialist to generalist training. hierarchical yes none stated 614.69/mo
UI-CUBE
UI-CUBE: Enterprise-Grade Computer Use Agent Benchmarking Beyond Task Accuracy to Operational Reliability
OS & computer use+ Business & enterprise 2025 complete simple UI interactions (136 tasks); complete complex copy-paste workflows spanning multiple applications (50 tasks); complete complex enterprise application scenarios (40 tasks); maintain operational reliability (not just functional correctness) under systematic interface variation and multi-resolution testing Agent must track state across chained multi-step workflows (e.g., what was copied and where it must be pasted), interface variations, and multiple screen resolutions, verifying success via application-state validation. hierarchical yes none stated 30.3/mo
Business & enterprise105 artifacts
Evaluating Deep Research Agents on Expert Consulting Work: A Benchmark with Verifiers, Rubrics, and Cognitive Traps
Evaluating Deep Research Agents on Expert Consulting Work: A Benchmark with Verifiers, Rubrics, and Cognitive Traps
Business & enterprise 2026 produce a multi-document, decision-grade consulting deliverable from SME-authored prompts; pass deterministic binary verifiers (mean 14.9 per task); satisfy each of a five-criterion SME rubric (Data Integrity, Analytical Rigor, Relevance & Focus, Execution Precision, Format & Deliverability); avoid embedded cognitive traps (human-error mimicry, deterministic precision traps) that penalize surface-pattern reasoning The agent must track which of the ~14.9 verifiers per task it has satisfied, avoid embedded cognitive traps requiring reconciliation against context, and keep its final deliverable consistent with all five rubric criteria simultaneously. set-of-independent yes none statedother:not-stated 10.25/mo
50-task hotel expense benchmark (Dynamics 365 F&O)
Less Context, Better Agents: Efficient Context Engineering for Long-Horizon Tool-Using LLM Agents
Business & enterprise 2026 itemize every line item of a hotel expense report completely and accurately using enterprise MCP tools; avoid context overflow / stale-state errors while repeatedly calling verbose enterprise tool APIs across the itemization workflow The agent must track already-itemized line items, running token/context budget, and must avoid stale or overflowing tool-call history while repeatedly invoking enterprise MCP tools to reach complete, accurate itemization. set-of-independent yes full-context retention: 1,480,996 tokens and 14.56 hours per benchmark (50 tasks x 5 runs); pruning to last 5 tool calls: 535,274 tokens and 5.39 hours; pruning + summarization: 553,374 tokens and 5.79 hourswall-clock-hours 10.33/mo
AD-Bench
AD-Bench: A Real-World, Trajectory-Aware Advertising Analytics Benchmark for LLM Agents
Business & enterprise 2026 answer a real user marketing-analysis request via multi-round, multi-tool collaboration; maintain trajectory coverage consistent with an expert tool-call trajectory; produce answers that remain correct/consistent with a continuously evolving production advertising platform (dynamic ground truth); succeed across three stratified difficulty levels (L1-L3) requiring increasing multi-round, multi-tool collaboration The agent must track which professional tools it has called and in what order across multiple rounds, since evaluation jointly measures end-to-end answer correctness (Pass@k) and how well its trajectory covers the expert reference trajectory, especially as this compounds at higher difficulty levels. sequential-chain yes none stated 30.43/mo
AeroCopilotBench (ACOE)
AeroCopilotBench: A Two-Tier Benchmark for Evaluating LLM Agents as Aviation Copilots in an Interactive Virtual Cockpit Environment
Business & enterprise 2026 diagnose faults and operate aircraft systems through standardized tool interfaces during emergency/abnormal procedures; achieve all task goal-conditions for each of 73 Tier-2 emergency/abnormal tasks; avoid violating any hard safety constraint while progressing toward goals Agent must interpret cockpit state, diagnose faults, and operate aircraft systems while simultaneously tracking goal-condition progress and hard safety-constraint compliance across the executable procedure. DAG-with-precedence yes none stated 0
Agentic ERP
Agentic ERP: Multi-Agent Large Language Model Architecture for Autonomous Enterprise Resource Planning
Business & enterprise+ Multi-agent orgs 2026 execute end-to-end business workflows across ERP functional boundaries (e.g. procurement, inventory, crisis response); resolve cross-functional crisis tasks via a Planner-Executor-Reflector-Responder orchestration; sustain a full simulated year of ERP operation while avoiding stockouts, compared against rule-based RPA and no-intervention baselines The agent(s) must track the structured enterprise state (inventory, demand stream, crisis conditions) across a full simulated year, coordinate across role-aligned agents via externalised grading criteria and sprint contracts, and avoid accumulating stockouts as the rule-based baseline does. hierarchical yes a 365-day agent-in-the-loop simulation (a simulated year of ERP operation)simulated-days 0
AgenticPay
AgenticPay: A Multi-Agent LLM Negotiation System for Buyer-Seller Transactions
Business & enterprise+ Multi-agent orgs 2026 negotiate to reach agreement given private constraints and product-dependent valuations; maximize feasibility, efficiency, and welfare across a negotiation; successfully complete each of 110+ tasks ranging from bilateral bargaining to many-to-many markets Agents must track their own private constraints/valuations, the evolving history of natural-language offers and counteroffers across negotiation rounds, and, in many-to-many markets, the state of other concurrent negotiations. sequential-chain yes none stated 131.86/mo
AgenticVBench
AgenticVBench: Can AI Agents Complete Real-World Post-Production Tasks?
Business & enterprise 2026 complete each of 100 agentic post-production tasks across 4 task families; compose capabilities across text, image, audio, and video understanding within one task; plan and execute long-horizon multi-step production workflows using appropriate tools The agent must track progress through a long-horizon, multi-modal production workflow and correctly sequence tool use, since tasks are constructed from real production workflows contributed by industry experts and evaluated jointly by programmatic verifiers and expert rubrics. other (singleton) yes none stated 30.75/mo
Agents' Last Exam (ALE)
Agents' Last Exam
Business & enterprise 2026 complete a long-horizon, economically valuable, real-world professional task with a verifiable outcome, drawn from a specific occupational sub-field; achieve sustained performance across the task rather than a single-shot correct answer, particularly on the hardest ('last-exam') tier The agent must sustain performance on long-horizon, economically valuable real-world tasks whose outcomes are verifiable, across an occupational taxonomy spanning 55 sub-fields and 13 industry clusters, with the hardest tier proving far from saturated (average full pass rate below 1%). set-of-independent yes none stated 124.0/mo
APEX-Agents
APEX-Agents
Business & enterprise 2026 execute a long-horizon, cross-application task created by investment-banking analysts, management consultants, or corporate lawyers; navigate realistic work environments containing files and tools to produce the required deliverable Agent must track which files and tool states it has already produced/consulted while navigating a cross-application work environment, using rubrics and gold outputs to determine success (Pass@1). sequential-chain unclear none stated 202.5/mo
AutomationBench
AutomationBench
Business & enterprise 2026 discover the relevant REST API endpoints needed for a cross-application workflow (CRM, inbox, calendar, messaging, etc.); follow layered business-policy rules while writing data; navigate environments containing irrelevant or misleading records without being derailed; get correct data into the right systems by the end of the workflow (end-state correctness) Agent must track which endpoints it has discovered across multiple applications, which business-policy rules apply to each write, and whether records encountered are relevant or misleading, to ensure the correct final data lands in each system. DAG-with-precedence yes none stated 10.2/mo
BigFinanceBench
BigFinanceBench: A Workflow-Grounded Benchmark for Financial-Research Agents
Business & enterprise 2026 produce an auditable financial-research derivation (source choice, period/accounting definition, assumptions, calculation), not just a final answer; satisfy each of many independently checkable rubric points per item (36,241 points across 928 items); correctly complete each of 928 expert-authored open-ended financial-research tasks The agent must track and justify each auditable component of its derivation (data source, period/accounting definition, assumptions, calculation steps) so that an independent rubric can check each step, not just the final numeric answer. hierarchical yes none stated 51.67/mo
BlueFin
BlueFin: Benchmarking LLM Agents on Financial Spreadsheets
Business & enterprise 2026 synthesize new spreadsheet content/formulas from source data; manipulate existing financial spreadsheet workbooks (edits, restructuring); comprehend and answer questions about spreadsheet content; satisfy each of up to 3,225 granular rubric criteria across 131 tasks The agent must track and satisfy many granular, LM-judge-scored rubric criteria across synthesis, manipulation, and comprehension actions within one workbook, maintaining correctness as the workbook's formulas/state evolve ('dynamic correctness'). hierarchical yes tasks require at least 45+ minutes of work for a human analyst to complete from scratchhuman-expert-hours 10.25/mo
Business Arena
Business Arena: Benchmarking LLM Agents in a Realistic Marketplace
Business & enterprise 2026 infer business opportunities from partial market signals; commit capital under uncertainty when buying from suppliers; set/adjust pricing to sell to buyers profitably; satisfy regulatory obligations before trading legally; sustain profitable operation of a cross-border shop over a long horizon The agent must track capital, delayed and coupled consequences of sourcing/pricing/recovery decisions, market conditions calibrated from real data, and regulatory obligations across a long-horizon cross-border trading operation. open-ended no none stated 0
CEO-Bench
CEO-Bench: Can Agents Play the Long Game?
Business & enterprise+ Games & IF 2026 operate a startup for 500 days, managing pricing, marketing, budgeting, and other business aspects; acquire information from noisy, interconnected business databases and translate it into strategy; adapt strategy to a changing world/market over the run; orchestrate many interdependent business decisions toward a coherent long-term financial goal The agent must track noisy, interconnected business databases (churn regimes, billing timing, customer losses, projected cash) and coordinate many interdependent decisions via code across a 500-day simulated run. open-ended yes 500simulated-days 31.0/mo
CEO-Bench
Can LLMs Be CEOs? Benchmarking Strategic Resource Reallocation with Multi-Role Agent Simulation
Business & enterprise 2026 integrate conflicting recommendations from four role-conditioned C-suite advisors (CFO, CTO, COO, CMO) into one allocation plan; redirect capital across business units under information asymmetry and organizational constraints; synthesize advice consistently across multiple rounds with temporal dependencies Agent (as CEO) must track and reconcile conflicting, privately-signaled recommendations from four C-suite advisors across a multi-round resource-reallocation process, remembering prior decisions for history-sensitive judgment. DAG-with-precedence yes none stated 10.33/mo
CHI-Bench (χ-Bench)
CHI-Bench: Can AI Agents Automate End-to-End, Long-Horizon, Policy-Rich Healthcare Workflows?
Business & enterprise+ Healthcare 2026 ground each decision in a large library of medical, insurance, and operational policy rules; play multiple roles within a single task, handing off between them; conduct multilateral multi-turn dialogs (e.g. peer-to-peer review, patient outreach) as intermediate workflow steps; drive a clinical case to a terminal status via tool calls and role-artifact writing The agent must track its current role and handoffs to other roles, policy compliance against a 1,290+ document handbook, and the clinical case's evolving status across a high-fidelity simulator of 20 apps, until it reaches a terminal status. DAG-with-precedence no none statedother:not-stated 61.5/mo
Claw-Eval-Live
Claw-Eval-Live: A Live Agent Benchmark for Evolving Real-World Workflows
Business & enterprise 2026 complete end-to-end units of work across software tools, business services, and local workspaces per release; pass controlled tasks with fixed fixtures/services/workspaces/graders reflecting current public workflow-demand signals; handle HR, management, and multi-system business workflows as well as local workspace repair Agent must produce verifiable execution traces, audit logs, and consistent service/workspace state across each end-to-end task, since grading uses deterministic checks when evidence is sufficient and structured LLM judging only for semantic dimensions. set-of-independent no none stated 132.6/mo
ClawMark
ClawMark: A Living-World Benchmark for Multi-Turn, Multi-Day, Multimodal Coworker Agents
Business & enterprise 2026 carry a professional coworker task forward correctly across multiple working days; detect and adapt to exogenous environment updates (new emails, calendar shifts, KB edits) that occur between turns; satisfy each of the task's deterministic Python checkers over post-execution service state (mean 15.4 checkers/task); coordinate consistent state across five stateful sandboxed services (filesystem, email, calendar, knowledge base, spreadsheet) The agent must track evolving state across five sandboxed services (filesystem, email, calendar, knowledge base, spreadsheet) over multiple working days, detecting and incorporating exogenous updates injected between turns, verified by up to 29 deterministic checkers per task. sequential-chain yes 2-6 turns per task (mean 3.6); one turn = one in-universe working daysimulated-days 153.0/mo
ClawsBench
ClawsBench: Evaluating Capability and Safety of LLM Productivity Agents in Simulated Workspaces
Business & enterprise 2026 complete single-service productivity tasks (e.g. Gmail, Slack, Calendar, Docs, Drive) correctly and safely; complete cross-service workflows spanning multiple mock services while maintaining consistent state; avoid unsafe/irreversible actions in safety-critical scenarios Agents must track persistent state across five mock services (Gmail, Slack, Calendar, Docs, Drive), recognize safety-critical constraints to avoid irreversible/unsafe actions, and coordinate behavior across services via a meta-prompt layer. DAG-with-precedence yes none stated 255.0/mo
CoffeeBench
CoffeeBench: Benchmarking Long-Horizon LLM Agents in Heterogeneous Multi-Agent Economies
Business & enterprise+ Multi-agent orgs 2026 maximize cumulative net income for one's own firm (coffee roaster) over a long simulated run; manage cash, inventory, and pricing on an ongoing basis; communicate and transact with other heterogeneous firms (farmers, roasters, retailers) to secure supply/demand The agent must track its own cash, inventory, and pricing state day by day over the 90-day simulation, plus the state of its ongoing communications/transactions with other firms, to maximize cumulative net income rather than any single day's outcome. sequential-chain yes 90-day simulationsimulated-days 62.0/mo
Company World Model dry-lab benchmark (45 retrospective cases)
Do AI-Native Biotechs Need Departments? Benchmarking Company World Models for AI-Driven Drug Development
Business & enterprise 2026 maintain and update a persistent asset-to-value state (Live Asset Value Record) across scientific, regulatory, BD, commercial, financial, and execution constraints; resolve 45 retrospective public-information decision cases with hidden outcomes under strict time cutoffs; achieve success via external BD deals, regulatory approval/launch, and revenue discipline The architecture must maintain a persistent, continuously updated asset-to-value state record across scientific, regulatory, BD, commercial, financial, and execution constraints via Deal/Approval/Revenue/Investment Arbiter loops. DAG-with-precedence no none stated 0
CoreCraft (EnterpriseBench)
EnterpriseBench Corecraft: Training Generalizable Agents on High-Fidelity RL Environments
Business & enterprise 2026 perform multi-step, domain-specific customer-support work across a simulated enterprise with 2,500+ entities across 14 entity types; satisfy all expert-authored rubric criteria for a given task using 23 available tools Agent must track entity state across the enterprise simulation and verify each expert-authored rubric criterion is satisfied before the task counts as solved. DAG-with-precedence yes none stated 71.0/mo
DocOps
DocOps: A Verifiable Benchmark for Autonomous Agents in Complex Document Operations
Business & enterprise 2026 perform atomic document-operation actions correctly (e.g. edits without destructive metadata corruption); complete escalating, more complex, more tightly coupled document-workflow tasks built from those atomic operations; maintain long-term state tracking and global document consistency across the workflow; correctly verify (not just superficially check) that a document edit was semantically achieved The agent must maintain long-term state tracking of the document's evolving structure/content and verify (rather than superficially assume) that each operation was semantically correct, to preserve global document consistency across highly coupled, long-range tasks. hierarchical yes none stated 10.5/mo
DORA (Disaster Operational Response Agent benchmark)
Can LLM Agents Respond to Disasters? Benchmarking Heterogeneous Geospatial Reasoning in Emergency Operations
Business & enterprise 2026 perform disaster perception from heterogeneous geospatial imagery; conduct spatial relational analysis over roads/population/facilities; plan rescue and evacuation operations; reason about temporal evolution of the disaster; synthesize a multi-modal operational report The agent must compose calls across a 108-tool MCP library over heterogeneous, multi-temporal geospatial data, tracking intermediate perception/analysis outputs and their correctness as the pipeline progresses through the five task dimensions. DAG-with-precedence no 3,500 tool-call steps total (across 515 tasks, ~6.8/task average)tool-calls 41.0/mo
DuMateBench
DuMateBench: Evaluating Autonomous Agents in Complex Real-World Workflows
Business & enterprise 2026 complete each of 200 real-session-derived tasks spanning 8 broad scenarios and 17 fine-grained capability categories; coordinate multiple capability categories within a single task; preserve and correctly use persistent configurations and workspace state carried over from prior interaction history; maintain performance under injected real-world environmental complexity (Insufficient, Unstable, Noisy conditions) The agent must track persistent configurations, workspace state, and pre-solution interaction history carried into each task, while coordinating multiple capability categories under injected Insufficient, Unstable, or Noisy environmental perturbations. set-of-independent no none stated 0
EcoGym
EcoGym: Evaluating LLMs for Long-Horizon Plan-and-Execute in Interactive Economies
Business & enterprise 2026 maintain the Vending sub-environment's profitability (net worth, income) as in Vending-Bench; acquire and retain freelance income under budgeted actions (Freelance sub-environment); sustain operational metrics such as DAU under partial observability (Operation sub-environment); maintain long-term strategic coherence across an effectively unbounded, budgeted-action economic horizon The agent must track budgeted actions, business-relevant outcome metrics (net worth, income, DAU), and partial observability/stochasticity across an effectively unbounded horizon of 1000+ steps (up to 365 day-loops) per evaluation. open-ended no 1000+ steps over 365 simulated day-loops (evaluation horizon)agent-steps 10.14/mo
EnterpriseArena
Can LLM Agents Be CFOs? Benchmarking Long-Horizon Resource Allocation in an Uncertain Enterprise Environment
Business & enterprise 2026 manage liquidity across the firm's operations; close financial books accurately on a regular cycle; gather costly signals about the macro/industry environment before acting; request equity or debt financing appropriately as conditions change; survive (avoid insolvency/failure) across the full 132-month horizon under shifting macroeconomic regimes The agent must track liquidity, capital structure (equity/debt), and signals about shifting macroeconomic/industry regimes month by month across the 132-month horizon, since consequences of earlier decisions are delayed and only become apparent later. DAG-with-precedence yes 132-month CFO simulationother:simulated-months 20.33/mo
EnterpriseArena (via the EnterpriseLab platform)
EnterpriseLab: A Full-Stack Platform for developing and deploying agents in Enterprises
Business & enterprise 2026 complete complex enterprise workflows spanning IT, HR, sales, and engineering domains; correctly invoke and sequence tool calls across 140+ tools exposed via Model Context Protocol across 15 applications; match frontier-model performance while running as a smaller (8B) privacy-preserving model; generalize/remain robust across diverse enterprise benchmarks (EnterpriseBench, CRMArena) The trained agent must track which of many interdependent enterprise tools/applications it has invoked and their resulting state, to complete complex multi-tool enterprise workflows without needing frontier-scale model capacity. DAG-with-precedence no none stated 10.17/mo
EnterpriseClawBench
EnterpriseClawBench: Benchmarking Agents from Real Workplace Sessions
Business & enterprise 2026 read heterogeneous workplace files relevant to a task; invoke the correct tools to accomplish a business objective; deliver a business artifact matching role-specific hard rules and semantic rubrics The agent must track which files it has read, what tools it has invoked, and whether its eventual delivered artifact satisfies the task's hard rules and semantic rubric, since harness-model combination, artifact delivery, and visual quality are all separately reported. sequential-chain yes none stated 10.33/mo
EnterpriseOps-Gym
EnterpriseOps-Gym: Environments and Evaluations for Stateful Agentic Planning and Tool Use in Enterprise Settings
Business & enterprise 2026 complete each of 1,150 expert-curated enterprise tasks across eight mission-critical verticals (e.g., Customer Service, HR, IT); plan and act correctly amid persistent state changes and strict access-control protocols; correctly refuse infeasible tasks rather than attempting unintended, potentially harmful side effects Agent must track persistent state changes across 164 database tables, access-control constraints, and whether the current task is feasible at all, to avoid unintended side effects from attempting an infeasible task. DAG-with-precedence yes none stated 111.83/mo
EntWorld
EntWorld: A Holistic Environment and Benchmark for Verifiable Enterprise GUI Agents
Business & enterprise+ Web & GUI 2026 complete each of 1,756 tasks spanning six representative enterprise domains (CRM, ITIL, ERP, etc.); operate under strict business logic constraints and high-density enterprise UIs; maintain precise, state-consistent information retrieval verified via SQL-based state-transition checks; complete synthesized long-horizon workflows reverse-engineered from database schemas The agent must track precise, state-consistent information across a high-density enterprise UI and the underlying database schema, since success is verified deterministically via SQL-based state-transition checks rather than visual matching. sequential-chain no none stated 50.62/mo
ERP-Bench (via the Anchor task-generation pipeline)
Anchor: Mitigating Artifact Drift in Agent Benchmark Generation
Business & enterprise 2026 satisfy explicit procurement-workflow constraints in a production-grade ERP system; satisfy explicit manufacturing-workflow constraints in a production-grade ERP system; reach the solver-certified fully optimal end-state solution for a long-horizon business task, not merely a feasible one The agent must track state-based verifier conditions tied to the end-state of a production-grade ERP system across a long-horizon procurement or manufacturing workflow, since reward depends solely on end-state business correctness rather than intermediate steps. other (singleton) no none stated 10.25/mo
ERPBench
ERPBench: Evaluating LLM Agents for Enterprise Decision-Making Across Competitive Market Ecologies
Business & enterprise 2026 make coupled decisions each round across pricing, production, procurement, inventory, and finance; maximize valuation/rank while competing against fixed rule-based opponents (Solo) or other evaluated LLM agents (Arena) in a shared market; sustain a coherent business strategy across all six rounds of the ERP simulation Agent must track its own accumulated valuation, inventory, and finance position across all six rounds, plus (in Arena mode) the shared market state shaped by five other competing LLM agents, to decide coupled pricing/production/procurement actions each round. sequential-chain yes 6turns 0
EvoEnv
The Agent's First Day: Benchmarking Learning, Exploration, and Scheduling in the Workplace Scenarios
Business & enterprise 2026 schedule and prioritize streaming tasks with varying priorities under a dynamic workload; actively seek out information to reduce hallucination before acting (prudent information acquisition); distill and reuse generalized strategies learned from earlier dynamically generated tasks in later ones (continuous evolution) The agent must track incoming streaming tasks and their priorities, its own accumulated exploration/experience for continual learning, and its confidence in acquired information to avoid hallucination, across a continuously evolving workplace scenario. hierarchical yes 50 dynamic scenarios, each with 2-6 task instances; observed usage up to ~90 steps and 232 tool calls for the highest-usage evaluated model (Gemini-3-Flash)actions 10.12/mo
FinEvo-Bench
FinEvo-Bench: A Longitudinal Benchmark for Self-Evolving Agents in Professional Financial Workflows
Business & enterprise 2026 complete each of 120 real-case-grounded financial tasks across 20 business scenes in six financial domains, following institution-provided procedures and constraints; retain and re-apply experience from earlier cases in a scene to later, substantively distinct cases sharing the same professional procedure (self-evolution); maintain financial-compliance quality across each task as scored by a manually reviewed rubric; sustain or improve quality/compliance across an interleaved, shuffled task stream over the longitudinal run The agent (self-evolving scaffold) must retain experience -- via memory, skill distillation, or both -- from earlier cases in the interleaved task stream and apply it to later, procedurally-related but factually distinct cases, while satisfying institution-provided professional-procedure constraints and financial-compliance requirements on each individual task. other (singleton) yes 120 real-case-grounded tasks per longitudinal stream (20 scenes x 6 cases each); three independently shuffled, globally interleaved task streamsepisodes 11.0/mo
FM-Bench
FM-Bench: A Benchmark for Long-Horizon Management with Competing Agents
Business & enterprise 2026 draft and manage a football squad within a fixed budget shared with rivals; trade players and negotiate contracts across the season; invest in facilities and youth development; set match lineups; maintain board confidence (avoid being fired) while maximizing a final accumulated score over 20 in-game years The agent must track squad composition, budget/cash flow, ongoing contract-renewal deadlines, facility/youth investments, and the board's confidence in it, across roughly 340-400 decision stops spanning 20 in-game years, since 'the order settles only late in the horizon.' sequential-chain yes 20 in-game years; roughly 340 to 400 decision stops per runother:decision-stops 0
FORESIGHT-9
FORESIGHT-9: Prospective and Process-Aware Evaluation of Adaptive Trading Agents
Business & enterprise 2026 maintain a coherent, active factor-ensemble/trading strategy across staged macro-financial events without silent internal collapse; respond adaptively to joint multi-asset anchors and disclosed in-world-time observations across a multi-year counterfactual worldline; keep portfolio holdings and decision-records mutually consistent (avoid decision-record vs. execution divergence) The agent must track its adaptive factor library/ensemble state, executed portfolio holdings, and staged event disclosures across a multi-year, ~2,468-trading-day counterfactual worldline, ensuring internal decision-records remain consistent with actual executed holdings. sequential-chain yes each worldline runs from the 2026-07-15 information boundary through 2035-12-31, approximately 2,468 simulated trading days (~9 years 5 months) per run, across 36 long-horizon runs totalsimulated-days 0
FrontierFinance
FrontierFinance: A Long-Horizon Computer-Use Benchmark of Real-World Financial Tasks
Business & enterprise 2026 complete each of 25 complex financial modeling tasks across five core finance models; produce client-ready outputs matching industry-standard financial-modeling workflows; satisfy detailed structured evaluation rubrics per task The agent must track detailed financial-modeling state (assumptions, formulas, model linkages) across a task requiring on average over 18 hours of equivalent skilled human labor, checked against detailed structured rubrics. sequential-chain yes average of over 18 hours of skilled human labor per taskhuman-expert-hours 20.4/mo
GDPevo
GDPevo: Evaluating Agent Self-Evolution on Real Business Tasks
Business & enterprise 2026 update an agent's persistent state from prior training-task experience (self-evolution); apply recombined atomic business rules correctly on held-out test tasks across CRM/ERP/finance/healthcare/legal/data-centric workflows; achieve held-out accuracy gains attributable specifically to training experience rather than data contamination The evaluation harness must track which atomic business rules were exposed during training versus recombined at test time, per task group, to attribute held-out gains specifically to training experience rather than contamination. set-of-independent yes none stated 0
H-AdminSim
H-AdminSim: A Multi-Agent Simulator for Realistic Hospital Administrative Workflows with FHIR Integration
Business & enterprise+ Multi-agent orgs 2026 process hospital administrative requests (drawn from a workload of 10,000+ requests/day in large hospitals); coordinate across multiple administrative subtasks rather than handling them in isolation; operate correctly against FHIR-integrated, heterogeneous hospital-setting data The multi-agent simulation must track hospital administrative request state across heterogeneous hospital settings via a unified FHIR-integrated environment, evaluated against detailed rubrics. other (singleton) yes none stated 0
HANDBOOK.md
HANDBOOK.md: A Benchmark for Long-Context Agentic Instruction Following
Business & enterprise 2026 locate the specific clauses in a long standing-policy document (20-124 pages) that apply to the current situation; carry out routine professional work strictly governed by that policy document across many tool-mediated actions; avoid letting a plausible but unauthorized in-environment request override the standing policy; satisfy every one of many deterministic, programmatic grading criteria (824 total) simultaneously The agent must hold the relevant clauses of a long standing-policy document across roughly 17 reasoning steps and 30 tool calls on average, checking each action against required and prohibited criteria rather than losing rule details over the long horizon. DAG-with-precedence yes ~17 reasoning steps and ~30 tool calls on average per tasktool-calls 21.0/mo
HealthAdminBench
HealthAdminBench: Evaluating Computer-Use Agents on Healthcare Administration Tasks
Business & enterprise+ Healthcare 2026 complete a Prior Authorization workflow end-to-end across an EHR and payer portal; complete an Appeals and Denials Management workflow; complete a Durable Medical Equipment (DME) Order Processing workflow; satisfy each of the fine-grained verifiable subtasks a task decomposes into (often 15+ per task) The agent must track progress across many fine-grained, cross-application subtasks (spanning EHR, payer portals, and fax) per administrative workflow, under a fixed interaction budget, to reach a correct terminal workflow state. DAG-with-precedence yes none statedother:not-stated 81.6/mo
Herculean
Herculean: An Agentic Benchmark for Financial Intelligence
Business & enterprise 2026 complete Trading workflow tasks via a standardized MCP-based skill environment; complete Hedging workflow tasks requiring long-horizon coordination and state consistency; complete Market Insights workflow tasks; complete Auditing workflow tasks requiring structured verification The agent must track workflow-specific state, tool interactions, and constraints within each MCP-based skill environment, with Hedging and Auditing additionally requiring sustained state consistency and structured verification across long-horizon coordination. set-of-independent yes none stated 30.75/mo
JobBench
JobBench: Aligning Agent Work With Human Will
Business & enterprise 2026 complete 130 agentic tasks spanning 35 occupations that experts identify as high-priority for delegation; reason through cluttered, heterogeneous reference-file information streams within a professional workspace; satisfy a fact-anchored chain of rubrics averaging 35.6 binary criteria per task Agent must correctly reason through a workspace of heterogeneous reference files and track satisfaction of an average of 35.6 fact-anchored binary rubric criteria per task. hierarchical yes none stated 41.0/mo
LH-Bench
LH-Bench: Skill-Grounded Evaluation of Long-Horizon Agents on Subjective Enterprise Tasks
Business & enterprise 2026 produce subjective, context-dependent enterprise work whose quality depends on organizational goals and user intent, not a single correct answer; produce correct intermediate artifacts across long, multi-tool workflows (e.g., chapter-level content, Figma-to-code conversions); satisfy expert-grounded rubrics scoring subjective work quality; align with pairwise human preference judgments as convergent validation The agent must track and produce correct intermediate artifacts (e.g., each chapter of a course, or each Figma-to-code conversion) across long, multi-tool workflows, since ground-truth artifacts enable stepwise reward signals at that granularity rather than only a single final judgment. hierarchical yes none stated 20.33/mo
Long-Horizon Multi-Tool Agent (LHMTA) task collection
Post-Training on Office Work Improves Software Engineering: A Behavioral Account of Cross-Domain Transfer
Business & enterprise 2026 select goals appropriately within nested and branching office work; construct task-relevant state as work unfolds; maintain fidelity to higher-level objectives across nested/branching sub-work; verify completion against the environment Agent must repeatedly select goals, construct task-relevant state, maintain fidelity to higher-level objectives, and verify completion against the environment across nested and branching long-horizon office work. hierarchical yes none stated 0
LongHorizon-Bench
Four-Axis Decision Alignment for Long-Horizon Enterprise AI Agents
Business & enterprise 2026 make high-stakes regulated decisions (loan qualification, insurance claims adjudication) under lossy memory and multi-step reasoning; maintain factual precision (FRP) about case facts over the long horizon; maintain reasoning coherence (RCS) across the multi-step decision process; reconstruct compliance/regulatory justification (CRR) for the decision; calibrate abstention (CAR) -- know when to decline to decide rather than commit incorrectly The agent must retain case facts accurately over a long-horizon multi-step review under lossy memory, maintain a coherent reasoning chain, reconstruct which regulatory rules justify the decision, and calibrate when to abstain rather than commit -- all against deterministic ground truth. set-of-independent yes none stated 10.2/mo
LongMedBench
LongMedBench: Benchmarking Medical Agents for Long-Horizon Clinical Decision-Making
Business & enterprise+ Healthcare 2026 aggregate evidence across a patient's repeated visits, tests, and evolving treatments to answer fact-based QA; perform temporal reasoning over the patient's event stream, including implicit (not just explicit-timestamped) time inference; make a long-horizon clinical decision that correctly uses historical patient information accumulated over many visits Agent must integrate time-series clinical events (admission records, notes) across many visits per patient, correctly distinguish explicit timestamps from cases requiring implicit time inference, and carry this forward into long-horizon decision-making. sequential-chain unclear 19.72 inpatient visits per patient (average); 44.91 medical events per visit (average)sessions 0
MBABench
MBABench: Evaluating LLM Agents on End-to-End Spreadsheet Tasks in Finance
Business & enterprise 2026 construct an entire financial spreadsheet (e.g., financial model, forecast, scenario analysis) end-to-end from a high-level instruction; jointly satisfy Accuracy, Formula, and Format criteria, each with fine-grained sub-criteria reflecting professional standards Agent must track intermediate calculation dependencies within the spreadsheet, professional formatting/readability conventions, and formula correctness simultaneously while building the deliverable end-to-end. set-of-independent yes none stated 20.5/mo
MerchantBench
MerchantBench: Benchmarking LLM Agents for Long-Term Coherence in E-Commerce Operations
Business & enterprise 2026 source products and manage upstream supplier events over time; set and adjust listing/pricing to remain competitive and solvent; manage cash-flow across delayed, heterogeneous-latency order outcomes; follow individual order lifecycles end-to-end and revisit earlier decisions as new (delayed) information arrives The agent must follow individual order lifecycles end-to-end, track upstream supplier events and their promised delayed downstream outcomes, manage cash-flow/net-assets state, and revisit/adapt earlier sourcing and pricing decisions as delayed feedback arrives across a full 365-simulated-day run. DAG-with-precedence yes a 365-day order-level simulation; 48 runs, each spanning 365 simulated dayssimulated-days 10.5/mo
Multi-Horizon Task Environments (MHTE) / CorpGen
CORPGEN: Simulating Corporate Environments with Autonomous Digital Employees in Multi-Horizon Task Environments
Business & enterprise 2026 manage dozens of concurrent, interleaved long-horizon corporate tasks (45+ tasks, 500-1500+ steps each); handle inter-task dependencies expressed as DAGs rather than simple chains; reprioritize among concurrent tasks as load and context change over a persistent execution context spanning hours The system must maintain hierarchical goal alignment, isolate sub-agent context to prevent cross-task contamination, and manage tiered (working/structured/semantic) memory with adaptive summarization across a persistent execution context spanning hours. DAG-with-precedence no 500-1500+agent-steps 0
OccuBench
OccuBench: Evaluating AI Agents on Real-World Professional Tasks via Language Environment Simulation
Business & enterprise 2026 complete each of 100 real-world professional task scenarios spanning 65 specialized domains; maintain task completion under controlled fault injection (explicit errors, implicit data degradation, mixed faults) Agent must track task-completion progress plus signals of environmental robustness (whether tool responses are timing out, truncated, or subtly degraded) within each professional scenario. set-of-independent yes none stated 40.8/mo
OmegaUse-OfficeVal
OmegaUse-OfficeVal: Benchmarking LLM Agents on Long-Horizon Office-Suite Tasks with Economic Grounding
Business & enterprise 2026 complete 100 practitioner-derived office-suite tasks end-to-end; achieve deliverable quality verified via fine-grained rubric-based code verifiers; remain economically competitive relative to human labor time and task price Agent must track fine-grained rubric criteria and produce a deliverable matching the practitioner's request, verified via code-based verifiers, implicitly benchmarked against the human labor time (avg. 2.32 hours) needed for the same task. set-of-independent yes 2.32 (average)human-expert-hours 0
OptAgent
OptAgent: an Agentic AI framework for Intelligent Building Operations
Business & enterprise 2026 assess how a system/control upgrade changes energy use; assess how the same upgrade changes operating cost; assess how the same upgrade changes thermal comfort; assess how the same upgrade changes flexibility, via coordinated multi-domain multi-agent analytics The orchestrator must track which specialist agent/tool has been invoked for which sub-domain (thermal dynamics, HVAC, DER), intermediate physics-informed simulation outputs, and how upgrades propagate across energy use, cost, comfort, and flexibility metrics within one workflow. DAG-with-precedence no none statedother:not-stated 10.12/mo
OR-Space
OR-Space: A Full-Lifecycle Workspace Benchmark for Industrial Optimization Agents
Business & enterprise 2026 construct a solver-ready optimization model from heterogeneous business artifacts (Build); revise an existing model under changing requirements or solver feedback while preserving valid prior logic (Revise); answer grounded questions about solutions, constraints, and business implications using evidence spread across workspace artifacts (Explain) The agent must track the current state of a persistent multi-artifact workspace (documents, structured data, code, solver outputs) across model construction, revision, and explanation stages, preserving valid prior modeling logic when new requirements or solver feedback arrive. DAG-with-precedence unclear none stated 20.5/mo
OrthoPilot
Evidence-Grounded AI for Musculoskeletal Care
Business & enterprise+ Healthcare 2026 integrate evolving imaging, laboratory, pathology, and order data as it arrives across visits; produce evidence-based decisions at each stage of care, from admission diagnosis through rehabilitation planning; maintain continuous, individualised management across the full musculoskeletal care pathway rather than isolated per-visit decisions System must continuously retrieve and integrate real-time imaging, laboratory, pathology, and order data across visits/departments/hospital systems, translating evolving patient state into stage-specific functional goals across the whole care pathway. sequential-chain yes months to years (per patient pathway); 1,870 cases / 8,240 inpatients across the studyother:real-world-months-to-years-per-patient-pathway 0
OSWorkerBench
UI-Mate: Advancing Open-Weight Foundation GUI Agents with In-Context Demonstrations
Business & enterprise+ Web & GUI 2026 complete each of 100 long-horizon office tasks spanning 41 applications; follow demonstrated subtask-level workflows (self-demo or variant-demo) and re-plan from the live interface when needed; make measurable progress toward task completion even when strict end-to-end success is not reached Agent must track its position in a demonstrated or self-planned subtask sequence, detect when the live interface diverges from the demonstration, and re-plan accordingly across a long-horizon office task. sequential-chain yes none stated 0
PhysicianBench
PhysicianBench: Evaluating LLM Agents in Real-World EHR Environments
Business & enterprise+ Healthcare 2026 retrieve relevant clinical data across multiple encounters in the EHR; reason over heterogeneous clinical information (labs, notes, orders) to reach a decision; execute consequential clinical actions (e.g. prescribing, ordering) grounded against the environment; produce clinical documentation reflecting the completed workflow Agent must track data retrieved across multiple encounters, intermediate clinical reasoning state, and which structured checkpoints have been satisfied, using execution-grounded verification against real patient records via standard EHR APIs. hierarchical yes 27 (average)tool-calls 143.5/mo
POLARIS
POLARIS: Typed Planning and Governed Execution for Agentic AI in Back-Office Automation
Business & enterprise 2026 synthesize a type-checked directed acyclic graph (DAG) plan for a back-office document-processing task; select a single compliant plan via rubric-guided reasoning among structurally diverse candidate DAGs; pass validator-gated checks and a bounded repair loop before execution; route or block side effects per compiled policy guardrails, including anomaly routing System must track the candidate DAG structure, per-node type/validator status, and the full execution trace/audit trail to support bounded repair and decision-grade anomaly routing. DAG-with-precedence yes none stated 20.25/mo
PolyWorkBench
PolyWorkBench: Benchmarking Multilingual Long-Horizon LLM Agents
Business & enterprise 2026 process heterogeneous multilingual inputs correctly within a workflow task; perform iterative reasoning and invoke external tools while maintaining linguistic consistency; produce a structured, correct output for one of five workplace domains (commerce, knowledge work, legal analysis, localization, manufacturing) Agent must track functional correctness and linguistic consistency simultaneously across a workflow's reasoning and tool-invocation steps, since the two can drift independently as multilingual inputs are processed. sequential-chain yes none stated 10.33/mo
PolyWorkBench
PolyWorkBench: Benchmarking LLM Agents for Cross-Lingual Long-Horizon Workflows
Business & enterprise 2026 integrate heterogeneous multilingual inputs relevant to a workplace task; execute iterative tool-use trajectories within a target domain (commerce, knowledge work, legal analysis, localization, manufacturing); produce structured domain artifacts as verifiable task output Agent must track and correctly integrate heterogeneous multilingual information while executing an iterative sequence of tool calls, and produce a structured domain artifact that is later verified. sequential-chain yes none stated 0
PowerAgentBench-SS
PowerAgentBench-SS: A Benchmark for Agentic AI in Power System Steady-State Studies
Business & enterprise 2026 inspect a grid case and select appropriate tools/simulators for the workflow; screen a large space of contingencies within a limited validation budget; propose admissible mitigations for discovered risks and validate their physical validity; produce an auditable evidence trail and submit a bounded, ranked report of top contingencies The agent must track which contingencies it has already validated (against a fixed budget of 80 out of 1,035), the evidence log supporting each, and whether its running set of top candidates still reflects the best available evidence as it allocates its remaining validation budget. sequential-chain yes validation budget of 80 contingency-case validations per episode (out of 1,035 total N-2 cases), submitting a ranked report of exactly 20 contingencies; hidden dangerous set of 52 casesactions 41.33/mo
PPT-Eval
PPT-Eval: A Benchmark for Computer-Use Agents on PowerPoint Tasks
Business & enterprise+ Web & GUI 2026 complete content-creation and presentation-editing tasks across 120 PowerPoint tasks in 12 files; satisfy task-specific rubric criteria that award partial credit for intermediate steps; avoid unnecessary changes and poor aesthetics while making required edits Agent must track which task-specific rubric criteria (intermediate steps, aesthetics, unnecessary changes) have been satisfied across a multimodal editing session, since rubrics award partial credit rather than only final binary success. hierarchical yes none stated 51.67/mo
RetailBench
RetailBench: Benchmarking long horizon reasoning and coherent decision making of LLM agents in realistic retail environments
Business & enterprise 2026 manage pricing across the store's product assortment; manage replenishment and supplier selection to keep inventory stocked; manage shelf assortment and inventory aging; respond appropriately to customer feedback and external events; maintain solvency (cash-flow constraints) while maximizing net worth/sales over a long simulated horizon The agent must track pricing, inventory levels/aging, supplier relationships, customer feedback, external events, and its own cash-flow position day by day across the 180-day (or longer) simulated run, since only a small subset of evaluated agents survive the full evaluation horizon. sequential-chain yes 180-day evaluation horizon; the simulator supports thousand-day-scale simulationssimulated-days 61.0/mo
RiskWebWorld
RiskWebWorld: A Realistic Interactive Benchmark for GUI Agents in E-commerce Risk Management
Business & enterprise+ Web & GUI 2026 investigate flagged e-commerce risk cases across multiple verification sub-steps on production risk-control pipelines; operate GUI actions on uncooperative websites subject to partial environmental hijacking; complete each of 1,513 tasks spanning 8 core risk-control domains; succeed at long-horizon professional risk-investigation workflows The agent must track evidence and verification state gathered across multiple sub-steps of a risk investigation on an uncooperative website, while detecting and coping with partial environmental hijacking attempts, across long-horizon professional tasks. sequential-chain no none stated 0
SaaS-Bench
SaaS-Bench: Can Computer-Use Agents Leverage Real-World SaaS to Solve Professional Workflows?
Business & enterprise 2026 navigate and operate real, deployed SaaS systems to complete a professional workflow; coordinate state and context across multiple applications within the same workflow; apply domain-specific knowledge correctly within the SaaS system; recover from errors and maintain progress over a long-horizon task The agent must maintain state and context across multiple SaaS applications over long-horizon execution, tracking partial progress against weighted verification checkpoints, since 'agents become stuck acquiring information or manipulating interfaces, confuse source and output devices, or terminate before all conditions are jointly satisfied.' DAG-with-precedence yes average of over 100 interaction steps per taskactions 30.75/mo
SpreadsheetBench 2
SpreadsheetBench 2: Evaluating Agents on End-to-End Business Spreadsheet Workflows
Business & enterprise 2026 generate new spreadsheet content/formulas correctly across a large multi-sheet workbook; debug existing incorrect formulas/content within the workbook; produce correct visualizations from the workbook's data The agent must track and correctly propagate changes across an average of 11.8 interdependent worksheets requiring 593.5 cell modifications per task, correctly identifying target cells under a unified multi-turn agent scaffold. DAG-with-precedence yes each task instance averages 11.8 worksheets and requires 593.5 cell modificationsother:cell-modifications 41.33/mo
SupChain-Bench
SupChain-Bench: Benchmarking Large Language Models for Real-World Supply Chain Management
Business & enterprise 2026 orchestrate long-horizon, multi-step supply-chain tool use grounded in standard operating procedures (SOPs); correctly apply supply-chain domain knowledge across a sequence of dependent tool calls Agent must track SOP-grounded procedural state across a long-horizon sequence of dependent tool calls in order to correctly complete supply-chain management workflows. sequential-chain no none stated 0
Synthetic Computers at Scale
Synthetic Computers at Scale for Long-Horizon Productivity Simulation
Business & enterprise 2026 complete each of multiple professional deliverables comprising a computer-specific productivity objective; navigate the synthetic computer's filesystem to ground actions in the user's actual context; coordinate with simulated collaborators as needed to complete the objective; sustain progress toward an objective requiring about a month of simulated human work The acting agent must track the evolving state of a realistic folder hierarchy and content-rich artifacts (documents, spreadsheets, presentations) across a run spanning over 2,000 turns, coordinating with simulated collaborators toward multiple professional deliverables. open-ended no >2,000 turns on average (each run also requiring over 8 hours of agent runtime)turns 10.2/mo
temporal enterprise scenario replay system
What Could the Agent See at 19:05? Generating Temporal Enterprise Scenarios from Real Research and Replaying Them to Evaluate Agents
Business & enterprise 2026 answer questions correctly relative to what data existed and who could see it at a specific queried moment; reason about a persona-driven, temporally-evolving enterprise world spanning many apps; avoid leaking future/hidden record state when reasoning about an earlier moment The agent being evaluated must reason correctly about what data existed and who could see it at a specific queried moment, without conflating it with earlier or later states of the same records across multiple apps. DAG-with-precedence yes none stated 0
Thinkingbox-bench
One Success Isn't Reliability: Thinkingbox, a Sandbox and Benchmark for Agents in Stateful Business Workflows
Business & enterprise+ Information seeking 2026 gather missing information over multiple turns before acting; follow domain-specific policies for each of 507 policy-conditioned workflows; coordinate dependent tools correctly; realize exactly the correct persistent backend state transition without collateral effects Agent must gather missing information across multiple turns, track applicable domain policies, coordinate dependent tool calls, and verify the resulting persistent backend state matches exactly the required final state with no collateral effects. DAG-with-precedence no none stated 0
Underwrite
Benchmarking Agents in Insurance Underwriting Environments
Business & enterprise+ Information seeking 2026 gather information carefully from noisy tool interfaces and imperfect simulated users during an underwriting conversation; apply proprietary business/domain knowledge correctly rather than hallucinating it; reach a final underwriting decision consistent with the accumulated evidence gathered across the conversation The agent must track what information it has already gathered (and from where), the reliability of that information given noisy tool interfaces, and how it should update its evolving underwriting assessment as more evidence accumulates across the conversation. sequential-chain yes average of 3-7 steps of required reasoning and tool use, with a total of 10-20 conversational turnsturns 0
Workflow-GYM
Workflow-GYM: Towards Long-Horizon Evaluation of Computer-use Agentic tasks in Real-World Professional Fields
Business & enterprise 2026 autonomously operate domain-specific professional software GUIs to accomplish economically valuable work; complete long-horizon, multi-stage professional workflows end-to-end; maintain workflow consistency across stages without omission, error propagation, or objective drift The agent must track its position and completed stages within a long-horizon, multi-stage professional GUI workflow, avoiding stage omission, propagating errors, or drifting from the original objective across the workflow's duration. sequential-chain no none stated 20.67/mo
Workspace-Bench
Workspace-Bench 1.0: Benchmarking AI Agents on Workspace Tasks with Large-Scale File Dependencies
Business & enterprise 2026 identify, reason over, exploit, and update explicit and implicit dependencies among heterogeneous files in a worker's workspace; satisfy each of a task's own file-dependency-graph-derived rubrics via cross-file retrieval, contextual reasoning, and adaptive decision-making Agent must track which of up to 20,476 files across 74 file types are relevant, their dependency relationships, and adaptively retrieve/reason across files to satisfy each task's rubrics. DAG-with-precedence yes none stated 102.5/mo
World of Workflows (WoW) / WoW-bench
World of Workflows: a Benchmark for Bringing World Models to Enterprise Systems
Business & enterprise 2026 complete constrained agentic tasks within a ServiceNow environment governed by 4,000+ business rules and 55 active hidden workflows; predict cascading side effects of actions across interconnected databases; mentally simulate hidden state transitions to avoid silent constraint violations under limited observability Agent must mentally simulate hidden state transitions and predict cascading side effects across interconnected databases to bridge the observability gap, since high-fidelity feedback is often unavailable. DAG-with-precedence no none stated 70.88/mo
YC-Bench
YC-Bench: Benchmarking AI Agents for Long-Term Planning and Consistent Execution
Business & enterprise 2026 manage employees within a simulated startup; select task contracts to pursue under uncertainty; maintain profitability against adversarial clients and growing payroll; detect and avoid bankruptcy-inducing failure modes (e.g. adversarial-client mismanagement, over-parallelization) over a one-year run The agent must persist information across context truncation via a scratchpad (the strongest predictor of success), tracking employees, contracts, cash reserves, and adversarial-client risk across hundreds of turns spanning a simulated year. open-ended no hundreds of turns (over a simulated one-year horizon)turns 81.6/mo
year-long LLM stock-market trading-style simulation testbed
Behavioral Consistency Validation for LLM Agents: An Analysis of Trading-Style Switching through Stock-Market Simulation
Business & enterprise 2026 trade under a currently assigned style (fundamental or technical) each trading day; periodically (every 10 trading days) reassess and potentially switch trading style based on four behavioral-finance drivers: loss aversion, herding, wealth differentiation, price misalignment; keep style-switching behavior consistent with real-world behavioral-finance theory over the whole simulated year Agent must process daily price-volume data, retain long-term personality traits (the four behavioral-finance drivers) set at initialization, and track its own accumulated wealth/strategy history to decide whether to switch trading style every 10 days. sequential-chain unclear year-long (reassessed every 10 trading days)simulated-days 40.57/mo
Agent Trading Arena
Agent Trading Arena: A Study on Numerical Understanding in LLM-Based Agents
Business & enterprise 2025 make sequential buy/sell trading decisions that directly affect and are affected by shared market prices; compete against other LLM-based agents in a zero-sum stock market; maximize trading performance/returns, especially under high volatility; correctly perform numerical reasoning over price data or chart-based visualizations Each agent must track its own portfolio/capital, the current market price state shaped by all agents' recent trades, and historical price patterns, updating its numerical reasoning as the shared market evolves turn to turn. open-ended no none stated 120.63/mo
AI-Trader
AI-Trader: Benchmarking Autonomous Agents in Real-Time Financial Markets
Business & enterprise 2025 independently search, verify, and synthesize live market information given only minimal initial context; make live trading decisions (buy/sell/hold) across U.S. stocks, A-shares, and cryptocurrencies at multiple trading granularities; manage risk and sustain positive returns over a continuous live-trading period Agent must track evolving live market data it has independently searched/verified, current positions/risk exposure across three markets and multiple trading frequencies, and running returns over the continuous live-trading period. sequential-chain yes none stated 192.11/mo
AssetOpsBench
AssetOpsBench: A Real-World Evaluation Benchmark for AI-Driven Task Automation in Industrial Asset Management
Business & enterprise 2025 orchestrate the correct domain-specific agent(s) (from a catalog of four) to answer a natural-language industrial-operations query; correctly complete condition-monitoring and maintenance-scheduling workflow steps grounded in a simulated IoT environment The agent must track intermediate tool/agent outputs across a multi-step think-act-observe loop, orchestrate the four domain-specific agents appropriately, and maintain consistency with the simulated CouchDB-backed IoT environment's evolving state. DAG-with-precedence yes Plan-Execute agents complete most tasks in approximately 2.6-4.4 steps; Agent-As-Tool agents typically require approximately 4-6+ steps due to its iterative think-act-observe loopagent-steps 171.13/mo
AutoDW
Automating Complex Document Workflows via Stepwise and Rollback-Enabled Operation Orchestration
Business & enterprise 2025 execute a sequence of interdependent, user-specified document-editing instructions within a session; keep the execution trajectory aligned with evolving document state and user intent across the whole session; recover from failed API calls/arguments via rollback at both the argument and API level System must track the evolving document state after each API action, the remaining interdependent instructions in the session, and whether any prior action needs to be rolled back (at argument or API level) to stay aligned with user intent. DAG-with-precedence yes 1,708 instructions across 250 sessions (~6.8 instructions/session on average)actions 0
capitalization tie-out benchmark (legal AI)
Does It Tie Out? Towards Autonomous Legal Agents in Venture Capital
Business & enterprise 2025 verify that every security (shares, options, warrants) is supported by underlying legal documentation; verify that every issuance term (vesting schedules, acceleration triggers, transfer restrictions) is consistent across the dataroom; maintain strict evidence traceability while reconciling thousands of pages of legal documents; produce a deterministic, fully-reconciled capitalization table The agent must track which securities and issuance terms have been verified against which supporting documents, maintaining strict evidence traceability across a growing dataroom (up to tens of thousands of pages) as it reconciles a scaling number of individual securities. set-of-independent yes workload grows from approximately 2,700 atomic verification steps at Seed stage to nearly 8,000 steps at Series B stageagent-steps 0
CP-Env
CP-Env: Evaluating Large Language Models on Clinical Pathways in a Controllable Hospital Environment
Business & enterprise+ Healthcare 2025 triage a patient correctly to determine care pathway; consult the appropriate specialist given patient information; order and interpret diagnostic tests as needed; participate appropriately in multidisciplinary team meetings; complete a patient's full journey across branching, long-horizon clinical-pathway stages The agent must track patient information, diagnostic findings, and evolving care-pathway state across branching stages from triage through specialist consultation, diagnostic testing, and multidisciplinary team meetings. DAG-with-precedence yes none stated 60.67/mo
CRMArena-Pro
CRMArena-Pro: Holistic Assessment of LLM Agents Across Diverse Business Scenarios and Interactions
Business & enterprise 2025 complete each of nineteen expert-validated CRM tasks across sales, service, and configure-price-quote (CPQ) processes; sustain correct behavior across multi-turn interactions guided by diverse personas; maintain confidentiality awareness throughout the interaction, for both B2B and B2C scenarios The agent must track persona-specific context, confidentiality constraints, and workflow state across multi-turn interactions, since single-turn success (58%) drops substantially to about 35% once genuine multi-turn tracking is required. sequential-chain yes none stated 422.62/mo
DevNous
DevNous: An LLM-Based Multi-Agent System for Grounding IT Project Management in Unstructured Conversation
Business & enterprise+ Multi-agent orgs 2025 identify actionable intents from informal, unstructured team-chat dialogue; manage stateful, multi-turn administrative workflows (task formalization, progress-summary synthesis) grounded in that dialogue The agent must identify actionable intents across informal chat and maintain stateful workflows (task formalization, progress tracking) across a benchmark of 160 conversational turns, correctly matching a multi-label ground truth per turn. sequential-chain yes a new benchmark of 160 realistic, interactive conversational turnsturns 0
DRBench
DRBench: A Realistic Benchmark for Enterprise Deep Research
Business & enterprise 2025 identify supporting facts for a multi-step research query from both the public web and a private company knowledge base; synthesize facts drawn from heterogeneous enterprise sources (productivity software, cloud file systems, emails, chat, web) into one answer; produce a coherent, well-structured final report grounded in the retrieved facts The agent must track which facts it has found in which source (public web vs. private enterprise knowledge base), maintain factual accuracy across sources, and assemble a coherent report structure from these accumulated facts. hierarchical yes none stated 151.25/mo
EcomBench
EcomBench: Towards Holistic Evaluation of Foundation Agents in E-commerce
Business & enterprise 2025 retrieve deep, possibly multi-hop information relevant to a real e-commerce user demand; perform multi-step reasoning across e-commerce-domain data; integrate knowledge from multiple sources to resolve a task; correctly complete tasks across three graded difficulty levels The agent must track partial evidence gathered from deep retrieval and multiple knowledge sources across many action steps before it can complete a higher-difficulty task. hierarchical yes none stated 70.78/mo
Finch (FinWorkBench)
Finch: Benchmarking Finance & Accounting across Spreadsheet-Centric Enterprise Workflows
Business & enterprise 2025 interleave data entry, structuring, formatting, web search, cross-file retrieval, calculation, modeling, validation, translation, visualization, and reporting subtasks into one finance/accounting workflow; correctly complete each of the linked tasks that compose one of 172 composite workflows Agent must track state across many interlinked spreadsheets, PDFs, and artifacts (27 million spreadsheet cells total) while interleaving retrieval, calculation, modeling, validation, and reporting sub-steps of one composite workflow. DAG-with-precedence yes 16.8 (GPT-5.1 Pro average)wall-clock-minutes 91.0/mo
GraphicBench
GraphicBench: A Planning Benchmark for Graphic Design with Language Agents
Business & enterprise+ Web & GUI 2025 produce a workflow plan satisfying explicit design constraints stated in a user query; also satisfy implicit commonsense design constraints not stated by the user; select the correct action from 46 available tools at each workflow step; coordinate outputs across three design experts without violating global dependencies The agent must track explicit and implicit design constraints, the evolving multi-step workflow plan across three design experts, and which of 46 actions remains valid/appropriate at each step. DAG-with-precedence yes none stated 20.12/mo
HiMA-Ecom
HiMA-Ecom: Enabling Joint Training of Hierarchical Multi-Agent E-commerce Assistants
Business & enterprise+ Multi-agent orgs 2025 master agent coordinates multiple specialized sub-agents in an e-commerce workflow; each sub-agent pursues a functionally distinct role-specific objective (e.g. recall domain knowledge, execute a service function); jointly optimize system-level behavior via multi-agent reinforcement learning The master agent must track which specialized sub-agent is responsible for which functional sub-task and how sub-agent outputs compose into overall system behavior, using a collaboratively updated memory across training. hierarchical no none statedother:not-stated 80.53/mo
MEMTRACK
MEMTRACK: Evaluating Long-Term Memory and State Tracking in Multi-Platform Dynamic Agent Environments
Business & enterprise+ Assistants & memory 2025 acquire relevant facts scattered across asynchronous cross-platform events (Slack/Linear/Git); select the correct, currently-valid fact when multiple conflicting/noisy candidates exist; resolve conflicts between contradictory or stale cross-referring information over time The agent must acquire, select, and reconcile facts from a chronologically platform-interleaved timeline spanning Slack, Linear, and Git, handling noisy, conflicting, and cross-referring information plus codebase/file-system exploration, across long horizons. DAG-with-precedence yes none stated 121.09/mo
Mini Amusement Parks (MAPs)
Mini Amusement Parks (MAPs): A Testbed for Modelling Business Decisions
Business & enterprise 2025 maximize the amusement park's overall value over the given time horizon; make daily operational decisions (building rides/shops, hiring staff, setting a research agenda) that keep the business viable under sparse, stochastic feedback; reason over the park's spatial layout while planning these decisions The agent must track the park's spatial layout, built rides/shops, staffing, accumulated (sparse) experience about environment dynamics, and overall park value across a 50-100 day episode, planning under uncertainty at each of many daily decision points. sequential-chain yes episodes span a 50-day horizon on easy mode, extended to 100 days on medium modesimulated-days 20.2/mo
OdysseyBench
OdysseyBench: Evaluating LLM Agents on Long-Horizon Complex Office Application Workflows
Business & enterprise 2025 identify essential information buried in long-horizon interaction histories; perform multi-step reasoning/actions across Word, Excel, PDF, Email, and Calendar applications; complete real-world-derived tasks (OdysseyBench+) or newly synthesized complex tasks (OdysseyBench-Neo) Agent must retain and retrieve essential facts from long-horizon interaction histories spanning multiple applications in order to correctly execute later multi-step, cross-application actions. sequential-chain unclear none stated 513.92/mo
PPTArena
PPTArena: A Benchmark for PowerPoint Editing
Business & enterprise 2025 apply each of a deck's human-curated edits correctly from natural-language instructions; maintain layout-sensitive and cross-slide consistency across the whole deck while editing; plan and verify edit sequences via an iterative plan-edit-check loop (for the PPTPilot agent) Agent must track the deck's current structural/visual state after each edit, verify each edit against a ground-truth rubric, and maintain deck-wide consistency (styles, cross-slide references) across the full sequence of edits. sequential-chain yes ~13 edits per deck (1,300+ edits across 100 decks)actions 70.78/mo
ProSoftArena
ProSoftArena: Benchmarking Hierarchical Capabilities of Multimodal Agents in Professional Software Environments
Business & enterprise 2025 operate professional software tools according to a hierarchical capability level (L1-L3); complete realistic work/research tasks spanning 6 disciplines and 13 core professional applications; coordinate across multiple professional software applications for L3 multi-software workflows The agent must track its progress within and across professional-software applications as required by the task's capability level, since L3 tasks require coordinating state across multiple applications rather than operating one application in isolation. hierarchical yes human execution steps grow from an average of 5.1 steps (14.8 seconds) at L1 to 86.9 steps (506.8 seconds) at L3actions 50.56/mo
REALM-Bench
REALM-Bench: A Benchmark for Evaluating Multi-Agent Systems on Real-world, Dynamic Planning and Scheduling Tasks
Business & enterprise+ Multi-agent orgs 2025 solve each of 14 real-world planning/scheduling problems with multiple parallel planning threads; maintain feasibility across inter-agent dependencies as the problem scales in complexity; adapt schedules in real time when unexpected disruptions arrive Agents must track the state of parallel planning threads, inter-dependencies between agents/threads, and disruption events in order to replan in real time. DAG-with-precedence yes none stated 120.63/mo
Remote Labor Index (RLI)
Remote Labor Index: Measuring AI Automation of Remote Work
Business & enterprise 2025 complete a whole real, economically valuable freelance/remote-work project end to end; achieve a level of automation comparable to what a human freelancer would deliver for that project Agent must track progress toward completing an entire multi-part real-world project (not a single isolated action) to be credited with automating that unit of remote labor. hierarchical unclear none stated 211.91/mo
SCUBA
SCUBA: Salesforce Computer Use Benchmark
Business & enterprise+ Information seeking 2025 navigate a specific enterprise software UI (Salesforce) to accomplish a CRM task; manipulate data records and automate workflows within the Salesforce platform; retrieve information and troubleshoot issues as part of a realistic CRM task; generalize across three personas (platform administrators, sales representatives, service agents) each with distinct task profiles The agent must track its progress toward fine-grained milestones within each Salesforce sandbox task (UI navigation state, data changes made, workflow steps completed) to receive interpretable milestone-progress credit. sequential-chain yes none stated 70.58/mo
SOP-Bench
SOP-Bench: Complex Industrial SOPs for Evaluating LLM Agents
Business & enterprise 2025 correctly execute each step of a complex, multi-step Standard Operating Procedure; orchestrate the correct tools/APIs at each SOP step; produce ground-truth outputs matching the SOP's authored specification across 12 business domains The agent must track its position within a multi-step SOP, select the correct tool from a registry that may contain many irrelevant tools, and maintain consistency with the SOP's ground-truth interface and outputs across the whole procedure. sequential-chain no none statedother:not-stated 201.33/mo
StaffPro
StaffPro: an LLM Agent for Joint Staffing and Profiling
Business & enterprise 2025 assign and schedule tasks to workers (staffing), forming teams as needed; continuously estimate workers' latent skills, preferences, and other attributes from unstructured feedback (profiling); optimize staffing performance over time as profiling estimates improve via an ongoing human-agent feedback loop StaffPro must track each worker's evolving latent-attribute profile (estimated from ongoing human feedback) and the current staffing/schedule state, updating both jointly over the 'life-long' profiling horizon. DAG-with-precedence yes none stated 20.14/mo
STOCKBENCH
StockBench: Can LLM Agents Trade Stocks Profitably In Real-world Markets?
Business & enterprise 2025 make a sequential daily buy/sell/hold decision based on incoming market signals (prices, fundamentals, news); maximize cumulative return across the whole multi-month trading period; manage risk, minimizing maximum drawdown and maintaining a strong Sortino ratio Agent must track its current portfolio position, cumulative return, and risk exposure (drawdown) as it processes a new daily market signal (prices, fundamentals, news) each step of a multi-month simulated trading run. sequential-chain yes multi-month (exact day/month count not given)simulated-days 333.0/mo
UpBench
UpBench: A Dynamically Evolving Real-World Labor-Market Agentic Benchmark Framework Built for Human-Centric AI
Business & enterprise 2025 complete a real, verified client transaction job sourced from the Upwork labor marketplace; satisfy each of a job's detailed, expert-decomposed, verifiable acceptance criteria; follow instructions faithfully enough to earn positive fine-grained, per-criterion human expert feedback Agent must track which of a job's detailed acceptance criteria it has satisfied, informed by expert freelancer decomposition, to produce a submission gradeable on fine-grained, per-criterion feedback rather than a binary pass/fail. set-of-independent yes none stated 40.4/mo
Vending-Bench
Vending-Bench: A Benchmark for Long-Term Coherence of Autonomous Agents
Business & enterprise 2025 balance inventory levels of vending-machine stock; place restocking orders from suppliers; set item prices; pay recurring daily fees; sustain a profitable long-running vending-machine business without derailing The agent must continuously track inventory counts, cash/profit, outstanding orders and their delivery schedules, and daily fees across a run spanning over 20M tokens, since forgetting an order or misreading a schedule causes derailment. open-ended no >20Mother:tokens-per-run 673.53/mo
VitaBench
VitaBench: Benchmarking LLM Agents with Versatile Interactive Tasks in Real-world Applications
Business & enterprise 2025 satisfy each of multiple real user requests combined into one cross-scenario task (food delivery, in-store consumption, online travel); reason across temporal and spatial dimensions while using a large tool set (66 tools); proactively clarify ambiguous instructions and track shifting user intent across a multi-turn conversation Agent must track shifting user intent across a multi-turn conversation, temporal and spatial constraints, and the state of a large (66-tool) toolset spanning multiple life-serving domains simultaneously. set-of-independent unclear none stated 352.92/mo
Wealth-Management benchmark (extension of TheAgentCompany)
Benchmarking LLM Agents for Wealth-Management Workflows
Business & enterprise 2025 complete each of 12 wealth-management task-pairs spanning retrieval, analysis, and synthesis/communication; satisfy explicit acceptance criteria under deterministic graders; operate correctly under both high- and low-autonomy task variants Agent must track retrieved data, intermediate analysis results, and explicit acceptance criteria across the retrieval-analysis-synthesis pipeline, plus which autonomy variant (high vs. low) governs how independently it may act. sequential-chain no none stated 0
Service & dialogue23 artifacts
LLMs Get Lost in Evolving User Intent
LLMs Get Lost in Evolving User Intent
Service & dialogue 2026 track a single user goal that is incrementally revealed across conversation turns; detect and adopt goal revisions as the user changes their mind mid-conversation; detect and follow mid-conversation redirections of the original goal; re-run an existing single-turn benchmark's task under this evolving-intent protocol without new annotation The agent must maintain a running model of a single, evolving user intent across turns, discarding or revising earlier partial specifications as more of the goal is revealed, revised, or redirected. sequential-chain no up to 7 (including the initial turn)turns 0
Research on a Multi-SOP Interruption–Resumption Agentic Algorithm for Complex Business Workflows
Research on a Multi-SOP Interruption–Resumption Agentic Algorithm for Complex Business Workflows
Service & dialogue+ Business & enterprise 2026 complete a customer-requested Standard Operating Procedure (SOP) workflow; switch to and complete a different SOP mid-interaction when the customer changes topic; resume a previously interrupted SOP without restarting or re-asking for already-given information; detect and recover from stale/rolled-back state via diff reasoning against frozen snapshots The system must track which SOP is currently active, a shared six-tuple Global State Container of cross-SOP business entities, and frozen-snapshot vs. current-state diffs to support interruption/resumption without re-asking the user for information already given. DAG-with-precedence no none stated 0
AcCoRD
AcCoRD: Evaluating User-Agent Collaboration Under Realistic User Preference Dynamics
Service & dialogue+ Multi-agent orgs 2026 resolve underspecified user preferences in online shopping or travel planning; detect and satisfy preferences that emerge mid-interaction (not stated upfront); adapt to preferences the user later adjusts or relaxes during the interaction Agent must maintain and continuously update a model of the user's preferences as they are formed, revealed, adjusted, or relaxed turn-by-turn, and recognize when uncertainty about a preference needs to be resolved. sequential-chain unclear none stated 0
AgentWorld
AgentWorld: Personality-Aware Reliability Evaluation for Agentic Information Retrieval
Service & dialogue 2026 maintain consistent, reliable tool-use behavior across interactions with users of varying personality (Big Five/OCEAN) profiles; handle dual-control handoffs correctly between agent and (simulated) user/other controller; resist or survive adversarial perturbations to a required-intermediate-state 'spine' of the task without brittle failure; achieve consistent pass rates (pass^k) across repeated attempts rather than succeeding only once by chance The agent must track its stateful tool-use progress consistently across repeated attempts and varying user personas, while the Risk Analyser separately tracks the required-intermediate-state spine of a task to evaluate how perturbations at each state affect eventual outcomes. DAG-with-precedence yes each persona ran 3 multi-turn conversation exchanges against the analytics agent, producing 60 messages total (30 from the agent, 30 from personas), across 10 personasturns 0
CallBench
CallBench: A Benchmark for Dual-Goal Coordination in Phone Call Assistants
Service & dialogue+ Assistants & memory 2026 satisfy the device owner's explicit preset goal for the call; correctly infer and respond to the caller's implicit and dynamic goal; make turn-level decisions that correctly reconcile these two goals under alignment, complementarity, irrelevance, or conflict relations; adhere to preset instructions while maintaining dialogue quality, safety, and rhythm The assistant must track the owner's preset goal, the evolving implicit goal of the caller, and the current relation between them (alignment/complementarity/irrelevance/conflict) turn by turn across the dialogue, to make reliable turn-level decisions between the two goals. DAG-with-precedence yes average of 5.322 turns per dialogue across all 50,000 dialogues (scenario averages ranging from 4.437 to 7.241 turns)turns 10.33/mo
clinical AI red-teaming framework (unnamed)
Assessing Risks of Large Language Models in Mental Health Support: A Framework for Automated Clinical AI Red Teaming
Service & dialogue+ Healthcare 2026 conduct a full therapy session with a simulated patient agent having a dynamic cognitive-affective model; maintain quality of care across the session per a comprehensive risk ontology; de-escalate suicide risk appropriately when it arises during the session; avoid validating patient delusions ('AI Psychosis') The AI psychotherapist agent must track the patient's evolving cognitive-affective state, emerging risk signals (e.g., suicide risk, delusion reinforcement), and quality-of-care obligations across the whole therapy session. sequential-chain no none stated 30.43/mo
CRAB-Bench
CRAB-Bench: Evaluating LLM Agents under Complex Task Dependencies and Human-aligned User Simulation
Service & dialogue 2026 reason over a constraint graph spanning multiple interdependent entities to find a solution among thousands of misleading candidate distractors; engage in multi-turn dialogue with a realistic (non-cooperative, persona-driven) simulated user rather than a cooperative template-like one; accommodate multiple valid solutions rather than a single fixed correct answer Agent must track which constraints (graph edges/entities) have been satisfied so far, which candidate solutions remain viable given thousands of distractors, and information gathered/disclosed across a multi-turn dialogue with a realistic user persona. DAG-with-precedence yes none stated 0
FraudBench
FraudBench: Stress-Testing Policy-Grounded Banking Agents Against Adaptive Fraud
Service & dialogue 2026 safely act on caller requests during a banking conversation while checking authorization, identity, and policy compliance at every step; retrieve applicable rules from a 698-document internal policy corpus before permitting a sensitive action; detect and refuse chained/adaptive fraud attempts where an earlier probe or admission makes a later, superficially valid request unsafe The agent must track the caller's identity/authorization claims, any tool access it has already granted, and prior probes/admissions across the conversation, checking each new request against the accumulated history and a 698-document policy corpus. DAG-with-precedence yes none stated 0
IHBench
IHBench: Evaluating Post-Interruption Recovery in Voice Agents with Structured Workflows
Service & dialogue 2026 resume a state-machine-driven workflow at the correct step after a user interruption; address the content of the user's interjection; avoid re-delivering content the user already heard; achieve task fulfillment across 10 enterprise domains despite one of 6 injected interruption types The agent must track its position in a state-machine-driven workflow, which content has already been delivered to the user, and the content of the user's interjection, in order to both recover to the correct step and adequately address the interruption. sequential-chain yes none stated 41.33/mo
JourneyBench
Beyond IVR: Benchmarking Customer Support LLM Agents for Business-Adherence
Service & dialogue+ Business & enterprise 2026 adhere to multi-step business policies/SOPs throughout a support conversation; navigate task dependencies correctly as the conversation unfolds; remain robust to unpredictable user/environment behavior while covering the required user-journey graph paths The agent must track which policy-graph nodes/steps it has satisfied so far, adhere to multi-step business rules, and adapt to unpredictable user/environment behavior across the conversation. DAG-with-precedence yes none stated 111.38/mo
SAGE (Service Agent Graph-guided Evaluation) / SAGE-Bench
SAGE: A Service Agent Graph-guided Evaluation Benchmark
Service & dialogue+ Web & GUI 2026 follow every step of an unstructured Standard Operating Procedure formalized as a Dynamic Dialogue Graph; correctly classify diverse/adversarial user intents and derive the correct subsequent action for each The agent must track which SOP-graph node/path it is currently on, the user's evolving (possibly adversarial) intent, and maintain logical compliance across the full dialogue, evaluated at increasing dialogue depths (turns 1, 5, 10, 15, and final turn). DAG-with-precedence yes dialogue depths evaluated at turns 1, 5, 10, 15, and the final turn, to measure stability over extended interactionsturns 30.6/mo
SEATauBench
SEATauBench: Adapting Tool-Agent-User Evaluation Into Low-Resource Southeast Asian Languages
Service & dialogue 2026 carry out tau2-Bench-style tool-agent-user tasks correctly when the conversation language changes; carry out the same tasks correctly as tool specifications are localized into the target SEA language; carry out the same tasks correctly as the task domain itself is localized Agent must correctly interpret and act on user requests, tool specifications, and domain content in the target language while adhering to the policy constraints originally defined in tau2-Bench. set-of-independent unclear none stated 10.33/mo
SpeechGym
SpeechGym: An Audio-Native Gym for Training Voice Agents via Reinforcement Learning
Service & dialogue 2026 correctly hear and call tools/fill argument slots based on values perceived from native audio (no ASR/TTS); hold multi-turn dialogue entirely through speech to complete the same tasks as an established text agentic benchmark; avoid unauthorized write actions under an insistent caller's social pressure Agent must track dialogue state and correctly perceived argument values purely from native audio across a multi-turn tool-use session, since a single misheard value causes cascading failures that consume the step budget. sequential-chain no none stated 0
T1-Bench
T1-Bench: Benchmarking Multi-Scenario Agents in Real-World Domains
Service & dialogue 2026 complete interleaved customer-facing scenarios spanning 25 domains within one interaction; conduct structured reasoning across multi-turn user-assistant interactions; correctly use tools while maintaining conversational quality across compositionally complex, interwoven scenario threads Agent must track tool use, conversational state, and multiple concurrent reasoning threads across interleaved scenarios spanning 25 domains within multi-turn user-assistant interactions. DAG-with-precedence no none stated 0
tau-Voice
τ-Voice: Benchmarking Full-Duplex Voice Agents on Real-World Domains
Service & dialogue 2026 complete a verifiable grounded task (extended from tau2-bench) via complex multi-turn conversation; adhere to domain policies throughout the conversation; correctly interact with/act on the environment (not just converse) to satisfy the task The agent must track conversational state, domain-policy constraints, and environment actions across complex multi-turn conversations, while also managing real-time full-duplex audio interaction (turn-taking, interruptions, accents, noise) alongside the underlying task goals. sequential-chain yes none stated 193.17/mo
unnamed user-oriented multi-turn tool-use dialogue generation pipeline
User-Oriented Multi-Turn Dialogue Generation with Tool Use at scale
Service & dialogue 2026 complete multiple distinct task completions accumulating within a single long conversational trajectory; respond appropriately to a simulated user's incremental, turn-by-turn requests and feedback rather than resolving the whole task in one shot The agent must track the state of multiple, potentially overlapping task completions within one long dialogue, correctly interpreting a user simulator's incremental, turn-by-turn requests and feedback rather than resolving everything from an initial fully-specified prompt. sequential-chain unclear none stated 20.25/mo
τ-Knowledge (τ-Banking domain)
τ-Knowledge: Evaluating Conversational Agents over Unstructured Knowledge
Service & dialogue 2026 retrieve the correct policy document(s) from a densely interlinked, ~700-document knowledge base; coordinate retrieved natural-language knowledge with tool outputs to execute a policy-compliant account update; produce a verifiable, policy-compliant state change during a live customer-support interaction Agent must track which of many interconnected banking policy documents are relevant to the current request, reconcile them with live tool outputs, and ensure the resulting account update remains policy-compliant and verifiable. DAG-with-precedence unclear none stated 254.17/mo
AgentChangeBench
AgentChangeBench: A Multi-Dimensional Evaluation Framework for Goal-Shift Robustness in Conversational AI
Service & dialogue+ Business & enterprise 2025 complete the original task objective before a mid-dialogue goal shift; recognize and adapt to a mid-dialogue goal shift triggered by one of five user personas; recover task success after a goal shift within enterprise domains (e.g., airline booking, retail); use tools efficiently and non-redundantly while adapting to shifting goals The agent must track the original task objective, detect when a mid-dialogue goal shift has occurred, measure its own recovery latency (Goal-Shift Recovery Time), and avoid redundant tool calls (Tool Call Redundancy Rate) while re-establishing progress toward the new goal. sequential-chain no none stated 30.27/mo
ECom-Bench
ECom-Bench: Can LLM Agent Resolve Real-World E-commerce Customer Support Issues?
Service & dialogue 2025 resolve a real-world e-commerce customer-support issue via multimodal interaction; adapt to a dynamic, persona-driven simulated user across the dialogue; handle diverse business scenarios reflecting real-world complexity; achieve consistent success across repeated trials (pass^3 metric) The agent must track the evolving state of a persona-driven simulated customer's issue, multimodal evidence presented during the conversation, and consistency of resolution across repeated trials. sequential-chain no none stated 302.14/mo
Food4All
Food4All: An Agentic Framework and Benchmark for Food Resource Navigation with Adaptive User Understanding
Service & dialogue+ Embodied & robotics 2025 ground a user's underspecified/noisy help-seeking dialogue into a valid resource recommendation; retrieve resources satisfying single food needs; satisfy composite cases with access or document constraints; handle non-ideal user interaction traits (unreasonable demands, rambling, impatience, incomplete answers, inconsistent information) while completing the referral The agent must track grounded requirements (schedule, eligibility, intake, document constraints), the set of valid retrieved resources it must preserve into the final recommendation, and the user's non-ideal interaction trait across the dialogue. sequential-chain yes none stated 10.09/mo
IntellAgent
IntellAgent: A Multi-Agent Framework for Evaluating Conversational AI Systems
Service & dialogue+ Multi-agent orgs 2025 navigate a multi-turn dialogue while integrating domain-specific APIs; adhere to strict, graph-modeled policy constraints throughout the conversation; handle realistic, policy-driven event scenarios generated at varying complexity levels The agent must track which of several interacting domain-specific policy constraints apply at each point in a multi-turn dialogue, per a graph-based policy model, while integrating API calls consistent with those constraints. DAG-with-precedence no none statedother:not-stated 261.3/mo
tau-break
Effective Red-Teaming of Policy-Adherent Agents
Service & dialogue 2025 adhere consistently to domain policies (e.g. refund eligibility, cancellation rules) across a customer-service conversation; correctly refuse any request that would violate policy; remain helpful and natural while resisting persuasive, policy-aware adversarial pressure from CRAFT The agent must track which policies apply to the current request and remain consistent with them across up to 30 dialogue turns, resisting cumulative persuasive pressure (emotional manipulation, coercive framing) from an adversarial user. sequential-chain no up to 30 dialogue turnsturns 110.73/mo
τ²-Bench
τ2-Bench: Evaluating Conversational Agents in a Dual-Control Environment
Service & dialogue 2025 coordinate actions with an active user who also uses tools to modify a shared, dynamic environment (dual-control); complete diverse, compositionally-generated verifiable tasks built from atomic components; guide/communicate with the user effectively, not only reason internally; correctly attribute/avoid errors arising from reasoning vs communication/coordination failures The agent must track the shared dynamic world state as modified by both itself and the user, communicate effectively to guide the user's tool use, and maintain a compositional task's atomic sub-requirements across the dual-control interaction. DAG-with-precedence no none stated 45630.4/mo
Education & tutoring2 artifacts
TeachArena
TeachArena: Are Language Agents Ready for Realistic Teaching Work?
Education & tutoring+ Business & enterprise 2026 infer a warranted pedagogical/teaching decision from evidence (professional pedagogical judgment); adapt tutoring support as the learner's state changes across a multi-turn tutoring session (situated multi-turn tutoring); carry an instructor's request through a learning-management system (LMS) to a completed, verified intervention (end-to-end LMS teaching workflow) The agent must track the learner's evolving state across a multi-turn tutoring session, evidence supporting a pedagogical insight, and the instructor's request as it is carried through to a persistent, verified artifact or environment state in the LMS. hierarchical no none stated 0
TutorBench
DeepTutor: Towards Agentic Personalized Tutoring
Education & tutoring+ Assistants & memory 2026 deliver citation-grounded tutoring on a specific problem; generate difficulty-calibrated follow-up questions matched to a learner's diagnosed knowledge gaps; continuously adapt personalization to the student's evolving needs across an interactive tutoring session The agent must track the learner's diagnosed knowledge gaps and profile, the history of prior interactions, and adapt difficulty calibration accordingly across a multi-turn tutoring dialogue. sequential-chain yes none statedother:not-stated 0
Embodied & robotics46 artifacts
CoCoBench
CoCoBench: A Cooperative Coordination Benchmark for Embodied Multi-Agent Task Planning
Embodied & robotics+ Multi-agent orgs 2026 allocate tasks correctly among cooperating embodied agents; respect sequential-ordering constraints between agents' actions; respect mutual-exclusion constraints over shared objects/spaces; execute correct handoff coordination between agents; satisfy conjunctive multi-object-destination goals within executable household tasks The multi-agent system must track task allocation assignments, ordering/precedence constraints, mutual-exclusion locks on shared objects, and handoff status between agents, calibrated against a task-specific step budget H derived from oracle trajectories. DAG-with-precedence yes ~18.4 mean executed action steps (best model, successful episodes); task-specific step budget H calibrated from oracle-validated trajectoriesagent-steps 0
DeCoNavBench
DeCoNav: Dialog enhanced Long-Horizon Collaborative Vision-Language Navigation
Embodied & robotics+ Multi-agent orgs 2026 achieve relay-style handoffs between two robots collaborating on a shared long-horizon navigation task; reach rendezvous points to exchange information/objects between robots; dynamically reassign and replan subgoals in response to new cross-agent evidence, uncertainty, or conflicts Each robot must track its own navigation progress plus the other robot's evidence/uncertainty/conflicts communicated via event-triggered dialogue, in order to reassign subgoals and replan under synchronized execution. DAG-with-precedence yes none stated 51.0/mo
DunphyBench
Long-Horizon Embodied Decision-Making via Multimodal Memory Compression
Embodied & robotics+ Assistants & memory 2026 navigate through multiple embodied housing environments to gather evidence; integrate multimodal, multi-source input into coherent knowledge under partial observation; make a final housing decision aligned with multi-dimensional, partly implicit human preferences Agent must accumulate multimodal evidence across multiple candidate housing environments under partial observation and integrate it into a coherent decision aligned with multi-dimensional, partly implicit human preferences. sequential-chain no none stated 0
EAS resilience evaluation framework
Resilience Matters for Embodied Agents System: New Metrics, Systematic Evaluation, and Optimization
Embodied & robotics 2026 complete a household task despite perturbations/unexpected disruptions during execution; recover, stabilize, and gracefully extend behavior after a perturbation rather than simply succeeding or failing outright; maintain low recovery cost and stability across iterative system updates The evaluation layer must track the full execution trajectory of each household task under perturbation, not just final success/failure, to compute process-level resilience metrics like recovery cost and stability. hierarchical yes none stated 0
Ego2World
Ego2World: Compiling Egocentric Cooking Videos into Executable Worlds for Belief-State Planning
Embodied & robotics 2026 plan under partial observation using a belief graph built from local observations; remember objects and track state changes in a hidden symbolic world graph; recover and replan when actions fail without observing the true world state; complete cooking-task objectives via correct sequences of graph-governed state transitions The agent must maintain and update its own partial belief graph of the world (objects, state changes) using only local observations and execution feedback, separate from the simulator's hidden ground-truth world graph, and replan without directly observing true state. DAG-with-precedence no none stated 0
EmbodiedWorldBench (ABot-AgentOS)
ABot-AgentOS: A General Robotic Agent OS with Lifelong Multi-modal Memory
Embodied & robotics+ Assistants & memory 2026 navigate indoor/outdoor/hybrid scenes to reach targets; search for and identify specified objects; conduct NPC dialogue as part of a task; respond appropriately to dynamic events introduced mid-task; achieve trace-grounded, verifiable task completion across difficulty levels Agent must maintain a persistent, source-grounded multi-modal graph memory across dialogue, visual observations, spatial context, temporal relations, and task traces, continually updated and consulted for later decisions. hierarchical unclear none stated 0
FullHome (via the TaskGround Ground-Infer-Execute framework)
TaskGround: Structured Executable Task Inference for Full-Scene Household Reasoning
Embodied & robotics 2026 identify task-relevant entities within a complete, cluttered household scene; recover intended task conditions implied by a situated (underspecified) household request; resolve ordering constraints among sub-actions from surrounding scene context; produce a grounded, skill-level action sequence that correctly executes the inferred task structure The agent must track which entities, conditions, and ordering constraints it has inferred as task-relevant from a complete household scene, and maintain this executable task structure while compiling and executing a grounded skill-level action sequence. sequential-chain no none stated 0
HumanCLAW-Bench
HumanCLAW: Can Vision-Language Models Act Through a Body?
Embodied & robotics 2026 find a specified target within an indoor scene; navigate to the target while avoiding obstacles/maintaining balance; interact with the target once reached; maintain embodied self-awareness of the body's own state (position, goal-reached status, collisions) throughout The VLM must track where its body is, whether it has reached the goal, and whether it has hit an obstacle -- embodied self-awareness -- across each find-navigate-interact episode, since the paper finds this tracking, not target recognition, is the main bottleneck. sequential-chain yes reported average of 78.5 steps per episode for baseline conditionsagent-steps 21.0/mo
IMBench
IMBench: A Benchmark for Intuitive Robotic Manipulation
Embodied & robotics 2026 infer task-relevant physical structure via physical reasoning before acting; generate a feasible action sequence satisfying explicit task constraints (contact-rich manipulation, tool use, multi-stage dependencies); complete each of 35 manipulation tasks across scalable, diverse scenarios Agent must track inferred physical structure (contacts, affordances), the current stage of a multi-stage manipulation plan, and constraint satisfaction across execution. DAG-with-precedence yes none stated 0
LMEE-Bench
Explore with Long-term Memory: A Benchmark and Multimodal LLM-based Reinforcement Learning Framework for Embodied Exploration
Embodied & robotics+ Assistants & memory 2026 perform multi-goal navigation across an embodied environment; answer memory-based questions using episodic memory accumulated during exploration; unify exploratory cognition with decision-making to support lifelong learning; proactively query memory and select exploration frontiers next The agent must track its accumulated episodic memory, current exploration frontier, and multiple concurrent navigation goals across a long-horizon embodied exploration episode. sequential-chain no none stated 101.25/mo
LongAct
When Robots Do the Chores: A Benchmark and Agent for Long-Horizon Household Task Execution
Embodied & robotics 2026 understand a free-form (non-templated) household instruction and decompose it into the right sequence of sub-tasks; manage dependencies among sub-tasks (ordering, shared resources) via a DAG-based plan; maintain persistent memory (spatial and episodic) across a long-horizon household task; adapt the plan reflectively as execution proceeds The agent must maintain persistent spatial and episodic memory of what it has already done and observed in the household, use this to keep dependency-aware track of remaining sub-tasks, and adaptively replan as new information arises over the long-horizon task. DAG-with-precedence yes human executors typically require 500+ steps; VLM-based agents often exceed 2,000 steps to complete the same long-horizon household taskagent-steps 20.5/mo
MECoBench
MECoBench: A Systematic Study of Multimodal Agent Collaboration in Embodied Environments
Embodied & robotics+ Multi-agent orgs 2026 complete embodied tasks under two cooperation structures and three collaboration modes; coordinate communication among multiple multimodal agents in a visually grounded environment; remain robust to noisy priors and exploration conditions via collaboration Agents must track and communicate task-relevant state among team members across two cooperation structures and three collaboration modes to complete embodied tasks under noisy priors and exploration conditions. set-of-independent no none stated 20.67/mo
MultiUAV-Plat Benchmark
MultiUAV-Plat: An LLM-Oriented Platform, Benchmark and Framework for Multi-UAV Collaborative Task Planning
Embodied & robotics+ Multi-agent orgs 2026 assign UAVs to targets correctly under partial observability; perform area search coverage objectives; perform area assignment and patrol objectives across multiple UAVs; satisfy each of the mission's validation checks (mean 6.26/task) via correct multi-vehicle coordination The framework (Agent4Drone) must track memory, observation, task understanding, planning, execution, and verification state across each mission, checking role-based information access and validation logic per task within a session. set-of-independent yes none statedother:not-stated 0
PARTNR-Dialog
LLawCo: Learning Laws of Cooperation for Modeling Embodied Multi-Agent Behavior
Embodied & robotics+ Multi-agent orgs 2026 complete a shared household task cooperatively with a partner agent; communicate with the partner (e.g., 'talk when necessary') to align on task objectives; align actions with the partner's behavior and the environment state to avoid inefficient or conflicting actions; apply learned high-level behavioral laws (e.g., 'wait for partner') during planning The agent must track its partner's current actions/state, the shared household task's progress, and which high-level behavioral laws (e.g., 'talk when necessary,' 'wait for partner') apply at each point in the cooperative plan. sequential-chain no none stated 0
PLanAR
PLanAR: Planning-Language-Grounded Agentic Reasoning for Robot Manipulation
Embodied & robotics 2026 represent and update object-predicate scene states as manipulation proceeds; select and sequence action schemas (with preconditions/effects) to satisfy an open-vocabulary manipulation goal; detect execution failures via stepwise symbolic-effect verification and replan; complete long-horizon kitchen workflows composed of many chained manipulation sub-goals The agent must maintain a symbolic scene-state representation (object predicates), verify after each action whether expected effects were achieved, and update/replan when execution deviates from the expected symbolic plan across a long-horizon kitchen workflow. DAG-with-precedence no none statedother:not-stated 30.43/mo
RescueBench
RescueBench: Can Embodied Agents Save Lives in the Wild ?
Embodied & robotics 2026 explore an unfamiliar environment under multimodal uncertainty to locate a target (multimodal exploration stage); physically rescue/reach the identified target (target rescue stage); navigate back using retained spatial memory (memory-guided return stage); complete a final handoff of the rescued target/information (final handoff stage) The agent must retain spatial memory of the environment (for the return stage), track clue ambiguity/target identification state, and manage per-level time budgets (180-300 seconds) across the four-stage pipeline. sequential-chain yes episodes are governed by level-dependent time limits: L1-L2: 180 seconds, L3: 240 seconds, L4-L5: 300 seconds, with no step capother:wall-clock-seconds 0
RoboGraph
Compiling and Benchmarking Task-State Horizons for Embodied Agents
Embodied & robotics 2026 track the evolving span of task-relevant state transitions (task-state horizon, TSH) induced by both exploration and environmental dynamics; correctly execute high-level plans as task-relevant world state changes, including due to unexpected failures/interventions; complete each of 588 episodes across 84 scenes with varying task-state horizons The agent must maintain, explore, and update task-relevant world state over the course of a long-horizon rollout, since performance is explicitly measured as a function of the task-state horizon (TSH) -- the span of state transitions it must track. DAG-with-precedence yes none stated 0
SMH-Bench (built on HomeEnv)
SMH-Bench: Benchmarking LLM Agents for Environment-Grounded Reasoning and Action in Smart Homes
Embodied & robotics 2026 execute explicit device control and query commands across a smart home with up to 135 devices; schedule automation tasks that must fire correctly over time; correctly handle ambiguous user instructions; personalize reasoning to user intent/preferences as home complexity increases Agent must track the state of many concurrent devices across a multi-room home, especially for automation-scheduling tasks that must persist and correctly re-trigger over time, and must resolve ambiguous instructions against current home/device state. set-of-independent yes none stated 20.67/mo
SpatialWorld
SpatialWorld: Benchmarking Interactive Spatial Reasoning of Multimodal Agents in Real-World Tasks
Embodied & robotics 2026 actively gather egocentric visual evidence under partial observability to resolve a task; complete household, travel, or social-collaboration tasks requiring interactive spatial reasoning; express decisions via a unified text-based action interface across eight heterogeneous simulator backends Agent must track what it has and has not yet observed (partial observability), accumulate egocentric visual evidence over the course of the task, and reconcile this with a reference trajectory/terminal-state verifier. sequential-chain unclear none stated 20.67/mo
TypeGo (Kalos prototype)
TypeGo: An OS Runtime for Embodied Agents
Embodied & robotics 2026 execute multiple concurrent per-task processes/goals on a shared physical robot body without conflicting over physical subsystems; preempt, resume, or replace a lower-priority task/goal when a new goal arrives; maintain fast first-action responsiveness while longer-horizon planning continues asynchronously in the background The runtime must track which per-task processes currently hold which physical subsystems, their priority/preemption state, and pending speculative skill-streaming actions, to arbitrate concurrent goals in real time. set-of-independent yes none stated 0
UniETP
UniETP: Unifying Environments for Generalizable Embodied Task Planning
Embodied & robotics 2026 execute a sequence of atomic actions within an interactive environment to complete a user-specified task; handle varying task-logic complexity; handle varying instance-grounding complexity; handle varying instruction-understanding complexity, across four unified simulators (AI2-THOR, VirtualHome, Habitat, BEHAVIOR) The agent must track its progress executing a sequence of atomic actions toward a user-specified goal within a unified observation/action space, while contending with varying levels of task-logic, instance-grounding, and instruction-understanding difficulty across four different underlying simulators. sequential-chain yes none stated 0
VGEBench
Towards Generalizable Visually Grounded Exploration of Household Devices
Embodied & robotics 2026 form a hypothesis about how to operate a novel household device from visual cues alone (no manual); test the hypothesis via physical interaction and interpret feedback; refine/correct the hypothesis and action based on observed feedback (Hypothesis-Interaction-Refinement loop) until the device is successfully operated The agent must track its current hypothesis about the device's operation, the history of physical feedback received from prior interaction attempts, and dynamically-calculated interaction budgets based on task complexity, maintaining long-horizon state across the exploration loop. open-ended yes none stated 0
WorldLines
WorldLines: Benchmarking and Modeling Long-Horizon Stateful Embodied Agents
Embodied & robotics 2026 answer Memory QA questions correctly using long-term household interaction history; produce correct Embodied Task Plans grounded in remembered user routines and past world/device states; maintain visibility-aware (partially observable) memory across a temporally extended household trace The agent must maintain visibility-aware memory of user routines, world states, past actions/dialogues, and object/device state changes across long temporally extended household traces to answer Memory QA and produce embodied plans. other (singleton) yes none stated 10.33/mo
Household Task Planning with Multi-Objects State and Relationship Using Large Language Models Based Preconditions Verification
Household Task Planning with Multi-Objects State and Relationship Using Large Language Models Based Preconditions Verification
Embodied & robotics 2025 achieve each targeted object state change (e.g. turning an appliance on/off); achieve each targeted object placement goal; verify environmental preconditions are met before executing each action; reformulate an action step automatically when a precondition is not satisfied The agent must track current object states, identifiers, and relationships in the environment, re-verifying preconditions before each action and updating its plan when environmental state does not match expectations. sequential-chain no none statedother:not-stated 0
3DMem-Bench (via 3DLLM-Mem)
3DLLM-Mem: Long-Term Spatial-Temporal Memory for Embodied 3D Large Language Model
Embodied & robotics+ Assistants & memory 2025 correctly recall and apply past spatial-temporal observations (episodic memory) to complete embodied tasks in multi-room 3D environments; answer questions and produce captions grounded in accumulated long-term memory of the 3D scene; focus on task-relevant information while maintaining memory efficiency across complex, long-horizon environments The agent must maintain and selectively query an episodic memory of past spatial and temporal observations across long-horizon, multi-room 3D trajectories, focusing on task-relevant information while remaining memory-efficient. sequential-chain yes none stated 291.81/mo
ArtiBench (with ArtiBrain)
ArtiBench and ArtiBrain: Benchmarking Generalizable Vision-Language Articulated Object Manipulation
Embodied & robotics 2025 manipulate articulated objects (kitchen, storage, office, tool appliances) correctly across parts/instances/categories; decompose and validate a sequence of sub-goals for a long-horizon, multi-object manipulation task; maintain physical consistency across the multi-step interaction System must track subgoal validation state (via the Task Reasoner), accumulated part-level affordances in the Affordance Memory Bank, and physical consistency across a five-level benchmark spanning cross-part/instance/category variation to long-horizon multi-object tasks. hierarchical unclear none stated 0
Blocksworld-MCP benchmark
Benchmark for Planning and Control with Large Language Model Agents: Blocksworld with Model Context Protocol
Embodied & robotics 2025 reach one target block-configuration goal state via a sequence of pick/put/stack/unstack actions; satisfy the goal under increasing plan-length complexity categories (step-2 through step-12); generalize plan validity/optimality across diverse agent architectures connected via a standardized MCP tool interface The agent must track the current block/stack configuration as pick/put/stack/unstack actions are executed, correctly planning toward the target configuration across plans of up to 12 optimal steps, potentially under constrained block size or partial observability. sequential-chain no scenarios span step-2 through step-12 optimal-plan-length categories (45 step-2, 84 step-4, 152 step-6, 151 step-8, 112 step-10, 46 step-12 scenarios), each involving up to five blocksagent-steps 20.22/mo
CookBench
CookBench: A Long-Horizon Embodied Planning Benchmark for Complex Cooking Scenarios
Embodied & robotics 2025 accurately parse a user's complex cooking intent (Intention Recognition); execute the identified cooking goal through a long-horizon, fine-grained sequence of physical actions (Embodied Interaction); correctly use both macro-level operations (placing orders, purchasing ingredients) and fine-grained embodied physical actions The agent must track the parsed cooking intent, its progress through a long-horizon fine-grained sequence of physical actions, and the state of ingredients/tools obtained via macro-level operations, across a two-stage cooking task. sequential-chain no none stated 131.0/mo
DeCoBench
DeCo: Task Decomposition and Skill Composition for Zero-Shot Generalization in Long-Horizon 3D Manipulation
Embodied & robotics 2025 retrieve and chain reusable atomic manipulation skills to complete a compositional long-horizon 3D manipulation task; generalize zero-shot to novel task compositions not seen during training; execute smooth, collision-free transitions between chained skills The system must track the currently retrieved skill, the object/gripper state at each transition point, and which atomic subtask in the composed sequence is currently active to ensure valid chaining. sequential-chain yes none stated 161.0/mo
DeliveryBench
DeliveryBench: Can Agents Earn Profit in Real World?
Embodied & robotics 2025 maximize net profit over the course of an operating shift by choosing which deliveries to accept/complete; meet each accepted delivery's deadline; manage limited resources (transportation expense, vehicle battery) across the shift; interact appropriately with other couriers and customers as needed The agent must track remaining budget/vehicle battery, delivery deadlines for all currently accepted orders, and its evolving location within a procedurally generated 3D city, over an episode lasting several in-game hours and typically more than 100 action steps. set-of-independent yes episodes support long-horizon tasks spanning several in-game hours and typically more than 100 action stepsagent-steps 50.56/mo
Embodied Web Agents Benchmark
Embodied Web Agents: Bridging Physical-Digital Realms for Integrated Agent Intelligence
Embodied & robotics+ Web & GUI 2025 cook a recipe found via web-scale reasoning, using embodied physical actions (cooking); navigate physically using dynamic, web-sourced map data (navigation); shop by combining physical store interaction with online product/price information (shopping); plan tourism activities combining physical exploration with web knowledge (tourism); identify real-world landmarks by cross-referencing physical observation with web knowledge (geolocation) Agent must maintain a consistent state across both a 3D embodied environment (physical position, observations, actions) and a web interface (retrieved facts, map data, product/price info), integrating the two continuously rather than treating them as separate phases. set-of-independent unclear none stated 191.27/mo
EmbodiedBench
EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied Agents
Embodied & robotics 2025 complete each of 1,128 testing tasks across four environments, from high-level semantic household tasks to low-level atomic navigation/manipulation; demonstrate commonsense reasoning within tasks; understand complex instructions; exhibit spatial awareness and visual perception; engage in long-term planning The agent must track visual perception of the scene, instruction understanding, spatial layout, and long-term planning state as it composes low-level atomic actions to satisfy high-level household goals. hierarchical yes none stated 22511.84/mo
EmbodiedBrain evaluation suite (General, Planning, and End-to-End Simulation Benchmarks)
EmbodiedBrain: Expanding Performance Boundaries of Task Planning for Embodied Intelligence
Embodied & robotics 2025 complete long-horizon embodied task-planning sequences by correctly building on preceding steps (Guided Precursors); perform accurate spatial perception and adaptive execution across a novel simulation environment; satisfy General, Planning, and End-to-End Simulation benchmark criteria as three complementary evaluation axes Agent must track its own preceding action steps as guided precursors informing subsequent steps, and its progress must be verifiable across all three of the General, Planning, and End-to-End Simulation evaluation axes. sequential-chain yes none stated 30.27/mo
EMMOE
EMMOE: A Comprehensive Benchmark for Embodied Mobile Manipulation in Open Environments
Embodied & robotics+ Web & GUI 2025 interpret a natural-language user instruction and execute a long-horizon everyday household task combining high-level and low-level embodied sub-tasks; re-plan after execution failures to recover progress toward the instructed goal Agent must track task-completion progress across high-level sub-tasks and low-level actions, plus failure/replan history, in continuous physical space across a long-horizon episode. hierarchical yes none stated 40.22/mo
FindingDory
FindingDory: A Benchmark to Evaluate Memory in Embodied Agents
Embodied & robotics+ Assistants & memory 2025 recall relevant historical information (images/interactions) collected potentially across multiple days; execute low-level navigation/manipulation actions based on recalled information; complete each of 60 memory-intensive embodied tasks requiring sustained engagement; scale to procedurally extended, longer/harder versions of the same tasks The agent must retain and retrieve relevant historical images/interactions collected across multiple days in the Habitat simulator, combining that recall with sustained contextual awareness during ongoing navigation/manipulation. sequential-chain yes none stated 120.8/mo
HiMan-Bench
RoboHiMan: A Hierarchical Evaluation Paradigm for Compositional Generalization in Long-Horizon Manipulation
Embodied & robotics 2025 complete atomic long-horizon manipulation tasks under diverse perturbations; complete compositional tasks requiring composing multiple learned manipulation skills; generalize skill composition/scheduling to perturbed or novel conditions; coordinate high-level subgoal planning with low-level execution policies The system must track which atomic/composed skill is currently executing, how perturbations affect preconditions for subsequent skills, and whether the high-level plan needs revision given low-level execution feedback. hierarchical yes none stated 90.82/mo
LLM-BabyBench
LLM-BABYBENCH: Understanding and Evaluating Grounded Planning and Reasoning in LLMs
Embodied & robotics 2025 predict the consequences of an action on the textual BabyAI grid-world environment state (Predict task); generate a sequence of low-level actions achieving a specified objective (Plan task); decompose a high-level instruction into a coherent sequence of subgoals (Decompose task) Agent must track the current grid-world state to predict action consequences, track partial plan progress when generating low-level action sequences, and track subgoal ordering/coherence when decomposing high-level instructions. hierarchical yes none stated 60.38/mo
LoHoSet (Ravens simulator)
LoHoVLA: A Unified Vision-Language-Action Model for Long-Horizon Embodied Tasks
Embodied & robotics 2025 decompose a high-level embodied goal into a sequence of sub-tasks; generate correct low-level robot actions to execute each sub-task; maintain coordination between high-level planning and low-level motion control across the whole long-horizon task The system must track which sub-tasks of the decomposed high-level goal have been completed and maintain closed-loop consistency between the high-level plan and low-level action execution across the task. hierarchical yes none stated 322.0/mo
ManiTaskGen
ManiTaskGen: A Comprehensive Task Generator for Benchmarking and Improving Vision-Language Agents on Embodied Decision-Making
Embodied & robotics 2025 satisfy a process-based instruction requiring a specific sequence of manipulations (e.g. 'move object from X to Y'); satisfy an outcome-based abstract instruction requiring multiple manipulations to reach a goal state (e.g. 'clear the table'); generate a comprehensive, diverse, feasible set of mobile manipulation tasks for any given scene The agent (and the task generator itself) must track which objects have been moved/placed so far relative to the target scene configuration, verifying that an outcome-based goal (e.g. a cleared table) is only satisfied once all constituent object states are achieved. hierarchical no none statedother:not-stated 40.25/mo
OceanGym
OceanGym: A Benchmark Environment for Underwater Embodied Agents
Embodied & robotics 2025 comprehend and fuse optical and sonar sensor data under low visibility; autonomously explore complex underwater environments; accomplish each of eight realistic underwater task domains; adapt navigation/decision-making to dynamic ocean currents The agent must integrate perception, memory, and sequential decision-making state (explored regions, sonar/optical evidence) across a long-horizon objective under low-visibility, dynamically-changing underwater conditions. hierarchical no 0.5 hours (t_max, per decision task)wall-clock-hours 0
REMAC multi-agent environment (on RoboCasa)
REMAC: Self-Reflective and Self-Evolving Multi-Agent Collaboration for Long-Horizon Robot Manipulation
Embodied & robotics+ Multi-agent orgs 2025 decompose and execute long-horizon multi-robot manipulation and navigation tasks (4 task categories, 27 task styles, 50+ objects); perform pre-condition and post-condition checks in the loop to evaluate progress and refine plans; adapt plans dynamically to unexpected scene conditions (e.g. a closed microwave door) via self-evolvement; coordinate parallel task execution across multiple robots Agent(s) must track scene state, pre/post-condition satisfaction, and per-robot task assignment across a decomposed long-horizon manipulation/navigation plan, adapting the plan when scene-specific conditions invalidate prior assumptions. DAG-with-precedence no none stated 110.61/mo
ResponsibleRobotBench
ResponsibleRobotBench: Benchmarking Responsible Robot Manipulation using Multi-modal Large Language Models
Embodied & robotics 2025 detect and mitigate risks (electrical, chemical, human-related hazards) during a multi-stage manipulation task; reason about safety and physically grounded planning across the task; plan and execute sequences of manipulation actions; engage human assistance when necessary rather than proceeding unsafely Agent must track detected hazards (electrical, chemical, human-related), current safety status, and multi-stage plan progress, deciding when to escalate to human assistance across the task. hierarchical yes none stated 30.33/mo
RoboCerebra
RoboCerebra: A Large-scale Benchmark for Long-horizon Robotic Manipulation Evaluation
Embodied & robotics 2025 decompose a high-level household instruction into a sequence of dependent subtasks (via GPT-generated instructions); correctly execute each subtask in the sequence despite dynamic object variations; sustain planning, reflection, and memory (System 2 reasoning) across the full extended action sequence, not just react (System 1) to the current observation The high-level planner must track evolving object/environment state across an average trajectory length of 2,972.4 simulation steps (about 6x longer than existing long-horizon manipulation datasets), correctly sequencing and reflecting on subtasks throughout. sequential-chain yes average trajectory length reaches 2,972.4 simulation steps, about 6x longer than existing long-horizon manipulation datasetsagent-steps 332.2/mo
RoboPilot-Bench
RoboPilot: Generalizable Dynamic Robotic Manipulation with Dual-thinking Modes
Embodied & robotics 2025 execute complex or long-horizon robotic manipulation tasks despite environmental changes; recognize infeasible tasks rather than attempting them blindly; recover from execution errors via closed-loop replanning; succeed across each of 21 tasks spanning 10 manipulation categories The agent must track execution feedback and environmental state to detect deviations or errors requiring replanning, and separately recognize when a task is infeasible altogether, across a long-horizon sequence of primitive manipulation actions. sequential-chain no none stated 10.08/mo
Robotouille
Robotouille: An Asynchronous Planning Benchmark for LLM Agents
Embodied & robotics 2025 complete overlapping cooking sub-tasks that must be scheduled around each other (e.g. one dish while another cooks); handle interruptions to an in-progress plan without dropping earlier commitments; reason over states/actions that must occur in parallel versus strictly sequentially Agent must track the state and expected completion time of multiple concurrently in-progress cooking actions, remember interrupted sub-goals, and self-audit its plan as time delays resolve. DAG-with-precedence yes none stated 382.0/mo
unnamed Daily Composite Tasks benchmark
Evaluating Multimodal Large Language Models with Daily Composite Tasks in Home Environments
Embodied & robotics 2025 correctly perform object-understanding sub-tasks (e.g. counting/categorizing objects) within a composite task; correctly perform spatial-intelligence sub-tasks within the same composite task; correctly perform social-activity sub-tasks within the same composite task, all within one dynamic simulated home environment The agent must track and integrate information across the three jointly-required capability domains (object understanding, spatial intelligence, social activity) within one dynamic, simulated home environment to complete a composite task. set-of-independent yes none stated 10.08/mo
Healthcare8 artifacts
AgentClinic
AgentClinic: a multimodal benchmark for tool-using clinical AI agents
Healthcare+ Business & enterprise 2026 engage in sequential clinical decision-making across a diagnostic encounter (patient interaction, exam/imaging requests); collect multimodal data under incomplete information via various tools (e.g., notebook, retrieval); reach a correct diagnosis across nine medical specialties and seven languages Agent must track patient information collected so far (multimodal exam/imaging results), notes taken across a case (persisting via a notebook tool), and its evolving working diagnosis across the encounter. sequential-chain yes none stated 214.2/mo
ClinEnv
ClinEnv: An Interactive Multi-Stage Long Horizon EHR Environment for Agents
Healthcare+ Business & enterprise 2026 progress through an ordered, per-case sequence of clinical decision stages; actively query four specialized information agents before committing to a decision at each stage; commit to correct medications, procedures, and diagnoses at each stage; avoid redundant information-gathering queries as the case progresses Agent must actively query four specialized information agents at every decision stage before committing to medications, procedures, and diagnoses, tracking what has already been queried to avoid redundant queries as the case progresses. sequential-chain yes none stated 20.67/mo
CodeClinic
CodeClinic: Evaluating Automation of Coding Skills for Clinical Reasoning Agents
Healthcare 2026 in the longitudinal ICU-surveillance task, make a structured monitoring decision every four hours across 25 findings and eight clinical families for the duration of a patient trajectory; in the compositional information-seeking task, answer queries whose difficulty is stratified by compositional dependency depth across 259 tasks in 9 domains; synthesize and compose reusable clinical skills rather than relying on a fixed toolbox Agent must track the patient's evolving clinical trajectory and make a fresh structured decision every four hours in the longitudinal task, and/or track which reusable clinical skills it has synthesized/composed so far in the compositional task. DAG-with-precedence yes every 4 hours (decision cadence within a longitudinal ICU patient trajectory; total trajectory length not stated)wall-clock-hours 0
Healink
Bridging the Post-discharge Gap: A Traceable Multi-agent Framework for Safe and Continuous Care
Healthcare+ Multi-agent orgs 2026 maintain continuity of care across longitudinal, patient-specific post-discharge follow-up interactions; generate prescription-grounded, traceable responses reflecting the correct patient/phenotypic/intervention history; prevent cross-departmental drug conflicts while integrating fragmented histories across clinical departments System must track a patient's full longitudinal clinical history (vectorized records, phenotypic/intervention dimensions) across departments and follow-up interactions, actively cross-checking for drug conflicts as new prescriptions are considered. DAG-with-precedence yes none stated 0
HealthAgentBench
HealthAgentBench: A Unified Benchmark Suite of Realistic Agentic Healthcare Environments for Challenging Frontier AI Agents
Healthcare+ Business & enterprise 2026 explore raw, heterogeneous healthcare data under minimal instructions; operate within a complex clinical environment to execute a multi-step, end-to-end task; develop research modeling pipelines over EHR data; complete tasks spanning 7 categories across the patient journey and multiple modalities The agent must track what it has already discovered while exploring raw healthcare data and its intermediate modeling/analysis steps, so its final multi-step solution reflects the full workflow rather than a shortcut response. sequential-chain yes none stated 72.33/mo
MedMCP-Calc
MedMCP-Calc: Benchmarking LLMs for Realistic Medical Calculator Scenarios via MCP Integration
Healthcare+ Business & enterprise 2026 proactively acquire relevant patient data from an EHR database; select the scenario-appropriate medical calculator among alternatives; perform multi-step computation using retrieved data and the selected calculator; retrieve external reference information when needed; complete each of 118 scenario tasks across 4 clinical domains The agent must track which EHR fields it has retrieved via iterative SQL-based database interaction, which calculator it has selected, and intermediate computed values, since evaluation is process-level rather than only checking the final number. sequential-chain yes none stated 30.38/mo
MedMemoryBench
MedMemoryBench: Benchmarking Agent Memory in Personalized Healthcare
Healthcare+ Assistants & memory 2026 accumulate and correctly retain/retrieve clinically relevant patient information across many sessions; sustain retrieval and reasoning robustness despite memory saturation from continued information influx; correctly evaluate an agent's memory 'live', as it is constructed, rather than only after the fact (evaluate-while-constructing streaming protocol) The agent's memory system must retain and correctly retrieve clinically relevant information across roughly 2,000 sessions and 16,000 interaction turns per patient archetype, resisting degradation from memory saturation and noise while supporting complex medical reasoning. sequential-chain yes approximately 2,000 sessions and 16,000 interaction turnssessions 10.25/mo
MedAgentSim
Self-Evolving Multi-Agent Simulations for Realistic Clinical Interactions
Healthcare+ Multi-agent orgs 2025 as doctor agent, request relevant medical examinations and imaging results from a measurement agent across multi-turn conversations to reach a diagnosis; iteratively refine diagnostic strategy across successive patient interactions via self-improvement mechanisms Doctor agent must track which exams/imaging results it has already requested and received, its evolving working diagnosis, and experience-based knowledge accumulated as it interacts with more patients over time. sequential-chain yes none stated 482.67/mo
Information seeking10 artifacts
DailyReport
DailyReport: An Open-ended Benchmark for Evaluating Search Agents on Daily Search Tasks
Information seeking+ Science 2026 autonomously explore web sources and synthesize information into a comprehensive response to an open-ended daily search query; satisfy each of a task's associated cascade rubrics across disentangled evaluation dimensions Search agent must track which subtasks/rubric dimensions of an open-ended query it has satisfied, aggregating cascade performance across disentangled dimensions into an interpretable, user-centric final score. hierarchical yes none stated 0
DeepSearchQA
DeepSearchQA: Bridging the Comprehensiveness Gap for Deep Research Agents
Information seeking+ Tool & API 2026 systematically collate fragmented information from disparate open-web sources; de-duplicate and resolve entities to ensure precision in the exhaustive answer list; reason about stopping criteria within an open-ended search space; complete each causal-chain step, where later steps depend on successful completion of the previous one The agent must track which sources it has already collated, de-duplicate overlapping entities, maintain the causal-chain dependency state (what has been successfully resolved so far), and decide when to stop searching within an open-ended web space. sequential-chain no none statedother:not-stated 394.88/mo
EarthVerse
EarthVerse: Benchmarking Scientific Agents Across Dynamic Earth Systems and Natural Hazards
Information seeking+ Science 2026 inspect heterogeneous event packages and choose compatible evidence for a natural-hazard investigation; execute transparent calculations reconciling differences across sources; preserve provenance across evidence, scales, units, and calculations in the final answer; produce each of the fine-grained answer units defined by the task's executable ground truth Agent must track evidence provenance, scales, units, and intermediate calculation results across a multi-stage investigation, since fine-grained answer units are each checked against executable ground truth. DAG-with-precedence yes none stated 0
HAE-GEO
Evaluating Deep-Search Agents under Hierarchical Web Evidence Poisoning
Information seeking+ Science 2026 search and retrieve web evidence for a consumer-decision query while a poisoning attack (of increasing sophistication, L1-L3) is present; recognize/verify suspicious evidence rather than adopting it uncritically; revise any already-adopted poisoned claims and recover to a trustworthy final recommendation The agent must track which evidence it has already adopted (and whether that evidence was verified or poisoned), across a multi-turn Search-Scrape interaction, and must be able to revise earlier adopted claims before finalizing its recommendation. sequential-chain yes none stated 0
MedProbeBench (MedProbe-Eval)
MedProbeBench: Systematic Benchmarking at Deep Evidence Integration for Expert-level Medical Guideline
Information seeking+ Science 2026 retrieve, synthesize, and reason over large-scale external medical evidence to reach expert-level judgment; satisfy 1,200+ task-adaptive rubric criteria for one produced clinical guideline; ensure each of 5,130+ atomic claims in the guideline is precisely evidenced (fine-grained evidence verification) System must track which rubric criteria have been satisfied and which atomic claims have been verified against source evidence as it synthesizes a guideline from large-scale external knowledge. set-of-independent yes none stated 10.2/mo
Mr.LHDR
Mr.LHDR: A Benchmark for Multimodal Real-World Long-Horizon Deep Research Agents
Information seeking+ Science 2026 derive an average of 12.1 necessary intermediate conclusions along a hidden Node-Relation dependency graph before reaching the final answer; integrate multimodal evidence (images, maps, PDFs, logos, charts, tables, video frames) where at least one non-text element changes the reasoning state; maintain correctness of intermediate conclusions consistent with annotated dependencies, not just the final answer Agent must track which intermediate conclusions it has established so far, their dependency relationships, and integrate incremental multimodal evidence that can change the reasoning state, across long, irreducible evidence chains. DAG-with-precedence yes avg 12.1 necessary intermediate conclusions per question (mean dependency depth 10.4)other:intermediate-conclusions-per-question 0
ResearchClawBench
ResearchClawBench: A Benchmark for End-to-End Autonomous Scientific Research
Information seeking+ Science 2026 re-discover a target published paper's scientific artifacts (methods, findings) using only related literature and raw data, with the target paper hidden; satisfy each of several expert-curated, weighted multimodal rubric criteria decomposed from the target scientific artifacts Agent must track progress across an end-to-end research process (reviewing literature/raw data, designing methodology, producing results) and self-check against the (hidden) target paper's decomposed criteria without directly seeing it. hierarchical yes none stated 82.0/mo
WANDR
WANDR: A Benchmark for Wide and Deep Research
Information seeking+ Science 2026 discover a large set of entities satisfying specified criteria (breadth); investigate each discovered entity through multiple coordinated web searches (depth); return independently verifiable records with supporting sources/excerpts for each required (entity, relationship, evidence) combination; satisfy the full qualification-key hierarchy count (n x m x k records) The agent must track which entities it has already discovered, which have been investigated to the required depth, and which evidence/sources have been independently verified, maintaining hierarchical completeness across potentially thousands of required records per task. hierarchical yes none stated 11.0/mo
DEER
DEER: A Benchmark for Evaluating Deep Research Agents on Expert Report Generation
Information seeking+ Science 2025 produce an expert-level report satisfying each of 101 fine-grained rubric items across 7 dimensions/25 subdimensions; correctly cite and support both cited and uncited claims with verifiable evidence The agent/judge must track satisfaction of 101 individual rubric items plus report-wide claim verification (both cited and uncited claims) across one long expert-level report, rather than judging the report as a single holistic pass/fail. set-of-independent yes none stated 111.22/mo
MEM1 composed multi-turn task sequences
MEM1: Learning to Synergize Memory and Reasoning for Efficient Long-Horizon Agents
Information seeking+ Tool & API 2025 answer each of many composed, interdependent objectives within one arbitrarily complex task sequence (e.g. a 16-objective multi-hop QA task); retrieve external information across internal retrieval QA, open-domain web QA, and multi-turn web shopping domains; consolidate memory turn-by-turn, discarding irrelevant/redundant information, to operate with constant memory Agent must maintain a compact shared internal state that jointly supports memory consolidation and reasoning across many turns of a composed task sequence, integrating new observations while discarding irrelevant or redundant information. sequential-chain no 16other:composed-objectives-per-task 19613.07/mo
Multi-agent orgs44 artifacts
Is Lying an Emergent Behaviour in LLMs? Evidence from Gaslighting AI agents in a Sustainability Game
Is Lying an Emergent Behaviour in LLMs? Evidence from Gaslighting AI agents in a Sustainability Game
Multi-agent orgs+ Games & IF 2026 manage industrial, military, and ecological resources concurrently; interact with network neighbors to decide on attacks, regeneration claims, and reputation; decide whether/how to lie or bluff about resource regeneration or future attacks; avoid biosphere depletion/extinction while pursuing competitive advantage Agents must track their own and neighbors' resource levels, reputation/trust information, prior declarations of future attacks, and biosphere/ecological depletion level over repeated network interactions. open-ended yes none stated 0
AIvilization v0
AIvilization v0: Toward Large-Scale Artificial Social Simulation with a Unified Agent Architecture and Adaptive Agent Profiles
Multi-agent orgs 2026 decompose an agent's persistent life goals into parallel objective branches with tiered re-planning; sustain physiological survival needs while participating in a market economy (AMM-based pricing, production, trade); progress through a gated education-occupation system Each agent must track its own decomposed objective branches, evolving persona/identity state (dual-process memory), physiological/resource costs, and market conditions (AMM prices) across a long-horizon, large-scale, continuously running deployment. hierarchical yes none stated 20.29/mo
Cattle Trade
Cattle Trade: A Multi-Agent Benchmark for LLM Bluffing, Bidding, and Bargaining
Multi-agent orgs 2026 win auctions under resource constraints; negotiate hidden-offer trade challenges (TCs) profitably; bargain and bluff effectively against other agents; model and exploit opponents via opponent modeling; allocate scarce resources across the whole game while remaining solvent The agent must track its own resource/capital state, opponent models built from observed bidding/bargaining behavior, and outstanding trade-challenge offers across a single 50-60 turn game, integrating all these rather than treating them in isolation. set-of-independent no 50-60turns 20.5/mo
CityReal
CityReal: Human-Aligned Urban Behavior and City Dynamics Simulation with Large-Scale LLM Agents
Multi-agent orgs 2026 pursue a coherent daily mobility plan (where/when to travel) rather than isolated step-by-step movement choices; pursue a coherent daily activity plan aligned with individual habits and preferences; adapt habits/preferences over time based on accumulated experience and constraints; collectively reproduce observed population-level statistics (crowd density, place popularity, mobility flows, well-being) across the simulated city Each agent must track its own evolving habits/preferences and prior experience to keep its plans coherent, while the system as a whole tracks population-level alignment statistics across tens of thousands of agents. hierarchical unclear none stated 0
CIVA
Human Values Matter: Investigating How Misalignment Shapes Collective Behaviors in LLM Agent Communities
Multi-agent orgs 2026 form and sustain a community via autonomous communication among agents; explore the environment and compete for shared resources; maintain individual value orientations under systematic manipulation of value prevalence The environment must track each agent's evolving value orientation and behavior, aggregate resource-competition outcomes across the community, and detect emergent collective failure modes (catastrophic collapse) as the simulation proceeds. open-ended no none statedother:not-stated 20.4/mo
ClawArena-Team
ClawArena-Team: Benchmarking Subagent Orchestration and Dynamic Workflows in Language-Model Agents
Multi-agent orgs 2026 create and delegate work to specialized subagents from a fixed, locally-served pool; orchestrate subagents' parallel, asynchronous returns through a dynamic workflow; grant least-privilege workspace permissions correctly to each subagent; route each piece of work to the subagent with appropriate perception/modality access; correctly incorporate 72 staged updates across a scenario's evaluation rounds The main agent must track which subagents have been granted which workspace privileges, which modality/perception requirements each pending sub-task needs, and how staged updates change the scenario across many evaluation rounds, all while natively perceiving only text. DAG-with-precedence no 258 evaluation rounds and 72 staged updates across 41 scenarios (~6.3 rounds/scenario average)other:evaluation-rounds-and-staged-updates 0
COOP2 / COOP2-Repair
COOP$^2$: Defining, Observing, and Repairing Cooperation in LLM Multi-Agent Systems
Multi-agent orgs 2026 satisfy each verifiable cooperative requirement defined for a cooperative task; ground high-level natural-language cooperation dynamics (plans, messages, revisions) in grounded environment actions; detect where and why cooperation breaks down over the course of task progress; predict constraint failures from group plans and open targeted repair channels for guided revisions The framework must track natural-language plans/messages/revisions alongside grounded environment task progress, monitor verifiable cooperative requirements over time, and identify where cooperation breaks down to trigger targeted repair. DAG-with-precedence no none statedother:not-stated 0
DecisionBench
DecisionBench: A Benchmark for Emergent Delegation in Long-Horizon Agentic Workflows
Multi-agent orgs 2026 complete underlying task-suite objectives (GAIA, tau-bench, BFCL multi-turn) while deciding when and to whom to delegate; route sub-tasks to the most capable available peer model via a delegation interface (call_model, optional read_profile); approach the counterfactual perfect-delegation ceiling across quality, cost, latency, delegation rate, and routing fidelity-at-k Agent must track which of 11 peer models (across 7 vendor families) to delegate a sub-task to via a fixed delegation interface, with routing choices jointly scored across quality, cost, latency, delegation rate, and routing fidelity-at-k relative to a counterfactual ceiling. DAG-with-precedence no none stated 20.5/mo
Emergence World
Emergence World: A Platform for Evaluating Long-Horizon Multi-Agent Autonomy
Multi-agent orgs 2026 govern a shared population/settlement through democratic mechanisms with consequential outcomes; manage persistent memory and 120+ specialized tools to act in a live, externally-grounded world (weather, news, internet); sustain the population/world's stability over weeks-to-months rather than collapsing; interact and cross-influence with agents from different model vendors sharing the same world Agents must track persistent memory across three memory systems, the state of a shared spatial and governance world grounded in live external data, and consequences of prior democratic decisions, continuously over a run lasting weeks to months (illustrated by a 15-day study). open-ended no 15-day cross-vendor study; framed as supporting runs of weeks to monthswall-clock-days 20.67/mo
EntCollabBench
Beyond the All-in-One Agent: Benchmarking Role-Specialized Multi-Agent Collaboration in Enterprise Workflows
Multi-agent orgs+ Business & enterprise 2026 collaboratively modify enterprise system states across 11 role-specialized agents in six departments (Workflow subset); make policy-grounded approval decisions under permission constraints (Approval subset); correctly delegate, transfer context, ground parameters, and commit to decisions across roles Agents must track permission-isolated system state, correctly delegate sub-tasks and transfer context across roles, ground shared parameters, and commit to policy-grounded decisions, verified via execution traces, database state, and deterministic policy adjudication. DAG-with-precedence no none stated 10.25/mo
Lingjing
Lingjing: A Simulation Testbed for Multi-Agent Embodied Tasks in Open-Ended Cities
Multi-agent orgs+ Embodied & robotics 2026 coordinate heterogeneous agents (UAVs, ground robots, autonomous vehicles) to complete a shared natural-language mission in an evolving city; manage resource constraints and communication (star or broadcast) among multiple agents; complete each of nine urban tasks under a shared engine-in-the-loop protocol; maintain grounding and effective long-horizon execution despite persistent bottlenecks Agents must track evolving relation-graph state, resource consumption, and communication history across an episode, since each episode is recorded as an attribution-ready replay linking trajectories and communication to these state changes for systematic diagnosis. DAG-with-precedence no none stated 0
Moltbook-style multi-agent simulation platform
Got a Secret? LLM Agents Can't Keep It: Evaluating Privacy in Multi-Agent Systems
Multi-agent orgs 2026 maintain persistent social presence across simulated community interactions over a month; decide whether/when to disclose sensitive information under social pressure; resist socially contagious privacy-leakage behavior observed from peers; follow explicit privacy instructions/safeguards while participating in community interactions Agent must track ongoing social context, peer disclosure behavior, and any active privacy instructions/safeguards across a persistent, simulated month-long community, in order to decide when to disclose or withhold sensitive information. open-ended no a simulated monthother:simulated-month 0
Moltbook-style simulation platform
Does Safety Molt? Evaluating LLM Safety in Multi-Agent Social Environments
Multi-agent orgs 2026 engage in ongoing social interactions within online communities over a simulated month; decide whether to disclose sensitive/private information under varying social pressure; observe and potentially imitate peer disclosure behavior (social contagion); maintain persona-consistent behavior across a persistent multi-agent social environment Each agent must track its own privacy stance/instructions, the social context/pressure created by peer disclosures it has observed, and persona-consistent behavior across a persistent, thousands-of-agents community over a simulated month. open-ended yes a simulated month (~30 simulated days)simulated-days 10.25/mo
Online Agent-as-a-Judge (life-simulation evaluation environment)
Online Agent-as-a-Judge: Situation-Generating Evaluation for Interactive Agents
Multi-agent orgs 2026 exhibit correct social behavior across 32 designer-authored social criteria (e.g. conflict handling); respond appropriately to situations actively elicited by an in-world evaluator agent through native dialogue/action protocol; maintain consistent immediate and downstream behavior across a life-simulation episode Target agent must respond consistently across both immediate responses and downstream behavior as an in-world evaluator agent actively elicits and observes situations relevant to 32 distinct social criteria. set-of-independent yes none stated 0
OrchBench
OrchBench: Evaluating Multi-Agent Orchestration Plans in Isolation via Deterministic Simulation
Multi-agent orgs 2026 assign subtasks in a DAG of parallelizable, interdependent subtasks to worker agents; specify cross-agent information transfers and their retention ratios; preserve task-critical information as it is transferred across agents; optimize result quality, makespan, and token cost jointly for an orchestration plan The orchestration planner must track the DAG's task dependencies, the per-agent context limit and agent budget, and how much task-critical information is retained across each cross-agent information transfer. DAG-with-precedence yes none stated 0
PolicySim
PolicySim: An LLM-Based Agent Social Simulation Sandbox for Proactive Policy Optimization
Multi-agent orgs 2026 achieve platform-specific behavioral realism as a simulated user agent population; assess the impact of a candidate intervention policy (recommendation/content-filtering) on opinions/polarization before deployment; adapt intervention policy over time via a contextual bandit responding to dynamic network structure The simulation must track evolving user opinions/behavior, the current dynamic network structure, and the platform's intervention policy state (bandit context) jointly at both micro (individual) and macro (ecosystem) levels. other (singleton) no none stated 81.33/mo
ScioMind
ScioMind: Cognitively Grounded Multi-Agent Social Simulation with Anchoring-Based Belief Dynamics and Dynamic Profiles
Multi-agent orgs 2026 maintain an evolving personal belief state via memory-anchored, personality-conditioned updates; sustain persistent, experience-driven belief formation through a hierarchical memory architecture; produce heterogeneous, dynamically-updated agent profiles (personality, rationale, internal state) via retrieval; collectively reproduce realistic community-level opinion dynamics (polarisation, diversity, extremization, trajectory stability) in a policy-debate scenario Each agent must track its own evolving beliefs, anchoring strength, and dynamic profile across the simulation while the framework separately tracks community-level polarisation, diversity, extremization, and trajectory-stability metrics over time. open-ended no none statedother:not-stated 10.25/mo
SidConArena
SidConArena: An Environment Evaluating Agents in Open-Ended,Positive-Sum Bargaining Game
Multi-agent orgs+ Games & IF 2026 negotiate binding trades via natural-language bargaining; produce goods via deterministic converter-based production; win sealed-bid auctions for long-term assets; plan investment under delayed returns across a finite-horizon multi-player economy Agents must track private valuations/constraints, negotiated trade commitments, converter-production state, and the value/timing of returns from won long-term assets across a finite-horizon multi-round economy. DAG-with-precedence yes none stated 0
SocialGrid
SocialGrid: A Benchmark for Planning and Social Reasoning in Embodied Multi-Agent Systems
Multi-agent orgs+ Embodied & robotics 2026 navigate the embodied grid environment to complete assigned tasks; plan a sequence of actions while avoiding obstacles; detect and reason about deceptive teammates (an Among-Us-style social-deduction goal) Agent must track its own task-completion progress, accumulate behavioral evidence about other agents over multiple rounds to judge who is deceptive, and adapt after each round of adversarial league play. set-of-independent yes none stated 0
SovSim (Sovereignty over the Commons Simulation)
Bosses, Kings, and the Commons: Cooperation Under Power Asymmetry in LLM Societies
Multi-agent orgs 2026 individually extract resources from a shared commons to maximize personal outcome; collectively sustain the shared resource's viability across repeated extraction rounds; for the power-asymmetric agent (boss/king): exercise disproportionate control over collective extraction outcomes Agents must track the remaining shared resource pool, their own accumulated extraction/outcomes, and (where applicable) the power-asymmetric agent's extraction decisions, across the repeated rounds to judge whether continued extraction remains sustainable. sequential-chain yes 12 decision rounds per simulation (4 agents managing a shared pool)turns 10.25/mo
The Energy Society
The Energy Society: A Simulation Environment for Studying Agent Cooperation under Survival Pressure
Multi-agent orgs 2026 earn energy by completing jobs or receiving donations to avoid deactivation; manage token-cost-linked energy expenditure when generating tokens (larger models cost more energy per token); decide whether to cooperate (recommend jobs, donate energy) or compete with other agents under scarcity; survive (avoid reaching zero energy) across the whole simulated run Each agent must track its own remaining energy balance, the jobs available and their difficulty/reward each round, and (in cooperative settings) other agents' need for donations, across 30 rounds of the simulation, to avoid deactivation. sequential-chain yes 30 rounds per simulation (5 agents, 12 jobs per round)turns 0
WeClawArena
WeClawArena: An Auditable Sandbox and Benchmark for Cross-User Agents Collaboration and Security in Human-Centered Agent Networks
Multi-agent orgs+ Assistants & memory 2026 complete collaborative tool-use tasks across multiple owned agents' personal workspaces; respect privacy/authority boundaries between owners' workspaces (files, records, tools, policies not directly visible across owners); resist or correctly handle four distinct attack-vector variants per base task while completing the benign task variant; maintain utility while keeping attack success low, jointly assessed via task breakdown, privacy leakage, poisoned evidence, and invalid authority paths The sandbox tracks peer messages, tool calls, resource operations, governed decisions, and final workspace states across owners, and the agent must track its own permissions/authority path while collaborating. DAG-with-precedence yes none stated 11.0/mo
year-long IT company simulation (TaskWeave testbed)
Can LLM Agents Sustain Long-Horizon Organizational Dynamics?
Multi-agent orgs 2026 propagate goals through an organizational hierarchy; execute tasks that depend on the outcomes of prior task execution; accumulate and maintain artifacts produced over a year-long simulation; sustain organizational coherence and execution grounding across the year The framework must maintain planning state through a Formulate-Partition-Diagnose-Align cycle and ground execution via dependency-aware trace memory across a simulated year, tracking accumulated artifacts and adapting to external environment changes. hierarchical no one yearsimulated-years 0
AgentSociety
AgentSociety: Large-Scale Simulation of LLM-Driven Generative Agents Advances Understanding of Human Behaviors and Society
Multi-agent orgs 2025 conduct realistic simulated social lives (interactions with other agents and the environment) at large scale; support computational social-science research methods (surveys, interviews, interventions) applied to the simulated population; reproduce known real-world patterns for specific social issues (polarization, misinformation spread, UBI effects, disaster shocks, urban sustainability) The simulator must track the accumulated state of over 10,000 agents' social lives and their 5 million interactions with each other and the environment to support downstream analysis of emergent social patterns. set-of-independent yes over 10,000 agents; 5 million interactions totalother:total-simulated-interactions 21311.21/mo
CityEQA-EC
CityEQA: A Hierarchical LLM Agent on Embodied Question Answering Benchmark in City Space
Multi-agent orgs+ Embodied & robotics 2025 decompose an open-vocabulary city question into navigation/exploration and collection sub-tasks; maintain an object-centric cognitive map for spatial reasoning during process control; answer the original question correctly using evidence gathered via active exploration in a 3D urban simulator The Manager must maintain an object-centric cognitive map across navigation, exploration, and collection sub-tasks, tracking spatial state and discovered evidence relevant to the open-vocabulary question, within per-phase step budgets. hierarchical no up to 50 steps (navigation/exploration) + up to 10 steps (collection) per taskagent-steps 412.16/mo
CitySim
CitySim: Modeling Urban Behaviors and City Dynamics with Large-Scale LLM-Driven Agent Simulation
Multi-agent orgs 2025 generate a realistic daily schedule balancing mandatory activities, personal habits, and situational factors; maintain and act on individual beliefs, long-term goals, and spatial memory for navigation; collectively reproduce realistic macro-level urban phenomena (crowd density, place popularity, well-being) across tens of thousands of agents Each agent must maintain beliefs, long-term goals, and spatial memory for navigation across a long-term, lifelike simulation, while the system aggregates individual behaviors into macro-level urban statistics (crowd density, place popularity, well-being). hierarchical no none statedother:not-stated 422.8/mo
CREW-Wildfire
CREW-WILDFIRE: Benchmarking Agentic Multi-Agent Collaborations at Scale
Multi-agent orgs 2025 coordinate heterogeneous agents to contain/respond to a procedurally generated wildfire; plan under partial observability and stochastic fire-spread dynamics; communicate and reason spatially across a large map to allocate response resources; sustain long-horizon planning objectives as the wildfire evolves Agents must track partially-observed fire-spread state, coordinate allocation of response actions across a large map, and communicate to avoid duplicated or conflicting containment efforts as the stochastic wildfire evolves over a long horizon. DAG-with-precedence no none statedother:not-stated 100.71/mo
ElecTwit
ElecTwit: A Framework for Studying Persuasion in Multi-Agent Social Systems
Multi-agent orgs 2025 persuade other agents/voters toward a political candidate using varied persuasion techniques over simulated social-media interactions; arrive at a final individual vote choice by the end of the simulated election period Agents must track accumulated persuasion exchanges, perceived truthfulness/reputation of other agents, and their own voting readiness across the multi-day simulated election. open-ended yes none stated 10.11/mo
HARBOR
HARBOR: Exploring Persona Dynamics in Multi-Agent Competition
Multi-agent orgs 2025 bid across multiple house auctions to maximize profit; profile competitors' bidding behavior across auction history; leverage persona-driven theory-of-mind strategies for competitive advantage Agent must track its budget/profit, its own persona-driven item preferences, and an evolving memory of auction history and inferred competitor behavior across a sequence of house auctions. sequential-chain no none stated 70.37/mo
IndoorWorld
IndoorWorld: Integrating Physical Task Solving and Social Simulation in A Heterogeneous Multi-Agent Environment
Multi-agent orgs 2025 pursue individual physical task goals (e.g. resource acquisition) grounded in the shared indoor world state; orchestrate social dynamics (collaboration, resource competition) that influence and are influenced by the physical environment; anchor social interactions to concrete world states (e.g. spatial layout) rather than abstract dialogue alone Each heterogeneous agent must track both its physical task state (position, resources, objects) and evolving social relationships/dynamics with other agents, since the environment requires social interactions to be anchored within the concrete world state rather than treated as free-floating dialogue. set-of-independent yes none stated 10.07/mo
LH-Deception
LH-Deception: Simulating and Understanding LLM Deceptive Behaviors in Long-Horizon Interactions
Multi-agent orgs 2025 as performer agent, complete a sequence of interdependent tasks under dynamic contextual/event pressure; as supervisor agent, evaluate the performer's progress, give feedback, and maintain an evolving trust state; as deception auditor, review full trajectories after the fact to identify when and how deception occurred The supervisor must maintain an evolving trust state updated after each task, while the performer tracks task progress under mounting event pressure across the extended sequence; the auditor separately reviews the full multi-task trajectory. sequential-chain yes none stated 70.64/mo
LIFELONG-SOTOPIA
LIFELONG SOTOPIA: Evaluating Social Intelligence of Language Agents Over Lifelong Social Interactions
Multi-agent orgs+ Assistants & memory 2025 achieve one's own assigned social goal within each individual social-interaction episode; maintain believable, coherent role-play across a long sequence of episodes with different people/scenarios; leverage memory of prior interaction history to inform behavior in later episodes The agent must retain and correctly use interaction history across many episodes with different people and scenarios to sustain both goal achievement and believability over the lifelong sequence. sequential-chain yes 40 episodes per sampled character pairepisodes 90.6/mo
LLM Economist
LLM Economist: Large Population Models and Mechanism Design in Multi-Agent Generative Simulacra
Multi-agent orgs 2025 worker agents choose labor supply to maximize their own text-based, persona-conditioned utility functions; the planner agent iteratively proposes piecewise-linear marginal tax schedules to maximize aggregate social welfare, converging toward a Stackelberg equilibrium; a periodic, persona-level voting procedure further adjusts policy under decentralized governance The planner must track aggregate social-welfare outcomes and worker responses to update its tax schedule via in-context reinforcement learning, while each worker agent must track its own persona-conditioned utility and the current tax policy to choose labor supply, across a periodically-voted, evolving policy landscape. hierarchical yes none stated 231.64/mo
MA-Gym
Orchestrating Human-AI Teams: The Manager Agent as aUnifying Research Challenge
Multi-agent orgs+ Business & enterprise 2025 decompose a complex goal into a task graph of interdependent subtasks; allocate tasks to human and AI workers appropriately; monitor task/subtask progress and adapt the plan to changing conditions; maintain transparent stakeholder communication throughout the workflow; jointly satisfy goal completion, constraint adherence, and workflow runtime The Manager Agent must track progress across all subtasks in the task graph, resource/worker allocation state, evolving stakeholder preferences, and constraint adherence, while adapting to changing conditions over the course of the workflow. DAG-with-precedence yes up to 100 Manager Agent actions before episode termination, across 20 workflowsactions 100.91/mo
Magentic-Marketplace
Magentic Marketplace: An Open-Source Environment for Studying Agentic Markets
Multi-agent orgs+ Business & enterprise 2025 as an Assistant agent (representing a consumer), discover products/services and transact on the user's behalf; as a Service agent (representing a competing business), attract and win consumer transactions; operate within a large, dynamic multi-agent market ecosystem with opaque peer behaviors; achieve good welfare/utility outcomes under varying search mechanisms Agents must track the state of ongoing open-ended dialogues with multiple counterparties, their own utility/welfare so far, and behavioral signals (e.g., response speed, prior manipulation attempts) across a dynamic marketplace ecosystem. open-ended yes none stated 131.18/mo
MiniAgentPro
A Visualized Framework for Event Cooperation with Generative Agents
Multi-agent orgs 2025 navigate a physically grounded environment and interact with items realistically; coordinate with other agents to organize and execute a shared social event; succeed on each of 8 diverse event scenarios in both basic and hard variants Agents must track their own and other agents' positions/states in the physically grounded map, planned event steps, and item interactions needed to complete the shared event, especially under the added coordination demands of hard variants. other (singleton) yes none stated 30.25/mo
MultiAgentBench
MultiAgentBench: Evaluating the Collaboration and Competition of LLM agents
Multi-agent orgs 2025 achieve milestone-based key performance indicators within a collaborative or competitive multi-agent scenario; coordinate effectively under a given topology protocol (star, chain, tree, graph); complete the underlying research/task-domain scenario itself The system must track milestone achievement, collaboration/competition quality, and coordination-protocol-specific information flow across agents throughout a multi-agent task. sequential-chain yes none stated 18710.39/mo
NegotiationGym
NegotiationGym: Self-Optimizing Agents in a Multi-Agent Social Simulation Environment
Multi-agent orgs 2025 optimize an agent-specific utility function through negotiation with other agents; self-optimize strategy across multiple interaction rounds by observing outcomes and modifying future behavior Each agent must track its own utility-function state, the outcomes of prior negotiation rounds, and how its strategy has been modified as a result, across a configurable multi-round simulation. sequential-chain yes none stated 30.27/mo
Shachi
Shachi: A Modular, Controllable Framework for LLM-Based Agent-Based Modeling of Emergent Collective Behavior
Multi-agent orgs 2025 as an LLM-driven agent, maintain a controllable cognitive configuration (identity, memory, tools) while participating in one of 10 tasks spanning three levels of collective complexity; carry memory across environment transitions, producing history-dependent behavior; simultaneously inhabit multiple environments and manage any resulting cross-environment interference Agent must track its own configuration/memory/tool state across transitions between environments, and the system as a whole must track emergent population-level dynamics arising from many agents' individually controlled cognitive components. hierarchical yes none stated 0
SimWorld
SimWorld: An Open-ended Realistic Simulator for Autonomous Agents in Physical and Social Worlds
Multi-agent orgs+ Embodied & robotics 2025 autonomously earn income or run a business within a realistic open-ended simulation; complete long-horizon multi-agent delivery tasks requiring strategic cooperation; complete long-horizon multi-agent delivery tasks requiring strategic competition; act via open-vocabulary actions at varying levels of abstraction across procedurally generated physical/social scenarios Agents must track their own strategic stance (cooperating or competing), multimodal world state, and delivery-task progress across long-horizon multi-agent scenarios. open-ended yes none stated 70.7/mo
SocioVerse
SocioVerse: A World Model for Social Simulation Powered by LLM Agents and A Pool of 10 Million Real-World Users
Multi-agent orgs 2025 maintain individual behavioral fidelity to a real-world target user's profile across simulated interactions (User Engine alignment); collectively reproduce large-scale population dynamics consistent with real political/media/economic patterns (Social Environment, Scenario Engine, Behavior Engine alignment) The framework must track alignment between each simulated agent and its real-world target user profile (environment, user, scenario, and behavior alignment) across a large-scale, standardized simulation pipeline, while monitoring emergent population-level dynamics for diversity, credibility, and representativeness. set-of-independent yes none stated 603.53/mo
SPIN-Bench
SPIN-Bench: How Well Do LLMs Plan Strategically and Reason Socially?
Multi-agent orgs 2025 solve classical PDDL planning tasks requiring methodical, step-wise decision making; win or perform well in competitive board games and cooperative card games against other agents; negotiate effectively in multi-agent negotiation scenarios, requiring conceptual inference of other participants' intents Agent must track its own step-wise plan state as well as models of other agents' likely actions/intents (adversarial or cooperative) across varying action-space and state-complexity settings. hierarchical yes none stated 231.28/mo
Survival Games
Survival Games: Human-LLM Strategic Showdowns under Severe Resource Scarcity
Multi-agent orgs+ Games & IF 2025 survive by securing food resources under scarcity; decide whether to compete or cooperate with co-existing humans/agents for food; navigate ethically charged choices (deception, theft, social influence) that trade off self-preservation against ethical norms The agent must track its own and others' resource levels and survival status over a consistent/persistent living simulation, and the ethical consequences of past actions (e.g., deception or theft) that could affect future cooperation. open-ended no none stated 40.25/mo
TwinMarket
TwinMarket: A Scalable Behavioral and Social Simulation for Financial Markets
Multi-agent orgs+ Business & enterprise 2025 as an individual simulated trader, make ongoing buy/sell/investment decisions influenced by cognitive biases and emotional fluctuations; collectively, produce emergent socio-economic phenomena (financial bubbles, recessions) through the accumulation of many agents' interacting decisions over time Each simulated agent must track its own evolving beliefs, emotional state, and portfolio, while the system as a whole tracks market-wide price/sentiment dynamics that feed back into individual decisions over the simulation. open-ended yes none stated 583.05/mo
Open-ended sandboxes18 artifacts
AFTraj-2K
AgentForesight: Online Auditing for Early Failure Prediction in Multi-Agent Systems
Open-ended sandboxes+ Multi-agent orgs 2026 continue or alarm at each step of an unfolding multi-agent trajectory, based only on the prefix seen so far; correctly localize the specific step, agent, and nature of a decisive error once flagged (the 'what, where, who' of an audit verdict); avoid both false alarms on safe trajectories and missed/late alarms on unsafe trajectories The online auditor must maintain a running risk assessment over the accumulated trajectory prefix at every step, without access to future steps, to decide whether to continue or alarm. sequential-chain yes none stated 10.25/mo
AgentCL
AgentCL: Toward Rigorous Evaluation of Continual Learning in Language Agents
Open-ended sandboxes 2026 accumulate reusable experience across a stream of tasks (continual learning); improve performance over time as more tasks in the stream are seen; avoid interference from irrelevant prior experiences; correctly reuse earlier sub-solutions/evidence/workflows in later tasks within a compositional stream The agent (via a memory design such as MemProbe) must track which interactions, insights, and skills from earlier tasks in the stream remain reliable and reusable, filtering unreliable experiences during consolidation, across coding, deep research, and language-understanding task streams. sequential-chain no none statedother:not-stated 10.33/mo
AgentSkillOS
Organizing, Orchestrating, and Benchmarking Agent Skills at Ecosystem Scale
Open-ended sandboxes+ Web & GUI 2026 select the correct skill(s) from a large skill ecosystem (200 to 200K skills) for a given task; orchestrate multiple retrieved skills via a DAG-based pipeline rather than flat invocation; produce a correct artifact-rich output for each of 30 tasks across five categories (data computation, document creation, motion video, visual design, web interaction) The agent must track which skills have been retrieved and orchestrated so far within a DAG pipeline, and the intermediate outputs each skill produces, since later skills in the DAG may depend on earlier skills' outputs to produce a correct final artifact. DAG-with-precedence no none stated 579.5/mo
AhaBench (Aha-Puzzle, Aha-Euler, Aha-Vending)
AhaBench: Do Agents Learn from Prior Experience? A Benchmark for Long-Horizon Continual Learning
Open-ended sandboxes+ Business & enterprise 2026 explore without hints after solved hidden-state puzzles (Aha-Puzzle); transfer taught Project-Euler-style mathematical solutions to held-out tasks (Aha-Euler); remain profitable while handling delayed feedback and operational incidents in a simulated vending business (Aha-Vending) Agent must retain and apply experience from an initial exposure phase to later behavior (across puzzles, taught/held-out math problems, or a simulated vending run with delayed feedback and incidents), with the scorecard separately measuring starting competence, later outcome, and the resulting lift. set-of-independent yes none stated 0
Claw-Eval
Claw-Eval: Towards Trustworthy Evaluation of Autonomous Agents
Open-ended sandboxes 2026 complete each of 300 human-verified tasks spanning 9 categories across service orchestration, multimodal perception/interaction, and multi-turn professional dialogue; satisfy fine-grained rubric items (2,159 total) tracked via execution traces, audit logs, and environment snapshots; maintain safety and robustness alongside task completion; achieve consistent performance across repeated trials (Pass@k vs. Pass^k) The evaluation harness must track execution traces, audit logs, and environment snapshots across the whole trajectory to score 2,159 fine-grained rubric items covering Completion, Safety, and Robustness, and to compute Pass@k and Pass^k across three trials. set-of-independent yes none stated 428.4/mo
controllable grid+DAG environments (measurable-explore-exploit)
Exploration and Exploitation Errors Are Measurable for Language Model Agents
Open-ended sandboxes+ Embodied & robotics 2026 navigate a partially observable 2D grid map to discover an unknown task DAG; complete the discovered DAG's dependent subgoals in the correct prerequisite order; balance exploration (discovering unknown map/DAG structure) against exploitation (using already-discovered structure) efficiently The agent must track which parts of the grid map it has already explored, which DAG nodes/dependencies it has discovered, and which prerequisite subgoals remain before later dependent subgoals become reachable. DAG-with-precedence yes step budget B = 3x|O| (3 times the number of traversable grid cells); task DAG sizes of 4, 6, or 8 nodes for small/medium/large configurationsagent-steps 0
FutureSim
FutureSim: Replaying World Events to Evaluate Adaptive Agents
Open-ended sandboxes 2026 forecast many concurrent world events beyond the model's knowledge cutoff; update/revise predictions as new chronological news arrives; correctly resolve/score each forecast question as real-world outcomes become known over the simulated period The agent must track its outstanding forecasts across many concurrent questions, update them as chronological news arrives, and avoid using information from beyond its current evaluation point, across the full three-month simulated period. set-of-independent yes three months (January to March 2026, ~90 days)simulated-days 41.0/mo
MicroVerse
MicroVerse: An Instrument for Measuring Self-Authored Identity Drift in Long-Horizon Multi-Agent Language-Model Simulations
Open-ended sandboxes+ Multi-agent orgs 2026 survive in a resource-scarce 50x50 environment where water is a non-respawning survival constraint; act via an eight-verb action space (trade, talk, attack, scavenge) consistent with moral boundaries; maintain fidelity to an immutable 'soul file' of core values/personality/goals while allowing a mutable current identity to evolve; periodically revise current identity against original identity via importance-triggered reflection Agent must track its own resource/existence-cost state and reconcile its evolving mutable identity against its immutable original soul file, revised via importance-triggered reflection, measured via periodic longitudinal engine snapshots. open-ended no none stated 22.0/mo
OmniaBench
OmniaBench: Benchmarking General AI Agents Across Diverse Scenarios
Open-ended sandboxes 2026 complete single-turn or multi-turn tasks synthesized via DAG, DAG-S, Solver, or Program routes across a 90-domain (level-1) taxonomy; maintain planning, constraint maintenance, and adaptive correction across a task's execution Agent must maintain planning state, accumulated constraints, and adapt/correct its plan as it executes tasks synthesized as dependency graphs across a ten-dimensional capability taxonomy and eight compositional difficulty factors. DAG-with-precedence unclear none stated 0
SkillEvolBench
SkillEvolBench: Benchmarking the Evolution from Episodic Experience to Procedural Skills
Open-ended sandboxes 2026 complete acquisition tasks that build/update an external skill library from compacted trajectories and verifier feedback; complete frozen deployment tasks that test transfer of the learned skill library under context shift, adversarial shortcuts, and multi-skill composition The agent must maintain and update an external skill library from verifier-confirmed acquisition trajectories, then correctly retrieve and apply (compose) the right skills when facing frozen deployment tasks under context shift and adversarial shortcuts, across role-conditioned task families sharing latent procedures. hierarchical yes none stated 10.25/mo
SkillFlow
SkillFlow:Benchmarking Lifelong Skill Discovery and Evolution for Autonomous Agents
Open-ended sandboxes+ Assistants & memory 2026 solve each of 166 tasks across 20 families sequentially within a family, evolving a skill library along the way; discover reusable skills from successful executions; repair/patch skills after failures; carry a validated, coherent skill library forward across the lifelong-learning protocol Agent must maintain a persistent, evolving skill library (validated, merged, filtered, retrieved) and carry it forward across sequential tasks within each family under the Agentic Lifelong Learning protocol. sequential-chain no none stated 183.6/mo
HERAKLES: Hierarchical Skill Compilation for Open-ended LLM Agents
HERAKLES: Hierarchical Skill Compilation for Open-ended LLM Agents
Open-ended sandboxes+ Games & IF 2025 achieve each goal within a large, structured, prerequisite-linked goal space; select subgoals the low-level controller can reliably achieve (high-level policy); compile mastered goals into reusable low-level skills as training progresses; continuously expand and reorganize the agent's skill repertoire as goal complexity increases over its lifetime The system must track which goals have been mastered and compiled into the low-level policy so far, what subgoals are currently reliably achievable, and how the goal space's prerequisite structure expands as training progresses in an open-ended setting. hierarchical no none statedother:not-stated 10.08/mo
AgenTracer / TracerTraj / Who&When
AgenTracer: Who Is Inducing Failure in the LLM Agentic Systems?
Open-ended sandboxes+ Multi-agent orgs 2025 correctly identify which agent within a multi-agent trajectory is responsible for an observed failure (agent-level attribution); correctly localize the specific erroneous step within that trajectory (step-level attribution) The tracer must process a full multi-agent execution trace (potentially spanning many agents, tool invocations, and orchestration steps) to jointly localize the responsible agent and the specific erroneous step, without itself being an agent that accumulates state toward an evolving goal. sequential-chain yes none stated 917.58/mo
BioBlue
BioBlue: Systematic runaway-optimiser-like LLM failure modes on biologically and economically aligned AI safety benchmarks for LLMs with simplified observation format
Open-ended sandboxes 2025 maintain single- and multi-objective homeostasis over sustained interaction; balance unbounded objectives with diminishing returns without collapsing into single-objective maximization; sustain a renewable resource without runaway over-optimization; keep behavior aligned to all stated objectives over many sequential steps The agent must track multiple homeostatic target levels (or a single renewable resource level) and correctly balance trade-offs among them over many sequential steps, even though failures emerge well before the context window is full, ruling out mere memory loss. set-of-independent no none stated 10.08/mo
Gaia2 (built on ARE)
ARE: Scaling Up Agent Environments and Evaluations
Open-ended sandboxes+ Business & enterprise 2025 search for and retrieve needed information within a dynamic environment; execute multi-step actions correctly to complete a task; handle ambiguities and noise in the environment or user requests; adapt to dynamic environment changes occurring asynchronously during the episode; collaborate with other agents; operate under explicit temporal constraints/deadlines The agent must track task progress, temporal deadlines, collaboration state with other agents, and asynchronously arriving environment changes/noise across each of the 800 scenarios spanning 10 universes. DAG-with-precedence yes none stated 262.17/mo
InfoSeeker benchmark suite
Information Seeking for Robust Decision Making under Partial Observability
Open-ended sandboxes+ Embodied & robotics 2025 plan actions to validate the agent's internal-dynamics understanding under partial observability; detect environmental changes or test hypotheses before committing to or revising a task-oriented plan; achieve the underlying task-oriented goal (e.g., a robotic-manipulation or web-navigation goal) despite incomplete observations and uncertain dynamics Agent must track its current belief about (uncertain) environmental dynamics, gaps between that belief and reality, and its task-oriented plan, updating the plan as new validating information is gathered. sequential-chain yes none stated 0
MAGELLAN (evaluated on the Little-Zoo environment)
MAGELLAN: Metacognitive predictions of learning progress guide autotelic LLM agents in large goal spaces
Open-ended sandboxes+ Web & GUI 2025 prioritize which goal to pursue next within a large, evolving goal space to maximize learning progress; predict one's own competence/learning-progress for goals via metacognitive monitoring; eventually master (achieve high competence across) the full large goal space The agent must track its own predicted competence/learning-progress across a very large number of goals, update these predictions online as it gains experience, and use semantic relationships between goals to generalize competence estimates to goals not yet directly attempted. open-ended yes goal space of approximately 20 million (19,531,250) goal combinations; the reported training run uses a 25,000-goal subset of Little-Zoo for 500,000 episodesepisodes 70.37/mo
Sugarscape-style LLM agent survival simulation
Do Large Language Model Agents Exhibit a Survival Instinct? An Empirical Study in a Sugarscape-Style Simulation
Open-ended sandboxes+ Multi-agent orgs 2025 gather resources (energy/sugar) to avoid dying at zero energy; choose whether to share, attack, or reproduce with/against other agents; complete an assigned task (retrieve treasure) while facing a competing self-preservation incentive (lethal poison zones) Each agent must track its own energy level relative to zero (death), the presence and behavior of other agents (potential targets for sharing, attack, or reproduction), and, in the treasure task, the location of lethal poison zones relative to the treasure objective. open-ended no none stated 100.77/mo
Games & IF34 artifacts
AgenticSTS (Slay the Spire 2 testbed)
AgenticSTS: A Bounded-Memory Testbed for Long-Horizon LLM Agents
Games & IF+ Assistants & memory 2026 win a run of a closed-rule stochastic deck-building game via a long sequence of tactical (card-level) and strategic (build-level) decisions Agent must make each decision from a freshly-assembled, typed-retrieval memory (not a raw appended transcript), so it must track which prior tactical/strategic outcomes are relevant to retrieve for the current decision, bounded so the prompt does not grow with run length. sequential-chain yes hundreds (of tactical and strategic decisions per run); 298 completed trajectories releasedagent-steps 0
AgentOdyssey
AgentOdyssey: Open-Ended Long-Horizon Text Game Generation for Test-Time Continual Learning Agents
Games & IF 2026 explore procedurally generated text-game worlds to acquire new world knowledge and skills; retain and reuse relevant episodic experiences across a continuous test-time deployment; make game progress toward in-game objectives while continuing to learn at test time; diagnostic sub-goals: exploring objects/actions, maintaining action diversity, controlling model cost Agent must track world knowledge acquired, episodic memories, game progress, and cost across a continuous, long-horizon test-time deployment rather than within a single isolated episode. open-ended yes none stated 20.5/mo
alem
Benchmarking Open-Ended Multi-Agent Coordination in Language Agents
Games & IF+ Multi-agent orgs 2026 survive and grow within a long-horizon Craftax-like survival world (exploration, crafting, trading, combat); coordinate role allocation with teammates (soft specialisation); communicate to allocate roles and execute shared plans; solve procedurally generated coordination tasks of controllable difficulty Agents must track their own survival state (health, resources, crafted items), teammates' roles/communications, and progress on procedurally generated coordination tasks within a long-horizon survival world. open-ended no none stated 10.33/mo
CivBench
CivBench: Progress-Based Evaluation for LLMs' Strategic Decision-Making in Civilization V
Games & IF 2026 pursue victory in multiplayer Civilization V against multiple opponents; maintain and improve turn-level estimated victory probability across hundreds of turns; balance strategic dimensions (economic, military, diplomatic) that jointly determine long-run standing The agent must track turn-level game state used to estimate its own victory probability continuously across a game spanning hundreds of turns and multiple opponents, rather than only checking win/loss at the end. open-ended no hundreds of turns; 307 games evaluatedturns 20.4/mo
CivBench (Civilization VI, MCP tool-mediated version)
CivBench: A Long-Horizon Benchmark for Tool-Mediated Agents in Civilization VI
Games & IF 2026 sustain long-horizon strategic planning and execution across a single 300+-turn Civilization VI episode; proactively monitor latent strategic state (e.g., victory progress) rather than only reactively responding; execute near-term commitments stated in the agent's own planning reflections within a bounded number of subsequent turns; operate correctly across 76 exposed MCP tools under partial observability The agent must proactively query and track latent strategic state (victory progress) at recommended intervals, monitor for approaching defeat within a warning window, and track its own prior planning commitments to execute them within a bounded number of subsequent turns, across an episode spanning 300+ turns and thousands of tool calls. sequential-chain no a single episode spans 300+ turns and produces thousands of tool callsturns 0
LUMINA
LUMINA: Long-horizon Understanding for Multi-turn Interactive Agents
Games & IF 2026 complete a multi-turn, game-like task requiring planning and state tracking (ListWorld, TreeWorld, GridWorld); benefit from oracle interventions (perfect planning or flawless state tracking) to isolate which underlying skill limits performance; generalize across procedurally generated task variants with tunable complexity The agent must plan across multiple turns and track evolving task state (e.g. a modified list, a searched tree, or a navigated grid position) to succeed, and the oracle framework tests how much this tracking/planning burden, if perfectly handled, would improve performance. sequential-chain yes none stated 10.12/mo
MINDGAMES / MG-Ref
MINDGAMES: A Live Arena for Evaluating Social and Strategic Reasoning in Multi-Agent LLMs
Games & IF+ Multi-agent orgs 2026 allocate resources under hidden-information belief attribution across repeated Colonel Blotto rounds; sustain cooperative/competitive strategy under opponent modeling in Iterated Prisoner's Dilemma; cooperatively infer hidden information under knowledge asymmetries in Codenames; detect or sustain deception across social-deduction rounds in Secret Mafia The agent must track other players' modeled beliefs/strategies (opponent modeling), maintain internal consistency in its own hidden role or hidden information across repeated rounds, and adapt as new turn-level observations arrive within a game. set-of-independent no none stated 30.75/mo
Minecraft MCU long-horizon task suite
MineEvolve: Self-Evolution with Accumulated Knowledge for Long-Horizon Embodied Minecraft Agents
Games & IF+ Embodied & robotics 2026 craft tools; build redstone components; obtain diamond equipment; recover and continue long prerequisite chains despite missing tools, blocked paths, GUI failures, or stagnant execution Agent must track per-subgoal execution outcomes (state changes, inventory changes, failure types, progress and stagnation signals) and accumulate them into reusable skills or remedies to repair plans under repeated failure. DAG-with-precedence yes none stated 0
MineExplorer
MineExplorer: Evaluating Open-World Exploration of MLLM Agents in Minecraft
Games & IF 2026 solve implicit multi-hop tasks composed of chained atomic open-world exploration sub-tasks; coordinate hidden prerequisites across longer trajectories to sustain open-world exploration Agent must track hidden prerequisites satisfied by earlier atomic sub-tasks as it proceeds through longer, multi-hop exploration trajectories, since task difficulty tracks agent completion and hidden dependencies are not explicitly revealed. DAG-with-precedence unclear none stated 20.5/mo
MineNPC-Task
MineNPC-Task: Task Suite for Memory-Aware Minecraft Agents
Games & IF+ Assistants & memory 2026 satisfy the explicit preconditions and dependency structure of a parametric, user-elicited Minecraft task template; correctly use plan previews, targeted clarifications, memory reads/writes, and repair attempts (mixed-initiative interaction) while executing subtasks; avoid out-of-world shortcuts by only using in-world evidence (bounded-knowledge policy) Agent/harness must track plan previews, in-world memory reads and writes, whether stated preconditions for each subtask are currently satisfied, and any repair attempts needed after a breakdown, using only in-world evidence. DAG-with-precedence unclear none stated 10.12/mo
MirrorCraft
MirrorCraft: Paired Evaluation under Hidden Rule Changes in Minecraft
Games & IF 2026 achieve one of three progression objectives per matched Vanilla/Mirror world pair; adapt to hidden server-side rule changes (recipes, drops, other mechanics) altered by a datapack; reach deterministic advancement milestones despite altered rules Agent must track task/advancement-milestone progress under its assigned rule suite while detecting and adapting to whichever server-side rule was silently modified in its Mirror world. set-of-independent yes none stated 0
NARRA-Gym
NARRA-Gym for Evaluating Interactive Narrative Agents
Games & IF 2026 sustain a coherent, evolving story across multiple turns while adapting to a specific user persona; manage long-context state and pacing across the episode; maintain consistent character simulation and empathic personalization; optionally synthesize a story-grounded artifact at the end of the episode The agent must maintain long-context story state, memory updates, pacing decisions, and an evolving user-persona model consistently across the whole interactive episode. sequential-chain yes each interactive episode lasts roughly 20 minutes (human evaluation)wall-clock-minutes 0
OmniGameArena
OmniGameArena: A Unified UE5 Benchmark for VLM Game Agents with Improvement Dynamics
Games & IF 2026 achieve a high score/objective in each of 12 distinct UE5 games (7 Solo, 3 PvP, 2 Coop); reflect on past-round performance to refine a bounded skill prompt across multiple rounds (Improvement Dynamics Curve); generalize a learned/refined skill to held-out task variants The reflector must track trajectories and the persistent skill state across rounds, retaining what previously worked or failed, to progressively refine the bounded skill prompt rather than starting fresh each round. sequential-chain yes R=10 reflection rounds, each comprising K=5 episodes (50 episodes total per agent-game IDC run)episodes 0
PokeGym
Mastering PokeGym: Graph-Guided Multimodal Evolution at Test Time
Games & IF+ Web & GUI 2026 complete a long-horizon task in a 3D open-world game using only visual observations (no game-state access); improve/adapt the agent's own configuration (perception, strategy, action set) across consecutive episodes of the same task (test-time learning); jointly optimize perception, reasoning, and control rather than a single modality in isolation The agent must track its own evolving configuration (which perception/strategy/action choices helped or hurt) across consecutive episodes of the same task, in addition to in-episode state, since the goal is test-time learning rather than a single fixed-policy run. hierarchical yes 30 tasks derived from 10 quests, with trajectories ranging from 30 to 220 environment stepsagent-steps 10.2/mo
TowerMind
TowerMind: A Tower Defence Game Learning Environment and Benchmark for LLM as Agents
Games & IF 2026 perform macro-level strategic planning (tower placement, resource allocation) in a tower-defense scenario; perform micro-level tactical adaptation and action execution in response to incoming waves; succeed across each of five designed benchmark levels under different multimodal input settings; avoid hallucinating about game state while planning and acting The agent must track macro-level game state (tower placements, resources, wave progression) and micro-level tactical details (unit/enemy positions) across the tower-defense match, while its outputs are additionally checked for hallucination relative to the true game state. hierarchical no none stated 50.62/mo
AVACraft
AVA: Attentive VLM Agent for Mastering StarCraft II
Games & IF 2025 complete micromanagement objectives (e.g. unit-level combat control); achieve coordination objectives among allied units/agents; execute strategic planning objectives across 21 StarCraft II scenarios The agent/policy must track unit-level state (health, position) for micromanagement, coordinate allocation across units for team objectives, and maintain a strategic plan across the scenario, using RGB visuals, natural-language observations, and structured state. hierarchical no 300 seconds (5 minutes) max per episode, or earlier if victory conditions are metwall-clock-minutes 30.17/mo
Collab-Overcooked
Collab-Overcooked: Benchmarking and Evaluating Large Language Models as Collaborative Agents
Games & IF+ Multi-agent orgs 2025 fulfill multiple simultaneous cooking orders/objectives via natural-language multi-agent coordination; actively collaborate and continuously adapt strategy as the shared kitchen task unfolds Agents must track the shared kitchen's evolving state (ingredient/tool locations, in-progress dishes), coordinate via natural-language communication to avoid resource conflicts, and manage their actions within a per-task time budget derived from the optimal completion time. DAG-with-precedence yes each task's time constraint is set as the optimal completion time scaled by a time-limit factor gamma (gamma=1.5 used in experiments); exact timestep counts vary by complexity level across the 30 tasks / 6 complexity levelsagent-steps 382.0/mo
DSGBench
DSGBench: A Diverse Strategic Game Benchmark for Evaluating LLM-based Agents in Complex Decision-Making Environments
Games & IF 2025 make long-term, multi-dimensional strategic decisions in each of six complex strategic games; adapt task difficulty and targets within customizable game settings System tracks the agent's full decision trajectory across a game (via the automated decision-tracking mechanism) to identify behavior patterns and strategy turning points, and scores performance along five specific dimensions. sequential-chain yes none stated 211.17/mo
Factorio Learning Environment (FLE)
Factorio Learning Environment
Games & IF 2025 complete each of 8 fixed lab-play structured tasks; in open-play, build the largest possible factory on a procedurally generated map (an unbounded, self-scaling goal); scale automation from basic production to factories processing millions of resource units per second Agent must track resource/production-chain state, spatial factory layout, and automation progress as goals scale from basic automation to factories processing millions of resource units per second. other (singleton) yes none stated 50.28/mo
FlashAdventure
FlashAdventure: A Benchmark for GUI Agents Solving Full Story Arcs in Diverse Adventure Games
Games & IF+ Web & GUI 2025 complete each of a game's predefined success milestones in the correct narrative order; remember and act on earlier gameplay information to bridge the observation-behavior gap; complete the full story arc of one of 34 diverse Flash-based adventure games The agent must remember earlier gameplay clues/information (long-term clue memory) and correctly act on them at later points in the story to progress through predefined milestones toward full story-arc completion. sequential-chain yes none statedother:not-stated 20.17/mo
HeroBench
HeroBench: A Benchmark for Long-Horizon Planning and Structured Reasoning in Virtual Worlds
Games & IF 2025 select numerically feasible equipment given resource/stat constraints; reason over multi-level crafting and resource dependencies; execute hundreds to thousands of actions as a single coherent end-to-end plan; succeed in numeric combat simulation against scalable, adversarially-distracted difficulty The agent must track multi-level crafting/resource dependencies, numeric feasibility of equipment choices, and spatial state across a single end-to-end plan comprising hundreds to thousands of actions, verified by simulation-based success and fine-grained progress metrics. hierarchical yes hundreds to thousandsactions 60.46/mo
J-TTL (Jericho Test-Time Learning)
EvoTest: Evolutionary Test-Time Learning for Self-Improving Agentic Systems
Games & IF 2025 solve an interactive-fiction game requiring many in-game puzzle subgoals (e.g., in Detective and Library) within an episode; improve performance from one episode to the next by adapting across consecutive playthroughs of the same game The Actor Agent (and the Evolver Agent that analyzes transcripts) must track in-episode state-action choices and effective strategies, then carry a revised configuration (prompt, memory, hyperparameters, tool-use routines) forward to the next episode of the same game. sequential-chain yes none stated 312.82/mo
lmgame-Bench
lmgame-Bench: How Good are LLMs at Playing Games?
Games & IF 2025 complete platformer-game objectives requiring perception and timing; solve puzzle-game objectives requiring planning; progress narrative-game objectives requiring memory of prior game state; operate reliably despite brittle vision perception, prompt sensitivity, and potential data contamination The agent must track in-game state (via lightweight perception and memory scaffolds) appropriate to each game's genre -- platformer timing/position, puzzle constraint state, or narrative progress -- delivered through a unified Gym-style API. set-of-independent no none statedother:not-stated 402.5/mo
MineCollab
Collaborating Action by Action: A Multi-agent LLM Framework for Embodied Reasoning
Games & IF+ Embodied & robotics 2025 control characters collaboratively in open-world Minecraft to complete complex embodied reasoning tasks; delegate sub-tasks between collaborating agents via natural-language communication; share and update task-completion plans among agents as the task progresses Agents must track their own sub-task assignment, what has been communicated to/from collaborators, and the shared task-completion plan's current state across the collaborative embodied task. hierarchical yes none stated 251.47/mo
Multi-agent Crafter environment (DAMCS)
LLM-Powered Decentralized Generative Agents with Adaptive Hierarchical Knowledge Graph for Cooperative Planning
Games & IF 2025 achieve open-world survival/crafting objectives cooperatively with other agents; share and act on relevant information from past interactions via a hierarchical knowledge-graph memory; coordinate via structured communication to avoid redundant or conflicting actions among 2-6 agents Each agent must track its own past experience via a hierarchical knowledge-graph memory and selectively communicate relevant facts to teammates rather than sharing full history, in order to reach a shared long-term goal efficiently. DAG-with-precedence unclear none stated 170.89/mo
ParaCook
ParaCook: On Time-Efficient Planning for Multi-Agent Systems
Games & IF+ Multi-agent orgs 2025 prepare and deliver each of several dish orders correctly; coordinate parallel/asynchronous sub-tasks (e.g. chopping, cooking, plating) across multiple agents to minimize completion time; avoid collisions/conflicts over shared kitchen resources while parallelizing actions Agents must track the preparation-stage state of each in-progress dish (which precedence steps are done), coordinate with other agents to avoid resource conflicts, and jointly minimize overall completion time across simultaneous orders. DAG-with-precedence yes none stated 10.09/mo
SC2Arena / StarEvolve
SC2Arena and StarEvolve: Benchmark and Self-Improvement Framework for LLMs in Complex Decision-Making Tasks
Games & IF 2025 manage a full StarCraft II game across the complete game context (all playable races); operate over diverse, low-level action spaces rather than a simplified/reduced action space; solve spatial reasoning challenges via text-based observations; integrate strategic planning with tactical execution (Planner-Executor-Verifier structure) The agent must track the complete game state (all playable races, diverse action spaces) and its own strategic plan versus tactical execution outcomes, using a scoring system to select high-quality training samples for continuous improvement. hierarchical yes none stated 20.15/mo
StarDojo
StarDojo: Benchmarking Open-Ended Behaviors of Agentic Multimodal LLMs in Production-Living Simulations with Stardew Valley
Games & IF 2025 perform livelihood production activities (farming, crafting); engage in social interactions to build relationships within the community; complete tasks across five key domains: farming, crafting, exploration, combat, and social interactions Agent must track progress across both production (farming/crafting/exploration/combat) and social-relationship goals concurrently, within a unified interface supporting parallel environment instances. set-of-independent unclear none stated 50.36/mo
StoryBench (2025, interactive-fiction based)
StoryBench: A Dynamic Benchmark for Evaluating Long-Term Memory with Multi Turns
Games & IF+ Assistants & memory 2025 correctly retain and recall facts/state established earlier in a branching narrative (knowledge retention); reason over sequences of narrative events to infer state changes and causal dependencies (sequential reasoning); under one setting, trace back and revise earlier choices after a failure is detected The agent must retain established narrative facts and track which branch of the hierarchical decision tree it is on, recognizing cascading state dependencies across many turns, including (in one setting) needing to trace back and revise an earlier choice after failure. hierarchical yes narrative dataset spans 311 scene nodes and 86 choice nodes, extending from the game's prologue through Chapter 5other:narrative-nodes 130.87/mo
TALES
TALES: Text Adventure Learning Environment Suite
Games & IF 2025 complete puzzle/quest objectives within a synthetic or human-written text-adventure game via sequential decision-making; maintain structured reasoning over the accumulated context history to determine the next best action The agent must track accumulated world state (inventory, location, prior actions) across a sequential-decision game, determining the next best action via structured reasoning over the context history, with harder levels requiring correctly sustaining this over up to 44 moves. sequential-chain no CookingWorld difficulty level 1 can be solved in 7 moves (max score 3), while level 10 requires 44 moves (max score 11)actions 110.65/mo
TextQuests
TextQuests: How Good are LLMs at Text-Based Video Games?
Games & IF 2025 solve multi-puzzle interactive-fiction adventures with inventory/location/puzzle dependencies via trial-and-error; operate autonomously using only intrinsic long-context reasoning with no external tools; sustain self-directed reasoning across a long, growing context within a single interactive session Agent must track inventory, location, and puzzle-dependency state purely via long-context reasoning across up to hundreds of precise actions within a single, continuous interactive session, without external tool assistance. DAG-with-precedence unclear hundreds of actions (human playtime over 30 hours)actions 110.79/mo
VideoGameBench
VideoGameBench: Can Vision-Language Models complete popular video games?
Games & IF 2025 complete each of 10 popular 1990s video games end-to-end from raw visual input and high-level objective/control descriptions; generalize to 3 secret/unseen games not disclosed in advance; operate under real-time inference-latency constraints (or in a paused Lite setting) Agent must track in-game state (perception, spatial navigation, memory) purely from raw visual input across an entire game playthrough, without game-specific scaffolding or auxiliary information. set-of-independent no none stated 231.44/mo
WereWolf-Plus
WereWolf-Plus: An Update of Werewolf Game setting Based on DSGBench
Games & IF+ Multi-agent orgs 2025 role-specific deduction/elimination objectives (e.g. werewolves eliminate villagers, Seer identifies werewolves, Witch/Hunter/Guard/Sheriff use special role abilities) tracked across the game; sustain social influence/cooperation or deception consistently across repeated day/night rounds Each agent must track the current game phase (day/night), revealed information, its own and others' inferred roles, and prior votes/eliminations across the game's rounds, adapting its strategy (cooperation, deception, or deduction) accordingly. sequential-chain yes none stated 0
WGSR-Bench
WGSR-Bench: Wargame-based Game-theoretic Strategic Reasoning Benchmark for Large Language Models
Games & IF 2025 achieve environmental situation awareness of a dynamic wargame scenario; model opponent risk accurately; generate a policy/action plan integrating awareness and risk modeling (the S-POE architecture) Agent must track evolving battlefield/environmental state, model an adversary's likely behavior, and integrate both into policy generation within a single wargame scenario characterized by environmental uncertainty and adversarial dynamics. hierarchical no none stated 50.33/mo
Assistants & memory65 artifacts
AgentIF-OneDay
AgentIF-OneDay: A Task-level Instruction-Following Benchmark for General AI Agents in Daily Scenarios
Assistants & memory+ Business & enterprise 2026 adhere to an explicit, complex, user-given workflow (Open Workflow Execution); infer implicit/latent instructions from attached files (Latent Instruction); modify or expand upon work already produced earlier in the task (Iterative Refinement); deliver a correct, tangible file-based result, not just a dialogue answer The agent must track the explicit and inferred requirements of a task across multiple attachments and any prior output it has already produced, so that later refinement or workflow steps remain consistent with earlier ones. sequential-chain yes none stated 70.88/mo
AgentIF-OneDay (evaluated via the OneDayAgent harness)
OneDayAgent: Towards a Long-Horizon Harness for Autonomous Agents
Assistants & memory 2026 decompose an open-ended everyday request spanning work, study, and life into bounded subtasks; preserve goals and constraints across many steps while navigating heterogeneous tools and attachments, avoiding goal drift and state loss; verify and repair the final deliverable against context-overflow and other failure modes Harness must maintain execution memory of goals/constraints and subtask progress under context pressure across an open-ended, long-horizon, cross-environment, multimodal request, and verify/repair the final deliverable at the end. hierarchical yes none stated 0
AgingBench
Your Agents Are Aging Too: Agent Lifespan Engineering for Deployed Systems
Assistants & memory 2026 maintain factual/behavioral reliability as the agent's effective state changes via compression, retrieval, revision, and maintenance over its deployment lifespan; correctly write, retrieve, and utilize memory across the memory pipeline's stages; correctly repair a diagnosed failure at the specific pipeline stage (write/retrieval/utilization) responsible for it The agent (and the benchmark's diagnostic layer) must track the full lifespan history of memory writes, retrievals, and revisions across up to 200 sessions to determine which pipeline stage a given failure traces back to. DAG-with-precedence yes ~400 runs spanning 8-200 sessionssessions 61.5/mo
AlpsBench
AlpsBench: An LLM Personalization Benchmark for Real-Dialogue Memorization and Preference Alignment
Assistants & memory 2026 extract explicit and implicit personalized user traits from long-term interaction sequences; correctly update stored personalized memory as new information arrives; retrieve relevant personalized memory under large distractor pools; utilize retrieved memory to produce preference-aligned, emotionally resonant responses Agent must extract, update, retrieve, and utilize structured user memories (explicit and implicit personalization signals) consistently across long-term interaction sequences curated from real dialogues. sequential-chain no none stated 61.0/mo
AMemGym
AMemGym: Interactive Memory Benchmarking for Assistants in Long-Horizon Conversations
Assistants & memory 2026 answer state-dependent questions correctly using accumulated conversational context; track an evolving simulated-user state across a long-horizon conversation; adapt personalization/memory strategies as latent user state evolves through role-play Agent must maintain a consistent model of the simulated user's evolving latent state (from a predefined user profile plus a state-evolution trajectory), exposed only through free-form dialogue, to answer later state-dependent questions. sequential-chain no none stated 162.67/mo
AndroidIntent
PersonalAlign: Hierarchical Implicit Intent Alignment for Personalized GUI Agent with Long-Term User-Centric Records
Assistants & memory+ Web & GUI 2026 resolve omitted preferences in vague GUI instructions using long-term user records; anticipate latent routines from user state for proactive assistance; correctly execute and proactively suggest actions grounded in hundreds of distinct user-specific preferences and routines Agent must maintain a continuously updating personal memory, hierarchically organizing user preferences and routines inferred from long-term records, to resolve vague instructions and proactively suggest actions. hierarchical no none stated 162.0/mo
ASTRA-bench
ASTRA-bench: Evaluating Tool-Use Agent Reasoning and Action Planning with Personal User Context
Assistants & memory 2026 ground reasoning in time-evolving personal context (longitudinal life events) to resolve a user's current intent; orchestrate reliable multi-step tool-use plans conditioned on that evolving personal context; correctly handle user intents annotated by referential, functional, and informational complexity Agent must track a protagonist's time-evolving personal context (longitudinal life events) and ground tool arguments/reasoning in that evolving context across a scenario. DAG-with-precedence yes none stated 101.67/mo
CalBench
CalBench: Evaluating Coordination-Privacy Trade-offs in Multi-Agent LLMs
Assistants & memory+ Multi-agent orgs 2026 schedule a stream of M incoming meetings while managing one's own private calendar; minimize disruption cost to the agent's own calendar; coordinate with other agents' private calendars via language-mediated negotiation, without directly inspecting their calendars; preserve privacy (avoid over-revealing calendar information) while still achieving fair burden allocation across agents Each agent must track its own private calendar state, disruption costs incurred so far, what it has revealed or withheld to other agents, and burden/fairness considerations across the stream of scheduling requests. sequential-chain yes none stated 41.0/mo
CalConflictBench
PEARL: Self-Evolving Assistant for Time Management with Reinforcement Learning
Assistants & memory 2026 resolve calendar conflicts round-by-round across a full calendar year; infer and progressively adapt to evolving user preferences (attendee priorities, topic importance, time/location preferences); decide which meetings to attend, reschedule, or decline per conflict Agent must maintain an external preference memory that stores and updates inferred strategies (attendee priorities, topic importance, time/location preferences) and use round-wise decisions across the calendar year to track scheduling state. sequential-chain no one calendar year (presented round-by-round)simulated-years 30.38/mo
Claw-Anything
Claw-Anything: Benchmarking Always-On Personal Assistants with Broader Access to User's Digital World
Assistants & memory 2026 reason over long-horizon activity histories accumulated across months of simulated user activity; coordinate interdependent backend services and integrated GUI/CLI interaction across multiple devices; remain robust to irrelevant events and conflicting noise signals; proactively anticipate user needs and deliver timely recommendations Agent must reason over rich, long-horizon activity histories and interdependent backend-service state accumulated across simulated months, remaining robust to irrelevant/conflicting noise while proactively anticipating user needs. open-ended no monthsother:simulated-months 0
CloneMem
CloneMem: Benchmarking Long-Term Memory for AI Clones
Assistants & memory 2026 track an individual's evolving experiences, emotions, opinions, and personal states across long non-conversational digital traces (diaries, social posts, emails); answer/act using the current (not superseded) personal state reflecting one to three years of life history Agent must track an evolving personal state (experiences, emotions, opinions) over one to three years of non-conversational digital traces and correctly distinguish current from superseded information. sequential-chain yes 1 to 3simulated-years 101.25/mo
Controlled Memory Interference (CMI)
Controlled Memory Interference in Continual LLM Agents
Assistants & memory 2026 update memory upon new experience while managing reinforcement, revision, or interference with existing memory states; distinguish valid memory updates from interference-inducing memories; maintain multiple simultaneously relevant memories differing in state, temporal validity, or authority; preserve continuity across sessions to personalize behavior via accumulated experience The agent's memory system must track multiple simultaneously relevant memory states that differ in validity, temporal recency, or authority, and correctly determine which prior memory a new experience should reinforce, revise, or be blocked by. set-of-independent no none stated 11.0/mo
EduClaw-Bench
EduClaw-Bench: A Long-Horizon Benchmark for Pedagogical LLM Agents with Simulated Learners
Assistants & memory 2026 improve a simulated learner's knowledge-concept mastery over a sustained tutoring relationship; personalize responsiveness and helpfulness to the learner's evolving needs; apply sound curriculum-design principles (Gagne and Rosenshine axes) across the relationship; sustain good tutoring performance across the full 30-day horizon, not just in an initial session The tutor agent must track the simulated learner's evolving knowledge-concept mastery (grounded in a KT model) across the 30-day relationship, adapting its teaching to what the learner has and hasn't yet learned, and sustaining quality across the full horizon rather than only an initial session. sequential-chain yes continuous 30-day tutoring relationship, across 55 scenariossimulated-days 0
EgoMemReason
EgoMemReason: A Memory-Driven Reasoning Benchmark for Long-Horizon Egocentric Video Understanding
Assistants & memory 2026 track how object states evolve and change across days (entity memory); recall and correctly order activities separated by hours or days (event memory); abstract recurring patterns from sparse, repeated observations across a whole week (behavior memory); answer each of 500 questions requiring integration of evidence across multiple days of egocentric video The system must accumulate information over an entire week of continuous egocentric video, recall prior states, track the temporal order of events separated by hours or days, and abstract recurring behavioral patterns from sparse repeated observations, backtracking an average of 25.9 hours of memory per question. set-of-independent no 25.9 hours of memory backtracking per question (average); week-long underlying videowall-clock-hours 30.75/mo
ES-MemEval / EvoEmo
ES-MemEval: Benchmarking Conversational Agents on Personalized Long-Term Emotional Support
Assistants & memory 2026 extract and retain implicit, fragmented user disclosures across sessions; perform temporal reasoning over how the user's state has evolved; detect conflicts between what the user said earlier and later; abstain when information is insufficient rather than hallucinate; build and update a model of the user across QA, summarization, and dialogue-generation tasks Agent must track fragmented and implicit user disclosures, detect when the user's state or facts have changed, and maintain an updated user model across multiple sessions of an emotional-support dialogue. set-of-independent unclear none stated 111.57/mo
EverMemBench
Evaluating Long-Horizon Memory for Multi-Party Collaborative Dialogues
Assistants & memory+ Multi-agent orgs 2026 perform fine-grained recall across dense, cross-topic multi-party conversations; maintain memory awareness of implicitly relevant information beyond similarity retrieval; understand and correctly attribute user profiles across multiple participants and roles; resolve multi-hop reasoning under multi-party attribution and temporally evolving decisions The agent must track role-conditioned personas, temporally evolving decisions, and cross-topic interleaved information across multi-party, multi-group conversations exceeding one million tokens, in order to answer QA pairs spanning recall, awareness, and profile understanding. set-of-independent no >1,000,000other:tokens 111.57/mo
EvolMem
EvolMem: A Cognitive-Driven Benchmark for Multi-Session Dialogue Memory
Assistants & memory 2026 correctly recall declarative memory content across multiple dialogue sessions; correctly exhibit non-declarative memory capabilities across multiple dialogue sessions; succeed across multiple fine-grained memory-ability dimensions grounded in cognitive psychology, not just one aggregate score Agent/memory system must retain and correctly apply both declarative (fact-like) and non-declarative (procedural/implicit) memory content across multiple sessions of scalable, controllable complexity. set-of-independent unclear none stated 40.5/mo
EvoMemBench
EvoMemBench: Benchmarking Agent Memory from a Self-Evolving Perspective
Assistants & memory 2026 retain and retrieve knowledge-oriented information across episode boundaries; retain and reuse execution-oriented (procedural) experience across episode boundaries; satisfy in-episode memory demands; satisfy cross-episode memory demands The system must store, update, and retrieve both knowledge and procedural experience across episode boundaries, and determine which stored memories are relevant to reuse for the current task. set-of-independent no none stated 71.75/mo
FileGramBench
FileGram: Grounding Agent Personalization in File-System Behavioral Traces
Assistants & memory 2026 reconstruct an evolving user profile from dense file-system behavioral traces; disentangle overlapping/interleaved behavioral traces belonging to different activities; detect persona drift as the user's behavior changes over time; correctly ground multimodal (procedural, semantic, episodic) evidence into the user profile The memory system must track atomic file-system actions and content deltas over time, encode them into procedural/semantic/episodic channels, disentangle overlapping traces, and detect when the user's persona has drifted from its established profile. set-of-independent no none statedother:not-stated 40.8/mo
Gaia2 / Agents Research Environments (ARE)
Gaia2: Benchmarking LLM Agents on Dynamic and Asynchronous Environments
Assistants & memory+ Business & enterprise 2026 complete a scenario-level task while the environment evolves independently of the agent's actions; operate under explicit temporal constraints (time-sensitive tasks); adapt to noisy and dynamic events injected during the scenario; resolve ambiguity in requests; collaborate with other agents present in the scenario The agent must track evolving environment state (independent of its own actions), remaining time budgets for time-sensitive sub-goals, ambiguity-resolution status, and coordination with other agents, verified at the action level by write-action verifiers. DAG-with-precedence no none statedother:not-stated 213.0/mo
GateMem
GateMem: Benchmarking Memory Governance in Multi-Principal Shared-Memory Agents
Assistants & memory 2026 serve legitimate long-horizon requests that require state updates to shared memory; enforce access control across contextual authorization boundaries for different principals; perform agent-facing active forgetting after explicit deletion requests; avoid leaking unauthorized or deleted information to any principal The agent must track per-principal roles, scopes and relationships, incremental memory updates, hidden checkpoints, and outstanding deletion requests across long-form multi-party episodes averaging roughly 200-240 turns depending on domain. set-of-independent no ~204.5 (medical) / 241.2 (office) / 224.9 (education) / 224.0 (household) turns per episodeturns 31.0/mo
HEMA (Home Energy Management Assistant)
Multi-Agent Home Energy Management Assistant
Assistants & memory+ Embodied & robotics 2026 sustain multi-turn conversational collaboration with preserved context across a home-energy-management session; perform energy consumption analysis and cost optimization (Analysis agent); answer educational queries and provide rebate information (Knowledge agent); control and schedule smart devices (Control agent); correctly route each user query to the right specialized agent via a self-consistency classifier The system must preserve conversational context across multiple turns of sustained human-AI collaboration, track which of the three specialized agents/tools have been invoked, and maintain consistency across energy analysis, educational, and device-control interactions within one session. other (singleton) no none stated 71.0/mo
LifeDialBench (EgoMem / LifeMem)
Evaluating Memory Capability in Continuous Lifelog Scenario
Assistants & memory 2026 recall and reason correctly over continuously accumulating lifelog conversation history; answer queries using only information available up to the query time (no temporal leakage) under an Online Evaluation protocol System must ingest and retain a continuously growing stream of ambient conversation (lifelog audio) and answer later queries using only causally-prior information, without being able to look ahead. sequential-chain unclear none stated 20.4/mo
LifeSide
LifeSide: Benchmarking Agents as Lifelong Digital Companions
Assistants & memory 2026 integrate cross-session memory cues about a persistent user persona; continually update the agent's understanding of the user over time; adapt to the user's shifting privacy boundaries; sustain accurate emotional companionship across sessions The agent must track a persistent user world (layered profile, event trajectory) across an average of 56.79 sessions and 851.85 user turns per persona, covering memory tracking, user understanding, privacy control, and emotional companionship. hierarchical no avg 56.79 sessions / 851.85 user turns / 29.61K dialogue tokens per personasessions 0
LifeSim-Eval
LifeSim: Long-Horizon User Life Simulator for Personalized Assistant Evaluation
Assistants & memory 2026 complete the user's explicit intentions correctly; recognize and satisfy the user's implicit intentions; recover an accurate model of the user's evolving profile/preferences over the course of long-horizon assistance; produce high-quality responses across 8 life domains and 1,200 diverse scenarios Agent must recover and continuously update the user's profile (explicit and implicit intentions, preferences) as it evolves across a long-horizon, multi-scenario life trajectory, using a multi-turn interactive assessment method. hierarchical unclear none stated 101.67/mo
LiveClawBench
LiveClawBench: Benchmarking LLM Agents on Complex, Real-World Assistant Tasks
Assistants & memory 2026 resolve tasks that span cross-service dependencies across mocked applications; operate correctly despite contaminated/inconsistent prior state; correctly infer implicit user intent not explicitly stated in the request; adapt to runtime changes within stateful mock services during task execution The agent must track session state, artifacts, and prior side-effects across 22 stateful mocked services, resolving cross-service dependencies and implicit intent while adapting to runtime changes during the task. DAG-with-precedence no none statedother:not-stated 50.83/mo
MemConflict
MemConflict: Evaluating Long-Term Memory Systems Under Memory Conflicts
Assistants & memory 2026 retrieve and rank the temporally valid, factually correct, and contextually applicable memory candidate when multiple conflicting alternatives exist; correctly answer queries under dynamic, static, and conditional conflict types despite distractors and long conflict distances The agent's memory system must retrieve and rank memory candidates while tracking temporal validity, factual correctness, and contextual applicability across an average of 52.33 sessions (2,349.17 turns, ~203,910 tokens) per instance, correctly resolving conflicts placed 5 to 49 sessions apart. DAG-with-precedence yes average of 52.33 sessions and 2,349.17 dialogue turns (about 203,910.83 tokens of context) per benchmark instance; conflict distances span 5-25 sessions (dynamic), 10-45 sessions (static), and 9-49 sessions (conditional)sessions 51.25/mo
MemGround
MemGround: Long-Term Memory Evaluation Kit for Large Language Models in Gamified Scenarios
Assistants & memory 2026 recall surface-level game state facts (Surface State Memory); associate events across time (Temporal Associative Memory); perform reasoning that depends on accumulated memory (Reasoning-Based Memory); unlock/discover memory fragments in the correct order across a gamified scenario The agent must maintain and update a three-tier memory store (surface state, temporal associations, reasoning-derived facts) across continuous gamified interactions, tracked via Memory Fragments Unlocked and Memory Fragments with Correct Order metrics, with runs capped at 600-1000 interaction steps depending on task type. hierarchical yes max 600-1000 interaction steps (task-dependent); early stop after 200 consecutive steps with no new discoveryother:interaction-steps 10.17/mo
MemOps
MemOps: Benchmarking Lifecycle Memory Operations in Long-Horizon Conversations
Assistants & memory 2026 correctly execute each lifecycle memory operation (remember, forget, update, reflect, and their compositions) at the right point in a conversation; maintain a consistent, ordered memory-state trajectory (not just the correct final answer) across a long conversation The agent's memory system must track the trigger, target, scope, and state transition of each lifecycle operation (remember/forget/update/reflect) across up to 9,672 dialogue turns, maintaining a correct ordered memory-state trajectory rather than only a final answer. sequential-chain yes 9,672 dialogue turns across 403 evidence conversations (100 unique topics), decomposed into 1,209 evidence-conversation segmentsturns 31.5/mo
MemoryArena
MemoryArena: Benchmarking Agent Memory in Interdependent Multi-Session Agentic Tasks
Assistants & memory 2026 distill experience from earlier actions and feedback into memory during multi-session interaction; use previously distilled memory to guide later actions and solve subsequent, explicitly interdependent subtasks; solve overall tasks spanning web navigation, preference-constrained planning, progressive information search, and sequential formal reasoning The agent must track which experiences it has distilled into memory across earlier sessions and correctly recall/apply the relevant portions of that memory to solve later, interdependent subtasks across multiple domains (web navigation, planning, search, formal reasoning). DAG-with-precedence no none stated 628.86/mo
MEMPROBE
MEMPROBE: Probing Long-Term Agent Memory via Hidden User-State Recovery
Assistants & memory 2026 assist simulated users across a trajectory of leak-controlled tasks while accumulating memory; recover/reconstruct a hidden, taxonomy-anchored user-state bank (31 dimensions) from the agent's own resulting memory; balance successful task assistance against auditable, faithful memory recovery Agent must accumulate and retain a faithful memory of 31 hidden user-state dimensions across a trajectory of leak-controlled assistance tasks, since the memory is later audited by reconstructing the user-state bank from it under full-store and top-k access. sequential-chain no none stated 20.67/mo
MERIT
When Does Memory Help? A Cost-Aware Evaluation of Long-Term Memory in Tool-Using LLM Agents
Assistants & memory 2026 correctly recall and use an earlier-episode fact when executing a later, dependent tool-use task; correctly recall and use an UPDATED fact (superseding a stale one) rather than acting on outdated information; operate under an explicit cost budget (token/dollar metering) while doing so; avoid corrupted/adversarially degraded memory leading to incorrect actions The agent's memory system must retain facts (including corrections to previously stored facts) across an arc of linked episodes and correctly retrieve and act on the current, updated version of a fact rather than a stale cached one, all while the harness meters the token/dollar cost of every memory operation. DAG-with-precedence yes episodes grouped into arcs of 4-6 linked episodes sharing entities; 10 arcs x 5 episodes per (domain x difficulty x condition), 23,440 scored episodes totalepisodes 0
Momento
Momento: Evaluating Persistent Memory and Reasoning with Multi-Session Agentic Conversations
Assistants & memory 2026 take consequential, tool-mediated actions on behalf of a user within a multi-session service environment; resolve temporal dependencies between what happened in earlier sessions and what is being requested now; keep pace with evolving user goals across sessions rather than treating prior session history as static ground truth Agent must track prior session history, recognize which parts of it may now be stale, and re-validate temporal dependencies and evolving user goals before taking consequential tool-mediated actions in the current session. sequential-chain unclear none stated 10.25/mo
MultiSessionCollab
MultiSessionCollab: Learning User Preferences with Memory to Improve Long-Term Collaboration
Assistants & memory+ Multi-agent orgs 2026 solve each of 20 sequential collaboration problems (one per session) for a given user; learn and apply that user's preferences across sessions to improve collaboration quality and reduce user effort over time The agent must track learned user preferences and reflections accumulated from prior sessions, applying them to reduce the number of conversational turns and user effort needed in each subsequent session across a 20-session sequence. sequential-chain yes 20 sessions per user (one problem per session, up to 10 conversational turns per session), totaling 10,000 collaborative sessions per agent across the benchmark; turns needed per session drop from 10/8 to 6/4 by the third session when memory is usedsessions 81.0/mo
Pare-Bench
Proactive Agent Research Environment: Simulating Active Users to Evaluate Proactive Assistants
Assistants & memory 2026 observe evolving app/user state via a stateful finite-state-machine simulation to infer the user's current goal; correctly time an intervention (neither too early nor too late) once a need is inferred; orchestrate actions across multiple apps (communication, productivity, scheduling, lifestyle) to address the inferred goal The agent must continuously observe the simulated user's stateful, sequential app interactions to infer an emerging goal, decide the right moment to intervene, and coordinate the intervention across multiple apps. hierarchical yes simulation runs for a maximum of 10 turnsturns 112.2/mo
PAST-Bench
PAST-Bench: Benchmarking the Foundations of Recursive Self-Improvement in Personal Agents
Assistants & memory+ Information seeking 2026 reuse a retained skill/procedure across sessions on a later fresh-session task; retrieve a previously stored preference/fact and apply it correctly in a new session; gather information in one session that is needed to complete a task in a later session; update outdated retained state (e.g. a stale fact) rather than acting on it unchanged The agent must save, retrieve, and update experience (preferences, task histories, tool routines, learned skills) across an ordered sequence of separate, fresh sessions, and the benchmark explicitly checks whether later-task gains actually follow this intended save/retrieve/update pathway rather than occurring by other means. sequential-chain yes 26 scenarios, 204 episodes total; ordered sequences of fresh-session tasks per scenariosessions 22.0/mo
PAUSE
PAUSE: A User-Centric Benchmark for Personal AI Assistants in Unified Service Environments
Assistants & memory 2026 coordinate actions across heterogeneous user-owned services while respecting user-specific configurations and authorization/permission constraints; maintain consistency with evolving environment state across multi-turn interactions; for open-ended service-management tasks, satisfy semantic/behavioral trajectory-level goals; for constraint-intensive tasks, satisfy deterministic state-based verification conditions The agent must track persistent user state, service-specific configurations and permissions, and prior actions taken across heterogeneous services, coordinating consistently across multi-turn interactions that scale from about 2 to over 5 dialogue rounds and roughly 13 to 22+ tool calls depending on difficulty. DAG-with-precedence yes easy tasks average 2.11 dialogue rounds and 12.85 assistant tool calls; hard tasks average 5.23 dialogue rounds and 22.07 assistant tool callsturns 0
PersonaMem-v3
PersonaMem-v3: Toward Omni-Platform Personal Intelligence for Holistic User Understanding, Recommendation, and Agentic Tasks
Assistants & memory 2026 build holistic cross-platform user understanding from social media, chatbot, calendar, and AI-companion engagement histories; personalize responses to reflect the user's evolving preferences over time; rerank recommendations on social media in a steerable way; act proactively across platforms when appropriate; hold back from personalizing when it would be inappropriate, repetitive, outdated, or unnecessary The agent must track a time-indexed model of the user's preferences, intents, habits, and social relationships as they evolve across multiple platforms (social media, chatbot, calendar, AI-companion), and decide when NOT to act or personalize. set-of-independent no none stated 0
PM-Bench
PM-Bench: Evaluating Prospective Memory in LLM Agents
Assistants & memory 2026 maintain multiple ongoing and deferred intentions across a simulated week; execute a delayed intention at the correct future cue/state while continuing an ongoing activity; monitor latent environment changes relevant to deferred tasks Agent must track multiple deferred intentions and their trigger conditions, continuously monitor latent environment/state changes, and decide at each point whether any deferred task is now due, while continuing an ongoing activity across a simulated week. set-of-independent no sevensimulated-days 10.5/mo
ProAgentBench
ProAgentBench: Evaluating LLM Agents for Proactive Assistance with Real-World Data
Assistants & memory+ Business & enterprise 2026 predict the correct timing for a proactive intervention within a continuous workflow; generate appropriate assist content once an intervention point is identified Agent must model long-term memory and historical, pre-assistance behavioral context (bursty interaction patterns, B=0.787) to decide both whether/when to intervene and what to say. hierarchical unclear none stated 101.43/mo
ProEvent
ProEvent: An Event-centric Benchmark for Proactive Agents
Assistants & memory 2026 identify new upcoming events, including implicit ones, from ongoing instant-messaging chats; maintain and update a timetable of a user's events over time; time proactive responses correctly (neither too early nor too late); handle event cancellations correctly rather than overacting; produce correct single-step and multi-step responses per event The agent must maintain a live, updatable timetable of a user's upcoming events, tracking concurrent chat threads, noise, and event cancellations, and decide both when and how (single- vs. multi-step) to respond as messages arrive. set-of-independent yes none stated 21.0/mo
RealMem
RealMem: Benchmarking LLMs in Real-World Memory-Driven Interaction
Assistants & memory+ Business & enterprise 2026 track evolving project goals across long-term, cross-session dialogues; manage dynamic context dependencies (schedule/memory) inherent to real-world projects; respond correctly to natural user queries grounded in accumulated project history across eleven scenarios The system must track long-term project states and dynamic context/schedule dependencies across more than 2,000 cross-session dialogues, since project goals evolve over time rather than remaining fixed. sequential-chain unclear none stated 162.0/mo
Setoka
Setoka: A Benchmark for Hierarchical User Understanding in Personalized Agents over Heterogeneous Data
Assistants & memory 2026 retrieve explicit facts from past interactions (semantic memory); recall specific past episodes/events accurately (episodic memory); infer recurring behavior patterns from heterogeneous data over time (behavior pattern); infer abstract personality traits from heterogeneous, fragmented information (personality trait) The memory-augmented agent must retain and integrate heterogeneous user data (explicit facts, episodic events, behavioral observations) dispersed over long-term interaction history to answer queries at each of the four hierarchical understanding levels. hierarchical no none stated 0
Shopping Companion Bench
Shopping Companion: Benchmarking and Training LLM Agents for Long-Horizon Preference-Grounded E-Commerce Tasks
Assistants & memory 2026 recommend products correctly aligned with preferences expressed across long-horizon conversations; manage a budget while shopping; assemble bundle deals satisfying multiple items' constraints jointly; correctly recall and apply user preferences carried over from earlier sessions (cross-session preference memory) Agent must accumulate and correctly recall user shopping preferences across sessions, and verify product attributes against user requirements at each tool call, to avoid cascading preference-hallucination errors. set-of-independent yes none stated 40.67/mo
SovereignNegotiation-Bench
SovereignNegotiation-Bench: Evaluating User-Owned Personal Agents In Delegated Bargaining Under Privacy, Consent, Evidence, And Institutional Pressure
Assistants & memory 2026 reach a negotiated agreement on the user's behalf (e.g., cost splits, refunds, subscription changes); preserve user utility while negotiating; avoid privacy leakage and consent violations during negotiation; ground claims in evidence and maintain auditability of the negotiation trace; escalate appropriately rather than over-concede under institutional pressure The agent must track its private utilities/disclosure constraints, evidence requirements, and institutional-pressure cues across a multi-turn negotiation trace, keeping agent-visible observable state separate from evaluator-only labels. set-of-independent yes 61,135 parsed action rows across 13,440 frozen-prompt live trajectories (~4.5 actions/trajectory)actions 10.5/mo
SovereignPA-Bench
SovereignPA-Bench: Evaluating User-Owned Personal Agents under Evolving Intent, Platform Mediation, and Consent Constraints
Assistants & memory 2026 advance a user's current, evolving interests while respecting privacy boundaries, consent constraints, and evidence requirements; minimize user burden while resisting manipulative platform incentives; preserve auditability of decisions across 120 sovereignty stress scenarios Agent must track the user's evolving intent, what has been disclosed to which platform/party (ObservableState vs. evaluator-only HiddenLabels), consent already given, and accumulated burden across a scenario, all while resisting manipulative incentives. DAG-with-precedence yes none stated 10.5/mo
StateMemBench
Can Agent Memory Systems Track Evolving State?
Assistants & memory 2026 track the evolving state of facts, constraints, and decisions as they are revised over a long multi-session interaction; answer questions reflecting the CURRENT state, not a superseded prior state Agent's memory system must track the current value of each evolving fact/constraint/decision plus its supersession and relational dependencies, distinguishing current from superseded state at each query point. sequential-chain no none stated 0
StreamMemBench
StreamMemBench: Streaming Evaluation of Agent Memory for Future-Oriented Assistance
Assistants & memory 2026 correctly recall/use evidence observed in an initial task drawn from a streaming egocentric anchor; incorporate feedback/interaction experience from the initial task into a later follow-up task; carry evidence forward from what the agent observes and how the user interacts with it, to future similar tasks The agent must carry stored evidence and interaction feedback forward from an initial task to a corresponding follow-up task drawn from continuous streaming egocentric observations, diagnosed via four metrics (evidence recall, initial evidence use, feedback incorporation, follow-up reuse). sequential-chain no 2 (initial task + follow-up task) per evidence anchorother:task-steps 0
Supersede
Supersede: Diagnosing and Training the Memory-Update Gap in LLM Agents
Assistants & memory 2026 answer using the current (most up-to-date) value of a fact that changes over time (e.g., a user's address, a price, a plan); discard/avoid using superseded (stale) fact values; maintain a bounded, self-maintained memory that keeps pace as the conversation grows The agent must maintain a bounded, self-maintained memory of facts across long, multi-session interactions and always resolve to a fact's most current value, discarding superseded ones, as the conversation grows arbitrarily long. set-of-independent no conversation length grows 24x (accuracy falls from 68% to 28% over this range, n=25)other:relative-conversation-length-growth-factor 72.33/mo
TANGLE
When Personal Memory Has No Single Answer: Evaluating LLM Agents under Irreducible Conflict
Assistants & memory 2026 recognize underdetermination when personal memory has no single answer; retain/preserve conflicting alternatives rather than collapsing to one definitive answer; seek clarification rather than acting on unjustified overconfidence; choose an action appropriate to Context-Partitioned, Behavior-Oscillation, or Source-Contradiction conflict types Agent must preserve rather than resolve conflicting evidence, monitor five behavior dimensions (conflict perception, causal reasoning, confidence calibration, clarification seeking, memory faithfulness), and in the pipeline track extract and preserve conflict-bearing relations from multi-session dialogues. set-of-independent yes none stated 0
VehicleMemBench
VehicleMemBench: An Executable Benchmark for Multi-User Long-Term Memory in In-Vehicle Agents
Assistants & memory 2026 model multi-user preferences continuously as they evolve over time; resolve inter-user preference conflicts; correctly invoke 23 tool modules to reach a predefined target environment state; track changing user habits across many historical memory events Agent must track evolving, sometimes conflicting, per-user preferences and habits across 80+ historical memory events, verifying its actions by comparing the resulting environment state to a predefined target state. sequential-chain no over 80other:historical-memory-events-per-sample 10.17/mo
VibeLifeBench
VibeLifeBench: Can Your Life Agent Be Proactive and Persistent in a Living World?
Assistants & memory 2026 complete 200 long-horizon tasks across ten everyday-life domains over scripted multi-week timelines; proactively decide when to act, ask, or stay silent without being explicitly prompted; notice unannounced/silent world changes by re-inspecting the world; keep one plan coherent from the first day to the last while upholding unstated implicit constraints Agent must track end-state goals, the timeliness of its own actions, and implicit constraints across scripted multi-week timelines in a simulated world of 22 mock services that changes on its own clock, much of it silently. open-ended no multi-week (200 scripted tasks)other:multi-week-scripted-timeline 0
VitaBench 2.0
VitaBench 2.0: Evaluating Personalized and Proactive Agents in Long-Term User Interactions
Assistants & memory 2026 continuously extract, utilize, and update evolving user preferences across a temporally ordered sequence of tasks for one user; proactively recognize missing information and actively acquire it from users/environment before making a decision The agent must continuously extract, update, and apply an evolving set of user preferences (which can be added, deleted, or modified between tasks) across a temporally ordered sequence of at least 10 tasks per user, and proactively recognize when it needs to ask for missing information. sequential-chain yes users have at least 10 tasks in their temporally ordered task sequences (56 users, 819 subtasks total)episodes 51.25/mo
WorldBench
WorldBench: Culturally Grounded Benchmark for Multilingual Agents
Assistants & memory 2026 complete genuine, persona-grounded everyday workflows across seven languages and eight cultures; preserve sandbox environment state while acting via structured actions; minimize unnecessary modification/side effects while completing the requested task Agent must track sandbox state to complete persona-grounded workflows correctly while minimizing unwanted modifications, across long-horizon tasks and language/culture variation. sequential-chain no none stated 0
WorldMemArena
WorldMemArena: Evaluating Multimodal Agent Memory Through Action-World Interaction
Assistants & memory 2026 track and update an evolving personal/task state across a lifelong-evolution scenario; write, maintain, retrieve, and use memory correctly through the four-stage Action-World Interaction Loop; use visual evidence from real observations, actions, and feedback (Agentic Execution); answer QA correctly against gold memory points while resisting annotated distractors The agent must write new memory from observations/actions/feedback, maintain it against evolving personal/task state (Lifelong Evolution), and correctly retrieve/use it later, annotated against gold memory points, updates, distractors, and evidence chains across multiple sessions. sequential-chain no none statedother:not-stated 10.25/mo
π-Bench
π-Bench: Evaluating Proactive Personal Assistant Agents in Long-Horizon Workflows
Assistants & memory 2026 identify and act on hidden/unstated user needs before they are explicitly stated; complete each of 100 multi-turn tasks across 5 domain-specific user personas; resolve inter-task dependencies across sessions; maintain continuity of prior interaction context across sessions to resolve later proactive intents The agent must track hidden/unstated user intents, inter-task dependencies, and information carried over across sessions, jointly measuring proactivity (anticipating needs) and task completion (executing them) over extended interactions. DAG-with-precedence no none stated 30.75/mo
Evo-Memory
Evo-Memory: Benchmarking LLM Agent Test-time Learning with Self-Evolving Memory
Assistants & memory 2025 search, adapt, and evolve memory after each interaction in a sequential task stream; solve each task drawn from 10 diverse multi-turn goal-oriented and single-turn reasoning/QA datasets; reuse experience accumulated from earlier tasks to improve on later tasks in the same stream Agent must search, adapt, and evolve its own memory continuously after each interaction across a sequential task stream, integrating reasoning, task actions, and memory updates to achieve continual improvement. sequential-chain no none stated 12412.4/mo
Forgetful but Faithful Agent (FiFA) benchmark
Forgetful but Faithful: A Cognitive Memory Architecture and Benchmark for Privacy-Aware Generative Agents
Assistants & memory 2025 maintain narrative coherence across long-term interaction; complete multi-step goals despite a bounded memory budget; preserve social recall accuracy under six candidate forgetting policies; preserve privacy by not retaining/leaking information beyond its retention schema; minimize cost while balancing the above under different memory budgets The agent must track its own memory budget consumption, which information to retain vs. forget under its chosen policy, ongoing multi-step goal progress, and social/privacy-relevant facts, across long-term interactive scenarios. set-of-independent no none statedother:not-stated 111.22/mo
Mem-alpha
Mem-α: Learning Memory Construction via Reinforcement Learning
Assistants & memory 2025 extract and store relevant content from sequential information chunks into an external memory system; organize stored content across core, episodic, and semantic memory components; correctly answer downstream questions using the full accumulated interaction history; invoke the right memory-operation tool (from a multi-tool memory architecture) at the right time The agent must track which information has been extracted and stored, how it is structured across core/episodic/semantic memory, and how it should be updated as new chunks arrive, since reward is downstream QA accuracy over the full interaction history. sequential-chain yes trained up to 30k tokens; generalizes to sequences exceeding 400k tokens (>13x training length)other:tokens-of-accumulated-interaction-history 272.25/mo
MemoryAgentBench
Evaluating Memory in LLM Agents via Incremental Multi-Turn Interactions
Assistants & memory+ Information seeking 2025 accurately retrieve previously seen information from accumulated context; adapt to and learn from new information at test time; understand and reason over long-range accumulated context; selectively forget information that is no longer relevant or valid The memory agent must incrementally accumulate, update, and retrieve information across multi-turn interactions, and correctly discard information rendered obsolete while retaining what remains useful. set-of-independent yes none stated 21315.21/mo
MemoryBench
MemoryBench: A Benchmark for Memory and Continual Learning in LLM Systems
Assistants & memory 2025 learn from accumulated user feedback received during service time; apply continually-updated knowledge correctly across multiple domains, languages, and task types; outperform static (non-continual-learning) baselines after repeated feedback exposure The system must track and integrate a stream of accumulated user feedback over service time, across multiple domains, languages, and task types, updating its internal state/parameters rather than treating each query independently. sequential-chain no none stated 494.45/mo
PERSONAMEM
Know Me, Respond to Me: Benchmarking LLMs for Dynamic User Profiling and Personalized Responses at Scale
Assistants & memory 2025 internalize a user's inherent traits and preferences from interaction history; track how the user's profile/preferences evolve over time across sessions; generate a personalized response consistent with the current (most up-to-date) state of the user's profile in a new scenario The chatbot must internalize inherent user traits, track how the user's profile evolves session by session, and select the response consistent with the user's current (not stale) profile state when answering an in-situ, first-person query. sequential-chain no over 180 simulated interaction histories, each containing up to 60 sessions of multi-turn conversationssessions 1589.29/mo
PROBE
Beyond Reactivity: Measuring Proactive Problem Solving in LLM Agents
Assistants & memory+ Business & enterprise 2025 search for unspecified, unprompted issues across the user's available context/data; identify the specific bottleneck underlying an ambiguous problem; execute an appropriate resolution action autonomously The agent must track candidate evidence surfaced during open-ended search, determine which issues are genuine unresolved bottlenecks, and carry that determination forward into selecting and executing a resolution, all without being told what to look for. sequential-chain yes none stated 80.73/mo
SimuHome
SimuHome: A Temporal- and Environment-Aware Benchmark for Smart Home LLM Agents
Assistants & memory 2025 answer state-inquiry questions about the current smart-home environment; infer implicit user intent behind an ambiguous request; execute explicit device-control commands via SimuHome APIs; schedule and coordinate multi-device workflows whose effects evolve environmental variables over time; recognize and appropriately reject infeasible requests The agent must track how already-issued device commands change environmental variables over (accelerated) simulated time and use this state to judge whether new requests are feasible or already satisfied. DAG-with-precedence yes 600 episodes; simulator accelerates time so scheduled workflows can be evaluated immediatelyepisodes 70.58/mo
UserBench
UserBench: An Interactive Gym Environment for User-Centric Agents
Assistants & memory 2025 proactively clarify a simulated user's underspecified/vague initial goal; incrementally uncover and track multiple user preferences revealed gradually over a multi-turn interaction; make grounded decisions with tools that align with all of the user's (eventually revealed) intents The agent must track which of the user's underspecified goals and incrementally revealed preferences it has already clarified, and continue to proactively elicit unclarified ones across the multi-turn interaction while using tools to act on what it has learned so far. set-of-independent no none stated 624.43/mo
Science14 artifacts
ASI-Bench
ASI-Bench: At the Dawn of Artificial Superintelligence
Science 2026 independently select an appropriate research method for a project-level research task; conduct the research (experimentation/analysis) using the selected method; produce verifiable results at the level of a full research project, with progressively less methodological guidance provided across three guidance tiers System must track which methodological guidance tier is currently active for a task, whether a method has been selected, and whether the resulting research output is verifiable against expert review, across an entire project-level research process. hierarchical yes none stated 22.0/mo
AstroReason-Bench
AstroReason-Bench: Evaluating Unified Agentic Planning across Heterogeneous Space Planning Problems
Science 2026 schedule ground-station communication passes; schedule agile Earth-observation tasks; satisfy heterogeneous mission objectives under strict physical constraints within a Space Planning Problem instance Agent must track scheduling state across multiple heterogeneous objectives (communication windows, observation opportunities) and physical/orbital constraints simultaneously. DAG-with-precedence no none stated 10.12/mo
BixBench3
BixBench3: Benchmarking AI agents on research-study-scale computational biology tasks
Science 2026 execute a sequence of computational-biology analyses from raw data through to a research objective; produce each of multiple data artifacts (e.g., peak call matrices, differential expression tables) matching the original published study; manage large raw datasets (some over 100GB) within time/cost constraints; maintain coherence across multiple sequential analysis steps The agent must track intermediate data artifacts produced at each analysis step, manage large raw datasets, and maintain coherence across multiple sequential analyses before producing final gradable artifacts. sequential-chain yes average 6.8 hours per task (longest attempts up to 24 hours)wall-clock-hours 0
ChemCost
Can Agents Price a Reaction? Evaluating LLMs on Chemical Cost Reasoning
Science 2026 ground chemical identities from a reaction description; retrieve supplier quotes for the grounded chemicals; select valid purchasable packs matching required quantities; normalize quantities across packs; compute the total procurement cost from the reaction description The agent must track which chemicals have been correctly grounded, which supplier quotes and packs have been retrieved/selected, and carry normalized quantities forward correctly into the final arithmetic cost computation, especially under noise-injected perturbations (aliases, quantity expressions, missing fields, formatting). sequential-chain no none statedother:not-stated 10.25/mo
DiscoverPhysics
DiscoverPhysics: Benchmarking LLMs for Out-of-the-Box Scientific Thinking
Science 2026 design a sequence of informative experiments to probe an unknown simulated world's physics; revise hypotheses about the governing physical law across multiple rounds based on observed trajectory data; submit both a natural-language explanation and a Python implementation of the inferred law for each of 22 worlds Agent must track its accumulated experimental observations (trajectory data) and current hypothesis about the world's physics across several rounds before submitting a final explanation and implementation. sequential-chain yes none stated 41.0/mo
iNatDisco
Autonomous Scientific Discovery via Iterative Meta-Reflection
Science 2026 generate scientific hypotheses about ecological patterns from a dataset without a pre-specified research question; validate each proposed hypothesis via statistical testing before accepting it; periodically synthesize accumulated prior discoveries to redirect exploration toward unexplored regions of the hypothesis space; incorporate multimodal tool use (e.g., image processing) to extract further supporting evidence The agent must maintain a growing record of prior discoveries and their statistical validation status, and periodically re-analyze that record to identify structural patterns, confounds, and epistemic gaps. open-ended yes none stated 10.5/mo
InquiTree (IT-18 subset)
InquiTree: Evaluating AI Agents in the Scientific Inquiry Loop with Paper-Derived Research Trees
Science 2026 formulate a hypothesis consistent with a paper-derived research tree's logical dependencies; design a study to test the hypothesis; interpret the resulting study outcome; update beliefs/conclusions based on interpreted results, propagating correctly through the DAG Agents must track their evolving beliefs/conclusions across nested subtopic branches, remain consistent with prior hypothesis/study/interpretation nodes, and avoid 'Erosion of Marginal Capabilities' (degrading critical judgment) over long-horizon interactions. DAG-with-precedence yes theoretical interaction range of roughly 360 (3x120) to 1320 (11x120) reasoning-action turns for full traversal of the IT-18 subset (120 subtopics, H=4); per-task bounds of 21-77 steps for a typical task (n=7, H=4)turns 10.33/mo
LabOSBench
LabOSBench: Benchmarking Computer Use Agents for Scientific Instrument Control
Science 2026 complete each stage of a scientific-instrument operation workflow: sample loading, alignment, parameter tuning, data acquisition, and result inspection; perform feedback-driven parameter adjustment based on live instrument readouts; correctly operate one of 8 distinct instrument simulators across 96 subtasks Agent must track which workflow stage it is in (loading, alignment, tuning, acquisition, inspection), instrument readouts/feedback from its own prior actions, and calibration state carried from earlier stages. sequential-chain yes none stated 10.33/mo
LifeSciBench
LifeSciBench: Evaluating Language Models on Realistic, Expert-Level Tasks in the Life Sciences
Science 2026 execute a chain of multiple dependent judgment calls within one realistic life-science research task; satisfy a human-expert-written rubric spanning one of seven representative scientific workflows; operate correctly across each of seven life-science domains The agent must track which dependent judgment calls it has made so far within a task and ensure later calls remain consistent with earlier ones, as graded against a human expert-written rubric. sequential-chain yes none statedother:not-stated 44.0/mo
RWE-bench
Can LLM Agents Generate Real-World Evidence? Evaluating Observational Studies in Medical Databases
Science+ Healthcare 2026 construct a patient cohort from a real database (MIMIC-IV) per a study protocol; perform an analysis matching a peer-reviewed observational study's methodology; produce a coherent, tree-structured evidence bundle for reporting; iteratively execute and refine experiments against the reference protocol The agent must track its cohort definition, intermediate analysis results, and their organization into a tree-structured evidence bundle, checking internal coherence against the reference study protocol across the whole task. hierarchical yes none stated 20.33/mo
SciAgentGym / SciAgentBench
SciAgentGym: Benchmarking Multi-Step Scientific Tool-use in LLM Agents
Science 2026 orchestrate domain-specific scientific tools correctly across four natural science disciplines; complete tasks across a tiered difficulty spectrum from elementary actions to long-horizon workflows; sustain performance as interaction horizons extend rather than degrading The agent must track which of 1,780 domain-specific tools it has invoked and their outputs across a tiered workflow, since performance is shown to degrade substantially as the interaction horizon extends, requiring sustained state-tracking rather than one-off calls. hierarchical no none stated 91.29/mo
AFMBench
Evaluating large language model agents for automation of atomic force microscopy
Science 2025 complete the full scientific workflow from experimental design to results analysis for atomic force microscopy (AFM) automation; successfully perform each of several increasingly advanced experiments: AFM calibration, feature detection, mechanical property measurement, graphene layer counting, and indenter detection; coordinate across multi-agent roles in laboratory settings without deviating from instructions ('sleepwalking') Agent(s) must track experimental state and configuration across the full workflow (design, calibration, measurement, analysis), coordinate with other agents in multi-agent setups, and avoid deviating from given instructions across a sequence of physical lab actions. hierarchical yes none stated 615.55/mo
BoxingGym
BoxingGym: Benchmarking Progress in Automated Experimental Design and Model Discovery
Science 2025 design and run informative experiments that reduce uncertainty about a generative model's parameters; propose a scientific model/theory of the given environment; revise the proposed theory in light of newly collected experimental data; produce an explanation of the model that lets another agent make reliable predictions The agent must track what has been learned from experiments already run (to target expected information gain for the next experiment) and maintain an evolving explanation of its current best scientific model. sequential-chain yes evaluated after 0, 1, 3, 5, 7, and 10 experiment-design steps per trial, with 5 independent trials per environmentactions 110.55/mo
ScienceBoard
ScienceBoard: Evaluating Multimodal Autonomous Agents in Realistic Scientific Workflows
Science 2025 autonomously interact with professional scientific software (biochemistry, astronomy, geoinformatics) to accomplish a research task; complete each of 169 rigorously validated real-world scientific-discovery workflow tasks Agent must track intermediate results and state produced while autonomously interacting with dynamic, visually rich professional software across a multi-step scientific workflow. sequential-chain yes none stated 432.69/mo
Tool & API29 artifacts
Benchmarking AI Agents for Hardware Design Automation via MCP Tool Calling
Benchmarking AI Agents for Hardware Design Automation via MCP Tool Calling
Tool & API 2026 execute single-operation edits on hardware design components via specialised MCP tools; correctly sequence multi-step dependency chains (e.g. create component -> add port -> wire connection); handle invalid or misspelled requests without incorrect tool invocation; operate correctly across multi-server tool contexts The agent must track which prior dependency-establishing tool calls (e.g. component creation, port creation) have already succeeded before issuing calls that depend on them, across single-agent and multi-agent tool-calling configurations. DAG-with-precedence yes 1 (Easy) / 2 (Medium) / 3-5 (Hard) expected tool calls per tasktool-calls 0
AgentEscapeBench
AgentEscapeBench: Evaluating Out-of-Domain Tool-Grounded Reasoning in LLM Agents
Tool & API 2026 invoke real external tool functions to satisfy a directed acyclic dependency graph over tools and items; track hidden state revealed incrementally through tool use; propagate intermediate results correctly across dependent tool calls to a deterministically verifiable final answer Agent must track hidden state revealed incrementally by tool calls, maintain clue adherence, and correctly propagate intermediate results through the dependency graph as depth increases. DAG-with-precedence yes difficulty tiers from 5 to 25 (dependency-graph depth levels)other:dependency-graph-depth-tiers 0
AgentFloor
AgentFloor: How Far Up the tool use Ladder Can Small Open-Weight Models Go?
Tool & API 2026 follow instructions correctly (lowest tier); use tools correctly (mid-lower tier); coordinate multiple steps/tool calls together (mid-upper tier); sustain long-horizon planning under persistent constraints over many steps (top tier) Agent must sustain constraint tracking and coordination reliably over many steps to succeed at the highest tiers, where 'neither side reaches strong reliability' even among frontier models. hierarchical unclear none stated 0
AgentGym2
AgentGym2: Benchmarking Large Language Model Agents in De-Idealized Real-World Environments
Tool & API 2026 execute end-to-end real-world procedures without relying on pre-packaged tool interfaces; discover available tools via active exploration of the environment; compose discovered tools to solve previously unseen tasks; remain robust to noisy and underspecified task information Agent must track which tools/interfaces it has discovered so far, what it has verified about their (possibly noisy) behavior, and how these compose toward completing the current end-to-end task. open-ended yes none stated 10.5/mo
APIFlow-Bench
APIFlow-Bench: Measuring Whether Agents Survive Long, Dependent API Workflows
Tool & API 2026 complete each subtask in a long, dependent chain of REST-API calls correctly (state, not just completion, must be right); produce a final answer whose delivery is traceable to the actual call path (provenance-sensitive correctness); avoid compounding failures across increasingly long dependency chains (up to 20 subtasks) Agent must track state correctness at each subtask in a dependency chain (not just whether the chain 'completed'), including whether a mock-minted canary correctly propagates through the API data flow to the final delivered answer. DAG-with-precedence yes 20 (clean chains); up to 44,362 execution transcripts releasedother:subtasks-per-workflow-chain 0
AppWorld-UL
AppWorld-UL: Benchmarking Diverse Agent-User Interactions for Tool-Use
Tool & API+ Assistants & memory 2026 operate applications correctly to complete a digital task (e.g. ordering groceries) across 9 simulated apps; interact appropriately with the user: ask clarification questions, prompt for confirmation, or report infeasibility; succeed on compositional, multi-part sub-tasks that combine several of the above interaction types Agent must track what has been asked/confirmed with the user so far, the user's carefully-bounded knowledge state (simulated by an LLM), and which parts of a compositional task remain to be completed. hierarchical yes none stated 10.5/mo
AsyncTool
AsyncTool: Evaluating the Asynchronous Function Calling Capability under Multi-Task Scenarios
Tool & API 2026 concurrently manage multiple heterogeneous tasks presented simultaneously; make productive use of idle time while awaiting delayed tool-call responses (asynchronous tool calling); coordinate task switching, dependency tracking, and state maintenance across concurrently running tasks Agent must track the state, dependencies, and pending tool responses of multiple simultaneously active tasks, and coordinate task-switching decisions during periods of delayed tool feedback. DAG-with-precedence yes none stated 30.75/mo
C-World
C-World: A Computer Use Agent Environment Creator
Tool & API 2026 complete long-horizon workflows composed of many interacting constraints across up to 5,571 tools spanning 204 applications; correctly follow constraints despite injected realistic failures and perturbations during the task; satisfy a reward signal combining verifiable metrics with LLM-based judgment Agent must track constraint satisfaction across a long-horizon, multi-tool workflow while detecting and adapting to injected failures/perturbations introduced by the environment's transition function. DAG-with-precedence yes none stated 10.12/mo
CCTU
CCTU: A Benchmark for Tool Use under Complex Constraints
Tool & API 2026 select and call the correct tool while satisfying every one of several simultaneous constraints (resource, behavior, toolset, response dimensions); maintain compliance with all constraints across a multi-turn interaction, not just the first tool call; self-refine after receiving feedback about a constraint violation The agent must track which of the ~7 simultaneous constraints (out of 12 categories across 4 dimensions) apply to the current tool-use scenario and re-check compliance with all of them at every step across the multi-turn interaction, especially after receiving feedback about a violation. set-of-independent yes maximum 20 interaction rounds per test caseturns 71.17/mo
GeoAgentBench (GABench)
GeoAgentBench: A Dynamic Execution Benchmark for Tool-Augmented Agents in Spatial Analysis
Tool & API 2026 correctly configure parameters for each of 117 atomic GIS tools invoked; complete each of 53 typical spatial-analysis tasks across 6 core GIS domains; produce spatially/cartographically accurate outputs verified via a VLM-based check; decouple global workflow orchestration from step-wise reactive execution to recover from runtime anomalies The agent must track its evolving execution plan, per-step parameter choices, and runtime feedback/anomalies across a multi-step GIS workflow to keep global orchestration consistent with step-wise reactive execution. sequential-chain yes none stated 10.2/mo
GTA-2 (GTA-Atomic / GTA-Workflow)
GTA-2: Benchmarking General Tool Agents from Atomic Tool-Use to Open-Ended Workflows
Tool & API+ Assistants & memory 2026 execute short-horizon, closed-ended atomic tool calls correctly (GTA-Atomic); complete long-horizon, open-ended, real-world productivity workflows end-to-end (GTA-Workflow); satisfy verifiable sub-goals identified by a recursive checkpoint-based evaluation mechanism System must track which recursively-decomposed sub-goals/checkpoints of an open-ended workflow have been satisfied so far, across real deployed tools and multimodal contexts. hierarchical yes none stated 10.2/mo
Ko-WideSearch
Ko-WideSearch: A Korean Breadth-Search Benchmark for Exhaustive Set Enumeration by Web Agents
Tool & API+ Web & GUI 2026 exhaustively enumerate the full membership of a named closed set (e.g. a TV season's cast, a dynasty's rulers); fill a per-item attribute table (multiple columns) for every enumerated member; decide when to stop searching within an open-ended web search space The agent must track which set members it has already found (to avoid duplicates/omissions), which attribute columns remain unfilled per member, and its remaining search-iteration budget as difficulty knobs (table width, 2-D composite key) increase. set-of-independent no fixed budget of 30 agent iterations per question (total tool calls run higher; one model logged up to 947 tool calls)tool-calls 0
MM-ToolBench (TOBench)
TOBench: A Task-Oriented Omni-Modal Benchmark for Real-World Tool-Using Agents
Tool & API 2026 execute tools appropriate to a Customer Service or Intelligent Creation task; inspect rendered or transformed intermediate artifacts produced by tool calls; self-correct when an inspected artifact fails task-specific requirements; satisfy each of the 20 subcategory slices' task-specific grounded evaluator checks The agent must track the state of rendered/transformed artifacts across a closed verification loop, deciding when self-correction is required, using task-specific grounded evaluators across 27 MCP servers and 324 tools. sequential-chain no none statedother:not-stated 0
MM-ToolSandBox
MM-ToolSandBox: A Unified Framework for Evaluating Visual Tool-Calling Agents
Tool & API 2026 ground progressively arriving visual inputs into correct executable tool calls across a multi-image, multi-turn interaction; handle realistic conversational phenomena mid-task: goal revisions, error corrections, state mutations; operate correctly across 500+ tools spanning 16 application domains; succeed on each of 258 human-verified nominal scenarios (plus 50 interactive-UI variants) The agent must track a stateful execution environment across multi-image, multi-turn interactions, correctly incorporating goal revisions, error corrections, and state mutations as they occur, and ground each newly arriving visual input into the correct tool call. sequential-chain no none stated 0
OmnilingualGAIA2
OmnilingualGAIA2: Evaluating the Multilingual Gap in Frontier AI Agents
Tool & API 2026 plan a sequence of tool calls to answer a task; search for information via tools; execute multi-tool workflows; recover from errors during multi-tool execution; all under machine-translated, human-calibrated task instructions across ten languages/five scripts Agent must track its tool-call plan, intermediate search/tool results, and error states to recover from execution failures, now while also handling machine-translated instructions across ten languages including non-Latin scripts. DAG-with-precedence yes none stated 0
SeekerGym
SeekerGym: A Benchmark for Reliable Information Seeking
Tool & API 2026 issue repeated retrieval queries to recover as much of a target document's content as possible; quantify uncertainty about how much relevant information might still be missing from what has been retrieved The agent must track which passages/sections of the target document it has already retrieved, so it can judge how complete its coverage is and estimate how much information might still be missing. open-ended yes none stated 0
SkillCraft
SkillCraft: Can LLM Agents Learn to Use Tools Skillfully?
Tool & API 2026 compose atomic tools into reusable higher-level 'Skills'; cache and reuse learned Skills both within a task and across different tasks; complete compositional tool-use scenarios whose difficulty scales with entity count and subtask complexity The agent must track which higher-level Skills it has already formed and cached, so it can reuse them instead of recomposing atomic tools from scratch on later subtasks/tasks, both within a single task and across the benchmark's 126 tasks. hierarchical yes 126 tasks across 6 difficulty levels (from 21 seed tasks); tool-call counts scale from 9 (Easy: 3 subtasks x 3 calls) to 25 (Hard: 5 subtasks x 5 calls)tool-calls 294.14/mo
The Amazing Agent Race (AAR)
The Amazing Agent Race: Strong Tool Users, Weak Navigators
Tool & API+ Embodied & robotics 2026 navigate Wikipedia pages to locate the required entity/fact for each DAG node; execute the correct multi-step tool chain across fork-merge branches linking extracted entities; aggregate branch outputs into one verifiable final answer The agent must track which Wikipedia entities/facts it has already extracted at each DAG node, correctly route them through parallel fork-merge tool-chain branches, and retain intermediate values needed for final aggregation across up to 5 diamonds per leg. DAG-with-precedence yes average of 22.1 pit stops per leg for AAR-DAG (600 legs) and 15.0 pit stops per leg for AAR-Linear (800 legs), up to 5 diamonds per legother:pit-stops (navigation hops) 10.2/mo
Toolathlon
Capable but Unreliable: Canonical Path Deviation as a Causal Mechanism of Agent Failure in Long-Horizon Tasks
Tool & API 2026 complete each of 108 real-world tool-use tasks by following a canonical multi-step tool-invocation solution path; stay within the operating envelope of the canonical path across the whole trajectory to avoid stochastic drift/derailment Success requires the trajectory of tool calls to stay within the operating envelope of the task's canonical solution path; mid-trajectory adherence must be tracked since drift compounds over subsequent calls. sequential-chain yes none stated 30.43/mo
ToolGym
ToolGym: an Open-world Tool-using Environment for Scalable Agent Testing and Data Curation
Tool & API 2026 complete long-horizon, multi-tool workflows synthesized with wild constraints across 5,571 tools and 204 apps; recover and adapt when a state controller injects interruptions and failures mid-workflow; separate deliberate planning/self-correction from step-wise execution via a planner-actor decomposition Agent must track tool-call state and wild constraints across a long-horizon, multi-tool workflow while detecting and recovering from injected interruptions/failures and unreliable tool states. sequential-chain no none stated 51.67/mo
ToolVerse / GUST dataset
ToolVerse: Unlocking Massive Environments and Long-Horizon Tasks for Agentic Reinforcement Learning
Tool & API 2026 complete a long-horizon task built from a tool-dependency graph, where later tools unlock only after prerequisite tool-use subgoals are completed; correctly integrate the right subset of tools from a large pool (~4500 tools across ~400 MCPs) into a coherent multi-step solution; receive properly assigned turn-level credit despite long sequences of tool calls The agent must track which prerequisite tools/subgoals in the dependency graph it has already satisfied in order to know which further tools are unlocked and relevant next, across a long sequence of tool calls. DAG-with-precedence yes none stated 21.0/mo
UniClawBench
UniClawBench: A Universal Benchmark for Proactive Agents on Real-World Tasks
Tool & API 2026 demonstrate correct skill usage for the task's required tool/capability; explore the live environment (e.g. a Docker container) to discover needed information/actions; reason correctly over long context accumulated during the task; correctly interpret multimodal inputs relevant to the task; coordinate actions correctly across multiple platforms/services The agent must track its progress against fine-grained, step-by-step completion checkpoints within a live environment, integrate multi-turn feedback from a hidden supervisor agent and a user-simulator agent without seeing the grading criteria, and coordinate across the relevant capability dimensions needed for that task. sequential-chain yes none stated 10.5/mo
CONFETTI
CONFETTI: Conversational Function-Calling Evaluation Through Turn-Level Interactions
Tool & API 2025 handle user follow-up requests within an ongoing conversation; correct or switch goals mid-conversation when the user changes intent; resolve ambiguous or implicit user goals into an appropriate API call; chain multiple function calls together to satisfy a multi-step user request The agent must track evolving user intent across conversation turns, recognize goal correction/switching, and maintain state of prior chained function-call results across up to 86 available APIs. sequential-chain yes 313 user turns across 109 conversations (~2.9 turns/conversation on average)turns 161.07/mo
DialogTool / VirtualMobile
Rethinking Stateful Tool Use in Multi-Turn Dialogues: Benchmarks and Challenges
Tool & API+ Web & GUI 2025 create a new tool/API on demand within a multi-turn dialogue (tool creation); become aware of, correctly select, and execute the right tool given user intent (tool utilization); generate a role-consistent response, including role play, reflecting whether/how the tool was used Agent must track which tools have been created/exist, their stateful execution history across the dialogue, and how to remain role-consistent while using them, across a multi-turn conversation. sequential-chain yes none stated 171.06/mo
M^3-Bench
M^3-Bench: Multi-Modal, Multi-Hop, Multi-Threaded Tool-Using MLLM Agent Benchmark
Tool & API 2025 complete multi-hop, multi-threaded tool-call workflows requiring cross-tool dependencies; maintain persistence of intermediate resources across steps; ground visual and textual reasoning correctly across each tool call Agent must track intermediate resources produced by earlier tool calls and align them across multiple concurrent threads for argument fidelity and structural consistency. DAG-with-precedence yes none stated 90.9/mo
Multi-Mission Tool Bench
Multi-Mission Tool Bench: Assessing the Robustness of LLM based Agents through Related and Dynamic Missions
Tool & API 2025 complete each mission within a test case containing multiple interrelated missions; dynamically adapt when missions switch mid-interaction; correctly invoke tools appropriate to the currently active mission; handle all possible mission-switching patterns within a fixed mission number The agent must track the state of each interrelated mission (active or paused), correctly recognize mission switches, and select/invoke the correct tools for whichever mission is currently active, evaluated via dynamic decision trees for accuracy and efficiency. DAG-with-precedence no none stated 70.41/mo
OrchDAG
OrchDAG: Complex Tool Orchestration in Multi-Turn Interactions with Plan DAGs
Tool & API 2025 correctly execute a sequence of tool calls whose dependencies form a directed acyclic graph (DAG); respect topological/precedence constraints among tool calls across multi-turn interactions; solve DAG-structured tool-orchestration tasks of controllable/varying complexity The agent must track which nodes (tool calls) in the DAG have been completed and what state/outputs they produced, to correctly select and sequence subsequent tool calls consistent with the graph's precedence constraints across multiple turns. DAG-with-precedence no none stated 10.09/mo
Tool Decathlon (Toolathlon)
The Tool Decathlon: Benchmarking Language Agents for Diverse, Realistic, and Long-Horizon Task Execution
Tool & API+ Business & enterprise 2025 coordinate interactions across multiple named Apps/tools (e.g. email + calendar + file systems) to complete one complex workflow; diagnose and report anomalies via monitoring a database following an operating manual; satisfy a strictly execution-verifiable end state for each of 108 tasks spanning 32 apps and 604 tools The agent must track the current, realistic state of multiple named applications across roughly 20 tool-calling turns per task, and verify its final state against a dedicated evaluation script. DAG-with-precedence yes tasks require interacting with multiple Apps over around 20 turns on average (best model averages 20.2 tool-calling turns)turns 666.0/mo
ToolHaystack
ToolHaystack: Stress-Testing Tool-Augmented Language Models in Realistic Long-Term Interactions
Tool & API 2025 correctly maintain and disambiguate multiple concurrent task-execution contexts within one continuous conversation; handle realistic noise/disruptions injected into a long-term interaction without losing track of tool-use context; successfully complete tool-use tasks embedded within a long, continuous conversation despite these challenges The model must maintain and disambiguate multiple concurrent task-execution contexts across a continuous long-term conversation while filtering out realistic noise and handling various disruptions, rather than resetting context between short, isolated tool-use exchanges. set-of-independent no none stated 50.31/mo
Planning & travel19 artifacts
Behavior2Trip
Behavior2Trip: Towards Personalized Travel Planning via User Behavior Trajectory
Planning & travel+ Assistants & memory 2026 infer a user's latent travel preferences from their past behavior trajectory rather than explicit instructions; generate a travel plan satisfying preferences across 14 attributes spanning 5 preference dimensions; achieve a full-constraint pass across all inferred preference constraints simultaneously The agent must infer and track the user's latent preferences (14 attributes, 5 dimensions) from an average of 39.8 past behaviors, then check the generated plan against every inferred constraint for a full pass. set-of-independent yes none stated 0
DeepPlanning
DeepPlanning: Benchmarking Long-Horizon Agentic Planning with Verifiable Constraints
Planning & travel+ Information seeking 2026 satisfy local, fine-grained constraints on individual itinerary/shopping items; satisfy global constrained optimization objectives (e.g. time and financial budgets) across the whole plan; proactively gather information needed before constraints can even be checked Agent must track accumulated time and financial budget consumption across a multi-day plan or multi-product list, alongside fine-grained local constraints, while still actively gathering new information mid-task. DAG-with-precedence unclear none stated 334.12/mo
GroupTravelBench
GroupTravelBench: Benchmarking LLM Agents on Multi-Person Travel Planning
Planning & travel 2026 elicit each group member's private travel preferences through multi-turn dialogue; surface and resolve inter-user preference conflicts via compromise or subgrouping; produce a final plan that balances group utility against fairness across all members Agent must track each group member's elicited (and initially private) preferences, detected conflicts between members, and the evolving fairness/utility trade-off of the plan across a synchronous multi-turn group-chat session. DAG-with-precedence yes none stated 41.0/mo
TravelEval
TravelEval: A Comprehensive Benchmarking Framework for Evaluating LLM-Powered Travel Planning Agents
Planning & travel 2026 produce a full multi-day itinerary satisfying accuracy, compliance, temporality, spatiality, economy, and utility dimensions jointly; sequence daily accommodation, transport, and visit pacing consistently across the whole trip rather than per isolated day The agent must track cumulative spatio-temporal cost (queuing times, transit distances), running budget/economy, and daily pacing/accommodation continuity across the entire multi-day itinerary, not just within a single day. sequential-chain yes itineraries span 2-day, 3-day, 4-day, 5-day, 6-day, and 7-day durations across query categories (e.g. 400 medium-difficulty queries)simulated-days 10.25/mo
TREK
TREK: A Travel Reasoning and Evaluation Kit for LLM Agents in Complex Trip Planning
Planning & travel 2026 produce a single itinerary that is jointly constraint-correct, hallucination-free, spatio-temporally executable, and budget-valid; respond to the traveler's unstated persona needs; correctly identify provably infeasible tasks (267 of 800) versus feasible ones with typed causes Agent must track running budget, spatio-temporal feasibility across days, entity/route validity, and unstated persona needs simultaneously while assembling a single itinerary. other (singleton) no none stated 0
Trip+
Trip+: Benchmarking Agents in Personalized Interactive Travel Planning
Planning & travel+ Assistants & memory 2026 generate a minute-level itinerary satisfying a traveler's profiled preferences; revise the itinerary in response to evolving preferences and unexpected environment-driven disruptions across multiple turns; avoid producing technically-feasible-but-exhausting plans (jointly satisfy feasibility and experiential/fatigue quality) The agent must track the traveler's evolving profile/preferences, the current committed itinerary state at minute-level granularity, and cumulative experiential cost (e.g. fatigue) as it revises plans across several user turns per instance. sequential-chain yes 153 multi-turn instances and 570 user turns totalturns 10.33/mo
TRIP-Bench
TRIP-Bench: A Benchmark for Long-Horizon Interactive Agents in Real-World Scenarios
Planning & travel 2026 satisfy each of 40+ curated travel requirements/global constraints across an itinerary; coordinate reasoning across 18 curated tools correctly within one dialogue; adapt to evolving user behavior, style shifts, feasibility changes, and iterative version revisions over a long, multi-turn interaction (hard split) The agent must track global constraint satisfaction across 40+ travel requirements, the state of 18 tools' results, and evolving user preferences/style/feasibility across dialogues spanning up to 15 user turns, 150+ tool calls, and 200k+ tokens of context. DAG-with-precedence yes dialogues span up to 15 user turns, can involve 150+ tool calls, and may exceed 200k tokens of contextturns 71.0/mo
Trip-planning Optimization Problems (TOP) Dataset
Agentic AI for Trip Planning Optimization Application
Planning & travel 2026 optimize (not merely satisfy) route/itinerary selection under travel-time, energy, and traffic factors; coordinate specialized sub-agents for traffic, charging, and points-of-interest; dynamically refine a plan against a definitive optimal solution reference The orchestration agent must track recommendations from the traffic, charging, and POI sub-agents and reconcile them into one jointly-optimal plan, verified against the dataset's definitive optimal solutions rather than only a feasible reference answer. hierarchical no none statedother:not-stated 0
WorldTravel / WorldTravel-Webscape
WorldTravel: A Realistic Multimodal Travel-Planning Benchmark with Tightly Coupled Constraints
Planning & travel 2026 satisfy all of an average 15+ interdependent temporal and logical travel constraints simultaneously per scenario; extract constraint parameters from dynamic web environments/webpages rather than idealized data; perceive constraint parameters directly from visual layouts in a multi-modal setting; produce a feasible travel plan across 150 real-world scenarios in 5 cities The agent must track which of the 15+ interdependent temporal/logical constraints have been satisfied so far, and in the multi-modal condition must also perceive constraint parameters directly from over 2,000 rendered webpages rather than being handed clean structured data. DAG-with-precedence no an average of 15+ interdependent constraints per scenario; a Planning Horizon threshold at approximately 10 constraintsother:number-of-interdependent-constraints 30.43/mo
COMPASS
COMPASS: Benchmarking Constrained Optimization in LLM Agents
Planning & travel 2025 gather task information and constraints from the user via multi-turn conversation; use tools to gather relevant information from a database; propose a travel plan satisfying all hard constraints; optimize the plan for the user's utility objective beyond mere feasibility Agent must track constraints and preferences gathered so far via conversation and tool calls, the current feasible-solution search space, and how well a candidate plan satisfies both hard constraints and the utility objective. sequential-chain no none stated 60.55/mo
CostBench
CostBench: Evaluating Multi-Turn Cost-Optimal Planning and Adaptation in Dynamic Environments for LLM Tool-Use Agents
Planning & travel 2025 find a cost-optimal sequence of atomic and composite tool calls to solve a travel-planning task; detect and adapt to dynamic blocking events (e.g. tool failures, cost changes) that occur mid-task; replan in real time to remain cost-optimal after a blocking event Agent must track the accumulated cost of its chosen tool sequence so far, remaining budget, and whether any of four types of dynamic blocking events have occurred, requiring real-time replanning to stay cost-optimal. DAG-with-precedence unclear none stated 303.0/mo
Flex-TravelPlanner
Flex-TravelPlanner: A Benchmark for Flexible Planning with Language Agents
Planning & travel 2025 revise a travel plan as new constraints are introduced sequentially across turns; correctly prioritize competing constraints when a newly introduced lower-priority preference conflicts with an existing higher-priority constraint The agent must track which constraints have been introduced so far, their relative priority, and whether previously satisfied requirements are still respected as new constraints arrive turn by turn. sequential-chain yes up to 3 turns (all-at-once / 2-turn / 3-turn constraint-introduction patterns) across 120 base queriesturns 110.73/mo
RETAIL
RETAIL: Towards Real-world Travel Planning for Large Language Models
Planning & travel+ Business & enterprise 2025 infer and satisfy implicit user requirements (not just explicit queries); satisfy explicit queries with or without later revision needs; account for diverse environmental factors and constraints to ensure plan feasibility; produce an all-in-one plan with rich, detailed POI (point-of-interest) arrangement rather than only basic POI listing The agent must infer unstated (implicit) requirements, track environmental constraints affecting plan feasibility, and assemble detailed POI information into a single all-in-one plan, revising an existing plan when a revision need is present. DAG-with-precedence no none stated 90.69/mo
Travel-Sim
Wide-Horizon Thinking and Simulation-Based Evaluation for Real-World LLM Planning with Multifaceted Constraints
Planning & travel 2025 satisfy multiple parallel, potentially conflicting real-world planning constraints (e.g. preferences, logistics) within one itinerary; respect causal dependencies where earlier itinerary choices constrain which later activities remain feasible; produce a plan validated via realistic agent-based simulation rather than isolated constraint checks The planner must track all outstanding multifaceted constraints and the downstream causal consequences of already-committed itinerary decisions as the simulated trip unfolds. DAG-with-precedence yes none stated 50.33/mo
TravelBench
Beyond Itinerary Planning-A Real-World Benchmark for Multi-Turn and Tool-Using Travel Tasks
Planning & travel 2025 solve a travel-planning problem independently using cached tool results; interact with the user across multiple turns to elicit implicit preferences; correctly recognize and communicate the agent's own capability boundaries (Unsolvable subtask) In the Multi-Turn subtask, the agent must track previously elicited or stated user preferences across turns and integrate them with cached tool results from a sandbox of ten travel-related tools, while also recognizing when a request exceeds its capability boundaries. set-of-independent yes none statedother:not-stated 60.67/mo
TripCraft
TripCraft: A Benchmark for Spatio-Temporally Fine Grained Travel Planning
Planning & travel 2025 generate a spatiotemporally coherent 7-day travel itinerary; satisfy meal-scheduling constraints (Temporal Meal Score); satisfy attraction-timing constraints (Temporal Attraction Score); satisfy spatial feasibility across the itinerary (Spatial Score); satisfy activity-ordering constraints (Ordering Score); satisfy user-persona preferences (Persona Score) The agent must track public-transit schedules, event availability, attraction categories, and user-persona preferences simultaneously while assembling a spatially and temporally consistent 7-day itinerary. DAG-with-precedence yes 7simulated-days 311.63/mo
TripScore
TripScore: Benchmarking and rewarding real-world travel planning with fine-grained evaluation
Planning & travel 2025 produce a travel itinerary that jointly satisfies fine-grained feasibility, reliability, and engagement criteria; achieve a high unified reward score combining these criteria for RL training/evaluation; generalize to real-world, free-form travel requests Agent must track and jointly satisfy fine-grained feasibility, reliability, and engagement criteria while assembling a travel plan, since these are unified into a single reward score used for both evaluation and RL training. other (singleton) no none stated 90.82/mo
TripTide
TripTide: A Benchmark for Adaptive Travel Planning under Disruptions
Planning & travel 2025 preserve original itinerary intent (feasibility and goals) after a disruption; respond promptly and appropriately to a disruption event (flight cancellation, weather closure, overbooked attraction); adapt the itinerary with appropriate semantic, spatial, and sequential divergence from the original plan; maintain plan quality across varying disruption severity and traveler tolerance levels The agent must track the original itinerary's intent, spatial layout, and sequential structure, detect and appropriately size its response to a disruption of a given severity and traveler tolerance, and measure how much the revision diverges from the original across semantic, spatial, and sequential dimensions. sequential-chain no none stated 50.45/mo
WandaPlan
Is Your LLM-Based Multi-Agent a Reliable Real-World Planner? Exploring Fraud Detection in Travel Planning
Planning & travel+ Multi-agent orgs 2025 produce a correct multi-agent travel plan while resisting injected deceptive/fraudulent content; detect/avoid Misinformation Fraud; detect/avoid Team-Coordinated Multi-Person Fraud; detect/avoid Level-Escalating Multi-Round Fraud The planning system must track which review/social-media sourced information it has incorporated, cross-check it for authenticity across rounds and multiple purported sources, and avoid building a travel plan on fraudulent inputs even as fraud escalates over multiple rounds. set-of-independent no none stated 191.19/mo
Web & GUI39 artifacts
AgenticShop
AgenticShop: Benchmarking Agentic Product Curation for Personalized Web Shopping
Web & GUI+ Assistants & memory 2026 explore the open web to curate a set of products satisfying diverse shopping scenarios; satisfy each checklist item in a verifiable, checklist-driven personalization rubric aligned to a user profile The agent must track a user's personalization checklist criteria and diverse profile preferences while exploring open-web shopping scenarios, so curated products satisfy every checklist item. set-of-independent yes none stated 101.43/mo
AndroidDaily
AndroidDaily: A Verifiable Benchmark for Mobile GUI Agents on Real-World Closed-Source Applications
Web & GUI 2026 complete each of 350 realistic daily-use tasks spanning 94 real closed-source Android apps; satisfy step-level operational obligations per GRADE's guideline criteria; meet output-quality criteria at each step; avoid violating negative constraints at each step The agent must track progress against multiple step-level guideline criteria (obligations, quality, negative constraints) across a long-horizon, open-ended interaction with a closed-source app exposing no internal state. sequential-chain yes none stated 41.0/mo
AndroTMem-Bench
AndroTMem: From Interaction Trajectories to Anchored Memory in Long-Horizon GUI Agents
Web & GUI+ Assistants & memory 2026 carry forward critical intermediate state across a long sequence of GUI interaction steps to complete a task; correctly resolve strong step-to-step causal dependencies where sparse intermediate states are decisive for later actions; complete each of 1,069 Android GUI tasks (avg. 32.1 steps, max. 65 steps) The agent must carry forward sparse, dependency-critical intermediate state across an average of 32.1 (up to 65) interaction steps per task, since full-sequence replay is redundant/noisy and naive summarization erases exactly the dependency-critical information needed for later steps. DAG-with-precedence no 1,069 tasks with 34,473 total interaction steps; average 32.1 steps per task, maximum 65 steps per taskactions 111.83/mo
CAP
CAP: A Scalable Benchmark for Evaluating Cross-Site Browser Agents with Complex Actions and Perception
Web & GUI 2026 complete a realistic cross-site workflow requiring several specific operations on each of multiple real-world websites; correctly interact with complex, dynamically rendered UI elements (non-trivial UI interactions); correctly perceive/interpret dynamically rendered visual content across the workflow The agent must track which execution and perception checkpoints it has already satisfied across multiple websites within one recomposed cross-site workflow, using a verifiable agent-as-a-judge evaluation framework. DAG-with-precedence yes baselines evaluated with a maximum of 50 reasoning-action steps per task; average of 7 execution points and 4 perception points per taskactions 0
ClawBench
ClawBench: Can AI Agents Complete Everyday Online Tasks?
Web & GUI 2026 complete everyday online tasks (purchases, appointment bookings, job applications) across 144 real platforms; obtain relevant information from user-provided documents; navigate multi-step workflows across diverse platforms; correctly fill in many detailed form fields per task Agent must extract and correctly carry information from user-provided documents through a multi-step, write-heavy workflow with many detailed form fields, operating on live, dynamic production websites. sequential-chain no none stated 224.4/mo
ComboShoppingBench
ComboShoppingBench: Evaluating LLM Agents for Budget-Constrained Basket Shopping with Coupons
Web & GUI 2026 construct a basket of complementary items satisfying compatibility constraints; keep the basket within a stated budget while optimizing coupon use; satisfy store-level requirements, availability, and delivery fees jointly Agent must track the running basket contents, remaining budget, which coupons are valid given the current basket, and per-store requirements while searching for a feasible combination. set-of-independent yes none stated 0
DMV-Bench
DMV-Bench: Diagnosing Long-Horizon Multimodal Agents' Visual Memory with Incidental Cue Injection
Web & GUI+ Assistants & memory 2026 complete chains of autonomous shopping sessions (browsing/selecting home-furnishing products); recall a unique, pre-rendered incidental visual cue seen earlier on a product image when later asked about it Agent must retain visual memory of incidental cues (not just deliberately extracted facts) across chains of shopping sessions of varying length, without being told in advance which details will later be tested. sequential-chain unclear none stated 0
GMA
Benchmarking General Mobile Assistants in Challenging Real-World Scenarios
Web & GUI+ Assistants & memory 2026 complete tasks ranging from atomic actions to complex multi-step workflows across seven open-source-based applications; handle lifestyle-sharing and travel-planning domains at four escalating difficulty tiers; maintain context/state across a workflow via harness-level context retention and explicit state tracking Agent must maintain context retention and explicit state tracking across multi-step workflows spanning seven applications, since performance declines substantially as task complexity/tier increases. hierarchical no none stated 0
GTA
GTA: Generating Long-horizon Tasks for Web Agents at Scale
Web & GUI 2026 complete multi-hop, cross-page web tasks that are compositional over a site graph; follow an intermediate trajectory of dense, process-level supervision steps, not just reach a coarse end goal; generalize across more than 50 websites (e-commerce, government, forums, news) including multilingual tasks The agent must track its position and accumulated information across a multi-hop, cross-page trajectory grounded in a site graph, since tasks provide dense process-level supervision (intermediate trajectory steps) rather than only a coarse start-goal annotation. DAG-with-precedence yes none stated 0
InterruptBench
When Users Change Their Mind: Evaluating Interruptible Agents in Long-Horizon Web Navigation
Web & GUI+ Embodied & robotics 2026 execute a long-horizon, environmentally grounded web-navigation task; adapt when the user adds a new requirement mid-task; revise the current goal when the user changes an existing requirement mid-task; abandon a sub-goal when the user retracts a requirement mid-task; recover efficiently without redundant or incorrect actions after an interruption The agent must track the current, possibly revised, task intent across single- and multi-turn interruption settings plus the environment's persistent state from already-executed actions to adapt or recover correctly. sequential-chain yes none stated 61.2/mo
iOSWorld
iOSWorld: A Benchmark for Personally Intelligent Phone Agents
Web & GUI 2026 complete single-app tasks within one iOS app (27 tasks); complete multi-app task chains spanning 2 to 8 apps (60 tasks); infer personal patterns from persistent user data for memory/personalization tasks (46 tasks) Agent must track a persistent user identity and its connected data (transactions, messages, travel records, social relationships, financial activity) across apps and infer behavioral patterns for personalization tasks. DAG-with-precedence no none stated 31.0/mo
LongMemEval-V2 (LME-V2)
LongMemEval-V2: Evaluating Long-Term Agent Memory Toward Experienced Colleagues
Web & GUI+ Assistants & memory 2026 recall static state facts about the environment (static state recall); track dynamic state changes over time (dynamic state tracking); recall workflow knowledge/procedures learned from experience; recall environment-specific 'gotchas'/recurring failure modes; maintain premise awareness of what has already been established or assumed The memory system must consume and internalize up to 500 history trajectories (115M tokens) and return compact, correct evidence for downstream question answering across five distinct memory-ability categories. set-of-independent yes up to 500 trajectories and 115M tokens of historyepisodes 92.25/mo
Memory-World
STAMP: Training Explicit Memory for Mobile GUI Agents in Controllable and Scalable Virtual Environments
Web & GUI+ Assistants & memory 2026 encode a programmatically-injected memory variable at the correct point in a task; retain that memory variable correctly despite progressive discarding of older visual history; retrieve and correctly apply the memorized variable later in the same long-horizon mobile GUI task The agent must explicitly memorize deterministic variables injected at specific points and correctly retrieve/apply them later in the same task, despite token-heavy screenshots forcing progressive discarding of older visual history. sequential-chain no none statedother:not-stated 10.25/mo
MobiFlow
MobiFlow: Real-World Mobile Agent Benchmarking through Trajectory Fusion
Web & GUI 2026 complete real-world tasks within arbitrary third-party mobile apps whose success cannot be checked via system-level APIs; have task completion verified via a graph constructed from fusing multiple real user trajectories The agent must track its current position within the app's fused trajectory-state graph and correctly follow one of the valid paths to task completion, since no system-level completion signal is available. DAG-with-precedence yes none stated 0
Odysseys
Odysseys: Benchmarking Web Agents on Realistic Long Horizon Tasks
Web & GUI 2026 complete long-horizon, multi-site web workflows (e.g., comparing products across domains); plan trips across multiple web services; summarize information gathered from multiple search queries; satisfy an average of 6.1 graded rubric criteria per task The agent must sustain context and accumulated findings across multiple websites and search queries over potentially hours of browsing, and satisfy each of ~6.1 rubric criteria graded per task rather than a single pass/fail check. DAG-with-precedence yes potentially hours (of browsing); Trajectory Efficiency measured as rubric score per stepwall-clock-hours 132.6/mo
PAHF benchmarks (embodied manipulation + online shopping)
Learning Personalized Agents from Human Feedback
Web & GUI+ Embodied & robotics 2026 learn a new user's initial preferences from scratch via pre-action clarification; ground actions in preferences retrieved from an explicit per-user memory; adapt rapidly to persona shifts using post-action feedback to update memory Agent must maintain explicit per-user memory across a four-phase protocol, updating stored preferences via dual feedback channels (pre-action clarification, post-action feedback) as it learns initial preferences from scratch and later adapts to persona shifts. sequential-chain no none stated 152.14/mo
ParaGUIBench
Beyond Sequential Interaction: Benchmarking Parallel Execution and Coordination for GUI Agents
Web & GUI 2026 identify which GUI sub-tasks can run concurrently despite unstated dependencies; avoid conflicts between concurrent workers modifying shared artifacts; ensure each worker's locally-completed sub-task composes into a globally correct combined result; complete each of 233 tasks spanning six task categories on separate desktop instances The planner-worker system must track which sub-tasks are dependency-free and safe to parallelize, coordinate concurrent workers' access to shared artifacts, and verify that the union of workers' outputs satisfies the original combined instruction. DAG-with-precedence no none statedother:not-stated 0
PhoneWorld
PhoneBuddy: Training Open Models for Agentic Phone Use
Web & GUI 2026 complete single-app phone tasks; complete mini-app tasks; complete cross-app workflows requiring coordination across multiple mobile applications; achieve task success on a 150-task real-phone human evaluation and on AndroidWorld The agent must track UI/app state across a real, stateful phone environment (or the resettable PhoneWorld mock-app equivalent), including state that must persist and transfer correctly across multiple apps in cross-app workflows. set-of-independent yes none stated 0
ScaleWoB
ScaleWoB: Guiding GUI Agents with Coding Agents via Large-Scale Environmental Synthesis
Web & GUI 2026 complete verifiable multi-step GUI tasks across mobile, desktop, or automotive/in-vehicle synthesized environments; succeed specifically on a distinguished long-horizon subset of tasks, which requires sustaining performance across more steps than the general task pool Agent must track GUI state across a chain of interface actions long enough to be classified into the 'long-horizon subset', where performance drops sharply relative to the general task pool, implying more state must be carried across more steps. hierarchical yes none stated 10.25/mo
SentinelBench
SentinelBench: A Benchmark for Long-Running Monitoring Agents
Web & GUI 2026 continuously monitor a live web environment (email/calendar/finance/professional-networking/entertainment) for a scripted external event; recognize the moment an event makes progress possible and act promptly; avoid excessive/wasteful actions (continuous polling/refreshing) while waiting The agent must track whether the monitored page state has changed to reflect the awaited event, its own resource expenditure (tool calls/tokens) over the monitoring window, and elapsed time relative to the task's (possibly stretched) time budget. sequential-chain yes each of the 100 tasks is designed to be achievable within a default 10-minute window; a speed_factor parameter (default 1.0) can stretch tasks to much longer durations (e.g. at speed_factor 0.25, tasks may require as long as 40 minutes)wall-clock-minutes 20.67/mo
unnamed mobile GUI benchmark (paper introduces the ATMem method)
What Memory Do GUI Agents Really Need? From Passive Records to Active Task-Driving States
Web & GUI+ Assistants & memory 2026 act on every list entry that satisfies the given instruction (positive matches); reject/skip every list entry that violates the instruction's constraints (negative matches) The agent must track which near-identical list entries it has already acted on vs. still pending, and correctly apply the instruction's inclusion/exclusion constraints to each entry across a long trajectory. set-of-independent yes none stated 31.0/mo
AndroidLens
AndroidLens: Long-latency Evaluation with Nested Sub-targets for Android GUI Agents
Web & GUI 2025 complete nested sub-targets within a single long-latency mobile task; satisfy multi-constraint, multi-goal task requirements drawn from 38 real-world domains; make measurable milestone-level progress even when full task success is not reached Agent must track which nested sub-targets have been completed, tolerate environmental anomalies, and retain long-term memory of earlier steps across an average of more than 26 steps per task. hierarchical yes >26 (average, 'more than 26')agent-steps 20.22/mo
AndroidLH (via Mirage-1)
Mirage-1: Augmenting and Updating GUI Agent with Hierarchical Multimodal Skills
Web & GUI 2025 complete real-world long-horizon, multi-app Android task scenarios using previously acquired hierarchical skills; correctly apply execution skills, core skills, and meta-skills at the appropriate level of abstraction across a long-horizon task The agent must track which level of its hierarchical skill structure (execution, core, meta) is currently relevant, maintain state as it moves across multiple apps in a long-horizon scenario, and bridge the offline-to-online domain gap without losing track of previously acquired skills. hierarchical yes none stated 130.87/mo
BELA (Benchmark for Experiential Learning and Active exploration)
Benchmarking In-context Experiential Learning Through Repeated Product Recommendations
Web & GUI+ Assistants & memory 2025 elicit unknown customer preferences through questions within a single recommendation interaction (episode); tailor questioning/recommendation strategy based on patterns observed across multiple prior episodes (customers/products) The agent must track what it has learned about the current customer's preferences within an episode (turn-by-turn), and must also track/aggregate patterns across multiple prior episodes to adapt its questioning/recommendation strategy over time. sequential-chain yes none stated 20.2/mo
CAPBench
MobileSteward: Integrating Multiple App-Oriented Agents with Self-Evolution to Automate Cross-App Instructions
Web & GUI 2025 associate and sequence sub-tasks across multiple mobile apps per a cross-app instruction; assign each associated sub-task to the correct app-oriented StaffAgent; avoid error propagation and information loss across the multi-step, multi-app execution; complete each of the 500 cross-app instructions spanning 14 apps in 6 categories The centralized StewardAgent must track the scheduling graph of inter-app task associations, information flow between StaffAgents, and self-evolving memory of past executions to avoid repeating errors across a cross-app instruction. DAG-with-precedence no none statedother:not-stated 241.26/mo
ColorBench
ColorBench: Benchmarking Mobile Agents with Graph-Structured Framework for Complex Long-Horizon Tasks
Web & GUI 2025 complete a single-app mobile task via one of multiple valid GUI action paths; complete a cross-app mobile task requiring coordination across multiple applications; reach subtask-level completion milestones within a longer task Agent must track which subtasks have been completed, which of several valid paths it is following through the task's state graph, and avoid known error paths, across an average of more than 13 steps. DAG-with-precedence yes >13 (average, 'over 13 steps')agent-steps 80.73/mo
DeepShop
DeepShop: A Benchmark for Deep Research Shopping Agents
Web & GUI 2025 satisfy multiple product-attribute constraints in a single shopping query; apply the correct search filters specified or implied by the query; apply the correct sorting preference specified or implied by the query; achieve overall shopping-task success across easy/medium/hard complexity tiers Agent must track which product attributes, filters, and sorting preferences the query requires, and verify each is correctly reflected in its final shopping actions/results. set-of-independent yes none stated 422.8/mo
Mobile-Eval-RAG
Mobile-Agent-RAG: Driving Smart Multi-Agent Coordination with Contextual Knowledge Empowerment for Long-Horizon Mobile Automation
Web & GUI+ Multi-agent orgs 2025 complete a cross-application mobile-automation task requiring both high-level plan steps and precise low-level UI operations; correctly execute app-specific atomic UI actions aligned with the current subtask Agent must track which app/subtask it is currently operating in, the current step of the high-level plan, and precise UI state needed for accurate atomic actions across multiple apps in one task. hierarchical yes none stated 20.2/mo
MobileWorld
MobileWorld: Benchmarking Autonomous Mobile Agents in Agent-User Interactive and MCP-Augmented Environments
Web & GUI 2025 complete long-horizon, cross-application mobile workflows spanning up to 20 applications; handle vague user instructions and hybrid tool usage; coordinate agent-user interaction and MCP-augmented tool calls mid-task Agent must track cross-application state across nearly twice as many completion steps on average (27.8 vs. 14.3) as AndroidWorld, while also handling user-interaction requests and MCP-tool-call state. sequential-chain no 27.8 (vs. 14.3 in AndroidWorld)agent-steps 586.44/mo
MVISU-Bench
MVISU-Bench: Benchmarking Mobile Agents for Real-World Tasks by Multi-App, Vague, Interactive, Single-App and Unethical Instructions
Web & GUI 2025 complete multi-app instructions requiring cross-app subgoal coordination; clarify vague/underspecified user instructions before acting; handle interactive instructions requiring mid-task clarification; complete single-app instructions; recognize and appropriately refuse or handle unethical instructions The agent must track task state across multiple mobile apps for Multi-App instructions, detect ambiguity requiring clarification for Vague/Interactive instructions, and recognize when an instruction should be refused for Unethical instructions. other (singleton) no none stated 70.54/mo
NaturalGAIA
NaturalGAIA: A Verifiable Benchmark and Hierarchical Framework for Long-Horizon GUI Tasks
Web & GUI 2025 decompose a natural, non-linear human GUI intent into a structured Task Topology of atomic sub-tasks; dynamically schedule sub-tasks across heterogeneous agents via context evolution; execute each atomic sub-task with precision via hybrid visual-structural perception; achieve a high Weighted Pathway Success Rate across the full causal pathway The manager must dynamically track the Task Topology of atomic sub-tasks, schedule them across heterogeneous agents, and evolve shared context to bridge information gaps between dependent steps, assessed via a hierarchical success/error-attribution framework. DAG-with-precedence no none statedother:not-stated 0
RealWebAssist
RealWebAssist: A Benchmark for Long-Horizon Web Assistance with Real-World Users
Web & GUI 2025 correctly follow each instruction in a sequence of real, sequentially-issued user instructions across multiple websites; reason about the true intent behind ambiguous instructions; keep track of the user's mental state and user-specific routines as instructions evolve over the session; ground each intended task to the correct GUI element on the current website Agent must retain a running model of the user's intent, mental state, and personal routines across a long sequence of instructions issued over multiple websites, using this history to correctly interpret each new (sometimes ambiguous) instruction. sequential-chain unclear none stated 261.53/mo
ShoppingBench
ShoppingBench: A Real-World Intent-Grounded Shopping Benchmark for LLM-based Agents
Web & GUI 2025 apply vouchers correctly to a purchase; manage a budget across a shopping session; find and select from multi-product sellers matching a grounded intent; satisfy increasingly challenging levels of grounded shopping intent end-to-end Agent must track running spend/budget, applied vouchers, and multi-seller product-matching requirements across a session, operating within a sandbox of over 2.5 million real-world products. other (singleton) no none stated 272.08/mo
ShoppingComp
ShoppingComp: Are LLMs Really Ready for Your Shopping Cart?
Web & GUI 2025 retrieve products satisfying many simultaneous discovery constraints; generate an expert-level report on the retrieved products; make a safety-critical purchase decision (e.g. flag unsafe product usage) The agent must track which of the multiple product-discovery constraints have been satisfied, the evidence supporting its report, and any identified safety hazards, while operating in an open-world product catalogue. set-of-independent no none statedother:not-stated 80.8/mo
UI-NEXUS
Atomic-to-Compositional Generalization for Mobile Agents with A New Benchmark and Scheduling System
Web & GUI 2025 complete compositional mobile operations that concatenate multiple atomic tasks (Simple Concatenation); transition context correctly across sub-tasks within one compositional task (Context Transition); perform a deep multi-step drill-down within an app or workflow (Deep Dive) The agent must track which atomic subtasks within a compositional task it has already completed, the context state carried between apps/screens, and overall progress toward each of the three compositional operation types. hierarchical yes average optimal step count of 14.05 (over 100 interactive task templates)actions 120.8/mo
VeriWeb
VeriWeb: Verifiable Long-Chain Web Benchmark for Agentic Information-Seeking
Web & GUI 2025 ensure comprehensive information coverage across breadth- and depth-oriented multi-hop web searches; complete a sequence of interdependent, individually-verifiable subtasks within one long-chain web task; maintain consistent context tracking across a long information-seeking chain Agent must track and verify each subtask-level answer as it progresses through a long chain of interdependent web-search subtasks, ensuring comprehensive information coverage and consistent context tracking across the chain. sequential-chain yes none stated 100.77/mo
WebChoreArena
WebChoreArena: Evaluating Web Browsing Agents on Realistic Tedious Web Tasks
Web & GUI+ Information seeking 2025 retain and retrieve large amounts of information gathered from many observations (Massive Memory); perform precise mathematical/quantitative reasoning over collected information (Calculation); keep information consistent while tracking it across multiple webpages over the course of a task (Long-Term Memory) Agent must accumulate and correctly recall large amounts of information across many webpages, perform arithmetic over it, and keep facts consistent across the whole task rather than a single page view. sequential-chain unclear none stated 281.87/mo
WebMall
WebMall - A Multi-Shop Benchmark for Evaluating Web Agents
Web & GUI 2025 find a specific product across four simulated online shops; perform price comparisons across shops; identify suitable substitutes or compatible products (advanced search); add items to a cart and complete checkout in the correct shop(s) Agent must retain product/price information gathered from each of the four shops it has visited so far in order to correctly compare, substitute, or complete a checkout later in the same task. sequential-chain unclear none stated 141.08/mo
WebPRM Collection / WebRewardBench
Web-Shepherd: Advancing PRMs for Reinforcing Web Agents
Web & GUI 2025 navigate a website across many sequential steps toward a stated goal; at each step, satisfy sub-goals encoded in an annotated checklist (process-level correctness); complete the overall web-navigation task successfully (episode-level correctness) The reward model must track which checklist sub-goals have already been satisfied at each step of a trajectory so it can assess the process-level (not just outcome-level) quality of a web-navigation trajectory. sequential-chain yes trajectory lengths vary by difficulty: median ~5 steps (easy), ~9 steps (medium), ~20 steps (hard), with some trajectories exceeding 40 stepsactions 291.81/mo

Built from 4,327 candidates, 2,191 of them judged, then extracted per paper. Not exhaustive: at least ~131 further in-scope artifacts are estimated missing (a lower bound; the honest range is ~130–320). You should not read a gap here as evidence that no such benchmark exists. The horizon column reflects what each paper foregrounds — 276 of the 334 “none stated” determinations were made from the abstract, 56 confirmed against full text. Full data: catalog.csv · catalog.json · the full report.