Less Context, Better Agents: Efficient Context Engineering for Long-Horizon Tool-Using LLM Agents
Goals to track itemize every line item of a hotel expense report completely and accurately using enterprise MCP tools; avoid context overflow / stale-state errors while repeatedly calling verbose enterprise tool APIs across the itemization workflow
Must keep track of The agent must track already-itemized line items, running token/context budget, and must avoid stale or overflowing tool-call history while repeatedly invoking enterprise MCP tools to reach complete, accurate itemization.
structure: set-of-independent · goals arrive: given-up-front · count: 50 tasks; five expense types grouped int · subgoal credit: yes · 1 citations · unit: wall-clock-hours
Horizonfull-context retention: 1,480,996 tokens and 14.56 hours per benchmark (50 tasks x 5 runs); pruning to last 5 tool calls: 535,274 tokens and 5.39 hours; pruning + summari
the paper’s own words
We evaluate four GPT-5 configurations on a 50-task hotel expense benchmark: no user model, full conversation history, context pruned to the last 5 tool call/response pairs, and pruning with automated summarization... The no-user-model baseline achieves only 8.0% complete itemization. Full-context retention improves completion to 71.0%... Adding summarization achieves the best result: 91.6% complete itemization and 99.64% average amount itemized.
Full-context retention improves completion to 71.0%, but consumes 1,480,996 tokens and 14.56 hours per benchmark. Pruning to the last 5 tool calls improves completion to 79.0% while reducing token use to 535,274 and runtime to 5.39 hours. Adding summarization achieves the best result: 91.6% complete itemization and 99.64% average amount itemized, with 553,374 tokens and 5.79 hours.
AD-Bench: A Real-World, Trajectory-Aware Advertising Analytics Benchmark for LLM Agents
Goals to track answer a real user marketing-analysis request via multi-round, multi-tool collaboration; maintain trajectory coverage consistent with an expert tool-call trajectory; produce answers that remain correct/consistent with a continuously evolving production advertising platform (dynamic ground truth); succeed across three stratified difficulty levels (L1-L3) requiring increasing multi-round, multi-tool collaboration
Must keep track of The agent must track which professional tools it has called and in what order across multiple rounds, since evaluation jointly measures end-to-end answer correctness (Pass@k) and how well its trajectory covers the expert reference trajectory, especially as this compounds at higher difficulty levels.
structure: sequential-chain · goals arrive: given-up-front · count: three difficulty levels (L1-L3) stratify · subgoal credit: yes · 3 citations
Horizonthe paper states none
the paper’s own words
a trajectory-aware evaluation that jointly measures end-to-end answer correctness (Pass@k) and trajectory coverage. Requests are stratified into three difficulty levels (L1-L3) to probe multi-round, multi-tool collaboration.
APEX-Agents
Goals to track execute a long-horizon, cross-application task created by investment-banking analysts, management consultants, or corporate lawyers; navigate realistic work environments containing files and tools to produce the required deliverable
Must keep track of Agent must track which files and tool states it has already produced/consulted while navigating a cross-application work environment, using rubrics and gold outputs to determine success (Pass@1).
structure: sequential-chain · goals arrive: given-up-front · count: 480 tasks (n=480) · subgoal credit: unclear · 20 citations
Horizonthe paper states none
the paper’s own words
APEX-Agents requires agents to navigate realistic work environments with files and tools. We test eight agents for the leaderboard using Pass@1.
APEX-Agents requires agents to execute long-horizon, cross-application tasks created by investment banking analysts, management consultants, and corporate lawyers
AeroCopilotBench: A Two-Tier Benchmark for Evaluating LLM Agents as Aviation Copilots in an Interactive Virtual Cockpit Environment
Goals to track diagnose faults and operate aircraft systems through standardized tool interfaces during emergency/abnormal procedures; achieve all task goal-conditions for each of 73 Tier-2 emergency/abnormal tasks; avoid violating any hard safety constraint while progressing toward goals
Must keep track of Agent must interpret cockpit state, diagnose faults, and operate aircraft systems while simultaneously tracking goal-condition progress and hard safety-constraint compliance across the executable procedure.
structure: DAG-with-precedence · goals arrive: given-up-front · count: 73 emergency/abnormal Tier-2 tasks (plus · subgoal credit: yes · 0 citations
Horizonthe paper states none
the paper’s own words
We establish a safety-gated evaluation framework in which a trajectory succeeds only when all task goals are achieved without violating any hard safety constraint, while safe goal progress and trajectory safety are measured separately.
Tier-2 comprises 73 emergency and abnormal tasks derived from the manufacturers'Pilot's Operating Handbooks (POHs) and instantiated in ACOE.
Agentic ERP: Multi-Agent Large Language Model Architecture for Autonomous Enterprise Resource Planning
Goals to track execute end-to-end business workflows across ERP functional boundaries (e.g. procurement, inventory, crisis response); resolve cross-functional crisis tasks via a Planner-Executor-Reflector-Responder orchestration; sustain a full simulated year of ERP operation while avoiding stockouts, compared against rule-based RPA and no-intervention baselines
Must keep track of The agent(s) must track the structured enterprise state (inventory, demand stream, crisis conditions) across a full simulated year, coordinate across role-aligned agents via externalised grading criteria and sprint contracts, and avoid accumulating stockouts as the rule-based baseline does.
structure: hierarchical · goals arrive: emitted-by-environment-over-time · count: a scenario-based task suite; a compariso · subgoal credit: yes · 0 citations · also Multi-agent organisations & societies · unit: simulated-days
Horizona 365-day agent-in-the-loop simulation (a simulated year of ERP operation)
the paper’s own words
the system is evaluated at three levels: a scenario-based task suite, a comprehensive comparison of six orchestration paradigms on cross-functional crisis tasks, and a 365-day agent-in-the-loop simulation against rule-based RPA and no-intervention baselines. Across these levels the proposed multi-agent method is significantly better than the baseline, and the system sustains a simulated year of operation with zero stockouts while the rule-based baseline accumulates hundreds under the same demand stream.
AgenticPay: A Multi-Agent LLM Negotiation System for Buyer-Seller Transactions
Goals to track negotiate to reach agreement given private constraints and product-dependent valuations; maximize feasibility, efficiency, and welfare across a negotiation; successfully complete each of 110+ tasks ranging from bilateral bargaining to many-to-many markets
Must keep track of Agents must track their own private constraints/valuations, the evolving history of natural-language offers and counteroffers across negotiation rounds, and, in many-to-many markets, the state of other concurrent negotiations.
structure: sequential-chain · goals arrive: given-up-front · count: over 110 tasks ranging from bilateral ba · subgoal credit: yes · 13 citations · also Multi-agent organisations & societies
Horizonthe paper states none
the paper’s own words
AgenticPay models markets in which buyers and sellers possess private constraints and product-dependent valuations, and must reach agreements through multi-round linguistic negotiation rather than numeric bidding alone. The framework supports a diverse suite of over 110 tasks ranging from bilateral bargaining to many-to-many markets, with structured action extraction and metrics for feasibility, efficiency, and welfare.
AgenticVBench: Can AI Agents Complete Real-World Post-Production Tasks?
Goals to track complete each of 100 agentic post-production tasks across 4 task families; compose capabilities across text, image, audio, and video understanding within one task; plan and execute long-horizon multi-step production workflows using appropriate tools
Must keep track of The agent must track progress through a long-horizon, multi-modal production workflow and correctly sequence tool use, since tasks are constructed from real production workflows contributed by industry experts and evaluated jointly by programmatic verifiers and expert rubrics.
structure: other:four-task-families-each-with-real-produc · goals arrive: given-up-front · count: 100 agentic tasks across 4 task families · subgoal credit: yes · 3 citations
Horizonthe paper states none
the paper’s own words
Tasks are paired with evaluation specifications that combine programmatic verifiers and expert rubrics.
they require composite capabilities across text, image, audio, and video understanding, along with long-horizon planning, and tool use
Agents' Last Exam
Goals to track complete a long-horizon, economically valuable, real-world professional task with a verifiable outcome, drawn from a specific occupational sub-field; achieve sustained performance across the task rather than a single-shot correct answer, particularly on the hardest ('last-exam') tier
Must keep track of The agent must sustain performance on long-horizon, economically valuable real-world tasks whose outcomes are verifiable, across an occupational taxonomy spanning 55 sub-fields and 13 industry clusters, with the hardest tier proving far from saturated (average full pass rate below 1%).
structure: set-of-independent · goals arrive: given-up-front · count: 1K+ tasks organized into 55 sub-fields g · subgoal credit: yes · 12 citations
Horizonthe paper states none
the paper’s own words
This paper introduces Agents' Last Exam (ALE), a benchmark designed to evaluate AI agents on long horizon, economically valuable, real world tasks with verifiable outcomes. Developed in collaboration with 250+ industry experts, ALE covers non-physical industries defined with reference to O*NET / SOC 2018 (the U.S. federal occupational taxonomy). It is organized around a task taxonomy with 55 sub fields grouped into 13 industry clusters covering 1K+ tasks. Current results show that the hardest tier remains far from saturated: across mainstream harness and backbone configurations, the average full pass rate is below 1%.
AutomationBench
Goals to track discover the relevant REST API endpoints needed for a cross-application workflow (CRM, inbox, calendar, messaging, etc.); follow layered business-policy rules while writing data; navigate environments containing irrelevant or misleading records without being derailed; get correct data into the right systems by the end of the workflow (end-state correctness)
Must keep track of Agent must track which endpoints it has discovered across multiple applications, which business-policy rules apply to each write, and whether records encountered are relevant or misleading, to ensure the correct final data lands in each system.
structure: DAG-with-precedence · goals arrive: implied-by-constraints · subgoal credit: yes · 1 citations
Horizonthe paper states none
the paper’s own words
Grading is programmatic and end-state only: whether the correct data ended up in the right systems.
tasks span Sales, Marketing, Operations, Support, Finance, and HR domains
BigFinanceBench: A Workflow-Grounded Benchmark for Financial-Research Agents
Goals to track produce an auditable financial-research derivation (source choice, period/accounting definition, assumptions, calculation), not just a final answer; satisfy each of many independently checkable rubric points per item (36,241 points across 928 items); correctly complete each of 928 expert-authored open-ended financial-research tasks
Must keep track of The agent must track and justify each auditable component of its derivation (data source, period/accounting definition, assumptions, calculation steps) so that an independent rubric can check each step, not just the final numeric answer.
structure: hierarchical · goals arrive: given-up-front · count: 928 items; 36,241 rubric points total (a · subgoal credit: yes · 5 citations
Horizonthe paper states none
the paper’s own words
We introduce BigFinanceBench, a 928-item expert-authored benchmark of open-ended financial-research tasks in which each item pairs a ground-truth reference answer with a point-weighted rubric that decomposes the derivation into independently checkable steps... Across 36,241 rubric points, the benchmark supports partial-credit evaluation and localization of failures across the analyst workflow.
which source was chosen, which period and accounting definition were used, which assumptions were made, and how the calculation was performed
BlueFin: Benchmarking LLM Agents on Financial Spreadsheets
Goals to track synthesize new spreadsheet content/formulas from source data; manipulate existing financial spreadsheet workbooks (edits, restructuring); comprehend and answer questions about spreadsheet content; satisfy each of up to 3,225 granular rubric criteria across 131 tasks
Must keep track of The agent must track and satisfy many granular, LM-judge-scored rubric criteria across synthesis, manipulation, and comprehension actions within one workbook, maintaining correctness as the workbook's formulas/state evolve ('dynamic correctness').
structure: hierarchical · goals arrive: given-up-front · count: 131 tasks containing 3,225 granular rubr · subgoal credit: yes · 1 citations · unit: human-expert-hours
Horizontasks require at least 45+ minutes of work for a human analyst to complete from scratch
the paper’s own words
we curate a set of 131 challenging, complex tasks with real-world relevance in the domain, containing 3,225 granular rubric criteria; notably, our rubric criteria and LM judge evaluations are validated by a team of expert human annotators.
at least 45+ minutes of work for an analyst to complete from scratch
Business Arena: Benchmarking LLM Agents in a Realistic Marketplace
Goals to track infer business opportunities from partial market signals; commit capital under uncertainty when buying from suppliers; set/adjust pricing to sell to buyers profitably; satisfy regulatory obligations before trading legally; sustain profitable operation of a cross-border shop over a long horizon
Must keep track of The agent must track capital, delayed and coupled consequences of sourcing/pricing/recovery decisions, market conditions calibrated from real data, and regulatory obligations across a long-horizon cross-border trading operation.
structure: open-ended · goals arrive: given-up-front · subgoal credit: no · 0 citations
Horizonthe paper states none
the paper’s own words
We introduce Business Arena, a controlled environment where an AI agent runs a cross-border shop, buying from suppliers and selling to buyers over a long horizon. ... action-level attribution identifies the sourcing, pricing, and recovery decisions that create or destroy value.
CEO-Bench: Can Agents Play the Long Game?
Goals to track operate a startup for 500 days, managing pricing, marketing, budgeting, and other business aspects; acquire information from noisy, interconnected business databases and translate it into strategy; adapt strategy to a changing world/market over the run; orchestrate many interdependent business decisions toward a coherent long-term financial goal
Must keep track of The agent must track noisy, interconnected business databases (churn regimes, billing timing, customer losses, projected cash) and coordinate many interdependent decisions via code across a 500-day simulated run.
structure: open-ended · goals arrive: given-up-front · subgoal credit: yes · 3 citations · also Games & interactive fiction · unit: simulated-days
Horizon500
the paper’s own words
We introduce CEO-Bench, which evaluates these capabilities together by simulating a representative real-world task: operating a startup for 500 days. An agent manages pricing, marketing, budgeting, and many other aspects of a fictional company through a programmable Python interface.
Can LLMs Be CEOs? Benchmarking Strategic Resource Reallocation with Multi-Role Agent Simulation
Goals to track integrate conflicting recommendations from four role-conditioned C-suite advisors (CFO, CTO, COO, CMO) into one allocation plan; redirect capital across business units under information asymmetry and organizational constraints; synthesize advice consistently across multiple rounds with temporal dependencies
Must keep track of Agent (as CEO) must track and reconcile conflicting, privately-signaled recommendations from four C-suite advisors across a multi-round resource-reallocation process, remembering prior decisions for history-sensitive judgment.
structure: DAG-with-precedence · goals arrive: mixed:given-up-front-scenario-with-conflicting · count: 13 scenarios; four role-conditioned advi · subgoal credit: yes · 1 citations
Horizonthe paper states none
the paper’s own words
LLM agents receive conflicting advice from four role-conditioned C-suite advisors (CFO, CTO, COO, CMO), each with private signals and distinct priorities, and must synthesize these into a concrete allocation plan evaluated along four dimensions: role integration, conditional boldness, history-sensitive judgment, and plan validity.
the process of redirecting capital across business units in a multi-round, constraint-rich organizational environment... Experiments across five frontier models on 13 scenarios reveal that all models achieve high structural validity but diverge sharply on strategic calibration
CHI-Bench: Can AI Agents Automate End-to-End, Long-Horizon, Policy-Rich Healthcare Workflows?
Goals to track ground each decision in a large library of medical, insurance, and operational policy rules; play multiple roles within a single task, handing off between them; conduct multilateral multi-turn dialogs (e.g. peer-to-peer review, patient outreach) as intermediate workflow steps; drive a clinical case to a terminal status via tool calls and role-artifact writing
Must keep track of The agent must track its current role and handoffs to other roles, policy compliance against a 1,290+ document handbook, and the clinical case's evolving status across a high-fidelity simulator of 20 apps, until it reaches a terminal status.
structure: DAG-with-precedence · goals arrive: given-up-front · count: long-horizon workflows across 3 domains, · subgoal credit: no · 6 citations · also Healthcare & clinical · unit: other:not-stated
Horizonthe paper states none
the paper’s own words
we introduce $\chi$-Bench, a benchmark of long-horizon healthcare workflows across three domains: provider prior authorization, payer utilization management, and care management. Each task hands the agent a clinical case in a high-fidelity simulator of 20 healthcare apps exposed via 87 MCP tools, which it must drive to a terminal status through tool calls and writing the role's artifacts, guided by a 1,290+ document managed-care operations handbook skill.
executing all tasks in a single session slumps the performance to 3.8%
Claw-Eval-Live: A Live Agent Benchmark for Evolving Real-World Workflows
Goals to track complete end-to-end units of work across software tools, business services, and local workspaces per release; pass controlled tasks with fixed fixtures/services/workspaces/graders reflecting current public workflow-demand signals; handle HR, management, and multi-system business workflows as well as local workspace repair
Must keep track of Agent must produce verifiable execution traces, audit logs, and consistent service/workspace state across each end-to-end task, since grading uses deterministic checks when evidence is sufficient and structured LLM judging only for semantic dimensions.
structure: set-of-independent · goals arrive: given-up-front · count: 105 tasks per release (ClawHub Top-500 s · subgoal credit: no · 13 citations
Horizonthe paper states none
the paper’s own words
For grading, Claw-Eval-Live records execution traces, audit logs, service state, and post-run workspace artifacts, using deterministic checks when evidence is sufficient and structured LLM judging only for semantic dimensions. The release contains 105 tasks spanning controlled business services and local workspace repair, and evaluates 13 frontier models under a shared public pass rule.
ClawMark: A Living-World Benchmark for Multi-Turn, Multi-Day, Multimodal Coworker Agents
Goals to track carry a professional coworker task forward correctly across multiple working days; detect and adapt to exogenous environment updates (new emails, calendar shifts, KB edits) that occur between turns; satisfy each of the task's deterministic Python checkers over post-execution service state (mean 15.4 checkers/task); coordinate consistent state across five stateful sandboxed services (filesystem, email, calendar,
Must keep track of The agent must track evolving state across five sandboxed services (filesystem, email, calendar, knowledge base, spreadsheet) over multiple working days, detecting and incorporating exogenous updates injected between turns, verified by up to 29 deterministic checkers per task.
structure: sequential-chain · goals arrive: emitted-by-environment-over-time · count: 100 tasks across 13 professional scenari · subgoal credit: yes · 15 citations · unit: simulated-days
Horizon2-6 turns per task (mean 3.6); one turn = one in-universe working day
the paper’s own words
The current release contains 100 tasks across 13 professional scenarios, executed against five stateful sandboxed services (filesystem, email, calendar, knowledge base, spreadsheet) and scored by 1537 deterministic Python checkers over post-execution service state; no LLM-as-judge is invoked during scoring. ... The strongest model reaches 75.8 weighted score, but the best strict Task Success is only 20.0%
Tasks range from two to six turns (mean 3.6) and 6 to 29 checkers (mean 15.4).
ClawsBench: Evaluating Capability and Safety of LLM Productivity Agents in Simulated Workspaces
Goals to track complete single-service productivity tasks (e.g. Gmail, Slack, Calendar, Docs, Drive) correctly and safely; complete cross-service workflows spanning multiple mock services while maintaining consistent state; avoid unsafe/irreversible actions in safety-critical scenarios
Must keep track of Agents must track persistent state across five mock services (Gmail, Slack, Calendar, Docs, Drive), recognize safety-critical constraints to avoid irreversible/unsafe actions, and coordinate behavior across services via a meta-prompt layer.
structure: DAG-with-precedence · goals arrive: given-up-front · count: 44 structured tasks spanning single-serv · subgoal credit: yes · 25 citations
Horizonthe paper states none
the paper’s own words
It includes five high-fidelity mock services (Gmail, Slack, Google Calendar, Google Docs, Google Drive) with full state management and deterministic snapshot/restore, along with 44 structured tasks covering single-service, cross-service, and safety-critical scenarios.
Experiments across 6 models, 4 agent harnesses, and 33 conditions show that with full scaffolding, agents achieve task success rates of 39-64% but exhibit unsafe action rates of 7-33%.
CoffeeBench: Benchmarking Long-Horizon LLM Agents in Heterogeneous Multi-Agent Economies
Goals to track maximize cumulative net income for one's own firm (coffee roaster) over a long simulated run; manage cash, inventory, and pricing on an ongoing basis; communicate and transact with other heterogeneous firms (farmers, roasters, retailers) to secure supply/demand
Must keep track of The agent must track its own cash, inventory, and pricing state day by day over the 90-day simulation, plus the state of its ongoing communications/transactions with other firms, to maximize cumulative net income rather than any single day's outcome.
structure: sequential-chain · goals arrive: given-up-front · subgoal credit: yes · 6 citations · also Multi-agent organisations & societies · unit: simulated-days
Horizon90-day simulation
the paper’s own words
two farmers, two roasters, and two retailers autonomously operate their businesses over a 90-day simulation, each seeking to maximize cumulative net income through communication and transactions while managing cash, inventory, and pricing. The evaluated model controls one coffee roaster, while the remaining firms are controlled by fixed reference agents.
Do AI-Native Biotechs Need Departments? Benchmarking Company World Models for AI-Driven Drug Development
Goals to track maintain and update a persistent asset-to-value state (Live Asset Value Record) across scientific, regulatory, BD, commercial, financial, and execution constraints; resolve 45 retrospective public-information decision cases with hidden outcomes under strict time cutoffs; achieve success via external BD deals, regulatory approval/launch, and revenue discipline
Must keep track of The architecture must maintain a persistent, continuously updated asset-to-value state record across scientific, regulatory, BD, commercial, financial, and execution constraints via Deal/Approval/Revenue/Investment Arbiter loops.
structure: DAG-with-precedence · goals arrive: given-up-front · count: 45 retrospective decision cases · subgoal credit: no · 0 citations
Horizonthe paper states none
the paper’s own words
The benchmark contains 45 retrospective public-information decision cases with strict time cutoffs, hidden outcomes, common schemas, automatic scoring, and blinded pairwise judging... The value-conversion architecture is a prompt-level approximation of a Company World Model: a Live Asset Value Record updated by Deal, Approval, Revenue, and Investment Arbiter loops.
EnterpriseBench Corecraft: Training Generalizable Agents on High-Fidelity RL Environments
Goals to track perform multi-step, domain-specific customer-support work across a simulated enterprise with 2,500+ entities across 14 entity types; satisfy all expert-authored rubric criteria for a given task using 23 available tools
Must keep track of Agent must track entity state across the enterprise simulation and verify each expert-authored rubric criterion is satisfied before the task counts as solved.
structure: DAG-with-precedence · goals arrive: given-up-front · count: expert-authored rubric criteria per task · subgoal credit: yes · 7 citations
Horizonthe paper states none
the paper’s own words
Frontier models such as GPT-5.2 and Claude Opus 4.6 solve fewer than 30% of tasks when all expert-authored rubric criteria must be satisfied.
CoreCraft is a fully operational enterprise simulation of a customer support organization, comprising over 2,500 entities across 14 entity types with 23 unique tools, designed to measure whether AI agents can perform the multi-step, domain-specific work that real jobs demand.
Can LLM Agents Respond to Disasters? Benchmarking Heterogeneous Geospatial Reasoning in Emergency Operations
Goals to track perform disaster perception from heterogeneous geospatial imagery; conduct spatial relational analysis over roads/population/facilities; plan rescue and evacuation operations; reason about temporal evolution of the disaster; synthesize a multi-modal operational report
Must keep track of The agent must compose calls across a 108-tool MCP library over heterogeneous, multi-temporal geospatial data, tracking intermediate perception/analysis outputs and their correctness as the pipeline progresses through the five task dimensions.
structure: DAG-with-precedence · goals arrive: given-up-front · count: 515 expert-authored tasks; gold trajecto · subgoal credit: no · 4 citations · unit: tool-calls
Horizon3,500 tool-call steps total (across 515 tasks, ~6.8/task average)
the paper’s own words
we introduce Disaster Operational Response Agent benchmark (DORA), the first agentic benchmark for end-to-end disaster response: 515 expert-authored tasks across 45 real-world disaster events spanning 10 types, paired with expert-verified, replayable gold trajectories totaling 3,500 tool-call steps. Tasks span five dimensions that cover the operational disaster-response pipeline: disaster perception, spatial relational analysis, rescue and evacuation planning, temporal evolution reasoning, and multi-modal report synthesis.
DocOps: A Verifiable Benchmark for Autonomous Agents in Complex Document Operations
Goals to track perform atomic document-operation actions correctly (e.g. edits without destructive metadata corruption); complete escalating, more complex, more tightly coupled document-workflow tasks built from those atomic operations; maintain long-term state tracking and global document consistency across the workflow; correctly verify (not just superficially check) that a document edit was semantically achieved
Must keep track of The agent must maintain long-term state tracking of the document's evolving structure/content and verify (rather than superficially assume) that each operation was semantically correct, to preserve global document consistency across highly coupled, long-range tasks.
structure: hierarchical · goals arrive: given-up-front · subgoal credit: yes · 1 citations
Horizonthe paper states none
the paper’s own words
we introduce DocOps, a deterministically verifiable evaluation framework underpinned by a hierarchical taxonomy that deconstructs document operations inspired by real-world practices into atomic dimensions and escalating workflow complexities... a fine-grained analysis of existing agents' manipulation behaviors uncovers 3 key failure modes: long-term state tracking collapse, shallow semantic verification, and destructive editing of structural metadata.
DuMateBench: Evaluating Autonomous Agents in Complex Real-World Workflows
Goals to track complete each of 200 real-session-derived tasks spanning 8 broad scenarios and 17 fine-grained capability categories; coordinate multiple capability categories within a single task; preserve and correctly use persistent configurations and workspace state carried over from prior interaction history; maintain performance under injected real-world environmental complexity (Insufficient, Unstable, Noisy conditions)
Must keep track of The agent must track persistent configurations, workspace state, and pre-solution interaction history carried into each task, while coordinating multiple capability categories under injected Insufficient, Unstable, or Noisy environmental perturbations.
structure: set-of-independent · goals arrive: given-up-front · count: 200 tasks spanning 8 broad scenarios and · subgoal credit: no · 0 citations
Horizonthe paper states none
the paper’s own words
The resulting benchmark comprises 200 tasks spanning 8 broad scenarios and 17 fine-grained capability categories, with most tasks requiring multiple capability coordination. We execute these tasks in isolated Docker containers injected with three forms of real-world environmental complexity: Insufficient, Unstable, and Noisy, and assess performance using a hybrid deterministic and LLM-as-Judge evaluation protocol.
Each task preserves the relevant pre-solution interaction history, persistent configurations, and workspace state, and is then validated through human verification.
Anchor: Mitigating Artifact Drift in Agent Benchmark Generation
Goals to track satisfy explicit procurement-workflow constraints in a production-grade ERP system; satisfy explicit manufacturing-workflow constraints in a production-grade ERP system; reach the solver-certified fully optimal end-state solution for a long-horizon business task, not merely a feasible one
Must keep track of The agent must track state-based verifier conditions tied to the end-state of a production-grade ERP system across a long-horizon procurement or manufacturing workflow, since reward depends solely on end-state business correctness rather than intermediate steps.
structure: other:constraint-optimization-program-with-con · goals arrive: given-up-front · count: 300 long-horizon tasks (ERP-Bench) spann · subgoal credit: no · 1 citations
Horizonthe paper states none
the paper’s own words
we find that generation parameters predict realized difficulty, and that frontier models satisfy explicit task constraints in 26.1% of trials but reach a fully optimal solution in only 17.4% of trials
AI agents are beginning to complete valuable, long-horizon business operations tasks
ERPBench: Evaluating LLM Agents for Enterprise Decision-Making Across Competitive Market Ecologies
Goals to track make coupled decisions each round across pricing, production, procurement, inventory, and finance; maximize valuation/rank while competing against fixed rule-based opponents (Solo) or other evaluated LLM agents (Arena) in a shared market; sustain a coherent business strategy across all six rounds of the ERP simulation
Must keep track of Agent must track its own accumulated valuation, inventory, and finance position across all six rounds, plus (in Arena mode) the shared market state shaped by five other competing LLM agents, to decide coupled pricing/production/procurement actions each round.
structure: sequential-chain · goals arrive: given-up-front · count: 100 fixed problems, each run for 6 round · subgoal credit: yes · 0 citations · unit: turns
Horizon6
the paper’s own words
We introduce ERPBench, an execution-instrumented benchmark for enterprise decision agents in a six-round Enterprise Resource Planning (ERP) simulation with coupled pricing, production, procurement, inventory, finance, and shared-market competition.
EcoGym: Evaluating LLMs for Long-Horizon Plan-and-Execute in Interactive Economies
Goals to track maintain the Vending sub-environment's profitability (net worth, income) as in Vending-Bench; acquire and retain freelance income under budgeted actions (Freelance sub-environment); sustain operational metrics such as DAU under partial observability (Operation sub-environment); maintain long-term strategic coherence across an effectively unbounded, budgeted-action economic horizon
Must keep track of The agent must track budgeted actions, business-relevant outcome metrics (net worth, income, DAU), and partial observability/stochasticity across an effectively unbounded horizon of 1000+ steps (up to 365 day-loops) per evaluation.
structure: open-ended · goals arrive: given-up-front · count: three sub-environments (Vending, Freelan · subgoal credit: no · 1 citations · unit: agent-steps
Horizon1000+ steps over 365 simulated day-loops (evaluation horizon)
the paper’s own words
EcoGym comprises three diverse environments: Vending (adapted from the closed-source Vending-Bench, with full open-source release), Freelance (new), and Operation (new), implemented in a unified decision-making process with standardized interfaces, and budgeted actions over an effectively unbounded horizon (1000+ steps if 365 day-loops for evaluation). The evaluation of EcoGym is based on business-relevant outcomes (e.g., net worth, income, and DAU), targeting long-term strategic coherence and robustness under partial observability and stochasticity.
EntWorld: A Holistic Environment and Benchmark for Verifiable Enterprise GUI Agents
Goals to track complete each of 1,756 tasks spanning six representative enterprise domains (CRM, ITIL, ERP, etc.); operate under strict business logic constraints and high-density enterprise UIs; maintain precise, state-consistent information retrieval verified via SQL-based state-transition checks; complete synthesized long-horizon workflows reverse-engineered from database schemas
Must keep track of The agent must track precise, state-consistent information across a high-density enterprise UI and the underlying database schema, since success is verified deterministically via SQL-based state-transition checks rather than visual matching.
structure: sequential-chain · goals arrive: given-up-front · count: 1,756 tasks across six enterprise domain · subgoal credit: no · 5 citations · also Web & GUI agents
Horizonthe paper states none
the paper’s own words
Experimental results demonstrate that state-of-the-art models (e.g., GPT-4.1) achieve 47.61% success rate on EntWorld, substantially lower than the human performance, highlighting a pronounced enterprise gap in current agentic capabilities.
enabling the synthesis of realistic, long-horizon workflows
Can LLM Agents Be CFOs? Benchmarking Long-Horizon Resource Allocation in an Uncertain Enterprise Environment
Goals to track manage liquidity across the firm's operations; close financial books accurately on a regular cycle; gather costly signals about the macro/industry environment before acting; request equity or debt financing appropriately as conditions change; survive (avoid insolvency/failure) across the full 132-month horizon under shifting macroeconomic regimes
Must keep track of The agent must track liquidity, capital structure (equity/debt), and signals about shifting macroeconomic/industry regimes month by month across the 132-month horizon, since consequences of earlier decisions are delayed and only become apparent later.
structure: DAG-with-precedence · goals arrive: given-up-front · count: 132-month simulated horizon; 23 LLMs and · subgoal credit: yes · 2 citations · unit: other:simulated-months
Horizon132-month CFO simulation
the paper’s own words
We introduce EnterpriseArena, a 132-month CFO simulator that evaluates long-horizon resource allocation under uncertainty in a FinTech lending firm. Agents must manage liquidity, close books, gather costly signals, and request equity or debt financing across changing macroeconomic regimes.
EnterpriseLab: A Full-Stack Platform for developing and deploying agents in Enterprises
Goals to track complete complex enterprise workflows spanning IT, HR, sales, and engineering domains; correctly invoke and sequence tool calls across 140+ tools exposed via Model Context Protocol across 15 applications; match frontier-model performance while running as a smaller (8B) privacy-preserving model; generalize/remain robust across diverse enterprise benchmarks (EnterpriseBench, CRMArena)
Must keep track of The trained agent must track which of many interdependent enterprise tools/applications it has invoked and their resulting state, to complete complex multi-tool enterprise workflows without needing frontier-scale model capacity.
structure: DAG-with-precedence · goals arrive: given-up-front · count: 15 applications; 140+ tools across IT, H · subgoal credit: no · 1 citations
Horizonthe paper states none
the paper’s own words
Our results demonstrate that 8B-parameter models trained within EnterpriseLab match GPT-4o's performance on complex enterprise workflows while reducing inference costs by 8-10x, and remain robust across diverse enterprise benchmarks, including EnterpriseBench (+10%) and CRMArena (+10%).
We validate the platform through EnterpriseArena, an instantiation with 15 applications and 140+ tools across IT, HR, sales, and engineering domains.
EnterpriseClawBench: Benchmarking Agents from Real Workplace Sessions
Goals to track read heterogeneous workplace files relevant to a task; invoke the correct tools to accomplish a business objective; deliver a business artifact matching role-specific hard rules and semantic rubrics
Must keep track of The agent must track which files it has read, what tools it has invoked, and whether its eventual delivered artifact satisfies the task's hard rules and semantic rubric, since harness-model combination, artifact delivery, and visual quality are all separately reported.
structure: sequential-chain · goals arrive: given-up-front · count: 852 reproducible tasks, each with recove · subgoal credit: yes · 1 citations
Horizonthe paper states none
the paper’s own words
Starting from a large archive of workplace sessions, the EnterpriseClawBench produces 852 reproducible tasks, each paired with recovered fixtures, rewritten prompts, role classes, skill subclasses, hard rules, and semantic rubrics.
EnterpriseOps-Gym: Environments and Evaluations for Stateful Agentic Planning and Tool Use in Enterprise Settings
Goals to track complete each of 1,150 expert-curated enterprise tasks across eight mission-critical verticals (e.g., Customer Service, HR, IT); plan and act correctly amid persistent state changes and strict access-control protocols; correctly refuse infeasible tasks rather than attempting unintended, potentially harmful side effects
Must keep track of Agent must track persistent state changes across 164 database tables, access-control constraints, and whether the current task is feasible at all, to avoid unintended side effects from attempting an infeasible task.
structure: DAG-with-precedence · goals arrive: given-up-front · count: 1,150 expert-curated tasks across eight · subgoal credit: yes · 11 citations
Horizonthe paper states none
the paper’s own words
EnterpriseOps-Gym features a containerized sandbox with 164 database tables and 512 functional tools to mimic real-world search friction. Within this environment, agents are evaluated on 1,150 expert-curated tasks across eight mission-critical verticals (including Customer Service, HR, and IT).
Evaluating Deep Research Agents on Expert Consulting Work: A Benchmark with Verifiers, Rubrics, and Cognitive Traps
Goals to track produce a multi-document, decision-grade consulting deliverable from SME-authored prompts; pass deterministic binary verifiers (mean 14.9 per task); satisfy each of a five-criterion SME rubric (Data Integrity, Analytical Rigor, Relevance & Focus, Execution Precision, Format & Deliverability); avoid embedded cognitive traps (human-error mimicry, deterministic precision traps) that penalize surface-pattern reas
Must keep track of The agent must track which of the ~14.9 verifiers per task it has satisfied, avoid embedded cognitive traps requiring reconciliation against context, and keep its final deliverable consistent with all five rubric criteria simultaneously.
structure: set-of-independent · goals arrive: given-up-front · count: mean 14.9 binary verifiers per task, plu · subgoal credit: yes · 1 citations · unit: other:not-stated
Horizonthe paper states none
the paper’s own words
we score two complementary layers: deterministic binary verifiers (mean 14.9 per task) and a five-criterion 0--3 SME rubric ... combined into a Verifier-Rubric Score (VRS, 0--100). Acceptance under a joint threshold (rubric mean >= 2.5 and verifier pass rate >= 80%) is uniformly low
a benchmark of 70 SME-authored management consulting prompts, each embedding cognitive traps that penalize surface-pattern reasoning
The Agent's First Day: Benchmarking Learning, Exploration, and Scheduling in the Workplace Scenarios
Goals to track schedule and prioritize streaming tasks with varying priorities under a dynamic workload; actively seek out information to reduce hallucination before acting (prudent information acquisition); distill and reuse generalized strategies learned from earlier dynamically generated tasks in later ones (continuous evolution)
Must keep track of The agent must track incoming streaming tasks and their priorities, its own accumulated exploration/experience for continual learning, and its confidence in acquired information to avoid hallucination, across a continuously evolving workplace scenario.
structure: hierarchical · goals arrive: emitted-by-environment-over-time · count: 50 dynamic scenarios, each with 2-6 task · subgoal credit: yes · 1 citations · unit: actions
Horizon50 dynamic scenarios, each with 2-6 task instances; observed usage up to ~90 steps and 232 tool calls for the highest-usage evaluated model (Gemini-3-Flash)
the paper’s own words
\method{} evaluates agents along three dimensions: (1) context-aware scheduling for streaming tasks with varying priorities; (2) prudent information acquisition to reduce hallucination via active exploration; and (3) continuous evolution by distilling generalized strategies from rule-based, dynamically generated tasks.
Each scenario encompasses 2 to 6 task instances... While Gemini-3-Flash uses substantially more steps (90) and tool calls (232) than the middle-tier models, this increased activity reflects the necessary complexity
FM-Bench: A Benchmark for Long-Horizon Management with Competing Agents
Goals to track draft and manage a football squad within a fixed budget shared with rivals; trade players and negotiate contracts across the season; invest in facilities and youth development; set match lineups; maintain board confidence (avoid being fired) while maximizing a final accumulated score over 20 in-game years
Must keep track of The agent must track squad composition, budget/cash flow, ongoing contract-renewal deadlines, facility/youth investments, and the board's confidence in it, across roughly 340-400 decision stops spanning 20 in-game years, since 'the order settles only late in the horizon.'
structure: sequential-chain · goals arrive: given-up-front · count: 20 in-game years; 26 tools; roughly 340- · subgoal credit: yes · 0 citations · unit: other:decision-stops
Horizon20 in-game years; roughly 340 to 400 decision stops per run
the paper’s own words
An LLM agent runs a football club for 20 in-game years through 26 tools and roughly 340 to 400 decision stops. It drafts a squad on the same budget as every rival, trades players, negotiates contracts, invests in facilities and youth, sets lineups, and answers to a board that can fire it, while a deterministic engine accumulates every year into one final score with no LLM judge or human rater.
FORESIGHT-9: Prospective and Process-Aware Evaluation of Adaptive Trading Agents
Goals to track maintain a coherent, active factor-ensemble/trading strategy across staged macro-financial events without silent internal collapse; respond adaptively to joint multi-asset anchors and disclosed in-world-time observations across a multi-year counterfactual worldline; keep portfolio holdings and decision-records mutually consistent (avoid decision-record vs. execution divergence)
Must keep track of The agent must track its adaptive factor library/ensemble state, executed portfolio holdings, and staged event disclosures across a multi-year, ~2,468-trading-day counterfactual worldline, ensuring internal decision-records remain consistent with actual executed holdings.
structure: sequential-chain · goals arrive: emitted-by-environment-over-time · count: nine worldlines x two agent frameworks x · subgoal credit: yes · 0 citations · unit: simulated-days
Horizoneach worldline runs from the 2026-07-15 information boundary through 2035-12-31, approximately 2,468 simulated trading days (~9 years 5 months) per run, across 36 long-ho
the paper’s own words
We evaluate two adaptive trading-agent frameworks with two foundation-model backbones across 36 long-horizon runs. Agent rankings vary substantially across worldlines and backbones, and a fixed equal-weight policy outperforms 31 of 36 runs. Process telemetry exposes failures that terminal returns conceal: in one high-return run, the live factor library collapsed while executed holdings converged to the equal-weight fallback, even though decision records continued to report an active factor ensemble.
each run terminates at 2035-12-31 ... approximately 2,468 daily marks
FinEvo-Bench: A Longitudinal Benchmark for Self-Evolving Agents in Professional Financial Workflows
Goals to track complete each of 120 real-case-grounded financial tasks across 20 business scenes in six financial domains, following institution-provided procedures and constraints; retain and re-apply experience from earlier cases in a scene to later, substantively distinct cases sharing the same professional procedure (self-evolution); maintain financial-compliance quality across each task as scored by a manually reviewed rubric;
Must keep track of The agent (self-evolving scaffold) must retain experience -- via memory, skill distillation, or both -- from earlier cases in the interleaved task stream and apply it to later, procedurally-related but factually distinct cases, while satisfying institution-provided professional-procedure constraints and financial-compliance requirements on each individual task.
structure: other:six-related-cases-per-scene-with-longitu · goals arrive: emitted-by-environment-over-time · count: 120 tasks; 20 business scenes across six · subgoal credit: yes · 1 citations · unit: episodes
Horizon120 real-case-grounded tasks per longitudinal stream (20 scenes x 6 cases each); three independently shuffled, globally interleaved task streams
the paper’s own words
Each scene contains six related but substantively distinct cases that share a professional procedure and a manually reviewed rubric for task quality and financial compliance... Letta achieves the highest evolved score (91.65) and fewest compliance issues (0.09 per task); Codex achieves the largest self-evolution gain (+19.37).
We introduce FinEvo-Bench, a longitudinal benchmark with 120 real-case-grounded tasks, 20 business scenes across six financial domains... We compare four self-evolving agent scaffolds using the same Qwen3.7-Max backbone and three independently shuffled, globally interleaved task streams.
FrontierFinance: A Long-Horizon Computer-Use Benchmark of Real-World Financial Tasks
Goals to track complete each of 25 complex financial modeling tasks across five core finance models; produce client-ready outputs matching industry-standard financial-modeling workflows; satisfy detailed structured evaluation rubrics per task
Must keep track of The agent must track detailed financial-modeling state (assumptions, formulas, model linkages) across a task requiring on average over 18 hours of equivalent skilled human labor, checked against detailed structured rubrics.
structure: sequential-chain · goals arrive: given-up-front · count: 25 complex financial modeling tasks acro · subgoal credit: yes · 2 citations · unit: human-expert-hours
Horizonaverage of over 18 hours of skilled human labor per task
the paper’s own words
we introduce FrontierFinance, a long-horizon benchmark of 25 complex financial modeling tasks across five core finance models, requiring an average of over 18 hours of skilled human labor per task to complete. Developed with financial professionals, the benchmark reflects industry-standard financial modeling workflows and is paired with detailed rubrics for structured evaluation.
GDPevo: Evaluating Agent Self-Evolution on Real Business Tasks
Goals to track update an agent's persistent state from prior training-task experience (self-evolution); apply recombined atomic business rules correctly on held-out test tasks across CRM/ERP/finance/healthcare/legal/data-centric workflows; achieve held-out accuracy gains attributable specifically to training experience rather than data contamination
Must keep track of The evaluation harness must track which atomic business rules were exposed during training versus recombined at test time, per task group, to attribute held-out gains specifically to training experience rather than contamination.
structure: set-of-independent · goals arrive: given-up-front · count: V1: 120 tasks in 12 groups (5 training + · subgoal credit: yes · 0 citations
Horizonthe paper states none
the paper’s own words
Its core mechanism, rule hybridization, decomposes each enterprise workflow into atomic business rules, distributes subsets of these rules across training tasks, and recombines them in held-out test tasks so that test-time gains are attributable.
V1 release contains 120 tasks in 12 groups, with five training and five held-out test tasks per group
H-AdminSim: A Multi-Agent Simulator for Realistic Hospital Administrative Workflows with FHIR Integration
Goals to track process hospital administrative requests (drawn from a workload of 10,000+ requests/day in large hospitals); coordinate across multiple administrative subtasks rather than handling them in isolation; operate correctly against FHIR-integrated, heterogeneous hospital-setting data
Must keep track of The multi-agent simulation must track hospital administrative request state across heterogeneous hospital settings via a unified FHIR-integrated environment, evaluated against detailed rubrics.
structure: other:multi-agent-simulated-administrative-wor · goals arrive: given-up-front · subgoal credit: yes · 0 citations · also Multi-agent organisations & societies
Horizonthe paper states none
the paper’s own words
These tasks are quantitatively evaluated using detailed rubrics, enabling systematic comparison of LLMs.
in large hospitals, process over 10,000 requests per day
HANDBOOK.md: A Benchmark for Long-Context Agentic Instruction Following
Goals to track locate the specific clauses in a long standing-policy document (20-124 pages) that apply to the current situation; carry out routine professional work strictly governed by that policy document across many tool-mediated actions; avoid letting a plausible but unauthorized in-environment request override the standing policy; satisfy every one of many deterministic, programmatic grading criteria (824 total) simultaneousl
Must keep track of The agent must hold the relevant clauses of a long standing-policy document across roughly 17 reasoning steps and 30 tool calls on average, checking each action against required and prohibited criteria rather than losing rule details over the long horizon.
structure: DAG-with-precedence · goals arrive: given-up-front · count: 65 agentic tasks; a rubric of 824 total · subgoal credit: yes · 2 citations · unit: tool-calls
Horizon~17 reasoning steps and ~30 tool calls on average per task
the paper’s own words
We present HANDBOOK_md, a benchmark of 65 agentic tasks modeled on how employees follow company handbooks. Each task places an agent in a self-contained company environment... and instructs it to carry out routine professional work governed by an expert-written standard operating procedure of 20-124 pages... each task carries a rubric of programmatic criteria (824 in total) that check both that required actions occurred and that prohibited actions did not.
Completing a task requires locating the clauses that apply, holding them across a horizon of roughly 17 reasoning steps and 30 tool calls on average.
HealthAdminBench: Evaluating Computer-Use Agents on Healthcare Administration Tasks
Goals to track complete a Prior Authorization workflow end-to-end across an EHR and payer portal; complete an Appeals and Denials Management workflow; complete a Durable Medical Equipment (DME) Order Processing workflow; satisfy each of the fine-grained verifiable subtasks a task decomposes into (often 15+ per task)
Must keep track of The agent must track progress across many fine-grained, cross-application subtasks (spanning EHR, payer portals, and fax) per administrative workflow, under a fixed interaction budget, to reach a correct terminal workflow state.
structure: DAG-with-precedence · goals arrive: given-up-front · count: 135 tasks yielding 1,698 evaluation poin · subgoal credit: yes · 8 citations · also Healthcare & clinical · unit: other:not-stated
Horizonthe paper states none
the paper’s own words
Each task is decomposed into fine-grained, verifiable subtasks, yielding 1,698 evaluation points. ... the best-performing agent (Claude Opus 4.6 CUA) achieves only 36.3 percent task success, while GPT-5.4 CUA attains the highest subtask success rate (82.8 percent).
Since many tasks involve 15 or more subtasks ... agents operate 'under a fixed interaction budget' and references 'maximum steps,' but does not specify an exact number.
Herculean: An Agentic Benchmark for Financial Intelligence
Goals to track complete Trading workflow tasks via a standardized MCP-based skill environment; complete Hedging workflow tasks requiring long-horizon coordination and state consistency; complete Market Insights workflow tasks; complete Auditing workflow tasks requiring structured verification
Must keep track of The agent must track workflow-specific state, tool interactions, and constraints within each MCP-based skill environment, with Hedging and Auditing additionally requiring sustained state consistency and structured verification across long-horizon coordination.
structure: set-of-independent · goals arrive: given-up-front · count: 4 representative financial workflows (Tr · subgoal credit: yes · 3 citations
Horizonthe paper states none
the paper’s own words
We introduce Herculean, the first skilled benchmark for agentic financial intelligence spanning four representative workflows, including Trading, Hedging, Market Insights, and Auditing. Each workflow is instantiated as a standardized MCP-based skill environment with its own tools, interaction dynamics, constraints, and success criteria, enabling consistent end-to-end assessment of heterogeneous agent systems.
where long-horizon coordination, state consistency, and structured verification are critical
JobBench: Aligning Agent Work With Human Will
Goals to track complete 130 agentic tasks spanning 35 occupations that experts identify as high-priority for delegation; reason through cluttered, heterogeneous reference-file information streams within a professional workspace; satisfy a fact-anchored chain of rubrics averaging 35.6 binary criteria per task
Must keep track of Agent must correctly reason through a workspace of heterogeneous reference files and track satisfaction of an average of 35.6 fact-anchored binary rubric criteria per task.
structure: hierarchical · goals arrive: given-up-front · count: 130 tasks across 35 occupations; averagi · subgoal credit: yes · 4 citations
Horizonthe paper states none
the paper’s own words
Each task is packaged as a workspace of heterogeneous reference files, requiring the agent to reason through the cluttered information streams of real professional work. Outputs are graded by a fact-anchored chain of rubrics, averaging 35.6 binary criteria per task.
JobBench covers 130 agentic tasks across 35 occupations.
LH-Bench: Skill-Grounded Evaluation of Long-Horizon Agents on Subjective Enterprise Tasks
Goals to track produce subjective, context-dependent enterprise work whose quality depends on organizational goals and user intent, not a single correct answer; produce correct intermediate artifacts across long, multi-tool workflows (e.g., chapter-level content, Figma-to-code conversions); satisfy expert-grounded rubrics scoring subjective work quality; align with pairwise human preference judgments as convergent validation
Must keep track of The agent must track and produce correct intermediate artifacts (e.g., each chapter of a course, or each Figma-to-code conversion) across long, multi-tool workflows, since ground-truth artifacts enable stepwise reward signals at that granularity rather than only a single final judgment.
structure: hierarchical · goals arrive: given-up-front · count: two environments: Figma-to-code (33 real · subgoal credit: yes · 2 citations
Horizonthe paper states none
the paper’s own words
The pillars are: (i) expert-grounded rubrics that give LLM judges the domain context needed to score subjective work, (ii) curated ground-truth artifacts that enable stepwise reward signals (e.g., chapter-level annotation for content tasks), and (iii) pairwise human preference evaluation for convergent validation.
Programmatic content (41 courses comprising 183 individually-evaluated chapters on a course platform serving 30+ daily users)
Post-Training on Office Work Improves Software Engineering: A Behavioral Account of Cross-Domain Transfer
Goals to track select goals appropriately within nested and branching office work; construct task-relevant state as work unfolds; maintain fidelity to higher-level objectives across nested/branching sub-work; verify completion against the environment
Must keep track of Agent must repeatedly select goals, construct task-relevant state, maintain fidelity to higher-level objectives, and verify completion against the environment across nested and branching long-horizon office work.
structure: hierarchical · goals arrive: mixed:given-up-front-tasks-with-self-directed- · count: 363 Long-Horizon Multi-Tool Agent (LHMTA · subgoal credit: yes · 0 citations
Horizonthe paper states none
the paper’s own words
We call this capability goal-directed execution (GDE): the repeated application of four behaviors, namely selecting goals, constructing task-relevant state, maintaining fidelity to higher-level objectives, and verifying completion against the environment... We test this by post-training Qwen3.5-122B-A10B on 363 Long-Horizon Multi-Tool Agent (LHMTA) tasks drawn from office workflows.
Four-Axis Decision Alignment for Long-Horizon Enterprise AI Agents
Goals to track make high-stakes regulated decisions (loan qualification, insurance claims adjudication) under lossy memory and multi-step reasoning; maintain factual precision (FRP) about case facts over the long horizon; maintain reasoning coherence (RCS) across the multi-step decision process; reconstruct compliance/regulatory justification (CRR) for the decision; calibrate abstention (CAR) -- know when to decline to decide rathe
Must keep track of The agent must retain case facts accurately over a long-horizon multi-step review under lossy memory, maintain a coherent reasoning chain, reconstruct which regulatory rules justify the decision, and calibrate when to abstain rather than commit -- all against deterministic ground truth.
structure: set-of-independent · goals arrive: given-up-front · count: four alignment axes (FRP, RCS, CRR, CAR) · subgoal credit: yes · 1 citations
Horizonthe paper states none
the paper’s own words
We propose that long-horizon decision behavior decomposes into four orthogonal alignment axes, each independently measurable and failable: factual precision (FRP), reasoning coherence (RCS), compliance reconstruction (CRR), and calibrated abstention (CAR).
Long-horizon enterprise agents make high-stakes decisions (loan underwriting, claims adjudication, clinical review, prior authorization) under lossy memory, multi-step reasoning, and binding regulatory constraints.
LongMedBench: Benchmarking Medical Agents for Long-Horizon Clinical Decision-Making
Goals to track aggregate evidence across a patient's repeated visits, tests, and evolving treatments to answer fact-based QA; perform temporal reasoning over the patient's event stream, including implicit (not just explicit-timestamped) time inference; make a long-horizon clinical decision that correctly uses historical patient information accumulated over many visits
Must keep track of Agent must integrate time-series clinical events (admission records, notes) across many visits per patient, correctly distinguish explicit timestamps from cases requiring implicit time inference, and carry this forward into long-horizon decision-making.
structure: sequential-chain · goals arrive: emitted-by-environment-over-time · count: 335 patients, avg 19.72 inpatient visits · subgoal credit: unclear · 0 citations · also Healthcare & clinical · unit: sessions
Horizon19.72 inpatient visits per patient (average); 44.91 medical events per visit (average)
the paper’s own words
It comprises 335 patients, with 19.72 inpatient visits per patient on average and 44.91 medical events per visit.
MBABench: Evaluating LLM Agents on End-to-End Spreadsheet Tasks in Finance
Goals to track construct an entire financial spreadsheet (e.g., financial model, forecast, scenario analysis) end-to-end from a high-level instruction; jointly satisfy Accuracy, Formula, and Format criteria, each with fine-grained sub-criteria reflecting professional standards
Must keep track of Agent must track intermediate calculation dependencies within the spreadsheet, professional formatting/readability conventions, and formula correctness simultaneously while building the deliverable end-to-end.
structure: set-of-independent · goals arrive: given-up-front · count: three top-level evaluation dimensions (A · subgoal credit: yes · 2 citations
Horizonthe paper states none
the paper’s own words
To reflect the multidimensional nature of solution quality, we develop an evaluation taxonomy comprising three dimensions: Accuracy, Formula, and Format, each comprising fine-grained criteria that reflect professional standards.
Evaluating over 18 agents, the benchmark reveals that even the strongest agents fall short of basic professional finance standards, and their performance degrade sharply as the difficulty increases beyond a few chained calculations.
MerchantBench: Benchmarking LLM Agents for Long-Term Coherence in E-Commerce Operations
Goals to track source products and manage upstream supplier events over time; set and adjust listing/pricing to remain competitive and solvent; manage cash-flow across delayed, heterogeneous-latency order outcomes; follow individual order lifecycles end-to-end and revisit earlier decisions as new (delayed) information arrives
Must keep track of The agent must follow individual order lifecycles end-to-end, track upstream supplier events and their promised delayed downstream outcomes, manage cash-flow/net-assets state, and revisit/adapt earlier sourcing and pricing decisions as delayed feedback arrives across a full 365-simulated-day run.
structure: DAG-with-precedence · goals arrive: emitted-by-environment-over-time · count: recurrent decisions across 4 named categ · subgoal credit: yes · 1 citations · unit: simulated-days
Horizona 365-day order-level simulation; 48 runs, each spanning 365 simulated days
the paper’s own words
Seller-side e-commerce provides a suitable setting for this evaluation through recurrent and interdependent decisions over Product Sourcing, Listing and Pricing Control, Cash-Flow Management, and Mixed-Latency Feedback Adaptation. We introduce MerchantBench, a 365-day order-level simulation grounded in 98,843 real e-commerce product records and equipped with 26 tools for agent interaction. MerchantBench couples promptly observable Upstream Supplier Events with delayed Downstream Order Outcomes, requiring agents to follow individual order lifecycles and revisit earlier decisions.
CORPGEN: Simulating Corporate Environments with Autonomous Digital Employees in Multi-Horizon Task Environments
Goals to track manage dozens of concurrent, interleaved long-horizon corporate tasks (45+ tasks, 500-1500+ steps each); handle inter-task dependencies expressed as DAGs rather than simple chains; reprioritize among concurrent tasks as load and context change over a persistent execution context spanning hours
Must keep track of The system must maintain hierarchical goal alignment, isolate sub-agent context to prevent cross-task contamination, and manage tiered (working/structured/semantic) memory with adaptive summarization across a persistent execution context spanning hours.
structure: DAG-with-precedence · goals arrive: mixed:given-up-front-with-reprioritization-emi · count: 45+ concurrent tasks, each requiring 500 · subgoal credit: no · 0 citations · unit: agent-steps
Horizon500-1500+
the paper’s own words
We identify four failure modes that cause baseline CUAs to degrade from 16.7% to 8.7% completion as load scales 25% to 100%, a pattern consistent across three independent implementations. These failure modes are context saturation (O(N) vs O(1) growth), memory interference, dependency complexity (DAGs vs. chains), and reprioritization overhead.
requiring coherent execution across dozens of interleaved tasks (45+, 500-1500+ steps) within persistent execution contexts spanning hours.
OR-Space: A Full-Lifecycle Workspace Benchmark for Industrial Optimization Agents
Goals to track construct a solver-ready optimization model from heterogeneous business artifacts (Build); revise an existing model under changing requirements or solver feedback while preserving valid prior logic (Revise); answer grounded questions about solutions, constraints, and business implications using evidence spread across workspace artifacts (Explain)
Must keep track of The agent must track the current state of a persistent multi-artifact workspace (documents, structured data, code, solver outputs) across model construction, revision, and explanation stages, preserving valid prior modeling logic when new requirements or solver feedback arrive.
structure: DAG-with-precedence · goals arrive: mixed:given-up-front-with-new-requirements-inj · count: 3 task modes (Build, Revise, Explain) · subgoal credit: unclear · 2 citations
Horizonthe paper states none
the paper’s own words
OR-Space defines three task modes: Build, where agents construct solver-ready optimization models from heterogeneous artifacts; Revise, where agents modify existing models under changing requirements or solver feedback while preserving valid prior logic; and Explain, where agents answer grounded questions about solutions, constraints, and business implications using evidence spread across workspace artifacts.
persistent multi-artifact workspaces and multi-stage task lifecycles
UI-Mate: Advancing Open-Weight Foundation GUI Agents with In-Context Demonstrations
Goals to track complete each of 100 long-horizon office tasks spanning 41 applications; follow demonstrated subtask-level workflows (self-demo or variant-demo) and re-plan from the live interface when needed; make measurable progress toward task completion even when strict end-to-end success is not reached
Must keep track of Agent must track its position in a demonstrated or self-planned subtask sequence, detect when the live interface diverges from the demonstration, and re-plan accordingly across a long-horizon office task.
structure: sequential-chain · goals arrive: given-up-front · count: 100 long-horizon office tasks across 41 · subgoal credit: yes · 0 citations · also Web & GUI agents
Horizonthe paper states none
the paper’s own words
OSWorkerBench Benchmark and Insights: A benchmark of 100 long-horizon office tasks across 41 applications that supports instruction-only and demonstration-guided evaluation.
OccuBench: Evaluating AI Agents on Real-World Professional Tasks via Language Environment Simulation
Goals to track complete each of 100 real-world professional task scenarios spanning 65 specialized domains; maintain task completion under controlled fault injection (explicit errors, implicit data degradation, mixed faults)
Must keep track of Agent must track task-completion progress plus signals of environmental robustness (whether tool responses are timing out, truncated, or subtly degraded) within each professional scenario.
structure: set-of-independent · goals arrive: given-up-front · count: 100 real-world professional task scenari · subgoal credit: yes · 4 citations
Horizonthe paper states none
the paper’s own words
OccuBench evaluates agents along two complementary dimensions: task completion across professional domains and environmental robustness under controlled fault injection (explicit errors, implicit data degradation, and mixed faults).
We introduce OccuBench, a benchmark covering 100 real-world professional task scenarios across 10 industry categories and 65 specialized domains, enabled by Language Environment Simulators (LESs) that simulate domain-specific environments through LLM-driven tool response generation.
OmegaUse-OfficeVal: Benchmarking LLM Agents on Long-Horizon Office-Suite Tasks with Economic Grounding
Goals to track complete 100 practitioner-derived office-suite tasks end-to-end; achieve deliverable quality verified via fine-grained rubric-based code verifiers; remain economically competitive relative to human labor time and task price
Must keep track of Agent must track fine-grained rubric criteria and produce a deliverable matching the practitioner's request, verified via code-based verifiers, implicitly benchmarked against the human labor time (avg. 2.32 hours) needed for the same task.
structure: set-of-independent · goals arrive: given-up-front · count: 100 tasks derived from practitioner offi · subgoal credit: yes · 0 citations · unit: human-expert-hours
Horizon2.32 (average)
the paper’s own words
To support stable evaluation, we develop code-based verifiers from fine-grained rubrics. We evaluate several frontier LLMs together with a human baseline. Although all evaluated LLMs are substantially cheaper and faster than human workers, they have not yet approached human-level deliverable quality.
The benchmark comprises 100 tasks derived from office-suite requests proposed by practitioners and adapted through a privacy-preserving process. On average, these tasks require 2.32 hours of human labor to complete.
OptAgent: an Agentic AI framework for Intelligent Building Operations
Goals to track assess how a system/control upgrade changes energy use; assess how the same upgrade changes operating cost; assess how the same upgrade changes thermal comfort; assess how the same upgrade changes flexibility, via coordinated multi-domain multi-agent analytics
Must keep track of The orchestrator must track which specialist agent/tool has been invoked for which sub-domain (thermal dynamics, HVAC, DER), intermediate physics-informed simulation outputs, and how upgrades propagate across energy use, cost, comfort, and flexibility metrics within one workflow.
structure: DAG-with-precedence · goals arrive: given-up-front · count: large-scale benchmark of about 4,000 run · subgoal credit: no · 1 citations · unit: other:not-stated
Horizonthe paper states none
the paper’s own words
an agentic AI layer with 11 specialist agents and 72 Model Context Protocol (MCP) tools that enable end-to-end execution of multi-step energy analytics. A representative case study demonstrates multi-domain, multi-agent coordination for assessing how system and control upgrades affect energy use, operating cost, thermal comfort, and flexibility.
a large-scale benchmark (about 4000 runs) systematically evaluates workflow performance in terms of accuracy, token consumption, execution time, and inference cost
Evidence-Grounded AI for Musculoskeletal Care
Goals to track integrate evolving imaging, laboratory, pathology, and order data as it arrives across visits; produce evidence-based decisions at each stage of care, from admission diagnosis through rehabilitation planning; maintain continuous, individualised management across the full musculoskeletal care pathway rather than isolated per-visit decisions
Must keep track of System must continuously retrieve and integrate real-time imaging, laboratory, pathology, and order data across visits/departments/hospital systems, translating evolving patient state into stage-specific functional goals across the whole care pathway.
structure: sequential-chain · goals arrive: emitted-by-environment-over-time · subgoal credit: yes · 0 citations · also Healthcare & clinical · unit: other:real-world-months-to-years-per-patient-pathway
Horizonmonths to years (per patient pathway); 1,870 cases / 8,240 inpatients across the study
the paper’s own words
Clinicians must repeatedly integrate evolving patient evidence, medical knowledge and stage-specific functional goals, yet evidence is often fragmented across visits, departments and hospital systems, disrupting continuous, individualised management.
recovery, remodelling and degeneration of bones, joints and related tissues unfold over months to years, care requires longitudinal management rather than isolated decisions
POLARIS: Typed Planning and Governed Execution for Agentic AI in Back-Office Automation
Goals to track synthesize a type-checked directed acyclic graph (DAG) plan for a back-office document-processing task; select a single compliant plan via rubric-guided reasoning among structurally diverse candidate DAGs; pass validator-gated checks and a bounded repair loop before execution; route or block side effects per compiled policy guardrails, including anomaly routing
Must keep track of System must track the candidate DAG structure, per-node type/validator status, and the full execution trace/audit trail to support bounded repair and decision-grade anomaly routing.
structure: DAG-with-precedence · goals arrive: given-up-front · subgoal credit: yes · 2 citations
Horizonthe paper states none
the paper’s own words
A planner proposes structurally diverse, type checked directed acyclic graphs (DAGs), a rubric guided reasoning module selects a single compliant plan, and execution is guarded by validator gated checks, a bounded repair loop, and compiled policy guardrails that block or route side effects before they occur.
Applied to document centric finance tasks, POLARIS produces decision grade artifacts and full execution traces while reducing human intervention.
PPT-Eval: A Benchmark for Computer-Use Agents on PowerPoint Tasks
Goals to track complete content-creation and presentation-editing tasks across 120 PowerPoint tasks in 12 files; satisfy task-specific rubric criteria that award partial credit for intermediate steps; avoid unnecessary changes and poor aesthetics while making required edits
Must keep track of Agent must track which task-specific rubric criteria (intermediate steps, aesthetics, unnecessary changes) have been satisfied across a multimodal editing session, since rubrics award partial credit rather than only final binary success.
structure: hierarchical · goals arrive: given-up-front · count: 120 PowerPoint tasks across 12 files, or · subgoal credit: yes · 5 citations · also Web & GUI agents
Horizonthe paper states none
the paper’s own words
we design a robust evaluation framework to help create task-specific rubrics for PowerPoint tasks... These rubrics award partial credit for intermediate steps, penalize unnecessary changes and poor aesthetics, and provide natural language feedback. This nuanced approach proves highly effective, achieving a Kendall's tau-b correlation of 0.77 with human judgments.
We introduce PPT-Eval, a benchmark of 120 PowerPoint tasks across 12 files that cover both content creation and presentation editing scenarios, organized by difficulty.
PhysicianBench: Evaluating LLM Agents in Real-World EHR Environments
Goals to track retrieve relevant clinical data across multiple encounters in the EHR; reason over heterogeneous clinical information (labs, notes, orders) to reach a decision; execute consequential clinical actions (e.g. prescribing, ordering) grounded against the environment; produce clinical documentation reflecting the completed workflow
Must keep track of Agent must track data retrieved across multiple encounters, intermediate clinical reasoning state, and which structured checkpoints have been satisfied, using execution-grounded verification against real patient records via standard EHR APIs.
structure: hierarchical · goals arrive: given-up-front · count: 100 long-horizon tasks; 670 checkpoints · subgoal credit: yes · 14 citations · also Healthcare & clinical · unit: tool-calls
Horizon27 (average)
the paper’s own words
Each task is decomposed into structured checkpoints (670 in total across the benchmark) capturing distinct stages of completion graded by task-specific scripts with execution-grounded verification.
requiring an average of 27 tool calls per task
PolyWorkBench: Benchmarking Multilingual Long-Horizon LLM Agents
Goals to track process heterogeneous multilingual inputs correctly within a workflow task; perform iterative reasoning and invoke external tools while maintaining linguistic consistency; produce a structured, correct output for one of five workplace domains (commerce, knowledge work, legal analysis, localization, manufacturing)
Must keep track of Agent must track functional correctness and linguistic consistency simultaneously across a workflow's reasoning and tool-invocation steps, since the two can drift independently as multilingual inputs are processed.
structure: sequential-chain · goals arrive: given-up-front · count: 67 tasks across 5 domains · subgoal credit: yes · 1 citations
Horizonthe paper states none
the paper’s own words
PolyWorkBench consists of 67 tasks across five domains, including commerce, knowledge work, legal analysis, localization, and manufacturing, where agents must process heterogeneous multilingual inputs, perform iterative reasoning, invoke external tools, and produce structured outputs.
PolyWorkBench: Benchmarking LLM Agents for Cross-Lingual Long-Horizon Workflows
Goals to track integrate heterogeneous multilingual inputs relevant to a workplace task; execute iterative tool-use trajectories within a target domain (commerce, knowledge work, legal analysis, localization, manufacturing); produce structured domain artifacts as verifiable task output
Must keep track of Agent must track and correctly integrate heterogeneous multilingual information while executing an iterative sequence of tool calls, and produce a structured domain artifact that is later verified.
structure: sequential-chain · goals arrive: given-up-front · count: 67 tasks across five core domains: comme · subgoal credit: yes · 0 citations
Horizonthe paper states none
the paper’s own words
we adopt Grade, a task-specific structural scoring rubric, as our primary ranking metric, and complement it with Pytest for executable state verification and LLM-as-Judge for semantic quality diagnostics.
enterprise workflows inherently require processing multilingual resources across extended trajectories. The interaction between multilinguality and long-horizon execution, however, remains underexplored.
PowerAgentBench-SS: A Benchmark for Agentic AI in Power System Steady-State Studies
Goals to track inspect a grid case and select appropriate tools/simulators for the workflow; screen a large space of contingencies within a limited validation budget; propose admissible mitigations for discovered risks and validate their physical validity; produce an auditable evidence trail and submit a bounded, ranked report of top contingencies
Must keep track of The agent must track which contingencies it has already validated (against a fixed budget of 80 out of 1,035), the evidence log supporting each, and whether its running set of top candidates still reflects the best available evidence as it allocates its remaining validation budget.
structure: sequential-chain · goals arrive: given-up-front · count: 1,035 total N-2 contingency cases per in · subgoal credit: yes · 4 citations · unit: actions
Horizonvalidation budget of 80 contingency-case validations per episode (out of 1,035 total N-2 cases), submitting a ranked report of exactly 20 contingencies; hidden dangerous
the paper’s own words
The benchmark exposes public case data, action constraints, a tool API, and a validation budget to an agent, while a hidden evaluator recomputes physical validity and scores the submitted report. We define the agent interface, tool contract, evidence log, and risk-sensitive metrics, including submitted recall, evidence-backed recall, found recall, false-safe penalties, severity regret, residual violation score, action cost, tool-use efficiency, and workflow diagnostics.
The validation budget is B=80 ... The report size is m=20 ... the top 5% of cases by severity
RetailBench: Benchmarking long horizon reasoning and coherent decision making of LLM agents in realistic retail environments
Goals to track manage pricing across the store's product assortment; manage replenishment and supplier selection to keep inventory stocked; manage shelf assortment and inventory aging; respond appropriately to customer feedback and external events; maintain solvency (cash-flow constraints) while maximizing net worth/sales over a long simulated horizon
Must keep track of The agent must track pricing, inventory levels/aging, supplier relationships, customer feedback, external events, and its own cash-flow position day by day across the 180-day (or longer) simulated run, since only a small subset of evaluated agents survive the full evaluation horizon.
structure: sequential-chain · goals arrive: given-up-front · count: 7 contemporary LLMs evaluated under repr · subgoal credit: yes · 6 citations · unit: simulated-days
Horizon180-day evaluation horizon; the simulator supports thousand-day-scale simulations
the paper’s own words
RetailBench models retail management as a partially observable decision process and is designed to support thousand-day-scale simulations. In this environment, agents must manage pricing, replenishment, supplier selection, shelf assortment, inventory aging, customer feedback, external events, and cash-flow constraints. We evaluate seven contemporary LLMs under representative agent frameworks over a 180-day evaluation horizon and compare them with a privileged oracle policy.
RiskWebWorld: A Realistic Interactive Benchmark for GUI Agents in E-commerce Risk Management
Goals to track investigate flagged e-commerce risk cases across multiple verification sub-steps on production risk-control pipelines; operate GUI actions on uncooperative websites subject to partial environmental hijacking; complete each of 1,513 tasks spanning 8 core risk-control domains; succeed at long-horizon professional risk-investigation workflows
Must keep track of The agent must track evidence and verification state gathered across multiple sub-steps of a risk investigation on an uncooperative website, while detecting and coping with partial environmental hijacking attempts, across long-horizon professional tasks.
structure: sequential-chain · goals arrive: given-up-front · count: 1,513 tasks sourced from production risk · subgoal credit: no · 0 citations · also Web & GUI agents
Horizonthe paper states none
the paper’s own words
Our evaluation across diverse models reveals a dramatic capability gap: top-tier generalist models achieve 49.1% success, while specialized open-weights GUI models lag at near-total failure.
This highlights that foundation model scale currently matters more than zero-shot interface grounding in long-horizon professional tasks.
SaaS-Bench: Can Computer-Use Agents Leverage Real-World SaaS to Solve Professional Workflows?
Goals to track navigate and operate real, deployed SaaS systems to complete a professional workflow; coordinate state and context across multiple applications within the same workflow; apply domain-specific knowledge correctly within the SaaS system; recover from errors and maintain progress over a long-horizon task
Must keep track of The agent must maintain state and context across multiple SaaS applications over long-horizon execution, tracking partial progress against weighted verification checkpoints, since 'agents become stuck acquiring information or manipulating interfaces, confuse source and output devices, or terminate before all conditions are jointly satisfied.'
structure: DAG-with-precedence · goals arrive: given-up-front · count: 106 tasks grounded in realistic work sce · subgoal credit: yes · 3 citations · unit: actions
Horizonaverage of over 100 interaction steps per task
the paper’s own words
SaaS-Bench, a benchmark built on 23 deployable SaaS systems across six professional domains, containing 106 tasks grounded in realistic work scenarios. These tasks require long-horizon execution, cover both text-only and multimodal settings, and are evaluated with weighted verification checkpoints that measure strict task completion and partial progress.
SaaS-Bench introduces long-horizon tasks with an average of over 100 interaction steps
SpreadsheetBench 2: Evaluating Agents on End-to-End Business Spreadsheet Workflows
Goals to track generate new spreadsheet content/formulas correctly across a large multi-sheet workbook; debug existing incorrect formulas/content within the workbook; produce correct visualizations from the workbook's data
Must keep track of The agent must track and correctly propagate changes across an average of 11.8 interdependent worksheets requiring 593.5 cell modifications per task, correctly identifying target cells under a unified multi-turn agent scaffold.
structure: DAG-with-precedence · goals arrive: given-up-front · count: 321 tasks; each instance averages 11.8 w · subgoal credit: yes · 4 citations · unit: other:cell-modifications
Horizoneach task instance averages 11.8 worksheets and requires 593.5 cell modifications
the paper’s own words
The benchmark contains 321 tasks; each instance averages 11.8 worksheets and requires 593.5 cell modifications, reflecting large multi-sheet workbooks with cross-sheet dependencies.
SupChain-Bench: Benchmarking Large Language Models for Real-World Supply Chain Management
Goals to track orchestrate long-horizon, multi-step supply-chain tool use grounded in standard operating procedures (SOPs); correctly apply supply-chain domain knowledge across a sequence of dependent tool calls
Must keep track of Agent must track SOP-grounded procedural state across a long-horizon sequence of dependent tool calls in order to correctly complete supply-chain management workflows.
structure: sequential-chain · goals arrive: given-up-front · subgoal credit: no · 0 citations
Horizonthe paper states none
the paper’s own words
supply chain workflows require reliable long-horizon, multi-step orchestration grounded in domain-specific procedures, which remains challenging for current models... we introduce SupChain-Bench, a unified real-world benchmark that assesses both supply chain domain knowledge and long-horizon tool-based orchestration grounded in standard operating procedures (SOPs).
Synthetic Computers at Scale for Long-Horizon Productivity Simulation
Goals to track complete each of multiple professional deliverables comprising a computer-specific productivity objective; navigate the synthetic computer's filesystem to ground actions in the user's actual context; coordinate with simulated collaborators as needed to complete the objective; sustain progress toward an objective requiring about a month of simulated human work
Must keep track of The acting agent must track the evolving state of a realistic folder hierarchy and content-rich artifacts (documents, spreadsheets, presentations) across a run spanning over 2,000 turns, coordinating with simulated collaborators toward multiple professional deliverables.
structure: open-ended · goals arrive: self-generated-by-agent · count: 1,000 synthetic computers; each run requ · subgoal credit: no · 1 citations · unit: turns
Horizon>2,000 turns on average (each run also requiring over 8 hours of agent runtime)
the paper’s own words
one agent creates productivity objectives that are specific to the computer's user and require multiple professional deliverables and about a month of human work; another agent then acts as that user and keeps working across the computer -- for example, navigating the filesystem for grounding, coordinating with simulated collaborators, and producing professional artifacts -- until these objectives are completed. ... each run requires over 8 hours of agent runtime and spans more than 2,000 turns on average.
One Success Isn't Reliability: Thinkingbox, a Sandbox and Benchmark for Agents in Stateful Business Workflows
Goals to track gather missing information over multiple turns before acting; follow domain-specific policies for each of 507 policy-conditioned workflows; coordinate dependent tools correctly; realize exactly the correct persistent backend state transition without collateral effects
Must keep track of Agent must gather missing information across multiple turns, track applicable domain policies, coordinate dependent tool calls, and verify the resulting persistent backend state matches exactly the required final state with no collateral effects.
structure: DAG-with-precedence · goals arrive: given-up-front · count: 507 policy-conditioned workflows across · subgoal credit: no · 0 citations · also Information seeking & deep research
Horizonthe paper states none
the paper’s own words
Each attempt is evaluated by task-specific executable checks that accept valid trajectories while rejecting wrong, missing, or extra effects; designated tasks additionally check required properties of the final response... the strongest model Claude Opus 5 achieves 66.50% pass@1, but only 47.53% pass^20.
Thinkingbox-bench contains 507 policy-conditioned workflows across numerous scenarios, including retail, hospitality, auto insurance, neobank internal IT, and consulting IT/HR support.
Benchmarking Agents in Insurance Underwriting Environments
Goals to track gather information carefully from noisy tool interfaces and imperfect simulated users during an underwriting conversation; apply proprietary business/domain knowledge correctly rather than hallucinating it; reach a final underwriting decision consistent with the accumulated evidence gathered across the conversation
Must keep track of The agent must track what information it has already gathered (and from where), the reliability of that information given noisy tool interfaces, and how it should update its evolving underwriting assessment as more evidence accumulates across the conversation.
structure: sequential-chain · goals arrive: given-up-front · subgoal credit: yes · 0 citations · also Information seeking & deep research · unit: turns
Horizonaverage of 3-7 steps of required reasoning and tool use, with a total of 10-20 conversational turns
the paper’s own words
We present Underwrite, an expert-first, multi-turn insurance underwriting benchmark designed in close collaboration with domain experts to capture real-world enterprise challenges. Underwrite introduces critical realism factors often absent in current benchmarks: proprietary business knowledge, noisy tool interfaces, and imperfect simulated users requiring careful information gathering.
We aimed for an average of 3-7 steps of required reasoning and tool use, with a total of 10-20 conversational turns.
Workflow-GYM: Towards Long-Horizon Evaluation of Computer-use Agentic tasks in Real-World Professional Fields
Goals to track autonomously operate domain-specific professional software GUIs to accomplish economically valuable work; complete long-horizon, multi-stage professional workflows end-to-end; maintain workflow consistency across stages without omission, error propagation, or objective drift
Must keep track of The agent must track its position and completed stages within a long-horizon, multi-stage professional GUI workflow, avoiding stage omission, propagating errors, or drifting from the original objective across the workflow's duration.
structure: sequential-chain · goals arrive: given-up-front · subgoal credit: no · 2 citations
Horizonthe paper states none
the paper’s own words
Through extensive experiments on state-of-the-art models, we find that even the strongest models achieve only slightly above 30% success rates, highlighting that professional long-horizon GUI workflows remain highly challenging for current GUI agents.
existing benchmarks rarely evaluate whether agents can operate graphical user interfaces to complete long-horizon, high-value professional workflows across diverse domains
Workspace-Bench 1.0: Benchmarking AI Agents on Workspace Tasks with Large-Scale File Dependencies
Goals to track identify, reason over, exploit, and update explicit and implicit dependencies among heterogeneous files in a worker's workspace; satisfy each of a task's own file-dependency-graph-derived rubrics via cross-file retrieval, contextual reasoning, and adaptive decision-making
Must keep track of Agent must track which of up to 20,476 files across 74 file types are relevant, their dependency relationships, and adaptively retrieve/reason across files to satisfy each task's rubrics.
structure: DAG-with-precedence · goals arrive: given-up-front · count: 388 tasks, each with its own file depend · subgoal credit: yes · 10 citations
Horizonthe paper states none
the paper’s own words
We construct realistic workspaces with 5 worker profiles, 74 file types, 20,476 files (up to 20GB) and curate 388 tasks, each with its own file dependency graph, evaluated across 7,399 total rubrics that require cross-file retrieval, contextual reasoning, and adaptive decision-making.
We further provide Workspace-Bench-Lite, a 100-task subset that preserves the benchmark distribution while reducing evaluation costs by about 70%.
World of Workflows: a Benchmark for Bringing World Models to Enterprise Systems
Goals to track complete constrained agentic tasks within a ServiceNow environment governed by 4,000+ business rules and 55 active hidden workflows; predict cascading side effects of actions across interconnected databases; mentally simulate hidden state transitions to avoid silent constraint violations under limited observability
Must keep track of Agent must mentally simulate hidden state transitions and predict cascading side effects across interconnected databases to bridge the observability gap, since high-fidelity feedback is often unavailable.
structure: DAG-with-precedence · goals arrive: given-up-front · count: 234 tasks (WoW-bench); 4,000+ business r · subgoal credit: no · 7 citations
Horizonthe paper states none
the paper’s own words
We introduce World of Workflows (WoW), a realistic ServiceNow-based environment incorporating 4,000+ business rules and 55 active workflows embedded in the system, alongside WoW-bench, a benchmark of 234 tasks evaluating constrained agentic task completion and enterprise dynamics modeling capabilities.
YC-Bench: Benchmarking AI Agents for Long-Term Planning and Consistent Execution
Goals to track manage employees within a simulated startup; select task contracts to pursue under uncertainty; maintain profitability against adversarial clients and growing payroll; detect and avoid bankruptcy-inducing failure modes (e.g. adversarial-client mismanagement, over-parallelization) over a one-year run
Must keep track of The agent must persist information across context truncation via a scratchpad (the strongest predictor of success), tracking employees, contracts, cash reserves, and adversarial-client risk across hundreds of turns spanning a simulated year.
structure: open-ended · goals arrive: given-up-front · subgoal credit: no · 8 citations · unit: turns
Horizonhundreds of turns (over a simulated one-year horizon)
the paper’s own words
we task an agent with running a simulated startup over a one-year horizon spanning hundreds of turns. The agent must manage employees, select task contracts, and maintain profitability in a partially observable environment where adversarial clients and growing payroll create compounding consequences for poor decisions.
we introduce YC-Bench, a benchmark that evaluates these capabilities by tasking an agent with running a simulated startup over a one-year horizon spanning hundreds of turns.
What Could the Agent See at 19:05? Generating Temporal Enterprise Scenarios from Real Research and Replaying Them to Evaluate Agents
Goals to track answer questions correctly relative to what data existed and who could see it at a specific queried moment; reason about a persona-driven, temporally-evolving enterprise world spanning many apps; avoid leaking future/hidden record state when reasoning about an earlier moment
Must keep track of The agent being evaluated must reason correctly about what data existed and who could see it at a specific queried moment, without conflating it with earlier or later states of the same records across multiple apps.
structure: DAG-with-precedence · goals arrive: emitted-by-environment-over-time · subgoal credit: yes · 0 citations
Horizonthe paper states none
the paper’s own words
Our system closes two gaps at once: it generates a realistic, persona-driven, temporally-evolving enterprise world from real research, and replays that world at any chosen moment to evaluate any pluggable agent.
Behavioral Consistency Validation for LLM Agents: An Analysis of Trading-Style Switching through Stock-Market Simulation
Goals to track trade under a currently assigned style (fundamental or technical) each trading day; periodically (every 10 trading days) reassess and potentially switch trading style based on four behavioral-finance drivers: loss aversion, herding, wealth differentiation, price misalignment; keep style-switching behavior consistent with real-world behavioral-finance theory over the whole simulated year
Must keep track of Agent must process daily price-volume data, retain long-term personality traits (the four behavioral-finance drivers) set at initialization, and track its own accumulated wealth/strategy history to decide whether to switch trading style every 10 days.
structure: sequential-chain · goals arrive: given-up-front · subgoal credit: unclear · 4 citations · unit: simulated-days
Horizonyear-long (reassessed every 10 trading days)
the paper’s own words
In year-long simulations, agents process daily price-volume data, trade under a designated style, and reassess their strategy every 10 trading days.
AI-Trader: Benchmarking Autonomous Agents in Real-Time Financial Markets
Goals to track independently search, verify, and synthesize live market information given only minimal initial context; make live trading decisions (buy/sell/hold) across U.S. stocks, A-shares, and cryptocurrencies at multiple trading granularities; manage risk and sustain positive returns over a continuous live-trading period
Must keep track of Agent must track evolving live market data it has independently searched/verified, current positions/risk exposure across three markets and multiple trading frequencies, and running returns over the continuous live-trading period.
structure: sequential-chain · goals arrive: emitted-by-environment-over-time · subgoal credit: yes · 19 citations
Horizonthe paper states none
the paper’s own words
Our benchmark implements a revolutionary fully autonomous minimal information paradigm where agents receive only essential context and must independently search, verify, and synthesize live market information without human intervention.
AI-Trader spans three major financial markets: U.S. stocks, A-shares, and cryptocurrencies, with multiple trading granularities to simulate live financial environments.
Agent Trading Arena: A Study on Numerical Understanding in LLM-Based Agents
Goals to track make sequential buy/sell trading decisions that directly affect and are affected by shared market prices; compete against other LLM-based agents in a zero-sum stock market; maximize trading performance/returns, especially under high volatility; correctly perform numerical reasoning over price data or chart-based visualizations
Must keep track of Each agent must track its own portfolio/capital, the current market price state shaped by all agents' recent trades, and historical price patterns, updating its numerical reasoning as the shared market evolves turn to turn.
structure: open-ended · goals arrive: given-up-front · subgoal credit: no · 12 citations
Horizonthe paper states none
the paper’s own words
we present the Agent Trading Arena, a virtual zero-sum stock market in which LLM-based agents engage in competitive multi-agent trading and directly impact price dynamics.
Existing approaches are limited to historical backtesting, where trading actions cannot influence market prices and agents train only on static data.
AssetOpsBench: A Real-World Evaluation Benchmark for AI-Driven Task Automation in Industrial Asset Management
Goals to track orchestrate the correct domain-specific agent(s) (from a catalog of four) to answer a natural-language industrial-operations query; correctly complete condition-monitoring and maintenance-scheduling workflow steps grounded in a simulated IoT environment
Must keep track of The agent must track intermediate tool/agent outputs across a multi-step think-act-observe loop, orchestrate the four domain-specific agents appropriately, and maintain consistency with the simulated CouchDB-backed IoT environment's evolving state.
structure: DAG-with-precedence · goals arrive: given-up-front · count: 140+ human-authored queries; Plan-Execut · subgoal credit: yes · 17 citations · unit: agent-steps
HorizonPlan-Execute agents complete most tasks in approximately 2.6-4.4 steps; Agent-As-Tool agents typically require approximately 4-6+ steps due to its iterative think-act-obs
the paper’s own words
AssetOpsBench provides a multimodal ecosystem comprising a catalog of four domain-specific agents, a curated dataset of 140+ human-authored natural-language queries grounded in real industrial scenarios, and a simulated, CouchDB-backed IoT environment. We introduce an automated evaluation framework that uses three key metrics to analyze architectural trade-offs between the Agent-As-Tool and Plan-Execute paradigms, along with a systematic procedure for the automated discovery of emerging failure modes.
Plan-Execute is consistently more step-efficient: most models complete tasks in ≈2.6-4.4 steps, whereas Agent-As-Tool typically requires ≈4-6+ steps due to its iterative think-act-observe loop.
Automating Complex Document Workflows via Stepwise and Rollback-Enabled Operation Orchestration
Goals to track execute a sequence of interdependent, user-specified document-editing instructions within a session; keep the execution trajectory aligned with evolving document state and user intent across the whole session; recover from failed API calls/arguments via rollback at both the argument and API level
Must keep track of System must track the evolving document state after each API action, the remaining interdependent instructions in the session, and whether any prior action needs to be rolled back (at argument or API level) to stay aligned with user intent.
structure: DAG-with-precedence · goals arrive: given-up-front · count: 1,708 human-annotated instructions acros · subgoal credit: yes · 0 citations · unit: actions
Horizon1,708 instructions across 250 sessions (~6.8 instructions/session on average)
the paper’s own words
AutoDW achieves 90% and 62% completion rates on instruction- and session-level tasks, respectively, outperforming strong baselines by 40% and 76%.
we construct a comprehensive benchmark of 250 sessions and 1,708 human-annotated instructions, reflecting realistic document processing scenarios with interdependent instructions
CP-Env: Evaluating Large Language Models on Clinical Pathways in a Controllable Hospital Environment
Goals to track triage a patient correctly to determine care pathway; consult the appropriate specialist given patient information; order and interpret diagnostic tests as needed; participate appropriately in multidisciplinary team meetings; complete a patient's full journey across branching, long-horizon clinical-pathway stages
Must keep track of The agent must track patient information, diagnostic findings, and evolving care-pathway state across branching stages from triage through specialist consultation, diagnostic testing, and multidisciplinary team meetings.
structure: DAG-with-precedence · goals arrive: mixed:overall-pathway-structure-given-up-front · subgoal credit: yes · 6 citations · also Healthcare & clinical
Horizonthe paper states none
the paper’s own words
CP-Env simulates a hospital ecosystem with patient and physician agents, constructing scenarios ranging from triage and specialist consultation to diagnostic testing and multidisciplinary team meetings for agent interaction. Following real hospital adaptive flow of healthcare, it enables branching, long-horizon task execution.
CRMArena-Pro: Holistic Assessment of LLM Agents Across Diverse Business Scenarios and Interactions
Goals to track complete each of nineteen expert-validated CRM tasks across sales, service, and configure-price-quote (CPQ) processes; sustain correct behavior across multi-turn interactions guided by diverse personas; maintain confidentiality awareness throughout the interaction, for both B2B and B2C scenarios
Must keep track of The agent must track persona-specific context, confidentiality constraints, and workflow state across multi-turn interactions, since single-turn success (58%) drops substantially to about 35% once genuine multi-turn tracking is required.
structure: sequential-chain · goals arrive: given-up-front · count: nineteen expert-validated tasks across s · subgoal credit: yes · 42 citations
Horizonthe paper states none
the paper’s own words
It distinctively incorporates multi-turn interactions guided by diverse personas and robust confidentiality awareness assessments... Experiments reveal leading LLM agents achieve only around 58% single-turn success on CRMArena-Pro, with performance dropping significantly to approximately 35% in multi-turn settings. While Workflow Execution proves more tractable for top agents (over 83% single-turn success), other evaluated business skills present greater challenges.
DRBench: A Realistic Benchmark for Enterprise Deep Research
Goals to track identify supporting facts for a multi-step research query from both the public web and a private company knowledge base; synthesize facts drawn from heterogeneous enterprise sources (productivity software, cloud file systems, emails, chat, web) into one answer; produce a coherent, well-structured final report grounded in the retrieved facts
Must keep track of The agent must track which facts it has found in which source (public web vs. private enterprise knowledge base), maintain factual accuracy across sources, and assemble a coherent report structure from these accumulated facts.
structure: hierarchical · goals arrive: given-up-front · count: 100 deep research tasks across 10 domain · subgoal credit: yes · 15 citations
Horizonthe paper states none
the paper’s own words
DRBench evaluates agents on multi-step queries ... that require identifying supporting facts from both the public web and private company knowledge base. Each task is grounded in realistic user personas and enterprise context, spanning a heterogeneous search space that includes productivity software, cloud file systems, emails, chat conversations, and the open web... agents are evaluated on their ability to recall relevant insights, maintain factual accuracy, and produce coherent, well-structured reports.
We release 100 deep research tasks across 10 domains, such as Sales, Cybersecurity, and Compliance.
DevNous: An LLM-Based Multi-Agent System for Grounding IT Project Management in Unstructured Conversation
Goals to track identify actionable intents from informal, unstructured team-chat dialogue; manage stateful, multi-turn administrative workflows (task formalization, progress-summary synthesis) grounded in that dialogue
Must keep track of The agent must identify actionable intents across informal chat and maintain stateful workflows (task formalization, progress tracking) across a benchmark of 160 conversational turns, correctly matching a multi-label ground truth per turn.
structure: sequential-chain · goals arrive: emitted-by-environment-over-time · count: 160 realistic, interactive conversationa · subgoal credit: yes · 0 citations · also Multi-agent organisations & societies · unit: turns
Horizona new benchmark of 160 realistic, interactive conversational turns
the paper’s own words
We introduce DevNous, a Large Language Model-based (LLM) multi-agent expert system, to automate this unstructured-to-structured translation process. DevNous integrates directly into team chat environments, identifying actionable intents from informal dialogue and managing stateful, multi-turn workflows for core administrative tasks like automated task formalization and progress summary synthesis. To quantitatively evaluate the system, we introduce a new benchmark of 160 realistic, interactive conversational turns.
EcomBench: Towards Holistic Evaluation of Foundation Agents in E-commerce
Goals to track retrieve deep, possibly multi-hop information relevant to a real e-commerce user demand; perform multi-step reasoning across e-commerce-domain data; integrate knowledge from multiple sources to resolve a task; correctly complete tasks across three graded difficulty levels
Must keep track of The agent must track partial evidence gathered from deep retrieval and multiple knowledge sources across many action steps before it can complete a higher-difficulty task.
structure: hierarchical · goals arrive: given-up-front · count: three difficulty levels; task categories · subgoal credit: yes · 7 citations
Horizonthe paper states none
the paper’s own words
It covers multiple task categories within e-commerce scenarios and defines three difficulty levels that evaluate agents on key capabilities such as deep information retrieval, multi-step reasoning, and cross-source knowledge integration.
Level 3 tasks cannot be solved in just a few action steps
Finch: Benchmarking Finance & Accounting across Spreadsheet-Centric Enterprise Workflows
Goals to track interleave data entry, structuring, formatting, web search, cross-file retrieval, calculation, modeling, validation, translation, visualization, and reporting subtasks into one finance/accounting workflow; correctly complete each of the linked tasks that compose one of 172 composite workflows
Must keep track of Agent must track state across many interlinked spreadsheets, PDFs, and artifacts (27 million spreadsheet cells total) while interleaving retrieval, calculation, modeling, validation, and reporting sub-steps of one composite workflow.
structure: DAG-with-precedence · goals arrive: given-up-front · count: 172 composite workflows with 384 tasks, · subgoal credit: yes · 9 citations · unit: wall-clock-minutes
Horizon16.8 (GPT-5.1 Pro average)
the paper’s own words
This yields 172 composite workflows with 384 tasks, involving 1,710 spreadsheets with 27 million cells, along with PDFs and other artifacts, capturing the intrinsically messy, long-horizon, knowledge-intensive, and collaborative nature of real-world enterprise work.
Under human evaluation, GPT-5.1 Pro spends an average of 16.8 minutes per workflow yet passes only 38.4% of workflows.
GraphicBench: A Planning Benchmark for Graphic Design with Language Agents
Goals to track produce a workflow plan satisfying explicit design constraints stated in a user query; also satisfy implicit commonsense design constraints not stated by the user; select the correct action from 46 available tools at each workflow step; coordinate outputs across three design experts without violating global dependencies
Must keep track of The agent must track explicit and implicit design constraints, the evolving multi-step workflow plan across three design experts, and which of 46 actions remains valid/appropriate at each step.
structure: DAG-with-precedence · goals arrive: mixed:explicit-constraints-given-up-front-plus · count: 1,079 user queries; 46 available actions · subgoal credit: yes · 2 citations · also Web & GUI agents
Horizonthe paper states none
the paper’s own words
We introduce GraphicBench, a new planning benchmark for graphic design that covers 1,079 user queries and input images across four design types. We further present GraphicTown, an LLM agent framework with three design experts and 46 actions (tools) to choose from for executing each step of the planned workflows in web environments.
HiMA-Ecom: Enabling Joint Training of Hierarchical Multi-Agent E-commerce Assistants
Goals to track master agent coordinates multiple specialized sub-agents in an e-commerce workflow; each sub-agent pursues a functionally distinct role-specific objective (e.g. recall domain knowledge, execute a service function); jointly optimize system-level behavior via multi-agent reinforcement learning
Must keep track of The master agent must track which specialized sub-agent is responsible for which functional sub-task and how sub-agent outputs compose into overall system behavior, using a collaboratively updated memory across training.
structure: hierarchical · goals arrive: given-up-front · subgoal credit: no · 8 citations · also Multi-agent organisations & societies · unit: other:not-stated
Horizonthe paper states none
the paper’s own words
HiMA-Ecom contains 22.8K instances, including agent-specific supervised fine-tuning samples with memory and system-level input-output pairs for joint multi-agent reinforcement learning. ... a master agent coordinates multiple specialized sub-agents
MEMTRACK: Evaluating Long-Term Memory and State Tracking in Multi-Platform Dynamic Agent Environments
Goals to track acquire relevant facts scattered across asynchronous cross-platform events (Slack/Linear/Git); select the correct, currently-valid fact when multiple conflicting/noisy candidates exist; resolve conflicts between contradictory or stale cross-referring information over time
Must keep track of The agent must acquire, select, and reconcile facts from a chronologically platform-interleaved timeline spanning Slack, Linear, and Git, handling noisy, conflicting, and cross-referring information plus codebase/file-system exploration, across long horizons.
structure: DAG-with-precedence · goals arrive: emitted-by-environment-over-time · subgoal credit: yes · 12 citations · also Personal assistants & long-term memory
Horizonthe paper states none
the paper’s own words
Consequently, our benchmark tests memory capabilities such as acquistion, selection and conflict resolution... We introduce pertinent metrics for Correctness, Efficiency, and Redundancy that capture the effectiveness of memory mechanisms beyond simple QA performance.
Each benchmark instance provides a chronologically platform-interleaved timeline, with noisy, conflicting, cross-referring information as well as potential codebase/file-system comprehension and exploration.
Mini Amusement Parks (MAPs): A Testbed for Modelling Business Decisions
Goals to track maximize the amusement park's overall value over the given time horizon; make daily operational decisions (building rides/shops, hiring staff, setting a research agenda) that keep the business viable under sparse, stochastic feedback; reason over the park's spatial layout while planning these decisions
Must keep track of The agent must track the park's spatial layout, built rides/shops, staffing, accumulated (sparse) experience about environment dynamics, and overall park value across a 50-100 day episode, planning under uncertainty at each of many daily decision points.
structure: sequential-chain · goals arrive: given-up-front · count: episodes span a 50-day horizon on easy m · subgoal credit: yes · 2 citations · unit: simulated-days
Horizonepisodes span a 50-day horizon on easy mode, extended to 100 days on medium mode
the paper’s own words
To this end, we introduce Mini Amusement Parks (MAPs), an amusement-park simulator designed to evaluate an agent's ability to model its environment, anticipate long-term consequences under uncertainty, and strategically operate a complex business. We provide expert human performance and a comprehensive evaluation of state-of-the-art agents, finding experts outperform these systems by 11.4x on easy mode and 15.3x on medium mode.
the horizon is extended from 50 to 100
OdysseyBench: Evaluating LLM Agents on Long-Horizon Complex Office Application Workflows
Goals to track identify essential information buried in long-horizon interaction histories; perform multi-step reasoning/actions across Word, Excel, PDF, Email, and Calendar applications; complete real-world-derived tasks (OdysseyBench+) or newly synthesized complex tasks (OdysseyBench-Neo)
Must keep track of Agent must retain and retrieve essential facts from long-horizon interaction histories spanning multiple applications in order to correctly execute later multi-step, cross-application actions.
structure: sequential-chain · goals arrive: given-up-front · count: 300 tasks (OdysseyBench+); 302 tasks (Od · subgoal credit: unclear · 51 citations
Horizonthe paper states none
the paper’s own words
Our benchmark comprises two complementary splits: OdysseyBench+ with 300 tasks derived from real-world use cases, and OdysseyBench-Neo with 302 newly synthesized complex tasks. Each task requires agent to identify essential information from long-horizon interaction histories and perform multi-step reasoning across various applications.
PPTArena: A Benchmark for PowerPoint Editing
Goals to track apply each of a deck's human-curated edits correctly from natural-language instructions; maintain layout-sensitive and cross-slide consistency across the whole deck while editing; plan and verify edit sequences via an iterative plan-edit-check loop (for the PPTPilot agent)
Must keep track of Agent must track the deck's current structural/visual state after each edit, verify each edit against a ground-truth rubric, and maintain deck-wide consistency (styles, cross-slide references) across the full sequence of edits.
structure: sequential-chain · goals arrive: given-up-front · count: over 1,300 human-curated edits across 10 · subgoal credit: yes · 7 citations · unit: actions
Horizon~13 edits per deck (1,300+ edits across 100 decks)
the paper’s own words
PPTArena features 100 decks with over 1,300 human-curated edits across 2,125 slides, spanning text, charts, animations, and professional master styles.
ProSoftArena: Benchmarking Hierarchical Capabilities of Multimodal Agents in Professional Software Environments
Goals to track operate professional software tools according to a hierarchical capability level (L1-L3); complete realistic work/research tasks spanning 6 disciplines and 13 core professional applications; coordinate across multiple professional software applications for L3 multi-software workflows
Must keep track of The agent must track its progress within and across professional-software applications as required by the task's capability level, since L3 tasks require coordinating state across multiple applications rather than operating one application in isolation.
structure: hierarchical · goals arrive: given-up-front · count: 436 realistic work and research tasks sp · subgoal credit: yes · 5 citations · unit: actions
Horizonhuman execution steps grow from an average of 5.1 steps (14.8 seconds) at L1 to 86.9 steps (506.8 seconds) at L3
the paper’s own words
We establish the first capability hierarchy tailored to agent use of professional software and construct a benchmark of 436 realistic work and research tasks spanning 6 disciplines and 13 core professional applications... Extensive experiments show that even the best-performing agent attains only a 24.4% success rate on L2 tasks and completely fails on L3 multi-software workflow.
human execution steps and time... from an average of 5.1 steps and 14.8 seconds at L1 to 86.9 steps and 506.8 seconds at L3
REALM-Bench: A Benchmark for Evaluating Multi-Agent Systems on Real-world, Dynamic Planning and Scheduling Tasks
Goals to track solve each of 14 real-world planning/scheduling problems with multiple parallel planning threads; maintain feasibility across inter-agent dependencies as the problem scales in complexity; adapt schedules in real time when unexpected disruptions arrive
Must keep track of Agents must track the state of parallel planning threads, inter-dependencies between agents/threads, and disruption events in order to replan in real time.
structure: DAG-with-precedence · goals arrive: mixed:given-up-front-problems-plus-injected-mi · count: 14 planning and scheduling problems · subgoal credit: yes · 12 citations · also Multi-agent organisations & societies
Horizonthe paper states none
the paper’s own words
The suite encompasses 14 designed planning and scheduling problems that progress from basic to highly complex, incorporating key aspects such as multi-agent coordination, inter-agent dependencies, and dynamic environmental disruptions.
Each problem can be scaled along three dimensions: the number of parallel planning threads, the complexity of inter-dependencies, and the frequency of unexpected disruptions requiring real-time adaptation.
Remote Labor Index: Measuring AI Automation of Remote Work
Goals to track complete a whole real, economically valuable freelance/remote-work project end to end; achieve a level of automation comparable to what a human freelancer would deliver for that project
Must keep track of Agent must track progress toward completing an entire multi-part real-world project (not a single isolated action) to be credited with automating that unit of remote labor.
structure: hierarchical · goals arrive: given-up-front · subgoal credit: unclear · 21 citations
Horizonthe paper states none
the paper’s own words
we introduce the Remote Labor Index (RLI), a broadly multi-sector benchmark comprising real-world, economically valuable projects designed to evaluate end-to-end agent performance in practical settings.
AIs have made rapid progress on research-oriented benchmarks of knowledge and reasoning, but it remains unclear how these gains translate into economic value and automation.
SCUBA: Salesforce Computer Use Benchmark
Goals to track navigate a specific enterprise software UI (Salesforce) to accomplish a CRM task; manipulate data records and automate workflows within the Salesforce platform; retrieve information and troubleshoot issues as part of a realistic CRM task; generalize across three personas (platform administrators, sales representatives, service agents) each with distinct task profiles
Must keep track of The agent must track its progress toward fine-grained milestones within each Salesforce sandbox task (UI navigation state, data changes made, workflow steps completed) to receive interpretable milestone-progress credit.
structure: sequential-chain · goals arrive: given-up-front · count: 300 task instances derived from real use · subgoal credit: yes · 7 citations · also Information seeking & deep research
Horizonthe paper states none
the paper’s own words
SCUBA operates in Salesforce sandbox environments with support for parallel execution and fine-grained evaluation metrics to capture milestone progress. We benchmark a diverse set of agents under both zero-shot and demonstration-augmented settings.
SCUBA contains 300 task instances derived from real user interviews, spanning three primary personas, platform administrators, sales representatives, and service agents.
SOP-Bench: Complex Industrial SOPs for Evaluating LLM Agents
Goals to track correctly execute each step of a complex, multi-step Standard Operating Procedure; orchestrate the correct tools/APIs at each SOP step; produce ground-truth outputs matching the SOP's authored specification across 12 business domains
Must keep track of The agent must track its position within a multi-step SOP, select the correct tool from a registry that may contain many irrelevant tools, and maintain consistency with the SOP's ground-truth interface and outputs across the whole procedure.
structure: sequential-chain · goals arrive: given-up-front · count: 2,000+ tasks across 12 business domains · subgoal credit: no · 20 citations · unit: other:not-stated
Horizonthe paper states none
the paper’s own words
LLM-based agents struggle to execute complex, multi-step Standard Operating Procedures (SOPs) that are fundamental to industrial automation. ... We introduce SOP-Bench, a benchmark of 2,000+ tasks from human expert-authored SOPs across 12 business domains ... yielding realistic tasks with executable interfaces and ground-truth outputs.
StockBench: Can LLM Agents Trade Stocks Profitably In Real-world Markets?
Goals to track make a sequential daily buy/sell/hold decision based on incoming market signals (prices, fundamentals, news); maximize cumulative return across the whole multi-month trading period; manage risk, minimizing maximum drawdown and maintaining a strong Sortino ratio
Must keep track of Agent must track its current portfolio position, cumulative return, and risk exposure (drawdown) as it processes a new daily market signal (prices, fundamentals, news) each step of a multi-month simulated trading run.
structure: sequential-chain · goals arrive: emitted-by-environment-over-time · subgoal credit: yes · 33 citations · unit: simulated-days
Horizonmulti-month (exact day/month count not given)
the paper’s own words
Agents receive daily market signals -- including prices, fundamentals, and news -- and make sequential buy, sell, or hold decisions.
STOCKBENCH, a contamination-free benchmark designed to evaluate LLM agents in realistic, multi-month stock trading environments. Agents receive daily market signals... and make sequential buy, sell, or hold decisions.
StaffPro: an LLM Agent for Joint Staffing and Profiling
Goals to track assign and schedule tasks to workers (staffing), forming teams as needed; continuously estimate workers' latent skills, preferences, and other attributes from unstructured feedback (profiling); optimize staffing performance over time as profiling estimates improve via an ongoing human-agent feedback loop
Must keep track of StaffPro must track each worker's evolving latent-attribute profile (estimated from ongoing human feedback) and the current staffing/schedule state, updating both jointly over the 'life-long' profiling horizon.
structure: DAG-with-precedence · goals arrive: emitted-by-environment-over-time · subgoal credit: yes · 2 citations
Horizonthe paper states none
the paper’s own words
By analyzing human feedback, our agent continuously estimates the latent features of workers, realizing life-long worker profiling and ensuring optimal staffing performance over time.
UpBench: A Dynamically Evolving Real-World Labor-Market Agentic Benchmark Framework Built for Human-Centric AI
Goals to track complete a real, verified client transaction job sourced from the Upwork labor marketplace; satisfy each of a job's detailed, expert-decomposed, verifiable acceptance criteria; follow instructions faithfully enough to earn positive fine-grained, per-criterion human expert feedback
Must keep track of Agent must track which of a job's detailed acceptance criteria it has satisfied, informed by expert freelancer decomposition, to produce a submission gradeable on fine-grained, per-criterion feedback rather than a binary pass/fail.
structure: set-of-independent · goals arrive: given-up-front · count: one job per task, each decomposed into d · subgoal credit: yes · 4 citations
Horizonthe paper states none
the paper’s own words
UpBench employs a rubric-based evaluation framework, in which expert freelancers decompose each job into detailed, verifiable acceptance criteria and assess AI submissions with per-criterion feedback. This structure enables fine-grained analysis of model strengths, weaknesses, and instruction-following fidelity beyond binary pass/fail metrics.
Each task corresponds to a verified client transaction, anchoring evaluation in genuine work activity and financial outcomes.
Vending-Bench: A Benchmark for Long-Term Coherence of Autonomous Agents
Goals to track balance inventory levels of vending-machine stock; place restocking orders from suppliers; set item prices; pay recurring daily fees; sustain a profitable long-running vending-machine business without derailing
Must keep track of The agent must continuously track inventory counts, cash/profit, outstanding orders and their delivery schedules, and daily fees across a run spanning over 20M tokens, since forgetting an order or misreading a schedule causes derailment.
structure: open-ended · goals arrive: given-up-front · subgoal credit: no · 67 citations · unit: other:tokens-per-run
Horizon>20M
the paper’s own words
Agents must balance inventories, place orders, set prices, and handle daily fees - tasks that are each simple but collectively, over long horizons (>20M tokens per run) stress an LLM's capacity for sustained, coherent decision-making.
VitaBench: Benchmarking LLM Agents with Versatile Interactive Tasks in Real-world Applications
Goals to track satisfy each of multiple real user requests combined into one cross-scenario task (food delivery, in-store consumption, online travel); reason across temporal and spatial dimensions while using a large tool set (66 tools); proactively clarify ambiguous instructions and track shifting user intent across a multi-turn conversation
Must keep track of Agent must track shifting user intent across a multi-turn conversation, temporal and spatial constraints, and the state of a large (66-tool) toolset spanning multiple life-serving domains simultaneously.
structure: set-of-independent · goals arrive: mixed:given-up-front-with-user-intent-shifting · count: 100 cross-scenario tasks (main results) · subgoal credit: unclear · 35 citations
Horizonthe paper states none
the paper’s own words
Each task is derived from multiple real user requests and requires agents to reason across temporal and spatial dimensions, utilize complex tool sets, proactively clarify ambiguous instructions, and track shifting user intent throughout multi-turn conversations.
yielding 100 cross-scenario tasks (main results) and 300 single-scenario tasks
Benchmarking LLM Agents for Wealth-Management Workflows
Goals to track complete each of 12 wealth-management task-pairs spanning retrieval, analysis, and synthesis/communication; satisfy explicit acceptance criteria under deterministic graders; operate correctly under both high- and low-autonomy task variants
Must keep track of Agent must track retrieved data, intermediate analysis results, and explicit acceptance criteria across the retrieval-analysis-synthesis pipeline, plus which autonomy variant (high vs. low) governs how independently it may act.
structure: sequential-chain · goals arrive: given-up-front · count: 12 task-pairs spanning retrieval, analys · subgoal credit: no · 0 citations
Horizonthe paper states none
the paper’s own words
We construct a benchmark of 12 task-pairs for wealth management assistants spanning retrieval, analysis, and synthesis/communication, with explicit acceptance criteria and deterministic graders.
This study introduces synthetic domain data, enriches colleague simulations, and prototypes an automatic task-generation pipeline.
Does It Tie Out? Towards Autonomous Legal Agents in Venture Capital
Goals to track verify that every security (shares, options, warrants) is supported by underlying legal documentation; verify that every issuance term (vesting schedules, acceleration triggers, transfer restrictions) is consistent across the dataroom; maintain strict evidence traceability while reconciling thousands of pages of legal documents; produce a deterministic, fully-reconciled capitalization table
Must keep track of The agent must track which securities and issuance terms have been verified against which supporting documents, maintaining strict evidence traceability across a growing dataroom (up to tens of thousands of pages) as it reconciles a scaling number of individual securities.
structure: set-of-independent · goals arrive: given-up-front · count: 184 to 1,292 individual securities requi · subgoal credit: yes · 0 citations · unit: agent-steps
Horizonworkload grows from approximately 2,700 atomic verification steps at Seed stage to nearly 8,000 steps at Series B stage
the paper’s own words
verifying that every security (for example, shares, options, warrants) and issuance term (for example, vesting schedules, acceleration triggers, transfer restrictions) is supported by large sets of underlying legal documentation. While LLMs continue to improve on legal benchmarks, specialized legal workflows, such as capitalization tie-out, remain out of reach even for strong agentic systems. The task requires multi-document reasoning, strict evidence traceability, and deterministic outputs.
Fig. 7 quantifies this burden by tracking the total number of atomic 'steps' executed by the counsel to complete the tie-out... We observe a near-tripling of workload, from approximately 2,700 steps at Seed to nearly 8,000 steps at Series B.