Agents' Last Exam
ALE benchmark evaluates AI on long-term, high-value real-world industry tasks; current top models achieve less than 1% success rate on hardest tasks.
Key Findings
Methodology
The Agents’ Last Exam (ALE) benchmark was developed through collaboration with over 250 industry experts, creating a task taxonomy covering 55 subdomains within 13 industry clusters. Tasks are sourced from real professional projects, ensuring representativeness and complexity. Multi-round expert reviews and automated quality controls verify authenticity and difficulty. The benchmark evaluates generalist AI agents—such as Claude Code and Codex—integrating multimodal perception, code execution, tool use, and long-term planning within a unified environment. Structured deliverables and milestone checks replace subjective human judgments, enabling objective, automatable validation. ALE emphasizes economic relevance, measuring not only task success but also industry impact, aiming to bridge the gap between benchmark performance and real-world deployment.
Key Results
- The strongest configuration tested—Codex combined with GPT-5.5—achieved 82% success on Terminal-Bench, but only around 50% on ALE’s easiest tasks and less than 10% on the hardest. Most mainstream models, including Claude Code, recorded near-zero pass rates at high difficulty levels, with an average full pass rate below 1%. This highlights the significant gap between current AI capabilities and industry-level long-term task performance.
- Performance varies considerably across industries; fields like electronics engineering and life sciences show relatively better results, but overall coverage remains limited. The experiments reveal that models struggle with multi-step, multimodal workflows, emphasizing the need for architectural and training innovations.
- Results demonstrate that models are far from matching human experts in complex, real-world tasks, especially those requiring planning, reasoning, and multimodal understanding. These findings underscore the importance of developing AI systems capable of sustained, industry-grade performance.
Significance
ALE serves as a pioneering benchmark targeting real-world, long-horizon industry tasks, addressing a critical gap in current AI evaluation frameworks. Unlike traditional benchmarks focused on short-term or synthetic tasks, ALE emphasizes authentic workflows with economic value, providing a meaningful measure of AI readiness for industrial deployment. Its collaborative design with industry experts ensures relevance and realism, fostering progress toward AI systems that can operate reliably in manufacturing, healthcare, finance, and other sectors. By establishing a rigorous, scalable, and evolving evaluation platform, ALE aims to accelerate AI research from laboratory success to tangible economic impact, ultimately contributing to productivity gains and industry transformation.
Technical Contribution
This work introduces a novel evaluation paradigm that integrates real industry workflows into a structured, verifiable benchmark. Key innovations include sourcing tasks directly from professional projects, multi-stage expert review, automated validation of heterogeneous outputs, and multimodal interaction modeling. The benchmark’s architecture supports continuous expansion and rolling evaluation, preventing overfitting and data contamination. The integration of long-term planning and multimodal perception within a unified evaluation environment pushes the boundaries of current AI capabilities, providing a comprehensive framework for assessing industrial-grade intelligence. These contributions establish a new standard for industry-relevant AI evaluation, bridging the gap between academic benchmarks and real-world applications.
Novelty
ALE’s primary innovation lies in its industry-grounded, long-term task design, combining multimodal interaction, structured deliverables, and expert-verified authenticity. Unlike prior benchmarks that focus on short-term, synthetic, or question-answer tasks, ALE emphasizes complex workflows that mirror real professional practices. Its multi-stage review process, continuous task pool expansion, and focus on economic impact distinguish it from existing evaluation systems, which often lack industry relevance and verification rigor. This approach sets a new benchmark for assessing AI’s readiness for industrial deployment, marking a shift from capability testing to practical, long-term performance measurement.
Limitations
- The current task pool, although extensive, still relies heavily on expert contributions, which may introduce subjective biases and limit coverage across all industries. Some sectors are underrepresented, and the complexity of tasks varies widely.
- Models still perform poorly on multi-step, multimodal workflows, indicating that current architectures lack the necessary reasoning, planning, and perception integration capabilities. Significant research is needed to close this gap.
- The evaluation process involves manual expert review at multiple stages, which limits scalability and introduces potential inconsistencies. Developing more advanced automated validation techniques remains an important future goal.
Future Work
Future directions include expanding the task pool to cover more industries and workflows, especially those with high economic impact. Incorporating reinforcement learning and meta-learning techniques could enhance models’ long-term planning and adaptability. Improving automated validation methods will increase evaluation scalability and objectivity. Strengthening industry collaborations will facilitate real-world deployment and feedback loops, accelerating progress toward industry-ready AI systems. Ultimately, the goal is to establish ALE as a dynamic, comprehensive, and industry-aligned evaluation platform that guides AI research toward tangible economic and societal benefits.
AI Executive Summary
In recent years, AI systems have achieved remarkable success in standardized benchmarks such as ImageNet, AlphaGo, and language models like GPT-4. These achievements have driven rapid progress in specific tasks, but their impact on real-world industry applications remains limited. The core challenge lies in the disconnect between short-term performance metrics and the ability of AI to perform complex, long-term, and economically valuable workflows in professional environments. Traditional benchmarks often evaluate isolated tasks—question answering, image classification, or synthetic simulations—that do not capture the intricacies of real industry work, which involves multi-step planning, multimodal perception, and domain-specific knowledge.
Recognizing this gap, the authors introduce Agents’ Last Exam (ALE), a novel benchmark designed to evaluate AI agents on authentic, long-horizon industry tasks. Developed through collaboration with over 250 industry experts, ALE encompasses over 1,000 tasks across 55 subdomains within 13 industry clusters, including manufacturing, healthcare, engineering, and more. These tasks are sourced directly from real professional projects, ensuring high relevance and complexity. Each task undergoes rigorous multi-stage review, including expert validation, engineering implementation, and automated verification, to guarantee authenticity and verifiability.
The evaluation framework centers on a generalist AI agent capable of multimodal perception, code execution, and long-term planning—similar to models like Claude Code and Codex. The benchmark employs structured deliverables and milestone checks to objectively assess task success, avoiding subjective human judgments. Results from current models reveal a significant performance gap: even the most advanced models achieve success rates below 50% on the easiest tasks and less than 10% on the hardest, with an average full pass rate below 1%. This stark contrast highlights the immense challenge of deploying AI in complex, real-world industry scenarios.
The significance of ALE extends beyond academic evaluation. It provides a practical, industry-aligned measure of AI readiness, guiding research toward capabilities that can truly transform industries. Its continuous, rolling evaluation mechanism ensures that the benchmark remains dynamic and relevant, fostering ongoing progress. The ultimate goal is to bridge the gap between benchmark success and tangible economic impact, accelerating AI’s integration into critical sectors like manufacturing, medicine, and finance. While current results underscore the difficulty ahead, ALE charts a clear path for future research, emphasizing the importance of long-term planning, multimodal understanding, and real-world verification in building truly industry-ready AI systems.
Deep Dive
Abstract
Recent AI systems have achieved strong results on a wide range of benchmarks, yet these gains have not translated into economically meaningful deployment across many professional domains. We argue that this gap is largely an evaluation problem: widely used benchmarks lack sustained performance measurement on real and economically valuable workflows. This paper introduces Agents' Last Exam (ALE), a benchmark designed to evaluate AI agents on long horizon, economically valuable, real world tasks with verifiable outcomes. Developed in collaboration with 250+ industry experts, ALE covers non-physical industries defined with reference to O*NET / SOC 2018 (the U.S. federal occupational taxonomy). It is organized around a task taxonomy with 55 sub fields grouped into 13 industry clusters covering 1K+ tasks. Current results show that the hardest tier remains far from saturated: across mainstream harness and backbone configurations, the average full pass rate is below 1%. ALE is designed as a living benchmark: its task pool grows continuously as new workflows and industries are onboarded. More broadly, ALE is intended not merely as another leaderboard, but as an instrument for closing the gap between benchmark success and GDP relevant impact.
References (20)
ReAct: Synergizing Reasoning and Acting in Language Models
Shunyu Yao, Jeffrey Zhao, Dian Yu et al.
The U.S. National Science Foundation
google,我,萨娜
Huafu Fang
Cursor
Kenneth A. Ross, C. S. Jensen, R. Snodgrass et al.
Artificial Intelligence Risk Management Framework (AI RMF 1.0)
AI Nist, Secretary Gina M. Raimondo, L. Locascio
The Twelfth International Conference on Learning Representations, ICLR 2024, Vienna, Austria, May 7-11, 2024
UNDERSTANDING WORK USING THE OCCUPATIONAL INFORMATION NETWORK (O*NET): IMPLICATIONS FOR PRACTICE AND RESEARCH
NORMAN G. Peterson, Michael D. Mumford, W. C. Borman et al.
ImageNet: A large-scale hierarchical image database
Jia Deng, Wei Dong, R. Socher et al.
Mastering the game of Go with deep neural networks and tree search
David Silver, Aja Huang, Chris J. Maddison et al.
Video Processing From Electro-Optical Sensors for Object Detection and Tracking in a Maritime Environment: A Survey
D. K. Prasad, D. Rajan, L. Rachmawati et al.
“Alibaba”
Ltd. v. Shenzhen Netac Technology Co. Ltd. and Guangzho Hangzhou Alibaba Advertising Co.
Measuring Massive Multitask Language Understanding
Dan Hendrycks, Collin Burns, Steven Basart et al.
Gym-Anything: Turn any Software into an Agent Environment
Pranjal Aggarwal, Graham Neubig, S. Welleck
Asynchronous Trajectory Matching-Based Multimodal Maritime Data Fusion for Vessel Traffic Surveillance in Inland Waterways
Yu Guo, Ryan Wen Liu, Jingxiang Qu et al.
WebArena: A Realistic Web Environment for Building Autonomous Agents
Shuyan Zhou, Frank F. Xu, Hao Zhu et al.
SWE-bench: Can Language Models Resolve Real-World GitHub Issues?
Carlos E. Jimenez, John Yang, Alexander Wettig et al.
GAIA: a benchmark for General AI Assistants
G. Mialon, Clémentine Fourrier, Craig Swift et al.
OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments
Tianbao Xie, Danyang Zhang, Jixuan Chen et al.
SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering
John Yang, Carlos E. Jimenez, Alexander Wettig et al.
OpenHands: An Open Platform for AI Software Developers as Generalist Agents
Xingyao Wang, Boxuan Li, Yufan Song et al.
Cited By (9)
Kimi K3: Open Frontier Intelligence
BrainPilot: Automating Brain Discovery with Agentic Research
FrontierChallenge: Evaluating Scientific Workflow Completion
What is Missing from AI Post-Training AI: An Empirical Analysis
StartupBench: Benchmarking General-Purpose Agents on Market-Validated End-to-End Workflows
FinProBench: Evaluating Financial AI Agents with Role-Grounded Rubrics Derived from Professional Deliverables
How Well Can Frontier Large Language Models Generate Structures? High Quality Prediction of Molecular Geometries with Help from Fine-Tuning
MobileForge: Annotation-Free Adaptation for Mobile GUI Agents with Hierarchical Feedback-Guided Policy Optimization
AGENTS4GEOS: agentic platform for open-source multi-physics simulation