Agents' Last Exam

TL;DR

ALE benchmark evaluates AI on long-term, high-value real-world industry tasks; current top models achieve less than 1% success rate on hardest tasks.

cs.AI 🔴 Advanced 2026-06-04 9 citations 89 views
Yiyou Sun Xinyang Han Weichen Zhang Yuanbo Pang Tianyu Wang Yuhan Cao Yixiao Huang Chris Duroiu Haoyun Zhang Jeffrey Lin Weishu Zhang Tyler Zeng Ying Yan Bo Liu Hanson Wen Mingyang Xu Xiaoyuan Liu Zimeng Chen Weiyan Shi Amanda Dsouza Vincent Sunn Chen Patrick Bryant Carl Boettiger Yamini Rangan Bradley Rothenberg Kyle Steinfeld Arvind Rao Tapio Schneider Georgios Yannakakis Laure Zanna Kaan Ozbay Ida Sim Tarek Zohdi George Em Karniadakis Jack Gallant Teresa Head-Gordon Yushan Li Wenxi Deng Tao Sun Huiqi Wang Zhun Wang Justin Xu Chris Yuhao Liu Yafei Cheng Rongwang Hu Aras Bacho Shengcao Cao Zengyi Qin Yixiong Chen Hengduan Fan Hao Liu Lin Zeng Shashank Muralidhar Bharadwaj Litian Gong Yingxuan Yang Maojia Song Ruheng Wang Zongzheng Zhang Honglin Bao Shuo Lu Jianhong Tu Zhonghua Wang Zheng Zhang Zijiao Chen Yanqiong Jiang Zhendong Li Bohan Lyu Chang Ma Peiran Xu Benran Zhang Shangding Gu Haoyue Hua Haoyang Li Wanzhe Liao Chengzhi Liu Junbo Peng Haoran Sun Zechen Xu Bo Chen Jiayi Cheng Yi Jiang Keying Kuang Yuan Li Youbang Pan Ziyan Rao Alexander Schubert Yifan Shen Vincent Siu Xiatao Sun Kangqi Zhang Xiaopan Zhang Yuchen Zhu Ishaan Singh Chandok Lei Ding Jingxuan Fan Andrew Glover Jiaming Hu Yiran Hu Wenbo Huang Zixin Jiang Haoran Jin Lukas Kim Ming Liu Yang Liu Alireza Rafiei Xuhuan Shen Kunyang Sun Sophia Sun Ting Sun Eric Wang Yixin Wang Hanwen Xing Sihan Xu Yuzheng Xu Zhongxing Xu Zhiling Yan Boqin Yuan Ruiqi Zhang Yifan Zhang Zibo Zhao Liana Santanu Bosu Antu Haoyue Bai Carlo Bosio Joseph Cavanagh Patricia Cavazos-Rehg Tianxing Chen Xuewen Chen Yipu Chen Chenyu Zhu Chen Dai Stefano De Castro Yunfu Deng Kaustubh Dhole Jiayuan Ding Chenchen Du Zhehang Du Hao Fan Run-Ze Fan Hengyu Fu Shi Gu Yifan Gu Charlie Guo Baihe Huang Baixiang Huang Rimika Jaiswal Zhihan Jiang Ran Jin Erin Kasson Xin Lan Joseph Lee Deren Lei Chenyu Li Daofeng Li Haitao Li Hongwei Li Jingyan Li Xiao Li Yi Li Yinsheng Li Yuangang Li Zhixu Li Wenyu Liang Longtai Liao Kevin Qinghong Lin Andy Zeyi Liu Che Liu Jiaming Liu Kaiyuan Liu Xuan Liu Pan Lu Wenbo Lv Yicheng Lyu Qiuyang Mang Kyle Montgomery Yuzhou Nie Ruoxi Ning Jorin Overwiening Xu Pan Layna Paraboschi Core Francisco Park Justin Purnomo Swati Rajwal Scott Rankin Bixuan Ren Yiren Rong HaoYang Shang Ventus Shaw Fiona Shen Jiawei Shen Minqi Shi Shi Qiu Huaxiu Yao Tianneng Shi Jonah So Vladislav Susoy Hannah Szlyk Haocheng Wang Jialu Wang Wei Wang Xinyu Wang Zehao Wang Dowling Wong Angela Wu Dehao Wu Fangyu Wu Mengyuan "Millie" Wu Yu Wu Yuchen Wu Yuhao Wu Qingpo Wuwu Weihang Xiao Yongyi Xiong Fan Xu Ruiling Xu Mingxuan Yan Benjamin Yang Jirong Yang Sen Yang Xiaoli Yang Yushi Yang Haoran Ye Xiaohu Yu Zhengming Yu Chenlong Zhang Chi Zhang Hanning Zhang Hanwen Zhang Junge Zhang Kunpeng Zhang Song Zhang Wenjin Zhang Wenshuo Zhang Ying Zhang Yizhi Zhang Brian Zhao Qijian Zhao Yimin Zhao Yuhaohua Zheng Liwei Zhou Tianyue Zhou Sichen Zhu Siqi Zhu Yan Zhu Yishu Zhu Jierui Zuo Chonghao Cai Helena Casademunt Wenjia Chen Cheng Cheng Nawen Deng Rao Fu Tianfu Fu Yifan Han He Ren Zhenyu He Qiao Jin Langlang Li Yuetai Li Sylvia Liu Lu Lu Luqing Zhou Subhabrata Mukherjee Yunqi Ouyang Yin Ren Dawei Shi Haoran Wu Zhiyue Wu Hannah Yao Zhuoran Yi Jenny Yu Rhea Zhan Hang Zhou Blake Zhu Junfan Zhu Alan Yuille Yang Liu Russell Alan Poldrack Jiachen Li Zhenglu Li Molei Tao Jing Huang Wenqi Shi Costas Spanos Lichao Sun Chenguang Wang Orson Xu Zhen Dong Hector Gomez Aylin Caliskan Ali Emami Haimin Hu Zhi Li Lihui Liu Murphy Niu Yi Shao Jianxin Sun Mikko Tolonen Ting Wang Sanjiv Das Yanjun Gao Wenbo Guo Erika J Schneider Zhiyong Lu Yian Ma Mark Mueller Radha Poovendran Somayeh Sojoudi Yinglun Zhu Dawn Song
AI evaluation industry application long-term tasks empirical validation industry workflows

Key Findings

Methodology

The Agents’ Last Exam (ALE) benchmark was developed through collaboration with over 250 industry experts, creating a task taxonomy covering 55 subdomains within 13 industry clusters. Tasks are sourced from real professional projects, ensuring representativeness and complexity. Multi-round expert reviews and automated quality controls verify authenticity and difficulty. The benchmark evaluates generalist AI agents—such as Claude Code and Codex—integrating multimodal perception, code execution, tool use, and long-term planning within a unified environment. Structured deliverables and milestone checks replace subjective human judgments, enabling objective, automatable validation. ALE emphasizes economic relevance, measuring not only task success but also industry impact, aiming to bridge the gap between benchmark performance and real-world deployment.

Key Results

  • The strongest configuration tested—Codex combined with GPT-5.5—achieved 82% success on Terminal-Bench, but only around 50% on ALE’s easiest tasks and less than 10% on the hardest. Most mainstream models, including Claude Code, recorded near-zero pass rates at high difficulty levels, with an average full pass rate below 1%. This highlights the significant gap between current AI capabilities and industry-level long-term task performance.
  • Performance varies considerably across industries; fields like electronics engineering and life sciences show relatively better results, but overall coverage remains limited. The experiments reveal that models struggle with multi-step, multimodal workflows, emphasizing the need for architectural and training innovations.
  • Results demonstrate that models are far from matching human experts in complex, real-world tasks, especially those requiring planning, reasoning, and multimodal understanding. These findings underscore the importance of developing AI systems capable of sustained, industry-grade performance.

Significance

ALE serves as a pioneering benchmark targeting real-world, long-horizon industry tasks, addressing a critical gap in current AI evaluation frameworks. Unlike traditional benchmarks focused on short-term or synthetic tasks, ALE emphasizes authentic workflows with economic value, providing a meaningful measure of AI readiness for industrial deployment. Its collaborative design with industry experts ensures relevance and realism, fostering progress toward AI systems that can operate reliably in manufacturing, healthcare, finance, and other sectors. By establishing a rigorous, scalable, and evolving evaluation platform, ALE aims to accelerate AI research from laboratory success to tangible economic impact, ultimately contributing to productivity gains and industry transformation.

Technical Contribution

This work introduces a novel evaluation paradigm that integrates real industry workflows into a structured, verifiable benchmark. Key innovations include sourcing tasks directly from professional projects, multi-stage expert review, automated validation of heterogeneous outputs, and multimodal interaction modeling. The benchmark’s architecture supports continuous expansion and rolling evaluation, preventing overfitting and data contamination. The integration of long-term planning and multimodal perception within a unified evaluation environment pushes the boundaries of current AI capabilities, providing a comprehensive framework for assessing industrial-grade intelligence. These contributions establish a new standard for industry-relevant AI evaluation, bridging the gap between academic benchmarks and real-world applications.

Novelty

ALE’s primary innovation lies in its industry-grounded, long-term task design, combining multimodal interaction, structured deliverables, and expert-verified authenticity. Unlike prior benchmarks that focus on short-term, synthetic, or question-answer tasks, ALE emphasizes complex workflows that mirror real professional practices. Its multi-stage review process, continuous task pool expansion, and focus on economic impact distinguish it from existing evaluation systems, which often lack industry relevance and verification rigor. This approach sets a new benchmark for assessing AI’s readiness for industrial deployment, marking a shift from capability testing to practical, long-term performance measurement.

Limitations

  • The current task pool, although extensive, still relies heavily on expert contributions, which may introduce subjective biases and limit coverage across all industries. Some sectors are underrepresented, and the complexity of tasks varies widely.
  • Models still perform poorly on multi-step, multimodal workflows, indicating that current architectures lack the necessary reasoning, planning, and perception integration capabilities. Significant research is needed to close this gap.
  • The evaluation process involves manual expert review at multiple stages, which limits scalability and introduces potential inconsistencies. Developing more advanced automated validation techniques remains an important future goal.

Future Work

Future directions include expanding the task pool to cover more industries and workflows, especially those with high economic impact. Incorporating reinforcement learning and meta-learning techniques could enhance models’ long-term planning and adaptability. Improving automated validation methods will increase evaluation scalability and objectivity. Strengthening industry collaborations will facilitate real-world deployment and feedback loops, accelerating progress toward industry-ready AI systems. Ultimately, the goal is to establish ALE as a dynamic, comprehensive, and industry-aligned evaluation platform that guides AI research toward tangible economic and societal benefits.

AI Executive Summary

In recent years, AI systems have achieved remarkable success in standardized benchmarks such as ImageNet, AlphaGo, and language models like GPT-4. These achievements have driven rapid progress in specific tasks, but their impact on real-world industry applications remains limited. The core challenge lies in the disconnect between short-term performance metrics and the ability of AI to perform complex, long-term, and economically valuable workflows in professional environments. Traditional benchmarks often evaluate isolated tasks—question answering, image classification, or synthetic simulations—that do not capture the intricacies of real industry work, which involves multi-step planning, multimodal perception, and domain-specific knowledge.

Recognizing this gap, the authors introduce Agents’ Last Exam (ALE), a novel benchmark designed to evaluate AI agents on authentic, long-horizon industry tasks. Developed through collaboration with over 250 industry experts, ALE encompasses over 1,000 tasks across 55 subdomains within 13 industry clusters, including manufacturing, healthcare, engineering, and more. These tasks are sourced directly from real professional projects, ensuring high relevance and complexity. Each task undergoes rigorous multi-stage review, including expert validation, engineering implementation, and automated verification, to guarantee authenticity and verifiability.

The evaluation framework centers on a generalist AI agent capable of multimodal perception, code execution, and long-term planning—similar to models like Claude Code and Codex. The benchmark employs structured deliverables and milestone checks to objectively assess task success, avoiding subjective human judgments. Results from current models reveal a significant performance gap: even the most advanced models achieve success rates below 50% on the easiest tasks and less than 10% on the hardest, with an average full pass rate below 1%. This stark contrast highlights the immense challenge of deploying AI in complex, real-world industry scenarios.

The significance of ALE extends beyond academic evaluation. It provides a practical, industry-aligned measure of AI readiness, guiding research toward capabilities that can truly transform industries. Its continuous, rolling evaluation mechanism ensures that the benchmark remains dynamic and relevant, fostering ongoing progress. The ultimate goal is to bridge the gap between benchmark success and tangible economic impact, accelerating AI’s integration into critical sectors like manufacturing, medicine, and finance. While current results underscore the difficulty ahead, ALE charts a clear path for future research, emphasizing the importance of long-term planning, multimodal understanding, and real-world verification in building truly industry-ready AI systems.

Deep Dive

Abstract

Recent AI systems have achieved strong results on a wide range of benchmarks, yet these gains have not translated into economically meaningful deployment across many professional domains. We argue that this gap is largely an evaluation problem: widely used benchmarks lack sustained performance measurement on real and economically valuable workflows. This paper introduces Agents' Last Exam (ALE), a benchmark designed to evaluate AI agents on long horizon, economically valuable, real world tasks with verifiable outcomes. Developed in collaboration with 250+ industry experts, ALE covers non-physical industries defined with reference to O*NET / SOC 2018 (the U.S. federal occupational taxonomy). It is organized around a task taxonomy with 55 sub fields grouped into 13 industry clusters covering 1K+ tasks. Current results show that the hardest tier remains far from saturated: across mainstream harness and backbone configurations, the average full pass rate is below 1%. ALE is designed as a living benchmark: its task pool grows continuously as new workflows and industries are onboarded. More broadly, ALE is intended not merely as another leaderboard, but as an instrument for closing the gap between benchmark success and GDP relevant impact.

cs.AI cs.CL cs.LG

References (20)

ReAct: Synergizing Reasoning and Acting in Language Models

Shunyu Yao, Jeffrey Zhao, Dian Yu et al.

2022 10647 citations View Analysis →

The U.S. National Science Foundation

1963 8 citations

google,我,萨娜

Huafu Fang

2006 9452 citations

Cursor

Kenneth A. Ross, C. S. Jensen, R. Snodgrass et al.

2009 197 citations

Artificial Intelligence Risk Management Framework (AI RMF 1.0)

AI Nist, Secretary Gina M. Raimondo, L. Locascio

2023 1276 citations

The Twelfth International Conference on Learning Representations, ICLR 2024, Vienna, Austria, May 7-11, 2024

2024 249 citations

UNDERSTANDING WORK USING THE OCCUPATIONAL INFORMATION NETWORK (O*NET): IMPLICATIONS FOR PRACTICE AND RESEARCH

NORMAN G. Peterson, Michael D. Mumford, W. C. Borman et al.

2001 469 citations

ImageNet: A large-scale hierarchical image database

Jia Deng, Wei Dong, R. Socher et al.

2009 75571 citations

Mastering the game of Go with deep neural networks and tree search

David Silver, Aja Huang, Chris J. Maddison et al.

2016 19210 citations

Video Processing From Electro-Optical Sensors for Object Detection and Tracking in a Maritime Environment: A Survey

D. K. Prasad, D. Rajan, L. Rachmawati et al.

2016 447 citations View Analysis →

“Alibaba”

Ltd. v. Shenzhen Netac Technology Co. Ltd. and Guangzho Hangzhou Alibaba Advertising Co.

2019 205 citations

Measuring Massive Multitask Language Understanding

Dan Hendrycks, Collin Burns, Steven Basart et al.

2020 9302 citations View Analysis →

Gym-Anything: Turn any Software into an Agent Environment

Pranjal Aggarwal, Graham Neubig, S. Welleck

2026 15 citations View Analysis →

Asynchronous Trajectory Matching-Based Multimodal Maritime Data Fusion for Vessel Traffic Surveillance in Inland Waterways

Yu Guo, Ryan Wen Liu, Jingxiang Qu et al.

2023 95 citations View Analysis →

WebArena: A Realistic Web Environment for Building Autonomous Agents

Shuyan Zhou, Frank F. Xu, Hao Zhu et al.

2023 1891 citations View Analysis →

SWE-bench: Can Language Models Resolve Real-World GitHub Issues?

Carlos E. Jimenez, John Yang, Alexander Wettig et al.

2023 3510 citations View Analysis →

GAIA: a benchmark for General AI Assistants

G. Mialon, Clémentine Fourrier, Craig Swift et al.

2023 1140 citations View Analysis →

OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments

Tianbao Xie, Danyang Zhang, Jixuan Chen et al.

2024 1081 citations View Analysis →

SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering

John Yang, Carlos E. Jimenez, Alexander Wettig et al.

2024 1689 citations View Analysis →

OpenHands: An Open Platform for AI Software Developers as Generalist Agents

Xingyao Wang, Boxuan Li, Yufan Song et al.

2024 968 citations View Analysis →

Cited By (9)

Kimi K3: Open Frontier Intelligence

2026 9 citations ⭐ Influential View Analysis →

BrainPilot: Automating Brain Discovery with Agentic Research

2026 1 citations ⭐ Influential View Analysis →

FrontierChallenge: Evaluating Scientific Workflow Completion

What is Missing from AI Post-Training AI: An Empirical Analysis

StartupBench: Benchmarking General-Purpose Agents on Market-Validated End-to-End Workflows

FinProBench: Evaluating Financial AI Agents with Role-Grounded Rubrics Derived from Professional Deliverables

2026 1 citations View Analysis →

How Well Can Frontier Large Language Models Generate Structures? High Quality Prediction of Molecular Geometries with Help from Fine-Tuning

MobileForge: Annotation-Free Adaptation for Mobile GUI Agents with Hierarchical Feedback-Guided Policy Optimization

2026 3 citations View Analysis →

AGENTS4GEOS: agentic platform for open-source multi-physics simulation