Humanity's Last Exam

TL;DR

HLE introduces a 2,500-question multi-modal academic benchmark, exposing current LLM limitations in complex reasoning and cross-disciplinary understanding.

cs.LG 🔴 Advanced 2025-01-24 56 views
Long Phan Alice Gatti Ziwen Han Nathaniel Li Josephina Hu Hugh Zhang Chen Bo Calvin Zhang Mohamed Shaaban John Ling Sean Shi Michael Choi Anish Agrawal Arnav Chopra Adam Khoja Ryan Kim Richard Ren Jason Hausenloy Oliver Zhang Mantas Mazeika Dmitry Dodonov Tung Nguyen Jaeho Lee Daron Anderson Mikhail Doroshenko Alun Cennyth Stokes Mobeen Mahmood Oleksandr Pokutnyi Oleg Iskra Jessica P. Wang John-Clark Levin Mstyslav Kazakov Fiona Feng Steven Y. Feng Haoran Zhao Michael Yu Varun Gangal Chelsea Zou Zihan Wang Serguei Popov Robert Gerbicz Geoff Galgon Johannes Schmitt Will Yeadon Yongki Lee Scott Sauers Alvaro Sanchez Fabian Giska Marc Roth Søren Riis Saiteja Utpala Noah Burns Gashaw M. Goshu Mohinder Maheshbhai Naiya Chidozie Agu Zachary Giboney Antrell Cheatom Francesco Fournier-Facio Sarah-Jane Crowson Lennart Finke Zerui Cheng Jennifer Zampese Ryan G. Hoerr Mark Nandor Hyunwoo Park Tim Gehrunger Jiaqi Cai Ben McCarty Alexis C Garretson Edwin Taylor Damien Sileo Qiuyu Ren Usman Qazi Lianghui Li Jungbae Nam John B. Wydallis Pavel Arkhipov Jack Wei Lun Shi Aras Bacho Chris G. Willcocks Hangrui Cao Sumeet Motwani Emily de Oliveira Santos Johannes Veith Edward Vendrow Doru Cojoc Kengo Zenitani Joshua Robinson Longke Tang Yuqi Li Joshua Vendrow Natanael Wildner Fraga Vladyslav Kuchkin Andrey Pupasov Maksimov Pierre Marion Denis Efremov Jayson Lynch Kaiqu Liang Aleksandar Mikov Andrew Gritsevskiy Julien Guillod Gözdenur Demir Dakotah Martinez Ben Pageler Kevin Zhou Saeed Soori Ori Press Henry Tang Paolo Rissone Sean R. Green Lina Brüssel Moon Twayana Aymeric Dieuleveut Joseph Marvin Imperial Ameya Prabhu Jinzhou Yang Nick Crispino Arun Rao Dimitri Zvonkine Gabriel Loiseau Mikhail Kalinin Marco Lukas Ciprian Manolescu Nate Stambaugh Subrata Mishra Tad Hogg Carlo Bosio Brian P Coppola Julian Salazar Jaehyeok Jin Rafael Sayous Stefan Ivanov Philippe Schwaller Shaipranesh Senthilkuma Andres M Bran Andres Algaba Kelsey Van den Houte Lynn Van Der Sypt Brecht Verbeken David Noever Alexei Kopylov Benjamin Myklebust Bikun Li Lisa Schut Evgenii Zheltonozhskii Qiaochu Yuan Derek Lim Richard Stanley Tong Yang John Maar Julian Wykowski Martí Oller Anmol Sahu Cesare Giulio Ardito Yuzheng Hu Ariel Ghislain Kemogne Kamdoum Alvin Jin Tobias Garcia Vilchis Yuexuan Zu Martin Lackner James Koppel Gongbo Sun Daniil S. Antonenko Steffi Chern Bingchen Zhao Pierrot Arsene Joseph M Cavanagh Daofeng Li Jiawei Shen Donato Crisostomi Wenjin Zhang Ali Dehghan Sergey Ivanov David Perrella Nurdin Kaparov Allen Zang Ilia Sucholutsky Arina Kharlamova Daniil Orel Vladislav Poritski Shalev Ben-David Zachary Berger Parker Whitfill Michael Foster Daniel Munro Linh Ho Shankar Sivarajan Dan Bar Hava Aleksey Kuchkin David Holmes Alexandra Rodriguez-Romero Frank Sommerhage Anji Zhang Richard Moat Keith Schneider Zakayo Kazibwe Don Clarke Dae Hyun Kim Felipe Meneguitti Dias Sara Fish Veit Elser Tobias Kreiman Victor Efren Guadarrama Vilchis Immo Klose Ujjwala Anantheswaran Adam Zweiger Kaivalya Rawal Jeffery Li Jeremy Nguyen Nicolas Daans Haline Heidinger Maksim Radionov Václav Rozhoň Vincent Ginis Christian Stump Niv Cohen Rafał Poświata Josef Tkadlec Alan Goldfarb Chenguang Wang Piotr Padlewski Stanislaw Barzowski Kyle Montgomery Ryan Stendall Jamie Tucker-Foltz Jack Stade T. Ryan Rogers Tom Goertzen Declan Grabb Abhishek Shukla Alan Givré John Arnold Ambay Archan Sen Muhammad Fayez Aziz Mark H Inlow Hao He Ling Zhang Younesse Kaddar Ivar Ängquist Yanxu Chen Harrison K Wang Kalyan Ramakrishnan Elliott Thornley Antonio Terpin Hailey Schoelkopf Eric Zheng Avishy Carmi Ethan D. L. Brown Kelin Zhu Max Bartolo Richard Wheeler Martin Stehberger Peter Bradshaw JP Heimonen Kaustubh Sridhar Ido Akov Jennifer Sandlin Yury Makarychev Joanna Tam Hieu Hoang David M. Cunningham Vladimir Goryachev Demosthenes Patramanis Michael Krause Andrew Redenti David Aldous Jesyin Lai Shannon Coleman Jiangnan Xu Sangwon Lee Ilias Magoulas Sandy Zhao Ning Tang Michael K. Cohen Orr Paradise Jan Hendrik Kirchner Maksym Ovchynnikov Jason O. Matos Adithya Shenoy Michael Wang Yuzhou Nie Anna Sztyber-Betley Paolo Faraboschi Robin Riblet Jonathan Crozier Shiv Halasyamani Shreyas Verma Prashant Joshi Eli Meril Ziqiao Ma Jérémy Andréoletti Raghav Singhal Jacob Platnick Volodymyr Nevirkovets Luke Basler Alexander Ivanov Seri Khoury Nils Gustafsson Marco Piccardo Hamid Mostaghimi Qijia Chen Virendra Singh Tran Quoc Khánh Paul Rosu Hannah Szlyk Zachary Brown Himanshu Narayan Aline Menezes Jonathan Roberts William Alley Kunyang Sun Arkil Patel Max Lamparth Anka Reuel Linwei Xin Hanmeng Xu Jacob Loader Freddie Martin Zixuan Wang Andrea Achilleos Thomas Preu Tomek Korbak Ida Bosio Fereshteh Kazemi Ziye Chen Biró Bálint Eve J. Y. Lo Jiaqi Wang Maria Inês S. Nunes Jeremiah Milbauer M Saiful Bari Zihao Wang Behzad Ansarinejad Yewen Sun Stephane Durand Hossam Elgnainy Guillaume Douville Daniel Tordera George Balabanian Hew Wolff Lynna Kvistad Hsiaoyun Milliron Ahmad Sakor Murat Eron Andrew Favre D. O. Shailesh Shah Xiaoxiang Zhou Firuz Kamalov Sherwin Abdoli Tim Santens Shaul Barkan Allison Tee Robin Zhang Alessandro Tomasiello G. Bruno De Luca Shi-Zhuo Looi Vinh-Kha Le Noam Kolt Jiayi Pan Emma Rodman Jacob Drori Carl J Fossum Niklas Muennighoff Milind Jagota Ronak Pradeep Honglu Fan Jonathan Eicher Michael Chen Kushal Thaman William Merrill Moritz Firsching Carter Harris Stefan Ciobâcă Jason Gross Rohan Pandey Ilya Gusev Adam Jones Shashank Agnihotri Pavel Zhelnov Mohammadreza Mofayezi Alexander Piperski David K. Zhang Kostiantyn Dobarskyi Roman Leventov Ignat Soroko Joshua Duersch Vage Taamazyan Andrew Ho Wenjie Ma William Held Ruicheng Xian Armel Randy Zebaze Mohanad Mohamed Julian Noah Leser Michelle X Yuan Laila Yacar Johannes Lengler Katarzyna Olszewska Claudio Di Fratta Edson Oliveira Joseph W. Jackson Andy Zou Muthu Chidambaram Timothy Manik Hector Haffenden Dashiell Stander Ali Dasouqi Alexander Shen Bita Golshani David Stap Egor Kretov Mikalai Uzhou Alina Borisovna Zhidkovskaya Nick Winter Miguel Orbegozo Rodriguez Robert Lauff Dustin Wehr Colin Tang Zaki Hossain Shaun Phillips Fortuna Samuele Fredrik Ekström Angela Hammon Oam Patel Faraz Farhidi George Medley Forough Mohammadzadeh Madellene Peñaflor Haile Kassahun Alena Friedrich Rayner Hernandez Perez Daniel Pyda Taom Sakal Omkar Dhamane Ali Khajegili Mirabadi Eric Hallman Kenchi Okutsu Mike Battaglia Mohammad Maghsoudimehrabani Alon Amit Dave Hulbert Roberto Pereira Simon Weber Handoko Anton Peristyy Stephen Malina Mustafa Mehkary Rami Aly Frank Reidegeld Anna-Katharina Dick Cary Friday Mukhwinder Singh Hassan Shapourian Wanyoung Kim Mariana Costa Hubeyb Gurdogan Harsh Kumar Chiara Ceconello Chao Zhuang Haon Park Micah Carroll Andrew R. Tawfeek Stefan Steinerberger Daattavya Aggarwal Michael Kirchhof Linjie Dai Evan Kim Johan Ferret Jainam Shah Yuzhou Wang Minghao Yan Krzysztof Burdzy Lixin Zhang Antonio Franca Diana T. Pham Kang Yong Loh Joshua Robinson Abram Jackson Paolo Giordano Philipp Petersen Adrian Cosma Jesus Colino Colin White Jacob Votava Vladimir Vinnikov Ethan Delaney Petr Spelda Vit Stritecky Syed M. Shahid Jean-Christophe Mourrat Lavr Vetoshkin Koen Sponselee Renas Bacho Zheng-Xin Yong Florencia de la Rosa Nathan Cho Xiuyu Li Guillaume Malod Orion Weller Guglielmo Albani Leon Lang Julien Laurendeau Dmitry Kazakov Fatimah Adesanya Julien Portier Lawrence Hollom Victor Souza Yuchen Anna Zhou Julien Degorre Yiğit Yalın Gbenga Daniel Obikoya Rai Filippo Bigi M. C. Boscá Oleg Shumar Kaniuar Bacho Gabriel Recchia Mara Popescu Nikita Shulga Ngefor Mildred Tanwie Thomas C. H. Lux Ben Rank Colin Ni Matthew Brooks Alesia Yakimchyk Huanxu Liu Stefano Cavalleri Olle Häggström Emil Verkama Joshua Newbould Hans Gundlach Leonor Brito-Santana Brian Amaro Vivek Vajipey Rynaa Grover Ting Wang Yosi Kratish Wen-Ding Li Sivakanth Gopi Andrea Caciolai Christian Schroeder de Witt Pablo Hernández-Cámara Emanuele Rodolà Jules Robins Dominic Williamson Vincent Cheng Brad Raynor Hao Qi Ben Segev Jingxuan Fan Sarah Martinson Erik Y. Wang Kaylie Hausknecht Michael P. Brenner Mao Mao Christoph Demian Peyman Kassani Xinyu Zhang David Avagian Eshawn Jessica Scipio Alon Ragoler Justin Tan Blake Sims Rebeka Plecnik Aaron Kirtland Omer Faruk Bodur D. P. Shinde Yan Carlos Leyva Labrador Zahra Adoul Mohamed Zekry Ali Karakoc Tania C. B. Santos Samir Shamseldeen Loukmane Karim Anna Liakhovitskaia Nate Resman Nicholas Farina Juan Carlos Gonzalez Gabe Maayan Earth Anderson Rodrigo De Oliveira Pena Elizabeth Kelley Hodjat Mariji Rasoul Pouriamanesh Wentao Wu Ross Finocchio Ismail Alarab Joshua Cole Danyelle Ferreira Bryan Johnson Mohammad Safdari Liangti Dai Siriphan Arthornthurasuk Isaac C. McAlister Alejandro José Moyano Alexey Pronin Jing Fan Angel Ramirez-Trinidad Yana Malysheva Daphiny Pottmaier Omid Taheri Stanley Stepanic Samuel Perry Luke Askew Raúl Adrián Huerta Rodríguez Ali M. R. Minissi Ricardo Lorena Krishnamurthy Iyer Arshad Anil Fasiludeen Ronald Clark Josh Ducey Matheus Piza Maja Somrak Eric Vergo Juehang Qin Benjámin Borbás Eric Chu Jack Lindsey Antoine Jallon I. M. J. McInnis Evan Chen Avi Semler Luk Gloor Tej Shah Marc Carauleanu Pascal Lauer Tran Đuc Huy Hossein Shahrtash Emilien Duc Lukas Lewark Assaf Brown Samuel Albanie Brian Weber Warren S. Vaz Pierre Clavier Yiyang Fan Gabriel Poesia Reis e Silva Long Lian Marcus Abramovitch Xi Jiang Sandra Mendoza Murat Islam Juan Gonzalez Vasilios Mavroudis Justin Xu Pawan Kumar Laxman Prasad Goswami Daniel Bugas Nasser Heydari Ferenc Jeanplong Thorben Jansen Antonella Pinto Archimedes Apronti Abdallah Galal Ng Ze-An Ankit Singh Tong Jiang Joan of Arc Xavier Kanu Priya Agarwal Mohammed Berkani Gang Zhang Zhehang Du Benedito Alves de Oliveira Junior Dmitry Malishev Nicolas Remy Taylor D. Hartman Tim Tarver Stephen Mensah Gautier Abou Loume Wiktor Morak Farzad Habibi Sarah Hoback Will Cai Javier Gimenez Roselynn Grace Montecillo Jakub Łucki Russell Campbell Asankhaya Sharma Khalida Meer Shreen Gul Daniel Espinosa Gonzalez Xavier Alapont Alex Hoover Gunjan Chhablani Freddie Vargus Arunim Agarwal Yibo Jiang Deepakkumar Patil David Outevsky Kevin Joseph Scaria Rajat Maheshwari Abdelkader Dendane Priti Shukla Ashley Cartwright Sergei Bogdanov Niels Mündler Sören Möller Luca Arnaboldi Kunvar Thaman Muhammad Rehan Siddiqi Prajvi Saxena Himanshu Gupta Tony Fruhauff Glen Sherman Mátyás Vincze Siranut Usawasutsakorn Dylan Ler Anil Radhakrishnan Innocent Enyekwe Sk Md Salauddin Jiang Muzhen Aleksandr Maksapetyan Vivien Rossbach Chris Harjadi Mohsen Bahaloohoreh Claire Sparrow Jasdeep Sidhu Sam Ali Song Bian John Lai Eric Singer Justine Leon Uro Greg Bateman Mohamed Sayed Ahmed Menshawy Darling Duclosel Dario Bezzi Yashaswini Jain Ashley Aaron Murat Tiryakioglu Sheeshram Siddh Keith Krenek Imad Ali Shah Jun Jin Scott Creighton Denis Peskoff Zienab EL-Wasif Ragavendran P Michael Richmond Joseph McGowan Tejal Patwardhan Hao-Yu Sun Ting Sun Nikola Zubić Samuele Sala Stephen Ebert Jean Kaddour Manuel Schottdorf Dianzhuo Wang Gerol Petruzella Alex Meiburg Tilen Medved Ali ElSheikh S Ashwin Hebbar Lorenzo Vaquero Xianjun Yang Jason Poulos Vilém Zouhar Sergey Bogdanik Mingfang Zhang Jorge Sanz-Ros David Anugraha Yinwei Dai Anh N. Nhu Xue Wang Ali Anil Demircali Zhibai Jia Yuyin Zhou Juncheng Wu Mike He Nitin Chandok Aarush Sinha Gaoxiang Luo Long Le Mickaël Noyé Michał Perełkiewicz Ioannis Pantidis Tianbo Qi Soham Sachin Purohit Letitia Parcalabescu Thai-Hoa Nguyen Genta Indra Winata Edoardo M. Ponti Hanchen Li Kaustubh Dhole Jongee Park Dario Abbondanza Yuanli Wang Anupam Nayak Diogo M. Caetano Antonio A. W. L. Wong Maria del Rio-Chanona Dániel Kondor Pieter Francois Ed Chalstrey Jakob Zsambok Dan Hoyer Jenny Reddish Jakob Hauser Francisco-Javier Rodrigo-Ginés Suchandra Datta Maxwell Shepherd Thom Kamphuis Qizheng Zhang Hyunjun Kim Ruiji Sun Jianzhu Yao Franck Dernoncourt Satyapriya Krishna Sina Rismanchian Bonan Pu Francesco Pinto Yingheng Wang Kumar Shridhar Kalon J. Overholt Glib Briia Hieu Nguyen David Soler Bartomeu Tony CY Pang Adam Wecker Yifan Xiong Fanfei Li Lukas S. Huber Joshua Jaeger Romano De Maddalena Xing Han Lù Yuhui Zhang Claas Beger Patrick Tser Jern Kon Sean Li Vivek Sanker Ming Yin Yihao Liang Xinlu Zhang Ankit Agrawal Li S. Yifei Zechen Zhang Mu Cai Yasin Sonmez Costin Cozianu Changhao Li Alex Slen Shoubin Yu Hyun Kyu Park Gabriele Sarti Marcin Briański Alessandro Stolfo Truong An Nguyen Mike Zhang Yotam Perlitz Jose Hernandez-Orallo Runjia Li Amin Shabani Felix Juefei-Xu Shikhar Dhingra Orr Zohar My Chiffon Nguyen Alexander Pondaven Abdurrahim Yilmaz Xuandong Zhao Chuanyang Jin Muyan Jiang Stefan Todoran Xinyao Han Jules Kreuer Brian Rabern Anna Plassart Martino Maggetti Luther Yap Robert Geirhos Jonathon Kean Dingsu Wang Sina Mollaei Chenkai Sun Yifan Yin Shiqi Wang Rui Li Yaowen Chang Anjiang Wei Alice Bizeul Xiaohan Wang Alexandre Oliveira Arrais Kushin Mukherjee Jorge Chamorro-Padial Jiachen Liu Xingyu Qu Junyi Guan Adam Bouyamourn Shuyu Wu Martyna Plomecka Junda Chen Mengze Tang Jiaqi Deng Shreyas Subramanian Haocheng Xi Haoxuan Chen Weizhi Zhang Yinuo Ren Haoqin Tu Sejong Kim Yushun Chen Sara Vera Marjanović Junwoo Ha Grzegorz Luczyna Jeff J. Ma Zewen Shen Dawn Song Cedegao E. Zhang Zhun Wang Gaël Gendron Yunze Xiao Leo Smucker Erica Weng Kwok Hao Lee Zhe Ye Stefano Ermon Ignacio D. Lopez-Miguel Theo Knights Anthony Gitter Namkyu Park Boyi Wei Hongzheng Chen Kunal Pai Ahmed Elkhanany Han Lin Philipp D. Siedler Jichao Fang Ritwik Mishra Károly Zsolnai-Fehér Xilin Jiang Shadab Khan Jun Yuan Rishab Kumar Jain Xi Lin Mike Peterson Zhe Wang Aditya Malusare Maosen Tang Isha Gupta Ivan Fosin Timothy Kang Barbara Dworakowska Kazuki Matsumoto Guangyao Zheng Gerben Sewuster Jorge Pretel Villanueva Ivan Rannev Igor Chernyavsky Jiale Chen Deepayan Banik Ben Racz Wenchao Dong Jianxin Wang Laila Bashmal Duarte V. Gonçalves Wei Hu Kaushik Bar Ondrej Bohdal Atharv Singh Patlan Shehzaad Dhuliawala Caroline Geirhos Julien Wist Yuval Kansal Bingsen Chen Kutay Tire Atak Talay Yücel Brandon Christof Veerupaksh Singla Zijian Song Sanxing Chen Jiaxin Ge Kaustubh Ponkshe Isaac Park Tianneng Shi Martin Q. Ma Joshua Mak Sherwin Lai Antoine Moulin Zhuo Cheng Zhanda Zhu Ziyi Zhang Vaidehi Patil Ketan Jha Qiutong Men Jiaxuan Wu Tianchi Zhang Bruno Hebling Vieira Alham Fikri Aji Jae-Won Chung Mohammed Mahfoud Ha Thi Hoang Marc Sperzel Wei Hao Kristof Meding Sihan Xu Vassilis Kostakos Davide Manini Yueying Liu Christopher Toukmaji Jay Paek Eunmi Yu Arif Engin Demircali Zhiyi Sun Ivan Dewerpe Hongsen Qin Roman Pflugfelder James Bailey Johnathan Morris Ville Heilala Sybille Rosset Zishun Yu Peter E. Chen Woongyeong Yeo Eeshaan Jain Ryan Yang Sreekar Chigurupati Julia Chernyavsky Sai Prajwal Reddy Subhashini Venugopalan Hunar Batra Core Francisco Park Hieu Tran Guilherme Maximiano Genghan Zhang Yizhuo Liang Hu Shiyu Rongwu Xu Rui Pan Siddharth Suresh Ziqi Liu Samaksh Gulati Songyang Zhang Peter Turchin Christopher W. Bartlett Christopher R. Scotese Phuong M. Cao Ben Wu Jacek Karwowski Davide Scaramuzza Aakaash Nattanmai Gordon McKellips Anish Cheraku Asim Suhail Ethan Luo Marvin Deng Jason Luo Ashley Zhang Kavin Jindel Jay Paek Kasper Halevy Allen Baranov Michael Liu Advaith Avadhanam David Zhang Vincent Cheng Brad Ma Evan Fu Liam Do Joshua Lass Hubert Yang Surya Sunkari Vishruth Bharath Violet Ai James Leung Rishit Agrawal Alan Zhou Kevin Chen Tejas Kalpathi Ziqi Xu Gavin Wang Tyler Xiao Erik Maung Sam Lee Ryan Yang Roy Yue Ben Zhao Julia Yoon Sunny Sun Aryan Singh Ethan Luo Clark Peng Tyler Osbey Taozhi Wang Daryl Echeazu Hubert Yang Timothy Wu Spandan Patel Vidhi Kulkarni Vijaykaarti Sundarapandiyan Ashley Zhang Andrew Le Zafir Nasim Srikar Yalam Ritesh Kasamsetty Soham Samal Hubert Yang David Sun Nihar Shah Abhijeet Saha Alex Zhang Leon Nguyen Laasya Nagumalli Kaixin Wang Alan Zhou Aidan Wu Jason Luo Anwith Telluri Steven Dillmann Zhengxiang Wang Junyu Luo Hugo Lunn Artem Gazizov Haitz Sáez de Ocáriz Borde Ivan Trus Morgan Hervault Zheyu Zhang Bo Chen Yuchen Wu Christopher J. Cordier Gün Kaynar Cansin Ayvaz Polina Avdiunina Johannes Brust Xingjian Diao K. D. Meaney Yifan Gu Chenyu Wang Chenzhuo Dong William Wright Simon Brave Owen Root Jiayuan Liu Chow Chun Lok Tianqin Li Shiyi Du Dailan He Lufeiya Liu Sina Jamalzadegan Anil Ramakrishna Xuanqing Xu Xin Qing Xin Luo Wenkai Li Shi Bo Filipp Gusev Maximos Skandalis Desheng Ma Chunhui Zhang Haoran Qiu Allen G Hart Rickard Brüel Gabrielsson Ido Akov Artem Lukoianov Summer Yue Alexandr Wang Dan Hendrycks
multi-modal academic benchmark large models knowledge frontier automated grading

Key Findings

Methodology

HLE features a globally curated set of 2,500 multi-modal questions spanning mathematics, sciences, and humanities, designed by experts. Questions include multiple-choice and short-answer formats, with clear, verifiable solutions that cannot be quickly retrieved online. Evaluation employs an automated grading system combined with multi-step reasoning and knowledge verification modules. The framework assesses accuracy, calibration, and reasoning paths, emphasizing tasks requiring deep understanding and cross-disciplinary knowledge integration. The system incorporates retrieval-augmented reasoning and multi-modal input processing, pushing models beyond simple pattern matching.

Key Results

  • On HLE, state-of-the-art models like GPT-4 achieve an average accuracy of 35%, significantly below human experts’ 90%+ performance. Performance varies across disciplines, with complex reasoning and multi-modal tasks dropping accuracy to 20-30%. Model calibration is poor, indicating unreliable confidence estimates. Comparing architectures reveals limited improvements, highlighting persistent gaps in reasoning and knowledge depth, especially in long-tail and cross-domain scenarios.
  • Models struggle notably with multi-modal questions involving charts, images, and multi-step inference, underscoring the challenge of genuine understanding. Ablation studies show that removing multi-modal inputs or multi-turn reasoning reduces accuracy by over 15%, confirming their importance.
  • The experiments demonstrate that current models are far from human-level performance on high-level academic tasks, emphasizing the need for new architectures and training paradigms to bridge this gap.

Significance

HLE serves as a comprehensive, challenging benchmark that pushes the frontier of AI understanding and reasoning. It exposes the limitations of existing models in high-order cognition, multi-modal comprehension, and interdisciplinary reasoning. Its broad coverage and difficulty make it a critical tool for guiding future research, benchmarking progress, and setting industry standards. By revealing the current gaps, HLE accelerates efforts toward models capable of genuine academic-level understanding, with implications for education, scientific research, and AI safety. It establishes a new gold standard for evaluating AI’s true cognitive capabilities.

Technical Contribution

This work introduces a multi-modal, multi-disciplinary benchmark with rigorous automatic evaluation, combining advanced reasoning modules, retrieval-augmented knowledge verification, and expert-designed questions. It innovates by integrating multi-turn reasoning and multi-modal input processing, setting a new standard for assessing deep understanding. The platform enables systematic comparison across architectures, fostering development of models with improved reasoning, knowledge integration, and calibration. It also provides an open, extensible evaluation framework that can evolve with future AI advancements.

Novelty

HLE is the first large-scale, multi-modal, cross-disciplinary academic benchmark designed explicitly to challenge models beyond simple pattern recognition. Its integration of expert-curated questions, multi-step reasoning, and multimodal inputs distinguishes it from existing benchmarks like MMLU or BIG-Bench, which focus mainly on text-based tasks. The emphasis on high difficulty and real-world knowledge makes it a pioneering effort in pushing AI towards human-level academic reasoning.

Limitations

  • Models still underperform significantly on multi-modal, multi-step, and long-tail knowledge tasks, indicating that current architectures lack the necessary reasoning depth and knowledge generalization. The automatic evaluation may not fully capture nuanced correctness, especially in open-ended questions. The benchmark’s difficulty might limit its immediate applicability in real-world scenarios, and the design may favor certain model types over others. Future work should incorporate human-in-the-loop assessments and expand multimodal capabilities.

Future Work

Future directions include expanding the question set, incorporating open-ended and interactive tasks, and enhancing multimodal understanding. Developing models with better reasoning, knowledge retrieval, and calibration will be prioritized. The authors also plan to integrate human feedback for more nuanced evaluation and explore real-world applications such as AI-assisted education and scientific discovery. Cross-model generalization and few-shot learning capabilities will be key focus areas, aiming to close the gap between AI and human expertise.

AI Executive Summary

The rapid progress of large language models (LLMs) has led to the widespread adoption of benchmarks like MMLU and BIG-Bench, yet these metrics are now insufficient to differentiate the true capabilities of the latest models. Recognizing this, the authors introduce 'Humanity’s Last Exam' (HLE), a comprehensive, multi-modal, 2,500-question benchmark designed to push the boundaries of AI understanding. Curated by global experts, HLE covers mathematics, sciences, humanities, and includes complex multi-modal questions involving images, charts, and multi-step reasoning. The goal is to evaluate whether models can truly comprehend and reason at a human expert level, especially in academic contexts.

The evaluation framework employs an automated scoring system combined with reasoning modules that verify the logical consistency and knowledge accuracy of model outputs. Experiments with GPT-4, PaLM 2, and Claude reveal that these models achieve only around 35% accuracy, far below human performance (>90%). The results highlight persistent gaps in reasoning depth, multi-modal comprehension, and cross-disciplinary knowledge integration. These findings underscore the need for new architectures and training paradigms to address the current limitations.

HLE’s significance lies in its ability to serve as a rigorous, challenging benchmark that exposes the true extent of AI’s cognitive abilities. It provides a clear target for future research and development, emphasizing the importance of multi-modal understanding, multi-step reasoning, and knowledge retrieval. The benchmark also offers a platform for systematic comparison of models, fostering innovation in AI architectures aimed at closing the gap with human experts.

Looking ahead, the authors plan to expand the question bank, incorporate more open-ended and interactive tasks, and improve multimodal processing capabilities. They advocate for integrating human-in-the-loop evaluation to refine assessment accuracy and reliability. Ultimately, HLE aims to catalyze the development of AI systems capable of genuine academic-level reasoning, with broad implications for education, scientific discovery, and AI safety. It marks a pivotal step toward understanding and enhancing AI’s cognitive frontier.

Deep Analysis

Background

Over recent years, large-scale pretraining of models such as GPT-4, PaLM 2, and Claude has revolutionized natural language understanding and generation. These models have achieved remarkable performance on benchmarks like MMLU and BIG-Bench, often surpassing human accuracy in many tasks. However, these benchmarks mainly test factual recall, simple reasoning, or pattern matching, and are limited in scope for evaluating genuine understanding. As models approach or exceed 90% accuracy, it becomes increasingly difficult to differentiate their true reasoning capabilities. This has led to a demand for more challenging, comprehensive evaluation frameworks that can reveal the remaining gaps in AI cognition. The authors propose HLE as a solution, aiming to push models beyond current limitations by designing questions that require deep reasoning, multi-modal comprehension, and cross-disciplinary knowledge integration.

Core Problem

Despite impressive progress, current models still struggle with high-level academic reasoning, especially in multi-modal contexts involving images, graphs, and complex multi-step inference. Existing benchmarks do not sufficiently challenge models in these aspects, leading to an overestimation of their true capabilities. This gap hampers the development of AI systems that can reliably assist in scientific research, education, and decision-making. The core problem is to create a benchmark that not only covers broad academic disciplines but also demands deep understanding, reasoning, and knowledge retrieval, thus exposing the real limits of current models and guiding future innovations.

Innovation

The paper introduces several key innovations: 1) a globally curated, expert-designed set of 2,500 multi-modal, high-difficulty questions covering diverse academic fields; 2) integration of multi-turn reasoning and knowledge verification modules to assess reasoning paths; 3) automatic, scalable scoring system that evaluates accuracy, calibration, and reasoning quality; 4) emphasis on questions that surpass simple retrieval, requiring genuine comprehension and inference. This approach differs from traditional benchmarks by combining multimodal inputs, multi-step reasoning, and expert-level difficulty, creating a more realistic and demanding testbed for AI cognition.

Methodology

  • �� Question Design: Experts from global institutions created questions across disciplines, ensuring high difficulty and diversity. • Multi-modal Inputs: Incorporate images, graphs, and text, requiring models to process and integrate different data types. • Automated Scoring: Use rule-based and model-based systems to evaluate correctness, reasoning coherence, and calibration. • Multi-turn Reasoning: Design questions that necessitate multiple inference steps, verifying logical consistency. • Knowledge Verification: Implement retrieval mechanisms to validate factual correctness during reasoning. • Evaluation Metrics: Measure accuracy, calibration, and reasoning path plausibility. • Difficulty Control: Ensure questions demand deep understanding, cross-disciplinary reasoning, and cannot be answered via simple retrieval.

Experiments

  • �� Dataset: 2,500 expert-curated questions spanning mathematics, physics, history, literature, and more. • Models: GPT-4, PaLM 2, Claude, and other leading LLMs. • Metrics: Accuracy, calibration error, reasoning path validity, and multimodal comprehension scores. • Hyperparameters: Temperature set to 0.7, reasoning rounds 3-5, retrieval steps enabled. • Protocol: Models answered questions with and without multimodal inputs; ablation studies removed components like multi-turn reasoning to assess impact. • Cross-model comparison: Analyzed performance gaps among architectures, identifying strengths and weaknesses.

Results

  • �� GPT-4 achieves only 35% accuracy on HLE, significantly below human experts’ 92%.• Multi-modal questions, especially involving graphs and images, drop accuracy to 20-30%, highlighting the difficulty of genuine multimodal understanding.• Calibration metrics reveal models are overconfident, with confidence scores poorly aligned with correctness.• Ablation studies show that removing multi-turn reasoning reduces accuracy by 15%, confirming its importance.• The results demonstrate that current models lack the depth of understanding needed for high-level academic reasoning, especially in complex, multimodal contexts.

Applications

  • �� Education: HLE can serve as a benchmark for developing AI tutors capable of high-level academic support. • Scientific research: Assists in automating cross-disciplinary hypothesis testing and knowledge synthesis. • Industry: Enhances AI systems for complex decision-making, legal analysis, and scientific discovery. • Policy: Provides a scientific basis for AI capability regulation and safety standards.

Limitations & Outlook

  • �� The benchmark’s difficulty may limit immediate practical deployment, as models still perform far below human levels. • Automatic evaluation might not fully capture nuanced reasoning or open-ended correctness. • The questions, while expert-designed, may still carry unintentional biases or gaps. • Future work should incorporate human-in-the-loop evaluation and expand multimodal capabilities to address these issues.

Plain Language Accessible to non-experts

Imagine a super tough quiz that tests everything you know—math, history, science, and even how to understand pictures and charts. This quiz is so hard that most people get only a few questions right. Now, think of AI as a student taking this quiz. Right now, AI can do okay on simple questions, but when it faces these tricky, multi-part puzzles with pictures, it struggles a lot. This ‘Humanity’s Last Exam’ is like a final boss battle for AI, designed to see if it truly understands complex ideas, can connect different subjects, and interpret images just like a human expert. It has 2,500 questions, carefully made by top teachers and scientists around the world. The goal is to push AI to its limits, showing us where it still needs to grow—like a coach testing a player’s skills before a big game. If AI can pass this exam, it means it’s really starting to think and learn like a human. This helps us build smarter AI that can help in schools, labs, and even in solving real-world problems. It’s like giving AI a final exam to see if it’s ready to join the team of human experts.

ELI14 Explained like you're 14

Imagine you’re taking a super hard test that covers everything—math, history, science, and even questions with pictures and charts. Some questions are simple, but many are tricky and need you to think carefully, connect different ideas, and understand images. Right now, AI is like a student who can do okay on easy questions, but when faced with these tough puzzles, it often gets lost. This ‘Humanity’s Last Exam’ is like a final boss in a video game, designed to see if AI can really understand complex stuff. It has 2,500 questions made by top teachers and scientists worldwide, so it’s super challenging. The test checks if AI can think step-by-step, understand pictures, and connect knowledge from different subjects—just like a human expert. If AI can pass this exam, it means it’s starting to think more like us, which is exciting because it could help in schools, labs, and solving big problems. It’s like giving AI a final test to see if it’s smart enough to join the team of real human experts. This pushes AI to become better and smarter, helping us build a future where machines understand the world just like we do.

Abstract

Benchmarks are important tools for tracking the rapid advancements in large language model (LLM) capabilities. However, benchmarks are not keeping pace in difficulty: LLMs now achieve over 90\% accuracy on popular benchmarks like MMLU, limiting informed measurement of state-of-the-art LLM capabilities. In response, we introduce Humanity's Last Exam (HLE), a multi-modal benchmark at the frontier of human knowledge, designed to be the final closed-ended academic benchmark of its kind with broad subject coverage. HLE consists of 2,500 questions across dozens of subjects, including mathematics, humanities, and the natural sciences. HLE is developed globally by subject-matter experts and consists of multiple-choice and short-answer questions suitable for automated grading. Each question has a known solution that is unambiguous and easily verifiable, but cannot be quickly answered via internet retrieval. State-of-the-art LLMs demonstrate low accuracy and calibration on HLE, highlighting a significant gap between current LLM capabilities and the expert human frontier on closed-ended academic questions. To inform research and policymaking upon a clear understanding of model capabilities, we publicly release HLE at https://lastexam.ai.

cs.LG cs.AI cs.CL