ASI-Bench: At the Dawn of Artificial Superintelligence
ASI-Bench evaluates AI's autonomous scientific exploration; scores drop from 50.91 to 26.62 as guidance decreases, showing reliance on human input.
Key Findings
Methodology
ASI-Bench, developed by over 40 experts with 31,000+ hours, features 60 complex research tasks across 11 domains. It employs a progressive guidance reduction strategy, from full procedural instructions (B1) to no guidance (B3), assessing AI’s ability to independently formulate methods, conduct experiments, and generate verifiable results. Tasks undergo rigorous expert review, AI-assisted auditing, sandbox execution, and scoring validation to ensure scientific validity. Across 18 state-of-the-art models, performance declines from an average of 50.91 under full guidance to 26.62 in autonomous mode, highlighting current limitations in AI’s scientific autonomy.
Key Results
- In 18 top configurations, the average score drops from 50.91 with full guidance to 29.10 when only method names are provided, and further to 26.62 when models must determine methods independently, indicating significant dependency on human guidance.
- Models like GPT-5.6 Ultra achieve a maximum score of 71.78 under full guidance but only 26.62 autonomously, illustrating the gap in autonomous scientific reasoning.
- The main bottleneck is transforming scientific methods into complete workflows; performance drops sharply when procedural details are removed, emphasizing the importance of operationalization capabilities.
Significance
This benchmark pioneers a comprehensive evaluation of AI’s ability to conduct autonomous scientific research, addressing a critical gap in existing assessments focused on knowledge recall or task execution. It provides a quantifiable measure of progress toward artificial superintelligence, encouraging development of models capable of independent exploration and innovation, thus advancing both AI research and scientific discovery.
Technical Contribution
The framework introduces a multi-stage guidance reduction paradigm, combining cross-disciplinary tasks with rigorous validation, to systematically measure autonomous research capabilities. It emphasizes the interaction between models and their operational systems, fostering innovations in model-operationalization and multi-stage reasoning. This approach sets a new standard for evaluating AI’s scientific autonomy and guides future system design.
Novelty
This is the first benchmark to systematically evaluate AI’s ability to independently formulate research strategies, select methods, and produce verifiable results across multiple scientific domains, moving beyond traditional knowledge-based assessments. Its progressive guidance reduction approach uniquely captures the transition from assisted to autonomous research, representing a significant innovation in AI evaluation.
Limitations
- Current systems still heavily depend on human guidance, especially in complex, open-ended tasks, indicating that true autonomous research remains distant.
- High computational costs and reliance on large-scale models limit practical deployment and scalability.
- Tasks, though diverse, do not yet encompass the full spectrum of scientific challenges, necessitating future expansion to include more complex and interdisciplinary problems.
Future Work
Future directions include enhancing models’ operationalization abilities, integrating multi-modal data, and improving reasoning and decision-making processes. Expanding task diversity and complexity will better reflect real-world scientific challenges. Community contributions of new tasks and evaluation methods will also be encouraged to refine and extend ASI-Bench’s capabilities, pushing AI closer to genuine autonomous scientific discovery.
AI Executive Summary
As artificial intelligence continues to evolve, the aspiration for autonomous scientific discovery has become a central goal. Traditional benchmarks primarily evaluate AI’s knowledge recall and procedural execution, but fall short in measuring genuine innovation and independence. Addressing this gap, ASI-Bench was developed as the first comprehensive, multi-domain benchmark designed to assess AI’s capacity for autonomous research.
Constructed over 31,000 hours by a team of experts, ASI-Bench encompasses 60 complex research tasks spanning fields like physics, chemistry, biology, and computer science. Each task simulates real scientific investigations, requiring the AI to understand problems, select appropriate methods, conduct experiments, and validate results. The benchmark employs a progressive guidance reduction strategy, starting with full procedural instructions and gradually removing human guidance, to evaluate how well AI can operate independently.
Experimental results reveal a stark performance decline as guidance diminishes. The best models achieve scores over 70 under full guidance but only around 26 in fully autonomous settings. This highlights a significant gap between current AI capabilities and the goal of autonomous scientific reasoning. The main challenge lies in operationalizing scientific methods—transforming procedural knowledge into complete workflows without human intervention.
This benchmark offers a crucial tool for tracking progress toward artificial superintelligence, emphasizing the importance of autonomous exploration and innovation. It encourages the AI community to develop systems capable of independent hypothesis generation, method formulation, and experimental validation. Although current systems show promise, substantial work remains to realize fully autonomous scientific discovery, making ASI-Bench a vital step in this journey. Future efforts will focus on improving models’ operationalization, expanding task diversity, and fostering collaborative community contributions to accelerate progress.
Deep Dive
Plain Language Accessible to non-experts
想象一个厨房里做菜,刚开始厨师(AI)需要厨师长详细指导每个步骤,比如切菜、炒菜、调味。随着时间推移,厨师长逐渐减少指导,只告诉他做菜的目标,比如“做一道意大利面”。最后,厨师自己决定用什么调料、怎么炒,自己完成整个过程。这个实验就像ASI-Bench测试的逐步减少指导,让我们知道这个厨师(AI)能不能自己独立完成复杂的菜肴。结果显示,刚开始他表现不错,但完全靠自己做时,还差得远。这说明,要让AI真正自主做科研,还需要很多训练和改进,就像厨师学做菜一样。
Abstract
Artificial superintelligence (ASI) requires AI to move beyond mastering existing knowledge toward exploring the unknown, creating new knowledge, and turning new ideas into verifiable results. However, the capabilities of today's AI systems are still largely built on learning, compressing, and applying existing human knowledge. Accordingly, existing benchmarks primarily test whether AI can produce correct answers based on learned knowledge, or whether it can complete tasks under extensive human guidance. We therefore introduce ASI-Bench, the first benchmark to jointly evaluate AI systems' capabilities of innovative exploration and autonomous scientific execution across general research domains, and the first to progressively withdraw human methodological guidance within the same research project to test how far AI can proceed on its own. Built by over 40 experts with the cost of 31,000+ human hours, ASI-Bench contains 60 project-level research tasks across 11 scientific domains and progressively reduces methodological guidance to test whether AI can independently select methods, conduct research, and produce verifiable results. All tasks undergo expert review, AI-assisted auditing, sandbox execution, and scorer validation. Across 18 state-of-the-art agent--model configurations, the average score drops from 50.91 with full methodological guidance to 29.10 with only the method specified and 26.62 when agents must determine the method themselves. This sharp decline shows that current systems remain heavily dependent on human guidance and are still far from autonomously conducting end-to-end, project-level scientific research. ASI-Bench is open to the world. We invite researchers and builders everywhere to contribute new tasks, challenge the limits of today's AI, and help accelerate humanity's collective path toward artificial superintelligence at https://asibench.apexin.ai/submit.