SciAgentGym: Benchmarking Multi-Step Scientific Tool-use in LLM Agents
SciAgentGym enhances scientific tool-use via SciForge; SciAgent-8B outperforms Qwen3-VL-235B-Instruct.
Key Findings
Methodology
SciAgentGym is an interactive environment with 1,780 domain-specific tools supporting multidisciplinary scientific reasoning. SciAgentBench evaluates capabilities from elementary actions to long-horizon workflows. SciForge generates logic-aware training trajectories via dependency graphs, enhancing model tool-use capabilities.
Key Results
- SciAgent-8B surpasses Qwen3-VL-235B-Instruct in cross-domain tool-use capabilities, achieving a 6.7% improvement.
- GPT-5's success rate in long-horizon tasks drops from 58.8% to 34.6%, highlighting tool-use bottlenecks.
- SciAgent-8B shows positive cross-domain transfer in multi-domain tasks.
Significance
This research provides a new perspective on automating scientific reasoning, filling gaps in existing benchmarks for tool-use capability evaluation. SciAgentGym allows researchers to better understand and enhance model performance in complex scientific tasks.
Technical Contribution
SciAgentGym offers an extensible environment supporting multidisciplinary tool use. SciForge synthesizes logic-aware training data, enhancing scientific tool-use capabilities. SciAgent-8B outperforms larger models in performance.
Novelty
This is the first systematic evaluation of scientific tool-use capabilities across multiple disciplines. SciForge enhances logical reasoning by generating training data via dependency graphs.
Limitations
- Current models perform poorly in long-horizon tasks, especially in complex tool-use scenarios.
- While SciAgentGym's toolkit is extensive, it may not cover all scientific domains' needs.
Future Work
Future work will focus on expanding the diversity and complexity of the toolkit and improving model performance in long-horizon tasks.
AI Executive Summary
Scientific reasoning requires integrating sophisticated toolkits, often overlooked by existing benchmarks. SciAgentGym fills this gap by providing an interactive environment with 1,780 domain-specific tools. SciAgentBench evaluates capabilities from elementary actions to long-horizon workflows. Our study finds that current state-of-the-art models still struggle with complex scientific tool-use, especially in long-horizon tasks. To address this, we propose SciForge, a data synthesis method that generates logic-aware training trajectories via dependency graphs. By fine-tuning on these trajectories, SciAgent-8B outperforms the significantly larger Qwen3-VL-235B-Instruct. Our results underscore the promising potential of next-generation autonomous scientific agents.
Deep Analysis
Background
Scientific reasoning increasingly relies on tool-assisted workflows, from molecular simulations to large-scale data analysis. Solving these scientific problems necessitates deploying tools, as solutions rarely emerge from direct inference but through extensive trial-and-error processes.
Core Problem
Existing scientific benchmarks predominantly target static question answering, failing to capture the interactive, tool-mediated nature of actual scientific workflows. With surging interest in developing capable scientific agents, there is an urgent need for an evaluation framework that mirrors real-world scientific reasoning.
Innovation
We introduce SciAgentGym, a hierarchical interactive environment designed for grounding LLM agents in multi-turn tool-use scientific reasoning tasks. The framework seamlessly integrates 1,780 domain-specific tools across Physics, Chemistry, Biology, and Materials Science.
Methodology
- �� SciAgentGym provides an interactive environment integrating 1,780 tools.
- �� SciAgentBench evaluates capabilities from elementary actions to long-horizon workflows.
- �� SciForge generates logic-aware training trajectories via dependency graphs.
Experiments
We evaluated multiple models on SciAgentBench, finding that even state-of-the-art models show significant performance drops in long-horizon tasks. The experiments used various datasets and benchmarks, focusing on tool-use efficiency and accuracy.
Results
SciAgent-8B surpasses Qwen3-VL-235B-Instruct in cross-domain tool-use capabilities, achieving a 6.7% improvement. GPT-5's success rate in long-horizon tasks drops from 58.8% to 34.6%.
Applications
SciAgentGym can be used to evaluate and enhance scientific agents' performance in complex tasks, particularly those requiring multi-step reasoning and tool use.
Limitations & Outlook
Current models perform poorly in long-horizon tasks, especially in complex tool-use scenarios. While SciAgentGym's toolkit is extensive, it may not cover all scientific domains' needs.
Plain Language Accessible to non-experts
Imagine you're in a kitchen cooking a meal, and SciAgentGym is like a kitchen full of various utensils and ingredients. You need to choose the right tools (scientific tools) based on the recipe (scientific task) to complete a dish (solve a problem). SciForge acts like a smart assistant, guiding you on which tool to use at each step, helping you complete the dish faster and better.
ELI14 Explained like you're 14
Imagine you're playing a complex game where you need different tools to solve puzzles. SciAgentGym is like a big warehouse in the game with all kinds of tools. SciForge is like a super guide that tells you how to use these tools to win. You can use these tools to complete tasks, just like beating the big boss in a game!
Glossary
SciAgentGym
An interactive environment with 1,780 domain-specific tools for scientific reasoning tasks.
Used to evaluate and enhance model performance in complex scientific tasks.
SciAgentBench
An evaluation suite for testing capabilities from elementary actions to long-horizon workflows.
Used to quantify the gap between tool availability and mastery.
SciForge
A data synthesis method generating logic-aware training trajectories via dependency graphs.
Enhances model performance in scientific tool-use.
Long-horizon tasks
Complex tasks requiring multi-step reasoning and tool use.
Used in SciAgentBench to test model capabilities.
Tool-use capability
The ability of models to effectively call and combine tools in scientific tasks.
Core evaluation metric in SciAgentGym and SciForge.
Open Questions Unanswered questions from this research
- 1 How to improve model performance in long-horizon tasks, especially in complex tool-use scenarios.
- 2 Whether SciAgentGym's toolkit is extensive enough to cover all scientific domains' needs.
Applications
Immediate Applications
Scientific Research
Researchers can use SciAgentGym to evaluate and enhance model performance in complex scientific tasks.
Long-term Vision
Automated Scientific Discovery
SciAgentGym could drive automation in scientific discovery, enabling more complex reasoning and tool use.
Abstract
Scientific reasoning inherently demands integrating sophisticated toolkits to navigate domain-specific knowledge. Yet, current benchmarks largely overlook agents' ability to orchestrate tools for such rigorous workflows. To bridge this gap, we introduce SciAgentGym, a scalable interactive environment featuring 1,780 domain-specific tools across four natural science disciplines, supported by a robust execution infrastructure. Complementing this, we present SciAgentBench, a tiered evaluation suite designed to stress-test agentic capabilities from elementary actions to long-horizon workflows. Our evaluation identifies a critical bottleneck: state-of-the-art models still struggle with complex scientific tool-use, and their performance degrades substantially as interaction horizons extend. To address this, we propose SciForge, a data synthesis method that models the tool action space as a dependency graph to generate logic-aware training trajectories. By fine-tuning on these trajectories, our SciAgent-8B outperforms the significantly larger Qwen3-VL-235B-Instruct while exhibiting positive cross-domain transfer of scientific tool-use capabilities. These results underscore the promising potential of next-generation autonomous scientific agents.