SkillsBench: Benchmarking How Well Agent Skills Work Across Diverse Tasks

TL;DR

SkillsBench benchmark shows curated skills improve task pass rate by 16.6 percentage points.

cs.AI 🔴 Advanced 2026-02-13 2 views
Xiangyi Li Yimin Liu Wenbo Chen Bingran You Zonglin Di Yifeng He Shenghan Zheng Kyoung Whan Choe Jiankai Sun Shuyi Wang Chujun Tao Binxu Li Xuandong Zhao Hejia Geng Xiaojun Wu Junwei Zhou Xiaokun Chen Hanwen Xing Yubo Li Qunhong Zeng Di Wang Yuanli Wang Roey Ben Chaim Penghao Jiang Haotian Shen Luyang Kong Xinyi Liu Runhui Wang Xuanqing Liu Jiachen Li Xin Lan Yueqian Lin Wengao Ye Junwei He Songlin Li Yue Zhang Yipeng Gao Yijiang Li Ze Ma Liqiang Jing Tianyu Wang Kaixin Li Yiqi Xue Haoran Lyu Yizhuo He Yuchen Tian Shutong Wu Bowei Wang Yixuan Gao Bo Chen Litong Liu Sikai Cheng Jiajun Bao Shuaicheng Tong Shuwen Xu Terry Yue Zhuo Tinghan Ye Qi Qi Miao Li Longtai Liao Zelin Tan Chang Shi Xilin Tang Srinath Tankasala Boqin Yuan Yaoyao Qian Jianhong Tu Chenguang Wang Yizhou Sun Wei Wang Aaron Taylor Ziyue Yang Changkun Guan Zhikang Dong Xinyu Zhang Steven Dillmann Han-chung Lee Dawn Song
benchmark large language model skill augmentation task pass rate AI

Key Findings

Methodology

The study uses the SkillsBench benchmark to evaluate 87 tasks under no-skills and curated-skills conditions. It analyzes the impact of skill packages on task pass rates using 18 model-harness configurations.

Key Results

  • Result 1: Curated skills increased average pass rate from 33.9% to 50.5%, a 16.6 percentage point improvement.
  • Result 2: In natural science, skill packages showed the most significant improvement, increasing by 28.8 percentage points.
  • Result 3: Small models with skills can match the performance of larger models without skills.

Significance

This study provides a standardized method to evaluate the effectiveness of large language models in complex tasks. SkillsBench allows researchers to systematically quantify the actual contribution of skill packages, advancing AI applications in specialized fields.

Technical Contribution

Introduced the first benchmark treating skills as a primary evaluation artifact, providing a paired evaluation framework that can assess other artifacts without confounding model and augmentation effects.

Novelty

SkillsBench is the first benchmark to systematically evaluate the effect of skill packages on task performance, differing from previous benchmarks that only assessed model capabilities.

Limitations

  • Limitation 1: In some tasks, skill packages may introduce additional overhead, reducing performance.
  • Limitation 2: The effectiveness of skill packages depends on specific model-harness configurations.

Future Work

Future research could explore skill package design across different domains, optimizing modularity and reusability to enhance AI performance in more tasks.

AI Executive Summary

SkillsBench is a new benchmark designed to evaluate the effectiveness of large language models' skills across diverse tasks. The study finds that curated skill packages significantly improve task pass rates, especially in fields requiring specialized knowledge like natural sciences. This benchmark provides a paired evaluation framework, allowing researchers to systematically assess the actual contribution of skill packages without confounding model and augmentation effects. This advancement opens new possibilities for AI applications in specialized domains.

Testing across 87 tasks, the study shows that using curated skill packages can increase average pass rates from 33.9% to 50.5%, a 16.6 percentage point improvement. Notably, in the natural sciences, the effect of skill packages is most pronounced, with a 28.8 percentage point increase. Additionally, the study finds that small models with skills can perform comparably to larger models without skills.

However, the study also highlights that the effectiveness of skill packages depends on specific model-harness configurations and may introduce additional overhead in some tasks. Future research can further optimize skill package design to improve modularity and reusability, enhancing AI performance across more tasks.

Deep Analysis

Background

As AI is increasingly applied in production workflows, effectively evaluating large language models' performance in complex tasks has become crucial. Previous studies focused on model capabilities, neglecting the actual contribution of skill packages in task execution.

Core Problem

There is a lack of systematic methods to quantify the actual contribution of skill packages in task execution, especially in complex tasks requiring specialized knowledge.

Innovation

SkillsBench provides a paired evaluation framework, allowing researchers to systematically assess the actual contribution of skill packages without confounding model and augmentation effects.

Methodology

  • �� Use the SkillsBench benchmark to evaluate 87 tasks under no-skills and curated-skills conditions.
  • �� Analyze the impact of skill packages on task pass rates using 18 model-harness configurations.
  • �� Employ a paired evaluation framework to ensure result reliability.

Experiments

The experimental design includes 87 tasks across 8 domains. It uses 18 model-harness configurations to compare task pass rates under no-skills and curated-skills conditions.

Results

The study finds that curated skill packages can significantly improve task pass rates, especially in fields requiring specialized knowledge like natural sciences.

Applications

SkillsBench can be used to evaluate AI applications in specialized fields, helping researchers optimize skill package design.

Limitations & Outlook

The effectiveness of skill packages depends on specific model-harness configurations and may introduce additional overhead in some tasks.

Plain Language Accessible to non-experts

Imagine you're cooking in a kitchen. A large language model is like a versatile chef who knows how to make many dishes but lacks specific steps for some complex recipes. Skill packages are like detailed cookbooks that provide specific instructions and tips for each dish. By using these cookbooks, the chef can better complete complex dishes. This is what SkillsBench does: it evaluates how effective these cookbooks are for different dishes.

ELI14 Explained like you're 14

Imagine you're playing a game with a super character, but some levels are really hard. Skill packages are like cheat codes in the game, telling you how to get through those tough levels. SkillsBench is a tool to test how effective these cheat codes are. It helps us know which cheat codes really work and which are just a waste of time.

Glossary

Skill Package

A modular procedural package used to enhance AI model task execution.

Used in the paper to improve task pass rates.

Benchmark

A standardized test used to evaluate the performance of models or algorithms.

SkillsBench is used to evaluate the effectiveness of skill packages.

Large Language Model (LLM)

A large-scale neural network model capable of processing and generating natural language.

Used in the study for task execution.

Task Pass Rate

The proportion of tasks successfully completed by a model.

Used to evaluate the improvement from skill packages.

Paired Evaluation

A method of comparing the performance of two different settings under the same conditions.

Used to compare task performance with and without skill packages.

Open Questions Unanswered questions from this research

  • 1 How to design more efficient skill packages for complex tasks across different domains?
  • 2 How to optimize skill package performance across different model-harness configurations?

Applications

Immediate Applications

AI Task Optimization

SkillsBench allows researchers to optimize AI performance in complex tasks, especially in fields requiring specialized knowledge.

Long-term Vision

Cross-Domain AI Applications

SkillsBench can help realize AI applications in more domains, promoting widespread adoption of AI technologies.

Abstract

Agent Skills are structured packages of procedural knowledge that augment large language model (LLM) agents at inference time. Despite rapid adoption, there is no standard way to measure whether they actually help. We present SkillsBench, a benchmark whose current inventory contains 87 tasks across 8 domains paired with curated Skills and deterministic verifiers. Our latest aggregate evaluation runs the 87-task benchmark under matched no-Skills and curated-Skills conditions for 18 model-harness configurations. Curated Skills raise the average pass rate from 33.9% to 50.5% (+16.6 percentage points; 25.5% normalized gain), with configuration-level gains ranging from +4.1 to +25.7 pp. Focused Skills with at most three modules outperform larger or exhaustive bundles, and smaller models with Skills can match larger models without them. SkillsBench establishes paired evaluation as the foundation for rigorous measurement of Skill efficacy on agentic, expertise-heavy work.

cs.AI