SkillCoach: Self-Evolving Rubrics for Evaluating and Enhancing Agentic Skill-Use
SkillCoach improves agentic skill-use evaluation with self-evolving rubrics, significantly enhancing assessment quality.
Key Findings
Methodology
SkillCoach derives skill-grounded process rubrics from real rollouts to evaluate agentic skill-use. It assesses trajectories along four dimensions: skill selection, skill following, skill composition, and skill-grounded reflection. The external verifier remains a separate signal, allowing process quality to be distinguished from accidental success.
Key Results
- Experiments show evolved rubrics significantly improve evaluation quality, revealing failures hidden by final accuracy, and provide stronger supervision signals than outcome-only filtering.
- Agents using SkillCoach demonstrate higher skill-use ability in skill-dependent tasks.
- SkillCoach better selects high-quality training trajectories in distractor-augmented libraries.
Significance
SkillCoach's impact lies in its ability to reveal hidden skill-use failures and provide stronger training signals, addressing long-standing pain points in reliable skill use within complex libraries.
Technical Contribution
SkillCoach fundamentally differs from existing methods by providing finer-grained process supervision in skill selection, execution, composition, and reflection through its self-evolving rubrics.
Novelty
SkillCoach is the first to use skill structure for defining process supervision in both evaluation and training, distinct from traditional methods focused solely on final outcomes.
Limitations
- Evolving rubrics in complex skill libraries may require significant computational resources.
- Dependence on rubrics may limit generalization to new tasks.
Future Work
Future directions include extending SkillCoach to support larger skill libraries and exploring its application potential across different domains.
AI Executive Summary
In modern enterprises, the growth of skill libraries makes reliable skill use challenging. SkillCoach introduces a self-evolving rubric framework to evaluate and enhance agentic skill-use. It derives skill-grounded process rubrics from real rollouts, assessing trajectories along four dimensions: skill selection, skill following, skill composition, and skill-grounded reflection. Experiments show that evolved rubrics significantly improve evaluation quality, revealing failures hidden by final accuracy and providing stronger supervision signals than outcome-only filtering.
SkillCoach's innovation lies in its ability to provide finer-grained process supervision in skill selection, execution, composition, and reflection. This method not only reveals hidden skill-use failures but also provides stronger training signals. Experimental results show that agents using SkillCoach demonstrate higher skill-use ability in skill-dependent tasks, especially in libraries with distractor skills.
Despite its significant advantages in skill-use evaluation and training, SkillCoach's application in complex skill libraries may require substantial computational resources. Additionally, dependence on rubrics may limit generalization to new tasks. Future research directions include extending SkillCoach to support larger skill libraries and exploring its application potential across different domains.
Deep Analysis
Background
In modern enterprises, the growth of skill libraries makes reliable skill use challenging. Traditional methods often assume that once a skill is generated or retrieved, the agent can use it correctly. However, as skill libraries expand, this assumption often breaks down. Agents may miss required skills, choose wrong ones, skip key steps, or forget final checks before submission.
Core Problem
Reliable skill use in complex skill libraries is a core problem. Agents need to select the right skills, follow steps, compose workflows correctly, and reflect on outputs. Existing methods often rely on final verifier success, which is too coarse for evaluation and training.
Innovation
SkillCoach introduces a self-evolving rubric framework to evaluate and enhance agentic skill-use. It derives skill-grounded process rubrics from real rollouts, assessing trajectories along four dimensions: skill selection, skill following, skill composition, and skill-grounded reflection. The external verifier remains a separate signal, allowing process quality to be distinguished from accidental success.
Methodology
- �� Derive skill-grounded process rubrics from real rollouts
- �� Assess trajectories along four dimensions: skill selection, skill following, skill composition, and skill-grounded reflection
- �� Keep the external verifier as a separate signal, allowing process quality to be distinguished from accidental success
- �� Use evolved rubrics to select high-quality training trajectories
Experiments
The experimental design includes testing SkillCoach's performance in skill-dependent tasks. Benchmarks used include no-skill, curated-skill, and self-generated-skill settings. Results show that evolved rubrics significantly improve evaluation quality, revealing failures hidden by final accuracy.
Results
Results show that agents using SkillCoach demonstrate higher skill-use ability in skill-dependent tasks, especially in libraries with distractor skills. Evolved rubrics significantly improve evaluation quality, revealing failures hidden by final accuracy and providing stronger supervision signals than outcome-only filtering.
Applications
SkillCoach can be applied in enterprise skill-dependent tasks, helping agents use skills more reliably. Its application in complex skill libraries can improve task success rates and provide stronger training signals.
Limitations & Outlook
Despite its significant advantages in skill-use evaluation and training, SkillCoach's application in complex skill libraries may require substantial computational resources. Additionally, dependence on rubrics may limit generalization to new tasks.
Plain Language Accessible to non-experts
Imagine a kitchen where a chef needs to choose the right tools and steps to complete a dish according to a recipe. SkillCoach acts like a smart assistant, helping the chef make the best choices among complex tools and steps. It focuses not only on whether the final dish is successful but also on whether each step is executed correctly. By continuously learning and adjusting, SkillCoach helps the chef improve cooking efficiency and success rates.
ELI14 Explained like you're 14
Imagine you're playing a game where you need to complete tasks. SkillCoach is like a super helper, guiding you to choose the right items and steps to finish the tasks. It cares not only about whether you complete the task but also about whether you did each step right. By learning continuously, SkillCoach helps you become better at the game!
Glossary
Skill Selection
Evaluates whether the agent invokes the appropriate gold skills while avoiding distractors.
Used as a dimension in SkillCoach to assess skill use.
Skill Following
Evaluates whether the agent follows the key procedures specified by the gold skills and reference solution.
Used as a dimension in SkillCoach to assess skill use.
Skill Composition
Evaluates whether the agent coordinates multiple skills or dependent subprocesses in a valid workflow.
Used as a dimension in SkillCoach to assess skill use.
Skill Reflection
Evaluates whether the agent performs explicit checks before final submission.
Used as a dimension in SkillCoach to assess skill use.
Self-Evolving Rubric
Skill-grounded process rubrics derived from real rollouts for evaluating and enhancing agentic skill-use.
A core innovation of SkillCoach.
Open Questions Unanswered questions from this research
- 1 How can SkillCoach be effectively applied in larger skill libraries?
- 2 What is the potential for SkillCoach's application across different domains?
Applications
Immediate Applications
Enterprise Skill Evaluation
SkillCoach can help enterprises evaluate and enhance agentic skill-use in complex skill libraries.
Long-term Vision
Cross-Domain Skill Application
SkillCoach has the potential to be applied across different domains, helping agents use skills more reliably in complex environments.
Abstract
Skills are becoming a reusable operational layer for LLM agents, encoding SOPs, domain rules, tool workflows, scripts, and validation routines. In realistic skill repositories, overlapping skills make reliable skill-use difficult. Final verifier success is too coarse for both evaluation and training, since an agent may pass through trial and error while selecting distractor skills, skipping required steps, composing workflows incorrectly or omitting final checks. We introduce SkillCoach, a self-evolving rubric framework for evaluating and enhancing agentic skill-use. SkillCoach derives skill-grounded process rubrics from real rollouts and evaluates trajectories along four dimensions: skill selection, skill following, skill composition, and skill-grounded reflection. It keeps the external verifier as a separate outcome signal, allowing process quality to be distinguished from accidental task success. The evolved rubrics further serve as process supervision for selecting high-quality training trajectories. Experiments show that evolved rubrics substantially improve evaluation quality, expose failures hidden by final accuracy, and provide stronger supervision signals than outcome-only filtering for enhancing agentic skill-use.