SkillTV-Bench: Benchmarking How Well Judges Perform on Skill-Augmented Agentic Execution

TL;DR

SkillTV-Bench benchmarks trajectory verification, improving accuracy by 14.8 points via automated JudgeSkill evolution.

cs.AI 🔴 Advanced 2026-08-06 43 views
Zhi Han Chenxi Zeng Liuhaichen Yang Zihan Guo Ming Zhou Yang Li
multi-domain trajectory verification skill-aware auto-evolution benchmark

Key Findings

Methodology

This work introduces SkillTV-Bench, a comprehensive benchmark with 681 real agent trajectories across 50 tasks and 11 domains, integrating task-time skills and inspectable artifacts. It supports multi-skill, evidence-grounded evaluation. The core innovation, SkillTV-Evolve, externalizes verification knowledge as a reusable JudgeSkill, guiding an agent judge to plan inspections, gather evidence, and issue verdicts. An automated evolution loop refines JudgeSkill by analyzing misjudged cases, leading to performance improvements. Evaluation involves multiple models like GPT-5.2 and Claude Sonnet 4.6, with metrics such as balanced accuracy and success rate, demonstrating significant gains after iterative evolution.

Key Results

  • On the SkillTV-Bench evaluation set, the refined JudgeSkill increased balanced accuracy from 56.8% to 58.6%, with success rate in trajectory selection rising from 22.9% to 45.5%. Multi-round rollouts further boosted success, confirming the effectiveness of the verification-guided selection strategy.
  • Across nine domains, the largest improvements were observed in media content (+36.7%) and manufacturing (+23.1%), indicating broad applicability. The approach effectively distinguishes successful from plausible-failed trajectories, especially in complex, skill-augmented scenarios.
  • The automated JudgeSkill evolution, driven by misjudged cases, systematically enhanced verification accuracy, validating the potential for continuous, benchmark-guided improvement of autonomous judgment systems.

Significance

This research addresses a critical gap in long-horizon task evaluation by incorporating task-time skills and evidence-grounded verification, significantly enhancing the reliability of AI agents in complex environments. The methodology offers a scalable, iterative framework for improving autonomous judgment, which is vital for deploying trustworthy AI in real-world applications such as robotics, autonomous vehicles, and industrial automation. By enabling continuous skill refinement without retraining models, it paves the way for more adaptable and dependable AI systems, ultimately contributing to safer and more efficient autonomous operations.

Technical Contribution

The paper introduces SkillTV-Bench, a large-scale, multi-domain benchmark for trajectory verification that emphasizes skill-aware, evidence-grounded evaluation. It proposes SkillTV-Evolve, an external, reusable JudgeSkill that guides inspection planning and evidence collection. The approach leverages an automated evolution loop to iteratively improve JudgeSkill based on misjudged cases, without updating underlying models. This combination of benchmark design, external knowledge externalization, and iterative optimization constitutes a novel paradigm for scalable, skill-aware trajectory verification, surpassing prior static or final-response-focused methods.

Novelty

This work is the first to systematically incorporate task-time skills into trajectory verification, using an external, evolvable JudgeSkill to guide inspections. Unlike previous benchmarks that focus on static responses or simple trajectories, SkillTV-Bench emphasizes dynamic, evidence-based validation across multiple domains. The automated evolution of verification knowledge, driven by misjudged cases, introduces a new paradigm for continuous improvement in AI judgment systems, setting a foundation for future research in skill-aware, scalable verification.

Limitations

  • The approach relies heavily on the availability of detailed task-time skills and inspectable artifacts, which may not be feasible in all real-world scenarios, especially where environment access is limited.
  • Automated evolution, while effective, incurs significant computational costs and may not be suitable for real-time applications.
  • Generalization to highly novel or extremely complex tasks remains uncertain, requiring further validation across diverse, unseen environments.

Future Work

Future directions include integrating multi-modal evidence sources, applying reinforcement learning for autonomous skill refinement, and extending the framework to multi-agent systems. Additionally, efforts will focus on reducing computational overhead and enhancing generalization to unseen tasks, aiming for real-time, scalable verification solutions that can adapt to evolving environments and task complexities.

AI Executive Summary

The rapid advancement of large language models (LLMs) has transformed AI from simple response generators into interactive agents capable of planning, tool use, and environment manipulation. However, evaluating these agents remains a challenge, especially for long-horizon, skill-augmented tasks. Traditional benchmarks focus on final outputs, neglecting the complex, multi-step processes involved in real-world tasks. This gap hampers the development of reliable, trustworthy AI systems.

To address this, the authors introduce SkillTV-Bench, a comprehensive benchmark comprising 681 real agent trajectories across 50 tasks and 11 domains. This benchmark uniquely incorporates task-time skills and inspectable artifacts, enabling a detailed, evidence-grounded evaluation of agent behavior. It supports multi-domain, skill-aware assessment, revealing the limitations of existing judges that often accept plausible failures or overlook critical errors.

Building on this, the paper proposes SkillTV-Evolve, a system that externalizes verification knowledge as a reusable JudgeSkill. This external knowledge guides an agent judge to plan inspections, gather relevant evidence, and issue verdicts. An automated evolution loop further refines JudgeSkill by analyzing misjudged cases, leading to continuous performance improvements. Experimental results demonstrate that this approach boosts judgment accuracy by 14.8 percentage points and significantly enhances trajectory selection success rates.

The broader impact of this work lies in its potential to improve the reliability and safety of autonomous systems across industries such as robotics, content moderation, and industrial automation. By enabling continuous, skill-aware verification, the framework offers a scalable pathway toward trustworthy AI deployment. Despite promising results, challenges remain in reducing computational costs and extending the approach to more complex, unseen environments. Future research will focus on multi-modal evidence integration, reinforcement learning-based skill refinement, and real-time verification, aiming to make autonomous systems more dependable and adaptable in dynamic settings.

Deep Dive

Plain Language Accessible to non-experts

想象你在厨房里做一道复杂的菜。这道菜需要按照很多步骤操作,比如切菜、调味、煮饭,每一步都很重要。有时候你会用不同的工具,比如刀、锅,还要遵守一些规则,比如火不能太大。整个做菜的过程就像一场长长的任务,不能只看最后的成品,要确保每一步都正确。这个研究就像是在教一个聪明的厨师助手,它不仅看成品,还会检查每个步骤是否按照规程操作,确保菜做得好吃又安全。它会用一些特别的规则,告诉你在哪些地方可能出错,帮你提前发现问题。通过不断练习和改正,这个助手变得越来越厉害,能更快帮你找到错误,让你做菜更顺利。这就像是让厨房变得更聪明、更可靠的秘密武器!

ELI14 Explained like you're 14

假设你在玩一款很复杂的游戏,里面有很多任务,比如建房子、打怪、收集宝藏。每个任务都需要你按顺序做很多事情,不能只看最后的奖励。这个研究就像是在教一个超级聪明的朋友,不仅帮你打怪,还能在你玩的时候检查你是不是每一步都做对了。它会用一些特别的规则,告诉你在哪些地方可能出错,帮你提前发现问题。通过不断练习和改正,这个朋友变得越来越厉害,能更快帮你找到错误,让你玩得更顺利。这就像是给你一个超级助手,让你在游戏中变得更厉害、更聪明!

Abstract

LLM agents increasingly execute long-horizon tasks through tool use and environment interaction, shifting evaluation from final-response scoring to verification of complete executions. For skill-augmented agents, verification additionally requires the procedural knowledge encoded in task-time skills, because this knowledge indicates what evidence to inspect and which failures are task-critical. However, existing judge benchmarks often expose final responses or static trajectories, and rarely combine task-time skills with directly inspectable artifacts and environments. We therefore introduce SkillTV-Bench, a 681-case benchmark of real agent trajectories from 50 tasks across eleven domains, designed to evaluate skill-aware trajectory verification for both LLM-as-a-Judge and Agent-as-a-Judge methods. Additionally, we propose SkillTV-Evolve, which externalizes verification knowledge as a reusable JudgeSkill that guides an agent judge to plan targeted inspections and issue evidence-grounded verdicts. On a disjoint development pool, an automated evolution loop further refines the JudgeSkill using misjudged cases. On SkillTV-Bench, the refined skill increases the same agent judge's accuracy by 14.8 percentage points. In offline rollout-pool selection, it increases selected-trajectory success from 22.9% with one rollout to 45.5% with ten rollouts. The code and data are available at https://github.com/HanZhi306/SkillTV-Bench

cs.AI