Do LLM-Generated Skills Make Better AI Data Scientists? A Component Ablation Across Data-Science Workflows
Study finds LLM-generated skills do not significantly improve performance in data science workflows.
Wei-Jung Huang
Study finds LLM-generated skills do not significantly improve performance in data science workflows.
Wei-Jung Huang
EvoSOP enables self-evolving LLM agents by iteratively synthesizing atomic actions into reusable SOPs, boosting success rates by 3-13% and reducing interaction rounds.
Haipeng Ding, Yuexiang Xie, Zhewei Wei et al.
Proposes a concept-based interpretability framework using unsupervised dictionary learning to improve transparency and performance of end-to-end autonomous driving models.
Franz Motzkus, Sebastian Bernhard
TurnOPD employs adaptive turn-depth control and progressive loss normalization to enhance long-horizon agent distillation efficiency, improving validation accuracy.
Yuhang Zhou, Kai Zheng, Haoling Li et al.
Proposes Heading-Specific Activation Steering to causally control tool invocation in five open-source models, validated through geometric and causal analysis.
Yuqi Chen, Vincent Siu, Yang Liu et al.
Proposes in-loop memory for language agents, reducing latency to 100μs and improving task efficiency and accuracy.
Yusuf Khan, Carlo Lipizzi
This study introduces M3Bench, a benchmark for evaluating model editing in medical VLMs, measuring reliability, locality, and generalization across 16,276 clinical questions.
Guli Zhu, Chenwei Wu, Liyue Shen
MetaSkill-Evolve introduces recursive self-improvement via two-timescale meta-skill evolution, boosting task accuracy by +23.54 points on OfficeQA.
Zefeng Wang, Minxi Yan, Jinhe Bi et al.
EvoAgentBench evaluates agent self-evolution via ability transfer across four domains, highlighting limitations of automatic methods.
Xingze Gao, Chuanrui Hu, Hongda Chen et al.
Proposes LLM-as-a-Tutor framework using pairwise comparison and atomic constraints to enhance non-verifiable RL.
Yujin Kim, Namgyu Ho, Sangmin Hwang et al.
The study explores GUI agents' reliance on pixels vs. structure, introducing the Perception-Fusion Gap metric.
Guijia Zhang, Yuxun Chen, Yuheng Qi et al.
Agentic IoT transforms IoT into cognitive systems via autonomous AI agents, enhancing real-time reasoning and adaptive planning.
Rümeysa Hilal Sevinç, Bahaeddin Türkoğlu, İbrahim Kök
BRAID models multi-modal reasoning as a unified MDP, jointly optimizing text and image generation, boosting multi-turn reasoning performance.
Zican Hu, Xuyang Hu, Yiming Liu et al.
Proposed cross-survey transfer framework; zero-shot LLMs achieve 52% accuracy on unseen items, nearing supervised models.
Chan-Tung Ku, Chan Hsu, Pei-Cing Huang et al.
Proposes Raven-Agent, a modular trading layer for prediction markets, achieving +15.9% risk-adjusted return in replay tests.
Yishu Wang, Yuxuan Wang, Jiaqi Deng et al.
Oyster-II uses Zero-RL and SERL for constructive safety, raising Chinese long-query safety by 14.65% over joint training.
Jiyang Guan, Yong Xie, Jun Chen et al.
Proposes EXT++ and PPR to enhance hypergraph RAG, significantly improving fact extraction and chunk retrieval.
Houda Khrouf, Pedro Fillastre, Sebastiao Correia
EvoPolicyGym benchmarks autonomous policy evolution; GPT-5.5 achieves top performance across 16 RL environments with limited interactions.
Zhilin Wang, Han Song, Runzhe Zhan et al.
ScopeEdit employs a dual-branch scope-aware mechanism with orthogonal low-rank geometry to control knowledge propagation in online multimodal model editing.
Siyuan Li, Youyuan Zhang, Ruitong Liu et al.
SkillCoach improves agentic skill-use evaluation with self-evolving rubrics, significantly enhancing assessment quality.
Jiayin Zhu, Kelong Mao, Yudong Guo et al.