Orchestra-o1: Omnimodal Agent Orchestration
Orchestra-o1 framework enhances multimodal agent collaboration, achieving 10.3% accuracy improvement on the OmniGAIA benchmark.
Fan Zhang, Vireo Zhang, Shengju Qian et al.
Orchestra-o1 framework enhances multimodal agent collaboration, achieving 10.3% accuracy improvement on the OmniGAIA benchmark.
Fan Zhang, Vireo Zhang, Shengju Qian et al.
This paper introduces feedback alignment in self-distillation, comparing three feedback types; structure-aligned critique outperforms others with +16.11% accuracy.
Semih Kara, Oğuzhan Ersoy
Proposes HiViG, a history-aware visually grounded test-time framework, boosting GUI task success rates by 5.8% (Qwen3-VL-32B) and 9% (Gemini-3-Flash) through macro-action history and visual error verification.
Jaewoo Lee, Zaid Khan, Archiki Prasad et al.
Proposes WorldKernel, a positive semidefinite coupling kernel capturing unobserved cross-world dependencies, addressing prediction's inability to express counterfactual uncertainty.
Fabio Rovai
Proposes skill coverage metric based on trajectory detection of skill behavior constraints, achieving 38.66%-45.51% coverage; improves failed task recovery by 16%.
Boyin Tan, Xiaowei Huang, Youcheng Sun
Trace2Policy transforms expert behavior into self-evolving decision agents using EISR, achieving 79.6% accuracy.
Junli Zha, Jinbo Wang, Chao Zhou et al.
Proposed Adaptive Regime Routing (ARR) to address LLM knowledge conflicts, improving resistance EM from 6 to 16-33.
Runze Jiang, Taiqiang Wu, Yan Wang et al.
Proposes a logic-guided data extraction framework combining ASP and LLMs, reducing calls by 40% while maintaining accuracy.
Mario Alviano, Lorenzo Grillo, Nicola Leone et al.
PRISM directly decodes instruction sets from model activations, outperforming activation-to-language baselines with 97%+ recall in security scenarios.
Gilad Gressel, Rahul Pankajakshan, Julia Diament et al.
Proposes STRP with tree convolution and inverse dilated convolution for fine-grained traffic prediction, improving accuracy by 20% and reducing training time by 30%.
Shuhao Li, Weidong Yang, Yue Cui et al.
REFLECT diagnoses error steps via controlled replay and outcome flip verification, achieving top localization accuracy in LLM traces.
Xiaofeng Lin, Yingxu Wang, Tung Sum Thomas Kwok et al.
Study finds LLM evaluators struggle to adapt to varying contexts and safety definitions, though they can learn from new information.
Anissa Alloula, Federico Licini, Ava Batchkala et al.
Using a task-based framework, real-world data from Perplexity shows AI agents significantly boost automation, efficiency, and task scope, with productivity gains of up to 87%.
Jeremy Yang, Kate Zyskowski, Noah Yonack et al.
This study compares declarative skill files and imperative state machines in AI tool use, showing that skills improve accuracy under high-quality knowledge retrieval.
M. Danish Lim, I. Danial Bin Sharudin, Wen Han Chen et al.
Skill creation via RWSA decomposition; W2S improves replay consistency by 10.5%.
Yuyang Zhang, Xinyuan Han, Xudong Jiang et al.
CaRE framework standardizes compute, metrics, and stochasticity in MDLM evaluation, revealing temperature and compute effects on strategy rankings.
Yash Shah, Abhijit Chakraborty, Vivek Gupta
OpenSkill framework enhances LLM agents' skill transfer in open-world settings without supervision, achieving top automated pass rates.
Zhiling Yan, Dingjie Song, Hanrong Zhang et al.
MLEvolve is a self-evolving multi-agent framework using LLMs for end-to-end machine learning algorithm discovery, achieving 65.3% medal rate within 12 hours.
Shangheng Du, Xiangchao Yan, Jinxin Shi et al.
Classified 10 agent memory systems, analyzed their system behaviors, and proposed 10 optimization strategies.
Yasmine Omri, Ziyu Gan, Zachary Broveak et al.
EvoSQL employs memory-augmented co-evolution with execution signals and LLM critique, boosting complex Text-to-SQL accuracy by up to 9.19%.
Jiawei Zhou, Jianwei Wang, Chenyu Zhou et al.