WebDancer: Towards Autonomous Information Seeking Agency
WebDancer, based on ReAct, achieves 64.1% Pass@1 on GAIA, significantly improving multi-step web reasoning.
Jialong Wu, Baixuan Li, Runnan Fang et al.
WebDancer, based on ReAct, achieves 64.1% Pass@1 on GAIA, significantly improving multi-step web reasoning.
Jialong Wu, Baixuan Li, Runnan Fang et al.
Combining orthogonality and variance losses enhances MoE expert specialization, achieving up to 23.79% performance gain.
Hongcan Guo, Haolang Lu, Guoshun Nan et al.
ACE-Step combines diffusion, DCAE, and linear Transformer to generate 4-minute music in 20 seconds, vastly improving speed and coherence.
Junmin Gong, Sean Zhao, Sen Wang et al.
BioHopR is a biomedical multi-hop, multi-answer reasoning benchmark based on PrimeKG, evaluating large language models' inference capabilities.
Yunsoo Kim, Yusuf Abdulle, Honghan Wu
LCoT2Tree converts long reasoning into trees, raising correctness prediction by 5.63% on average over length baselines.
Gangwei Jiang, Yahui Liu, Zhaoyi Li et al.
AudioGenie: A training-free multi-agent framework for multimodality-to-multiaudio generation, achieving SOTA performance.
Yan Rong, Jinting Wang, Guangzhi Lei et al.
EPiC generates high-precision anchor videos via masking, eliminating the need for point cloud or pose estimation, enabling efficient 3D-informed camera control.
Zun Wang, Jaemin Cho, Jialu Li et al.
BehaviorSFT enhances clinical agent proactivity using behavioral tokens, achieving 97.3% Macro F1 on BehaviorBench.
Yubin Kim, Zhiyuan Hu, Hyewon Jeong et al.
Introduces ViewSpatial-Bench, a multi-viewpoint spatial localization benchmark, with 46.24% performance gain after fine-tuning.
Dingming Li, Hongxing Li, Zixuan Wang et al.
This study evaluates LLMs in medical summarization, emphasizing vocabulary adaptation's role in high OOV scenarios.
Gunjan Balde, Soumyadeep Roy, Mainack Mondal et al.
This study demonstrates how confidence gaps and information formats influence herd behavior in LLM multi-agent systems, with flip rates up to 0.63 and implications for adaptive collaboration.
Young-Min Cho, Sharath Chandra Guntuku, Lyle Ungar
LLaMEA-BO uses large language models to automatically generate Bayesian optimization algorithms, improving performance on 19 BBOB functions.
Wenhu Li, Niki van Stein, Thomas Bäck et al.
Proposed Fork-Merge Decoding enhances AV-LLMs' multimodal understanding without additional training.
Chaeyoung Jung, Youngjoon Jang, Jongmin Choi et al.
AVCD reduces hallucinations in audio-visual LLMs via contrastive decoding, improving accuracy on AVHBench dataset.
Chaeyoung Jung, Youngjoon Jang, Joon Son Chung
Privacy-preserving P2P energy trading using hybrid secure computations with Paillier encryption and secret sharing.
Junhong Liu, Qinfei Long, Rong-Peng Liu et al.
Automated pipeline SWE-rebench extracts 21,336 interactive Python tasks from GitHub, enabling scalable, decontaminated evaluation of software engineering models.
Ibragim Badertdinov, Alexander Golubev, Maksim Nekrashevich et al.
EgoZero leverages egocentric human demonstrations captured via Project Aria glasses to learn robot manipulation policies with zero robot data, using point cloud-based state-action representations.
Vincent Liu, Ademi Adeniji, Haotian Zhan et al.
Alita uses minimal predefined tools and self-evolution, achieving 75.15% pass@1, outperforming complex systems.
Jiahao Qiu, Xuan Qi, Tongcheng Zhang et al.
Introduces ImgEdit dataset with 1.2M high-quality image pairs, trains ImgEdit-E1 model, and proposes ImgEdit-Bench for comprehensive evaluation, outperforming existing models.
Yang Ye, Xianyi He, Zongjian Li et al.
Proposed a logit-based ensemble method matching human evaluation of LLM consistency.
Xiaoyuan Wu, Weiran Lin, Omer Akgul et al.