Goal2Skill: Long-Horizon Manipulation with Adaptive Planning and Reflection
Goal2Skill uses a dual-system with VLM-based planning and diffusion-based control, achieving 32.4% success in long-horizon tasks.
Zhen Liu, Xinyu Ning, Zhe Hu et al.
Goal2Skill uses a dual-system with VLM-based planning and diffusion-based control, achieving 32.4% success in long-horizon tasks.
Zhen Liu, Xinyu Ning, Zhe Hu et al.
ASTER generates pseudo-anomalies in latent space, combining Transformer and pre-trained LLMs for unsupervised time-series anomaly detection, outperforming existing methods.
Romain Hermary, Samet Hicsonmez, Dan Pineau et al.
Doc-V* enhances multi-page document VQA by sequential evidence aggregation, outperforming RAG baseline by 47.9%.
Yuanlei Zheng, Pei Fu, Hang Li et al.
Proposed BVE framework achieves high-quality 3D editing using 3D masks and self-constructed data, significantly improving editing performance.
Yizhao Xu, Hongyuan Zhu, Caiyun Liu et al.
Proves that a single-layer linear Transformer is mathematically equivalent to Ordinary Least Squares (OLS) regression via spectral decomposition, enabling one-pass statistical inference.
Xiaojun Tan, Yuchen Zhao
Proposes Proxy Compression Hypothesis (PCH) to explain reward hacking, emphasizing goal compression, optimization amplification, and evaluator-policy co-adaptation.
Xiaohua Wang, Muzhao Tian, Yuqi Zeng et al.
RiskWebWorld offers 1,513 tasks showcasing challenges for GUI agents in e-commerce risk management, with top models achieving 49.1% success.
Renqi Chen, Zeyin Tao, Jianming Guo et al.
LAMO framework enhances lightweight GUI agents' task scalability via multi-role orchestration, achieving 52.6% accuracy on ScreenSpot-pro.
Ziwei Wang, Junjie Zheng, Leyang Yang et al.
Propose a lightweight framework using KL divergence to analyze quantization sensitivity in mixed-precision SSM-Transformer models.
Jason Kong, Nilesh Prasad Pandey, Flavio Ponzina et al.
Proposes dual macro-level framework distinguishing functional and ontological creativity in AI agents, highlighting current limitations and future paths.
Giorgio Franceschelli, Mirco Musolesi
Proposes InfiniteScienceGym, a seed-based, verifiable scientific benchmark for evaluating reasoning and tool use, with models achieving only ~50% accuracy.
Oliver Bentham, Vivek Srikumar
This paper systematically analyzes the training dynamics of on-policy distillation (OPD) for large language models, revealing that success hinges on thinking-pattern alignment and new capabilities.
Yaxuan Li, Yuxin Zuo, Bingxiang He et al.
Introduces Touch Dreaming-enhanced HTD model, achieving 90.9% success in humanoid contact-rich manipulation via multimodal Transformer.
Yaru Niu, Zhenlong Fang, Binghong Chen et al.
Parcae, a stable looped language model constrained by spectral norm limits, reduces perplexity by 6.3% and surpasses Transformer baselines at 1.3B parameters.
Hayden Prairie, Zachary Novack, Taylor Berg-Kirkpatrick et al.
Tree Learning employs hierarchical parameter inheritance to enable multi-skill continual learning in humanoid robots, achieving 100% skill retention and smooth switching.
Yifei Yan, Linqi Ye
Proposes CAMFusion, a multiview transformer for fusing vision-language embeddings, achieving state-of-the-art 3D semantic classification.
Tomas Berriel Martins, Martin R. Oswald, Javier Civera
Proposes Fun-TSG, a function-driven multivariate time series generator with variable-level anomaly labels, enabling controllable, transparent synthetic data creation.
Pierre Lotte, André Péninou, Olivier Teste
SD-Zero transforms binary rewards into dense token-level supervision via self-revision and self-distillation, boosting reasoning performance by over 10%.
Yinghui He, Simran Kaur, Adithya Bhaskar et al.
V-Nutri fuses final-dish and cooking keyframes; HD-EPIC experiments show process cues can improve nutrition estimation.
Chengkun Yue, Chuanzhi Xu, Jiangpeng He
Meerkat combines clustering with agentic search, finding roughly 4× more CyBench reward-hacking cases than prior audits.
Adam Stein, Davis Brown, Hamed Hassani et al.