ECHO: Prune To Act, Trace To Learn With Selective Turn Memory In Agentic RL
ECHO achieves 43.4% accuracy in Agentic RL using selective turn memory, outperforming GRPO and SUPO.
Zijun Xie, Binbin Zheng, Enlei Gong et al.
ECHO achieves 43.4% accuracy in Agentic RL using selective turn memory, outperforming GRPO and SUPO.
Zijun Xie, Binbin Zheng, Enlei Gong et al.
DualBrep encodes CAD models via dual scalar fields in a unified latent space, enabling end-to-end geometry and topology generation.
Yilin Liu, Pradeep Jayaraman, Chinthala Reddy et al.
Localized Conformal Prediction (LCP) combined with vision-language models and nonlinear cosine similarity transformation reduces average prediction set size while maintaining coverage.
Clément Fuchs, Tim Bary, Benoît Macq
MV-GEL employs multi-view ranking and VLM reasoning to localize geometric entities on meshes, achieving up to 1.7× IoU improvement.
Kartik Bali, Roland Aydin
SimpleSearch-VL employs FAR for efficient sampling, integrates evidence verification for reliability, and maintains lightweight tools, significantly boosting multimodal agentic search performance.
Ming Dai, Zhihong Lu, Jinjie Gu et al.
ChronoFlow-Policy unifies past-current-future interaction flow via sparse 3D keypoints, boosting robot manipulation success by 20%.
Bokai Lin, Yifu Xu, Xinyu Zhan et al.
SAGE employs Multi-Hypothesis Failure Attribution to boost autonomous research reliability from 42% to 92%.
Jie Ma, Binfei Chu, Jie Gao et al.
Xiaomi-GUI-0 achieves 72.0% success on real devices, significantly improving stability.
Wanxia Cao, Chengzhen Duan, Pei Fu et al.
Learning from Failure: Using failed trajectories for inference-time self-improvement, achieving a 6.6% success rate increase.
Xueqiao Sun, Xiaohan Wang, Ludwig Schmidt et al.
Embodied CAD integrates LLM planning with geometric solvers via a hierarchical skill library, enabling reliable parametric assembly modeling.
Fumin Liu, Haoyu Zhou, Fei Hao et al.
Delta-JEPA introduces a latent difference decoder to enhance action-sensitive world modeling, achieving over 89% planning success in continuous control tasks.
Zhenghao Zhang, Yuanxiang Wang, Zhenyu Guan et al.
ForgeDrive achieves 90.3 EPDMS on NAVSIM by unifying driving simulation and planning via visual-action cross-conditioning.
Xuchang Zhong, He Zheng, Chenxu Zhao et al.
RosettaSim employs structured autoregressive modeling with attention mechanisms and semantic retrieval, achieving state-of-the-art long-term traffic simulation with a correlation of r=0.83.
Lingyu Xiao, Zexin Feng, Xintao Yan
AETDICE unifies AET framework for offline nonlinear multi-objective RL, bridging SER and ESR paradigms with a novel decomposition approach.
Woosung Kim, Youngjun Suh, Jinho Lee et al.
Proposes View-PNDF, a neuron-level fine-tuning method that enhances multi-view X-ray report consistency, achieving state-of-the-art results.
Yucheng Chen, Jinjing Zhu, Yang Yu et al.
Proposes a faithfulness evaluation framework combining Lean compilation and semantic judging, revealing a 30% gap between compilation success and semantic accuracy on 400 statements.
Ke Zhang, Patricio Gallardo Candela, Sudhir Murthy et al.
Proposes Turing-inspired Label Imitation Game (LIG) with Transformer-based TTN for zero-shot pseudo-label pruning, improving detection F1 by up to 44%.
Brent A. Griffin, Jason J. Corso
GradeSQL framework enhances Text-to-SQL reliability using ORM, achieving a 4.33% gain on BIRD.
Mattia Tritto, Giuseppe Farano, Dario Di Palma et al.
Synthesis for Time Window Temporal Logic via MILP enhances control input robustness.
Philip Smith, Ahmad Ahmad, Kevin Leahy
Proposes DERAIL attack exploiting scoring head vulnerability in generative E2E autonomous driving, causing collision rates up to 50%.
Halima Bouzidi, Mboutidem Ekemini Mkpong, Haoyu Liu et al.