Relational Semantic Reasoning on 3D Scene Graphs for Open World Interactive Object Search
SCOUT method uses 3D scene graphs for open-world interactive object search, enhancing efficiency.
Imen Mahdi, Matteo Cassinelli, Fabien Despinoy et al.
SCOUT method uses 3D scene graphs for open-world interactive object search, enhancing efficiency.
Imen Mahdi, Matteo Cassinelli, Fabien Despinoy et al.
The study reveals massive activations and attention sinks in Transformers as architectural artifacts.
Shangwen Sun, Alfredo Canziani, Yann LeCun et al.
Built on Mip-NeRF with adaptive weighted MSE, this method reconstructs 3D gas plumes from LWIR hyperspectral images, reducing training data by 50%.
Scout Jarman, Zigfried Hampel-Arias, Adra Carr et al.
WebChain is the largest real-world web interaction dataset with multi-modal alignment, enabling state-of-the-art web agent training.
Sicheng Fan, Rui Wan, Yifei Leng et al.
Introduced Multilingual Cloud Corpus, a structured, parallel, multimodal dataset of 42 Bangladeshi minority languages, enabling low-resource NLP applications.
Mohammad Mamun Or Rashid
Proposes VideoHV-Agent, a hypothesis-verification multi-agent framework, achieving state-of-the-art accuracy in long video question answering.
Zheng Wang, Haoran Chen, Haoxuan Qin et al.
Guiding diffusion-based reconstruction with contrastive signals significantly enhances balanced visual representation, improving both D-Ability and P-Ability.
Boyu Han, Qianqian Xu, Shilong Bao et al.
HiMAP-Travel solves long-horizon travel planning with budget and diversity constraints, achieving 52.65% FPR on the test set.
The Viet Bui, Wenjun Li, Yong Liu
Proposed IF-RewardBench, a comprehensive benchmark using preference graphs for instruction-following evaluation, revealing significant deficiencies in current judge models.
Bosi Wen, Yilin Niu, Cunxiang Wang et al.
GOLF leverages group-level natural language feedback to boost RL exploration, achieving 2.2× sample efficiency.
Lei Huang, Xiang Cheng, Chenxiao Zhao et al.
Pointer-CAD addresses B-rep entity selection issues via pointer-based commands, reducing segmentation error significantly.
Dacheng Qi, Chenyu Wang, Jingwei Xu et al.
Using Lw1+∞ Ward identities and Berends-Giele recursion, the paper demonstrates non-zero single-minus graviton amplitudes in specific configurations.
Alfredo Guevara, Alexandru Lupsasca, David Skinner et al.
Proposed perception-aware time-optimal trajectory planning integrating nonlinear dynamics and visual constraints, enabling high-speed quadrotor racing with improved robustness.
Chao Qin, Jiaxu Xing, Rudolf Reiter et al.
CodeTaste benchmark evaluates LLMs' ability to perform human-like code refactorings, combining correctness and preference alignment, revealing current models' gaps.
Alex Thillen, Niels Mündler, Veselin Raychev et al.
SORT integrates local attention, request-centric sampling, and generative pre-training, achieving 6.35% order increase and halving latency in industrial recommenders.
Chunqi Wang, Bingchao Wu, Taotian Pang et al.
Combining evolutionary game theory and large language models, this study models the co-evolution of social behaviors in human-AI hybrid societies.
The Anh Han, Joel Z. Leibo, Tom Lenaerts et al.
Proposes DistriVoting with GMM and SelfStepConf to improve confidence calibration, outperforming SOTA by 3.5%.
Xizhong Yang, Haotian Zhang, Huiming Wang et al.
VAS quantifies visual attention; AVAR improves multimodal reasoning by 7% without retraining.
Ruilin Luo, Chufan Shi, Yizhen Zhang et al.
DSRM-HRL achieves recommendation fairness via state purification and hierarchical decision-making, enhancing utility and exposure equity.
Yun Lu, Xiaoyu Shi, Hong Xie et al.
DEVS-based framework uses natural language to generate and verify long-horizon discrete-event world models, ensuring consistency.
Zheyu Chen, Huiteng Zhuang, Zhuohuan Li et al.