GRACE: Generating Socially Appropriate Robot Actions Leveraging LLMs and Human Explanations
GRACE combines LLMs and human explanations, boosting socially appropriate robot actions.
Fethiye Irmak Dogan, Umut Ozyurt, Gizem Cinar et al.
GRACE combines LLMs and human explanations, boosting socially appropriate robot actions.
Fethiye Irmak Dogan, Umut Ozyurt, Gizem Cinar et al.
Decentralized multi-agent RL with soft collision avoidance achieves near-time-optimal drone flight, max speed 13.65 m/s, collision rate 5.9%.
Xian Wang, Jin Zhou, Yuanli Feng et al.
Proposed a Pose-Guided Motion Model to generate fine-grained, motion-consistent sign language videos, significantly improving detail and temporal consistency.
Tongkai Shi, Lianyu Hu, Fanhua Shang et al.
Introduces human-defined correctness and ICE metrics for data filtering, improving tool-using LLMs performance with high-quality synthetic data.
Shadi Iskander, Nachshon Cohen, Zohar Karnin et al.
TiM4Rec integrates time-awareness with SSD to enhance low-dimensional recommendation performance.
Hao Fan, Mengyi Zhu, Yanrong Hu et al.
Proposes a standardized framework for tool integration in LLMs, combining fine-tuning and in-context learning to enhance complex task performance.
Zhuocheng Shen
Introduces DERS, a robustness benchmark for endoscopic depth estimation under synthetic corruption, evaluated on SCARED dataset.
An Wang, Haochen Yin, Beilei Cui et al.
Divergence-based calibration (DC-PDD) outperforms existing methods in detecting training data, improving AUC by 8.6%.
Weichao Zhang, Ruqing Zhang, Jiafeng Guo et al.
DROP employs online sampling predictive control with vision-based pose estimation to reorient objects, achieving performance comparable to RL methods without extensive training.
Albert H. Li, Preston Culbertson, Vince Kurtz et al.
Proposed an efficient modified MUSIC algorithm leveraging RIS symmetry for near-field multi-user localization, reducing complexity by 184x.
Parisa Ramezani, Alva Kosasih, Emil Björnson
OneBEV uses a single panoramic image to achieve bird's-eye-view semantic mapping, reaching 51.1% mIoU, with a novel Mamba-based transformation.
Jiale Wei, Junwei Zheng, Ruiping Liu et al.
Introduces a bottom-up class-agnostic image segmentation method using mean-shift clustering, achieving 31% improvement on ADE20K.
Sebastian Dille, Ari Blondal, Sylvain Paris et al.
ReMEmbR employs long-term spatio-temporal memory with retrieval mechanisms for robot navigation QA.
Abrar Anwar, John Welsh, Joydeep Biswas et al.
Covariance-based deep learning model for buried object classification, improving robustness with SPDNet.
Douba Jafuno, Ammar Mian, Guillaume Ginolhac et al.
Proposes LSCS architecture combining model parameters, explicit memory, knowledge graphs, and text storage for lifelong experience absorption and accurate recall.
Yu Wang, Chi Han, Tongtong Wu et al.
Oryx supports arbitrary resolution spatial-temporal understanding via OryxViT and dynamic compression, enabling efficient long video and 3D scene processing.
Zuyan Liu, Yuhao Dong, Ziwei Liu et al.
Proposes MMSearch with MMSEARCH-ENGINE pipeline, achieving top performance on 14 subfields, surpassing commercial search engines in multimodal tasks.
Dongzhi Jiang, Renrui Zhang, Ziyu Guo et al.
FRAMES dataset evaluates LLMs' factuality, retrieval, and reasoning via multi-hop questions, showing a 50%+ performance boost with multi-step retrieval.
Satyapriya Krishna, Kalpesh Krishna, Anhad Mohananey et al.
LogicPro synthesizes complex logical reasoning data using program-guided learning, enhancing model performance.
Jin Jiang, Yuchen Yan, Yang Liu et al.
Proposes SciLead, an LLM-based system for automated scientific leaderboard construction, achieving 55% F1 in full scenario, addressing manual errors.
Furkan Şahinuç, Thy Thy Tran, Yulia Grishina et al.