ZeroVO: Visual Odometry with Minimal Assumptions
ZeroVO achieves zero-shot generalization across environments, improving performance by over 30%.
Lei Lai, Zekai Yin, Eshed Ohn-Bar
ZeroVO achieves zero-shot generalization across environments, improving performance by over 30%.
Lei Lai, Zekai Yin, Eshed Ohn-Bar
Fast ECoT accelerates embodied reasoning by 7.5× via caching high-level thoughts and parallel generation, enabling real-time robotic control.
Zhekai Duan, Yuan Zhang, Shikai Geng et al.
SongBloom combines autoregressive sketching and diffusion refinement to generate high-quality songs, outperforming existing methods.
Chenyu Yang, Shuai Wang, Hangting Chen et al.
ATRADA framework enhances aircraft trajectory datasets by generating high-quality synthetic data in latent space.
Seokbin Yoon, Keumjin Lee
LeVo framework significantly enhances song generation quality through multi-preference alignment.
Shun Lei, Yaoxun Xu, Zhiwei Lin et al.
SpatialLM integrates point cloud encoding with large language models, achieving state-of-the-art indoor scene layout estimation and 3D detection.
Yongsen Mao, Junhao Zhong, Chuan Fang et al.
Analyzed 40 LLM uncertainty quantification methods, highlighting evaluation on non-realistic benchmarks and advocating human-centered assessment.
Siddartha Devic, Tejas Srinivasan, Jesse Thomason et al.
Instruction-based text embedding framework using Mistral-7B, combining soft supervision and adaptive hard-negative mining for multi-task performance.
Jooyoung Choi, Hyun Kim, Hansol Jang et al.
ReCogDrive combines VLM and diffusion planning to generate smooth, safe trajectories, achieving SOTA on NAVSIM with PDMS 90.8.
Yongkang Li, Kaixin Xiong, Xiangyu Guo et al.
Introduced Audio-Aware Decoding (AAD), improving LALMs' F1 score by 0.046 to 0.428 on object hallucination datasets.
Tzu-wen Hsu, Ke-Han Lu, Cheng-Han Chiang et al.
Introduces SamS, an adaptive sample scheduling algorithm that improves large language model preference alignment by over 12% using model state feedback.
Zixuan Huang, Yikun Ban, Lean Fu et al.
Lingshu: a unified multimodal medical foundation model with extensive knowledge integration and state-of-the-art performance.
LASA Team, Weiwen Xu, Hou Pong Chan et al.
Proposes GenMinds framework integrating cognitive science for structured belief modeling; introduces RECAP for causal reasoning evaluation, advancing from mimicry to thought simulation.
Chance Jiajie Li, Jiayi Wu, Zhenze Mo et al.
Proposes VLM4TS with ViT4TS for zero-shot time series anomaly detection, achieving 24.6% F1 improvement and 36x token efficiency.
Zelin He, Sarah Alnegheimish, Matthew Reimherr
Proposes GTRS, a unified framework combining diffusion-based trajectory generation and vocabulary regularization, significantly improving end-to-end multimodal planning robustness.
Zhenxin Li, Wenhao Yao, Zi Wang et al.
Proposes terrain-aware outdoor 3D scene graph generation combining LiDAR and semantic mapping, enabling multi-layer environment understanding.
Chad R Samuelson, Timothy W McLain, Joshua G Mangelson
Proposes a five-dimensional framework distinguishing AI agents from LLM chatbots, based on evolutionary analysis of environment and capabilities.
Jiachen Zhu, Menghui Zhu, Renting Rui et al.
Introduces Cartridges, a self-study trained lightweight KV cache that replicates large context models, reducing memory by 38.6× and increasing throughput 26.4×.
Sabri Eyuboglu, Ryan Ehrlich, Simran Arora et al.
Proposes a unified mathematical framework for LLM hallucinations, analyzing their roots, detection, and mitigation, with experimental validation.
Chaozhuo Li, Pengbo Wang, Chenxu Wang et al.
EF-VFM extends variational flow matching with exponential family distributions, achieving state-of-the-art tabular data generation.
Andrés Guzmán-Cordero, Floor Eijkelboom, Jan-Willem van de Meent