CogNav: Cognitive Process Modeling for Object Goal Navigation with LLMs
CogNav employs LLM-based cognitive modeling with a heterogeneous map, boosting ObjectNav success by at least 14%.
Yihan Cao, Jiazhao Zhang, Zhinan Yu et al.
CogNav employs LLM-based cognitive modeling with a heterogeneous map, boosting ObjectNav success by at least 14%.
Yihan Cao, Jiazhao Zhang, Zhinan Yu et al.
Transforms slow bidirectional video diffusion into fast autoregressive models, achieving 9.4 FPS with high quality via distillation and ODE initialization.
Tianwei Yin, Qiang Zhang, Richard Zhang et al.
Introduces 3DSRBench, a benchmark with 2772 annotated QA pairs, to evaluate large multimodal models' 3D spatial reasoning, revealing current limitations especially in uncommon viewpoints.
Wufei Ma, Haoyu Chen, Guofeng Zhang et al.
MV-DUSt3R+ reconstructs scenes from sparse views in 2 seconds, significantly improving speed and accuracy.
Zhenggang Tang, Yuchen Fan, Dilin Wang et al.
SAME employs a state-adaptive Mixture of Experts for multi-task visual navigation, outperforming task-specific models with dynamic expert routing.
Gengze Zhou, Yicong Hong, Zun Wang et al.
UniScene unifies occupancy, video, and LiDAR generation via an occupancy-centric progressive framework.
Bohan Li, Jiazhe Guo, Hongsi Liu et al.
InternVL 2.5 achieves over 70% on MMMU by model, data, and test-time scaling, surpassing previous open-source limits.
Zhe Chen, Weiyun Wang, Yue Cao et al.
NVILA employs a 'scale-then-compress' approach, boosting high-res image and long video processing efficiency, reducing training costs by 1.9-5.1×.
Zhijian Liu, Ligeng Zhu, Baifeng Shi et al.
VisionZip selects high-information visual tokens, reduces redundancy, boosts inference speed by 8x, with performance drop under 5%.
Senqiao Yang, Yukang Chen, Zhuotao Tian et al.
Proposes NoiseRefine, mapping noise to enable guidance-free high-quality image synthesis, trained with only 50K text-image pairs.
Donghoon Ahn, Jiwon Kang, Sanghyun Lee et al.
MIDI extends pre-trained single-object 3D models into multi-instance diffusion, capturing spatial relations with attention mechanisms.
Zehuan Huang, Yuan-Chen Guo, Xingqiao An et al.
PaliGemma 2 integrates SigLIP-So400m and Gemma 2, employing a three-stage training process across multiple scales, achieving state-of-the-art transfer performance on diverse tasks.
Andreas Steiner, André Susano Pinto, Michael Tschannen et al.
Proposes BTimer, a real-time monocular dynamic scene reconstruction model using bullet-time representation, achieving 150ms high-quality scene synthesis.
Hanxue Liang, Jiawei Ren, Ashkan Mirzaei et al.
HunyuanVideo is an open-source large-scale video generation model with over 1.3 billion parameters, matching or surpassing top industry closed-source models.
Weijie Kong, Qi Tian, Zijian Zhang et al.
SceneFactor employs factored latent diffusion models for controllable large-scale 3D scene generation and editing, enabling intuitive object manipulation with minimal clicks.
Alexey Bokhovkin, Quan Meng, Shubham Tulsiani et al.
LamRA leverages large multimodal models with two-stage training and lightweight LoRA modules to unify multi-task retrieval and reranking, achieving strong zero-shot performance.
Yikun Liu, Pingan Chen, Jiayin Cai et al.
Proposes Structured LATent (SLAT) for scalable 3D generation, integrating sparse grids with multiview features, enabling multi-format outputs with up to 2 billion parameters.
Jianfeng Xiang, Zelong Lv, Sicheng Xu et al.
DynSUP uses two unposed images to achieve dynamic Gaussian splatting, significantly enhancing dynamic scene synthesis.
Weihang Li, Weirong Chen, Shenhan Qian et al.
Proposes Video-3D LLM, integrating 3D position encoding into video representations, achieving SOTA on 5 3D scene benchmarks with 58.1% [email protected] on ScanRefer.
Duo Zheng, Shijia Huang, Liwei Wang
AMO sampler enhances text rendering in diffusion models by adaptive overshooting, improving accuracy by 35.9% without extra training.
Xixi Hu, Keyang Xu, Bo Liu et al.