cs.CV 2412.07825

3DSRBench: A Comprehensive 3D Spatial Reasoning Benchmark

Introduces 3DSRBench, a benchmark with 2772 annotated QA pairs, to evaluate large multimodal models' 3D spatial reasoning, revealing current limitations especially in uncommon viewpoints.

Wufei Ma, Haoyu Chen, Guofeng Zhang et al.

2024-12-11 30
cs.CV 2412.04468

NVILA: Efficient Frontier Visual Language Models

NVILA employs a 'scale-then-compress' approach, boosting high-res image and long video processing efficiency, reducing training costs by 1.9-5.1×.

Zhijian Liu, Ligeng Zhu, Baifeng Shi et al.

2024-12-06 32
cs.CV 2412.03895

A Noise is Worth Diffusion Guidance

Proposes NoiseRefine, mapping noise to enable guidance-free high-quality image synthesis, trained with only 50K text-image pairs.

Donghoon Ahn, Jiwon Kang, Sanghyun Lee et al.

2024-12-05 33
cs.CV 2412.03555

PaliGemma 2: A Family of Versatile VLMs for Transfer

PaliGemma 2 integrates SigLIP-So400m and Gemma 2, employing a three-stage training process across multiple scales, achieving state-of-the-art transfer performance on diverse tasks.

Andreas Steiner, André Susano Pinto, Michael Tschannen et al.

2024-12-05 37
cs.CV 2412.01506

Structured 3D Latents for Scalable and Versatile 3D Generation

Proposes Structured LATent (SLAT) for scalable 3D generation, integrating sparse grids with multiview features, enabling multi-format outputs with up to 2 billion parameters.

Jianfeng Xiang, Zelong Lv, Sicheng Xu et al.

2024-12-02 892 citations 45