cs.CV 2409.02389

Multi-modal Situated Reasoning in 3D Scenes

Introduced MSQA, a large-scale multi-modal dataset with 251K QA pairs, leveraging 3D scene graphs and interleaved inputs for improved scene reasoning.

Xiongkun Linghu, Jiangyong Huang, Xuesong Niu et al.

2024-09-04 38
cs.CV 2408.10539

Training Matting Models without Alpha Labels

Train matting models without alpha labels using DDC loss, achieving excellent results on AM-2K and P3M-10K datasets.

Wenze Liu, Zixuan Ye, Hao Lu et al.

2024-08-20 32
cs.CV 2408.06693

DC3DO: Diffusion Classifier for 3D Objects

DC3DO combines LION and diffusion models for zero-shot 3D shape classification, achieving 12.5% improvement over multi-view methods.

Nursena Koprucu, Meher Shashwat Nigam, Shicheng Xu et al.

2024-08-13 34
cs.CV 2408.01800

MiniCPM-V: A GPT-4V Level MLLM on Your Phone

MiniCPM-V employs architecture optimization to enable GPT-4V-level performance on mobile devices, supports 30+ languages, and surpasses prior models in accuracy.

Yuan Yao, Tianyu Yu, Ao Zhang et al.

2024-08-03 37
cs.CV 2408.00714

SAM 2: Segment Anything in Images and Videos

SAM 2 employs streaming memory within a transformer framework for promptable image and video segmentation, achieving 3x fewer interactions and 6x faster speeds than prior models.

Nikhila Ravi, Valentin Gabeur, Yuan-Ting Hu et al.

2024-08-02 40
cs.CV 2408.00712

MotionFix: Text-Driven 3D Human Motion Editing

Proposes MotionFix dataset and TMED diffusion model for text-driven 3D human motion editing, achieving significant improvements in editing accuracy.

Nikos Athanasiou, Alpár Cseke, Markos Diomataris et al.

2024-08-02 44