OpenMask3D: Open-Vocabulary 3D Instance Segmentation
OpenMask3D achieves zero-shot open-vocabulary 3D instance segmentation, excelling on ScanNet200's long-tail classes.
Ayça Takmaz, Elisabetta Fedele, Robert W. Sumner et al.
OpenMask3D achieves zero-shot open-vocabulary 3D instance segmentation, excelling on ScanNet200's long-tail classes.
Ayça Takmaz, Elisabetta Fedele, Robert W. Sumner et al.
Proposed MME benchmark evaluates 30 multimodal large models across 14 subtasks, measuring perception and cognition with manual instruction design.
Chaoyou Fu, Peixian Chen, Yunhang Shen et al.
Proposed PSF-aware Transformer (PART) for minimalist high-quality panoramic imaging, significantly improving image restoration.
Qi Jiang, Shaohua Gao, Yao Gao et al.
MotionGPT fine-tunes large language models with 0.4% parameters to generate human motion from multimodal signals like text and poses.
Yaqi Zhang, Di Huang, Bin Liu et al.
R2A framework surpasses Flamingo-80B in zero-shot video QA with only 1.3B parameters.
Junting Pan, Ziyi Lin, Yuying Ge et al.
DreamSim leverages synthetic data to develop a holistic perceptual similarity metric, outperforming existing pixel-based metrics and generalizing well to real images.
Stephanie Fu, Netanel Tamir, Shobhita Sundaram et al.
LVLM-eHub evaluates 8 large vision-language models, revealing overfitting and object hallucination issues, proposing a multi-turn reasoning framework.
Peng Xu, Wenqi Shao, Kaipeng Zhang et al.
NAVI is a category-agnostic dataset with high-quality 3D models and near-perfect camera parameters, enabling precise multi-scene 3D reconstruction and correspondence tasks.
Varun Jampani, Kevis-Kokitsi Maninis, Andreas Engelhardt et al.
Introduces ARGO1M dataset and CIR method for cross-scenario, cross-location action recognition, outperforming prior approaches.
Chiara Plizzari, Toby Perrett, Barbara Caputo et al.
AssistGPT employs PEIL framework with structured code and tools for autonomous multi-modal reasoning.
Difei Gao, Lei Ji, Luowei Zhou et al.
TF++ leverages attention-based feature pooling and uncertainty-aware speed classification to address biases in end-to-end driving models, boosting Longest6 score by 11 points.
Bernhard Jaeger, Kashyap Chitta, Andreas Geiger
Valley integrates ViT-L/14 with three temporal modules, constructs 702k video-text and 73k instruction datasets, significantly enhancing video understanding.
Ruipu Luo, Ziwang Zhao, Min Yang et al.
Introduces Aria Digital Twin (ADT), a dataset with 200 egocentric sequences, supporting 3D detection, tracking, and scene understanding tasks.
Xiaqing Pan, Nicholas Charron, Yongqian Yang et al.
A multi-modal foundation model system enables zero-shot stylized 3D asset generation from abstract scene descriptions, outperforming baselines in semantic fidelity.
Ian Huang, Vrishab Krishna, Omoruyi Atekha et al.
M$^3$IT dataset optimizes vision-language models with 40 datasets and 80 languages.
Lei Li, Yuwei Yin, Shicheng Li et al.
DIFT leverages pre-trained diffusion models' implicit features for unsupervised semantic, geometric, and temporal correspondence, outperforming weakly-supervised methods.
Luming Tang, Menglin Jia, Qianqian Wang et al.
RAM model achieves high-accuracy zero-shot image tagging using large-scale image-text pair training.
Youcai Zhang, Xinyu Huang, Jinyu Ma et al.
VideoComposer employs motion vectors and spatio-temporal encoding to enable multi-modal controllable video synthesis with high temporal consistency.
Xiang Wang, Hangjie Yuan, Shiwei Zhang et al.
DiffLL enhances low-light images using Wavelet-based Conditional Diffusion Model, improving efficiency by 70x.
Hai Jiang, Ao Luo, Songchen Han et al.
FlowCam trains generalizable 3D radiance fields without camera poses using pixel-aligned scene flow, enhancing 3D reconstruction from video data.
Cameron Smith, Yilun Du, Ayush Tewari et al.