cs.CV 2305.05665

ImageBind: One Embedding Space To Bind Them All

ImageBind learns a unified embedding space for six modalities, enabling zero-shot cross-modal retrieval, composition, detection, and generation, surpassing specialized models.

Rohit Girdhar, Alaaeldin El-Nouby, Zhuang Liu et al.

2023-05-10 1776 citations 30
cs.CV 2305.04789

AvatarReX: Real-time Expressive Full-body Avatars

AvatarReX employs NeRF with structured local implicit fields and geometry-appearance disentanglement for real-time, expressive full-body avatar synthesis, achieving high fidelity and speed.

Zerong Zheng, Xiaochen Zhao, Hongwen Zhang et al.

2023-05-08 36
cs.CV 2305.01115

In-Context Learning Unlocked for Diffusion Models

Prompt Diffusion enables in-context learning in diffusion models, trained on six vision-language tasks for strong generalization.

Zhendong Wang, Yifan Jiang, Yadong Lu et al.

2023-05-02 126 citations 62
cs.CV 2304.14108

DataComp: In search of the next generation of multimodal datasets

Introduces DataComp benchmark, utilizing 128B image-text pairs with filtering strategies, achieving 79.2% ImageNet zero-shot accuracy with ViT-L/14, outperforming prior datasets.

Samir Yitzhak Gadre, Gabriel Ilharco, Alex Fang et al.

2023-04-27 754 citations 36
cs.CV 2304.08485

Visual Instruction Tuning

LLaVA: leveraging GPT-4 generated multimodal instruction data, achieves 85.1% relative score, with 92.53% accuracy on Science QA, pioneering multimodal instruction tuning.

Haotian Liu, Chunyuan Li, Qingyang Wu et al.

2023-04-18 10948 citations 40
cs.CV 2304.08463

Learning to Render Novel Views from Wide-Baseline Stereo Pairs

Proposes a multi-view transformer-based method with epipolar sampling for high-quality novel view synthesis from a single wide-baseline stereo pair, outperforming prior sparse observation approaches.

Yilun Du, Cameron Smith, Ayush Tewari et al.

2023-04-18 45
cs.CV 2304.05497

Revisiting Single-gated Mixtures of Experts

Single-gate MoE achieves comparable efficiency and accuracy to complex models, outperforming non-mixture baselines.

Amelie Royer, Ilia Karmanov, Andrii Skliar et al.

2023-04-12 36
cs.CV 2304.02643

Segment Anything

SAM model supports prompt-based image segmentation with over 1 billion masks, achieving state-of-the-art zero-shot performance.

Alexander Kirillov, Eric Mintun, Nikhila Ravi et al.

2023-04-06 34
cs.CV 2303.18246

3D Human Pose Estimation via Intuitive Physics

IPMAN integrates biomechanical stability via pressure heatmaps, CoP, and CoM, improving 3D human pose accuracy and physical plausibility by 15% on standard datasets.

Shashank Tripathi, Lea Müller, Chun-Hao P. Huang et al.

2023-04-01 43
cs.CV 2303.17604

Token Merging for Fast Stable Diffusion

This paper introduces ToMe for Stable Diffusion, reducing tokens by 60%, doubling generation speed, and saving 5.6× memory without retraining.

Daniel Bolya, Judy Hoffman

2023-03-31 281 citations 32
cs.CV 2303.17603

NeRF-Supervised Deep Stereo

Proposes NeRF-supervised deep stereo training, generating synthetic data from single-camera images, achieving 30-40% performance gains without ground-truth.

Fabio Tosi, Alessio Tonioni, Daniele De Gregorio et al.

2023-03-31 26