cs.CV 2303.15389

EVA-CLIP: Improved Training Techniques for CLIP at Scale

EVA-CLIP combines pre-trained EVA, LAMB optimizer, and masking techniques to achieve 82.0% zero-shot ImageNet accuracy with 5.0B parameters, trained on 9B samples.

Quan Sun, Yuxin Fang, Ledell Wu et al.

2023-03-28 32
cs.CV 2303.13508

DreamBooth3D: Subject-Driven Text-to-3D Generation

DreamBooth3D combines DreamBooth and DreamFusion to generate personalized 3D models from 3-6 images, achieving high-quality, editable assets.

Amit Raj, Srinivas Kaza, Ben Poole et al.

2023-03-24 292 citations 36
cs.CV 2303.11797

CAT-Seg: Cost Aggregation for Open-Vocabulary Semantic Segmentation

Proposes CAT-Seg, a cost aggregation-based open-vocabulary semantic segmentation framework, fine-tuning CLIP encoders to improve pixel-level recognition of seen and unseen classes, outperforming state-of-the-art.

Seokju Cho, Heeseong Shin, Sunghwan Hong et al.

2023-03-21 297 citations 48
cs.CV 2303.11328

Zero-1-to-3: Zero-shot One Image to 3D Object

Zero-1-to-3 leverages pre-trained diffusion models for zero-shot single-image 3D view synthesis and reconstruction, with strong generalization.

Ruoshi Liu, Rundi Wu, Basile Van Hoorick et al.

2023-03-21 43
cs.CV 2303.09295

DIRE for Diffusion-Generated Image Detection

Proposes DIRE, leveraging diffusion model reconstruction error for high-accuracy detection of diffusion-generated images, outperforming existing methods.

Zhendong Wang, Jianmin Bao, Wengang Zhou et al.

2023-03-16 37