cs.CV 2305.01115

In-Context Learning Unlocked for Diffusion Models

Prompt Diffusion enables in-context learning in diffusion models, trained on six vision-language tasks for strong generalization.

Zhendong Wang, Yifan Jiang, Yadong Lu et al.

2023-05-02 126 citations 63
cs.CV 2304.14108

DataComp: In search of the next generation of multimodal datasets

Introduces DataComp benchmark, utilizing 128B image-text pairs with filtering strategies, achieving 79.2% ImageNet zero-shot accuracy with ViT-L/14, outperforming prior datasets.

Samir Yitzhak Gadre, Gabriel Ilharco, Alex Fang et al.

2023-04-27 754 citations 37
cs.CL 2304.10346

Interventional Probing in High Dimensions: An NLI Case Study

Using natural logic intermediate features, the study applies Amnesic and Mnestic interventions to analyze causal roles in high-dimensional model representations, revealing limitations and improvements.

Julia Rozanova, Marco Valentino, Lucas Cordeiro et al.

2023-04-20 53
cs.GR 2304.10320

Neurosymbolic Models for Computer Graphics

Neurosymbolic models combine symbolic programs and neural networks for 2D/3D shape and material generation.

Daniel Ritchie, Paul Guerrero, R. Kenny Jones et al.

2023-04-20 28
cs.CV 2304.08485

Visual Instruction Tuning

LLaVA: leveraging GPT-4 generated multimodal instruction data, achieves 85.1% relative score, with 92.53% accuracy on Science QA, pioneering multimodal instruction tuning.

Haotian Liu, Chunyuan Li, Qingyang Wu et al.

2023-04-18 10948 citations 46
cs.CV 2304.08463

Learning to Render Novel Views from Wide-Baseline Stereo Pairs

Proposes a multi-view transformer-based method with epipolar sampling for high-quality novel view synthesis from a single wide-baseline stereo pair, outperforming prior sparse observation approaches.

Yilun Du, Cameron Smith, Ayush Tewari et al.

2023-04-18 50