Shap-E: Generating Conditional 3D Implicit Functions
Shap·E employs conditional diffusion to directly generate 3D implicit function parameters, enabling rapid, diverse 3D asset creation.
Heewoo Jun, Alex Nichol
Shap·E employs conditional diffusion to directly generate 3D implicit function parameters, enabling rapid, diverse 3D asset creation.
Heewoo Jun, Alex Nichol
Prompt Diffusion enables in-context learning in diffusion models, trained on six vision-language tasks for strong generalization.
Zhendong Wang, Yifan Jiang, Yadong Lu et al.
Introduces DataComp benchmark, utilizing 128B image-text pairs with filtering strategies, achieving 79.2% ImageNet zero-shot accuracy with ViT-L/14, outperforming prior datasets.
Samir Yitzhak Gadre, Gabriel Ilharco, Alex Fang et al.
AMT achieves efficient video frame interpolation using bidirectional pixel correlations and multi-field transforms, enhancing PSNR.
Zhen Li, Zuo-Liang Zhu, Ling-Hao Han et al.
NeuralField-LDM employs hierarchical latent diffusion models to generate complex, realistic 3D scenes, outperforming state-of-the-art methods with significant improvements in FID and FVD scores.
Seung Wook Kim, Bradley Brown, Kangxue Yin et al.
LLaVA: leveraging GPT-4 generated multimodal instruction data, achieves 85.1% relative score, with 92.53% accuracy on Science QA, pioneering multimodal instruction tuning.
Haotian Liu, Chunyuan Li, Qingyang Wu et al.
Proposes a multi-view transformer-based method with epipolar sampling for high-quality novel view synthesis from a single wide-baseline stereo pair, outperforming prior sparse observation approaches.
Yilun Du, Cameron Smith, Ayush Tewari et al.
Proposed certified black-box defense using Robust UNet denoiser with 35% improvement on CIFAR-10.
Astha Verma, A V Subramanyam, Siddhesh Bangar et al.
InterGen employs diffusion models with cooperative transformers and a large-scale multimodal dataset to generate diverse two-person interaction motions guided by text.
Han Liang, Wenqian Zhang, Wenxuan Li et al.
Single-gate MoE achieves comparable efficiency and accuracy to complex models, outperforming non-mixture baselines.
Amelie Royer, Ilia Karmanov, Andrii Skliar et al.
StillFast introduces an end-to-end model for short-term object interaction anticipation, achieving 13.29% Top-5 mAP on EGO4D v2, surpassing SOTA.
Francesco Ragusa, Giovanni Maria Farinella, Antonino Furnari
SAM model supports prompt-based image segmentation with over 1 billion masks, achieving state-of-the-art zero-shot performance.
Alexander Kirillov, Eric Mintun, Nikhila Ravi et al.
GINA-3D generates tri-plane implicit assets from Waymo sensor data, achieving FID 59.5 on WOD-Vehicle.
Bokui Shen, Xinchen Yan, Charles R. Qi et al.
EGC integrates energy-based and diffusion models, achieving state-of-the-art performance in both image classification (78.9% accuracy on CIFAR-10) and high-fidelity image generation (FID 6.05 on ImageNet-1k).
Qiushan Guo, Chuofan Ma, Yi Jiang et al.
IPMAN integrates biomechanical stability via pressure heatmaps, CoP, and CoM, improving 3D human pose accuracy and physical plausibility by 15% on standard datasets.
Shashank Tripathi, Lea Müller, Chun-Hao P. Huang et al.
This paper introduces ToMe for Stable Diffusion, reducing tokens by 60%, doubling generation speed, and saving 5.6× memory without retraining.
Daniel Bolya, Judy Hoffman
Proposes NeRF-supervised deep stereo training, generating synthetic data from single-camera images, achieving 30-40% performance gains without ground-truth.
Fabio Tosi, Alessio Tonioni, Daniele De Gregorio et al.
LLaMA-Adapter uses zero-initialized attention with only 1.2M parameters to achieve instruction tuning comparable to full fine-tuning.
Renrui Zhang, Jiaming Han, Chris Liu et al.
CARTO achieves category- and joint-agnostic 3D reconstruction from a single stereo image, improving mAP 3D IOU50 by 20.4%.
Nick Heppert, Muhammad Zubair Irshad, Sergey Zakharov et al.
Proposes Anti-DreamBooth using adversarial noise to disrupt personalized diffusion model training, reducing fake image generation by over 85%.
Thanh Van Le, Hao Phung, Thuan Hoang Nguyen et al.