Content-Based Search for Deep Generative Models
Proposes content-based deep generative model retrieval using multi-modal feature contrastive learning and probability optimization.
Daohan Lu, Sheng-Yu Wang, Nupur Kumari et al.
Proposes content-based deep generative model retrieval using multi-modal feature contrastive learning and probability optimization.
Daohan Lu, Sheng-Yu Wang, Nupur Kumari et al.
Phenaki generates variable-length videos from text using causal attention and a bidirectional masked transformer.
Ruben Villegas, Mohammad Babaeizadeh, Pieter-Jan Kindermans et al.
Bayesian Prompt Learning regularizes prompt distributions through variational inference and improves unseen-prompt generalization across 15 benchmarks.
Mohammad Mahdi Derakhshani, Enrique Sanchez, Adrian Bulat et al.
VICRegL combines global and local feature learning, boosting detection and segmentation while maintaining classification accuracy.
Adrien Bardes, Jean Ponce, Yann LeCun
Spotlight model uses vision-language integration for mobile UI understanding, surpassing existing methods.
Gang Li, Yang Li
DINOSAUR leverages self-supervised feature reconstruction with Slot Attention, outperforming existing models, scalable to COCO and PASCAL VOC datasets.
Maximilian Seitzer, Max Horn, Andrii Zadaianchuk et al.
Make-A-Video generates videos from text using text-image data and unsupervised video learning.
Uriel Singer, Adam Polyak, Thomas Hayes et al.
GET3D combines differentiable surface modeling with 2D GANs to generate complex, high-detail textured 3D meshes directly from images.
Jun Gao, Tianchang Shen, Zian Wang et al.
Proposes U3HS for open-set panoptic segmentation, detecting unknown objects via uncertainty and embedding clustering, without prior knowledge.
Stefano Gasperini, Alvaro Marcos-Ramiro, Michael Schmidt et al.
Proposes AiRLoc, a reinforcement learning-based aerial goal localization model, outperforming heuristics with high generalization in disaster scenarios.
Aleksis Pirinen, Anton Samuelsson, John Backsund et al.
MapTR employs structured Transformer with permutation-equivalent modeling, achieving real-time high-precision HD map construction, 8× faster, 5.0 mAP higher.
Bencheng Liao, Shaoyu Chen, Xinggang Wang et al.
DreamBooth fine-tunes diffusion models with few images, enabling personalized subject generation with high fidelity.
Nataniel Ruiz, Yuanzhen Li, Varun Jampani et al.
Cold diffusion: Inverts arbitrary image transforms without noise, enhancing generative model diversity.
Arpit Bansal, Eitan Borgnia, Hong-Min Chu et al.
Using ViT with the 8-Point Algorithm for relative pose prediction, improving performance in limited data scenarios.
Chris Rockwell, Justin Johnson, David F. Fouhey
Proposes RAN models with adjustable non-linearity, achieving hardware-efficient deep networks; RAN-e and RAN-i outperform baselines on ImageNet and hardware benchmarks.
Kartikeya Bhardwaj, James Ward, Caleb Tung et al.
RelPose predicts probabilistic relative rotations using an energy model, improving 3D reconstruction from sparse images.
Jason Y. Zhang, Deva Ramanan, Shubham Tulsiani
MonoViT combines Transformer and CNN for self-supervised monocular depth estimation, achieving state-of-the-art results with Abs Rel 0.099 on KITTI.
Chaoqiang Zhao, Youmin Zhang, Matteo Poggi et al.
Proposes Cross Attention Control for Prompt-to-Prompt image editing, enabling text-only localized and global modifications without masks, with high fidelity.
Amir Hertz, Ron Mokady, Jay Tenenbaum et al.
Proposes Textual Inversion, optimizing 3-5 images to learn new pseudo-words in frozen models, enabling personalized generation.
Rinon Gal, Yuval Alaluf, Yuval Atzmon et al.
ViP3D employs 3D agent queries for end-to-end visual trajectory prediction, outperforming traditional pipelines with significant improvements.
Junru Gu, Chenxu Hu, Tianyuan Zhang et al.