VAST: A Vision-Audio-Subtitle-Text Omni-Modality Foundation Model and Dataset
VAST model uses the VAST-27M dataset to achieve omni-modality video understanding, setting 22 new SOTA results.
Sihan Chen, Handong Li, Qunbo Wang et al.
VAST model uses the VAST-27M dataset to achieve omni-modality video understanding, setting 22 new SOTA results.
Sihan Chen, Handong Li, Qunbo Wang et al.
NavGPT leverages large language models for explicit reasoning in vision-and-language navigation, demonstrating zero-shot planning with path decomposition and landmark recognition.
Gengze Zhou, Yicong Hong, Qi Wu
Banana network uses Banach fixed-point for pointcloud segmentation with inter-part equivariance, enhancing segmentation accuracy.
Congyue Deng, Jiahui Lei, Bokui Shen et al.
ViTMatte combines hybrid attention and lightweight convolutions, achieving state-of-the-art image matting performance surpassing prior methods.
Jingfeng Yao, Xinggang Wang, Shusheng Yang et al.
BLIP-Diffusion combines pre-trained subject representations with diffusion models for zero-shot and fast fine-tuning controllable text-to-image generation.
Dongxu Li, Junnan Li, Steven C. H. Hoi
Proposes Diffusion Hyperfeatures to fuse multi-scale, multi-timestep features from diffusion models, improving semantic keypoint correspondence with 85.3% accuracy on SPair-71k.
Grace Luo, Lisa Dunlap, Dong Huk Park et al.
Proposes Temporal Contrastive Learning (TCL) and Siamese STCL to boost low-latency SNNs, achieving SOTA on CIFAR-10/100.
Haonan Qiu, Zeyin Song, Yanqi Chen et al.
ControlVideo employs full cross-frame attention and hierarchical sampling to enable training-free, high-quality controllable text-to-video generation, outperforming state-of-the-art methods.
Yabo Zhang, Yuxiang Wei, Dongsheng Jiang et al.
Reinforcement learning with latent motion spaces enables diverse, realistic human motion synthesis in complex 3D indoor scenes.
Kaifeng Zhao, Yan Zhang, Shaofei Wang et al.
TextDiffuser combines Transformer layout prediction with latent diffusion models, enabling high-quality, controllable text image synthesis.
Jingye Chen, Yupan Huang, Tengchao Lv et al.
Paxion enhances video-language models' action knowledge understanding from 50% to 80% using the DVDM objective.
Zhenhailong Wang, Ansel Blume, Sha Li et al.
Proposed PYoCo noise prior fine-tunes pretrained image diffusion models for high-quality video synthesis, outperforming SOTA benchmarks.
Songwei Ge, Seungjun Nah, Guilin Liu et al.
PMC-VQA employs visual instruction tuning, constructing a large-scale medical VQA dataset, significantly improving free-form answer accuracy.
Xiaoman Zhang, Chaoyi Wu, Ziheng Zhao et al.
ULIP-2 leverages large multimodal models to automatically generate detailed 3D shape descriptions, enabling scalable, annotation-free pre-training with state-of-the-art zero-shot classification results.
Le Xue, Ning Yu, Shu Zhang et al.
Proposes MMG-Ego4D, a multimodal egocentric action recognition dataset and framework, with Transformer fusion, contrastive alignment, and prototypical loss, improving generalization in missing and zero-shot scenarios.
Xinyu Gong, Sreyas Mohan, Naina Dhingra et al.
InstructBLIP uses instruction tuning on BLIP-2, achieving SOTA zero-shot performance across 13 unseen vision-language datasets.
Wenliang Dai, Junnan Li, Dongxu Li et al.
VideoChat integrates video foundation models with LLMs via learnable interfaces, enabling advanced spatiotemporal reasoning and causal inference.
KunChang Li, Yinan He, Yi Wang et al.
Proposes ThinkTwice scalable decoder for end-to-end autonomous driving, combining coarse prediction, scene imagination, region retrieval, and multi-layer refinement, outperforming SOTA.
Xiaosong Jia, Penghao Wu, Li Chen et al.
ImageBind learns a unified embedding space for six modalities, enabling zero-shot cross-modal retrieval, composition, detection, and generation, surpassing specialized models.
Rohit Girdhar, Alaaeldin El-Nouby, Zhuang Liu et al.
AvatarReX employs NeRF with structured local implicit fields and geometry-appearance disentanglement for real-time, expressive full-body avatar synthesis, achieving high fidelity and speed.
Zerong Zheng, Xiaochen Zhao, Hongwen Zhang et al.