cs.CV 2305.13077

ControlVideo: Training-free Controllable Text-to-Video Generation

ControlVideo employs full cross-frame attention and hierarchical sampling to enable training-free, high-quality controllable text-to-video generation, outperforming state-of-the-art methods.

Yabo Zhang, Yuxiang Wei, Dongsheng Jiang et al.

2023-05-22 379 citations 42
cs.CV 2305.10855

TextDiffuser: Diffusion Models as Text Painters

TextDiffuser combines Transformer layout prediction with latent diffusion models, enabling high-quality, controllable text image synthesis.

Jingye Chen, Yupan Huang, Tengchao Lv et al.

2023-05-18 38
cs.CV 2305.07214

MMG-Ego4D: Multi-Modal Generalization in Egocentric Action Recognition

Proposes MMG-Ego4D, a multimodal egocentric action recognition dataset and framework, with Transformer fusion, contrastive alignment, and prototypical loss, improving generalization in missing and zero-shot scenarios.

Xinyu Gong, Sreyas Mohan, Naina Dhingra et al.

2023-05-12 38 citations 47
cs.CV 2305.06355

VideoChat: Chat-Centric Video Understanding

VideoChat integrates video foundation models with LLMs via learnable interfaces, enabling advanced spatiotemporal reasoning and causal inference.

KunChang Li, Yinan He, Yi Wang et al.

2023-05-11 40
cs.CV 2305.05665

ImageBind: One Embedding Space To Bind Them All

ImageBind learns a unified embedding space for six modalities, enabling zero-shot cross-modal retrieval, composition, detection, and generation, surpassing specialized models.

Rohit Girdhar, Alaaeldin El-Nouby, Zhuang Liu et al.

2023-05-10 1776 citations 30
cs.CV 2305.04789

AvatarReX: Real-time Expressive Full-body Avatars

AvatarReX employs NeRF with structured local implicit fields and geometry-appearance disentanglement for real-time, expressive full-body avatar synthesis, achieving high fidelity and speed.

Zerong Zheng, Xiaochen Zhao, Hongwen Zhang et al.

2023-05-08 36