cs.CV 2305.13077

ControlVideo: Training-free Controllable Text-to-Video Generation

ControlVideo employs full cross-frame attention and hierarchical sampling to enable training-free, high-quality controllable text-to-video generation, outperforming state-of-the-art methods.

Yabo Zhang, Yuxiang Wei, Dongsheng Jiang et al.

2023-05-22 379 citations 42
cs.CV 2305.10855

TextDiffuser: Diffusion Models as Text Painters

TextDiffuser combines Transformer layout prediction with latent diffusion models, enabling high-quality, controllable text image synthesis.

Jingye Chen, Yupan Huang, Tengchao Lv et al.

2023-05-18 39
cs.CV 2305.07214

MMG-Ego4D: Multi-Modal Generalization in Egocentric Action Recognition

Proposes MMG-Ego4D, a multimodal egocentric action recognition dataset and framework, with Transformer fusion, contrastive alignment, and prototypical loss, improving generalization in missing and zero-shot scenarios.

Xinyu Gong, Sreyas Mohan, Naina Dhingra et al.

2023-05-12 38 citations 50
cs.CV 2305.06355

VideoChat: Chat-Centric Video Understanding

VideoChat integrates video foundation models with LLMs via learnable interfaces, enabling advanced spatiotemporal reasoning and causal inference.

KunChang Li, Yinan He, Yi Wang et al.

2023-05-11 43