cs.CV 2502.17863

A Survey: Spatiotemporal Consistency in Video Generation

This survey frames video generation as sequential sampling from a high-dimensional spatiotemporal distribution and compares VAE, AR, diffusion, and flow models.

Zhiyu Yin, Kehai Chen, Xuefeng Bai et al.

2025-02-25 23
cs.CV 2502.04896

Goku: Flow Based Video Generative Foundation Models

Goku employs rectified flow Transformer for joint image-video generation, achieving top-tier performance with 0.76 on GenEval and 84.85 on VBench.

Shoufa Chen, Chongjian Ge, Yuqi Zhang et al.

2025-02-07 43
cs.CV 2502.04507

Fast Video Generation with Sliding Tile Attention

Introduces Sliding Tile Attention (STA), achieving 2.8-17× speedup in video diffusion models with minimal quality loss, based on local 3D attention patterns.

Peiyuan Zhang, Yongqi Chen, Runlong Su et al.

2025-02-07 24