cs.CV 2411.17470

Towards Precise Scaling Laws for Video Diffusion Transformers

This paper proposes a precise scaling law for video diffusion Transformers, integrating optimal hyperparameter prediction to enhance performance and reduce inference costs by 40.1%.

Yuanyang Yin, Yaqi Zhao, Mingwu Zheng et al.

2024-11-26 21 citations 48
cs.CV 2411.00639

Event-guided Low-light Video Semantic Segmentation

EVSNet leverages event-based motion features with a lightweight fusion framework, achieving 34.1% mIoU on low-light VSPW, 11× parameter efficiency over SOTA.

Zhen Yao, Mooi Choo Chuah

2024-11-01 23 citations 45