cs.CV 2503.10589

Long Context Tuning for Video Generation

Long Context Tuning (LCT) extends pre-trained video diffusion models' context window, enabling scene-level multi-shot generation with high consistency.

Yuwei Guo, Ceyuan Yang, Ziyan Yang et al.

2025-03-14 21
cs.CV 2503.08507

Referring to Any Person

This paper introduces RexSeek, a model combining multimodal large language models and object detection, achieving superior multi-person referring understanding on the HumanRef dataset.

Qing Jiang, Lin Wu, Zhaoyang Zeng et al.

2025-03-11 24 citations 26
cs.CV 2503.07314

Automated Movie Generation via Multi-Agent CoT Planning

Proposes MovieAgent, a multi-agent hierarchical CoT framework for automated movie generation, achieving superior narrative coherence and character consistency.

Weijia Wu, Zeyu Zhu, Mike Zheng Shou

2025-03-10 33
cs.CV 2503.01262

Object-Aware Video Matting with Cross-Frame Guidance

Object-aware video matting with cross-frame guidance achieves SOTA, reducing reliance on manual trimaps, with MAD 4.23 and MSE 0.31 on RVM.

Huayu Zhang, Dongyue Wu, Yuanjie Shao et al.

2025-03-03 31