cs.CV 2512.22096

Yume-1.5: A Text-Controlled Interactive World Generation Model

Yume-1.5 employs joint spatiotemporal-channel modeling (TSCM) to enable real-time long video generation with improved speed and stability, supporting interactive world creation.

Xiaofeng Mao, Zhen Li, Chuanhao Li et al.

2025-12-27 64 citations 36
cs.CV 2512.20619

SemanticGen: Video Generation in Semantic Space

SemanticGen uses a two-stage diffusion framework in semantic space, boosting long video generation speed and quality, with a 30% faster convergence and better long-term consistency.

Jianhong Bai, Xiaoshi Wu, Xintao Wang et al.

2025-12-24 22
cs.CV 2512.19539

StoryMem: Multi-shot Long Video Storytelling with Memory

StoryMem employs Memory-to-Video (M2V) to convert pretrained single-shot diffusion models into multi-shot storytelling tools, significantly improving cross-shot consistency with 28.7% gain on ST-Bench.

Kaiwen Zhang, Liming Jiang, Angtian Wang et al.

2025-12-23 23
cs.CV 2512.16978

A Benchmark for Omni-Modal Reasoning in Long Videos

Introduces LongShOTBench, a long-video multi-modal reasoning benchmark with rubric-based scoring, evaluating 105 models, with top score of 66.64%.

Mohammed Irfan Kurpath, Jaseel Muhammad Kaithakkodan, Jinxing Zhou et al.

2025-12-19 34