cs.CV 2512.19539

StoryMem: Multi-shot Long Video Storytelling with Memory

StoryMem employs Memory-to-Video (M2V) to convert pretrained single-shot diffusion models into multi-shot storytelling tools, significantly improving cross-shot consistency with 28.7% gain on ST-Bench.

Kaiwen Zhang, Liming Jiang, Angtian Wang et al.

2025-12-23 31
stat.ML 2512.20685

Diffusion Models in Simulation-Based Inference: A Tutorial Review

This review discusses diffusion models in simulation-based inference, focusing on training, inference, and evaluation, highlighting guidance, score composition, and flow matching techniques.

Jonas Arruda, Niels Bracher, Ullrich Köthe et al.

2025-12-22 45
cs.CV 2512.16978

A Benchmark for Omni-Modal Reasoning in Long Videos

Introduces LongShOTBench, a long-video multi-modal reasoning benchmark with rubric-based scoring, evaluating 105 models, with top score of 66.64%.

Mohammed Irfan Kurpath, Jaseel Muhammad Kaithakkodan, Jinxing Zhou et al.

2025-12-19 36