MOSS-Video-Preview: Toward Real-Time Video Understanding via Cross-Attention
MOSS-Video-Preview achieves real-time video understanding via cross-attention, boosting speed by 5x.
Pengyu Wang, Chenkun Tan, Shaojun Zhou et al.
MOSS-Video-Preview achieves real-time video understanding via cross-attention, boosting speed by 5x.
Pengyu Wang, Chenkun Tan, Shaojun Zhou et al.
Proposes a unified, representation- and geometry-guided discrete tokenizer for autonomous driving, improving scene reconstruction and planning performance.
Ziyang Yao, Zeyu Zhu, YunCheng Jiang et al.
TLG improves video QA accuracy by 24.5% to 71.37% using temporal logic grounding.
Ali Alavi
Proposes an interactive video world modeling framework integrating diffusion, VAE, and memory, achieving real-time multi-modal control with 95% scene accuracy.
Jiuming Liu, Chaojun Ni, Mengmeng Liu et al.
Decoupled Residual Denoising Diffusion (DRDD) introduces a two-stage process—stochastic noise diffusion for domain harmonization and deterministic residual diffusion for semantic mapping—achieving unified, data-efficient image-to-image translation with superior performance.
Ziyue Lin, Jiahe Hou, Hongyu Xia et al.
MBench evaluates long-term memory in video world models via entity, environment, and causal dimensions with 12 sub-metrics, using real long videos and VLM.
Shengjun Zhang, Zhang Zhang, Simin Huang et al.
Proposes Wan2.2 video diffusion model compression via few-step distillation and low-bit quantization, achieving superior efficiency.
Jinyang Du, Shenghao Jin, Ziqian Xu et al.
Introduces pause-and-think dataset and a 4B model fine-tuned with structured reasoning, achieving 58% accuracy with 59× fewer parameters than state-of-the-art models.
Shivam Singh, Saptarshi Majumder, Pratik Prabhanjan Brahma et al.
CAFOSat employs deep learning and GradCAM localization to refine annotations, covering 45,000 image patches across 20 US states for infrastructure-aware CAFO mapping.
Oishee Bintey Hoque, Nibir Chandra Mandal, Mandy L Wilson et al.
Zamba2-VL combines SSM and Transformer, achieving 10x faster inference at 1.2B-7B parameters.
Hassan Shapourian, Kasra Hejazi, Olabode M. Sule et al.
StressDream steers video world models by optimizing initial noise to enhance robust policy evaluation.
Junwon Seo, Sushant Veer, Ran Tian et al.
Lumos-Nexus employs a two-stage training and UPFB to bridge frequencies, boosting video fidelity and reasoning-driven generation.
Jiazheng Xing, Hangjie Yuan, Lingling Cai et al.
Introduces SOCO benchmark with over 1 million keypoint pairs across 100 categories, revealing that vision foundation models encode strong semantics but poorly transfer correspondences across categories.
Olaf Dünkel, Basavaraj Sunagad, Haoran Wang et al.
C4G introduces timestamp-conditioned learnable Gaussian tokens with transformer decoding, enabling efficient 4D scene reconstruction from monocular video without scene-specific optimization, reducing Gaussian count by orders of magnitude.
Mungyeom Kim, Minkyeong Jeon, Honggyu An et al.
VolFill employs a latent diffusion model with TUDF grids to reconstruct complete 3D scenes from a single RGB image, achieving state-of-the-art accuracy.
Tuan Duc Ngo, Chuang Gan, Evangelos Kalogerakis
Introduces Neural Token Reconstruction (NTR) with masked latent reconstruction to enhance scene token representations, achieving 8.0461 RFS on Waymo E2E.
Jiahui Li, Jiawei Sun, Zixiang Ren et al.
iVGR internalizes visual localization into textual reasoning using reinforcement learning, significantly enhancing MLLMs' performance.
Chang-Bin Zhang, Yujie Zhong, Qiang Zhang et al.
VideoMLA introduces low-rank latent KV cache, reducing memory by 92.7% for minute-scale video diffusion while maintaining high quality.
Hidir Yesiltepe, Jiazhen Hu, Tuna Han Salih Meral et al.
NeuROK employs a transformer-based encoder-decoder to learn a low-dimensional latent space for 4D object dynamics, trained on large-scale geometric trajectories, bypassing predefined physical models.
Chen Geng, Guangzhao He, Yue Gao et al.
YoCausal employs a two-level causality benchmark using real-world videos and natural reversal, evaluating 13 SOTA video diffusion models' causal understanding via RSI and CCI metrics.
You-Zhe Xie, Yu-Hsuan Li, Jie-Ying Lee et al.