Out of Sight, Out of Mind? Evaluating State Evolution in Video World Models
STEVO-Bench evaluates video world models' ability to evolve state during observation interruptions, revealing limitations.
Ziqi Ma, Mengzhan Liufu, Georgia Gkioxari
STEVO-Bench evaluates video world models' ability to evolve state during observation interruptions, revealing limitations.
Ziqi Ma, Mengzhan Liufu, Georgia Gkioxari
InterEdit uses Semantic-Aware Plan Token Alignment and Interaction-Aware Frequency Token Alignment for multi-human 3D motion editing.
Yebin Yang, Di Wen, Lei Qi et al.
coDrawAgents framework improves compositional text-to-image generation with 94% overall accuracy on GenEval via multi-agent collaboration.
Chunhan Li, Qifeng Wu, Jia-Hui Pan et al.
OARS framework uses COMPASS reward for real-time image super-resolution, enhancing perceptual quality and fidelity.
Shijie Zhao, Xuanyu Zhang, Bin Chen et al.
VGGT-World predicts scene geometry evolution via autoregressive modeling of frozen GFM features, achieving 21% better depth accuracy with 0.43B params, 3.6-5× faster.
Xiangyu Sun, Shijie Wang, Fengyi Zhang et al.
Introduces Alternating Gradient Flow (AGF) to prevent structural collapse under 75% compression on ImageNet-1K.
Tianhao Qian, Zhuoxuan Li, Jinde Cao et al.
VQQA employs multi-agent QA with semantic gradients to improve video quality by +11.57%.
Yiwen Song, Tomas Pfister, Yale Song
EVATok achieves efficient visual autoregressive generation with adaptive video tokenization, saving 24.4% tokens on average.
Tianwei Xiong, Jun Hao Liew, Zilong Huang et al.
MM-CondChain uses VPIR for visually grounded deep compositional reasoning, with top model achieving only 53.33 Path F1.
Haozhan Shen, Shilin Yan, Hongwei Xue et al.
OmniStream achieves perception, reconstruction, and action in visual streams using causal spatiotemporal attention and 3D-RoPE, excelling across 29 datasets.
Yibin Yan, Jilan Xu, Shangzhe Di et al.
DreamVideo-Omni achieves multi-subject video customization with latent identity reinforcement learning, enhancing identity fidelity and motion control precision.
Yujie Wei, Xinyu Liu, Shiwei Zhang et al.
AutoGaze autoregressively selects multi-scale video patches, reducing redundancy and enhancing efficiency, enabling 1K-frame 4K video processing.
Baifeng Shi, Stephanie Fu, Long Lian et al.
EndoCoT activates MLLMs' reasoning potential, achieving 92.1% accuracy, 8.3% higher than the baseline.
Xuanlang Dai, Yujie Zhou, Long Xing et al.
BiGain enhances diffusion models by frequency separation, improving classification accuracy by 7.15% and FID by 0.34.
Jiacheng Liu, Shengkun Tang, Jiacheng Cui et al.
RDNet enhances salient object detection in optical remote sensing images using dynamic adaptive modules.
Bin Wan, Runmin Cong, Xiaofei Zhou et al.
FlashMotion introduces a three-stage training framework combining diffusion and adversarial objectives, achieving 47× faster controllable video generation with high quality.
Quanhao Li, Zhen Xing, Rui Wang et al.
O3N framework achieves state-of-the-art performance on QuadOcc and Human360Occ benchmarks using polar-spiral topology for 360° spatial representation.
Mengfei Duan, Hao Shi, Fei Teng et al.
HATS employs hardness-aware trajectory synthesis, enhancing GUI agents' generalization in ambiguous interactions.
Rui Shao, Ruize Gao, Bin Xie et al.
Introduced UniCAC benchmark to evaluate 24 algorithms under various optical aberrations.
Xiaolong Qian, Qi Jiang, Yao Gao et al.
Node-RF integrates Neural ODE with NeRF for continuous-time scene dynamics, achieving superior long-range extrapolation and generalization.
Hiran Sarkar, Liming Kuang, Yordanka Velikova et al.