Advances in 4D Representation: Geometry, Motion, and Interaction
Using NeRFs and 3DGS for 4D representation generation and reconstruction, enhancing dynamic scene understanding.
Mingrui Zhao, Sauradip Nag, Kai Wang et al.
Using NeRFs and 3DGS for 4D representation generation and reconstruction, enhancing dynamic scene understanding.
Mingrui Zhao, Sauradip Nag, Kai Wang et al.
VFM-VAE leverages frozen VFM as encoder with novel decoder, achieving gFID 1.62 in 640 epochs, 10× faster than prior methods.
Tianci Bi, Xiaoyi Zhang, Yan Lu et al.
Glyph renders long texts into images processed by VLMs, achieving 3-4× token compression with performance comparable to Qwen3-8B.
Jiale Cheng, Yusen Liu, Xinyu Zhang et al.
Proposes a context-aware pseudo-label scoring framework for zero-shot video summarization, achieving 57.58 F1 on SumMe.
Yuanli Wu, Long Zhang, Yue Du et al.
VISTA employs multi-agent self-improvement, achieving up to 60% win rate in video quality enhancement.
Do Xuan Long, Xingchen Wan, Hootan Nakhost et al.
XModBench evaluates Gemini 2.5 Pro's cross-modal consistency, revealing less than 60% accuracy in spatial and temporal reasoning.
Xingrui Wang, Jiang Liu, Chao Huang et al.
Proposes NEO, a native vision-language model trained on 390M image-text pairs, achieving pixel-word alignment and surpassing modular models in various benchmarks.
Haiwen Diao, Mingxuan Li, Silei Wu et al.
WithAnyone model achieves controllable and ID-consistent image generation using contrastive loss and the MultiID-2M dataset.
Hengyuan Xu, Wei Cheng, Peng Xing et al.
Proposes Generative Universal Verifier, trained on ViVerBench, improving visual verification by 8.3 points with OmniVerifier-7B and TTS strategies.
Xinchen Zhang, Xiaoying Zhang, Youbin Wu et al.
FusionNet and Transformer-based models achieved top PSNR of 26.35 in NTIRE 2025 low-light enhancement, demonstrating multi-model fusion effectiveness.
Xiaoning Liu, Zongwei Wu, Florin-Alexandru Vasluianu et al.
Mask-GRPO introduces reinforcement learning into masked generative models, significantly improving text-to-image generation with a 0.73 score on GenEval and FID of 8.32, surpassing SOTA.
Yifu Luo, Xinhao Hu, Keyu Fan et al.
DepthVLA integrates a pretrained depth module into a mixture-of-transformers framework, significantly improving spatial reasoning and manipulation success rates (e.g., 78.5% in real-world tasks).
Tianyuan Yuan, Yicheng Liu, Chenhao Lu et al.
EgoSocial leverages multimodal cues to improve proactive intervention detection in OLMMs, boosting timing accuracy by 45.6% on Phi-4.
Xijun Wang, Tanay Sharma, Achin Kulshrestha et al.
DriveVLA-W0 employs world modeling to predict future images, providing dense self-supervision that significantly enhances autonomous driving models' scalability, outperforming baselines on large datasets.
Yingyan Li, Shuyao Shang, Weisong Liu et al.
AnyUp is a universal feature upsampling method that generalizes to any feature type at inference, outperforming state-of-the-art.
Thomas Wimmer, Prune Truong, Marie-Julie Rakotosaona et al.
DrivingScene employs a static-to-dynamic two-stage training with residual scene flow, achieving real-time high-fidelity 4D scene reconstruction from two images.
Qirui Hou, Wenzhang Sun, Chang Zeng et al.
Proposes a novel ML framework for detecting corner cases in autonomous driving, enhancing detection accuracy.
Sebastian Schmidt, Julius Körner, Stephan Günnemann
Proposes a domain-conditioned, image-free instance generation method that boosts ILR performance by 15% on average across seven benchmarks.
Yankun Wu, Zakaria Laskar, Giorgos Kordopatis-Zilos et al.
FlexTraj introduces a point-based trajectory control framework enabling multi-granularity, alignment-agnostic image-to-video synthesis with improved efficiency.
Zhiyuan Zhang, Can Wang, Dongdong Chen et al.
Proposed a guided diffusion model with Transformer for hyperspectral data augmentation, boosting forest classification accuracy by 4.8%.
Mattia Ferrari, Lorenzo Bruzzone