Muskie: Multi-view Masked Image Modeling for 3D Vision Pre-training
Muskie learns multi-view geometry through masked completion, reaching 2.38 cm 3D ATE on NAVI.
Wenyu Li, Sidun Liu, Peng Qiao et al.
Muskie learns multi-view geometry through masked completion, reaching 2.38 cm 3D ATE on NAVI.
Wenyu Li, Sidun Liu, Peng Qiao et al.
Deming-cycle-based multi-agent system SciEducator achieves superior scientific video understanding and multimodal education, outperforming SOTA models with 65.31% relevance and 81.88% accuracy.
Zhiyu Xu, Weilong Yan, Yufei Shi et al.
SketchVerify enhances motion planning quality for physics-aware video generation using sketch-guided verification.
Yidong Huang, Zun Wang, Han Lin et al.
Proposes TwiG framework enabling real-time textual reasoning during visual generation, significantly enhancing semantic richness.
Ziyu Guo, Renrui Zhang, Hongyu Li et al.
VLA-Fool framework combines textual, visual, and cross-modal attacks, revealing vulnerabilities in embodied VLA models under multimodal perturbations.
Yuping Yan, Yuhan Xie, Yixin Zhang et al.
VisPlay uses self-evolving RL to improve vision-language models' reasoning via unlabeled image data.
Yicheng He, Chengsong Huang, Zongxia Li et al.
ARC-Chapter leverages large-scale multimodal data and hierarchical annotations, achieving 14% F1 and 11.3% SODA improvements on long video chaptering.
Junfu Pu, Teng Wang, Yixiao Ge et al.
PhysX-Anything uses VLM and a novel 3D tokenization to generate high-quality, physically grounded assets from a single image, enabling direct simulation deployment.
Ziang Cao, Fangzhou Hong, Zhaoxi Chen et al.
This study introduces SpatialSky-Bench and Sky-VLM, achieving SOTA in 13 UAV spatial reasoning tasks with 53.3 average score, surpassing baselines by over 130%.
Lingfeng Zhang, Yuchen Zhang, Hongsheng Li et al.
GeoX-Bench benchmarks large multimodal models for cross-view geo-localization and pose estimation, showing strong localization but challenges in pose accuracy.
Yushuo Zheng, Jiangyong Ying, Huiyu Duan et al.
DGS-Net uses gradient decomposition and distillation to improve CLIP fine-tuning for AI image detection, achieving a 6.6% accuracy boost.
Jiazhen Yan, Ziqiang Li, Fan Wang et al.
Uni-Inter synthesizes 3D human motion across diverse interaction contexts using a Unified Interactive Volume (UIV).
Sheng Liu, Yuanzhi Liang, Jiepeng Wang et al.
Fourier series-based tampering synthesis (FSTS) models invisible tampering parameters, significantly improving real-world forgery localization generalization.
Zeqin Yu, Haotao Xie, Jian Zhang et al.
ReaSon employs causal information bottleneck and reinforcement learning to select keyframes, significantly improving video reasoning accuracy under limited frame budgets.
Yuan Zhou, Litao Hua, Shilong Jin et al.
GCAgent employs structured schematic and narrative memory, boosting long-video understanding by 23.5% accuracy, using a multi-stage perception-action-reflection cycle.
Jeong Hun Yeo, Sangyun Chung, Sungjune Park et al.
MonkeyOCR v1.5 achieves robust document parsing via visual consistency RL and a two-stage pipeline, improving OmniDocBench performance by 2.34%.
Jiarui Zhang, Yuliang Liu, Zijun Wu et al.
AHA! uses 3D Gaussian Splatting for animating humans in diverse scenes, enhancing geometric consistency and free-viewpoint rendering.
Aymen Mir, Jian Wang, Riza Alp Guler et al.
FIBO leverages structured long captions and DimFusion to significantly improve controllability and expressiveness in text-to-image generation.
Eyal Gutflaish, Eliran Kachlon, Hezi Zisman et al.
DeepEyesV2 employs a two-stage training and dynamic tool invocation to enhance multimodal reasoning, achieving 63.7% accuracy on RealX-Bench.
Jack Hong, Chenxiao Zhao, ChengLin Zhu et al.
Proposes 'Thinking with Video' paradigm using Sora-2 to unify multimodal reasoning, achieving 92% accuracy on MATH and 69.2% on MMMU benchmarks.
Jingqi Tong, Yurong Mou, Hangcheng Li et al.