cs.CV 2505.14683

Emerging Properties in Unified Multimodal Pretraining

BAGEL employs large-scale interleaved multimodal pretraining with Mixture-of-Transformers, achieving emergent reasoning capabilities surpassing open-source models.

Chaorui Deng, Deyao Zhu, Kunchang Li et al.

2025-05-21 38
cs.CV 2505.12620

BusterX: MLLM-Powered AI-Generated Video Forgery Detection and Explanation

BusterX leverages MLLM and reinforcement learning to detect AI-generated videos, utilizing a 200K high-quality dataset and producing interpretable reasoning chains, significantly improving accuracy and explainability.

Haiquan Wen, Yiwei He, Zhenglin Huang et al.

2025-05-19 28 citations 71
cs.CV 2505.10565

Depth Anything with Any Prior

Proposes Prior Depth Anything, integrating sparse measurements and depth prediction to produce dense, metric depth maps with strong zero-shot generalization.

Zehan Wang, Siyu Chen, Lihe Yang et al.

2025-05-16 35
cs.CV 2505.07818

DanceGRPO: Unleashing GRPO on Visual Generation

This paper introduces DanceGRPO, leveraging stable Group Relative Policy Optimization (GRPO) to significantly improve reinforcement learning for visual content generation, with up to 181% performance gains on benchmarks.

Zeyue Xue, Jie Wu, Yu Gao et al.

2025-05-13 362 citations 31