cs.CV 2505.07062

Seed1.5-VL Technical Report

Seed1.5-VL combines a 532M vision encoder with a 20B MoE LLM, achieving SOTA on 38 of 60 benchmarks, excelling in multimodal reasoning.

Dong Guo, Faming Wu, Feida Zhu et al.

2025-05-12 40
cs.CV 2505.05071

FG-CLIP: Fine-Grained Visual and Textual Alignment

FG-CLIP leverages 1.6 billion long caption-image pairs, 12M region annotations, and 10M hard negatives to enhance fine-grained alignment.

Chunyu Xie, Bin Wang, Fanjing Kong et al.

2025-05-08 33
cs.CV 2504.19724

RepText: Rendering Visual Text via Replicating

RepText employs a copying mechanism within a diffusion framework to accurately replicate multilingual visual text without understanding, achieving 85.4% accuracy on ICDAR.

Haofan Wang, Yujia Xu, Yimeng Li et al.

2025-04-28 44