From Pixels to Words -- Towards Native One-Vision Models at Scale
NEO-ov, a fully native end-to-end vision-language model, supports multi-image and video understanding with superior fine-grained perception and spatial reasoning.
Haiwen Diao, Jiahao Wang, Penghao Wu et al.