Multimodal RewardBench 2: Evaluating Omni Reward Models for Interleaved Text and Image
Introduces MMRB2 benchmark, evaluating omni reward models across four tasks; Gemini 3 Pro achieves 75-80% accuracy, surpassing GPT-4o's 59%.
Yushi Hu, Reyhane Askari-Hemmat, Melissa Hall et al.