Judge Anything: MLLM as a Judge Across Any Modality
This paper introduces JudgeAnything, a benchmark leveraging multimodal large language models (MLLMs) as automated judges across 15 diverse tasks, revealing strong performance in understanding but limitations in generation.
Shu Pu, Yaochen Wang, Dongping Chen et al.