cs.CL 2503.17489

Judge Anything: MLLM as a Judge Across Any Modality

This paper introduces JudgeAnything, a benchmark leveraging multimodal large language models (MLLMs) as automated judges across 15 diverse tasks, revealing strong performance in understanding but limitations in generation.

Shu Pu, Yaochen Wang, Dongping Chen et al.

2025-03-22 33 citations 58
cs.AI 2503.16416

Survey on Evaluation of LLM-based Agents

This survey systematically analyzes evaluation methods for LLM-based agents, covering core capabilities, application benchmarks, evaluation frameworks, and key dimensions.

Asaf Yehudai, Lilach Eden, Alan Li et al.

2025-03-21 214 citations 47