JudgeAnything
Emerging1papers using it
2025first seen
JudgeAnything is a benchmark designed to evaluate the judging capabilities of multimodal large language models (MLLMs) across various tasks by incorporating human judgments and detailed rubrics for pair comparison and score evaluation.