← all datasets

JudgeAnything

Emerging
1papers using it
2025first seen

JudgeAnything is a benchmark designed to evaluate the judging capabilities of multimodal large language models (MLLMs) across various tasks by incorporating human judgments and detailed rubrics for pair comparison and score evaluation.

Papers using JudgeAnything (1)

JudgeAnything dataset β€” papers, benchmarks & downloads Β· Large Language Models