← all papers · overview

Chef: A Comprehensive Evaluation Framework For Standardized Assessment Of Multimodal Large Language Models

Abstract

Multimodal Large Language Models (MLLMs) have shown impressive abilities in interacting with visual content with myriad potential downstream tasks. However, even though a list of benchmarks has been proposed, the capabilities and limitations of MLLMs are still not comprehensively understood, due to a lack of a standardized and holistic evaluation framework. To this end, we present the first Compre

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).