← all papers · overview

Beyond Task Performance: Evaluating And Reducing The Flaws Of Large Multimodal Models With In-context Learning

Abstract

Following the success of Large Language Models (LLMs), Large Multimodal Models (LMMs), such as the Flamingo model and its subsequent competitors, have started to emerge as natural steps towards generalist agents. However, interacting with recent LMMs reveals major limitations that are hardly captured by the current evaluation benchmarks. Indeed, task performances (e.g., VQA accuracy) alone do not

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).