Detect, Describe, Discriminate: Moving Beyond VQA For MLLM Evaluation
2024 Β· Manu Gaur, Darshan Singh S, Makarand Tapaswi
Abstract
Visual Question Answering (VQA) with multiple choice questions enables a vision-centric evaluation of Multimodal Large Language Models (MLLMs). Although it reliably checks the existence of specific visual abilities, it is easier for the model to select an answer from multiple choices (VQA evaluation) than to generate the answer itself. In this work, we offer a novel perspective: we evaluate how well an MLLM understands a specific visual concept by its ability to uniquely describe two extremely similar images that differ only in the targeted visual concept. Specifically, we assess the ability of MLLMs to capture specific points of visual differences using self-retrieval, i.e., by retrieving the target image using its generated caption against the other image in the pair serving as the distractor. We curate 247 highly similar image pairs as part of the D3 benchmark. For each image pair, the model is prompted to: (1) Detect a specific visual difference, and (2) Describe the target image u
Authors
(none)
Tags
Stats
Related papers
- Vision-deepresearch Benchmark: Rethinking Visual And Textual Search For Multimodal Large Language Models (2026)7.27
- VQA4CIR: Boosting Composed Image Retrieval With Visual Question Answering (2023)5.24
- Visual Haystacks: A Vision-centric Needle-in-a-haystack Benchmark (2024)0.00
- From Known To The Unknown: Transferring Knowledge To Answer Questions About Novel Visual And Semantic Concepts (2018)8.82
- Fine-grained Late-interaction Multi-modal Retrieval For Retrieval Augmented Visual Question Answering (2023)5.24
- Leveraging Visual Question Answering For Image-caption Ranking (2016)12.10
- Mrag-bench: Vision-centric Evaluation For Retrieval-augmented Multimodal Models (2024)0.00
- Pixel-grounded Retrieval For Knowledgeable Large Multimodal Models (2026)0.00