VCR
Emerging12papers using it
2019first seen
The VCR (Visual Commonsense Reasoning) dataset is used to evaluate a model's reasoning ability in understanding the semantics of visual content and natural language through tasks that require fine-grained visual and textual information.
Papers using VCR (12)
- CoRGI: Verified Chain-of-Thought Reasoning with Post-hoc Visual GroundingCoherent Multimodal Reasoning with Iterative Self-Evaluation for Vision-Language ModelsUniFine: A Unified and Fine-grained Approach for Zero-shot Vision-Language UnderstandingVisualBERT: A Simple and Performant Baseline for Vision and LanguageVL-BERT: Pre-training of Generic Visual-Linguistic RepresentationsKVL-BERT: Knowledge Enhanced Visual-and-Linguistic BERT for Visual
Commonsense ReasoningIdealGPT: Iteratively Decomposing Vision and Language Reasoning via
Large Language ModelsOn Advances in Text Generation from Images Beyond Captioning: A Case
Study in Self-RationalizationUnderstanding ME? Multimodal Evaluation for Fine-grained Visual
CommonsenseImproving Vision-and-Language Reasoning via Spatial Relations ModelingFrom Wrong To Right: A Recursive Approach Towards Vision-Language
ExplanationDo Vision-Language Transformers Exhibit Visual Commonsense? An Empirical
Study of VCR