← all datasets

TextVQA

Canonical
20papers using it
1,682HF downloads
38HF likes
2020first seen

TextVQA requires models to read and reason about text in images to answer questions about them. Specifically, models need to incorporate a new modality of text present in the images and reason over it to answer TextVQA questions. TextVQA dataset contains 45,336 questions over 28,408 images from the OpenImages dataset.

Papers using TextVQA (20)

TextVQA dataset β€” papers, benchmarks & downloads Β· Multimodal