DocVQA
Emerging13papers using it
2020first seen
The 'DocVQA' dataset is a benchmark that contains documents and associated questions used to evaluate the performance of models in understanding and extracting information from visual documents.
Papers using DocVQA (11)
- Chain-of-Thought Compression Should Not Be Blind: V-Skip for Efficient Multimodal Reasoning via Dual-Path AnchoringSimple Vision-language Math Reasoning Via Rendered TextQianfan-vl: Domain-enhanced Universal Vision-language ModelsDescribe Anything Model for Visual Question Answering on Text-rich ImagesSpatially Grounded Explanations in Vision Language Models for Document Visual Question AnsweringConstructive Distortion: Improving Mllms With Attention-guided Image WarpingInterpret, Prune And Distill Donut : Towards Lightweight Vlms For VQA On DocumentMGA-VQA: Secure And Interpretable Graph-augmented Visual Question Answering With Memory-guided Protection Against Unauthorized Knowledge UseWilddoc: How Far Are We From Achieving Comprehensive And Robust Document Understanding In The Wild?LayoutLMv2: Multi-modal Pre-training for Visually-Rich Document
UnderstandingLLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via
Hierarchical Window Transformer