← all datasets

NExT-GQA

Emerging
5papers using it
2023first seen

Can I Trust Your Answer? Visually Grounded Video Question Answering Introduction We study visually grounded VideoQA by forcing vision-language models (VLMs) to answer questions and simultaneously ground the relevant video moments as visual evidences. We show that this task is easy for human yet is extremely challenging

Papers using NExT-GQA (5)

NExT-GQA dataset β€” papers, benchmarks & downloads Β· Multimodal