NExT-GQA
Emerging5papers using it
2023first seen
Can I Trust Your Answer? Visually Grounded Video Question Answering Introduction We study visually grounded VideoQA by forcing vision-language models (VLMs) to answer questions and simultaneously ground the relevant video moments as visual evidences. We show that this task is easy for human yet is extremely challenging
Papers using NExT-GQA (5)
- Chrono: A Simple Blueprint for Representing Time in MLLMsLeAdQA: LLM-Driven Context-Aware Temporal Grounding for Video Question AnsweringMUPA: Towards Multi-Path Agentic Reasoning for Grounded Video Question AnsweringCan I Trust Your Answer? Visually Grounded Video Question AnsweringLanguage Repository for Long Video Understanding