NExT-QA
Emerging22papers using it
2022first seen
The 'NExT-QA' dataset/benchmark is designed for evaluating Video Question Answering (VideoQA) systems by providing a structured framework for assessing their ability to identify critical moments in videos and reason about causal relationships to answer complex questions.
Papers using NExT-QA (22)
- Bridging VideoQA and Video-Guided Agentic Tasks via Generalized Keyframe ExtractionViqagent: Zero-shot Video Question Answering Via Agent With Open-vocabulary Grounding ValidationLeAdQA: LLM-Driven Context-Aware Temporal Grounding for Video Question AnsweringGrid-logat: Grid Based Local And Global Area Transcription For Video Question AnsweringBuilding a Mind Palace: Structuring Environment-Grounded Semantic Graphs for Effective Long Video Analysis with LLMsProgressive Video Condensation with MLLM Agent for Long-form Video UnderstandingBoxTuning: Directly Injecting the Object Box for Multimodal Model Fine-TuningClue Matters: Leveraging Latent Visual Clues to Empower Video ReasoningHORNet: Task-Guided Frame Selection for Video Question Answering with Vision-Language ModelsReMoRa: Multimodal Large Language Model based on Refined Motion Representation for Long-Video UnderstandingStructured Over Scale: Learning Spatial Reasoning from Educational VideoBridging Vision Language Models and Symbolic Grounding for Video Question AnsweringExplicit Abstention Knobs For Predictable Reliability In Video Question AnsweringCotasks: Chain-of-thought Based Video Instruction Tuning TasksVideoMultiAgents: A Multi-Agent Framework for Video Question AnsweringAsk and Remember: A Questions-Only Replay Strategy for Continual Visual Question AnsweringVideoAgent: Long-form Video Understanding with Large Language Model as
AgentDistilling Vision-Language Models on Millions of VideosRethinking Multi-Modal Alignment in Video Question Answering from
Feature and Sample PerspectivesMIST: Multi-modal Iterative Spatial-Temporal Transformer for Long-form
Video Question AnsweringCan I Trust Your Answer? Visually Grounded Video Question AnsweringLanguage Repository for Long Video Understanding