ScanQA
Emerging8papers using it
2024first seen
ScanQA is a dataset that contains question-answer pairs related to 3D scenes, used to evaluate the performance of models in embodied question answering tasks.
Papers using ScanQA (8)
- SpatialPrompting: Keyframe-driven Zero-Shot Spatial Reasoning with
Off-the-Shelf Multimodal Large Language Models3D Question Answering via only 2D Vision-Language ModelsNot All Modalities Are Equal: Instruction-Aware Gating for Multimodal VideosChat-Scene++: Exploiting Context-Rich Object Identification for 3D LLMChain of Questions: Guiding Multimodal Curiosity in Language Models3D-MoRe: Unified Modal-Contextual Reasoning for Embodied Question AnsweringBridging the Gap between 2D and 3D Visual Question Answering: A Fusion
Approach for 3D VQAEvaluating Zero-Shot GPT-4V Performance on 3D Visual Question Answering
Benchmarks