Awesome Multimodal
๐
Papers
๐งญ
Topics
๐ฅ
Trending
๐บ๏ธ
Map
๐
Leaderboards
๐
Learn
๐ค
Ask AI
โฏ
More
๐ฅ
Authors
๐
Reading Packs
๐
Datasets
๐ ๏ธ
Tools
๐ฐ
News
๐
Blogs
โ๏ธ
Newsletter
๐ฏ
Research Radar
๐
Saved
+ Add Paper
โพ
โ
โ authors
ยท
overview
Loading authorโฆ
๐ค
Ask AI
Mohit Bansal โ most-cited papers & profile ยท Multimodal
โ authors
ยท
overview
Mohit Bansal
41
papers ยท
1552
citations ยท
61
h-index
University of North Carolina at Chapel Hill ยท University of North Carolina Health Care ยท King's College London ยท King's College Hospital NHS Foundation Trust ยท ABES Engineering College
Google Scholar โ
Semantic Scholar โ
OpenAlex โ
Most-cited papers
LXMERT: Learning Cross-Modality Encoder Representations from Transformers
2019 ยท 223 citations
How Much Can CLIP Benefit Vision-and-Language Tasks?
2021 ยท 153 citations
Unifying Vision-and-Language Tasks via Text Generation
2021 ยท 64 citations
PanoGen: Text-Conditioned Panoramic Environment Generation for Vision-and-Language Navigation
2023 ยท 18 citations
Contrastive Region Guidance: Improving Grounding in Vision-Language Models without Training
2024 ยท 13 citations
WinoGAViL: Gamified Association Benchmark to Challenge Vision-and-Language Models
2022 ยท 6 citations
MLP Architectures for Vision-and-Language Modeling: An Empirical Study
2021 ยท 5 citations
Improving Cross-Modal Alignment in Vision Language Navigation via Syntactic Information
2021 ยท 4 citations
CoDi-2: In-Context, Interleaved, and Interactive Any-to-Any Generation
2023 ยท 3 citations
Vision-and-Language Navigation Today and Tomorrow: A Survey in the Era of Foundation Models
2024 ยท 3 citations
Object Ordering with Bidirectional Matchings for Visual Reasoning
2018 ยท 1 citations
Rephrase, Augment, Reason: Visual Grounding of Questions for Vision-Language Models
2023 ยท 1 citations
CREMA: Generalizable and Efficient Video-Language Reasoning via Multimodal Modular Fusion
2024 ยท 1 citations
Knowledge-Aware Reasoning over Multimodal Semi-structured Tables
2024 ยท 1 citations
Prune-Then-Plan: Step-Level Calibration for Stable Frontier Exploration in Embodied Question Answering
2025
Top co-authors
Hao Tan
ยท 5
Elias Stengel-Eskin
ยท 2
Jaemin Cho
ยท 2
Jialu Li
ยท 2
Shoubin Yu
ยท 2
Anna Rohrbach
ยท 1
Archiki Prasad
ยท 1
Chenguang Zhu
ยท 1
Chenguang Zhu
ยท 1
Dan Roth
ยท 1
David Wan
ยท 1
Gabriel Stanovsky
ยท 1
Topics
Vision-Language Models
Visual QA & Reasoning
Video-Language
Benchmarks
Embodied & Agents
Audio-Visual
Instruction Tuning