Awesome Multimodal
📄
Papers
🧭
Topics
🔥
Trending
🗺️
Map
🏆
Leaderboards
🎓
Learn
🤖
Ask AI
⋯
More
👥
Authors
📚
Reading Packs
📊
Datasets
🛠️
Tools
📰
News
📝
Blogs
✉️
Newsletter
🎯
Research Radar
🔖
Saved
+ Add Paper
☾
☀
← authors
·
overview
Loading author…
🤖
Ask AI
Trevor Darrell — most-cited papers & profile · Multimodal
← authors
·
overview
Trevor Darrell
41
papers ·
19060
citations ·
139
h-index
University of California, Berkeley
Google Scholar ↗
Semantic Scholar ↗
OpenAlex ↗
Most-cited papers
Multimodal Compact Bilinear Pooling for Visual Question Answering and Visual Grounding
2016 · 395 citations
Large Language Models are Visual Reasoning Coordinators
2023 · 14 citations
Are You Looking? Grounding to Multiple Modalities in Vision-and-Language Navigation
2019 · 13 citations
Multitask Vision-Language Prompt Tuning
2022 · 13 citations
Compositional Chain-of-Thought Prompting for Large Multimodal Models
2023 · 4 citations
Multimodal Explanations: Justifying Decisions and Pointing to the Evidence
2018 · 2 citations
Latent Implicit Visual Reasoning
2025
DAVE: A VLM Vision Encoder For Document Understanding And Web Agents
2025
Reconstruction Alignment Improves Unified Multimodal Models
2025
Puzzled by Puzzles: When Vision-Language Models Can't Take a Hint
2025
Generate, but Verify: Reducing Hallucination in Vision-Language Models with Retrospective Resampling
2025
Modular Visual Question Answering via Code Generation
2023
From Wrong To Right: A Recursive Approach Towards Vision-Language Explanation
2023
Vision-Language Models Create Cross-Modal Task Representations
2024
Enhancing Few-Shot Vision-Language Classification with Large Multimodal Model Features
2024
Top co-authors
Roei Herzig
· 4
Anna Rohrbach
· 3
Brandon Huang
· 3
Jiaxin Ge
· 3
Chancharik Mitra
· 2
Dan Klein
· 2
David M. Chan
· 2
Dong Huk Park
· 2
Heekyung Lee
· 2
Joseph E. Gonzalez
· 2
Kurt Keutzer
· 2
Leonid Karlinsky
· 2
Topics
Vision-Language Models
Visual QA & Reasoning
Video-Language
Benchmarks
Embodied & Agents
Instruction Tuning
Audio-Visual