Awesome Multimodal
📄
Papers
🧭
Topics
🔥
Trending
🗺️
Map
🏆
Leaderboards
🎓
Learn
🤖
Ask AI
⋯
More
👥
Authors
📚
Reading Packs
📊
Datasets
🛠️
Tools
📰
News
📝
Blogs
✉️
Newsletter
🎯
Research Radar
🔖
Saved
+ Add Paper
☾
☀
← authors
·
overview
Loading author…
🤖
Ask AI
William Yang Wang — most-cited papers & profile · Multimodal
← authors
·
overview
William Yang Wang
42
papers ·
1421
citations ·
65
h-index
Massachusetts Institute of Technology
Google Scholar ↗
Semantic Scholar ↗
OpenAlex ↗
Most-cited papers
Reinforced Cross-Modal Matching and Self-Supervised Imitation Learning for Vision-Language Navigation
2018 · 37 citations
GPT-4V(ision) as a Generalist Evaluator for Vision-Language Tasks
2023 · 13 citations
Generate then Select: Open-ended Visual Question Answering Guided by World Knowledge
2023 · 11 citations
Visual Chain of Thought: Bridging Logical Gaps with Multimodal Infillings
2023 · 10 citations
Learning to Stop: A Simple yet Effective Approach to Urban Vision-Language Navigation
2020 · 2 citations
CLIP also Understands Text: Prompting CLIP for Phrase Understanding
2022 · 2 citations
Discffusion: Discriminative Diffusion Models as Few-shot Vision and Language Learners
2023 · 1 citations
ImaginE: An Imagination-Based Automatic Evaluation Metric for Natural Language Generation
2021
Text as Images: Can Multimodal Large Language Models Follow Printed Instructions in Pixels?
2023
Losing Visual Needles in Image Haystacks: Vision Language Models are Easily Distracted in Short and Long Contexts
2024
Top co-authors
Yujie Lu
· 4
An Yan
· 2
Michael Saxon
· 2
Wanrong Zhu
· 2
Xin Eric Wang
· 2
Aditya Sharma
· 1
Alexander Hanbo Li
· 1
Alex Mei
· 1
Andy Ouyang
· 1
An Yan
· 1
Arjun Akula
· 1
Asli Celikyilmaz
· 1
Topics
Vision-Language Models
Video-Language
Image-Text Retrieval
Benchmarks
Visual QA & Reasoning
Embodied & Agents
Audio-Visual
Instruction Tuning