Awesome Multimodal
📄
Papers
🧭
Topics
🔥
Trending
🗺️
Map
🏆
Leaderboards
🎓
Learn
🤖
Ask AI
⋯
More
👥
Authors
📚
Reading Packs
📊
Datasets
🛠️
Tools
📰
News
📝
Blogs
✉️
Newsletter
🎯
Research Radar
🔖
Saved
+ Add Paper
☾
☀
← authors
·
overview
Loading author…
🤖
Ask AI
Hao Wu — most-cited papers & profile · Multimodal
← authors
·
overview
Hao Wu
140
papers ·
321
citations ·
23
h-index
Anhui University · China Power Engineering Consulting Group (China)
Google Scholar ↗
Semantic Scholar ↗
OpenAlex ↗
Most-cited papers
Learning the Best Pooling Strategy for Visual Semantic Embedding
2020 · 23 citations
Teampath: Building Multimodal Pathology Experts With Reasoning AI Copilots
2025 · 1 citations
Semi-supervised Multi-modal Medical Image Segmentation For Complex Situations
2025 · 1 citations
Orchestra-o1: Omnimodal Agent Orchestration
2026
ViCA: Efficient Multimodal LLMs with Vision-Only Cross-Attention
2026
Stop Wandering: Efficient Vision-Language Navigation via Metacognitive Reasoning
2026
InterCoG: Towards Spatially Precise Image Editing with Interleaved Chain-of-Grounding Reasoning
2026
Think-as-You-See: Streaming Chain-of-Thought Reasoning for Large Vision-Language Models
2026
HiDrop: Hierarchical Vision Token Reduction in MLLMs via Late Injection, Concave Pyramid Pruning, and Early Exit
2026
UTPTrack: Towards Simple and Unified Token Pruning for Visual Tracking
2026
Speak While Watching: Unleashing TRUE Real-Time Video Understanding Capability of Multimodal Large Language Models
2026
SaFeR-VLM: Toward Safety-aware Fine-grained Reasoning in Multimodal Models
2025
Fusion to Enhance: Fusion Visual Encoder to Enhance Multimodal Language Model
2025
MME-Emotion: A Holistic Evaluation Benchmark for Emotional Intelligence in Multimodal Large Language Models
2025
SCA3D: Enhancing Cross-modal 3D Retrieval via 3D Shape and Caption Paired Data Augmentation
2025
Top co-authors
Xiaoyu Shen
· 5
Junlong Tong
· 4
Yunpu Ma
· 4
Junyan Lin
· 3
Fan Zhang
· 2
Haoxuan Li
· 2
Kun Wang
· 2
XuDong Wang
· 2
Yingqi Fan
· 2
Zhihong Zhu
· 2
Anhao Zhao
· 1
Bin Chen
· 1
Topics
Vision-Language Models
Benchmarks
Video-Language
Visual QA & Reasoning
Embodied & Agents
Image-Text Retrieval
Audio-Visual
eess.IV
q-bio.QM