Awesome Multimodal
📄
Papers
🧭
Topics
🔥
Trending
🗺️
Map
🏆
Leaderboards
🎓
Learn
🤖
Ask AI
⋯
More
👥
Authors
📚
Reading Packs
📊
Datasets
🛠️
Tools
📰
News
📝
Blogs
✉️
Newsletter
🎯
Research Radar
🔖
Saved
+ Add Paper
☾
☀
← authors
·
overview
Loading author…
🤖
Ask AI
Yan Huang — most-cited papers & profile · Multimodal
← authors
·
overview
Yan Huang
62
papers ·
228
citations ·
44
h-index
China Mobile (China) · Dongguan University of Technology · South China University of Technology
Google Scholar ↗
Semantic Scholar ↗
OpenAlex ↗
Most-cited papers
BEVBert: Multimodal Map Pre-training for Language-guided Navigation
2022 · 11 citations
Open-Vocabulary Octree-Graph for 3D Scene Understanding
2024 · 9 citations
Frequency-domain Multi-modal Fusion for Language-guided Medical Image Segmentation
2025 · 2 citations
Neighbor-view Enhanced Model for Vision and Language Navigation
2021 · 2 citations
Improving Vision-Language-Action Model Fine-Tuning with Structured Stage and Keyframe Supervision
2026
Multi-View Video Diffusion Policy: A 3D Spatio-Temporal-Aware Video Action Model
2026
UAOR: Uncertainty-aware Observation Reinjection for Vision-Language-Action Models
2026
Ec-flow: Enabling Versatile Robotic Manipulation From Action-unlabeled Videos Via Embodiment-centric Flow
2025
Bridgevla: Input-output Alignment For Efficient 3D Manipulation Learning With Vision-language Models
2025
CMF: Cascaded Multi-model Fusion for Referring Image Segmentation
2021
ETPNav: Evolving Topological Planning for Vision-Language Navigation in Continuous Environments
2023
Efficient Token-Guided Image-Text Retrieval with Consistent Multimodal Contrastive Training
2023
Top co-authors
Liang Wang
· 9
Peiyan Li
· 5
Tieniu Tan
· 4
Dong An
· 3
Jiabing Yang
· 3
Bin Zhao
· 2
Jing Liu
· 2
Kai Wang
· 2
Tao Kong
· 2
Zichen Wen
· 2
Angang Du
· 1
Bohong Yin
· 1
Topics
Vision-Language Models
Embodied & Agents
Video-Language
Benchmarks
Instruction Tuning
Image-Text Retrieval