Awesome Multimodal
📄
Papers
🧭
Topics
🔥
Trending
🗺️
Map
🏆
Leaderboards
🎓
Learn
🤖
Ask AI
⋯
More
👥
Authors
📚
Reading Packs
📊
Datasets
🛠️
Tools
📰
News
📝
Blogs
✉️
Newsletter
🎯
Research Radar
🔖
Saved
+ Add Paper
☾
☀
← authors
·
overview
Loading author…
🤖
Ask AI
Yu Wang — most-cited papers & profile · Multimodal
← authors
·
overview
Yu Wang
331
papers ·
3277
citations ·
28
h-index
China University of Geosciences (Beijing) · University of Florida
Google Scholar ↗
Semantic Scholar ↗
OpenAlex ↗
Most-cited papers
Mitigating Visual Knowledge Forgetting in MLLM Instruction-tuning via Modality-decoupled Gradient Descent
2025 · 1 citations
Urban Socio-Semantic Segmentation with Vision-Language Reasoning
2026
Fedmgp: Personalized Federated Learning With Multi-group Text-visual Prompts
2025
Fighting Fire With Fire (F3): A Training-free And Efficient Visual Adversarial Example Purification Method In Lvlms
2025
Cosmos 3: Omnimodal World Models for Physical AI
2026
STAR: Semantic-Temporal Adaptive Representation Learning for Few-Shot Action Recognition
2026
MHSA: A Lightweight Framework for Mitigating Hallucinations via Steered Attention in LVLMs
2026
Emotrans: A Benchmark For Understanding, Reasoning, And Predicting Emotion Transitions In Multimodal Llms
2026
CLEAR: Null-Space Projection for Cross-Modal De-Redundancy in Multimodal Recommendation
2026
Exposing Cross-Modal Consistency for Fake News Detection in Short-Form Videos
2026
StreamingVLA: Streaming Vision-Language-Action Model with Action Flow Matching and Adaptive Early Observation
2026
Q-probe: Scaling Image Quality Assessment To High Resolution Via Context-aware Agentic Probing
2026
SimpleVLA-RL: Scaling VLA Training via Reinforcement Learning
2025
GLM-4.5V and GLM-4.1V-Thinking: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning
2025
Gatefusion: Hierarchical Gated Cross-modal Fusion For Active Speaker Detection
2025
Top co-authors
Ruobing Xie
· 3
Xingwu Sun
· 3
Yanfeng Wang
· 3
Di Wang
· 2
Hongcheng Liu
· 2
Xinyu Zhang
· 2
Yanpeng Sun
· 2
Yan Wang
· 2
Yiqing Huang
· 2
Zhanhui Kang
· 2
Aarti Basant
· 1
Akash Gokul
· 1
Topics
Vision-Language Models
Benchmarks
Video-Language
Visual QA & Reasoning
Embodied & Agents
Audio-Visual
Image-Text Retrieval
Instruction Tuning
cs.MM