Awesome Multimodal
📄
Papers
🧭
Topics
🔥
Trending
🗺️
Map
🏆
Leaderboards
🎓
Learn
🤖
Ask AI
⋯
More
👥
Authors
📚
Reading Packs
📊
Datasets
🛠️
Tools
📰
News
📝
Blogs
✉️
Newsletter
🎯
Research Radar
🔖
Saved
+ Add Paper
☾
☀
← authors
·
overview
Loading author…
🤖
Ask AI
Wei Zhang — most-cited papers & profile · Multimodal
← authors
·
overview
Wei Zhang
337
papers ·
1424
citations ·
14
h-index
Xiamen University
Google Scholar ↗
Semantic Scholar ↗
OpenAlex ↗
Most-cited papers
Co-attending Free-form Regions and Detections with Multi-modal Multiplicative Feature Embedding for Visual Question Answering
2017 · 18 citations
Scale, Don't Fine-tune: Guiding Multimodal Llms For Efficient Visual Place Recognition At Test-time
2025 · 3 citations
ERNIE 5.0 Technical Report
2026 · 2 citations
Guiding Cross-Modal Representations with MLLM Priors via Preference Alignment
2025 · 1 citations
Illuminating Unified Multimodal Model for Free-form Interleaved Text-Image Generation
2026
UniFit: Towards Universal Virtual Try-on with MLLM-Guided Semantic Alignment
2025
Vldrive: Vision-augmented Lightweight Mllms For Efficient Language-grounded Autonomous Driving
2025
Beyond Isolated Utterances: Cue-Guided Interaction for Context-Dependent Conversational Multimodal Understanding
2026
VLA-IAP: Training-Free Visual Token Pruning via Interaction Alignment for Vision-Language-Action Models
2026
CoT4Det: A Chain-of-Thought Framework for Perception-Oriented Vision-Language Tasks
2025
JavisGPT: A Unified Multi-modal LLM for Sounding-Video Comprehension and Generation
2025
Seeing is Believing: Rich-Context Hallucination Detection for MLLMs via Backward Visual Grounding
2025
$\mathcalVisi\mathcalPruner$: Decoding Discontinuous Cross-Modal Dynamics for Efficient Multimodal LLMs
2025
EarthGPT-X: A Spatial MLLM for Multi-level Multi-Source Remote Sensing Imagery Understanding with Visual Prompting
2025
HiRes-LLaVA: Restoring Fragmentation Input in High-Resolution Large Vision-Language Models
2024
Top co-authors
Chunwei Wang
· 3
Hang Xu
· 3
Lu Hou
· 3
Pan Lu
· 3
Runhui Huang
· 3
Tong Zhang
· 3
Fan Li
· 2
Guansong Lu
· 2
Hongsheng Li
· 2
Jianhua Han
· 2
Jianyong Wang
· 2
Jingdong Wang
· 2
Topics
Vision-Language Models
Video-Language
Visual QA & Reasoning
Benchmarks
Image-Text Retrieval
Instruction Tuning
Audio-Visual
Embodied & Agents