Awesome Multimodal
📄
Papers
🧭
Topics
🔥
Trending
🗺️
Map
🏆
Leaderboards
🎓
Learn
🤖
Ask AI
⋯
More
👥
Authors
📚
Reading Packs
📊
Datasets
🛠️
Tools
📰
News
📝
Blogs
✉️
Newsletter
🎯
Research Radar
🔖
Saved
+ Add Paper
☾
☀
← authors
·
overview
Loading author…
🤖
Ask AI
Lu Hou — most-cited papers & profile · Multimodal
← authors
·
overview
Lu Hou
22
papers ·
285
citations ·
0
h-index
Google Scholar ↗
Semantic Scholar ↗
OpenAlex ↗
Most-cited papers
FILIP: Fine-grained Interactive Language-Image Pre-Training
2021 · 206 citations
Wukong: A 100 Million Large-scale Chinese Cross-modal Pre-training Benchmark
2022 · 29 citations
Enabling Multimodal Generation on CLIP via Vision-Language Knowledge Distillation
2022 · 4 citations
Drivepi: Spatial-aware 4D MLLM For Unified Autonomous Driving Understanding, Perception, Prediction And Planning
2025
The Synergy Dilemma of Long-CoT SFT and RL: Investigating Post-Training Techniques for Reasoning VLMs
2025
HiRes-LLaVA: Restoring Fragmentation Input in High-Resolution Large Vision-Language Models
2024
MoPE-CLIP: Structured Pruning for Efficient Vision-Language Models with Module-wise Pruning Error Metric
2024
ILLUME: Illuminating Your LLMs to See, Draw, and Self-Enhance
2024
Top co-authors
Runhui Huang
· 5
Hang Xu
· 4
Guansong Lu
· 3
Lewei Yao
· 3
Wei Zhang
· 3
Xiaodan Liang
· 3
Xin Jiang
· 3
Chunjing Xu
· 2
Chunwei Wang
· 2
Haoli Bai
· 2
Hengshuang Zhao
· 2
Jianhua Han
· 2
Topics
Vision-Language Models
Benchmarks
Video-Language
Visual QA & Reasoning
Image-Text Retrieval
Embodied & Agents