Awesome Multimodal
📄
Papers
🧭
Topics
🔥
Trending
🗺️
Map
🏆
Leaderboards
🎓
Learn
🤖
Ask AI
⋯
More
👥
Authors
📚
Reading Packs
📊
Datasets
🛠️
Tools
📰
News
📝
Blogs
✉️
Newsletter
🎯
Research Radar
🔖
Saved
+ Add Paper
☾
☀
← authors
·
overview
Loading author…
🤖
Ask AI
Xiaoye Qu — most-cited papers & profile · Multimodal
← authors
·
overview
Xiaoye Qu
53
papers ·
30
citations ·
26
h-index
Huazhong University of Science and Technology Hospital · Huazhong University of Science and Technology
Google Scholar ↗
Semantic Scholar ↗
OpenAlex ↗
Most-cited papers
Fine-grained Iterative Attention Network for TemporalLanguage Localization in Videos
2020 · 13 citations
Reducing the Vision and Language Bias for Temporal Sentence Grounding
2022 · 3 citations
Mitigating Multilingual Hallucination in Large Vision-Language Models
2024 · 1 citations
Alleviating Hallucination in Large Vision-Language Models with Active Retrieval Augmentation
2024 · 1 citations
Look, Compare, Decide: Alleviating Hallucination in Large Vision-Language Models via Multi-View Multi-Path Reasoning
2024 · 1 citations
Persistent Visual Memory: Sustaining Perception For Deep Generation In Lvlms
2026
Spotlight on Token Perception for Multimodal Reinforcement Learning
2025
Advancing Multimodal Reasoning: From Optimized Cold Start to Staged Reinforcement Learning
2025
Diffthinker: Towards Generative Multimodal Reasoning With Diffusion Models
2025
Framethinker: Learning To Think With Long Videos Via Multi-turn Frame Spotlighting
2025
SATORI-R1: Incentivizing Multimodal Reasoning through Explicit Visual Anchoring
2025
Progressively Guide to Attend: An Iterative Alignment Framework for Temporal Sentence Grounding
2021
Unified Multi-modal Unsupervised Representation Learning for Skeleton-based Action Understanding
2023
SURf: Teaching Large Vision-Language Models to Selectively Utilize Retrieved Information
2024
Top co-authors
Yu Cheng
· 8
Daizong Liu
· 5
Yafu Li
· 5
Wei Wei
· 4
Zefeng He
· 4
Jiashuo Sun
· 2
Pan Zhou
· 2
Siyuan Huang
· 2
Zhaochen Su
· 2
Jiacheng Chen
· 1
Jiayu Chen
· 1
Jihai Zhang
· 1
Topics
Vision-Language Models
Visual QA & Reasoning
Video-Language
Benchmarks
Image-Text Retrieval
Audio-Visual
cs.MM
cs.RO