Awesome Multimodal
📄
Papers
🧭
Topics
🔥
Trending
🗺️
Map
🏆
Leaderboards
🎓
Learn
🤖
Ask AI
⋯
More
👥
Authors
📚
Reading Packs
📊
Datasets
🛠️
Tools
📰
News
📝
Blogs
✉️
Newsletter
🎯
Research Radar
🔖
Saved
+ Add Paper
☾
☀
← authors
·
overview
Loading author…
🤖
Ask AI
Jie Tang — most-cited papers & profile · Multimodal
← authors
·
overview
Jie Tang
112
papers ·
5582
citations ·
83
h-index
University of Leeds · South China University of Technology · Tsinghua University
Google Scholar ↗
Semantic Scholar ↗
OpenAlex ↗
Most-cited papers
NavOne: One-Step Global Planning for Vision-Language Navigation on Top-Down Maps
2026
Action Draft and Verify: A Self-Verifying Framework for Vision-Language-Action Model
2026
GLAD: Generative Language-Assisted Visual Tracking for Low-Semantic Templates
2026
Boosting Multi-modal Keyphrase Prediction with Dynamic Chain-of-Thought in Vision-Language Models
2025
GLM-4.5V and GLM-4.1V-Thinking: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning
2025
Top co-authors
Aohan Zeng
· 1
Baoxu Wang
· 1
Bin Chen
· 1
Bin Xu
· 1
Boyan Shi
· 1
Changyu Pang
· 1
Chao Feng
· 1
Chenhui Zhang
· 1
Chen Zhao
· 1
Da Yin
· 1
Debing Liu
· 1
Dingkang Yang
· 1
Topics
Vision-Language Models
Video-Language
Embodied & Agents
Benchmarks
Instruction Tuning
Visual QA & Reasoning