Awesome Multimodal
📄
Papers
🧭
Topics
🔥
Trending
🗺️
Map
🏆
Leaderboards
🎓
Learn
🤖
Ask AI
⋯
More
👥
Authors
📚
Reading Packs
📊
Datasets
🛠️
Tools
📰
News
📝
Blogs
✉️
Newsletter
🎯
Research Radar
🔖
Saved
+ Add Paper
☾
☀
← authors
·
overview
Loading author…
🤖
Ask AI
Ting Yao — most-cited papers & profile · Multimodal
← authors
·
overview
Ting Yao
33
papers ·
1098
citations ·
58
h-index
Nanjing University of Posts and Telecommunications · Peptidream (Japan)
Google Scholar ↗
Semantic Scholar ↗
OpenAlex ↗
Most-cited papers
X-Linear Attention Networks for Image Captioning
2020 · 689 citations
Scheduled Sampling in Vision-Language Pretraining with Decoupled Encoder-Decoder Network
2021 · 8 citations
Semantic-Conditional Diffusion Networks for Image Captioning
2022 · 6 citations
CoCo-BERT: Improving Video-Language Pre-training with Contrastive Cross-modal Matching and Denoising
2021 · 2 citations
Uni-EDEN: Universal Encoder-Decoder Network by Multi-Granular Vision-Language Pre-training
2022 · 1 citations
Unleashing Text-to-Image Diffusion Prior for Zero-Shot Image Captioning
2025
Top co-authors
Yingwei Pan
· 6
Hongyang Chao
· 3
Tao Mei
· 3
Jianjie Luo
· 2
Jianlin Feng
· 2
Tao Mei
· 2
Yehao Li
· 2
Jiahao Fan
· 1
Jianjie Luo
· 1
Jingwen Chen
· 1
Jingwen Chen
· 1
Weiyao Lin
· 1
Topics
Vision-Language Models
Visual QA & Reasoning
Video-Language
Image-Text Retrieval
Benchmarks