Awesome Multimodal
📄
Papers
🧭
Topics
🔥
Trending
🗺️
Map
🏆
Leaderboards
🎓
Learn
🤖
Ask AI
⋯
More
👥
Authors
📚
Reading Packs
📊
Datasets
🛠️
Tools
📰
News
📝
Blogs
✉️
Newsletter
🎯
Research Radar
🔖
Saved
+ Add Paper
☾
☀
← authors
·
overview
Loading author…
🤖
Ask AI
Yu-Gang Jiang — most-cited papers & profile · Multimodal
← authors
·
overview
Yu-Gang Jiang
100
papers ·
153
citations ·
0
h-index
Google Scholar ↗
Semantic Scholar ↗
OpenAlex ↗
Most-cited papers
Learning Modality Interaction for Temporal Sentence Localization and Event Captioning in Videos
2020 · 11 citations
Unified Multimodal Pre-training and Prompt-based Tuning for Vision-Language Understanding and Generation
2021 · 8 citations
FoodLMM: A Versatile Food Assistant using Large Multi-modal Model
2023 · 6 citations
INST-IT: Boosting Instance Understanding via Explicit Visual Prompt Instruction Tuning
2024 · 1 citations
A Safety Report on GPT-5.2, Gemini 3 Pro, Qwen3-VL, Grok 4.1 Fast, Nano Banana Pro, and Seedream 4.5
2026
Unison: Benchmarking Unified Multimodal Models via Synergistic Understanding and Generation
2026
VLA-Pro: Cross-Task Procedural Memory Transfer for Vision-Language-Action Models
2026
RoboOmni: Proactive Robot Manipulation in Omni-modal Context
2025
NAP-Tuning: Neural Augmented Prompt Tuning for Adversarially Robust Vision-Language Models
2025
Daily-Omni: Towards Audio-Visual Reasoning with Temporal Alignment across Modalities
2025
EAGLE: Towards Efficient Arbitrary Referring Visual Prompts Comprehension for Multimodal Large Language Models
2024
Adversarial Prompt Distillation for Vision-Language Models
2024
Top co-authors
Zuxuan Wu
· 6
Jingjing Chen
· 5
Xingjun Ma
· 3
Xin Wang
· 3
Bin Zhu
· 2
Shaoxiang Chen
· 2
Xipeng Qiu
· 2
Bojia Zi
· 1
Bo Li
· 1
Chong-Wah Ngo
· 1
Hang Xu
· 1
Henghui Ding
· 1
Topics
Benchmarks
Vision-Language Models
Video-Language
Instruction Tuning
Visual QA & Reasoning
Embodied & Agents
Audio-Visual
Image-Text Retrieval