Awesome Multimodal
📄
Papers
🧭
Topics
🔥
Trending
🗺️
Map
🏆
Leaderboards
🎓
Learn
🤖
Ask AI
⋯
More
👥
Authors
📚
Reading Packs
📊
Datasets
🛠️
Tools
📰
News
📝
Blogs
✉️
Newsletter
🎯
Research Radar
🔖
Saved
+ Add Paper
☾
☀
← authors
·
overview
Loading author…
🤖
Ask AI
Dacheng Tao — most-cited papers & profile · Multimodal
← authors
·
overview
Dacheng Tao
234
papers ·
1914
citations ·
155
h-index
The University of Sydney · Nanyang Technological University · UNSW Sydney
Google Scholar ↗
Semantic Scholar ↗
OpenAlex ↗
Most-cited papers
Multi-modal Factorized Bilinear Pooling with Co-Attention Learning for Visual Question Answering
2017 · 102 citations
Deep Modular Co-Attention Networks for Visual Question Answering
2019 · 99 citations
Multimodal Unified Attention Networks for Vision-and-Language Interactions
2019 · 34 citations
Unified Discrete Diffusion for Simultaneous Vision-Language Generation
2022 · 8 citations
CLAMP: Prompt-based Contrastive Learning for Connecting Language and Animal Pose
2022 · 3 citations
ESceme: Vision-and-Language Navigation with Episodic Scene Memory
2023 · 1 citations
FORGE: Fine-grained Multimodal Evaluation for Manufacturing Scenarios
2026
Learn to Think: Improving Multimodal Reasoning through Vision-Aware Self-Improvement Training
2026
VideoTIR: Accurate Understanding for Long Videos with Efficient Tool-Integrated Reasoning
2026
From Hindsight to Foresight: Self-Encouraged Hindsight Distillation for Knowledge-based Visual Question Answering
2025
Review Of Hallucination Understanding In Large Language And Vision Models
2025
Hyper-modal Imputation Diffusion Embedding with Dual-Distillation for Federated Multimodal Knowledge Graph Completion
2025
Adavideorag: Omni-contextual Adaptive Retrieval-augmented Efficient Long Video Understanding
2025
MLLM-Guided VLM Fine-Tuning with Joint Inference for Zero-Shot Composed Image Retrieval
2025
Optmerge: Unifying Multimodal LLM Capabilities And Modalities Via Model Merging
2025
Top co-authors
Li Shen
· 5
Jun Yu
· 4
Zhou Yu
· 4
Liang Ding
· 3
Siyuan Liang
· 3
Aishan Liu
· 2
Baohang Zhou
· 2
Dadong Wang
· 2
Jianping Fan
· 2
Jing Zhang
· 2
Qi Tian
· 2
Qi Zheng
· 2
Topics
Vision-Language Models
Visual QA & Reasoning
Benchmarks
Video-Language
Image-Text Retrieval
Instruction Tuning
Audio-Visual
Embodied & Agents