Awesome Multimodal
π
Papers
π§
Topics
π₯
Trending
πΊοΈ
Map
π
Leaderboards
π
Learn
π€
Ask AI
β―
More
π₯
Authors
π
Reading Packs
π
Datasets
π οΈ
Tools
π°
News
π
Blogs
βοΈ
Newsletter
π―
Research Radar
π
Saved
+ Add Paper
βΎ
β
β authors
Β·
overview
Loading authorβ¦
π€
Ask AI
Gengyuan Zhang β most-cited papers & profile Β· Multimodal
β authors
Β·
overview
Gengyuan Zhang
5
papers Β·
0
citations
Google Scholar β
Semantic Scholar β
OpenAlex β
Most-cited papers
Avila: Asynchronous Vision-language Agent For Streaming Multimodal Data Interaction
2025
Can Vision-Language Models be a Good Guesser? Exploring VLMs for Times and Location Reasoning
2023
Perceive, Query & Reason: Enhancing Video QA with Question-Guided Temporal Queries
2024
Top co-authors
Hermine Kleiner
Β· 1
Kerui Zhang
Β· 1
Lorenzo Baraldi
Β· 1
Rajat Koner
Β· 1
Rita Cucchiara
Β· 1
Roberto Amoroso
Β· 1
Tanveer Hannan
Β· 1
Volker Tresp
Β· 1
Volker Tresp
Β· 1
Yurui Zhang
Β· 1
Topics
Video-Language
Benchmarks
Vision-Language Models
Visual QA & Reasoning
Embodied & Agents