Awesome Multimodal
📄
Papers
🧭
Topics
🔥
Trending
🗺️
Map
🏆
Leaderboards
🎓
Learn
🤖
Ask AI
⋯
More
👥
Authors
📚
Reading Packs
📊
Datasets
🛠️
Tools
📰
News
📝
Blogs
✉️
Newsletter
🎯
Research Radar
🔖
Saved
+ Add Paper
☾
☀
← authors
·
overview
Loading author…
🤖
Ask AI
Gedas Bertasius — most-cited papers & profile · Multimodal
← authors
·
overview
Gedas Bertasius
6
papers ·
8
citations ·
26
h-index
University of North Carolina at Chapel Hill · University of North Carolina Health Care
Google Scholar ↗
Semantic Scholar ↗
OpenAlex ↗
Most-cited papers
VX2TEXT: End-to-End Learning of Video-Based Text Generation From Multimodal Inputs
2021 · 8 citations
WatchAct: A Benchmark for Behavior-Grounded Robot Manipulation
2026
Exact: A Video-language Benchmark For Expert Action Analysis
2025
MuMUR : Multilingual Multimodal Universal Retrieval
2022
ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos
2024
Top co-authors
Avinash Madasu
· 1
Baiqi Li
· 1
Ce Zhang
· 1
Devi Parikh
· 1
Estelle Aflalo
· 1
Feihong He
· 1
Gabriela Ben Melech Stan
· 1
Jindong Gu
· 1
Jue Wang
· 1
Lorenzo Torresani
· 1
Md Mohaiminul Islam
· 1
Mingyu Ding
· 1
Topics
Video-Language
Benchmarks
Vision-Language Models
Visual QA & Reasoning
Embodied & Agents
Audio-Visual
Image-Text Retrieval