Awesome Multimodal
๐
Papers
๐งญ
Topics
๐ฅ
Trending
๐บ๏ธ
Map
๐
Leaderboards
๐
Learn
๐ค
Ask AI
โฏ
More
๐ฅ
Authors
๐
Reading Packs
๐
Datasets
๐ ๏ธ
Tools
๐ฐ
News
๐
Blogs
โ๏ธ
Newsletter
๐ฏ
Research Radar
๐
Saved
+ Add Paper
โพ
โ
โ authors
ยท
overview
Loading authorโฆ
๐ค
Ask AI
Yu Qiao โ most-cited papers & profile ยท Multimodal
โ authors
ยท
overview
Yu Qiao
31
papers ยท
476
citations ยท
109
h-index
Kyung Hee University ยท Beijing Academy of Artificial Intelligence ยท Shanghai Artificial Intelligence Laboratory
Google Scholar โ
Semantic Scholar โ
OpenAlex โ
Most-cited papers
LLaMA-Adapter V2: Parameter-Efficient Visual Instruction Model
2023 ยท 119 citations
Are We on the Right Way for Evaluating Large Vision-Language Models?
2024 ยท 44 citations
EmbodiedGPT: Vision-Language Pre-Training via Embodied Chain of Thought
2023 ยท 41 citations
InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition
2023 ยท 31 citations
LVLM-eHub: A Comprehensive Evaluation Benchmark for Large Vision-Language Models
2023 ยท 20 citations
JourneyDB: A Benchmark for Generative Image Understanding
2023 ยท 11 citations
InternLM-XComposer2: Mastering Free-form Text-Image Composition and Comprehension in Vision-Language Large Model
2024 ยท 11 citations
Structured Triplet Learning with POS-tag Guided Attention for Visual Question Answering
2018 ยท 10 citations
MM-Interleaved: Interleaved Image-Text Generative Modeling via Multi-modal Feature Synchronizer
2024 ยท 6 citations
MMT-Bench: A Comprehensive Multimodal Benchmark for Evaluating Large Vision-Language Models Towards Multitask AGI
2024 ยท 6 citations
ChartAssisstant: A Universal Chart Multimodal Language Model via Chart-to-Table Pre-training and Multitask Instruction Tuning
2024 ยท 4 citations
Uni-Perceiver v2: A Generalist Model for Large-Scale Vision and Vision-Language Tasks
2022 ยท 3 citations
B-AVIBench: Towards Evaluating the Robustness of Large Vision-Language Model on Black-box Adversarial Visual-Instructions
2024 ยท 3 citations
TinyLVLM-eHub: Towards Comprehensive and Efficient Evaluation for Large Vision-Language Models
2023 ยท 2 citations
MLLMs-Augmented Visual-Language Representation Learning
2023 ยท 1 citations
Top co-authors
Hongsheng Li
ยท 8
Wenqi Shao
ยท 7
Ping Luo
ยท 6
Jifeng Dai
ยท 5
Fanqing Meng
ยท 4
Haodong Duan
ยท 4
Kaipeng Zhang
ยท 4
Conghui He
ยท 3
Dahua Lin
ยท 3
Wenhai Wang
ยท 3
Xizhou Zhu
ยท 3
Aojun Zhou
ยท 2
Topics
Vision-Language Models
Video-Language
Benchmarks
Visual QA & Reasoning
Instruction Tuning
Image-Text Retrieval
Embodied & Agents
Audio-Visual