Awesome Multimodal
📄
Papers
🧭
Topics
🔥
Trending
🗺️
Map
🏆
Leaderboards
🎓
Learn
🤖
Ask AI
⋯
More
👥
Authors
📚
Reading Packs
📊
Datasets
🛠️
Tools
📰
News
📝
Blogs
✉️
Newsletter
🎯
Research Radar
🔖
Saved
+ Add Paper
☾
☀
← authors
·
overview
Loading author…
🤖
Ask AI
Shijian Lu — most-cited papers & profile · Multimodal
← authors
·
overview
Shijian Lu
41
papers ·
603
citations ·
68
h-index
Mohamed bin Zayed University of Artificial Intelligence
Google Scholar ↗
Semantic Scholar ↗
OpenAlex ↗
Most-cited papers
Mitigating Object Hallucinations in Large Vision-Language Models through Visual Contrastive Decoding
2023 · 2 citations
Mitigating Object Hallucinations in Large Vision-Language Models with Assembly of Global and Local Attention
2024 · 2 citations
Longvt: Incentivizing "thinking With Long Videos" Via Native Tool Calling
2025
Referring Multiple Regions with Large Multimodal Models via Contextual Latent Steering
2026
Skeleton-to-Image Encoding: Enabling Skeleton Representation Learning via Vision-Pretrained Models
2026
On the Generalization Capacities of MLLMs for Spatial Intelligence
2026
E.M.Ground: A Temporal Grounding Vid-LLM with Holistic Event Perception and Matching
2026
Direction-aware 3D Large Multimodal Models
2026
A Comprehensive Study on Visual Token Redundancy for Discrete Diffusion-based Multimodal Large Language Models
2025
EarthGPT-X: A Spatial MLLM for Multi-level Multi-Source Remote Sensing Imagery Understanding with Visual Prompting
2025
MMRel: Benchmarking Relation Understanding in Multi-Modal Large Language Models
2024
LongHalQA: Long-Context Hallucination Evaluation for MultiModal Large Language Models
2024
The Curse of Multi-Modalities: Evaluating Hallucinations of Large Multimodal Models across Language, Visual, and Audio
2024
Top co-authors
Sicong Leng
· 4
Lidong Bing
· 3
Chunyan Miao
· 2
Deli Zhao
· 2
Hang Zhang
· 2
Xin Li
· 2
Zuhao Yang
· 2
Chong Wang
· 1
Feng Tian
· 1
Guanzheng Chen
· 1
Han Qiu
· 1
Hao Cheng
· 1
Topics
Vision-Language Models
Visual QA & Reasoning
Benchmarks
Video-Language
Audio-Visual
Instruction Tuning