Awesome Multimodal
π
Papers
π§
Topics
π₯
Trending
πΊοΈ
Map
π
Leaderboards
π
Learn
π€
Ask AI
β―
More
π₯
Authors
π
Reading Packs
π
Datasets
π οΈ
Tools
π°
News
π
Blogs
βοΈ
Newsletter
π―
Research Radar
π
Saved
+ Add Paper
βΎ
β
β authors
Β·
overview
Loading authorβ¦
π€
Ask AI
Yining Li β most-cited papers & profile Β· Multimodal
β authors
Β·
overview
Yining Li
17
papers Β·
59
citations Β·
0
h-index
Google Scholar β
Semantic Scholar β
OpenAlex β
Most-cited papers
Open-Vocabulary SAM: Segment and Recognize Twenty-thousand Classes Interactively
2024 Β· 43 citations
An Open and Comprehensive Pipeline for Unified Object Grounding and Detection
2024 Β· 7 citations
Towards Language-Driven Video Inpainting via Multimodal Large Language Models
2024 Β· 2 citations
InternLM-XComposer2-4KHD: A Pioneering Large Vision-Language Model Handling Resolutions from 336 Pixels to 4K HD
2024 Β· 2 citations
Achieving Sample and Computational Efficient Reinforcement Learning by Action Space Reduction via Grouping
2023 Β· 1 citations
OMG-Seg: Is One Model Good Enough For All Segmentation?
2024 Β· 1 citations
GTA: A Benchmark for General Tool Agents
2024 Β· 1 citations
DataChef: Cooking Up Optimal Data Recipes for LLM Adaptation via Reinforcement Learning
2026
Provable Last-iterate Convergence For Multi-objective Safe LLM Alignment Via Optimistic Primal-dual
2026
MIG: Automatic Data Selection for Instruction Tuning by Maximizing Information Gain in Semantic Space
2025
InternLM-XComposer2: Mastering Free-form Text-Image Composition and Comprehension in Vision-Language Large Model
2024
MG-LLaVA: Towards Multi-Granularity Visual Instruction Tuning
2024
Auto Cherry-Picker: Learning from High-quality Generative Data Driven by Language
2024
InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output
2024
Topics
Evaluation
Vision-Language
Model Architecture
Training Techniques
Fine-Tuning
Visual Language
Object Detection
Segmentation
Video Understanding
patch