← all papers · overview

HERM: Benchmarking And Enhancing Multimodal Llms For Human-centric Understanding

Abstract

The significant advancements in visual understanding and instruction following from Multimodal Large Language Models (MLLMs) have opened up more possibilities for broader applications in diverse and universal human-centric scenarios. However, existing image-text data may not support the precise modality alignment and integration of multi-grained information, which is crucial for human-centric visu

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).