← all papers · overview

Easier Painting Than Thinking: Can Text-to-image Models Set The Stage, But Not Direct The Play?

Abstract

Text-to-image (T2I) generation aims to synthesize images from textual prompts, which jointly specify what must be shown and imply what can be inferred, which thus correspond to two core capabilities: \textbf\{\textit\{composition\}\} and \textbf\{\textit\{reasoning\}\}. Despite recent advances of T2I models in both composition and reasoning, existing benchmarks remain limited in evaluation. They not only fail to provide comprehensive coverage across and within both capabilities, but also largely restrict evaluation to low scene density and simple one-to-one reasoning. To address these limitations, we propose \textbf\{\textsc\{T2I-CoReBench\}\}, a comprehensive and complex benchmark that evaluates both composition and reasoning capabilities of T2I models. To ensure comprehensiveness, we structure composition around scene graph elements (\textit\{instance\}, \textit\{attribute\}, and \textit\{relation\}) and reasoning around the philosophical framework of inference (\textit\{deductive\},

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).