Abstract
Large Vision-Language Models (LVLMs) have experienced significant advancements in recent years. However, their performance still falls short in tasks requiring deep visual perception, such as identifying subtle differences between images. A potential cause is the scarcity of visual knowledge in popular instruction-tuning corpora, resulting in inadequate visual perception and reasoning capabilities. To address this challenge, we introduce a self-improvement framework grounded in a novel visual knowledge-intensive task, \underline\{C\}ausality-driven \underline\{V\}isual object \underline\{C\}ompletion (CVC). This task requires LVLMs to infer the masked object in an image based on its \textit\{causal\} relationships with the other visible information. We first obtain rich examples cheaply through our automated instance construction pipeline, without relying on sophisticated LVLMs (\textit\{e.g.\}, GPT-4V) or human assistance. Then, LVLMs effectively self-improve through trial and error l