Abstract
While current LLM chatbots like GPT-4V bridge the gap between human instructions and visual representations to enable text-image generations, they still lack efficient alignment methods for high-fidelity performance on multiple downstream tasks. In this paper, we propose \textbf\{\}, a novel unified multimodal LLM framework for generating interleaved text-image conversation across v