← all papers · overview

Image-of-thought Prompting For Visual Reasoning Refinement In Multimodal Large Language Models

Abstract

Recent advancements in Chain-of-Thought (CoT) and related rationale-based works have significantly improved the performance of Large Language Models (LLMs) in complex reasoning tasks. With the evolution of Multimodal Large Language Models (MLLMs), enhancing their capability to tackle complex multimodal reasoning problems is a crucial frontier. However, incorporating multimodal rationales in CoT ha

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).