← all papers · overview

Image First Or Text First? Optimising The Sequencing Of Modalities In Large Language Model Prompting And Reasoning Tasks

Abstract

This paper examines how the sequencing of images and text within multi-modal prompts influences the reasoning performance of large language models (LLMs). We performed empirical evaluations using three commercial LLMs. Our results demonstrate that the order in which modalities are presented can significantly affect performance, particularly in tasks of varying complexity. For simpler tasks involvi

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).