← all papers · overview

IWISDM: Assessing Instruction Following In Multimodal Models At Scale

Abstract

The ability to perform complex tasks from detailed instructions is a key to many remarkable achievements of our species. As humans, we are not only capable of performing a wide variety of tasks but also very complex ones that may entail hundreds or thousands of steps to complete. Large language models and their more recent multimodal counterparts that integrate textual and visual inputs have achie

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).