← all papers · overview

Are You Sure? Challenging Llms Leads To Performance Drops In The Flipflop Experiment

Abstract

The interactive nature of Large Language Models (LLMs) theoretically allows models to refine and improve their answers, yet systematic analysis of the multi-turn behavior of LLMs remains limited. In this paper, we propose the FlipFlop experiment: in the first round of the conversation, an LLM completes a classification task. In a second round, the LLM is challenged with a follow-up phrase like "Ar

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).