← all papers · overview

When Llms Meet Cunning Texts: A Fallacy Understanding Benchmark For Large Language Models

Abstract

Recently, Large Language Models (LLMs) make remarkable evolutions in language understanding and generation. Following this, various benchmarks for measuring all kinds of capabilities of LLMs have sprung up. In this paper, we challenge the reasoning and understanding abilities of LLMs by proposing a FaLlacy Understanding Benchmark (FLUB) containing cunning texts that are easy for humans to understa

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).