← all papers · overview

Brainbench: Exposing The Commonsense Reasoning Gap In Large Language Models

Abstract

Large language models (LLMs) achieve impressive scores on standard benchmarks yet routinely fail questions that any human would answer correctly in seconds. We introduce BrainBench, a benchmark of 100 brainteaser questions spanning 20 carefully designed categories, each targeting a specific commonsense reasoning failure mode in LLMs. Categories range from implicit physical constraints ("Should I w

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).