← all papers · overview

See What Llms Cannot Answer: A Self-challenge Framework For Uncovering LLM Weaknesses

Abstract

The impressive performance of Large Language Models (LLMs) has consistently surpassed numerous human-designed benchmarks, presenting new challenges in assessing the shortcomings of LLMs. Designing tasks and finding LLMs' limitations are becoming increasingly important. In this paper, we investigate the question of whether an LLM can discover its own limitations from the errors it makes. To this en

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).