← all papers · overview

On Limitations Of The Transformer Architecture

Abstract

What are the root causes of hallucinations in large language models (LLMs)? We use Communication Complexity to prove that the Transformer layer is incapable of composing functions (e.g., identify a grandparent of a person in a genealogy) if the domains of the functions are large enough; we show through examples that this inability is already empirically present when the domains are quite small. We

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).