← all papers · overview

Out-of-distribution Generalization Via Composition: A Lens Through Induction Heads In Transformers

Abstract

Large language models (LLMs) such as GPT-4 sometimes appear to be creative, solving novel tasks often with a few demonstrations in the prompt. These tasks require the models to generalize on distributions different from those from training data -- which is known as out-of-distribution (OOD) generalization. Despite the tremendous success of LLMs, how they approach OOD generalization remains an open

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).