← all papers · overview

It Ain't That Bad: Understanding The Mysterious Performance Drop In OOD Generalization For Generative Transformer Models

Abstract

Large language models (LLMs) have achieved remarkable proficiency on solving diverse problems. However, their generalization ability is not always satisfying and the generalization problem is common for generative transformer models in general. Researchers take basic mathematical tasks like n-digit addition or multiplication as important perspectives for investigating their generalization behavior

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).