← all papers · overview

Pretraining Data Mixtures Enable Narrow Model Selection Capabilities In Transformer Models

Abstract

Transformer models, notably large language models (LLMs), have the remarkable ability to perform in-context learning (ICL) -- to perform new tasks when prompted with unseen input-output examples without any explicit model training. In this work, we study how effectively transformers can bridge between their pretraining data mixture, comprised of multiple distinct task families, to identify and lea

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).