← all papers · overview

Capacity-aware Mixture Law Enables Efficient LLM Data Optimization

Abstract

A data mixture refers to how different data sources are combined to train large language models, and selecting an effective mixture is crucial for optimal downstream performance. Existing methods either conduct costly searches directly on the target model or rely on mixture scaling laws that fail to extrapolate well to large model sizes. We address these limitations by introducing a compute-effici

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).