← all papers · overview

D-CPT Law: Domain-specific Continual Pre-training Scaling Law For Large Language Models

Abstract

Continual Pre-Training (CPT) on Large Language Models (LLMs) has been widely used to expand the model's fundamental understanding of specific downstream domains (e.g., math and code). For the CPT on domain-specific LLMs, one important question is how to choose the optimal mixture ratio between the general-corpus (e.g., Dolma, Slim-pajama) and the downstream domain-corpus. Existing methods usually

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).