← all papers · overview

CMR Scaling Law: Predicting Critical Mixture Ratios For Continual Pre-training Of Language Models

Abstract

Large Language Models (LLMs) excel in diverse tasks but often underperform in specialized fields due to limited domain-specific or proprietary corpus. Continual pre-training (CPT) enhances LLM capabilities by imbuing new domain-specific or proprietary knowledge while replaying general corpus to prevent catastrophic forgetting. The data mixture ratio of general corpus and domain-specific corpus, ho

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).