← all papers · overview

Davir: Data Selection Via Implicit Reward For Large Language Models

Abstract

We introduce DavIR, a model-based data selection method for post-training Large Language Models. DavIR generalizes Reducible Holdout Loss to core-set selection problem of causal language modeling, and quantifies the learnability of a given datum with respect to a pre-trained LLM based on relative reduction in loss during fine-tuning, a metric we show to be closely related to the implicit reward mo

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).