← all papers · overview

Selection Of LLM Fine-tuning Data Based On Orthogonal Rules

Abstract

High-quality training data is critical to the performance of large language models (LLMs). Recent work has explored using LLMs to rate and select data based on a small set of human-designed criteria (rules), but these approaches often rely heavily on heuristics, lack principled metrics for rule evaluation, and generalize poorly to new tasks. We propose a novel rule-based data selection framework t

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).