← all papers · overview

Explainable Token-level Noise Filtering For LLM Fine-tuning Datasets

Abstract

Large Language Models (LLMs) have seen remarkable advancements, achieving state-of-the-art results in diverse applications. Fine-tuning, an important step for adapting LLMs to specific downstream tasks, typically involves further training on corresponding datasets. However, a fundamental discrepancy exists between current fine-tuning datasets and the token-level optimization mechanism of LLMs: mos

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).