← all papers · overview

Data-efficient LLM Fine-tuning For Code Generation

Abstract

Large language models (LLMs) have demonstrated significant potential in code generation tasks. However, there remains a performance gap between open-source and closed-source models. To address this gap, existing approaches typically generate large amounts of synthetic data for fine-tuning, which often leads to inefficient training. In this work, we propose a data selection strategy in order to imp

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).