← all papers · overview

Entropy-guided Sequence Weighting For Efficient Exploration In Rl-based LLM Fine-tuning

Abstract

We introduce Entropy-Guided Sequence Weighting (EGSW), a novel approach that enhances the exploration-exploitation tradeoff by dynamically assigning weights to generated outputs based on their advantage and entropy for Reinforcement Learning-based Large Language Model fine-tuning. EGSW integrates entropy regularization with advantage-based weighting to balance policy updates, enabling efficient ex

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).