← all papers · overview

Reinforcement Learning With Promising Tokens For Large Language Models

Abstract

Reinforcement learning (RL) has emerged as a key paradigm for aligning and optimizing large language models (LLMs). Standard approaches treat the LLM as the policy and apply RL directly over the full vocabulary space. However, this formulation includes the massive tail of contextually irrelevant tokens in the action space, which could distract the policy from focusing on decision-making among the

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).