← all papers · overview

Upskill: Mutual Information Skill Learning For Structured Response Diversity In Llms

Abstract

Reinforcement Learning with Verifiable Rewards (RLVR) has improved the reasoning abilities of large language models (LLMs) on mathematics and programming tasks, but standard approaches that optimize single-attempt accuracy can inadvertently suppress response diversity across repeated attempts, narrowing exploration and overlooking underrepresented strategies. We introduce UpSkill, a training time

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).