← all papers · overview

Improving The Language Understanding Capabilities Of Large Language Models Using Reinforcement Learning

Abstract

Instruction-fine-tuned large language models (LLMs) under 14B parameters continue to underperform on natural language understanding (NLU) tasks, often trailing smaller models like BERT-base on benchmarks such as GLUE and SuperGLUE. Motivated by the success of reinforcement learning in reasoning tasks (e.g., DeepSeek), we explore Proximal Policy Optimization (PPO) as a framework to improve the NLU

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).