← all papers · overview

Reinforcement Learning Fine-tuning Of Language Models Is Biased Towards More Extractable Features

Abstract

Many capable large language models (LLMs) are developed via self-supervised pre-training followed by a reinforcement-learning fine-tuning phase, often based on human or AI feedback. During this stage, models may be guided by their inductive biases to rely on simpler features which may be easier to extract, at a cost to robustness and generalisation. We investigate whether principles governing indu

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).