← all papers · overview

Interpreting Learned Feedback Patterns In Large Language Models

Abstract

Reinforcement learning from human feedback (RLHF) is widely used to train large language models (LLMs). However, it is unclear whether LLMs accurately learn the underlying preferences in human feedback data. We coin the term \textit\{Learned Feedback Pattern\} (LFP) for patterns in an LLM's activations learned during RLHF that improve its performance on the fine-tuning task. We hypothesize that LL

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).