← all papers · overview

Chatglm-rlhf: Practices Of Aligning Large Language Models With Human Feedback

Abstract

ChatGLM is a free-to-use AI service powered by the ChatGLM family of large language models (LLMs). In this paper, we present the ChatGLM-RLHF pipeline -- a reinforcement learning from human feedback (RLHF) system -- designed to enhance ChatGLM's alignment with human preferences. ChatGLM-RLHF encompasses three major components: the collection of human preference data, the training of the reward mod

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).