← all papers · overview

Rescue: Ranking LLM Responses With Partial Ordering To Improve Response Generation

Abstract

Customizing LLMs for a specific task involves separating high-quality responses from lower-quality ones. This skill can be developed using supervised fine-tuning with extensive human preference data. However, obtaining a large volume of expert-annotated data is costly for most tasks. In this paper, we explore a novel method to optimize LLMs using ranking metrics. This method trains the model to pr

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).