← all papers · overview

Porover: Improving Safety And Reducing Overrefusal In Large Language Models With Overgeneration And Preference Optimization

Abstract

Achieving both high safety and high usefulness simultaneously in large language models has become a critical challenge in recent years.Models often exhibit unsafe behavior or adopt an overly cautious approach leading to frequent overrefusal of benign prompts, which reduces their usefulness. A major factor underlying these behaviors is how the models are finetuned and aligned, particularly the natu

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).