← all papers · overview

Self-judge: Selective Instruction Following With Alignment Self-evaluation

Abstract

Pre-trained large language models (LLMs) can be tailored to adhere to human instructions through instruction tuning. However, due to shifts in the distribution of test-time data, they may not always execute instructions accurately, potentially generating factual errors or misaligned content when acting as chat assistants. To enhance the reliability of LLMs in following instructions, we propose the

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).