safety
loadingβ¦
loadingβ¦
safety is one of the most active areas in Awesome Large Language Models β 30 papers in this collection, evaluated on datasets like VL-RewardBench, Multimodal RewardBench. A strong starting point is "COMPASS: A Framework for Evaluating Organization-Specific Policy Alignment in LLMs".