← all papers · overview

Evolving Diverse Red-team Language Models In Multi-round Multi-agent Games

Abstract

The primary challenge in deploying Large Language Model (LLM) is ensuring its harmlessness. Red team can identify vulnerabilities by attacking LLM to attain safety. However, current efforts heavily rely on single-round prompt designs and unilateral red team optimizations against fixed blue teams. These static approaches lead to significant reductions in generation diversity, known as the mode coll

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).