Reinforcement learning (RL) policies may exhibit unsafe behavior and are hard
to explain. We use counterfactual large language model reasoning to enhance RL
policy safety post-training. We show that our approach improves and helps to
explain the RL policy safety.
Related papers
Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).