← all papers · overview

Amplification Effects In Test-time Reinforcement Learning: Safety And Reasoning Vulnerabilities

Abstract

Test-time training (TTT) has recently emerged as a promising method to improve the reasoning abilities of large language models (LLMs), in which the model directly learns from test data without access to labels. However, this reliance on test data also makes TTT methods vulnerable to harmful prompt injections. In this paper, we investigate safety vulnerabilities of TTT methods, where we study a re

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).