← all papers · overview

Taro: Token-level Adaptive Routing For LLM Test-time Alignment

Abstract

Large language models (LLMs) exhibit strong reasoning capabilities but typically require expensive post-training to reach high performance. Recent test-time alignment methods offer a lightweight alternative, but have been explored mainly for preference alignment rather than reasoning. To bridge this gap, we propose, Token-level Adaptive Routing (TARo), which steers frozen LLMs toward structured re

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).