← all papers · overview

Autojudge: Judge Decoding Without Manual Annotation

Abstract

We introduce AutoJudge, a method that accelerates large language model (LLM) inference with task-specific lossy speculative decoding. Instead of matching the original model output distribution token-by-token, we identify which of the generated tokens affect the downstream quality of the response, relaxing the distribution match guarantee so that the "unimportant" tokens can be generated faster. Ou

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).