← all papers · overview

Pay For Hints, Not Answers: LLM Shepherding For Cost-efficient Inference

Abstract

Large Language Models (LLMs) deliver state-of-the-art performance on complex reasoning tasks, but their inference costs limit deployment at scale. Small Language Models (SLMs) offer dramatic cost savings yet lag substantially in accuracy. Existing approaches - routing and cascading - treat the LLM as an all-or-nothing resource: either the query bypasses the LLM entirely, or the LLM generates a com

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).