← all papers · overview

Language Model Cascades: Token-level Uncertainty And Beyond

Abstract

Recent advances in language models (LMs) have led to significant improvements in quality on complex NLP tasks, but at the expense of increased inference costs. Cascading offers a simple strategy to achieve more favorable cost-quality tradeoffs: here, a small model is invoked for most "easy" instances, while a few "hard" instances are deferred to the large model. While the principles underpinning c

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).