← all papers · overview

BATON: Enhancing Batch-wise Inference Efficiency For Large Language Models Via Dynamic Re-batching

Abstract

The advanced capabilities of Large Language Models (LLMs) have inspired the development of various interactive web services or applications, such as ChatGPT, which offer query inference services for users. Unlike traditional DNN model, the inference of LLM entails different iterations of forward computation for different queries, which result in efficiency challenges for existing run-to-completion

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).