← all papers · overview

Plato: Plan To Efficiently Decode For Large Language Model Inference

Abstract

Large language models (LLMs) have achieved remarkable success in natural language tasks, but their inference incurs substantial computational and memory overhead. To improve efficiency, parallel decoding methods like Skeleton-of-Thought (SoT) decompose prompts into sub-problems for concurrent processing. However, these methods significantly compromise answer quality by treating semantically linked

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).