← all papers · overview

Dovetail: A CPU/GPU Heterogeneous Speculative Decoding For LLM Inference

Abstract

With the continuous advancement in the performance of large language models (LLMs), their demand for computational resources and memory has significantly increased, which poses major challenges for efficient inference on consumer-grade devices and legacy servers. These devices typically feature relatively weaker GPUs and stronger CPUs. Although techniques such as parameter offloading and partial o

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).