← all papers · overview

Cross-family Speculative Prefill: Training-free Long-context Compression With Small Draft Models

Abstract

Prompt length is a major bottleneck in agentic large language model (LLM) workloads, where repeated inference steps and multi-call loops incur substantial prefill cost. Recent work on speculative prefill demonstrates that attention-based token importance estimation can enable training-free prompt compression, but this assumes the existence of a draft model that shares the same tokenizer as the tar

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).