← all papers · overview

Doc-to-lora: Learning To Instantly Internalize Contexts

Abstract

Long input sequences are central to in-context learning, document understanding, and multi-step reasoning of Large Language Models (LLMs). However, the quadratic attention cost of Transformers makes inference memory-intensive and slow. While context distillation (CD) can transfer information into model parameters, per-prompt distillation is impractical due to training costs and latency. To address

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).