← all papers · overview

Infinitehip: Extending Language Model Context Up To 3 Million Tokens On A Single GPU

Abstract

In modern large language models (LLMs), handling very long context lengths presents significant challenges as it causes slower inference speeds and increased memory costs. Additionally, most existing pre-trained LLMs fail to generalize beyond their original training sequence lengths. To enable efficient and practical long-context utilization, we introduce InfiniteHiP, a novel, and practical LLM in

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).