← all papers · overview

L3: DIMM-PIM Integrated Architecture And Coordination For Scalable Long-context LLM Inference

Abstract

Large Language Models (LLMs) increasingly require processing long text sequences, but GPU memory limitations force difficult trade-offs between memory capacity and bandwidth. While HBM-based acceleration offers high bandwidth, its capacity remains constrained. Offloading data to host-side DIMMs improves capacity but introduces costly data swapping overhead. We identify that the critical memory bot

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).