← all papers · overview

Llmservingsim 2.0: A Unified Simulator For Heterogeneous And Disaggregated LLM Serving Infrastructure

Abstract

Large language model (LLM) serving infrastructures are undergoing a shift toward heterogeneity and disaggregation. Modern deployments increasingly integrate diverse accelerators and near-memory processing technologies, introducing significant hardware heterogeneity, while system software increasingly separates computation, memory, and model components across distributed resources to improve scalab

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).