← all papers · overview

Parallelization Strategies For Dense LLM Deployment: Navigating Through Application-specific Tradeoffs And Bottlenecks

Abstract

Breakthroughs in the generative AI domain have fueled an explosion of large language model (LLM)-powered applications, whose workloads fundamentally consist of sequences of inferences through transformer architectures. Within this rapidly expanding ecosystem, dense LLMs--those that activate all model parameters for each token generation--form the foundation for advanced expert-based variants. Dens

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).