← all papers · overview

DWDP: Distributed Weight Data Parallelism For High-performance LLM Inference On NVL72

Abstract

Large language model (LLM) inference increasingly depends on multi-GPU execution, yet existing inference parallelization strategies require layer-wise inter-rank synchronization, making end-to-end performance sensitive to workload imbalance. We present DWDP (Distributed Weight Data Parallelism), an inference parallelization strategy that preserves data-parallel execution while offloading MoE weigh

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).