← all papers · overview

Dynamic Expert Sharing: Decoupling Memory From Parallelism In Mixture-of-experts Diffusion Llms

Abstract

Among parallel decoding paradigms, diffusion large language models (dLLMs) have emerged as a promising candidate that balances generation quality and throughput. However, their integration with Mixture-of-Experts (MoE) architectures is constrained by an expert explosion: as the number of tokens generated in parallel increases, the number of distinct experts activated grows nearly linearly. This re

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).