← all papers · overview

Moeless: Efficient Moe LLM Serving Via Serverless Computing

Abstract

Large Language Models (LLMs) have become a cornerstone of AI, driving progress across diverse domains such as content creation, search and recommendation systems, and AI-assisted workflows. To alleviate extreme training costs and advancing model scales, Mixture-of-Experts (MoE) has become a popular backbone for modern LLMs, which are commonly served in distributed deployment using expert paralleli

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).