← all papers · overview

Ada-k Routing: Boosting The Efficiency Of Moe-based Llms

Abstract

In the era of Large Language Models (LLMs), Mixture-of-Experts (MoE) architectures offer a promising approach to managing computational costs while scaling up model parameters. Conventional MoE-based LLMs typically employ static Top-K routing, which activates a fixed and equal number of experts for each token regardless of their significance within the context. In this paper, we propose a novel Ad

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).