← all papers · overview

C3PO: Critical-layer, Core-expert, Collaborative Pathway Optimization For Test-time Expert Re-mixing

Abstract

Mixture-of-Experts (MoE) Large Language Models (LLMs) suffer from severely sub-optimal expert pathways-our study reveals that naive expert selection learned from pretraining leaves a surprising 10-20% accuracy gap for improvement. Motivated by this observation, we develop a novel class of test-time optimization methods to re-weight or "re-mixing" the experts in different layers jointly for each te

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).