← all papers · overview

XFT: Unlocking The Power Of Code Instruction Tuning By Simply Merging Upcycled Mixture-of-experts

Abstract

We introduce XFT, a simple yet powerful training scheme, by simply merging upcycled Mixture-of-Experts (MoE) to unleash the performance limit of instruction-tuned code Large Language Models (LLMs). While vanilla sparse upcycling fails to improve instruction tuning, XFT introduces a shared expert mechanism with a novel routing weight normalization strategy into sparse upcycling, which significantly

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).