← all papers · overview

Channel Merging: Preserving Specialization For Merged Experts

Abstract

Lately, the practice of utilizing task-specific fine-tuning has been implemented to improve the performance of large language models (LLM) in subsequent tasks. Through the integration of diverse LLMs, the overall competency of LLMs is significantly boosted. Nevertheless, traditional ensemble methods are notably memory-intensive, necessitating the simultaneous loading of all specialized models into

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).