← all papers · overview

Performance Characterization Of Expert Router For Scalable LLM Inference

Abstract

Large Language Models (LLMs) have experienced widespread adoption across scientific and industrial domains due to their versatility and utility for diverse tasks. Nevertheless, deploying and serving these models at scale with optimal throughput and latency remains a significant challenge, primarily because of LLMs' high computational and memory demands. Specialized models optimized for specific ta

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).