← all papers · overview

Data Driven Optimization Of GPU Efficiency For Distributed LLM Adapter Serving

Abstract

Large Language Model (LLM) adapters enable low-cost model specialization, but introduce complex caching and scheduling challenges in distributed serving systems where hundreds of adapters must be hosted concurrently. While prior work has largely focused on latency minimization, resource efficiency through throughput maximization remains underexplored. This paper presents a data-driven pipeline tha

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).