← all papers · overview

Chameleon: A Heterogeneous And Disaggregated Accelerator System For Retrieval-augmented Language Models

Abstract

A Retrieval-Augmented Language Model (RALM) combines a large language model (LLM) with a vector database to retrieve context-specific knowledge during text generation. This strategy facilitates impressive generation quality even with smaller models, thus reducing computational demands by orders of magnitude. To serve RALMs efficiently and flexibly, we propose Chameleon, a heterogeneous accelerator

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).