← all papers · overview

The Ups And Downs Of Large Language Model Inference With Vocabulary Trimming By Language Heuristics

Abstract

Deploying large language models (LLMs) encounters challenges due to intensive computational and memory requirements. Our research examines vocabulary trimming (VT) inspired by restricting embedding entries to the language of interest to bolster time and memory efficiency. While such modifications have been proven effective in tasks like machine translation, tailoring them to LLMs demands specific

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).