← all papers · overview

Quickllama: Query-aware Inference Acceleration For Large Language Models

Abstract

The capacity of Large Language Models (LLMs) to comprehend and reason over long contexts is pivotal for advancements in diverse fields. Yet, they still stuggle with capturing long-distance dependencies within sequences to deeply understand semantics. To address this issue, we introduce Query-aware Inference for LLMs (Q-LLM), a system designed to process extensive sequences akin to human cognition.

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).