← all papers · overview

Selfcp: Compressing Over-limit Prompt Via The Frozen Large Language Model Itself

Abstract

Long prompt leads to huge hardware costs when using transformer-based Large Language Models (LLMs). Unfortunately, many tasks, such as summarization, inevitably introduce long documents, and the wide application of in-context learning easily makes the prompt length explode. This paper proposes a Self-Compressor (SelfCP), which employs the target LLM itself to compress over-limit prompts into dense

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).