← all papers · overview

A Text Is Worth Several Tokens: Text Embedding From Llms Secretly Aligns Well With The Key Tokens

Abstract

Text embeddings from large language models (LLMs) have achieved excellent results in tasks such as information retrieval, semantic textual similarity, etc. In this work, we show an interesting finding: when feeding a text into the LLM-based embedder, the obtained text embedding will be able to be aligned with the key tokens in the input text. We first fully analyze this phenomenon on eight LLM-bas

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).