← all papers · overview

One Model Is Enough: Native Retrieval Embeddings From LLM Agent Hidden States

Bo Jiang·2026

Abstract

LLM agents that retrieve external knowledge typically generate a search query as text, then run a separate embedding model to encode it into a vector. This two-model pipeline adds infrastructure complexity and latency, yet is redundant: the LLM already encodes the full conversational context in its hidden states. We propose equipping LLM agents with native retrieval capability by adding a lightwei

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).