← all papers · overview

Deep De Finetti: Recovering Topic Distributions From Large Language Models

Abstract

Large language models (LLMs) can produce long, coherent passages of text, suggesting that LLMs, although trained on next-word prediction, must represent the latent structure that characterizes a document. Prior work has found that internal representations of LLMs encode one aspect of latent structure, namely syntax; here we investigate a complementary aspect, namely the document's topic structure.

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).