← all papers · overview

Memhunter: Automated And Verifiable Memorization Detection At Dataset-scale In Llms

Abstract

Large language models (LLMs) have been shown to memorize and reproduce content from their training data, raising significant privacy concerns, especially with web-scale datasets. Existing methods for detecting memorization are primarily sample-specific, relying on manually crafted or discretely optimized memory-inducing prompts generated on a per-sample basis, which become impractical for dataset-

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).