← all papers · overview

Tag&tab: Pretraining Data Detection In Large Language Models Using Keyword-based Membership Inference Attack

Abstract

Large language models (LLMs) have become essential tools for digital task assistance. Their training relies heavily on the collection of vast amounts of data, which may include copyright-protected or sensitive information. Recent studies on detecting pretraining data in LLMs have primarily focused on sentence- or paragraph-level membership inference attacks (MIAs), usually involving probability an

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).