← all papers · overview

A Semantic Invariant Robust Watermark For Large Language Models

Abstract

Watermark algorithms for large language models (LLMs) have achieved extremely high accuracy in detecting text generated by LLMs. Such algorithms typically involve adding extra watermark logits to the LLM's logits at each generation step. However, prior algorithms face a trade-off between attack robustness and security robustness. This is because the watermark logits for a token are determined by a

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).