← all papers · overview

Quantifying Perturbation Impacts For Large Language Models

Abstract

We consider the problem of quantifying how an input perturbation impacts the outputs of large language models (LLMs), a fundamental task for model reliability and post-hoc interpretability. A key obstacle in this domain is disentangling the meaningful changes in model responses from the intrinsic stochasticity of LLM outputs. To overcome this, we introduce Distribution-Based Perturbation Analysis

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).