← all papers · overview

Exploiting Voice Activity Detection Vulnerabilities: A Universal Adversarial Perturbation Framework for Speech Pipeline Disruption

Abstract

Modern speech-controlled systems rely on Voice Activity Detection (VAD) as the critical gatekeeper in speech processing pipelines. Although adversarial attacks on Automatic Speech Recognition (ASR) have been extensively studied, VAD security remains largely unexplored, exposing a fundamental vulnerability. This paper introduces Silent Deception, a Universal Adversarial Perturbation (UAP) framework designed to force VAD models to misclassify active speech as silence. By targeting VAD rather than ASR, the proposed attack achieves effective speech pipeline disruption, creating a silent denial-of-service (DoS) condition where downstream components never receive valid input. The UAPs are crafted using gradient-based optimization on Silero VAD and WebRTC VAD, maximizing the False Negative Rate (FNR) while strictly preserving perceptual quality. Evaluation demonstrates a 90% bypass success rate and significant ASR degradation, measured via Word Error Rate (WER). This work highlights the urgent need for adversarial robustness in VAD systems as a primary defense capability in next- generation speech pipelines.

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).