Abstract
Alzheimer’s disease and related dementia (ADRD) remain underdiagnosed due to limited clinical resources, motivating scalable, speech-based screening tools. Transformer-based pipelines have shown promise in detecting acoustic and linguistic markers of cognitive decline, but their performance is limited by scarce patient speech data. In this work, we propose two generative speech augmentation pipelines to address this limitation: (1) a text-to-speech pipeline using fine-tuned large language models and zero-shot synthesis to generate language specific, demographically targeted utterances, and (2) a voice conversion pipeline that diversifies speaker profiles while preserving linguistic content. We evaluate these pipelines with multimodal (SpeechCARE-AGF) and acoustic-only (SpeechCARE-Whisper) transformer architectures. Results show consistent diagnostic improvements, with text-to-speech achieving the largest gains (SpeechCARE-Whisper: Micro-F1 = 90.1, F1-ADRD = 90.4), outperforming baseline (Micro-F1 = 80.2, F1-ADRD = 82.9). This establishes state-of-the-art performance in ADRD detection from spontaneous speech and demonstrates that generative synthesis provides clinically relevant variability beyond conventional perturbations.