Periodgrad: Towards Pitch-controllable Neural Vocoder Based On A Diffusion Probabilistic Model
2024 Β· Yukiya Hono, Kei Hashimoto, Yoshihiko Nankaku, et al.
Abstract
This paper presents a neural vocoder based on a denoising diffusion probabilistic model (DDPM) incorporating explicit periodic signals as auxiliary conditioning signals. Recently, DDPM-based neural vocoders have gained prominence as non-autoregressive models that can generate high-quality waveforms. The neural vocoders based on DDPM have the advantage of training with a simple time-domain loss. In practical applications, such as singing voice synthesis, there is a demand for neural vocoders to generate high-fidelity speech waveforms with flexible pitch control. However, conventional DDPM-based neural vocoders struggle to generate speech waveforms under such conditions. Our proposed model aims to accurately capture the periodic structure of speech waveforms by incorporating explicit periodic signals. Experimental results show that our model improves sound quality and provides better pitch control than conventional DDPM-based neural vocoders.
Authors
(none)
Tags
Stats
Related papers
- Specgrad: Diffusion Probabilistic Model Based Neural Vocoder With Adaptive Noise Spectral Shaping (2022)11.49
- Quasi-periodic Parallel Wavegan Vocoder: A Non-autoregressive Pitch-dependent Dilated Convolution Model For Parametric Speech Generation (2020)3.58
- Quasi-periodic Wavenet Vocoder: A Pitch Dependent Dilated Convolution Model For Parametric Speech Generation (2019)7.50
- Wavefit: An Iterative And Non-autoregressive Neural Vocoder Based On Fixed-point Iteration (2022)9.41
- Neuraldps: Neural Deterministic Plus Stochastic Model With Multiband Excitation For Noise-controllable Waveform Generation (2022)0.00
- BDDM: Bilateral Denoising Diffusion Models For Fast And High-quality Speech Synthesis (2022)4.76
- Infergrad: Improving Diffusion Models For Vocoder By Considering Inference In Training (2022)9.41
- Diffar: Denoising Diffusion Autoregressive Model For Raw Speech Waveform Generation (2023)0.00