Awesome Speech Audio
📄
Papers
🧭
Topics
🔥
Trending
🗺️
Map
🏆
Leaderboards
🎓
Learn
🤖
Ask AI
⋯
More
👥
Authors
📚
Reading Packs
📊
Datasets
🛠️
Tools
📰
News
📝
Blogs
✉️
Newsletter
🎯
Research Radar
🔖
Saved
+ Add Paper
☾
☀
← authors
·
overview
Loading author…
🤖
Ask AI
Yuki Saito — most-cited papers & profile · Speech Audio
← authors
·
overview
Yuki Saito
18
papers ·
365
citations ·
17
h-index
Bunkyo University · The University of Tokyo
Google Scholar ↗
Semantic Scholar ↗
OpenAlex ↗
Most-cited papers
Statistical Parametric Speech Synthesis Incorporating Generative Adversarial Networks
2017 · 228 citations
Voice Conversion Using Sequence-to-Sequence Learning of Context Posterior Probabilities
2017 · 54 citations
JVS corpus: free Japanese multi-speaker voice corpus
2019 · 41 citations
The T05 System for The VoiceMOS Challenge 2024: Transfer Learning from Deep Image Classifier to Naturalness MOS Prediction of High-Quality Synthetic Speech
2024 · 22 citations
DNN-based Speaker Embedding Using Subjective Inter-speaker Similarity for Multi-speaker Modeling in Speech Synthesis
2019 · 5 citations
Coco-Nut: Corpus of Japanese Utterance and Voice Characteristics Description for Prompt-based Control
2023 · 5 citations
Phase reconstruction from amplitude spectrograms based on von-Mises-distribution deep neural network
2018 · 3 citations
V2S attack: building DNN-based voice conversion from automatic speaker verification
2019 · 3 citations
Lifter Training and Sub-band Modeling for Computationally Efficient and High-Quality Voice Conversion Using Spectral Differentials
2020 · 3 citations
Generative Moment Matching Network-based Random Modulation Post-filter for DNN-based Singing Voice Synthesis and Neural Double-tracking
2019 · 1 citations
Speaker-Conditioned Phrase Break Prediction for Text-to-Speech with Phoneme-Level Pre-trained Language Model
2025
Analysing the Language of Neural Audio Codecs
2025
StyleCap: Automatic Speaking-Style Captioning from Speech Based on Speech and Language Self-supervised Learning Models
2023
Frame-Wise Breath Detection with Self-Training: An Exploration of Enhancing Breath Naturalness in Text-to-Speech
2024
Building speech corpus with diverse voice characteristics for its prompt-based representation
2024
Top co-authors
Hiroshi Saruwatari
· 11
Shinnosuke Takamichi
· 11
Wataru Nakata
· 5
Tomoki Koriyama
· 4
Detai Xin
· 3
Dong Yang
· 3
Aya Watanabe
· 2
Kazuki Yamauchi
· 2
Takaaki Saeki
· 2
Yusuke Ijima
· 2
and Hiroshi Saruwatari
· 1
Daichi Kitamura
· 1
Topics
Audio Generation
Text-to-Speech
Speaker Analysis
Voice Cloning
Speech Enhancement
Speech Recognition
Audio Understanding
Music Generation
Multimodal Audio