Awesome Speech Audio
📄
Papers
🧭
Topics
🔥
Trending
🗺️
Map
🏆
Leaderboards
🎓
Learn
🤖
Ask AI
⋯
More
👥
Authors
📚
Reading Packs
📊
Datasets
🛠️
Tools
📰
News
📝
Blogs
✉️
Newsletter
🎯
Research Radar
🔖
Saved
+ Add Paper
☾
☀
← authors
·
overview
Loading author…
🤖
Ask AI
Daisuke Saito — most-cited papers & profile · Speech Audio
← authors
·
overview
Daisuke Saito
18
papers ·
70
citations ·
0
h-index
Yokohama National University · Kyorin University · Tharawal Aboriginal · Kyorin University Hospital · The University of Tokyo
Google Scholar ↗
Semantic Scholar ↗
OpenAlex ↗
Most-cited papers
The Voice Conversion Challenge 2018: Promoting Development of Parallel and Nonparallel Methods
2018 · 66 citations
Benchmarking Prosody Encoding in Discrete Speech Tokens
2025 · 2 citations
Wasserstein GAN and Waveform Loss-based Acoustic Model Training for Multi-speaker Text-to-Speech Synthesis Systems Using a WaveNet Vocoder
2018 · 1 citations
Simulating Native Speaker Shadowing for Nonnative Speech Assessment with Latent Speech Representations
2024 · 1 citations
Self-supervised Speech Comparison for L2 Phone, Rhythm, and Intonation Scoring
2026
Leveraging Soft Distributions of SSL-Derived Discrete Speech Tokens for Downstream Inference
2026
Kanade: A Simple Disentangled Tokenizer for Spoken Language Modeling
2026
SSL-GMMVC: Interpretable Voice Conversion via Locally Linear GMM Transforms in Self-Supervised Representation Space
2026
Prosodic ABX: A Language-Agnostic Method for Measuring Prosodic Contrast in Speech Representations
2026
Advanced Modeling of Interlanguage Speech Intelligibility Benefit with L1-L2 Multi-Task Learning Using Differentiable K-Means for Accent-Robust Discrete Token-Based ASR
2026
MixedG2P-T5: G2P-free Speech Synthesis for Mixed-script texts using Speech Self-Supervised Learning and Language Model
2025
Discrete Tokens Exhibit Interlanguage Speech Intelligibility Benefit: an Analytical Study Towards Accent-robust ASR Only with Native Speech Data
2025
Prosodically Enhanced Foreign Accent Simulation by Discrete Token-based Resynthesis Only with Native Speech Corpora
2025
A Perception-Based L2 Speech Intelligibility Indicator: Leveraging a Rater's Shadowing and Sequence-to-sequence Voice Conversion
2025
A Pilot Study of GSLM-based Simulation of Foreign Accentuation Only Using Native Speech Corpora
2024
Top co-authors
Nobuaki Minematsu
· 16
Stephen McIntosh
· 3
Eunjung Yeo
· 1
Herman Kamper
· 1
Kwanghee Choi
· 1
Tomoki Toda
· 1
Yi Zhao
· 1
Zhenhua Ling
· 1
Zhijie Huang
· 1
Topics
Speech Recognition
Audio Generation
cs.SD
eess.AS
Text-to-Speech
cs.CL
Voice Cloning
Audio Understanding
Speaker Analysis
cs.LG