Awesome Speech Audio
📄
Papers
🧭
Topics
🔥
Trending
🗺️
Map
🏆
Leaderboards
🎓
Learn
🤖
Ask AI
⋯
More
👥
Authors
📚
Reading Packs
📊
Datasets
🛠️
Tools
📰
News
📝
Blogs
✉️
Newsletter
🎯
Research Radar
🔖
Saved
+ Add Paper
☾
☀
← authors
·
overview
Loading author…
🤖
Ask AI
Yinghao Aaron Li — most-cited papers & profile · Speech Audio
← authors
·
overview
Yinghao Aaron Li
11
papers ·
65
citations ·
7
h-index
Columbia University
Google Scholar ↗
Semantic Scholar ↗
OpenAlex ↗
Most-cited papers
StyleTTS 2: Towards Human-Level Text-to-Speech through Style Diffusion and Adversarial Training with Large Speech Language Models
2023 · 23 citations
StyleTTS: A Style-Based Generative Model for Natural and Diverse Text-to-Speech Synthesis
2022 · 16 citations
Phoneme-Level BERT for Enhanced Prosody of Text-to-Speech with Grapheme Predictions
2023 · 14 citations
Improved Decoding of Attentional Selection in Multi-Talker Environments with Self-Supervised Learned Speech Representation
2023 · 4 citations
StarGANv2-VC: A Diverse, Unsupervised, Non-parallel Framework for Natural-Sounding Voice Conversion
2021 · 3 citations
SLMGAN: Exploiting Speech Language Model Representations for Unsupervised Zero-Shot Voice Conversion in GANs
2023 · 2 citations
StyleTTS-VC: One-Shot Voice Conversion by Knowledge Transfer from Style-Based TTS Models
2022 · 1 citations
HiFTNet: A Fast High-Quality Neural Vocoder with Harmonic-plus-Noise Filter and Inverse Short Time Fourier Transform
2023 · 1 citations
Style-Talker: Finetuning Audio Language Model and Style-Based Text-to-Speech Model for Fast Spoken Dialogue Generation
2024 · 1 citations
Speech Slytherin: Examining the Performance and Efficiency of Mamba for Speech Separation, Recognition, and Synthesis
2024
StyleTTS-ZS: Efficient High-Quality Zero-Shot Text-to-Speech Synthesis with Distilled Time-Varying Style Diffusion
2024
Top co-authors
Nima Mesgarani
· 10
Xilin Jiang
· 5
Cong Han
· 4
Cong Han
· 3
Cong Han
· 2
Adrian Nicolas Florea
· 1
Ali Zare
· 1
and Nima Mesgarani
· 1
Gavin Mischler
· 1
Ge Zhu
· 1
Jordan Darefsky
· 1
Vinay S. Raghavan
· 1
Topics
Text-to-Speech
Audio Generation
Speech Enhancement
Voice Cloning
Speech Recognition
Audio Understanding
Speaker Analysis
Speech Translation
Music Generation
Multimodal Audio