Awesome Multimodal
📄
Papers
🧭
Topics
🔥
Trending
🗺️
Map
🏆
Leaderboards
🎓
Learn
🤖
Ask AI
⋯
More
👥
Authors
📚
Reading Packs
📊
Datasets
🛠️
Tools
📰
News
📝
Blogs
✉️
Newsletter
🎯
Research Radar
🔖
Saved
+ Add Paper
☾
☀
← authors
·
overview
Loading author…
🤖
Ask AI
Haizhou Li — most-cited papers & profile · Multimodal
← authors
·
overview
Haizhou Li
159
papers ·
455
citations ·
74
h-index
National University of Singapore · Shenzhen Research Institute of Big Data · Chinese University of Hong Kong, Shenzhen
Google Scholar ↗
Semantic Scholar ↗
OpenAlex ↗
Most-cited papers
An Overview of Voice Conversion and its Challenges: From Statistical Modeling to Deep Learning
2020 · 26 citations
Speaker Extraction with Co-Speech Gestures Cue
2022 · 25 citations
Overview of Speaker Modeling and Its Applications: From the Lens of Deep Speaker Representation Learning
2024 · 24 citations
Expressive TTS Training with Frame and Style Reconstruction Loss
2020 · 20 citations
Leveraging Acoustic and Linguistic Embeddings from Pretrained speech and language Models for Intent Classification
2021 · 19 citations
SpEx+: A Complete Time Domain Speaker Extraction Network
2020 · 18 citations
S2S-Arena: Evaluating Paralinguistic Instruction Following in Speech-to-Speech Models
2025 · 17 citations
SongBloom: Coherent Song Generation via Interleaved Autoregressive Sketching and Diffusion Refinement
2025 · 16 citations
Converting Anyone's Emotion: Towards Speaker-Independent Emotional Voice Conversion
2020 · 15 citations
Seen and Unseen emotional style transfer for voice conversion with a new emotional speech dataset
2020 · 14 citations
On the End-to-End Solution to Mandarin-English Code-switching Speech Recognition
2018 · 13 citations
GraphSpeech: Syntax-Aware Graph Attention Network For Neural Speech Synthesis
2020 · 13 citations
Pre-Training Transformer Decoder for End-to-End ASR Model with Unpaired Speech Data
2022 · 13 citations
Generative x-vectors for text-independent speaker verification
2018 · 11 citations
A Modularized Neural Network with Language-Specific Output Layers for Cross-lingual Voice Conversion
2019 · 11 citations
Topics
Speech Recognition
Audio Understanding
Speaker Analysis
Audio Generation
Text-to-Speech
Speech Enhancement
Speech Translation
Multimodal Audio
Voice Cloning
eess.AS