Awesome Speech Audio
📄
Papers
🧭
Topics
🔥
Trending
🗺️
Map
🏆
Leaderboards
🎓
Learn
🤖
Ask AI
⋯
More
👥
Authors
📚
Reading Packs
📊
Datasets
🛠️
Tools
📰
News
📝
Blogs
✉️
Newsletter
🎯
Research Radar
🔖
Saved
+ Add Paper
☾
☀
← authors
·
overview
Loading author…
🤖
Ask AI
Soumi Maiti — most-cited papers & profile · Speech Audio
← authors
·
overview
Soumi Maiti
20
papers ·
42
citations ·
13
h-index
Carnegie Mellon University
Google Scholar ↗
Semantic Scholar ↗
OpenAlex ↗
Most-cited papers
Improving Massively Multilingual ASR With Auxiliary CTC Objectives
2023 · 26 citations
Speech denoising by parametric resynthesis
2019 · 4 citations
Unsupervised Data Selection for TTS: Using Arabic Broadcast News as a Case Study
2023 · 4 citations
EEND-SS: Joint End-to-End Neural Speaker Diarization and Speech Separation for Flexible Number of Speakers
2022 · 2 citations
Exploring Speech Recognition, Translation, and Understanding with Discrete Speech Units: A Comparative Study
2023 · 2 citations
Generating Multilingual Voices Using Speaker Space Translation Based on Bilingual Speaker Data
2020 · 1 citations
SpeechLMScore: Evaluating speech generation using speech language model
2022 · 1 citations
TMT: Tri-Modal Translation between Speech, Image, and Text by Processing Different Modalities as Different Languages
2024 · 1 citations
Text-To-Speech Synthesis In The Wild
2024 · 1 citations
Speaker independence of neural vocoders and their effect on parametric resynthesis speech enhancement
2019
End-to-End Diarization for Variable Number of Speakers with Local-Global Networks and Discriminative Speaker Embeddings
2021
Learning to Speak from Text: Zero-Shot Multilingual Text-to-Speech with Unsupervised Text Pretraining
2023
ESPnet-ST-v2: Multipurpose Spoken Language Translation Toolkit
2023
Voxtlm: unified decoder-only models for consolidating speech recognition/synthesis and speech/text continuation tasks
2023
Reproducing Whisper-Style Training Using an Open-Source Toolkit and Publicly Available Data
2023
Top co-authors
Shinji Watanabe
· 16
Yifan Peng
· 7
Jee-weon Jung
· 6
Jinchuan Tian
· 5
Wangyou Zhang
· 5
Xuankai Chang
· 5
Shinnosuke Takamichi
· 4
Brian Yan
· 3
Jiatong Shi
· 3
Roshan Sharma
· 3
Takaaki Saeki
· 3
William Chen
· 3
Topics
Speech Recognition
Text-to-Speech
Speech Translation
Speech Enhancement
Audio Understanding
Speaker Analysis
Audio Generation
Music Generation
Voice Cloning
Multimodal Audio