Awesome Speech Audio
📄
Papers
🧭
Topics
🔥
Trending
🗺️
Map
🏆
Leaderboards
🎓
Learn
🤖
Ask AI
⋯
More
👥
Authors
📚
Reading Packs
📊
Datasets
🛠️
Tools
📰
News
📝
Blogs
✉️
Newsletter
🎯
Research Radar
🔖
Saved
+ Add Paper
☾
☀
← authors
·
overview
Loading author…
🤖
Ask AI
Atsushi Ando — most-cited papers & profile · Speech Audio
← authors
·
overview
Atsushi Ando
13
papers ·
2
citations ·
24
h-index
Sumitomo Chemical (Japan) · NTT (Japan)
Google Scholar ↗
Semantic Scholar ↗
OpenAlex ↗
Most-cited papers
Speaker consistency loss and step-wise optimization for semi-supervised joint training of TTS and ASR using unpaired text data
2022 · 1 citations
Recursive Attentive Pooling for Extracting Speaker Embeddings from Multi-Speaker Recordings
2024 · 1 citations
Guided Speaker Embedding
2024 · 1 citations
Tight Boundary Prediction in Speaker Diarization Using Causal-Anticausal Consistency
2026
Can We Really Repurpose Multi-Speaker ASR Corpus for Speaker Diarization?
2025
Mitigating Non-Target Speaker Bias in Guided Speaker Embedding
2025
Pretraining Multi-Speaker Identification for Neural Speaker Diarization
2025
Microphone Array Geometry Independent Multi-Talker Distant ASR: NTT System for the DASR Task of the CHiME-8 Challenge
2025
End-to-End Joint Target and Non-Target Speakers ASR
2023
NTT speaker diarization system for CHiME-7: multi-domain, multi-microphone End-to-end and vector clustering diarization
2023
Speech Rhythm-Based Speaker Embeddings Extraction from Phonemes and Phoneme Duration for Multi-Speaker Speech Synthesis
2024
SpeakerBeam-SS: Real-time Target Speaker Extraction with Lightweight Conv-TasNet and State Space Modeling
2024
NTT Multi-Speaker ASR System for the DASR Task of CHiME-8 Challenge
2024
Top co-authors
Marc Delcroix
· 10
Naohiro Tawara
· 9
Shota Horiguchi
· 9
Takanori Ashihara
· 8
Hiroshi Sato
· 6
Takafumi Moriya
· 6
Atsunori Ogawa
· 3
Masato Mimura
· 3
Tsubasa Ochiai
· 3
Alexis Plaquet
· 2
Kohei Matsuura
· 2
Naoki Makishima
· 2
Topics
Speaker Analysis
Speech Recognition
Audio Understanding
eess.AS
cs.SD
Speech Enhancement
Text-to-Speech
Audio Generation