Awesome Speech Audio
📄
Papers
🧭
Topics
🔥
Trending
🗺️
Map
🏆
Leaderboards
🎓
Learn
🤖
Ask AI
⋯
More
👥
Authors
📚
Reading Packs
📊
Datasets
🛠️
Tools
📰
News
📝
Blogs
✉️
Newsletter
🎯
Research Radar
🔖
Saved
+ Add Paper
☾
☀
← authors
·
overview
Loading author…
🤖
Ask AI
Pankaj Wasnik — most-cited papers & profile · Speech Audio
← authors
·
overview
Pankaj Wasnik
19
papers ·
16
citations ·
9
h-index
Sony (Taiwan) · Sony Corporation (United States)
Google Scholar ↗
Semantic Scholar ↗
OpenAlex ↗
Most-cited papers
DubWise: Video-Guided Speech Duration Control in Multimodal LLM-based Text-to-Speech for Dubbing
2024 · 1 citations
Graph-Based Phonetic Error Correction of Noisy ASR
2026
CrossAccent-TTS: Cross-Lingual Accent-Intensity Controllable Text-to-Speech via Disentangled Speaker and Accent Representations
2026
Adaptive Oscillatory Inductive Bias for Modeling Sharp Prosodic Dynamics in Diffusion-Based TTS
2026
Windowed SummaryMixing: An Efficient Fine-Tuning of Self-Supervised Learning Models for Low-resource Speech Recognition
2026
Attention Is Not Always the Answer: Optimizing Voice Activity Detection with Simple Feature Fusion
2025
LASPA: Language Agnostic Speaker Disentanglement with Prefix-Tuned Cross-Attention
2025
Isometric Neural Machine Translation using Phoneme Count Ratio Reward-based Reinforcement Learning
2024
Efficient infusion of self-supervised representations in Automatic Speech Recognition
2024
VECL-TTS: Voice identity and Emotional style controllable Cross-Lingual Text-to-Speech
2024
Enhancing Whisper's Accuracy and Speed for Indian Languages through Prompt-Tuning and Tokenization
2024
EmoReg: Directional Latent Vector Modeling for Emotional Intensity Regularization in Diffusion-based Voice Conversion
2024
Top co-authors
Rajiv Ratn Shah
· 4
Ashishkumar Gudmalwar
· 3
Kumud Tripathi
· 3
Nirmesh J. Shah
· 3
Nirmesh Shah
· 3
Ashishkumar P. Gudmalwar
· 2
Mohammadi Zaki
· 2
Aditya Srinivas Menon
· 1
Aditya Srinivas Menon
· 1
Aneesh Mukkamala
· 1
Ankit Tatawat
· 1
Ashishkumar P. Gudmalwar
· 1
Topics
Speech Recognition
Text-to-Speech
Voice Cloning
Audio Generation
Speech Translation
Audio Understanding
Speaker Analysis
Multimodal Audio