Awesome Multimodal
📄
Papers
🧭
Topics
🔥
Trending
🗺️
Map
🏆
Leaderboards
🎓
Learn
🤖
Ask AI
⋯
More
👥
Authors
📚
Reading Packs
📊
Datasets
🛠️
Tools
📰
News
📝
Blogs
✉️
Newsletter
🎯
Research Radar
🔖
Saved
+ Add Paper
☾
☀
← authors
·
overview
Loading author…
🤖
Ask AI
Wei Han — most-cited papers & profile · Multimodal
← authors
·
overview
Wei Han
51
papers ·
4150
citations ·
19
h-index
Nanjing University of Science and Technology · Guangzhou Railway Polytechnic
Google Scholar ↗
Semantic Scholar ↗
OpenAlex ↗
Most-cited papers
Conformer: Convolution-augmented Transformer for Speech Recognition
2020 · 2748 citations
W2v-BERT: Combining Contrastive Learning and Masked Language Modeling for Self-Supervised Speech Pre-Training
2021 · 262 citations
ContextNet: Improving Convolutional Neural Networks for Automatic Speech Recognition with Global Context
2020 · 257 citations
Pushing the Limits of Semi-Supervised Learning for Automatic Speech Recognition
2020 · 200 citations
Seq-NMS for Video Object Detection
2016 · 154 citations
Google USM: Scaling Automatic Speech Recognition Beyond 100 Languages
2023 · 112 citations
A comparison of end-to-end models for long-form speech recognition
2019 · 85 citations
Noise2Music: Text-conditioned Music Generation with Diffusion Models
2023 · 50 citations
Noise2Music: Text-conditioned Music Generation with Diffusion Models
2023 · 50 citations
AudioPaLM: A Large Language Model That Can Speak and Listen
2023 · 41 citations
Co-training Transformer with Videos and Images Improves Action Recognition
2021 · 31 citations
SLM: Bridge the thin gap between speech and text foundation models
2023 · 28 citations
Dual-mode ASR: Unify and Improve Streaming ASR with Full-context Modeling
2020 · 24 citations
Efficient Domain Adaptation for Speech Foundation Models
2023 · 19 citations
Exploring Targeted Universal Adversarial Perturbations to End-to-end ASR Models
2021 · 13 citations
Topics
Speech Recognition
Speech Translation
Text-to-Speech
Audio Understanding
Audio Generation
Speech Enhancement
Multimodal Audio
Value-Based
Multi-Agent
Model-Based RL