Awesome Multimodal
📄
Papers
🧭
Topics
🔥
Trending
🗺️
Map
🏆
Leaderboards
🎓
Learn
🤖
Ask AI
⋯
More
👥
Authors
📚
Reading Packs
📊
Datasets
🛠️
Tools
📰
News
📝
Blogs
✉️
Newsletter
🎯
Research Radar
🔖
Saved
+ Add Paper
☾
☀
← authors
·
overview
Loading author…
🤖
Ask AI
Zhijie Yan — most-cited papers & profile · Multimodal
← authors
·
overview
Zhijie Yan
26
papers ·
238
citations ·
15
h-index
Beijing Sport University
Google Scholar ↗
Semantic Scholar ↗
OpenAlex ↗
Most-cited papers
MinMo: A Multimodal Large Language Model for Seamless Voice Interaction
2025 · 86 citations
Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models
2023 · 22 citations
LauraGPT: Listen, Attend, Understand, and Regenerate Audio with GPT
2023 · 18 citations
A Real-time Speaker Diarization System Based on Spatial Spectrum
2021 · 17 citations
Speaker Overlap-aware Neural Diarization for Multi-party Meeting Analysis
2022 · 15 citations
Deep-FSMN for Large Vocabulary Continuous Speech Recognition
2018 · 13 citations
Streaming Chunk-Aware Multihead Attention for Online End-to-End Speech Recognition
2020 · 13 citations
Linear networks based speaker adaptation for speech synthesis
2018 · 12 citations
M2MeT: The ICASSP 2022 Multi-Channel Multi-Party Meeting Transcription Challenge
2021 · 11 citations
Automatic Spelling Correction with Transformer for CTC-based End-to-End Speech Recognition
2019 · 9 citations
Deep Feed-forward Sequential Memory Networks for Speech Synthesis
2018 · 8 citations
Neural Zero-Inflated Quality Estimation Model For Automatic Speech Recognition System
2019 · 3 citations
CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens
2024 · 3 citations
CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models
2024 · 3 citations
Paraformer: Fast and Accurate Parallel Transformer for Non-autoregressive End-to-End Speech Recognition
2022 · 2 citations
Topics
Speech Recognition
Text-to-Speech
Audio Understanding
Speech Translation
Speaker Analysis
Audio Generation
Multimodal Audio
Voice Cloning
Speech Enhancement
Manipulation