Awesome Multimodal
📄
Papers
🧭
Topics
🔥
Trending
🗺️
Map
🏆
Leaderboards
🎓
Learn
🤖
Ask AI
⋯
More
👥
Authors
📚
Reading Packs
📊
Datasets
🛠️
Tools
📰
News
📝
Blogs
✉️
Newsletter
🎯
Research Radar
🔖
Saved
+ Add Paper
☾
☀
← authors
·
overview
Loading author…
🤖
Ask AI
Jiaming Zhou — most-cited papers & profile · Multimodal
← authors
·
overview
Jiaming Zhou
36
papers ·
10
citations ·
2
h-index
Nankai University · China Institute of Water Resources and Hydropower Research
Google Scholar ↗
Semantic Scholar ↗
OpenAlex ↗
Most-cited papers
Diversifying Spatial-Temporal Perception for Video Domain Generalization
2023 · 4 citations
PB-LRDWWS System for the SLT 2024 Low-Resource Dysarthria Wake-Up Word Spotting Challenge
2024 · 3 citations
StreamMel: Real-Time Zero-shot Text-to-Speech via Interleaved Continuous Autoregressive Modeling
2025 · 1 citations
RA-CLAP: Relation-Augmented Emotional Speaking Style Contrastive Language-Audio Pretraining For Speech Retrieval
2025 · 1 citations
CKGConv: General Graph Convolution with Continuous Kernels
2024 · 1 citations
DIFFA-2: A Practical Diffusion Large Language Model for General Audio Understanding
2026
H-SAGE: Holistic Speaker-Aware Guided Experts for MoE-based Multi-Talker ASR
2026
CosyEdit2: Speech-Editing-Oriented Reinforcement Learning Unlocks Better Zero-Shot TTS
2026
Next Forcing: Causal World Modeling with Multi-Chunk Prediction
2026
CosyEdit: Unlocking End-to-End Speech Editing Capability from Zero-Shot Text-to-Speech Models
2026
SpeechLLM-as-Judges: Towards General and Interpretable Speech Quality Evaluation
2025
GLAD: Global-Local Aware Dynamic Mixture-of-Experts for Multi-Talker ASR
2025
From Watch to Imagine: Steering Long-horizon Manipulation via Human Demonstration and Future Envisionment
2025
End-to-End Humanoid Robot Safe and Comfortable Locomotion Policy
2025
RealTalk-CN: A Realistic Chinese Speech-Text Dialogue Benchmark With Cross-Modal Interaction Analysis
2025
Topics
Speech Recognition
cs.SD
Manipulation
Human-Robot Interaction
Speech Translation
eess.AS
Control
Perception
Speaker Analysis
Text-to-Speech