Awesome Multimodal
📄
Papers
🧭
Topics
🔥
Trending
🗺️
Map
🏆
Leaderboards
🎓
Learn
🤖
Ask AI
⋯
More
👥
Authors
📚
Reading Packs
📊
Datasets
🛠️
Tools
📰
News
📝
Blogs
✉️
Newsletter
🎯
Research Radar
🔖
Saved
+ Add Paper
☾
☀
← authors
·
overview
Loading author…
🤖
Ask AI
Junhyeok Lee — most-cited papers & profile · Multimodal
← authors
·
overview
Junhyeok Lee
20
papers ·
25
citations ·
0
h-index
Johns Hopkins University
Google Scholar ↗
Semantic Scholar ↗
OpenAlex ↗
Most-cited papers
CapSpeech: Enabling Downstream Applications in Style-Captioned Text-to-Speech
2025 · 15 citations
Direct Preference-based Policy Optimization without Reward Modeling
2023 · 5 citations
Assem-VC: Realistic Voice Conversion by Assembling Modern Speech Synthesis Techniques
2021 · 3 citations
Super Monotonic Alignment Search
2024 · 1 citations
PITS: Variational Pitch Inference without Fundamental Frequency for End-to-End Pitch-controllable TTS
2023 · 1 citations
Reconstruct! Don't Encode: Self-Supervised Representation Reconstruction Loss for High-Intelligibility and Low-Latency Streaming Neural Audio Codec
2026
Motif-2-12.7B-Reasoning: A Practitioner's Guide to RL Training Recipes
2025
Expanding Foundational Language Capabilities in Open-Source LLMs through a Korean Case Study
2025
MaskVCT: Masked Voice Codec Transformer for Zero-Shot Voice Conversion With Increased Controllability via Multiple Guidances
2025
Improving Test-Time Performance of RVQ-based Neural Codecs
2025
Motif 2.6B Technical Report
2025
Motif 2 12.7B technical report
2025
Controllable and Interpretable Singing Voice Decomposition via Assem-VC
2021
Talking Face Generation with Multilingual TTS
2022
NU-Wave 2: A General Neural Audio Upsampling Model for Various Sampling Rates
2022
Topics
Audio Generation
Text-to-Speech
Speech Enhancement
Training Techniques
Model Architecture
Voice Cloning
eess.AS
cs.AI
Evaluation
Fine-Tuning