Awesome Generative Models
π
Papers
π§
Topics
π₯
Trending
πΊοΈ
Map
π
Leaderboards
π
Learn
π€
Ask AI
β―
More
π₯
Authors
π
Reading Packs
π
Datasets
π οΈ
Tools
π°
News
π
Blogs
βοΈ
Newsletter
π―
Research Radar
π
Saved
+ Add Paper
βΎ
β
β all topics
overview
Sound
loadingβ¦
π€
Ask AI
Awesome Sound β curated papers, datasets & benchmarks Β· Awesome Generative Models
β all topics
overview
Sound
22 papers tagged Sound β re-sort below
Papers
π₯ Trending (default)
π Most cited
π Newest first
π€ A β Z by title
22 papers Β· trending (default)
numbers = π₯ heat
Pmeta-TLA: Backdoor Attacks for Speech Classification Models via Meta-Learning with Timbre Leakage Attack
(2026)
Yueming Huang et al.
5.01
Validating the Single Item Kawaii Measure
(2026)
Katie Seaborn et al.
5.01
UT-AISTimprt submission for ICME 2026 Grand Challenge on Academic Text-to-Music Generation
(2026)
Shunsuke Yoshida et al.
4.39
Clean2FX: Label-conditioned modeling for clean-to-effect guitar audio transformations
(2026)
Oliverio Bombicci Pontelli et al.
4.39
Optimal Transport-based Semantic Alignment for LLM-based Audio-Visual Speech Recognition
(2026)
Xugang Lu et al.
4.39
Rethinking Speech Foundation Model Fine-tuning: Better SFT or Better Match?
(2026)
Wangjin Zhou et al.
4.39
AV-JEPA: Extending LeJEPA to Audio-Visual Self-Supervised Learning
(2026)
Benjamin Robson et al.
4.39
Segmental DTW: A Parallelizable Alternative to Dynamic Time Warping
(2026)
TJ Tsai
4.39
A Geometry-Limited Identification Floor and Its Consequences for Voice-Clone Attribution in Professional Voice Actors
(2026)
Shuhei Kato
4.39
AuEmoChat: Authentic Emotion Understanding and Rendering for Conversational Speech Synthesis
(2026)
Zhenqi Jia et al.
4.39
Constrained Hebbian Learning Supports Efficient Representational Allocation under Structural Constraints
(2026)
Patrick Inoue et al.
4.39
Addressing Limited Data in Auditory Attention Decoding with Diffusion Generative Models
(2026)
David Rannaleet et al.
4.39
What the Waveform Knows: Transparent-first Speech and Audio Intelligence with Caption Studio
(2026)
Cheng Siong Chin et al.
4.39
CAPS: A Cascaded Reconstruction Model to Power Saving in Hearables Using Sub-Nyquist Sampling with Bandwidth Extension
(2026)
Tarikul Islam Tamiti et al.
4.39
RPPNet: Perceptually-Grouped Rhythm-Pitch Primitives for Long-Term Structure Melody Generation via Boundary-Aware Modeling
(2026)
Tieyao Zhang et al.
4.39
StellarTTS: Sparse Temporal Embedding for Low-Latency and Robust Speech Synthesis
(2026)
Kaicheng Luo et al.
4.39
Scalable Keyword Spotting via Modular Network Expansion
(2026)
Viktor Khaymonenko et al.
4.39
Layer-Wise Decision Fusion for Fake Audio Detection Using XLS-R
(2026)
Yixuan Xiao et al.
4.39
Cumsum-Composable Phase Transport for Low-Cost Streaming Keyword Spotting
(2026)
Mahesh Godavarti
4.39
Audio-Zero: Label-Free Self-Evolution for Fine-Grained Audio Reasoning
(2026)
Siqian Tong et al.
4.39
Pushing the Frontier of Full-Song Generation: Hierarchical Autoregressive Planning Meets Flow-Matching Rendering
(2026)
Junyu Dai et al.
4.39
Auto-AEG: Scalable Data Construction for Open-Vocabulary Audio Event Grounding
(2026)
Zihan Zhang et al.
2.00