SUPERB
Emerging52papers using it
2021first seen
SUPERB is a benchmark dataset that contains a variety of speech processing tasks and is used to evaluate the performance of speech foundation models.
Papers using SUPERB (52)
- Rethinking Speech Foundation Model Fine-tuning: Better SFT or Better Match?Fast Speech Foundation Model Distillation Using Interleaved StackingCodec2Vec: Self-Supervised Speech Representation Learning Using Neural Speech CodecsWavCube: Unifying Speech Representation for Understanding and Generation via Semantic-Acoustic Joint ModelingSPEAR: A Unified SSL Framework for Learning Speech and Audio RepresentationsAn Exploration of Mamba for Speech Self-Supervised ModelsUSAD: Universal Speech and Audio Representation via DistillationTask-Agnostic Structured Pruning of Speech Representation ModelsFindings of the 2023 ML-SUPERB Challenge: Pre-Training and Evaluation
over More Languages and BeyondSelf-Supervised Representation Learning for Speech Using Visual
Grounding and Masked Language ModelingEnsemble knowledge distillation of self-supervised speech modelsSpeechLM: Enhanced Speech Pre-Training with Unpaired Textual DataWhat Do Self-Supervised Speech and Speaker Models Learn? New Findings
From a Cross Model Layer-Wise AnalysisUniSpeech-SAT: Universal Speech Representation Learning with Speaker
Aware Pre-TrainingAn Exploration of Self-Supervised Pretrained Representations for
End-to-End Speech RecognitionFitHuBERT: Going Thinner and Deeper for Knowledge Distillation of Speech
Self-Supervised LearningEfficient Speech Representation Learning with Low-Bit QuantizationDon't speak too fast: The impact of data bias on self-supervised speech
modelsSUPERB-SG: Enhanced Speech processing Universal PERformance Benchmark
for Semantic and Generative CapabilitiesImproving the Robustness of DistilHuBERT to Unseen Noisy Conditions via
Data Augmentation, Curriculum Learning, and Multi-Task EnhancementSCORE: Self-supervised Correspondence Fine-tuning for Improved Content
RepresentationsCocktail HuBERT: Generalized Self-Supervised Pre-training for Mixture
and Single-Source SpeechSelf-supervised Neural Factor Analysis for Disentangling Utterance-level
Speech RepresentationsImproving Distortion Robustness of Self-supervised Speech Processing
Tasks with Domain AdaptationDeep versus Wide: An Analysis of Student Architectures for Task-Agnostic
Knowledge Distillation of Self-Supervised Speech ModelsEvaluating context-invariance in unsupervised speech representationsML-SUPERB: Multilingual Speech Universal PERformance BenchmarkRecycle-and-Distill: Universal Compression Strategy for
Transformer-based Speech SSL Models with Attention Map Reusing and Masking
DistillationOn the Transferability of Whisper-based Representations for
"In-the-Wild" Cross-Task Downstream Speech ApplicationsDPHuBERT: Joint Distillation and Pruning of Self-Supervised Speech
ModelsAre Paralinguistic Representations all that is needed for Speech Emotion
Recognition?A Large-Scale Evaluation of Speech Foundation ModelsLightHuBERT: Lightweight and Configurable Speech Representation Learning
with Once-for-All Hidden-Unit BERTCoBERT: Self-Supervised Speech Representation Learning Through Code
Representation LearningSUPERB @ SLT 2022: Challenge on Generalization and Efficiency of
Self-Supervised Speech Representation LearningExploring Effective Fusion Algorithms for Speech Based Self-Supervised
Learning ModelsMasked Modeling Duo for Speech: Specializing General-Purpose Audio
Representation to Speech using Denoising DistillationMCR-Data2vec 2.0: Improving Self-supervised Speech Pre-training via
Model-level Consistency RegularizationSelective HuBERT: Self-Supervised Pre-Training for Target Speaker in
Clean and Mixture SpeechAn Experimental Study: Assessing the Combined Framework of WavLM and
BEST-RQ for Text-to-Speech SynthesisSTaR: Distilling Speech Temporal Relation for Lightweight Speech
Self-Supervised Learning ModelsCan you Remove the Downstream Model for Speaker Recognition with
Self-Supervised Speech Features?SKILL: Similarity-aware Knowledge distILLation for Speech
Self-Supervised LearningMulti-Stage Multi-Modal Pre-Training for Automatic Speech RecognitionRemoving Speaker Information from Speech Representation using
Variable-Length Soft PoolingRefining Self-Supervised Learnt Speech Representation using Brain
ActivationsLASER: Learning by Aligning Self-supervised Representations of Speech
for Improving Content-related TasksGenDistiller: Distilling Pre-trained Language Models based on an
Autoregressive Generative ModelJOOCI: a Framework for Learning Comprehensive Speech RepresentationsEH-MAM: Easy-to-Hard Masked Acoustic Modeling for Self-Supervised Speech
Representation LearningWavLM: Large-Scale Self-Supervised Pre-Training for Full Stack Speech
ProcessingAn Adapter-Based Unified Model for Multiple Spoken Language Processing
Tasks