Convolution-based Channel-frequency Attention For Text-independent Speaker Verification
2022 Β· Jingyu Li, Yusheng Tian, Tan Lee
Abstract
Deep convolutional neural networks (CNNs) have been applied to extracting speaker embeddings with significant success in speaker verification. Incorporating the attention mechanism has shown to be effective in improving the model performance. This paper presents an efficient two-dimensional convolution-based attention module, namely C2D-Att. The interaction between the convolution channel and frequency is involved in the attention calculation by lightweight convolution layers. This requires only a small number of parameters. Fine-grained attention weights are produced to represent channel and frequency-specific information. The weights are imposed on the input features to improve the representation ability for speaker modeling. The C2D-Att is integrated into a modified version of ResNet for speaker embedding extraction. Experiments are conducted on VoxCeleb datasets. The results show that C2DAtt is effective in generating discriminative attention maps and outperforms other attention me
Authors
(none)
Tags
Stats
Related papers
- Frequency And Temporal Convolutional Attention For Text-independent Speaker Recognition (2019)0.00
- Duality Temporal-channel-frequency Attention Enhanced Speaker Representation Learning (2021)5.24
- Multi-frequency Information Enhanced Channel Attention Module For Speaker Representation Learning (2022)0.00
- End-to-end Attention Based Text-dependent Speaker Verification (2017)14.87
- ECAPA-TDNN: Emphasized Channel Attention, Propagation And Aggregation In TDNN Based Speaker Verification (2020)23.07
- Dynamic Kernels And Channel Attention For Low Resource Speaker Verification (2022)0.00
- MFA: TDNN With Multi-scale Frequency-channel Attention For Text-independent Speaker Verification With Short Utterances (2022)13.79
- Double Multi-head Attention For Speaker Verification (2020)8.09