Real-time single-channel speech enhancement based on causal attention mechanism. (December 2022)
- Record Type:
- Journal Article
- Title:
- Real-time single-channel speech enhancement based on causal attention mechanism. (December 2022)
- Main Title:
- Real-time single-channel speech enhancement based on causal attention mechanism
- Authors:
- Fan, Junyi
Yang, Jibin
Zhang, Xiongwei
Yao, Yao - Abstract:
- Highlights: A real-time speech enhancement model based on causal attention mechanism is proposed. Attention weights masked by the upper triangle matrix enable causal attention. Single-side relative position representation enables attention to local information. Combined time and frequency domain loss function guides the training of models. Experimental results show the superiority of the model over other existing models. Abstract: To achieve real-time single-channel speech enhancement, i.e., enhancing with no or low latency, this paper proposes a causal speech enhancement model with an attention mechanism based on Transformer. The model uses a causal codec with a U- net -like structure as the backbone network, which is improved with an upper triangle mask matrix and a single-side relative position representation on the basis of ensuring the causality. The mask matrix preserves the attentional focus on the historical global information and the single-side relative position representation focuses more on the information that needs attention in the local information. In addition, the weighted loss function in both time and frequency domains is used to guide the optimization direction of the training. Exhaustive comparison experiments are conducted on the Voice-Bank Demand dataset, and the experimental results show that the proposed causal model, compared with existing real-time single-channel speech enhancement models, not only possesses better enhancement results but also hasHighlights: A real-time speech enhancement model based on causal attention mechanism is proposed. Attention weights masked by the upper triangle matrix enable causal attention. Single-side relative position representation enables attention to local information. Combined time and frequency domain loss function guides the training of models. Experimental results show the superiority of the model over other existing models. Abstract: To achieve real-time single-channel speech enhancement, i.e., enhancing with no or low latency, this paper proposes a causal speech enhancement model with an attention mechanism based on Transformer. The model uses a causal codec with a U- net -like structure as the backbone network, which is improved with an upper triangle mask matrix and a single-side relative position representation on the basis of ensuring the causality. The mask matrix preserves the attentional focus on the historical global information and the single-side relative position representation focuses more on the information that needs attention in the local information. In addition, the weighted loss function in both time and frequency domains is used to guide the optimization direction of the training. Exhaustive comparison experiments are conducted on the Voice-Bank Demand dataset, and the experimental results show that the proposed causal model, compared with existing real-time single-channel speech enhancement models, not only possesses better enhancement results but also has faster training speed and fewer trainable parameters. … (more)
- Is Part Of:
- Applied acoustics. Volume 201(2022)
- Journal:
- Applied acoustics
- Issue:
- Volume 201(2022)
- Issue Display:
- Volume 201, Issue 2022 (2022)
- Year:
- 2022
- Volume:
- 201
- Issue:
- 2022
- Issue Sort Value:
- 2022-0201-2022-0000
- Page Start:
- Page End:
- Publication Date:
- 2022-12
- Subjects:
- Attention mechanism -- Causality -- Single-channel -- Single-side relative position representation -- Speech enhancement
DNN Deep Neural Network -- CRN Convolutional Recurrent Network -- LSTM Long Short-Term Memory -- TCM Temporal Convolution Module -- CNN Convolutional Neural Network -- DCN Dilated Convolutional Network -- MulCA Multi-scale time sensitive Channel Attention -- SRPR Single-side Relative Position Representation -- MSE Mean Square Error -- MAE Mean Absolute Error -- CMAM Causal Multi-Head Attention Mechanism -- MAM Multi-Head Attention Mechanism -- CSAM Causal Synthetic Attention Mechanism -- SAM Synthetic Attention Mechanism -- RPR Relative Position Representation -- CMAM_SP Causal Multi-Head Attention Mechanism with SRPR -- CSAM_SP Causal Synthetic Attention Mechanism with SRPR -- DFT Discrete Fourier Transform -- SNR Signal-to-Noise Ratios -- PESQ Perceptual Evaluation of Speech Quality -- STOI Short-Time Objective Intelligibility -- LSD Log Spectral Distance -- RTF Real-Time Factor
Acoustical engineering -- Periodicals
Periodicals
620.2 - Journal URLs:
- http://www.sciencedirect.com/science/journal/0003682X ↗
http://www.elsevier.com/journals ↗
http://www.elsevier.com/homepage/elecserv.htt ↗ - DOI:
- 10.1016/j.apacoust.2022.109084 ↗
- Languages:
- English
- ISSNs:
- 0003-682X
- Deposit Type:
- Legaldeposit
- View Content:
- Available online (eLD content is only available in our Reading Rooms) ↗
- Physical Locations:
- British Library DSC - 1571.400000
British Library DSC - BLDSS-3PM
British Library HMNTS - ELD Digital store - Ingest File:
- 24456.xml