MFF-SAug: Multi feature fusion with spectrogram augmentation of speech emotion recognition using convolution neural network. (September 2022)
- Record Type:
- Journal Article
- Title:
- MFF-SAug: Multi feature fusion with spectrogram augmentation of speech emotion recognition using convolution neural network. (September 2022)
- Main Title:
- MFF-SAug: Multi feature fusion with spectrogram augmentation of speech emotion recognition using convolution neural network
- Authors:
- Jothimani, S.
Premalatha, K. - Abstract:
- Abstract: The Speech Emotion Recognition (SER) is a complex task because of the feature selections that reflect the emotion from the human speech. The SER plays a vital role and is very challenging in Human-Computer Interaction (HCI). Traditional methods provide inconsistent feature extraction for emotion recognition. The primary motive of this paper is to improve the accuracy of the classification of eight emotions from the human voice. The proposed MFF-SAug research, Enhance the emotion prediction from the speech by Noise Removal, White Noise Injection, and Pitch Tuning. On pre-processed speech signals, the feature extraction techniques Mel Frequency Cepstral Coefficients (MFCC), Zero Crossing Rate (ZCR), and Root Mean Square (RMS) are applied and combined to achieve substantial performance used for emotion recognition. The augmentation applies to the raw speech for a contrastive loss that maximizes agreement between differently augmented samples in the latent space and reconstructs the loss of input representation for better accuracy prediction. A state-of-the-art Convolution Neural Network (CNN) is proposed for enhanced speech representation learning and voice emotion classification. Further, this MFF-SAug method is compared with the CNN + LSTM model. The experimental analysis was carried out using the RAVDESS, CREMA, SAVEE, and TESS datasets. Thus, the classifier achieved a robust representation for speech emotion recognition with an accuracy of 92.6 %, 89.9, 84.9 %,Abstract: The Speech Emotion Recognition (SER) is a complex task because of the feature selections that reflect the emotion from the human speech. The SER plays a vital role and is very challenging in Human-Computer Interaction (HCI). Traditional methods provide inconsistent feature extraction for emotion recognition. The primary motive of this paper is to improve the accuracy of the classification of eight emotions from the human voice. The proposed MFF-SAug research, Enhance the emotion prediction from the speech by Noise Removal, White Noise Injection, and Pitch Tuning. On pre-processed speech signals, the feature extraction techniques Mel Frequency Cepstral Coefficients (MFCC), Zero Crossing Rate (ZCR), and Root Mean Square (RMS) are applied and combined to achieve substantial performance used for emotion recognition. The augmentation applies to the raw speech for a contrastive loss that maximizes agreement between differently augmented samples in the latent space and reconstructs the loss of input representation for better accuracy prediction. A state-of-the-art Convolution Neural Network (CNN) is proposed for enhanced speech representation learning and voice emotion classification. Further, this MFF-SAug method is compared with the CNN + LSTM model. The experimental analysis was carried out using the RAVDESS, CREMA, SAVEE, and TESS datasets. Thus, the classifier achieved a robust representation for speech emotion recognition with an accuracy of 92.6 %, 89.9, 84.9 %, and 99.6 % for RAVDESS, CREMA, SAVEE, and TESS datasets, respectively. Highlights: Speech emotion recognition Feature fusion – time and frequency domain feature extraction Augmentation techniques Convolution Neural Networks Improved accuracy of emotion classification … (more)
- Is Part Of:
- Chaos, solitons and fractals. Volume 162(2022)
- Journal:
- Chaos, solitons and fractals
- Issue:
- Volume 162(2022)
- Issue Display:
- Volume 162, Issue 2022 (2022)
- Year:
- 2022
- Volume:
- 162
- Issue:
- 2022
- Issue Sort Value:
- 2022-0162-2022-0000
- Page Start:
- Page End:
- Publication Date:
- 2022-09
- Subjects:
- Augmentation -- Contrastive loss -- MFCC -- RMS -- Speech emotion recognition -- ZCR
Chaotic behavior in systems -- Periodicals
Solitons -- Periodicals
Fractals -- Periodicals
Chaotic behavior in systems
Fractals
Solitons
Periodicals
003.7 - Journal URLs:
- http://www.elsevier.com/journals ↗
http://www.sciencedirect.com/science/journal/09600779 ↗ - DOI:
- 10.1016/j.chaos.2022.112512 ↗
- Languages:
- English
- ISSNs:
- 0960-0779
- Deposit Type:
- Legaldeposit
- View Content:
- Available online (eLD content is only available in our Reading Rooms) ↗
- Physical Locations:
- British Library DSC - 3129.716000
British Library DSC - BLDSS-3PM
British Library HMNTS - ELD Digital store - Ingest File:
- 23288.xml