HiLAM-aligned kernel discriminant analysis for text-dependent speaker verification. (15th November 2021)
- Record Type:
- Journal Article
- Title:
- HiLAM-aligned kernel discriminant analysis for text-dependent speaker verification. (15th November 2021)
- Main Title:
- HiLAM-aligned kernel discriminant analysis for text-dependent speaker verification
- Authors:
- Laskar, Mohammad Azharuddin
Laskar, Rabul Hussain - Abstract:
- Highlights: The main highlights of the work may be stated as follows. Kernel Discriminant Analysis for Text-dependent speaker verification task. A new speaker-text class is defined using Hierarchical multi-Layer Acoustic Model. Integration of new speaker-text class definition with i-vector and xvector systems. The systems have been validated on Part 1 of the RSR2015 database. Abstract: Probabilistic Linear Discriminant Analysis (PLDA) has been a commonly used backend classifier for many text-dependent speaker verification (TDSV) systems. Lately, PLDA projections have been integrated with the traditional Dynamic Time Warping (DTW) template matching framework, resulting in the DTW Online i-vector/PLDA system. The system is shown to achieve state-of-the-art performance for TDSV task. PLDA model serves to train a subspace that compensates for channel and session variabilities. It assumes linear separability between speaker-phrase information and other components. However, this relationship is known to be non-linear. The non-linearity is more prominent in case of short speech extracts, as in the case of the online i-vectors. This results in loss of vital speaker-phrase information at PLDA modeling. To this end, this work explores Kernel Discriminant Analysis (KDA) for TDSV task. It further proposes to use Hierarchical Multi-Layer Acoustic Model (HiLAM) to complement KDA with a more effective speaker-text class definition. The proposed system is hypothesized to benefit on threeHighlights: The main highlights of the work may be stated as follows. Kernel Discriminant Analysis for Text-dependent speaker verification task. A new speaker-text class is defined using Hierarchical multi-Layer Acoustic Model. Integration of new speaker-text class definition with i-vector and xvector systems. The systems have been validated on Part 1 of the RSR2015 database. Abstract: Probabilistic Linear Discriminant Analysis (PLDA) has been a commonly used backend classifier for many text-dependent speaker verification (TDSV) systems. Lately, PLDA projections have been integrated with the traditional Dynamic Time Warping (DTW) template matching framework, resulting in the DTW Online i-vector/PLDA system. The system is shown to achieve state-of-the-art performance for TDSV task. PLDA model serves to train a subspace that compensates for channel and session variabilities. It assumes linear separability between speaker-phrase information and other components. However, this relationship is known to be non-linear. The non-linearity is more prominent in case of short speech extracts, as in the case of the online i-vectors. This results in loss of vital speaker-phrase information at PLDA modeling. To this end, this work explores Kernel Discriminant Analysis (KDA) for TDSV task. It further proposes to use Hierarchical Multi-Layer Acoustic Model (HiLAM) to complement KDA with a more effective speaker-text class definition. The proposed system is hypothesized to benefit on three counts — non-linear modeling ability of KDA, speaker idiosyncrasy information associated with HiLAM-defined speaker-text units and modeling of the exact context of the pass-phrase, as offered by HiLAM. It shows a relative Equal Error Rate (EER) reduction of up to 50.63% on Part 1 of the RSR2015 database when compared to the baseline DTW Online i-vector/PLDA system. … (more)
- Is Part Of:
- Expert systems with applications. Volume 182(2021)
- Journal:
- Expert systems with applications
- Issue:
- Volume 182(2021)
- Issue Display:
- Volume 182, Issue 2021 (2021)
- Year:
- 2021
- Volume:
- 182
- Issue:
- 2021
- Issue Sort Value:
- 2021-0182-2021-0000
- Page Start:
- Page End:
- Publication Date:
- 2021-11-15
- Subjects:
- Text-dependent speaker verification -- Kernel Discriminant Analysis -- HiLAM -- Online i-vector/PLDA -- X-vector
Expert systems (Computer science) -- Periodicals
Systèmes experts (Informatique) -- Périodiques
Electronic journals
006.33 - Journal URLs:
- http://www.sciencedirect.com/science/journal/09574174 ↗
http://www.elsevier.com/journals ↗ - DOI:
- 10.1016/j.eswa.2021.115281 ↗
- Languages:
- English
- ISSNs:
- 0957-4174
- Deposit Type:
- Legaldeposit
- View Content:
- Available online (eLD content is only available in our Reading Rooms) ↗
- Physical Locations:
- British Library DSC - 3842.004220
British Library DSC - BLDSS-3PM
British Library HMNTS - ELD Digital store - Ingest File:
- 18482.xml