Speaker-adapted confidence measures for speech recognition of video lectures. (May 2016)
- Record Type:
- Journal Article
- Title:
- Speaker-adapted confidence measures for speech recognition of video lectures. (May 2016)
- Main Title:
- Speaker-adapted confidence measures for speech recognition of video lectures
- Authors:
- Sanchez-Cortina, Isaias
Andrés-Ferrer, Jesús
Sanchis, Alberto
Juan, Alfons - Abstract:
- Abstract : Highlights: A new, particular logistic regression model is proposed to improve confidence measures for automatic speech recognition. Speaker-adapted models are proposed to further improve confidence measures. Empirical results are provided showing that speaker-adapted models outperform their non-adapted counterparts. The improvement of confidence measures shown to be useful on an interactive speech transcription application. Abstract: Automatic speech recognition applications can benefit from a confidence measure (CM) to predict the reliability of the output. Previous works showed that a word-dependent naïve Bayes (NB) classifier outperforms the conventional word posterior probability as a CM. However, a discriminative formulation usually renders improved performance due to the available training techniques. Taking this into account, we propose a logistic regression (LR) classifier defined with simple input functions to approximate to the NB behaviour. Additionally, as a main contribution, we propose to adapt the CM to the speaker in cases in which it is possible to identify the speakers, such as online lecture repositories. The experiments have shown that speaker-adapted models outperform their non-adapted counterparts on two difficult tasks from English (videoLectures.net) and Spanish (poliMedia) educational lectures. They have also shown that the NB model is clearly superseded by the proposed LR classifier.
- Is Part Of:
- Computer speech & language. Volume 37(2016)
- Journal:
- Computer speech & language
- Issue:
- Volume 37(2016)
- Issue Display:
- Volume 37, Issue 2016 (2016)
- Year:
- 2016
- Volume:
- 37
- Issue:
- 2016
- Issue Sort Value:
- 2016-0037-2016-0000
- Page Start:
- 11
- Page End:
- 23
- Publication Date:
- 2016-05
- Subjects:
- Confidence measures -- Speech recognition -- Speaker adaptation -- Log-linear models -- Online video lectures
Speech processing systems -- Periodicals
Automatic speech recognition -- Periodicals
Computers -- Periodicals
Linguistics -- Periodicals
Speech-Language Pathology -- Periodicals
Traitement automatique de la parole -- Périodiques
Reconnaissance automatique de la parole -- Périodiques
Automatic speech recognition
Speech processing systems
Electronic journals
Periodicals
006.454 - Journal URLs:
- http://www.journals.elsevier.com/computer-speech-and-language/ ↗
http://www.elsevier.com/journals ↗ - DOI:
- 10.1016/j.csl.2015.10.003 ↗
- Languages:
- English
- ISSNs:
- 0885-2308
- Deposit Type:
- Legaldeposit
- View Content:
- Available online (eLD content is only available in our Reading Rooms) ↗
- Physical Locations:
- British Library DSC - 3394.276600
British Library DSC - BLDSS-3PM
British Library HMNTS - ELD Digital store - Ingest File:
- 1216.xml