Weighted finite-state transducers for normalization of historical texts. (March 2019)
- Record Type:
- Journal Article
- Title:
- Weighted finite-state transducers for normalization of historical texts. (March 2019)
- Main Title:
- Weighted finite-state transducers for normalization of historical texts
- Authors:
- Etxeberria, Izaskun
Alegria, Iñaki
Uria, Larraitz - Abstract:
- Abstract: This paper presents a study about methods for normalization of historical texts. The aim of these methods is learning relations between historical and contemporary word forms. We have compiled training and test corpora for different languages and scenarios, and we have tried to read the results related to the features of the corpora and languages. Our proposed method, based on weighted finite-state transducers, is compared to previously published ones. Our method learns to map phonological changes using a noisy channel model; it is a simple solution that can use a limited amount of supervision in order to achieve adequate performance. The compiled corpora are ready to be used for other researchers in order to compare results. Concerning the amount of supervision for the task, we investigate how the size of training corpus affects the results and identify some interesting factors to anticipate the difficulty of the task.
- Is Part Of:
- Natural language engineering. Volume 25:Part 2(2019)
- Journal:
- Natural language engineering
- Issue:
- Volume 25:Part 2(2019)
- Issue Display:
- Volume 25, Issue 2, Part 2 (2019)
- Year:
- 2019
- Volume:
- 25
- Issue:
- 2
- Part:
- 2
- Issue Sort Value:
- 2019-0025-0002-0002
- Page Start:
- 307
- Page End:
- 321
- Publication Date:
- 2019-03
- Subjects:
- lexical normalization, -- historical texts, -- computational morphology, -- phonology
Natural language processing (Computer science) -- Periodicals
Software engineering -- Periodicals
006.35 - Journal URLs:
- http://journals.cambridge.org/action/displayJournal?jid=NLE ↗
- DOI:
- 10.1017/S1351324918000505 ↗
- Languages:
- English
- ISSNs:
- 1351-3249
- Deposit Type:
- Legaldeposit
- View Content:
- Available online (eLD content is only available in our Reading Rooms) ↗
- Physical Locations:
- British Library HMNTS - ELD Digital store
- Ingest File:
- 13002.xml