Using comparable corpora for under-resourced areas of machine translation. (2019)
- Record Type:
- Book
- Title:
- Using comparable corpora for under-resourced areas of machine translation. (2019)
- Main Title:
- Using comparable corpora for under-resourced areas of machine translation
- Further Information:
- Note: Inguna Skadiņa, Robert Gaizauskas, Bogdan Babych, Nikola Ljubešić, Dan Tufiş, Andrejs Vasiļjevs, editors.
- Editors:
- Skadina, Inguna
Gaizauskas, Robert
Babych, Bogdan
Ljubešić, Nikola
Tufiș, Dan
Vasiļjevs, Andrejs - Contents:
- Intro; Contents; Chapter 1: Introduction; 1.1 Parallel Data; 1.2 Comparable Corpora and Comparability; 1.3 Acquisition of Parallel Data from Comparable Corpora; 1.4 Comparable Corpora in Machine Translation; 1.5 Summary of the Book; 1.6 The ACCURAT Project; References; Chapter 2: Cross-Language Comparability and Its Applications for MT; 2.1 Introduction: Definition and Use of the Concept of Comparability; 2.2 Development and Calibration of Comparability Metrics on Parallel Corpora; 2.2.1 Application of Corpus Comparability: Selecting Coherent Parallel Corpora for Domain-Specific MT Training 2.2.2 Methodology2.2.2.1 Description of Calculation Method; 2.2.2.2 Symmetric vs. Asymmetric Calculation of Distance; 2.2.2.3 Calibrating the Distance Metric; 2.2.3 Validation of the Scores: Cross-Language Agreement for Source vs. Target Sides of TMX Files; 2.2.4 Discussion; 2.3 Exploration of Comparability Features in Document-Aligned Comparable Corpora: Wikipedia; 2.3.1 Overview: Wikipedia as a Source of Comparable Corpora; 2.3.2 Previous Work on Using Wikipedia as a Linguistic Resource; 2.3.3 Methodology; 2.3.3.1 Document Pre-processing; 2.3.3.2 Similarity Measures 2.3.3.3 Eliciting Human Judgements2.3.4 Results and Analysis; 2.3.4.1 Responses to the Questionnaire; 2.3.4.2 Inter-assessor Agreement; 2.3.4.3 Correlation of Similarity Measures to Human Judgements; 2.3.4.4 Classification Task; 2.3.5 Discussion; 2.3.5.1 Features of `Similar ́Articles; 2.3.5.2 Measuring Cross-LanguageIntro; Contents; Chapter 1: Introduction; 1.1 Parallel Data; 1.2 Comparable Corpora and Comparability; 1.3 Acquisition of Parallel Data from Comparable Corpora; 1.4 Comparable Corpora in Machine Translation; 1.5 Summary of the Book; 1.6 The ACCURAT Project; References; Chapter 2: Cross-Language Comparability and Its Applications for MT; 2.1 Introduction: Definition and Use of the Concept of Comparability; 2.2 Development and Calibration of Comparability Metrics on Parallel Corpora; 2.2.1 Application of Corpus Comparability: Selecting Coherent Parallel Corpora for Domain-Specific MT Training 2.2.2 Methodology2.2.2.1 Description of Calculation Method; 2.2.2.2 Symmetric vs. Asymmetric Calculation of Distance; 2.2.2.3 Calibrating the Distance Metric; 2.2.3 Validation of the Scores: Cross-Language Agreement for Source vs. Target Sides of TMX Files; 2.2.4 Discussion; 2.3 Exploration of Comparability Features in Document-Aligned Comparable Corpora: Wikipedia; 2.3.1 Overview: Wikipedia as a Source of Comparable Corpora; 2.3.2 Previous Work on Using Wikipedia as a Linguistic Resource; 2.3.3 Methodology; 2.3.3.1 Document Pre-processing; 2.3.3.2 Similarity Measures 2.3.3.3 Eliciting Human Judgements2.3.4 Results and Analysis; 2.3.4.1 Responses to the Questionnaire; 2.3.4.2 Inter-assessor Agreement; 2.3.4.3 Correlation of Similarity Measures to Human Judgements; 2.3.4.4 Classification Task; 2.3.5 Discussion; 2.3.5.1 Features of `Similar ́Articles; 2.3.5.2 Measuring Cross-Language Similarity; 2.3.6 Section Conclusions; 2.4 Metrics for Identifying Comparability Levels in Non-aligned Documents; 2.4.1 Using Parallel and Comparable Corpora for MT; 2.4.2 Related Work; 2.4.3 Comparability Metrics; 2.4.3.1 Lexical Mapping Based Metric 2.4.3.2 Keyword-Based Metric2.4.3.3 Machine Translation (MT)-Based Metrics; 2.4.4 Experiments and Evaluation; 2.4.4.1 Data Sources; 2.4.4.2 Experimental Results; 2.4.5 Metric Application to Equivalent Extraction; 2.4.6 Discussion; 2.4.6.1 Advantages and Disadvantages of the Metrics; 2.4.6.2 Using Semi-parallel Equivalents in MT Systems; 2.4.7 Conclusion; References; Chapter 3: Collecting Comparable Corpora; 3.1 Introduction; 3.2 Previous Work in Collecting Comparable Corpora; 3.2.1 Web Crawling; 3.2.2 Identifying Comparable Text; 3.3 ACCURAT Techniques to Collect Comparable Documents 3.3.1 Comparable Corpora Collection from Wikipedia3.3.1.1 Extracting Comparable Articles; 3.3.1.2 Measuring Similarity in Inter-language Linked Documents; 3.3.2 Comparable Corpora Collection from News Articles; 3.3.3 Comparable Corpora Collection from Narrow Domains; 3.3.3.1 Acquiring Comparable Documents; 3.3.3.2 Aligning Comparable Document Pairs; References; Chapter 4: Extracting Data from Comparable Corpora; 4.1 Introduction; 4.2 Term Extraction, Tagging and Mapping for Under-Resourced Languages; 4.2.1 Related Work; 4.2.2 Term Extraction, Tagging and Mapping with the ACCURAT Toolkit … (more)
- Publisher Details:
- Cham, Switzerland : Springer
- Publication Date:
- 2019
- Extent:
- 1 online resource (vi, 323 pages), illustrations (some color)
- Subjects:
- 418.1/88
Corpora (Linguistics)
Machine translating
Electronic books
Electronic books - Languages:
- English
- ISBNs:
- 9783319990040
3319990047 - Related ISBNs:
- 9783319990033
- Notes:
- Note: Online resource; title from PDF title page (SpringerLink, viewed February 14, 2019).
- Access Rights:
- Legal Deposit; Only available on premises controlled by the deposit library and to one user at any one time; The Legal Deposit Libraries (Non-Print Works) Regulations (UK).
- Access Usage:
- Restricted: Printing from this resource is governed by The Legal Deposit Libraries (Non-Print Works) Regulations (UK) and UK copyright law currently in force.
- View Content:
- Available online (eLD content is only available in our Reading Rooms) ↗
- Physical Locations:
- British Library HMNTS - ELD.DS.386243
- Ingest File:
- 02_375.xml