A large reproducible benchmark of ontology-based methods and word embeddings for word similarity. Issue 96 (February 2021)
- Record Type:
- Journal Article
- Title:
- A large reproducible benchmark of ontology-based methods and word embeddings for word similarity. Issue 96 (February 2021)
- Main Title:
- A large reproducible benchmark of ontology-based methods and word embeddings for word similarity
- Authors:
- Lastra-Díaz, Juan J.
Goikoetxea, Josu
Hadj Taieb, Mohamed Ali
Garcia-Serrano, Ana
Ben Aouicha, Mohamed
Agirre, Eneko
Sánchez, David - Abstract:
- Abstract: This work is a companion reproducibility paper of the experiments and results reported in Lastra-Diaz et al. (2019a), which is based on the evaluation of a companion reproducibility dataset with the HESML V1R4 library and the long-term reproducibility tool called Reprozip. Human similarity and relatedness judgements between concepts underlie most of cognitive capabilities, such as categorization, memory, decision-making and reasoning. For this reason, the research on methods for the estimation of the degree of similarity and relatedness between words and concepts has received a lot of attention in the fields of artificial intelligence and cognitive sciences. However, despite the huge research effort done, there is a lack of a self-contained, reproducible and extensible collection of benchmarks which being amenable to become a de facto standard for large scale experimentation in this line of research. In order to bridge this reproducibility gap, this work introduces a set of reproducible experiments on word similarity and relatedness by providing a detailed reproducibility protocol together with a set of software tools and a self-contained reproducibility dataset, which allow that all experiments and results in our aforementioned work to be reproduced exactly. Our aforementioned primary work introduces the largest, most detailed and reproducible experimental survey on word similarity and relatedness reported in the literature, which is based on the implementation ofAbstract: This work is a companion reproducibility paper of the experiments and results reported in Lastra-Diaz et al. (2019a), which is based on the evaluation of a companion reproducibility dataset with the HESML V1R4 library and the long-term reproducibility tool called Reprozip. Human similarity and relatedness judgements between concepts underlie most of cognitive capabilities, such as categorization, memory, decision-making and reasoning. For this reason, the research on methods for the estimation of the degree of similarity and relatedness between words and concepts has received a lot of attention in the fields of artificial intelligence and cognitive sciences. However, despite the huge research effort done, there is a lack of a self-contained, reproducible and extensible collection of benchmarks which being amenable to become a de facto standard for large scale experimentation in this line of research. In order to bridge this reproducibility gap, this work introduces a set of reproducible experiments on word similarity and relatedness by providing a detailed reproducibility protocol together with a set of software tools and a self-contained reproducibility dataset, which allow that all experiments and results in our aforementioned work to be reproduced exactly. Our aforementioned primary work introduces the largest, most detailed and reproducible experimental survey on word similarity and relatedness reported in the literature, which is based on the implementation of all evaluated methods into the same software platform. Our reproducible experiments evaluate most of methods in the families of ontology-based semantic similarity measures and word embedding models. We also detail how to extend our experiments to evaluate other unconsidered experimental setups. Finally, we provide a corrigendum for a mismatch in the MC28 similarity scores used in our original experiments. Highlights: A reproducible benchmark of ontology-based similarity measures and word embeddings. The largest known set of reproducible experiments on word similarity and relatedness. A detailed reproducibility protocol based on the HESML library and a self-contained dataset. Our experiment file can be used as a template to evaluate other experimental setups. Providing a long-time reproducibility protocol based on the Reprozip tool. … (more)
- Is Part Of:
- Information systems. Issue 96(2021)
- Journal:
- Information systems
- Issue:
- Issue 96(2021)
- Issue Display:
- Volume 96, Issue 96 (2021)
- Year:
- 2021
- Volume:
- 96
- Issue:
- 96
- Issue Sort Value:
- 2021-0096-0096-0000
- Page Start:
- Page End:
- Publication Date:
- 2021-02
- Subjects:
- Ontology-based semantic similarity measures -- Word embeddings -- Information Content models -- Reproducible benchmark -- HESML -- Reprozip
Database management -- Periodicals
Electronic data processing -- Periodicals
Bases de données -- Gestion -- Périodiques
Informatique -- Périodiques
Database management
Electronic data processing
Periodicals
005.7 - Journal URLs:
- http://www.sciencedirect.com/science/journal/03064379 ↗
http://www.elsevier.com/journals ↗ - DOI:
- 10.1016/j.is.2020.101636 ↗
- Languages:
- English
- ISSNs:
- 0306-4379
- Deposit Type:
- Legaldeposit
- View Content:
- Available online (eLD content is only available in our Reading Rooms) ↗
- Physical Locations:
- British Library DSC - 4496.367300
British Library DSC - BLDSS-3PM
British Library HMNTS - ELD Digital store - Ingest File:
- 15001.xml