Building a morpho-semantic knowledge graph for Arabic information retrieval. Issue 6 (November 2020)
- Record Type:
- Journal Article
- Title:
- Building a morpho-semantic knowledge graph for Arabic information retrieval. Issue 6 (November 2020)
- Main Title:
- Building a morpho-semantic knowledge graph for Arabic information retrieval
- Authors:
- Bounhas, Ibrahim
Soudani, Nadia
Slimani, Yahya - Abstract:
- Highlights: A morpho-semantic knowledge graph CAMS-KG is built from vocalized Classical Arabic corpus. CAMS-KG combines tools for morphological analysis and disambiguation, and implements a concordance builder tool, and KG representation. KG stores the extracted morpho-semantic knowledge: representing morphological categories and both morphological and semantic relations. BM25 ranking is used for retrieving related documents for a given query. CAMS-KG is evaluated on two datasets (Tashkeela, and ZAD). Several query expansion strategies are experimented on 25 queries from ZAD dataset. Abstract: In this paper, we propose to build a morpho-semantic knowledge graph from Arabic vocalized corpora. Our work focuses on classical Arabic as it has not been deeply investigated in related works. We use a tool suite which allows analyzing and disambiguating Arabic texts, taking into account short diacritics to reduce ambiguities. At the morphological level, we combine Ghwanmeh stemmer and MADAMIRA which are adapted to extract a multi-level lexicon from Arabic vocalized corpora. At the semantic level, we infer semantic dependencies between tokens by exploiting contextual knowledge extracted by a concordancer. Both morphological and semantic links are represented through compressed graphs, which are accessed through lazy methods. These graphs are mined using a measure inspired from BM25 to compute one-to-many similarity. Indeed, we propose to evaluate the morpho-semantic Knowledge Graph inHighlights: A morpho-semantic knowledge graph CAMS-KG is built from vocalized Classical Arabic corpus. CAMS-KG combines tools for morphological analysis and disambiguation, and implements a concordance builder tool, and KG representation. KG stores the extracted morpho-semantic knowledge: representing morphological categories and both morphological and semantic relations. BM25 ranking is used for retrieving related documents for a given query. CAMS-KG is evaluated on two datasets (Tashkeela, and ZAD). Several query expansion strategies are experimented on 25 queries from ZAD dataset. Abstract: In this paper, we propose to build a morpho-semantic knowledge graph from Arabic vocalized corpora. Our work focuses on classical Arabic as it has not been deeply investigated in related works. We use a tool suite which allows analyzing and disambiguating Arabic texts, taking into account short diacritics to reduce ambiguities. At the morphological level, we combine Ghwanmeh stemmer and MADAMIRA which are adapted to extract a multi-level lexicon from Arabic vocalized corpora. At the semantic level, we infer semantic dependencies between tokens by exploiting contextual knowledge extracted by a concordancer. Both morphological and semantic links are represented through compressed graphs, which are accessed through lazy methods. These graphs are mined using a measure inspired from BM25 to compute one-to-many similarity. Indeed, we propose to evaluate the morpho-semantic Knowledge Graph in the context of Arabic Information Retrieval (IR). Several scenarios of document indexing and query expansion are assessed. That is, we vary indexing units for Arabic IR based on different levels of morphological knowledge, a challenging issue which is not yet resolved in previous research. We also experiment several combinations of morpho-semantic query expansion. This permits to validate our resource and to study its impact on IR based on state-of-the art evaluation metrics. … (more)
- Is Part Of:
- Information processing & management. Volume 57:Issue 6(2020:Nov.)
- Journal:
- Information processing & management
- Issue:
- Volume 57:Issue 6(2020:Nov.)
- Issue Display:
- Volume 57, Issue 6 (2020)
- Year:
- 2020
- Volume:
- 57
- Issue:
- 6
- Issue Sort Value:
- 2020-0057-0006-0000
- Page Start:
- Page End:
- Publication Date:
- 2020-11
- Subjects:
- Morpho-semantic knowledge extraction -- Classical Arabic text mining -- Arabic information retrieval -- Graph-based knowledge representation
Information storage and retrieval systems -- Periodicals
Information science -- Periodicals
Systèmes d'information -- Périodiques
Sciences de l'information -- Périodiques
Information science
Information storage and retrieval systems
Periodicals
658.4038 - Journal URLs:
- http://www.sciencedirect.com/science/journal/03064573 ↗
http://www.elsevier.com/journals ↗ - DOI:
- 10.1016/j.ipm.2019.102124 ↗
- Languages:
- English
- ISSNs:
- 0306-4573
- Deposit Type:
- Legaldeposit
- View Content:
- Available online (eLD content is only available in our Reading Rooms) ↗
- Physical Locations:
- British Library DSC - 4493.893000
British Library DSC - BLDSS-3PM
British Library HMNTS - ELD Digital store - Ingest File:
- 14754.xml