Indexing Arabic texts using association rule data mining. Issue 1 (20th June 2018)
- Record Type:
- Journal Article
- Title:
- Indexing Arabic texts using association rule data mining. Issue 1 (20th June 2018)
- Main Title:
- Indexing Arabic texts using association rule data mining
- Authors:
- Haraty, Ramzi A.
Nasrallah, Rouba - Abstract:
- Abstract : Purpose: The purpose of this paper is to propose a new model to enhance auto-indexing Arabic texts. The model denotes extracting new relevant words by relating those chosen by previous classical methods to new words using data mining rules. Design/methodology/approach: The proposed model uses an association rule algorithm for extracting frequent sets containing related items – to extract relationships between words in the texts to be indexed with words from texts that belong to the same category. The associations of words extracted are illustrated as sets of words that appear frequently together. Findings: The proposed methodology shows significant enhancement in terms of accuracy, efficiency and reliability when compared to previous works. Research limitations/implications: The stemming algorithm can be further enhanced. In the Arabic language, we have many grammatical rules. The more we integrate rules to the stemming algorithm, the better the stemming will be. Other enhancements can be done to the stop-list. This is by adding more words to it that should not be taken into consideration in the indexing mechanism. Also, numbers should be added to the list as well as using the thesaurus system because it links different phrases or words with the same meaning to each other, which improves the indexing mechanism. The authors also invite researchers to add more pre-requisite texts to have better results. Originality/value: In this paper, the authors present a fullAbstract : Purpose: The purpose of this paper is to propose a new model to enhance auto-indexing Arabic texts. The model denotes extracting new relevant words by relating those chosen by previous classical methods to new words using data mining rules. Design/methodology/approach: The proposed model uses an association rule algorithm for extracting frequent sets containing related items – to extract relationships between words in the texts to be indexed with words from texts that belong to the same category. The associations of words extracted are illustrated as sets of words that appear frequently together. Findings: The proposed methodology shows significant enhancement in terms of accuracy, efficiency and reliability when compared to previous works. Research limitations/implications: The stemming algorithm can be further enhanced. In the Arabic language, we have many grammatical rules. The more we integrate rules to the stemming algorithm, the better the stemming will be. Other enhancements can be done to the stop-list. This is by adding more words to it that should not be taken into consideration in the indexing mechanism. Also, numbers should be added to the list as well as using the thesaurus system because it links different phrases or words with the same meaning to each other, which improves the indexing mechanism. The authors also invite researchers to add more pre-requisite texts to have better results. Originality/value: In this paper, the authors present a full text-based auto-indexing method for Arabic text documents. The auto-indexing method extracts new relevant words by using data mining rules, which has not been investigated before. The method uses an association rule mining algorithm for extracting frequent sets containing related items to extract relationships between words in the texts to be indexed with words from texts that belong to the same category. The benefits of the method are demonstrated using empirical work involving several Arabic texts. … (more)
- Is Part Of:
- Library hi tech. Volume 37:Issue 1(2019)
- Journal:
- Library hi tech
- Issue:
- Volume 37:Issue 1(2019)
- Issue Display:
- Volume 37, Issue 1 (2019)
- Year:
- 2019
- Volume:
- 37
- Issue:
- 1
- Issue Sort Value:
- 2019-0037-0001-0000
- Page Start:
- 101
- Page End:
- 117
- Publication Date:
- 2018-06-20
- Subjects:
- Precision -- Recall -- Arabic text -- Auto-indexing -- Frequent sets -- Rule-based data mining
Library science -- Technological innovations -- Periodicals
Libraries -- Automation -- Periodicals
Information science -- Periodicals
025.00285 - Journal URLs:
- http://www.emeraldinsight.com/0737-8831.htm ↗
http://www.emeraldinsight.com/ ↗
http://firstsearch.oclc.org ↗ - DOI:
- 10.1108/LHT-07-2017-0147 ↗
- Languages:
- English
- ISSNs:
- 0737-8831
- Deposit Type:
- Legaldeposit
- View Content:
- Available online (eLD content is only available in our Reading Rooms) ↗
- Physical Locations:
- British Library DSC - 5198.870000
British Library DSC - BLDSS-3PM
British Library HMNTS - ELD Digital store - Ingest File:
- 22107.xml