A comparative study of effective approaches for Arabic sentiment analysis. Issue 2 (March 2021)
- Record Type:
- Journal Article
- Title:
- A comparative study of effective approaches for Arabic sentiment analysis. Issue 2 (March 2021)
- Main Title:
- A comparative study of effective approaches for Arabic sentiment analysis
- Authors:
- Abu Farha, Ibrahim
Magdy, Walid - Abstract:
- Highlights: We survey state-of-the-art methods for Arabic sentiment analysis. We replicate the most effective methods (over 10 methods) for Arabic SA reported in the literature and introduce new ones, and compare their performance on the most popular publicly available benchmark Arabic SA datasets. We built the largest Arabic word-embeddings trained on 250 million unique tweets, covering multiple Arabic dialects that exist on social media. We apply BERT-based models for Arabic SA and compare their performance with all the existing state-of-the-art Arabic SA approaches, showing the superiority of using Arabic specific BERT. We conduct an extensive error analysis for the different approaches, which included the reannotation of the existing Arabic SA datasets to assess the subjectivity of the task and the presence of sarcasm. We empirically show the challenges that sarcasm imposes on sentiment analysis systems. Based on our comprehensive analysis, we suggest the most important future research directions in Arabic sentiment analysis in specific and related NLP tasks in general. Abstract: Sentiment analysis (SA) is a natural language processing (NLP) application that aims to analyse and identify sentiment within a piece of text. Arabic SA started to receive more attention in the last decade with many approaches showing some effectiveness for detecting sentiment on multiple datasets. While there have been some surveys summarising some of the approaches for Arabic SA in literature,Highlights: We survey state-of-the-art methods for Arabic sentiment analysis. We replicate the most effective methods (over 10 methods) for Arabic SA reported in the literature and introduce new ones, and compare their performance on the most popular publicly available benchmark Arabic SA datasets. We built the largest Arabic word-embeddings trained on 250 million unique tweets, covering multiple Arabic dialects that exist on social media. We apply BERT-based models for Arabic SA and compare their performance with all the existing state-of-the-art Arabic SA approaches, showing the superiority of using Arabic specific BERT. We conduct an extensive error analysis for the different approaches, which included the reannotation of the existing Arabic SA datasets to assess the subjectivity of the task and the presence of sarcasm. We empirically show the challenges that sarcasm imposes on sentiment analysis systems. Based on our comprehensive analysis, we suggest the most important future research directions in Arabic sentiment analysis in specific and related NLP tasks in general. Abstract: Sentiment analysis (SA) is a natural language processing (NLP) application that aims to analyse and identify sentiment within a piece of text. Arabic SA started to receive more attention in the last decade with many approaches showing some effectiveness for detecting sentiment on multiple datasets. While there have been some surveys summarising some of the approaches for Arabic SA in literature, most of these approaches are reported on different datasets, which makes it difficult to identify the most effective approaches among those. In addition, those approaches do not cover the recent advances in NLP that use transformers. This paper presents a comprehensive comparative study on the most effective approaches used for Arabic sentiment analysis. We re-implement most of the existing approaches for Arabic SA and test their effectiveness on three of the most popular benchmark datasets for Arabic SA. Further, we examine the use of transformer-based language models for Arabic SA and show their superior performance compared to the existing approaches, where the best model achieves F-score scores of 0.69, 0.76, and 0.92 on the SemEval, ASTD, and ArSAS benchmark datasets. We also apply an extensive analysis of the possible reasons for failures, which show the limitations of the existing annotated Arabic SA datasets, and the challenge of sarcasm that is prominent in Arabic dialects. Finally, we highlight the main gaps in Arabic sentiment analysis research and suggest the most in-need future research directions in this area. … (more)
- Is Part Of:
- Information processing & management. Volume 58:Issue 2(2021)
- Journal:
- Information processing & management
- Issue:
- Volume 58:Issue 2(2021)
- Issue Display:
- Volume 58, Issue 2 (2021)
- Year:
- 2021
- Volume:
- 58
- Issue:
- 2
- Issue Sort Value:
- 2021-0058-0002-0000
- Page Start:
- Page End:
- Publication Date:
- 2021-03
- Subjects:
- Arabic -- Sentiment Analysis -- Sarcasm
Information storage and retrieval systems -- Periodicals
Information science -- Periodicals
Systèmes d'information -- Périodiques
Sciences de l'information -- Périodiques
Information science
Information storage and retrieval systems
Periodicals
658.4038 - Journal URLs:
- http://www.sciencedirect.com/science/journal/03064573 ↗
http://www.elsevier.com/journals ↗ - DOI:
- 10.1016/j.ipm.2020.102438 ↗
- Languages:
- English
- ISSNs:
- 0306-4573
- Deposit Type:
- Legaldeposit
- View Content:
- Available online (eLD content is only available in our Reading Rooms) ↗
- Physical Locations:
- British Library DSC - 4493.893000
British Library DSC - BLDSS-3PM
British Library HMNTS - ELD Digital store - Ingest File:
- 15543.xml