Characterising text mining: a systematic mapping review of the Portuguese language. Issue 2 (1st April 2018)
- Record Type:
- Journal Article
- Title:
- Characterising text mining: a systematic mapping review of the Portuguese language. Issue 2 (1st April 2018)
- Main Title:
- Characterising text mining: a systematic mapping review of the Portuguese language
- Authors:
- Souza, Ellen
Costa, Danilo
Castro, Dayvid W.
Vitório, Douglas
Teles, Ingryd
Almeida, Rafaela
Alves, Tiago
Oliveira, Adriano L.I.
Gusmão, Cristine - Abstract:
- Abstract : Documents written in natural language constitute a major part of the artefacts produced during the software engineering life cycle. Studies indicate that more than 80% of enterprise data is stored in some sort of unstructured form, mainly as text. Therefore, the growth of user‐generated content, especially from social media, provides a huge amount of data which allows discovering the experiences, opinions, and feelings of users. Text mining refers to the set of tools, techniques, and algorithms adopted to extract useful information from unstructured data. Considering that Portuguese ranks among the ten most spoken languages, and it is the second most common in Twitter, this study aims to map current primary studies that relate to the application of text mining for Portuguese. A systematic mapping method was applied and 6075 primary studies were retrieved up to the year 2014. A total of 203 studies were included, from which more than 60% analyse texts written in Brazilian variant. The majority of studies focus on the text classification task. Support vector machine and Naïve Bayes appear as main the algorithms. Folha de São Paulo and Público newspapers appear as main corpora, followed by the Portuguese Attorney General's Office corpus and Twitter.
- Is Part Of:
- IET software. Volume 12:Issue 2(2018)
- Journal:
- IET software
- Issue:
- Volume 12:Issue 2(2018)
- Issue Display:
- Volume 12, Issue 2 (2018)
- Year:
- 2018
- Volume:
- 12
- Issue:
- 2
- Issue Sort Value:
- 2018-0012-0002-0000
- Page Start:
- 49
- Page End:
- 75
- Publication Date:
- 2018-04-01
- Subjects:
- text analysis -- data mining -- natural language processing -- social networking (online) -- Bayes methods -- software engineering -- pattern classification -- support vector machines
text mining characterization -- systematic mapping review -- Portuguese language -- documents -- natural language -- software engineering life cycle -- enterprise data -- user-generated content -- social media -- information extraction -- unstructured data -- Twitter -- systematic mapping method -- text classification task -- support vector machine -- naïve Bayes -- Folha de São Paulo newspapers -- Público newspapers -- Portuguese Attorney General Office corpus
Computer software -- Periodicals
Software engineering -- Periodicals
005.1 - Journal URLs:
- http://digital-library.theiet.org/content/journals/iet-sen ↗
http://ieeexplore.ieee.org/servlet/opac?punumber=4124007 ↗
https://ietresearch.onlinelibrary.wiley.com/journal/17518814 ↗
http://www.theiet.org/ ↗
http://scitation.aip.org/dbt/dbt.jsp?KEY=ISEOB7&Volume=CURVOL&Issue=CURISS ↗ - DOI:
- 10.1049/iet-sen.2016.0226 ↗
- Languages:
- English
- ISSNs:
- 1751-8806
- Deposit Type:
- Legaldeposit
- View Content:
- Available online (eLD content is only available in our Reading Rooms) ↗
- Physical Locations:
- British Library DSC - 4363.253550
British Library DSC - BLDSS-3PM
British Library HMNTS - ELD Digital store - Ingest File:
- 16455.xml