Automated retrieval of information on threatened species from online sources using machine learning. Issue 7 (6th May 2021)
- Record Type:
- Journal Article
- Title:
- Automated retrieval of information on threatened species from online sources using machine learning. Issue 7 (6th May 2021)
- Main Title:
- Automated retrieval of information on threatened species from online sources using machine learning
- Authors:
- Kulkarni, Ritwik
Di Minin, Enrico - Abstract:
- Abstract: As resources for conservation are limited, gathering and analysing information from digital platforms can help investigate the global biodiversity crisis in a cost‐efficient manner. Development and application of methods for automated content analysis of digital data sources are especially important in the context of investigating human–nature interactions. In this study, we introduce novel application methods to automatically collect and analyse textual data on species of conservation concern from digital platforms. An end‐to‐end pipeline is constructed that begins from searching and downloading news articles about species listed in Appendix I of the Convention on International Trade in Endangered Species of Wild Fauna and Flora (CITES) along with news articles from specific Twitter handles and proceeds with implementing natural language processing and machine learning methods to filter and retain only relevant articles. A crucial aspect here is the automatic annotation of training data, which can be challenging in many machine learning applications. A Named Entity Recognition model is then used to extract additional relevant information for each article. The data collected over a 1‐month period included 15, 088 articles focusing on 585 species listed in Appendix I of CITES. The accuracy of the neural network to detect relevant articles was 95.91% while the Named Entity recognition model helped extract information on prices, location and quantities of tradedAbstract: As resources for conservation are limited, gathering and analysing information from digital platforms can help investigate the global biodiversity crisis in a cost‐efficient manner. Development and application of methods for automated content analysis of digital data sources are especially important in the context of investigating human–nature interactions. In this study, we introduce novel application methods to automatically collect and analyse textual data on species of conservation concern from digital platforms. An end‐to‐end pipeline is constructed that begins from searching and downloading news articles about species listed in Appendix I of the Convention on International Trade in Endangered Species of Wild Fauna and Flora (CITES) along with news articles from specific Twitter handles and proceeds with implementing natural language processing and machine learning methods to filter and retain only relevant articles. A crucial aspect here is the automatic annotation of training data, which can be challenging in many machine learning applications. A Named Entity Recognition model is then used to extract additional relevant information for each article. The data collected over a 1‐month period included 15, 088 articles focusing on 585 species listed in Appendix I of CITES. The accuracy of the neural network to detect relevant articles was 95.91% while the Named Entity recognition model helped extract information on prices, location and quantities of traded animals and plants. A regularly updated database, which can be queried and analysed for various research purposes and to inform conservation decision making, is generated by the system. The results demonstrate that natural language processing can be used successfully to extract information from digital text content. The proposed methods can be applied to multiple digital data platforms at the same time and used to investigate human–nature interactions in conservation science and practice. … (more)
- Is Part Of:
- Methods in ecology and evolution. Volume 12:Issue 7(2021)
- Journal:
- Methods in ecology and evolution
- Issue:
- Volume 12:Issue 7(2021)
- Issue Display:
- Volume 12, Issue 7 (2021)
- Year:
- 2021
- Volume:
- 12
- Issue:
- 7
- Issue Sort Value:
- 2021-0012-0007-0000
- Page Start:
- 1226
- Page End:
- 1239
- Publication Date:
- 2021-05-06
- Subjects:
- biodiversity -- digital conservation -- machine learning -- natural language processing -- wildlife trade
Ecology -- Periodicals
Evolution -- Periodicals
577 - Journal URLs:
- http://onlinelibrary.wiley.com/journal/10.1111/(ISSN)2041-210X ↗
http://onlinelibrary.wiley.com/ ↗ - DOI:
- 10.1111/2041-210X.13608 ↗
- Languages:
- English
- ISSNs:
- 2041-210X
- Deposit Type:
- Legaldeposit
- View Content:
- Available online (eLD content is only available in our Reading Rooms) ↗
- Physical Locations:
- British Library DSC - BLDSS-3PM
British Library HMNTS - ELD Digital store - Ingest File:
- 17454.xml