Citizen Science for Mining the Biomedical Literature. (31st December 2016)
- Record Type:
- Journal Article
- Title:
- Citizen Science for Mining the Biomedical Literature. (31st December 2016)
- Main Title:
- Citizen Science for Mining the Biomedical Literature
- Authors:
- Tsueng, Ginger
Nanis, Steven M.
Fouquier, Jennifer
Good, Benjamin M.
Su, Andrew I. - Abstract:
- Biomedical literature represents one of the largest and fastest growing collections of unstructured biomedical knowledge. Finding critical information buried in the literature can be challenging. To extract information from free-flowing text, researchers need to: 1. identify the entities in the text (named entity recognition), 2. apply a standardized vocabulary to these entities (normalization), and 3. identify how entities in the text are related to one another (relationship extraction). Researchers have primarily approached these information extraction tasks through manual expert curation and computational methods. We have previously demonstrated that named entity recognition (NER) tasks can be crowdsourced to a group of non-experts via the paid microtask platform, Amazon Mechanical Turk (AMT), and can dramatically reduce the cost and increase the throughput of biocuration efforts. However, given the size of the biomedical literature, even information extraction via paid microtask platforms is not scalable. With our web-based application Mark2Cure (http://mark2cure.org ), we demonstrate that NER tasks also can be performed by volunteer citizen scientists with high accuracy. We apply metrics from the Zooniverse Matrices of Citizen Science Success and provide the results here to serve as a basis of comparison for other citizen science projects. Further, we discuss design considerations, issues, and the application of analytics for successfully moving a crowdsourcing workflowBiomedical literature represents one of the largest and fastest growing collections of unstructured biomedical knowledge. Finding critical information buried in the literature can be challenging. To extract information from free-flowing text, researchers need to: 1. identify the entities in the text (named entity recognition), 2. apply a standardized vocabulary to these entities (normalization), and 3. identify how entities in the text are related to one another (relationship extraction). Researchers have primarily approached these information extraction tasks through manual expert curation and computational methods. We have previously demonstrated that named entity recognition (NER) tasks can be crowdsourced to a group of non-experts via the paid microtask platform, Amazon Mechanical Turk (AMT), and can dramatically reduce the cost and increase the throughput of biocuration efforts. However, given the size of the biomedical literature, even information extraction via paid microtask platforms is not scalable. With our web-based application Mark2Cure (http://mark2cure.org ), we demonstrate that NER tasks also can be performed by volunteer citizen scientists with high accuracy. We apply metrics from the Zooniverse Matrices of Citizen Science Success and provide the results here to serve as a basis of comparison for other citizen science projects. Further, we discuss design considerations, issues, and the application of analytics for successfully moving a crowdsourcing workflow from a paid microtask platform to a citizen science platform. To our knowledge, this study is the first application of citizen science to a natural language processing task. … (more)
- Is Part Of:
- Citizen science. Volume 1:Number 2(2016)
- Journal:
- Citizen science
- Issue:
- Volume 1:Number 2(2016)
- Issue Display:
- Volume 1, Issue 2 (2016)
- Year:
- 2016
- Volume:
- 1
- Issue:
- 2
- Issue Sort Value:
- 2016-0001-0002-0000
- Page Start:
- Page End:
- Publication Date:
- 2016-12-31
- Subjects:
- information extraction -- citizen science -- microtask -- biocuration -- natural language processing -- biomedical literature
Science -- Citizen participation -- Periodicals
Volunteer workers in science -- Periodicals
507.2 - Journal URLs:
- http://theoryandpractice.citizenscienceassociation.org/articles/ ↗
- DOI:
- 10.5334/cstp.56 ↗
- Languages:
- English
- ISSNs:
- 2057-4991
- Deposit Type:
- Legaldeposit
- View Content:
- Available online (eLD content is only available in our Reading Rooms) ↗
- Physical Locations:
- British Library HMNTS - ELD Digital store
- Ingest File:
- 14678.xml