Maximizing clinical cohort size using free text queries. (1st May 2015)
- Record Type:
- Journal Article
- Title:
- Maximizing clinical cohort size using free text queries. (1st May 2015)
- Main Title:
- Maximizing clinical cohort size using free text queries
- Authors:
- Gundlapalli, Adi V.
Redd, Doug
Gibson, Bryan Smith
Carter, Marjorie
Korhonen, Chris
Nebeker, Jonathan
Samore, Matthew H.
Zeng-Treitler, Qing - Abstract:
- Abstract: Background: Cohort identification is important in both population health management and research. In this project we sought to assess the use of text queries for cohort identification. Specifically we sought to determine the incremental value of unstructured data queries when added to structured queries for the purpose of patient cohort identification. Methods: Three cohort identification tasks were evaluated: identification of individuals taking gingko biloba and warfarin simultaneously (Gingko/Warfarin), individuals who were overweight, and individuals with uncontrolled diabetes (UCD). We assessed the increase in cohort size when unstructured data queries were added to structured data queries. The positive predictive value of unstructured data queries was assessed by manual chart review of a random sample of 500 patients. Results: For Gingko/Warfarin, text query increased the cohort size from 9 to 28, 924 over the cohort identified by query of pharmacy data only. For the weight-related tasks, text search increased the cohort by 5–29% compared to the cohort identified by query of the vitals table. For the UCD task, text query increased the cohort size by 2–43% compared to the cohort identified by query of laboratory results or ICD codes. The positive predictive values for text searches were 52% for Gingko/Warfarin, 19–94% for the weight cohort and 44% for UCD. Discussion: This project demonstrates the value and limitation of free text queries in patient cohortAbstract: Background: Cohort identification is important in both population health management and research. In this project we sought to assess the use of text queries for cohort identification. Specifically we sought to determine the incremental value of unstructured data queries when added to structured queries for the purpose of patient cohort identification. Methods: Three cohort identification tasks were evaluated: identification of individuals taking gingko biloba and warfarin simultaneously (Gingko/Warfarin), individuals who were overweight, and individuals with uncontrolled diabetes (UCD). We assessed the increase in cohort size when unstructured data queries were added to structured data queries. The positive predictive value of unstructured data queries was assessed by manual chart review of a random sample of 500 patients. Results: For Gingko/Warfarin, text query increased the cohort size from 9 to 28, 924 over the cohort identified by query of pharmacy data only. For the weight-related tasks, text search increased the cohort by 5–29% compared to the cohort identified by query of the vitals table. For the UCD task, text query increased the cohort size by 2–43% compared to the cohort identified by query of laboratory results or ICD codes. The positive predictive values for text searches were 52% for Gingko/Warfarin, 19–94% for the weight cohort and 44% for UCD. Discussion: This project demonstrates the value and limitation of free text queries in patient cohort identification from large data sets. The clinical domain and prevalence of the inclusion and exclusion criteria in the patient population influence the utility and yield of this approach. Highlights: We demonstrate the value of free text queries in cohort identification. Incremental value is added compared to structured data queries alone. We determine the value of free text using 3 disparate use cases. Use case specific values and limitations are identified in large data sets. Exploratory value of a direct search tool in contrast with heavier NLP systems. … (more)
- Is Part Of:
- Computers in biology and medicine. Volume 60(2015)
- Journal:
- Computers in biology and medicine
- Issue:
- Volume 60(2015)
- Issue Display:
- Volume 60, Issue 2015 (2015)
- Year:
- 2015
- Volume:
- 60
- Issue:
- 2015
- Issue Sort Value:
- 2015-0060-2015-0000
- Page Start:
- 1
- Page End:
- 7
- Publication Date:
- 2015-05-01
- Subjects:
- Gingko -- Warfarin -- Overweight -- Diabetes -- Text query -- Structured data -- Cohort identification -- Unstructured data -- Clinical notes
Medicine -- Data processing -- Periodicals
Biology -- Data processing -- Periodicals
610.285 - Journal URLs:
- http://www.sciencedirect.com/science/journal/00104825/ ↗
http://www.elsevier.com/journals ↗ - DOI:
- 10.1016/j.compbiomed.2015.01.008 ↗
- Languages:
- English
- ISSNs:
- 0010-4825
- Deposit Type:
- Legaldeposit
- View Content:
- Available online (eLD content is only available in our Reading Rooms) ↗
- Physical Locations:
- British Library DSC - 3394.880000
British Library DSC - BLDSS-3PM
British Library HMNTS - ELD Digital store - Ingest File:
- 6340.xml