Development and validation of various phenotyping algorithms for Diabetes Mellitus using data from electronic health records. (December 2017)
- Record Type:
- Journal Article
- Title:
- Development and validation of various phenotyping algorithms for Diabetes Mellitus using data from electronic health records. (December 2017)
- Main Title:
- Development and validation of various phenotyping algorithms for Diabetes Mellitus using data from electronic health records
- Authors:
- Esteban, Santiago
Rodríguez Tablado, Manuel
Peper, Francisco E.
Mahumud, Yamila S.
Ricci, Ricardo I.
Kopitowski, Karin S.
Terrasa, Sergio A. - Abstract:
- Highlights: Performance comparison of different electronic phenotyping algorithms using real data extracted from EHRs from Argentina. Stacked generalization showed to be the best performing strategy in the validation set. These algorithms can facilitate the development of local studies, pontetially reducing research costs. Abstract: Background and Objective: Recent progression towards precision medicine has encouraged the use of electronic health records (EHRs) as a source for large amounts of data, which is required for studying the effect of treatments or risk factors in more specific subpopulations. Phenotyping algorithms allow to automatically classify patients according to their particular electronic phenotype thus facilitating the setup of retrospective cohorts. Our objective is to compare the performance of different classification strategies (only using standardized problems, rule-based algorithms, statistical learning algorithms (six learners) and stacked generalization (five versions)), for the categorization of patients according to their diabetic status (diabetics, not diabetics and inconclusive; Diabetes of any type) using information extracted from EHRs. Methods: Patient information was extracted from the EHR at Hospital Italiano de Buenos Aires, Buenos Aires, Argentina. For the derivation and validation datasets, two probabilistic samples of patients from different years (2005: n = 1663; 2015: n = 800) were extracted. The only inclusion criterion was age (≥40Highlights: Performance comparison of different electronic phenotyping algorithms using real data extracted from EHRs from Argentina. Stacked generalization showed to be the best performing strategy in the validation set. These algorithms can facilitate the development of local studies, pontetially reducing research costs. Abstract: Background and Objective: Recent progression towards precision medicine has encouraged the use of electronic health records (EHRs) as a source for large amounts of data, which is required for studying the effect of treatments or risk factors in more specific subpopulations. Phenotyping algorithms allow to automatically classify patients according to their particular electronic phenotype thus facilitating the setup of retrospective cohorts. Our objective is to compare the performance of different classification strategies (only using standardized problems, rule-based algorithms, statistical learning algorithms (six learners) and stacked generalization (five versions)), for the categorization of patients according to their diabetic status (diabetics, not diabetics and inconclusive; Diabetes of any type) using information extracted from EHRs. Methods: Patient information was extracted from the EHR at Hospital Italiano de Buenos Aires, Buenos Aires, Argentina. For the derivation and validation datasets, two probabilistic samples of patients from different years (2005: n = 1663; 2015: n = 800) were extracted. The only inclusion criterion was age (≥40 & <80 years). Four researchers manually reviewed all records and classified patients according to their diabetic status (diabetic: diabetes registered as a health problem or fulfilling the ADA criteria; non-diabetic: not fulfilling the ADA criteria and having at least one fasting glycemia below 126 mg/dL; inconclusive: no data regarding their diabetic status or only one abnormal value). The best performing algorithms within each strategy were tested on the validation set. Results: The standardized codes algorithm achieved a Kappa coefficient value of 0.59 (95% CI 0.49, 0.59) in the validation set. The Boolean logic algorithm reached 0.82 (95% CI 0.76, 0.88). A slightly higher value was achieved by the Feedforward Neural Network (0.9, 95% CI 0.85, 0.94). The best performing learner was the stacked generalization meta-learner that reached a Kappa coefficient value of 0.95 (95% CI 0.91, 0.98). Conclusions: The stacked generalization strategy and the feedforward neural network showed the best classification metrics in the validation set. The implementation of these algorithms enables the exploitation of the data of thousands of patients accurately. … (more)
- Is Part Of:
- Computer methods and programs in biomedicine. Volume 152(2017)
- Journal:
- Computer methods and programs in biomedicine
- Issue:
- Volume 152(2017)
- Issue Display:
- Volume 152, Issue 2017 (2017)
- Year:
- 2017
- Volume:
- 152
- Issue:
- 2017
- Issue Sort Value:
- 2017-0152-2017-0000
- Page Start:
- 53
- Page End:
- 70
- Publication Date:
- 2017-12
- Subjects:
- Electronic phenotyping algorithms -- Stacked generalization -- Electronic health records -- Diabetes Mellitus
Medicine -- Computer programs -- Periodicals
Biology -- Computer programs -- Periodicals
Computers -- Periodicals
Medicine -- Periodicals
Médecine -- Logiciels -- Périodiques
Biologie -- Logiciels -- Périodiques
Biology -- Computer programs
Medicine -- Computer programs
Periodicals
Electronic journals
610.28 - Journal URLs:
- http://www.sciencedirect.com/science/journal/01692607 ↗
http://www.elsevier.com/journals ↗ - DOI:
- 10.1016/j.cmpb.2017.09.009 ↗
- Languages:
- English
- ISSNs:
- 0169-2607
- Deposit Type:
- Legaldeposit
- View Content:
- Available online (eLD content is only available in our Reading Rooms) ↗
- Physical Locations:
- British Library DSC - 3394.095000
British Library DSC - BLDSS-3PM
British Library HMNTS - ELD Digital store - Ingest File:
- 4896.xml