Using electronic health records to identify candidates for human immunodeficiency virus pre‐exposure prophylaxis: An application of super learning to risk prediction when the outcome is rare. (24th June 2020)
- Record Type:
- Journal Article
- Title:
- Using electronic health records to identify candidates for human immunodeficiency virus pre‐exposure prophylaxis: An application of super learning to risk prediction when the outcome is rare. (24th June 2020)
- Main Title:
- Using electronic health records to identify candidates for human immunodeficiency virus pre‐exposure prophylaxis: An application of super learning to risk prediction when the outcome is rare
- Authors:
- Gruber, Susan
Krakower, Douglas
Menchaca, John T.
Hsu, Katherine
Hawrusik, Rebecca
Maro, Judith C.
Cocoros, Noelle M.
Kruskal, Benjamin A.
Wilson, Ira B.
Mayer, Kenneth H.
Klompas, Michael - Abstract:
- Abstract : Human immunodeficiency virus (HIV) pre‐exposure prophylaxis (PrEP) protects high risk patients from becoming infected with HIV. Clinicians need help to identify candidates for PrEP based on information routinely collected in electronic health records (EHRs). The greatest statistical challenge in developing a risk prediction model is that acquisition is extremely rare. Methods : Data consisted of 180 covariates (demographic, diagnoses, treatments, prescriptions) extracted from records on 399 385 patient (150 cases) seen at Atrius Health (2007‐2015), a clinical network in Massachusetts. Super learner is an ensemble machine learning algorithm that uses k ‐fold cross validation to evaluate and combine predictions from a collection of algorithms. We trained 42 variants of sophisticated algorithms, using different sampling schemes that more evenly balanced the ratio of cases to controls. We compared super learner's cross validated area under the receiver operating curve (cv‐AUC) with that of each individual algorithm. Results : The least absolute shrinkage and selection operator (lasso) using a 1:20 class ratio outperformed the super learner (cv‐AUC = 0.86 vs 0.84). A traditional logistic regression model restricted to 23 clinician‐selected main terms was slightly inferior (cv‐AUC = 0.81). Conclusion : Machine learning was successful at developing a model to predict 1‐year risk of acquiring HIV based on a physician‐curated set of predictors extracted from EHRs.
- Is Part Of:
- Statistics in medicine. Volume 39:Number 23(2020)
- Journal:
- Statistics in medicine
- Issue:
- Volume 39:Number 23(2020)
- Issue Display:
- Volume 39, Issue 23 (2020)
- Year:
- 2020
- Volume:
- 39
- Issue:
- 23
- Issue Sort Value:
- 2020-0039-0023-0000
- Page Start:
- 3059
- Page End:
- 3073
- Publication Date:
- 2020-06-24
- Subjects:
- EHR -- machine learning -- predictive modeling -- PrEP -- risk score prediction -- super learner
Medical statistics -- Periodicals
Statistique médicale -- Périodiques
Statistiques médicales -- Périodiques
610.727 - Journal URLs:
- http://onlinelibrary.wiley.com/ ↗
- DOI:
- 10.1002/sim.8591 ↗
- Languages:
- English
- ISSNs:
- 0277-6715
- Deposit Type:
- Legaldeposit
- View Content:
- Available online (eLD content is only available in our Reading Rooms) ↗
- Physical Locations:
- British Library DSC - 8453.576000
British Library DSC - BLDSS-3PM
British Library STI - ELD Digital store - Ingest File:
- 14602.xml