Estimation of required sample size for external validation of risk models for binary outcomes. (October 2021)
- Record Type:
- Journal Article
- Title:
- Estimation of required sample size for external validation of risk models for binary outcomes. (October 2021)
- Main Title:
- Estimation of required sample size for external validation of risk models for binary outcomes
- Authors:
- Pavlou, Menelaos
Qu, Chen
Omar, Rumana Z
Seaman, Shaun R
Steyerberg, Ewout W
White, Ian R
Ambler, Gareth - Abstract:
- Risk-prediction models for health outcomes are used in practice as part of clinical decision-making, and it is essential that their performance be externally validated. An important aspect in the design of a validation study is choosing an adequate sample size. In this paper, we investigate the sample size requirements for validation studies with binary outcomes to estimate measures of predictive performance (C-statistic for discrimination and calibration slope and calibration in the large). We aim for sufficient precision in the estimated measures. In addition, we investigate the sample size to achieve sufficient power to detect a difference from a target value. Under normality assumptions on the distribution of the linear predictor, we obtain simple estimators for sample size calculations based on the measures above. Simulation studies show that the estimators perform well for common values of the C-statistic and outcome prevalence when the linear predictor is marginally Normal. Their performance deteriorates only slightly when the normality assumptions are violated. We also propose estimators which do not require normality assumptions but require specification of the marginal distribution of the linear predictor and require the use of numerical integration. These estimators were also seen to perform very well under marginal normality. Our sample size equations require a specified standard error (SE) and the anticipated C-statistic and outcome prevalence. The sample sizeRisk-prediction models for health outcomes are used in practice as part of clinical decision-making, and it is essential that their performance be externally validated. An important aspect in the design of a validation study is choosing an adequate sample size. In this paper, we investigate the sample size requirements for validation studies with binary outcomes to estimate measures of predictive performance (C-statistic for discrimination and calibration slope and calibration in the large). We aim for sufficient precision in the estimated measures. In addition, we investigate the sample size to achieve sufficient power to detect a difference from a target value. Under normality assumptions on the distribution of the linear predictor, we obtain simple estimators for sample size calculations based on the measures above. Simulation studies show that the estimators perform well for common values of the C-statistic and outcome prevalence when the linear predictor is marginally Normal. Their performance deteriorates only slightly when the normality assumptions are violated. We also propose estimators which do not require normality assumptions but require specification of the marginal distribution of the linear predictor and require the use of numerical integration. These estimators were also seen to perform very well under marginal normality. Our sample size equations require a specified standard error (SE) and the anticipated C-statistic and outcome prevalence. The sample size requirement varies according to the prognostic strength of the model, outcome prevalence, choice of the performance measure and study objective. For example, to achieve an SE < 0.025 for the C-statistic, 60–170 events are required if the true C-statistic and outcome prevalence are between 0.64–0.85 and 0.05–0.3, respectively. For the calibration slope and calibration in the large, achieving SE < 0.15 would require 40–280 and 50–100 events, respectively. Our estimators may also be used for survival outcomes when the proportion of censored observations is high. … (more)
- Is Part Of:
- Statistical methods in medical research. Volume 30:Number 10(2021)
- Journal:
- Statistical methods in medical research
- Issue:
- Volume 30:Number 10(2021)
- Issue Display:
- Volume 30, Issue 10 (2021)
- Year:
- 2021
- Volume:
- 30
- Issue:
- 10
- Issue Sort Value:
- 2021-0030-0010-0000
- Page Start:
- 2187
- Page End:
- 2206
- Publication Date:
- 2021-10
- Subjects:
- Sample size calculation -- prediction model -- C-statistic -- discrimination -- calibration
Medicine -- Research -- Statistical methods -- Periodicals
Research -- Periodicals
Review Literature -- Periodicals
Statistics -- methods -- Periodicals
Médecine -- Recherche -- Méthodes statistiques -- Périodiques
610.727 - Journal URLs:
- http://smm.sagepub.com/ ↗
http://www.ingentaselect.com/rpsv/cw/arn/09622802/contp1.htm ↗
http://www.uk.sagepub.com/home.nav ↗
http://firstsearch.oclc.org ↗
http://firstsearch.oclc.org/journal=0962-2802;screen=info;ECOIP ↗ - DOI:
- 10.1177/09622802211007522 ↗
- Languages:
- English
- ISSNs:
- 0962-2802
- Deposit Type:
- Legaldeposit
- View Content:
- Available online (eLD content is only available in our Reading Rooms) ↗
- Physical Locations:
- British Library DSC - BLDSS-3PM
British Library HMNTS - ELD Digital store - Ingest File:
- 17371.xml