Performance of Regression Models as a Function of Experiment Noise. Issue 15 (June 2021)
- Record Type:
- Journal Article
- Title:
- Performance of Regression Models as a Function of Experiment Noise. Issue 15 (June 2021)
- Main Title:
- Performance of Regression Models as a Function of Experiment Noise
- Authors:
- Li, Gang
Zrimec, Jan
Ji, Boyang
Geng, Jun
Larsbrink, Johan
Zelezniak, Aleksej
Nielsen, Jens
Engqvist, Martin KM - Abstract:
- Background: A challenge in developing machine learning regression models is that it is difficult to know whether maximal performance has been reached on the test dataset, or whether further model improvement is possible. In biology, this problem is particularly pronounced as sample labels (response variables) are typically obtained through experiments and therefore have experiment noise associated with them. Such label noise puts a fundamental limit to the metrics of performance attainable by regression models on the test dataset. Results: We address this challenge by deriving an expected upper bound for the coefficient of determination ( R 2 ) for regression models when tested on the holdout dataset. This upper bound depends only on the noise associated with the response variable in a dataset as well as its variance. The upper bound estimate was validated via Monte Carlo simulations and then used as a tool to bootstrap performance of regression models trained on biological datasets, including protein sequence data, transcriptomic data, and genomic data. Conclusions: The new method for estimating upper bounds for model performance on test data should aid researchers in developing ML regression models that reach their maximum potential. Although we study biological datasets in this work, the new upper bound estimates will hold true for regression models from any research field or application area where response variables have associated noise.
- Is Part Of:
- Bioinformatics and biology insights. Volume 2021:Issue 15(2021)
- Journal:
- Bioinformatics and biology insights
- Issue:
- Volume 2021:Issue 15(2021)
- Issue Display:
- Volume 2021, Issue 15 (2021)
- Year:
- 2021
- Volume:
- 2021
- Issue:
- 15
- Issue Sort Value:
- 2021-2021-0015-0000
- Page Start:
- Page End:
- Publication Date:
- 2021-06
- Subjects:
- machine learning -- experiment noise -- label noise -- regression models -- upper bound
Bioinformatics -- Periodicals
Biology -- Data processing -- Periodicals
570.285 - Journal URLs:
- http://insights.sagepub.com/journal-bioinformatics-and-biology-insights-j39 ↗
http://www.uk.sagepub.com/home.nav ↗ - DOI:
- 10.1177/11779322211020315 ↗
- Languages:
- English
- ISSNs:
- 1177-9322
- Deposit Type:
- Legaldeposit
- View Content:
- Available online (eLD content is only available in our Reading Rooms) ↗
- Physical Locations:
- British Library DSC - BLDSS-3PM
British Library HMNTS - ELD Digital store - Ingest File:
- 19277.xml