Evaluation of variable selection methods for random forests and omics data sets. (16th October 2017)

Record Type:: Journal Article
Title:: Evaluation of variable selection methods for random forests and omics data sets. (16th October 2017)
Main Title:: Evaluation of variable selection methods for random forests and omics data sets
Authors:: Degenhardt, Frauke
Seifert, Stephan
Szymczak, Silke
Abstract:: Abstract: Machine learning methods and in particular random forests are promising approaches for prediction based on high dimensional omics data sets. They provide variable importance measures to rank predictors according to their predictive power. If building a prediction model is the main goal of a study, often a minimal set of variables with good prediction performance is selected. However, if the objective is the identification of involved variables to find active networks and pathways, approaches that aim to select all relevant variables should be preferred. We evaluated several variable selection procedures based on simulated data as well as publicly available experimental methylation and gene expression data. Our comparison included the Boruta algorithm, the Vita method, recurrent relative variable importance, a permutation approach and its parametric variant (Altmann) as well as recursive feature elimination (RFE). In our simulation studies, Boruta was the most powerful approach, followed closely by the Vita method. Both approaches demonstrated similar stability in variable selection, while Vita was the most robust approach under a pure null model without any predictor variables related to the outcome. In the analysis of the different experimental data sets, Vita demonstrated slightly better stability in variable selection and was less computationally intensive than Boruta. In conclusion, we recommend the Boruta and Vita approaches for the analysis of … (more)
Is Part Of:: Briefings in bioinformatics. Volume 20:Number 2(2019)
Journal:: Briefings in bioinformatics
Issue:: Volume 20:Number 2(2019)
Issue Display:: Volume 20, Issue 2 (2019)
Year:: 2019
Volume:: 20
Issue:: 2
Issue Sort Value:: 2019-0020-0002-0000
Page Start:: 492
Page End:: 503
Publication Date:: 2017-10-16
Subjects:: machine learning -- random forest -- feature selection -- high dimensional data -- relevant variables
Genetics -- Data processing -- Periodicals
Molecular biology -- Data processing -- Periodicals
Genomes -- Data processing -- Periodicals
572.80285
Journal URLs:: http://bib.oxfordjournals.org ↗
http://www.oxfordjournals.org/content?genre=journal&issn=1477-4054 ↗
http://ukcatalogue.oup.com/ ↗
http://firstsearch.oclc.org ↗
DOI:: 10.1093/bib/bbx124 ↗
Languages:: English
ISSNs:: 1467-5463
Deposit Type:: Legaldeposit
View Content:: Available online (eLD content is only available in our Reading Rooms) ↗
Physical Locations:: British Library DSC - 2283.958363
British Library DSC - BLDSS-3PM
British Library HMNTS - ELD Digital store
Ingest File:: 20848.xml