Minimum redundancy maximum relevance feature selection approach for temporal gene expression data. Issue 1 (December 2017)
- Record Type:
- Journal Article
- Title:
- Minimum redundancy maximum relevance feature selection approach for temporal gene expression data. Issue 1 (December 2017)
- Main Title:
- Minimum redundancy maximum relevance feature selection approach for temporal gene expression data
- Authors:
- Radovic, Milos
Ghalwash, Mohamed
Filipovic, Nenad
Obradovic, Zoran - Abstract:
- Abstract Background Feature selection, aiming to identify a subset of features among a possibly large set of features that are relevant for predicting a response, is an important preprocessing step in machine learning. In gene expression studies this is not a trivial task for several reasons, including potential temporal character of data. However, most feature selection approaches developed for microarray data cannot handle multivariate temporal data without previous data flattening, which results in loss of temporal information. We propose a temporal minimum redundancy - maximum relevance (TMRMR) feature selection approach, which is able to handle multivariate temporal data without previous data flattening. In the proposed approach we compute relevance of a gene by averaging F-statistic values calculated across individual time steps, and we compute redundancy between genes by using a dynamical time warping approach. Results The proposed method is evaluated on three temporal gene expression datasets from human viral challenge studies. Obtained results show that the proposed method outperforms alternatives widely used in gene expression studies. In particular, the proposed method achieved improvement in accuracy in 34 out of 54 experiments, while the other methods outperformed it in no more than 4 experiments. Conclusion We developed a filter-based feature selection method for temporal gene expression data based on maximum relevance and minimum redundancy criteria. TheAbstract Background Feature selection, aiming to identify a subset of features among a possibly large set of features that are relevant for predicting a response, is an important preprocessing step in machine learning. In gene expression studies this is not a trivial task for several reasons, including potential temporal character of data. However, most feature selection approaches developed for microarray data cannot handle multivariate temporal data without previous data flattening, which results in loss of temporal information. We propose a temporal minimum redundancy - maximum relevance (TMRMR) feature selection approach, which is able to handle multivariate temporal data without previous data flattening. In the proposed approach we compute relevance of a gene by averaging F-statistic values calculated across individual time steps, and we compute redundancy between genes by using a dynamical time warping approach. Results The proposed method is evaluated on three temporal gene expression datasets from human viral challenge studies. Obtained results show that the proposed method outperforms alternatives widely used in gene expression studies. In particular, the proposed method achieved improvement in accuracy in 34 out of 54 experiments, while the other methods outperformed it in no more than 4 experiments. Conclusion We developed a filter-based feature selection method for temporal gene expression data based on maximum relevance and minimum redundancy criteria. The proposed method incorporates temporal information by combining relevance, which is calculated as an average F-statistic value across different time steps, with redundancy, which is calculated by employing dynamical time warping approach. As evident in our experiments, incorporating the temporal information into the feature selection process leads to selection of more discriminative features. … (more)
- Is Part Of:
- BMC bioinformatics. Volume 18:Issue 1(2017)
- Journal:
- BMC bioinformatics
- Issue:
- Volume 18:Issue 1(2017)
- Issue Display:
- Volume 18, Issue 1 (2017)
- Year:
- 2017
- Volume:
- 18
- Issue:
- 1
- Issue Sort Value:
- 2017-0018-0001-0000
- Page Start:
- 1
- Page End:
- 14
- Publication Date:
- 2017-12
- Subjects:
- Feature selection -- Gene expression -- Temporal data
Bioinformatics -- Periodicals
Computational biology -- Periodicals
570.285 - Journal URLs:
- http://www.biomedcentral.com/bmcbioinformatics/ ↗
http://www.pubmedcentral.nih.gov/tocrender.fcgi?journal=13 ↗
http://link.springer.com/ ↗ - DOI:
- 10.1186/s12859-016-1423-9 ↗
- Languages:
- English
- ISSNs:
- 1471-2105
- Deposit Type:
- Legaldeposit
- View Content:
- Available online (eLD content is only available in our Reading Rooms) ↗
- Physical Locations:
- British Library DSC - BLDSS-3PM
British Library HMNTS - Digital store
British Library HMNTS - ELD Digital store - Ingest File:
- 9978.xml