Estimating gene expression from high-dimensional DNA methylation levels in cancer data: A bimodal unsupervised dimension reduction algorithm. (April 2019)
- Record Type:
- Journal Article
- Title:
- Estimating gene expression from high-dimensional DNA methylation levels in cancer data: A bimodal unsupervised dimension reduction algorithm. (April 2019)
- Main Title:
- Estimating gene expression from high-dimensional DNA methylation levels in cancer data: A bimodal unsupervised dimension reduction algorithm
- Authors:
- Damgacioglu, Haluk
Celik, Emrah
Celik, Nurcin - Abstract:
- Highlights: A bi-modal dimension reduction is proposed for unsupervised methylation analysis → 83 characters. A curve-fitting based filtering is developed to identify informative biomarkers → 82 characters. Pearson correlation is used to eliminate redundant biomarkers → 65 characters. The performance of our algorithm is shown using nine real cancer datasets → 76 characters. Our algorithm outperforms two benchmark algorithms in eight datasets out of nine → 83 characters. Abstract: Recent molecular and genetic studies have revealed the importance of DNA methylation, a key epigenetic mark, in regulating gene expression and the abnormal profiles of DNA methylation in various diseases including cancer. Here, unsupervised learning methods that are geared towards high-throughput DNA methylation analysis are used to extract useful information from high-dimensional genome wide methylation data in order to provide crucial insights for accurate early diagnosis and treatment of cancer. Herein, these methods are highly dependent on the performance of an earlier step of dimension reduction that aims to find the best subset of attributes to be retained for learning. Widely used algorithms in the literature commonly suffer from resulting in trivial cluster structures and failing to shed light on the relationship between DNA methylation and cancer types due to their myopic and arbitrary search mechanisms. Addressing this issue, we introduce a bimodal unsupervised dimension reductionHighlights: A bi-modal dimension reduction is proposed for unsupervised methylation analysis → 83 characters. A curve-fitting based filtering is developed to identify informative biomarkers → 82 characters. Pearson correlation is used to eliminate redundant biomarkers → 65 characters. The performance of our algorithm is shown using nine real cancer datasets → 76 characters. Our algorithm outperforms two benchmark algorithms in eight datasets out of nine → 83 characters. Abstract: Recent molecular and genetic studies have revealed the importance of DNA methylation, a key epigenetic mark, in regulating gene expression and the abnormal profiles of DNA methylation in various diseases including cancer. Here, unsupervised learning methods that are geared towards high-throughput DNA methylation analysis are used to extract useful information from high-dimensional genome wide methylation data in order to provide crucial insights for accurate early diagnosis and treatment of cancer. Herein, these methods are highly dependent on the performance of an earlier step of dimension reduction that aims to find the best subset of attributes to be retained for learning. Widely used algorithms in the literature commonly suffer from resulting in trivial cluster structures and failing to shed light on the relationship between DNA methylation and cancer types due to their myopic and arbitrary search mechanisms. Addressing this issue, we introduce a bimodal unsupervised dimension reduction algorithm (BOUNDER) that identifies the best subset of loci for downstream analysis considering the variability and redundancy across all the samples using bimodal modeling before it feeds into the learning method. BOUNDER models each locus as a bimodal representation using a piecewise linear function with two segments and filters the informative loci based on the fitted line characteristics. To the best of our knowledge, the work presented here is the first study that uses bimodal modeling in unsupervised learning in DNA methylation analysis. BOUNDER is tailored for DNA methylation analysis using a detailed parameter tuning analysis. The performance of BOUNDER is benchmarked against those of widely used conventional algorithms using real lung, breast, kidney, and urological cancer datasets obtained from Gene Expression Omnibus in terms of their accuracies in hierarchical clustering and k-means clustering. Computational experiments reveal that BOUNDER outperforms the PCA and filtering based approach by providing the highest accuracy in 6 out of 9 datasets while providing more interpretable results through a correlation analysis. The BOUNDER algorithm is also shown to be more robust when compared to multiple other conventional dimension reduction algorithms across different datasets. … (more)
- Is Part Of:
- Computers & industrial engineering. Volume 130(2019)
- Journal:
- Computers & industrial engineering
- Issue:
- Volume 130(2019)
- Issue Display:
- Volume 130, Issue 2019 (2019)
- Year:
- 2019
- Volume:
- 130
- Issue:
- 2019
- Issue Sort Value:
- 2019-0130-2019-0000
- Page Start:
- 348
- Page End:
- 357
- Publication Date:
- 2019-04
- Subjects:
- DNA methylation data -- Values -- Dimension reduction algorithm -- Big data analytics -- Beta distribution -- Piece linear curve fitting
Engineering -- Data processing -- Periodicals
Industrial engineering -- Periodicals
620.00285 - Journal URLs:
- http://www.sciencedirect.com/science/journal/03608352 ↗
http://www.elsevier.com/journals ↗ - DOI:
- 10.1016/j.cie.2019.02.038 ↗
- Languages:
- English
- ISSNs:
- 0360-8352
- Deposit Type:
- Legaldeposit
- View Content:
- Available online (eLD content is only available in our Reading Rooms) ↗
- Physical Locations:
- British Library DSC - 3394.713000
British Library DSC - BLDSS-3PM
British Library HMNTS - ELD Digital store - Ingest File:
- 9839.xml