RMisbeta: A robust missing value imputation approach in transcriptomics and metabolomics data. (November 2021)
- Record Type:
- Journal Article
- Title:
- RMisbeta: A robust missing value imputation approach in transcriptomics and metabolomics data. (November 2021)
- Main Title:
- RMisbeta: A robust missing value imputation approach in transcriptomics and metabolomics data
- Authors:
- Shahjaman, Md.
Rahman, Md. Rezanur
Islam, Tania
Auwul, Md. Rabiul
Moni, Mohammad Ali
Mollah, Md. Nurul Haque - Abstract:
- Abstract: Transcriptomics and metabolomics data often contain missing values or outliers due to limitations of the data acquisition techniques. Most of the statistical methods require complete datasets for downstream analysis. A number of methods have been developed for missing value imputation using the classical mean and variance based on maximum likelihood estimators, which are not robust against outliers. Consequently, the performance of these methods deteriorates in the presence of outliers. Hence precise imputation of missing values and outliers handling are both concurrently important. Therefore, in this paper, we developed a robust iterative approach using robust estimators based on the minimum beta divergence method, which simultaneously impute missing values and outliers. We investigate the performance of the proposed method in a comparison with six frequently used missing value imputation methods such as Zero, KNN, robust SVD, EM, random forest (RF) and weighted least square approach (WLSA) through feature selection using both simulated and real datasets. Ten performance indices were used to explore the optimal method such as Frobenius norm (FOBN), accuracy (ACC), sensitivity (SN), specificity (SP), positive predictive value (PPV), negative predictive value (NPV), detection rate (DR), misclassification error rate (MER), the area under the ROC curve (AUC) and computational runtime. Evaluation based on both simulated and real data suggests the superiority of theAbstract: Transcriptomics and metabolomics data often contain missing values or outliers due to limitations of the data acquisition techniques. Most of the statistical methods require complete datasets for downstream analysis. A number of methods have been developed for missing value imputation using the classical mean and variance based on maximum likelihood estimators, which are not robust against outliers. Consequently, the performance of these methods deteriorates in the presence of outliers. Hence precise imputation of missing values and outliers handling are both concurrently important. Therefore, in this paper, we developed a robust iterative approach using robust estimators based on the minimum beta divergence method, which simultaneously impute missing values and outliers. We investigate the performance of the proposed method in a comparison with six frequently used missing value imputation methods such as Zero, KNN, robust SVD, EM, random forest (RF) and weighted least square approach (WLSA) through feature selection using both simulated and real datasets. Ten performance indices were used to explore the optimal method such as Frobenius norm (FOBN), accuracy (ACC), sensitivity (SN), specificity (SP), positive predictive value (PPV), negative predictive value (NPV), detection rate (DR), misclassification error rate (MER), the area under the ROC curve (AUC) and computational runtime. Evaluation based on both simulated and real data suggests the superiority of the proposed method over the other traditional methods in terms of various rates of outliers and missing values. The suggested approach also keeps almost equal performance in absence of outliers with the other methods. The proposed method is accurate, simple, and consumes lower computational time compared to the other methods. Therefore, our recommendation is to apply the proposed procedure for large-scale transcriptomics and metabolomics data analysis. The computational tool has been implemented in an R package, which is publicly available from https://CRAN.R-project.org/package=rMisbeta . Highlights: We develop a new robust missing value imputation approach for transcriptomics and metabolomics data analysis. The proposed algorithm has been implemented in a R package available in https://CRAN.R-project.org/package=rMisbeta . From breast cancer dataset rMisbeta identified 6 outlying DEGs that were not detected by the other stat-of arts methods. From GC-MS metabolomics dataset rMisbeta identified 2 additional metabolites that were not detected by the other methods. rMisbeta is accurate, simple and fast, in the sense that it requires lower computational time. … (more)
- Is Part Of:
- Computers in biology and medicine. Volume 138(2021)
- Journal:
- Computers in biology and medicine
- Issue:
- Volume 138(2021)
- Issue Display:
- Volume 138, Issue 2021 (2021)
- Year:
- 2021
- Volume:
- 138
- Issue:
- 2021
- Issue Sort Value:
- 2021-0138-2021-0000
- Page Start:
- Page End:
- Publication Date:
- 2021-11
- Subjects:
- Transcriptomics data -- GC-MS metabolomics Data -- Missing values -- Outliers -- Robustness -- And beta weight function
Medicine -- Data processing -- Periodicals
Biology -- Data processing -- Periodicals
610.285 - Journal URLs:
- http://www.sciencedirect.com/science/journal/00104825/ ↗
http://www.elsevier.com/journals ↗ - DOI:
- 10.1016/j.compbiomed.2021.104911 ↗
- Languages:
- English
- ISSNs:
- 0010-4825
- Deposit Type:
- Legaldeposit
- View Content:
- Available online (eLD content is only available in our Reading Rooms) ↗
- Physical Locations:
- British Library DSC - 3394.880000
British Library DSC - BLDSS-3PM
British Library HMNTS - ELD Digital store - Ingest File:
- 19802.xml