Constrained Standardization of Count Data from Massive Parallel Sequencing. Issue 11 (28th May 2021)
- Record Type:
- Journal Article
- Title:
- Constrained Standardization of Count Data from Massive Parallel Sequencing. Issue 11 (28th May 2021)
- Main Title:
- Constrained Standardization of Count Data from Massive Parallel Sequencing
- Authors:
- Van Houtven, Joris
Cuypers, Bart
Meysman, Pieter
Hooyberghs, Jef
Laukens, Kris
Valkenborg, Dirk - Abstract:
- Graphical abstract: Highlights: Normalization method for proteomics and transcriptomics (and possibly other -omics). Library size correction, bias removal, magnitude (and thus variance) re-scaling. Performance comparable with DESeq2, but faster and easier to use. Enables experiments designs with many parallel runs, protocols, instruments and labs. Immediately after each instrument run, allowing quick triage and quality control. Abstract: In high-throughput omics disciplines like transcriptomics, researchers face a need to assess the quality of an experiment prior to an in-depth statistical analysis. To efficiently analyze such voluminous collections of data, researchers need triage methods that are both quick and easy to use. Such a normalization method for relative quantitation, CONSTANd, was recently introduced for isobarically-labeled mass spectra in proteomics. It transforms the data matrix of abundances through an iterative, convergent process enforcing three constraints: (I) identical column sums; (II) each row sum is fixed (across matrices) and (III) identical to all other row sums. In this study, we investigate whether CONSTANd is suitable for count data from massively parallel sequencing, by qualitatively comparing its results to those of DESeq2. Further, we propose an adjustment of the method so that it may be applied to identically balanced but differently sized experiments for joint analysis. We find that CONSTANd can process large data sets at well over 1Graphical abstract: Highlights: Normalization method for proteomics and transcriptomics (and possibly other -omics). Library size correction, bias removal, magnitude (and thus variance) re-scaling. Performance comparable with DESeq2, but faster and easier to use. Enables experiments designs with many parallel runs, protocols, instruments and labs. Immediately after each instrument run, allowing quick triage and quality control. Abstract: In high-throughput omics disciplines like transcriptomics, researchers face a need to assess the quality of an experiment prior to an in-depth statistical analysis. To efficiently analyze such voluminous collections of data, researchers need triage methods that are both quick and easy to use. Such a normalization method for relative quantitation, CONSTANd, was recently introduced for isobarically-labeled mass spectra in proteomics. It transforms the data matrix of abundances through an iterative, convergent process enforcing three constraints: (I) identical column sums; (II) each row sum is fixed (across matrices) and (III) identical to all other row sums. In this study, we investigate whether CONSTANd is suitable for count data from massively parallel sequencing, by qualitatively comparing its results to those of DESeq2. Further, we propose an adjustment of the method so that it may be applied to identically balanced but differently sized experiments for joint analysis. We find that CONSTANd can process large data sets at well over 1 million count records per second whilst mitigating unwanted systematic bias and thus quickly uncovering the underlying biological structure when combined with a PCA plot or hierarchical clustering. Moreover, it allows joint analysis of data sets obtained from different batches, with different protocols and from different labs but without exploiting information from the experimental setup other than the delineation of samples into identically processed sets (IPSs). CONSTANd's simplicity and applicability to proteomics as well as transcriptomics data make it an interesting candidate for integration in multi-omics workflows. … (more)
- Is Part Of:
- Journal of molecular biology. Volume 433:Issue 11(2021)
- Journal:
- Journal of molecular biology
- Issue:
- Volume 433:Issue 11(2021)
- Issue Display:
- Volume 433, Issue 11 (2021)
- Year:
- 2021
- Volume:
- 433
- Issue:
- 11
- Issue Sort Value:
- 2021-0433-0011-0000
- Page Start:
- Page End:
- Publication Date:
- 2021-05-28
- Subjects:
- normalization -- RNA-seq -- transcriptomics -- proteomics -- multi-omics
IPS identically processed (sub)set -- IPFP iterative proportional fitting procedure -- GLM generalized linear model
Molecular biology -- Periodicals
Biology -- Periodicals
Biochemistry -- Periodicals
Bacteriology -- Periodicals
Molecular Biology -- Periodicals
Biochemistry -- Periodicals
Biologie moléculaire -- Périodiques
Biologie -- Périodiques
Biochimie -- Périodiques
Moleculaire biologie
Biochemistry
Biology
Molecular biology
Periodicals
572.805 - Journal URLs:
- http://www.sciencedirect.com/science/journal/00222836 ↗
http://www.elsevier.com/journals ↗ - DOI:
- 10.1016/j.jmb.2021.166966 ↗
- Languages:
- English
- ISSNs:
- 0022-2836
- Deposit Type:
- Legaldeposit
- View Content:
- Available online (eLD content is only available in our Reading Rooms) ↗
- Physical Locations:
- British Library DSC - 5020.700000
British Library DSC - BLDSS-3PM
British Library HMNTS - ELD Digital store - Ingest File:
- 16771.xml