Performance of methods that separate common and distinct variation in multiple data blocks. (9th October 2018)
- Record Type:
- Journal Article
- Title:
- Performance of methods that separate common and distinct variation in multiple data blocks. (9th October 2018)
- Main Title:
- Performance of methods that separate common and distinct variation in multiple data blocks
- Authors:
- Måge, Ingrid
Smilde, Age K.
van der Kloet, Frans M. - Abstract:
- Abstract: In many areas of science, multiple sets of data are collected from the samples. Such data sets can be analyzed by multiblock (or data fusion) methods. The aim is usually to get a holistic understanding of the system or better prediction of some response. Lately, several scientific groups have developed methods for separating common and distinct variation between multiple data blocks. Although the objective is the same, the strategies and algorithms are completely different for these methods. In this paper, we investigate the practical properties of the four most popular methods for separating common and distinct variation: JIVE, DISCO, PCA‐GCA, and OnPLS. The main barrier complicating the use of any of these methods is model selection and validation. Especially when the numbers of blocks is more than two. By the use of extensive simulations, we have elucidated the three properties that are important for assessing the validity of the results: The ability to identify the correct model, the ability to estimate the true, underlying subspaces, and the robustness towards misspecification of the model. The simulated data sets mimic a range of "real life" data, with different dimensionalities and variance structures. We are thus able to identify which methods work best for different types of data structures, and pinpoint weak spots for each method. The results show that PCA‐GCA works best for model selection, while JIVE and DISCO give the best estimates of the subspacesAbstract: In many areas of science, multiple sets of data are collected from the samples. Such data sets can be analyzed by multiblock (or data fusion) methods. The aim is usually to get a holistic understanding of the system or better prediction of some response. Lately, several scientific groups have developed methods for separating common and distinct variation between multiple data blocks. Although the objective is the same, the strategies and algorithms are completely different for these methods. In this paper, we investigate the practical properties of the four most popular methods for separating common and distinct variation: JIVE, DISCO, PCA‐GCA, and OnPLS. The main barrier complicating the use of any of these methods is model selection and validation. Especially when the numbers of blocks is more than two. By the use of extensive simulations, we have elucidated the three properties that are important for assessing the validity of the results: The ability to identify the correct model, the ability to estimate the true, underlying subspaces, and the robustness towards misspecification of the model. The simulated data sets mimic a range of "real life" data, with different dimensionalities and variance structures. We are thus able to identify which methods work best for different types of data structures, and pinpoint weak spots for each method. The results show that PCA‐GCA works best for model selection, while JIVE and DISCO give the best estimates of the subspaces and are most robust towards model misspecification. Abstract : In this work, we compare the four most popular methods for separating common and distinct variation: JIVE, DISCO, PCA‐GCA, and OnPLS. By the use of extensive simulations, we investigate three important properties: the ability to identify the correct model, the ability to estimate the true, underlying subspaces, and the robustness towards model misspecification. The results show that PCA‐GCA works best for model selection, while JIVE and DISCO are most robust and give the best estimates of the subspaces. … (more)
- Is Part Of:
- Journal of chemometrics. Volume 33:Number 1(2019)
- Journal:
- Journal of chemometrics
- Issue:
- Volume 33:Number 1(2019)
- Issue Display:
- Volume 33, Issue 1 (2019)
- Year:
- 2019
- Volume:
- 33
- Issue:
- 1
- Issue Sort Value:
- 2019-0033-0001-0000
- Page Start:
- n/a
- Page End:
- n/a
- Publication Date:
- 2018-10-09
- Subjects:
- DISCO -- JIVE -- multiblock -- OnPLS -- PCA‐GCA
Chemistry -- Mathematics -- Periodicals
Chemistry -- Statistical methods -- Periodicals
542.85 - Journal URLs:
- http://onlinelibrary.wiley.com/ ↗
- DOI:
- 10.1002/cem.3085 ↗
- Languages:
- English
- ISSNs:
- 0886-9383
- Deposit Type:
- Legaldeposit
- View Content:
- Available online (eLD content is only available in our Reading Rooms) ↗
- Physical Locations:
- British Library DSC - 4957.380000
British Library DSC - BLDSS-3PM
British Library HMNTS - ELD Digital store - Ingest File:
- 9447.xml