Multivariate feature selection and autoencoder embeddings of ovarian cancer clinical and genetic data. (15th November 2022)
- Record Type:
- Journal Article
- Title:
- Multivariate feature selection and autoencoder embeddings of ovarian cancer clinical and genetic data. (15th November 2022)
- Main Title:
- Multivariate feature selection and autoencoder embeddings of ovarian cancer clinical and genetic data
- Authors:
- Bote-Curiel, Luis
Ruiz-Llorente, Sergio
Muñoz-Romero, Sergio
Yagüe-Fernández, Mónica
Barquín, Arantzazu
García-Donas, Jesús
Rojo-Álvarez, José Luis - Abstract:
- Abstract: Although certain genetic alterations have been defined as predictive and prognostic biomarkers in the context of ovarian cancer (OC), data science methods represent alternative approaches to identify novel correlations and define relevant markers in these gynecological tumors. Considering this potential, our work focused both on clinical and genomic data information collected from patients with OC to identify relationships between clinical and genetic factors and disease progression-related variables. For this aim, we proposed two analyses: (1) a nonlinear exploration of an OC dataset using autoencoders, a type of neural network that can be used as a feature extraction tool to represent a dataset in 3-dimensional latent space, so that we could assess whether there are intrinsic or natural nonlinear separability patterns between disease progression groups (in our case, platinum-sensitive and platinum-resistant groups); and (2) the identification of relevant variable relationships by means of an adaptation of the informative variable identifier (IVI), a feature selection method that labels each input feature as informative or noisy with respect to the task at hand, identifies the relationships among features, and builds a ranking of features, allowing us to study which input features and relationships may be most informative for the OC disease progression classification to define new biomarkers involved in disease progression. Our interest has been in clinical andAbstract: Although certain genetic alterations have been defined as predictive and prognostic biomarkers in the context of ovarian cancer (OC), data science methods represent alternative approaches to identify novel correlations and define relevant markers in these gynecological tumors. Considering this potential, our work focused both on clinical and genomic data information collected from patients with OC to identify relationships between clinical and genetic factors and disease progression-related variables. For this aim, we proposed two analyses: (1) a nonlinear exploration of an OC dataset using autoencoders, a type of neural network that can be used as a feature extraction tool to represent a dataset in 3-dimensional latent space, so that we could assess whether there are intrinsic or natural nonlinear separability patterns between disease progression groups (in our case, platinum-sensitive and platinum-resistant groups); and (2) the identification of relevant variable relationships by means of an adaptation of the informative variable identifier (IVI), a feature selection method that labels each input feature as informative or noisy with respect to the task at hand, identifies the relationships among features, and builds a ranking of features, allowing us to study which input features and relationships may be most informative for the OC disease progression classification to define new biomarkers involved in disease progression. Our interest has been in clinical and genetic factors and in the combination of clinical features and genetic profile. Results with autoencoders suggest a pattern of separability between disease progression groups in the clinical part and for the combination of genes and clinical features of the OC dataset, that is increased via supervised fine tuning. In the genetic part, this pattern of separability is not observed, but it is more defined when a supervised fine tuning is performed. Results of the IVI-mediated feature selection method show significance for relevant clinical variables (such as type of surgery and neoadjuvant chemotherapy), some mutation genes, and low-risk genetic features. These results highlight the efficacy of the considered approaches to better understand the clinical course of OC. Highlights: Data science methods are suitable for identifying biomarkers in OC. Feature selection methods show predictive roles of variables in an OC dataset. Features extraction methods reveal some patterns of separability in an OC dataset. … (more)
- Is Part Of:
- Expert systems with applications. Volume 206(2022)
- Journal:
- Expert systems with applications
- Issue:
- Volume 206(2022)
- Issue Display:
- Volume 206, Issue 2022 (2022)
- Year:
- 2022
- Volume:
- 206
- Issue:
- 2022
- Issue Sort Value:
- 2022-0206-2022-0000
- Page Start:
- Page End:
- Publication Date:
- 2022-11-15
- Subjects:
- Autoencoders -- Clinical data -- Feature extraction -- Feature selection -- Genetic data -- Ovarian cancer -- Platinum-resistant -- Platinum-sensitive
Expert systems (Computer science) -- Periodicals
Systèmes experts (Informatique) -- Périodiques
Electronic journals
006.33 - Journal URLs:
- http://www.sciencedirect.com/science/journal/09574174 ↗
http://www.elsevier.com/journals ↗ - DOI:
- 10.1016/j.eswa.2022.117865 ↗
- Languages:
- English
- ISSNs:
- 0957-4174
- Deposit Type:
- Legaldeposit
- View Content:
- Available online (eLD content is only available in our Reading Rooms) ↗
- Physical Locations:
- British Library DSC - 3842.004220
British Library DSC - BLDSS-3PM
British Library HMNTS - ELD Digital store - Ingest File:
- 23554.xml