Sparse Proximal Support Vector Machines for feature selection in high dimensional datasets. Issue 23 (15th December 2015)
- Record Type:
- Journal Article
- Title:
- Sparse Proximal Support Vector Machines for feature selection in high dimensional datasets. Issue 23 (15th December 2015)
- Main Title:
- Sparse Proximal Support Vector Machines for feature selection in high dimensional datasets
- Authors:
- Pappu, Vijay
Panagopoulos, Orestis P.
Xanthopoulos, Petros
Pardalos, Panos M. - Abstract:
- Highlights: Sparse Proximal Support Vector Machines is an embedded feature selection method. sPSVMs removes more than 98% of features in many high dimensional datasets. An efficient alternating optimization technique is proposed. sPSVMs induces class-specific local sparsity. Abstract: Classification of High Dimension Low Sample Size (HDLSS) datasets is a challenging task in supervised learning. Such datasets are prevalent in various areas including biomedical applications and business analytics. In this paper, a new embedded feature selection method for HDLSS datasets is introduced by incorporating sparsity in Proximal Support Vector Machines (PSVMs). Our method, called Sparse Proximal Support Vector Machines (sPSVMs), learns a sparse representation of PSVMs by first casting it as an equivalent least squares problem and then introducing the l 1 -norm for sparsity. An efficient algorithm based on alternating optimization techniques is proposed. sPSVMs remove more than 98% of features in many high dimensional datasets without compromising on generalization performance. Stability in the feature selection process of sPSVMs is also studied and compared with other univariate filter techniques. Additionally, sPSVMs offer the advantage of interpreting the selected features in the context of the classes by inducing class-specific local sparsity instead of global sparsity like other embedded methods. sPSVMs appear to be robust with respect to data dimensionality. Moreover, sPSVMs areHighlights: Sparse Proximal Support Vector Machines is an embedded feature selection method. sPSVMs removes more than 98% of features in many high dimensional datasets. An efficient alternating optimization technique is proposed. sPSVMs induces class-specific local sparsity. Abstract: Classification of High Dimension Low Sample Size (HDLSS) datasets is a challenging task in supervised learning. Such datasets are prevalent in various areas including biomedical applications and business analytics. In this paper, a new embedded feature selection method for HDLSS datasets is introduced by incorporating sparsity in Proximal Support Vector Machines (PSVMs). Our method, called Sparse Proximal Support Vector Machines (sPSVMs), learns a sparse representation of PSVMs by first casting it as an equivalent least squares problem and then introducing the l 1 -norm for sparsity. An efficient algorithm based on alternating optimization techniques is proposed. sPSVMs remove more than 98% of features in many high dimensional datasets without compromising on generalization performance. Stability in the feature selection process of sPSVMs is also studied and compared with other univariate filter techniques. Additionally, sPSVMs offer the advantage of interpreting the selected features in the context of the classes by inducing class-specific local sparsity instead of global sparsity like other embedded methods. sPSVMs appear to be robust with respect to data dimensionality. Moreover, sPSVMs are able to perform feature selection and classification in one step, eliminating the need for dimensionality reduction on the data. To that end, sPSVMs can be used for preprocessing free classification tasks. … (more)
- Is Part Of:
- Expert systems with applications. Volume 42:Issue 23(2015)
- Journal:
- Expert systems with applications
- Issue:
- Volume 42:Issue 23(2015)
- Issue Display:
- Volume 42, Issue 23 (2015)
- Year:
- 2015
- Volume:
- 42
- Issue:
- 23
- Issue Sort Value:
- 2015-0042-0023-0000
- Page Start:
- 9183
- Page End:
- 9191
- Publication Date:
- 2015-12-15
- Subjects:
- Embedded feature selection -- Sparsity -- Regularization -- Class-specific feature selection -- High dimensional datasets
Expert systems (Computer science) -- Periodicals
Systèmes experts (Informatique) -- Périodiques
Electronic journals
006.33 - Journal URLs:
- http://www.sciencedirect.com/science/journal/09574174 ↗
http://www.elsevier.com/journals ↗ - DOI:
- 10.1016/j.eswa.2015.08.022 ↗
- Languages:
- English
- ISSNs:
- 0957-4174
- Deposit Type:
- Legaldeposit
- View Content:
- Available online (eLD content is only available in our Reading Rooms) ↗
- Physical Locations:
- British Library DSC - 3842.004220
British Library DSC - BLDSS-3PM
British Library HMNTS - ELD Digital store - Ingest File:
- 8959.xml