PUMA: Parallel subspace clustering of categorical data using multi-attribute weights. (15th July 2019)
- Record Type:
- Journal Article
- Title:
- PUMA: Parallel subspace clustering of categorical data using multi-attribute weights. (15th July 2019)
- Main Title:
- PUMA: Parallel subspace clustering of categorical data using multi-attribute weights
- Authors:
- Pang, Ning
Zhang, Jifu
Zhang, Chaowei
Qin, Xiao
Cai, Jianghui - Abstract:
- Highlights: We propose a way of calculating attribute weights through co-occurrence probabilities of attribute values among multiple dimensions. We design a subspace clustering algorithm driven by co-occurrence frequencies of multiple attributes of categorical data. We implement the two-stage clustering algorithm using the MapReduce programming model. Abstract: There are two main reasons why traditional clustering schemes are incompetent for high-dimensional categorical data. First, traditional methods usually represent each cluster by all dimensions without difference; and second, traditional clustering methods only rely on an individual dimension of projection as an attribute's weight ignoring relevance among attributes. We solve these two problems by a MapReduce-based subspace clustering algorithm (called PUMA ) using multi-attribute weights. The attribute subspaces are constructed in our PUMA by calculating an attribute-value weight based on the co-occurrence probability of attribute values among different dimensions. PUMA obtains sub-clusters corresponding to respective attribute subspaces from each computing node in parallel. Lastly, PUMA measures various scale clusters by applying the hierarchical clustering method to iteratively merge sub-clusters. We implement PUMA on a 24-node Hadoop cluster. Experimental results reveal that using multi-attribute weights with subspace clustering can achieve better clustering accuracy on both synthetic and real-world highHighlights: We propose a way of calculating attribute weights through co-occurrence probabilities of attribute values among multiple dimensions. We design a subspace clustering algorithm driven by co-occurrence frequencies of multiple attributes of categorical data. We implement the two-stage clustering algorithm using the MapReduce programming model. Abstract: There are two main reasons why traditional clustering schemes are incompetent for high-dimensional categorical data. First, traditional methods usually represent each cluster by all dimensions without difference; and second, traditional clustering methods only rely on an individual dimension of projection as an attribute's weight ignoring relevance among attributes. We solve these two problems by a MapReduce-based subspace clustering algorithm (called PUMA ) using multi-attribute weights. The attribute subspaces are constructed in our PUMA by calculating an attribute-value weight based on the co-occurrence probability of attribute values among different dimensions. PUMA obtains sub-clusters corresponding to respective attribute subspaces from each computing node in parallel. Lastly, PUMA measures various scale clusters by applying the hierarchical clustering method to iteratively merge sub-clusters. We implement PUMA on a 24-node Hadoop cluster. Experimental results reveal that using multi-attribute weights with subspace clustering can achieve better clustering accuracy on both synthetic and real-world high dimensional datasets. Experimental results also show that PUMA achieves high performance in terms of extensibility, scalability and the nearly linear speedup with respect to number of nodes. Additionally, experimental results demonstrate that PUMA is reasonable, effective, and practical to expert systems such as knowledge acquisition, word sense disambiguation, automatic abstracting and recommender systems. … (more)
- Is Part Of:
- Expert systems with applications. Volume 126(2019)
- Journal:
- Expert systems with applications
- Issue:
- Volume 126(2019)
- Issue Display:
- Volume 126, Issue 2019 (2019)
- Year:
- 2019
- Volume:
- 126
- Issue:
- 2019
- Issue Sort Value:
- 2019-0126-2019-0000
- Page Start:
- 233
- Page End:
- 245
- Publication Date:
- 2019-07-15
- Subjects:
- Parallel subspace clustering -- Multi-attribute weights -- High dimension -- Categorical data -- MapReduce
Expert systems (Computer science) -- Periodicals
Systèmes experts (Informatique) -- Périodiques
Electronic journals
006.33 - Journal URLs:
- http://www.sciencedirect.com/science/journal/09574174 ↗
http://www.elsevier.com/journals ↗ - DOI:
- 10.1016/j.eswa.2019.02.030 ↗
- Languages:
- English
- ISSNs:
- 0957-4174
- Deposit Type:
- Legaldeposit
- View Content:
- Available online (eLD content is only available in our Reading Rooms) ↗
- Physical Locations:
- British Library DSC - 3842.004220
British Library DSC - BLDSS-3PM
British Library HMNTS - ELD Digital store - Ingest File:
- 9682.xml