Improving spherical k-means for document clustering: Fast initialization, sparse centroid projection, and efficient cluster labeling. (15th July 2020)
- Record Type:
- Journal Article
- Title:
- Improving spherical k-means for document clustering: Fast initialization, sparse centroid projection, and efficient cluster labeling. (15th July 2020)
- Main Title:
- Improving spherical k-means for document clustering: Fast initialization, sparse centroid projection, and efficient cluster labeling
- Authors:
- Kim, Hyunjoong
Kim, Han Kyul
Cho, Sungzoon - Abstract:
- Highlights: Spherical k-means for document clustering is improved to overcome its weaknesses. Our method ensures dispersed initial points with faster computation time. Our method preserves sparsity of centroid vectors for better interpretability. We provide unsupervised document cluster labeling method. Abstract: Due to its simplicity and intuitive interpretability, spherical k-means is often used for clustering a large number of documents. However, there exist a number of drawbacks that need to be addressed for much effective document clustering. Without well-dispersed initial points, spherical k-means fails to converge quickly, which is critical for clustering a large number of documents. Furthermore, its dense centroid vectors needlessly incorporate the impact of infrequent and less-informative words, thereby distorting the distance calculation between the document vectors. In this paper, we propose practical improvements on spherical k-means to overcome these issues during document clustering. Our proposed initialization method not only guarantees dispersed initial points, but is also up to 1000 times faster than previously well-known initialization method such as k-means++. Furthermore, we enforce sparsity on the centroid vectors by using a data-driven threshold that is capable of dynamically adjusting its value depending on the clusters. Additionally, we propose an unsupervised cluster labeling method that effectively extracts meaningful keywords to describe eachHighlights: Spherical k-means for document clustering is improved to overcome its weaknesses. Our method ensures dispersed initial points with faster computation time. Our method preserves sparsity of centroid vectors for better interpretability. We provide unsupervised document cluster labeling method. Abstract: Due to its simplicity and intuitive interpretability, spherical k-means is often used for clustering a large number of documents. However, there exist a number of drawbacks that need to be addressed for much effective document clustering. Without well-dispersed initial points, spherical k-means fails to converge quickly, which is critical for clustering a large number of documents. Furthermore, its dense centroid vectors needlessly incorporate the impact of infrequent and less-informative words, thereby distorting the distance calculation between the document vectors. In this paper, we propose practical improvements on spherical k-means to overcome these issues during document clustering. Our proposed initialization method not only guarantees dispersed initial points, but is also up to 1000 times faster than previously well-known initialization method such as k-means++. Furthermore, we enforce sparsity on the centroid vectors by using a data-driven threshold that is capable of dynamically adjusting its value depending on the clusters. Additionally, we propose an unsupervised cluster labeling method that effectively extracts meaningful keywords to describe each cluster. We have tested our improvements on seven different text datasets that include both new and publicly available datasets. Based on our experiments on these datasets, we have found that our proposed improvements successfully overcome the drawbacks of spherical k-means in significantly reduced computation time. Furthermore, we have qualitatively verified the performance of the proposed cluster labeling method by extracting descriptive keywords of the clusters from these datasets. … (more)
- Is Part Of:
- Expert systems with applications. Volume 150(2020)
- Journal:
- Expert systems with applications
- Issue:
- Volume 150(2020)
- Issue Display:
- Volume 150, Issue 2020 (2020)
- Year:
- 2020
- Volume:
- 150
- Issue:
- 2020
- Issue Sort Value:
- 2020-0150-2020-0000
- Page Start:
- Page End:
- Publication Date:
- 2020-07-15
- Subjects:
- Spherical k-means -- Document clustering -- k-means initialization -- Sparse vector projection -- Clustering labeling
Expert systems (Computer science) -- Periodicals
Systèmes experts (Informatique) -- Périodiques
Electronic journals
006.33 - Journal URLs:
- http://www.sciencedirect.com/science/journal/09574174 ↗
http://www.elsevier.com/journals ↗ - DOI:
- 10.1016/j.eswa.2020.113288 ↗
- Languages:
- English
- ISSNs:
- 0957-4174
- Deposit Type:
- Legaldeposit
- View Content:
- Available online (eLD content is only available in our Reading Rooms) ↗
- Physical Locations:
- British Library DSC - 3842.004220
British Library DSC - BLDSS-3PM
British Library HMNTS - ELD Digital store - Ingest File:
- 13500.xml