Ground truth bias in external cluster validity indices. (May 2017)
- Record Type:
- Journal Article
- Title:
- Ground truth bias in external cluster validity indices. (May 2017)
- Main Title:
- Ground truth bias in external cluster validity indices
- Authors:
- Lei, Yang
Bezdek, James C.
Romano, Simone
Vinh, Nguyen Xuan
Chan, Jeffrey
Bailey, James - Abstract:
- Abstract: External cluster validity indices (CVIs) are used to quantify the quality of a clustering by comparing the similarity between the clustering and a ground truth partition. However, some external CVIs show a biased behavior when selecting the most similar clustering. Users may consequently be misguided by such results. Recognizing and understanding the bias behavior of CVIs is therefore crucial. It has been noticed that, some external CVIs exhibit a preferential bias towards a larger or smaller number of clusters which is monotonic (directly or inversely) in the number of clusters in candidate partitions. This type of bias is caused by the functional form of the CVI model. For example, the popular Rand Index (RI) exhibits a monotone increasing (NCinc) bias, while the Jaccard Index (JI) index suffers from a monotone decreasing (NCdec) bias. This type of bias has been previously recognized in the literature. In this work, we identify a new type of bias arising from the distribution of the ground truth (reference) partition against which candidate partitions are compared. We call this new type of bias ground truth (GT) bias. This type of bias occurs if a change in the reference partition causes a change in the bias status (e.g., NCinc, NCdec) of a CVI. For example, NCinc bias in the RI can be changed to NCdec bias by skewing the distribution of clusters in the ground truth partition. It is important for users to be aware of this new type of biased behavior, since it mayAbstract: External cluster validity indices (CVIs) are used to quantify the quality of a clustering by comparing the similarity between the clustering and a ground truth partition. However, some external CVIs show a biased behavior when selecting the most similar clustering. Users may consequently be misguided by such results. Recognizing and understanding the bias behavior of CVIs is therefore crucial. It has been noticed that, some external CVIs exhibit a preferential bias towards a larger or smaller number of clusters which is monotonic (directly or inversely) in the number of clusters in candidate partitions. This type of bias is caused by the functional form of the CVI model. For example, the popular Rand Index (RI) exhibits a monotone increasing (NCinc) bias, while the Jaccard Index (JI) index suffers from a monotone decreasing (NCdec) bias. This type of bias has been previously recognized in the literature. In this work, we identify a new type of bias arising from the distribution of the ground truth (reference) partition against which candidate partitions are compared. We call this new type of bias ground truth (GT) bias. This type of bias occurs if a change in the reference partition causes a change in the bias status (e.g., NCinc, NCdec) of a CVI. For example, NCinc bias in the RI can be changed to NCdec bias by skewing the distribution of clusters in the ground truth partition. It is important for users to be aware of this new type of biased behavior, since it may affect the interpretations of CVI results. The objective of this article is to study the empirical and theoretical implications of GT bias. To the best of our knowledge, this is the first extensive study of such a property for external CVIs. Our computational experiments show that 5 of 26 pair-counting based CVIs studied in this paper, which are all functions of the RI, exhibit GT bias. Following the numerical examples, we provide a theoretical analysis of GT bias based on the relationship between the RI and quadratic entropy. Specifically, we prove that the quadratic entropy of the ground truth partition provides a computable test which predicts the NC bias status of the RI. Highlights: We identify the GT bias effect for external validation measures, and explain its importance. We test and discuss NC bias for 26 popular pair-counting based validation measures. We prove that the RI and related 4 indices suffer from GT bias. We provide theoretical explanations for understanding why and when GT bias happens. We present experimental results that support our analysis. We present an empirical example to show that the ARI also suffers from a modified GT bias. … (more)
- Is Part Of:
- Pattern recognition. Volume 65(2017:May)
- Journal:
- Pattern recognition
- Issue:
- Volume 65(2017:May)
- Issue Display:
- Volume 65 (2017)
- Year:
- 2017
- Volume:
- 65
- Issue Sort Value:
- 2017-0065-0000-0000
- Page Start:
- 58
- Page End:
- 70
- Publication Date:
- 2017-05
- Subjects:
- External cluster validity indices -- Rand index -- Ground truth bias -- Quadratic entropy
Pattern perception -- Periodicals
Perception des structures -- Périodiques
Patroonherkenning
006.4 - Journal URLs:
- http://www.sciencedirect.com/science/journal/00313203 ↗
http://www.sciencedirect.com/ ↗ - DOI:
- 10.1016/j.patcog.2016.12.003 ↗
- Languages:
- English
- ISSNs:
- 0031-3203
- Deposit Type:
- Legaldeposit
- View Content:
- Available online (eLD content is only available in our Reading Rooms) ↗
- Physical Locations:
- British Library DSC - BLDSS-3PM
British Library HMNTS - ELD Digital store - Ingest File:
- 7658.xml