Interpretable classification models for recidivism prediction. (5th September 2016)
- Record Type:
- Journal Article
- Title:
- Interpretable classification models for recidivism prediction. (5th September 2016)
- Main Title:
- Interpretable classification models for recidivism prediction
- Authors:
- Zeng, Jiaming
Ustun, Berk
Rudin, Cynthia - Abstract:
- Summary: We investigate a long‐debated question, which is how to create predictive models of recidivism that are sufficiently accurate, transparent and interpretable to use for decision making. This question is complicated as these models are used to support different decisions, from sentencing, to determining release on probation to allocating preventative social services. Each case might have an objective other than classification accuracy, such as a desired true positive rate TPR or false positive rate FPR. Each (TPR, FPR) pair is a point on the receiver operator characteristic (ROC) curve. We use popular machine learning methods to create models along the full ROC curve on a wide range of recidivism prediction problems. We show that many methods (support vector machines, stochastic gradient boosting and ridge regression) produce equally accurate models along the full ROC curve. However, methods that are designed for interpretability (classification and regression trees and C5.0) cannot be tuned to produce models that are accurate and/or interpretable. To handle this shortcoming, we use a recent method called supersparse linear integer models to produce accurate, transparent and interpretable scoring systems along the full ROC curve. These scoring systems can be used for decision making for many different use cases, since they are just as accurate as the most powerful black box machine learning models for many applications, but completely transparent, and highlySummary: We investigate a long‐debated question, which is how to create predictive models of recidivism that are sufficiently accurate, transparent and interpretable to use for decision making. This question is complicated as these models are used to support different decisions, from sentencing, to determining release on probation to allocating preventative social services. Each case might have an objective other than classification accuracy, such as a desired true positive rate TPR or false positive rate FPR. Each (TPR, FPR) pair is a point on the receiver operator characteristic (ROC) curve. We use popular machine learning methods to create models along the full ROC curve on a wide range of recidivism prediction problems. We show that many methods (support vector machines, stochastic gradient boosting and ridge regression) produce equally accurate models along the full ROC curve. However, methods that are designed for interpretability (classification and regression trees and C5.0) cannot be tuned to produce models that are accurate and/or interpretable. To handle this shortcoming, we use a recent method called supersparse linear integer models to produce accurate, transparent and interpretable scoring systems along the full ROC curve. These scoring systems can be used for decision making for many different use cases, since they are just as accurate as the most powerful black box machine learning models for many applications, but completely transparent, and highly interpretable. … (more)
- Is Part Of:
- Journal of the Royal Statistical Society. Volume 180:Number 3(2017:Jul.)
- Journal:
- Journal of the Royal Statistical Society
- Issue:
- Volume 180:Number 3(2017:Jul.)
- Issue Display:
- Volume 180, Issue 3 (2017)
- Year:
- 2017
- Volume:
- 180
- Issue:
- 3
- Issue Sort Value:
- 2017-0180-0003-0000
- Page Start:
- 689
- Page End:
- 722
- Publication Date:
- 2016-09-05
- Subjects:
- Binary classification -- Interpretability -- Machine learning -- Recidivism -- Scoring systems
Social sciences -- Statistical methods -- Periodicals
Statistics -- Periodicals
300.15195 - Journal URLs:
- http://rss.onlinelibrary.wiley.com/hub/journal/10.1111/(ISSN)1467-985X/ ↗
https://academic.oup.com/jrsssa ↗
http://onlinelibrary.wiley.com/ ↗ - DOI:
- 10.1111/rssa.12227 ↗
- Languages:
- English
- ISSNs:
- 0964-1998
- Deposit Type:
- Legaldeposit
- View Content:
- Available online (eLD content is only available in our Reading Rooms) ↗
- Physical Locations:
- British Library DSC - 4866.000000
British Library DSC - BLDSS-3PM
British Library STI - ELD Digital store - Ingest File:
- 1311.xml