Lessons learned from data stream classification applied to credit scoring. (30th December 2020)
- Record Type:
- Journal Article
- Title:
- Lessons learned from data stream classification applied to credit scoring. (30th December 2020)
- Main Title:
- Lessons learned from data stream classification applied to credit scoring
- Authors:
- Barddal, Jean Paul
Loezer, Lucas
Enembreck, Fabrício
Lanzuolo, Riccardo - Abstract:
- Abstract: The financial credibility of a person is a factor used to determine whether a loan should be approved or not, and this is quantified by a 'credit score, ' which is calculated using a variety of factors, including past performance on debt obligations, profiling, amongst others. Machine learning has been widely applied to automate the development of effective credit scoring models over the years. Yet, studies show that the development of robust credit scoring models may take longer than a year, and thus, if the behavior of customers changes over time, the model will be outdated even before its deployment. In this paper, we made 3 anonymized real-world credit scoring datasets available alongside the results obtained. In each of these datasets, we verify whether the credit scoring task should be thought as an ephemeral scenario since many of the variables may drift over time, and thus, data stream mining techniques should be used since they were tailored for incremental learning and to detect and adapt to changes in the data distribution. Therefore, we compare both traditional batch machine learning algorithms with data stream algorithms in different validation schemes using both Kolmogorov–Smirnov and Population Stability Index metrics. Furthermore, we also provide insights on the importance of features according to their Information Value, Mean Decrease Impurity, and Mean Positional Gain metrics, such that the last depicts changes in the importance of features overAbstract: The financial credibility of a person is a factor used to determine whether a loan should be approved or not, and this is quantified by a 'credit score, ' which is calculated using a variety of factors, including past performance on debt obligations, profiling, amongst others. Machine learning has been widely applied to automate the development of effective credit scoring models over the years. Yet, studies show that the development of robust credit scoring models may take longer than a year, and thus, if the behavior of customers changes over time, the model will be outdated even before its deployment. In this paper, we made 3 anonymized real-world credit scoring datasets available alongside the results obtained. In each of these datasets, we verify whether the credit scoring task should be thought as an ephemeral scenario since many of the variables may drift over time, and thus, data stream mining techniques should be used since they were tailored for incremental learning and to detect and adapt to changes in the data distribution. Therefore, we compare both traditional batch machine learning algorithms with data stream algorithms in different validation schemes using both Kolmogorov–Smirnov and Population Stability Index metrics. Furthermore, we also provide insights on the importance of features according to their Information Value, Mean Decrease Impurity, and Mean Positional Gain metrics, such that the last depicts changes in the importance of features over time. For 2 of the 3 tested datasets, the results obtained by data stream learners are comparable to predictive models currently in use, thus showing the efficiency of data stream classification for the credit scoring task. Highlights: Three real-world anonymized credit scoring datasets are made publicly available. Assessment of batch and data stream mining classifiers in credit scoring datasets. Results show that data stream learners are comparable to predictive models currently in use. … (more)
- Is Part Of:
- Expert systems with applications. Volume 162(2020)
- Journal:
- Expert systems with applications
- Issue:
- Volume 162(2020)
- Issue Display:
- Volume 162, Issue 2020 (2020)
- Year:
- 2020
- Volume:
- 162
- Issue:
- 2020
- Issue Sort Value:
- 2020-0162-2020-0000
- Page Start:
- Page End:
- Publication Date:
- 2020-12-30
- Subjects:
- Credit scoring -- Machine learning datasets -- Data stream classification
Expert systems (Computer science) -- Periodicals
Systèmes experts (Informatique) -- Périodiques
Electronic journals
006.33 - Journal URLs:
- http://www.sciencedirect.com/science/journal/09574174 ↗
http://www.elsevier.com/journals ↗ - DOI:
- 10.1016/j.eswa.2020.113899 ↗
- Languages:
- English
- ISSNs:
- 0957-4174
- Deposit Type:
- Legaldeposit
- View Content:
- Available online (eLD content is only available in our Reading Rooms) ↗
- Physical Locations:
- British Library DSC - 3842.004220
British Library DSC - BLDSS-3PM
British Library HMNTS - ELD Digital store - Ingest File:
- 14542.xml