Is poverty predictable with machine learning? A study of DHS data from Kyrgyzstan. (June 2022)
- Record Type:
- Journal Article
- Title:
- Is poverty predictable with machine learning? A study of DHS data from Kyrgyzstan. (June 2022)
- Main Title:
- Is poverty predictable with machine learning? A study of DHS data from Kyrgyzstan
- Authors:
- Li, Qing
Yu, Shuai
Échevin, Damien
Fan, Min - Abstract:
- Abstract: A prerequisite for eliminating poverty is to accurately identify and target the households in poverty. While some factors such as asset holdings are well recognized as relevant for assessing and predicting poverty, a priori selected indicators are not sufficient conditions for poverty and the key factors may vary from one case to another. Researchers have begun to apply machine learning algorithms to predict poor households. This paper uses the accuracy of prediction as the standard to study the application of machine learning algorithms. Using the DHS data of 8040 households in Kyrgyzstan, we apply a state-of-the-art algorithm (XGBoost) to explore the full dataset, profiting from the algorithm's ability in handling many variables, and compare the results with the a priori selected variables. We also compare XGBoost with generalized linear model (GLM), the latter being viewed as an approach in between traditional models and modern machine learning algorithms. The results imply that the inclusion of more variables is not necessarily preferable for prediction; a few important variables selected by the algorithms may also perform well. Different algorithms may select different variables as the important ones for prediction. XGBoost performs better than GLM in most cases, and machine learning is useful for variable selection. Additionally, XGBoost is particularly preferable when using a priori variables. Highlights: We apply a state-of-the-art machine learningAbstract: A prerequisite for eliminating poverty is to accurately identify and target the households in poverty. While some factors such as asset holdings are well recognized as relevant for assessing and predicting poverty, a priori selected indicators are not sufficient conditions for poverty and the key factors may vary from one case to another. Researchers have begun to apply machine learning algorithms to predict poor households. This paper uses the accuracy of prediction as the standard to study the application of machine learning algorithms. Using the DHS data of 8040 households in Kyrgyzstan, we apply a state-of-the-art algorithm (XGBoost) to explore the full dataset, profiting from the algorithm's ability in handling many variables, and compare the results with the a priori selected variables. We also compare XGBoost with generalized linear model (GLM), the latter being viewed as an approach in between traditional models and modern machine learning algorithms. The results imply that the inclusion of more variables is not necessarily preferable for prediction; a few important variables selected by the algorithms may also perform well. Different algorithms may select different variables as the important ones for prediction. XGBoost performs better than GLM in most cases, and machine learning is useful for variable selection. Additionally, XGBoost is particularly preferable when using a priori variables. Highlights: We apply a state-of-the-art machine learning algorithm (XGBoost) to predict poverty at household level, making full use of available data to explore possibilities for improvement. We compare the performances of XGBoost with GLM, the latter being viewed as an approach in between traditional models and machine learning algorithms. XGBoost works better than GLM in most cases. It is harmless to include more variables for prediction. Machine learning is useful for variable selection, and XGBoost is particularly preferable for a priori variables. … (more)
- Is Part Of:
- Socio-economic planning sciences. Number 81(2022)
- Journal:
- Socio-economic planning sciences
- Issue:
- Number 81(2022)
- Issue Display:
- Volume 81, Issue 81 (2022)
- Year:
- 2022
- Volume:
- 81
- Issue:
- 81
- Issue Sort Value:
- 2022-0081-0081-0000
- Page Start:
- Page End:
- Publication Date:
- 2022-06
- Subjects:
- Poverty prediction -- Machine learning -- XGBoost -- Generalized linear model
AUC Area Under receiver operating characteristic Curve -- CART Classification And Regression Tree -- DHS Demographic and Health Survey -- GIS Geographic Information Systems -- GLM Generalized Linear Model -- XGBoost eXtreme Gradient Boosting
Planning -- Periodicals
Economic policy -- Periodicals
Social policy -- Periodicals
Planification -- Périodiques
Politique économique -- Périodiques
Politique sociale -- Périodiques
ECONOMIC PLANNING
SOCIAL PLANNING
DECISION-MAKING
361 - Journal URLs:
- http://www.sciencedirect.com/science/journal/00380121 ↗
http://www.elsevier.com/journals ↗ - DOI:
- 10.1016/j.seps.2021.101195 ↗
- Languages:
- English
- ISSNs:
- 0038-0121
- Deposit Type:
- Legaldeposit
- View Content:
- Available online (eLD content is only available in our Reading Rooms) ↗
- Physical Locations:
- British Library DSC - 8319.576000
British Library DSC - BLDSS-3PM
British Library HMNTS - ELD Digital store - Ingest File:
- 21582.xml