A comparison of statistical learning methods for deriving determining factors of accident occurrence from an imbalanced high resolution dataset. (June 2019)
- Record Type:
- Journal Article
- Title:
- A comparison of statistical learning methods for deriving determining factors of accident occurrence from an imbalanced high resolution dataset. (June 2019)
- Main Title:
- A comparison of statistical learning methods for deriving determining factors of accident occurrence from an imbalanced high resolution dataset
- Authors:
- Schlögl, Matthias
Stütz, Rainer
Laaha, Gregor
Melcher, Michael - Abstract:
- Highlights: A comparative assessment of methods for modeling imbalanced data is presented. 45 different variables available at high resolution are considered. Data are balanced using a combination of over- and undersampling techniques. Tree-based methods prevail in terms of model quality. Findings demonstrate merits and challenges of big data in road safety analysis. Abstract: One of the main aims of accident data analysis is to derive the determining factors associated with road traffic accident occurrence. While current studies mainly use variants of count data regression to achieve this aim, the problem can also be considered as a binary classification task, with the dichotomous target variable indicating events (accidents) and non-events (no accidents). The effects of 45 variables – describing road condition and geometry, traffic volume and regulations, weather, and accident time – are analyzed using a dataset in high temporal (1 h) and spatial (250 m) resolution, covering the whole highway network of Austria over the period of four consecutive years. A combination of synthetic minority oversampling and maximum dissimilarity undersampling is used to balance the training dataset. We employ and compare a series of statistical learning techniques with respect to their predictive performance and discuss the importance of determining factors of accident occurrence from the ensemble of models. Findings substantiate that a trade-off between accuracy and sensitivity is inherentHighlights: A comparative assessment of methods for modeling imbalanced data is presented. 45 different variables available at high resolution are considered. Data are balanced using a combination of over- and undersampling techniques. Tree-based methods prevail in terms of model quality. Findings demonstrate merits and challenges of big data in road safety analysis. Abstract: One of the main aims of accident data analysis is to derive the determining factors associated with road traffic accident occurrence. While current studies mainly use variants of count data regression to achieve this aim, the problem can also be considered as a binary classification task, with the dichotomous target variable indicating events (accidents) and non-events (no accidents). The effects of 45 variables – describing road condition and geometry, traffic volume and regulations, weather, and accident time – are analyzed using a dataset in high temporal (1 h) and spatial (250 m) resolution, covering the whole highway network of Austria over the period of four consecutive years. A combination of synthetic minority oversampling and maximum dissimilarity undersampling is used to balance the training dataset. We employ and compare a series of statistical learning techniques with respect to their predictive performance and discuss the importance of determining factors of accident occurrence from the ensemble of models. Findings substantiate that a trade-off between accuracy and sensitivity is inherent to imbalanced classification problems. Results show satisfying performance of tree-based methods which exhibit accuracies between 75% and 90% while exhibiting sensitivities between 30% and 50%. Overall, this analysis emphasizes the merits of using high-resolution data in the context of accident analysis. … (more)
- Is Part Of:
- Accident analysis and prevention. Volume 127(2019)
- Journal:
- Accident analysis and prevention
- Issue:
- Volume 127(2019)
- Issue Display:
- Volume 127, Issue 2019 (2019)
- Year:
- 2019
- Volume:
- 127
- Issue:
- 2019
- Issue Sort Value:
- 2019-0127-2019-0000
- Page Start:
- 134
- Page End:
- 149
- Publication Date:
- 2019-06
- Subjects:
- 62-07
Statistical learning -- Imbalanced data -- Binary classification -- Accident analysis -- Road safety
Accidents -- Prevention -- Periodicals
Accident Prevention -- Periodicals
Accidents -- Prévention -- Périodiques
363.106 - Journal URLs:
- http://www.sciencedirect.com/science/journal/00014575 ↗
http://www.elsevier.com/journals ↗ - DOI:
- 10.1016/j.aap.2019.02.008 ↗
- Languages:
- English
- ISSNs:
- 0001-4575
- Deposit Type:
- Legaldeposit
- View Content:
- Available online (eLD content is only available in our Reading Rooms) ↗
- Physical Locations:
- British Library DSC - 0573.130000
British Library DSC - BLDSS-3PM
British Library HMNTS - ELD Digital store - Ingest File:
- 16393.xml