A predictive model of recreational water quality based on adaptive synthetic sampling algorithms and machine learning. (15th June 2020)
- Record Type:
- Journal Article
- Title:
- A predictive model of recreational water quality based on adaptive synthetic sampling algorithms and machine learning. (15th June 2020)
- Main Title:
- A predictive model of recreational water quality based on adaptive synthetic sampling algorithms and machine learning
- Authors:
- Xu, Tingting
Coco, Giovanni
Neale, Martin - Abstract:
- Abstract: Predicting recreational water quality is one of the most difficult tasks in water management with major implications for humans and society. Many data-driven models have been used to predict water quality indicators to allow a real time assessment of public health risk. This assessment is most commonly based on Faecal Indicator Bacteria (FIB), with the value of FIB compared with thresholds published in guidelines. However, FIB values usually tend to be unbalanced within water quality datasets, with small proportions of data exceeding guideline thresholds and far larger numbers that do not. This can be a limiting factor in the uptake of model predictions since, even if the overall accuracy is high, the sensitivity of the predictions can be low. To address this issue, this paper proposes an adaptive synthetic sampling algorithm (ADASYN) to generate synthetic above-threshold FIB instances and test the validity of the approach for the prediction of recreational water quality. The models in this paper are based on four machine learning techniques: k-mean nearest neighbour, boosting decision tree, support vector machine, and multi-layer perceptron artificial neural network and are applied to five different locations in Auckland, New Zealand. Aside from support vector machine, all models provide favourable predictions with relatively high sensitivity (around 75%) and overall accuracy (over 90%), indicating that both the compliant and exceedance conditions can beAbstract: Predicting recreational water quality is one of the most difficult tasks in water management with major implications for humans and society. Many data-driven models have been used to predict water quality indicators to allow a real time assessment of public health risk. This assessment is most commonly based on Faecal Indicator Bacteria (FIB), with the value of FIB compared with thresholds published in guidelines. However, FIB values usually tend to be unbalanced within water quality datasets, with small proportions of data exceeding guideline thresholds and far larger numbers that do not. This can be a limiting factor in the uptake of model predictions since, even if the overall accuracy is high, the sensitivity of the predictions can be low. To address this issue, this paper proposes an adaptive synthetic sampling algorithm (ADASYN) to generate synthetic above-threshold FIB instances and test the validity of the approach for the prediction of recreational water quality. The models in this paper are based on four machine learning techniques: k-mean nearest neighbour, boosting decision tree, support vector machine, and multi-layer perceptron artificial neural network and are applied to five different locations in Auckland, New Zealand. Aside from support vector machine, all models provide favourable predictions with relatively high sensitivity (around 75%) and overall accuracy (over 90%), indicating that both the compliant and exceedance conditions can be effectively predicted through the use of more sophisticated model training which involves artificial data. Considering the model accuracy and stability, boosting decision trees (BDT) and multi-layer perceptron artificial neural (MLP-ANN) network are the best two models and the multi-layer perceptron is the most efficient with the shortest computation time. Graphical abstract: Image 1 Highlights: Applied ADASYN to address the problem of unbalanced water quality datasets. Ssessed the ability of machine learning models when using a balanced dataset. Compared results at different locations using different machine learning models. … (more)
- Is Part Of:
- Water research. Volume 177(2020)
- Journal:
- Water research
- Issue:
- Volume 177(2020)
- Issue Display:
- Volume 177, Issue 2020 (2020)
- Year:
- 2020
- Volume:
- 177
- Issue:
- 2020
- Issue Sort Value:
- 2020-0177-2020-0000
- Page Start:
- Page End:
- Publication Date:
- 2020-06-15
- Subjects:
- Water quality -- Adaptive synthetic sampling algorithm -- Nearest neighbour -- Boosting decision tree -- Support vector machine -- Artificial neural network
Water -- Pollution -- Research -- Periodicals
363.7394 - Journal URLs:
- http://catalog.hathitrust.org/api/volumes/oclc/1769499.html ↗
http://www.sciencedirect.com/science/journal/00431354 ↗
http://www.elsevier.com/journals ↗ - DOI:
- 10.1016/j.watres.2020.115788 ↗
- Languages:
- English
- ISSNs:
- 0043-1354
- Deposit Type:
- Legaldeposit
- View Content:
- Available online (eLD content is only available in our Reading Rooms) ↗
- Physical Locations:
- British Library DSC - 9273.400000
British Library DSC - BLDSS-3PM
British Library HMNTS - ELD Digital store - Ingest File:
- 13369.xml