Use of generalized additive modelling techniques to create synthetic daily temperature networks for benchmarking homogenization algorithms. Issue 1 (25th June 2019)
- Record Type:
- Journal Article
- Title:
- Use of generalized additive modelling techniques to create synthetic daily temperature networks for benchmarking homogenization algorithms. Issue 1 (25th June 2019)
- Main Title:
- Use of generalized additive modelling techniques to create synthetic daily temperature networks for benchmarking homogenization algorithms
- Authors:
- Killick, Rachel E
Jolliffe, Ian T
Willett, Kate M - Abstract:
- Abstract: Background: The removal of non-climatic artefacts, inhomogeneities, from observed station data for the purpose of climate research is an ongoing task. Progress on homogenization algorithms is limited by a lack of suitable test data that are sufficiently realistic and where the 'truth' is known a priori. Objectives: This article describes a new method to create realistic, synthetic, daily temperature data, where the truth is known completely, thus allowing them to be used as a benchmark against which to test the performance of homogenization algorithms. Methods: Observations, reanalysis data and the power of statistical modelling, specifically gamma generalized additive models (GAMs), were combined to produce clean synthetic, daily mean temperature time series. The created clean data were corrupted with realistic inhomogeneities using both constant offsets and time-varying offsets produced by perturbing some of the GAM's input variables. When assessing the created clean and corrupted data, particular focus was given to variability, inter-station correlations, temporal autocorrelations and extreme values. Results: This is the first daily benchmarking study at this scale, bringing improvements over monthly or annual studies as it is at the daily level that the extremes of climate are felt. The created inhomogeneities mimic real-world features, as identified from real-world data by the homogenization community. They take the form of both steps and trends and exploreAbstract: Background: The removal of non-climatic artefacts, inhomogeneities, from observed station data for the purpose of climate research is an ongoing task. Progress on homogenization algorithms is limited by a lack of suitable test data that are sufficiently realistic and where the 'truth' is known a priori. Objectives: This article describes a new method to create realistic, synthetic, daily temperature data, where the truth is known completely, thus allowing them to be used as a benchmark against which to test the performance of homogenization algorithms. Methods: Observations, reanalysis data and the power of statistical modelling, specifically gamma generalized additive models (GAMs), were combined to produce clean synthetic, daily mean temperature time series. The created clean data were corrupted with realistic inhomogeneities using both constant offsets and time-varying offsets produced by perturbing some of the GAM's input variables. When assessing the created clean and corrupted data, particular focus was given to variability, inter-station correlations, temporal autocorrelations and extreme values. Results: This is the first daily benchmarking study at this scale, bringing improvements over monthly or annual studies as it is at the daily level that the extremes of climate are felt. The created inhomogeneities mimic real-world features, as identified from real-world data by the homogenization community. They take the form of both steps and trends and explore changes in the mean and variance. Clean and corrupted data are created for four regions in the USA: Wyoming, the North East, the South East and the South West. These four regions encompass a diverse range of climates, from a snow climate in the North East, to desert climates in the South West. Four test scenarios were created to allow the assessment of algorithm ability for different inhomogeneity and data structures. Scenario 1 was a best guess for the real world. The other three scenarios were constructed in ways that allowed the effects of station density, step versus trend inhomogeneities and varying temporal autocorrelation to be investigated. The closeness to reality of the created clean and corrupted data was assessed by comparisons with the observed data, noting that the observed data contain some level of both systematic and random error and are therefore not perfect themselves. Generally, the created clean data had higher interstation correlations in deseasonalized series (~0.9 versus ~0.7) and lower temporal autocorrelations in deseasonalized difference series (~0.01 versus ~0.10) than their real-world counterparts. The addition of inhomogeneities to create the corrupted data resulted in higher temporal autocorrelations in the deseasonalized difference series (~0.20 versus ~0.10) than those seen in the observed data. Despite these differences, the created levels of correlation are able to address issues of signal-to-noise ratio for detection algorithms, as in real-world data. Conclusions: These created clean and corrupted data provide a valuable first daily surface temperature data set that can be used in homogenisation benchmarking studies. It is anticipated that they will serve as a baseline to be built upon in the future. … (more)
- Is Part Of:
- Dynamics and statistics of the climate system. Volume 3:Issue 1(2018)
- Journal:
- Dynamics and statistics of the climate system
- Issue:
- Volume 3:Issue 1(2018)
- Issue Display:
- Volume 3, Issue 1 (2018)
- Year:
- 2018
- Volume:
- 3
- Issue:
- 1
- Issue Sort Value:
- 2018-0003-0001-0000
- Page Start:
- Page End:
- Publication Date:
- 2019-06-25
- Subjects:
- Homogenization -- benchmarking -- Generalized Additive Model -- synthetic data
Climatology -- Periodicals
Climatology -- Statistical methods -- Periodicals
Climatology -- Mathematical models -- Periodicals
551.6 - Journal URLs:
- http://www.oxfordjournals.org/ ↗
https://academic.oup.com/climatesystem ↗ - DOI:
- 10.1093/climsys/dzz001 ↗
- Languages:
- English
- ISSNs:
- 2059-6987
- Deposit Type:
- Legaldeposit
- View Content:
- Available online (eLD content is only available in our Reading Rooms) ↗
- Physical Locations:
- British Library DSC - BLDSS-3PM
British Library HMNTS - ELD Digital store - Ingest File:
- 12401.xml