How to make more from exposure data? An integrated machine learning pipeline to predict pathogen exposure. Issue 10 (19th August 2019)
- Record Type:
- Journal Article
- Title:
- How to make more from exposure data? An integrated machine learning pipeline to predict pathogen exposure. Issue 10 (19th August 2019)
- Main Title:
- How to make more from exposure data? An integrated machine learning pipeline to predict pathogen exposure
- Authors:
- Fountain‐Jones, Nicholas M.
Machado, Gustavo
Carver, Scott
Packer, Craig
Recamonde‐Mendoza, Mariana
Craft, Meggan E. - Editors:
- Fenton, Andy
- Abstract:
- Abstract: Predicting infectious disease dynamics is a central challenge in disease ecology. Models that can assess which individuals are most at risk of being exposed to a pathogen not only provide valuable insights into disease transmission and dynamics but can also guide management interventions. Constructing such models for wild animal populations, however, is particularly challenging; often only serological data are available on a subset of individuals and nonlinear relationships between variables are common. Here we provide a guide to the latest advances in statistical machine learning to construct pathogen‐risk models that automatically incorporate complex nonlinear relationships with minimal statistical assumptions from ecological data with missing data. Our approach compares multiple machine learning algorithms in a unified environment to find the model with the best predictive performance and uses game theory to better interpret results. We apply this framework on two major pathogens that infect African lions: canine distemper virus (CDV) and feline parvovirus. Our modelling approach provided enhanced predictive performance compared to more traditional approaches, as well as new insights into disease risks in a wild population. We were able to efficiently capture and visualize strong nonlinear patterns, as well as model complex interactions between variables in shaping exposure risk from CDV and feline parvovirus. For example, we found that lions were more likely toAbstract: Predicting infectious disease dynamics is a central challenge in disease ecology. Models that can assess which individuals are most at risk of being exposed to a pathogen not only provide valuable insights into disease transmission and dynamics but can also guide management interventions. Constructing such models for wild animal populations, however, is particularly challenging; often only serological data are available on a subset of individuals and nonlinear relationships between variables are common. Here we provide a guide to the latest advances in statistical machine learning to construct pathogen‐risk models that automatically incorporate complex nonlinear relationships with minimal statistical assumptions from ecological data with missing data. Our approach compares multiple machine learning algorithms in a unified environment to find the model with the best predictive performance and uses game theory to better interpret results. We apply this framework on two major pathogens that infect African lions: canine distemper virus (CDV) and feline parvovirus. Our modelling approach provided enhanced predictive performance compared to more traditional approaches, as well as new insights into disease risks in a wild population. We were able to efficiently capture and visualize strong nonlinear patterns, as well as model complex interactions between variables in shaping exposure risk from CDV and feline parvovirus. For example, we found that lions were more likely to be exposed to CDV at a young age but only in low rainfall years. When combined with our data calibration approach, our framework helped us to answer questions about risk of pathogen exposure that are difficult to address with previous methods. Our framework not only has the potential to aid in predicting disease risk in animal populations, but also can be used to build robust predictive models suitable for other ecological applications such as modelling species distribution or diversity patterns. Abstract : The authors provide a practical guide to integrating the latest advances in data science and machine learning to better quantify disease risk. Importantly, many of the methods employed in the guide can be easily adjusted to address broader ecological questions involving, but not limited to, big data or complex non‐linear interactions. … (more)
- Is Part Of:
- Journal of animal ecology. Volume 88:Issue 10(2019)
- Journal:
- Journal of animal ecology
- Issue:
- Volume 88:Issue 10(2019)
- Issue Display:
- Volume 88, Issue 10 (2019)
- Year:
- 2019
- Volume:
- 88
- Issue:
- 10
- Issue Sort Value:
- 2019-0088-0010-0000
- Page Start:
- 1447
- Page End:
- 1461
- Publication Date:
- 2019-08-19
- Subjects:
- boosted regression trees -- disease ecology -- gradient boosting models -- machine learning -- model‐agnostic methods -- random forests -- serology -- support vector machines
Animal ecology -- Periodicals
591.7 - Journal URLs:
- http://www.jstor.org/journals/00218790.html ↗
http://www3.interscience.wiley.com/journal/117960113/home ↗
http://onlinelibrary.wiley.com/ ↗
http://firstsearch.oclc.org ↗
http://firstsearch.oclc.org/journal=0021-8790;screen=info;ECOIP ↗ - DOI:
- 10.1111/1365-2656.13076 ↗
- Languages:
- English
- ISSNs:
- 0021-8790
- Deposit Type:
- Legaldeposit
- View Content:
- Available online (eLD content is only available in our Reading Rooms) ↗
- Physical Locations:
- British Library DSC - 4936.000000
British Library DSC - BLDSS-3PM
British Library HMNTS - ELD Digital store - Ingest File:
- 11857.xml