skip to main content
US FlagAn official website of the United States government
dot gov icon
Official websites use .gov
A .gov website belongs to an official government organization in the United States.
https lock icon
Secure .gov websites use HTTPS
A lock ( lock ) or https:// means you've safely connected to the .gov website. Share sensitive information only on official, secure websites.


Title: Improving area of occupancy estimates for parapatric species using distribution models and support vector machines
Abstract As geographic range estimates for the IUCN Red List guide conservation actions, accuracy and ecological realism are crucial. IUCN’s extent of occurrence (EOO) is the general region including the species’ range, while area of occupancy (AOO) is the subset of EOO occupied by the species. Data‐poor species with incomplete sampling present particular difficulties, but species distribution models (SDMs) can be used to predict suitable areas. Nevertheless, SDMs typically employ abiotic variables (i.e., climate) and do not explicitly account for biotic interactions that can impose range constraints. We sought to improve range estimates for data‐poor, parapatric species by masking out areas under inferred competitive exclusion. We did so for two South American spiny pocket mice:Heteromys australis(Least Concern) andHeteromys teleus(Vulnerable due to especially poor sampling), whose ranges appear restricted by competition. For both species, we estimated EOO using SDMs and AOO with four approaches: occupied grid cells, abiotic SDM prediction, and this prediction masked by approximations of the areas occupied by each species’ congener. We made the masks using support vector machines (SVMs) fit with two data types: occurrence coordinates alone; and coordinates along with SDM predictions of suitability. Given the uncertainty in calculating AOO for low‐data species, we made estimates for the lower and upper bounds for AOO, but only make recommendations forH. teleusas its full known range was considered. The SVM approaches (especially the second one) had lower classification error and made more ecologically realistic delineations of the contact zone. ForH. teleus, the lower AOO bound (a strongly biased underestimate) corresponded to Endangered (occupied grid cells), while the upper bounds (other approaches) led to Near Threatened. As we currently lack data to determine the species’ true occupancy within the post‐processed SDM prediction, we recommend that an updated listing forH. teleusinclude these bounds for AOO. This study advances methods for estimating the upper bound of AOO and highlights the need for better ways to produce unbiased estimates of lower bounds. More generally, the SVM approaches for post‐processing SDM predictions hold promise for improving range estimates for other uses in biogeography and conservation.  more » « less
Award ID(s):
1661510
PAR ID:
10375277
Author(s) / Creator(s):
 ;  ;  ;  ;  
Publisher / Repository:
Wiley Blackwell (John Wiley & Sons)
Date Published:
Journal Name:
Ecological Applications
Volume:
31
Issue:
1
ISSN:
1051-0761
Format(s):
Medium: X
Sponsoring Org:
National Science Foundation
More Like this
  1. Environmental conditions are dynamic, and plants respond to those dynamics on multiple time scales. Disequilibrium occurs when a response occurs more slowly than the driving environmental changes. We review evidence regarding disequilibrium in plant distributions, including their responses to paleoclimate changes, recent climate change and new species introductions. There is strong evidence that plant species distributions are often in some disequilibrium with their environmental conditions.This disequilibrium poses a challenge when projecting future species distributions using species distribution models (SDMs). Classically, SDMs assume that the set of species occurrences is an unbiased sample of the suitable environmental conditions. However, a species in disequilibrium with the environment may have higher‐than‐expected occurrence probabilities (e.g. due to extinction debts) or lower‐than‐expected occurrence probabilities (e.g. due to dispersal limitation) in different areas. If unaccounted for, this will lead to biased estimates of the environmental suitability.We review methods for avoiding such biases in SDMs, ranging from simple thinning of the occurrence dataset to complex dynamic and process‐based models. Such models require large data inputs, natural history knowledge and technical expertise, so implementing them can be challenging. Despite this, we advocate for their increased use, since process‐based models provide the best potential to account for biases in model training data and to then represent the dynamics of species occupancy as ranges shift.Synthesis. Occurrence records for a species are often in disequilibrium with climate. SDMs trained on such data will produce biased estimates of a species' niche unless this disequilibrium is addressed in the modelling. A range of tools, spanning a wide gradient of complexity and realism, can resolve this bias. 
    more » « less
  2. Dar, Kamran Shaukat (Ed.)
    Species distribution models (SDMs) are increasingly popular tools for profiling disease risk in ecology, particularly for infectious diseases of public health importance that include an obligate non-human host in their transmission cycle. SDMs can create high-resolution maps of host distribution across geographical scales, reflecting baseline risk of disease. However, as SDM computational methods have rapidly expanded, there are many outstanding methodological questions. Here we address key questions about SDM application, using schistosomiasis risk in Brazil as a case study. Schistosomiasis is transmitted to humans through contact with the free-living infectious stage ofSchistosomaspp. parasites released from freshwater snails, the parasite’s obligate intermediate hosts. In this study, we compared snail SDM performance across machine learning (ML) approaches (MaxEnt, Random Forest, and Boosted Regression Trees), geographic extents (national, regional, and state), types of presence data (expert-collected and publicly-available), and snail species (Biomphalaria glabrata,B.straminea, andB.tenagophila). We used high-resolution (1km) climate, hydrology, land-use/land-cover (LULC), and soil property data to describe the snails’ ecological niche and evaluated models on multiple criteria. Although all ML approaches produced comparable spatially cross-validated performance metrics, their suitability maps showed major qualitative differences that required validation based on local expert knowledge. Additionally, our findings revealed varying importance of LULC and bioclimatic variables for different snail species at different spatial scales. Finally, we found that models using publicly-available data predicted snail distribution with comparable AUC values to models using expert-collected data. This work serves as an instructional guide to SDM methods that can be applied to a range of vector-borne and zoonotic diseases. In addition, it advances our understanding of the relevant environment and bioclimatic determinants of schistosomiasis risk in Brazil. 
    more » « less
  3. Abstract Species distribution models (SDMs) have become increasingly popular for making ecological inferences, as well as predictions to inform conservation and management. In predictive modeling, practitioners often use correlative SDMs that only evaluate a single spatial scale and do not account for differences in life stages. These modeling decisions may limit the performance of SDMs beyond the study region or sampling period. Given the increasing desire to develop transferable SDMs, a robust framework is necessary that can account for known challenges of model transferability. Here, we propose a comparative framework to develop transferable SDMs, which was tested using satellite telemetry data from green turtles (Chelonia mydas). This framework is characterized by a set of steps comparing among different models based on (1) model algorithm (e.g., generalized linear model vs. Gaussian process regression) and formulation (e.g., correlative model vs. hybrid model), (2) spatial scale, and (3) accounting for life stage. SDMs were fitted as resource selection functions and trained on data from the Gulf of Mexico with bathymetric depth, net primary productivity, and sea surface temperature as covariates. Independent validation datasets from Brazil and Qatar were used to assess model transferability. A correlative SDM using a hierarchical Gaussian process regression (HGPR) algorithm exhibited greater transferability than a hybrid SDM using HGPR, as well as correlative and hybrid forms of hierarchical generalized linear models. Additionally, models that evaluated habitat selection at the finest spatial scale and that did not account for life stage proved to be the most transferable in this study. The comparative framework presented here may be applied to a variety of species, ecological datasets (e.g., presence‐only, presence‐absence, mark‐recapture), and modeling frameworks (e.g., resource selection functions, step selection functions, occupancy models) to generate transferable predictions of species–habitat associations. We expect that SDM predictions resulting from this comparative framework will be more informative management tools and may be used to more accurately assess climate change impacts on a wide array of taxa. 
    more » « less
  4. Dainton, John (Ed.)
    Improving models of species' distributions is essential for conservation, especially in light of global change. Species distribution models (SDMs) often rely on mean environmental conditions, yet species distributions are also a function of environmental heterogeneity and filtering acting at multiple spatial scales. Geodiversity, which we define as the variation of abiotic features and processes of Earth's entire geosphere (inclusive of climate), has potential to improve SDMs and conservation assessments, as they capture multiple abiotic dimensions of species niches, however they have not been sufficiently tested in SDMs. We tested a range of geodiversity variables computed at varying scales using climate and elevation data. We compared predictive performance of MaxEnt SDMs generated using CHELSA bioclimatic variables to those also including geodiversity variables for 31 mammalian species in Colombia. Results show the spatial grain of geodiversity variables affects SDM performance. Some variables consistently exhibited an increasing or decreasing trend in variable importance with spatial grain, showing slight scale-dependence and indicating that some geodiversity variables are more relevant at particular scales for some species. Incorporating geodiversity variables into SDMs, and doing so at the appropriate spatial scales, enhances the ability to model species-environment relationships, thereby contributing to the conservation and management of biodiversity. This article is part of the Theo Murphy meeting issue ‘Geodiversity for science and society’. 
    more » « less
  5. Abstract AimSpecies distribution models (SDMs) that integrate presence‐only and presence–absence data offer a promising avenue to improve information on species' geographic distributions. The use of such ‘integrated SDMs’ on a species range‐wide extent has been constrained by the often limited presence–absence data and by the heterogeneous sampling of the presence‐only data. Here, we evaluate integrated SDMs for studying species ranges with a novel expert range map‐based evaluation. We build new understanding about how integrated SDMs address issues of estimation accuracy and data deficiency and thereby offer advantages over traditional SDMs. LocationSouth and Central America. Time Period1979–2017. Major Taxa StudiedHummingbirds. MethodsWe build integrated SDMs by linking two observation models – one for each data type – to the same underlying spatial process. We validate SDMs with two schemes: (i) cross‐validation with presence–absence data and (ii) comparison with respect to the species' whole range as defined with IUCN range maps. We also compare models relative to the estimated response curves and compute the association between the benefit of the data integration and the number of presence records in each data set. ResultsThe integrated SDM accounting for the spatially varying sampling intensity of the presence‐only data was one of the top performing models in both model validation schemes. Presence‐only data alleviated overly large niche estimates, and data integration was beneficial compared to modelling solely presence‐only data for species which had few presence points when predicting the species' whole range. On the community level, integrated models improved the species richness prediction. Main ConclusionsIntegrated SDMs combining presence‐only and presence–absence data are successfully able to borrow strengths from both data types and offer improved predictions of species' ranges. Integrated SDMs can potentially alleviate the impacts of taxonomically and geographically uneven sampling and to leverage the detailed sampling information in presence–absence data. 
    more » « less