Addressing Data Fragmentation in Biodiversity: A Workflow for Integrated Species Distribution Models
Sam Wenaas Perrin, Kwaku Peprah Adjei, Philip Stanley Mostert, Ron Ronald Togunov, Ivar Herfindal, Joachim Paul Töpper, John‐Arvid Grytnes, Joseph Chipperfield, Robert Brian O'Hara, Anders Gravbrøt FinstadABSTRACT
Aim
A comprehensive understanding of the spatial distribution of biodiversity is hindered by fragmented datasets, sampling biases, and inconsistent observation protocols. Here, we present a workflow that integrates disparate datasets to produce large scale maps of biodiversity metrics as a basis for management‐relevant information tools. We use integrated Species Distribution modelling (iSDM) to account for sampling biases and disparate data collection techniques, taking advantage of the vast numbers of open datasets available in data aggregators like GBIF.
Location
Norway (excluding Svalbard and Jan Mayen).
Taxon
Vascular plants.
Methods
The workflow consists of four main steps: data acquisition, data integration, integrated species distribution modelling (iSDM), and the production of derived outputs. Input data include structured surveys, opportunistic observations, and environmental covariates. These are standardized and integrated into a point‐processed based iSDM framework to produce species richness maps, associated uncertainties, and sampling effort maps. The outputs are further processed to identify biodiversity hotspots or to summarize species–environment relationships. The workflow used vascular plant data from Norway, combining occurrence‐only and detection/non‐detection datasets with environmental covariates. Outputs were generated at a spatial resolution of 500 × 500 m, balancing accuracy, computational feasibility and relevance for management decisions. High‐performance computing resources were utilized for model fitting and predictions. A subset of available data were used to validate the species richness maps.
Results
We produced detailed maps of species richness, uncertainties and sampling intensity across Norway's heterogeneous landscape, incorporating 1218 species in our final results. The species richness patterns display patterns consistent with previous mapping efforts, highlighting higher richness in the southern and coastal regions of Norway and distinctly lower richness in the north. Validation showed a modest increase in model accuracy, improving on past modelling efforts while integrating a much wider variety of datasets through the iSDM framework. The workflow highlights limitations in the infrastructure of the currently openly accessible data, particularly the need for more structured detection/non‐detection datasets and standardized metadata.
Main Conclusions
This study underscores the potential of workflows that integrate disparate datasets for biodiversity modelling. To maximize accuracy and utility, future efforts should focus on improving data standardization, the publication and collection of more structured data, and fostering data‐sharing collaborations. Advances in the workflow itself, including optimizing modelling covariates and integrating more comprehensive spatio‐temporal aspects, will also increase the relevance of the outputs. These advances will increase our ability to estimate species richness with a precision and accuracy that can reliably inform conservation and management decisions.