Spatial machine learning modelling reveals that soil indicators and tree type best explain shallow landslide release
Denise Christina Rüther, Kristine Flacké Haualand, Iris Louisa Johanna Peeters, Mark Andrew Kusk GillespieThe exploration of shallow landslide susceptibility is often impaired by biased landslide inventories, and by over-optimistic performance metrics linked to inadequate models. Here, we use a systematically mapped event inventory of 571 shallow landslides triggered in southern Norway and apply 32 gradient boosted decision tree models to rigorously test the effects of (1) a nested vs. simple cross-validation strategy, (2) spatial vs. non-spatial models, (3) four different cross-validation sampling strategies which were applied on (4) full vs. forest-only datasets. Model evaluation shows that models with random cross-validation provided the highest test performance metrics but did not account for autocorrelation and were likely over-optimistic. The spatial models with spatial cross-validation reduced autocorrelation in the model residuals at the expense of predictive power. Although no model emerged as optimal, findings across models suggest that important explanatory factors like elevation, aspect and bedrock weatherability serve as soil indicators, illustrating a need for improved datasets for soil thickness and heterogeneity. In the forest-only models, tree type was consistently an important predictor, with higher landslide susceptibility in deciduous forest, illustrating the potential of forest variables and forest-specific threshold values in shallow landslide susceptibility mapping.