Objectives: Soil erosion is currently recognized as a critical global challenge with severe social, economic, and environmental consequences. Accurate estimation of soil erodibility is a pivotal factor in the precise simulation of erosion processes and sustainable land management. In recent years, Digital Soil Mapping (DSM) coupled with Machine Learning (ML) algorithms has emerged as a powerful approach for handling complex environmental datasets. The primary objective of this study was to evaluate the influence of three groups of environmental covariates—soil properties, geomorphometric indices derived from a Digital Elevation Model (DEM), and multi-sensor remote sensing data (Landsat 8 and Sentinel-2)—on the spatial modeling of water erodibility (Factor K) and wind erodibility (Erodible Fraction - EF) in the Sistan Plain. Specifically, four widely used ML algorithms, including Regression Tree (rpart), Support Vector Machine (SVM), Cubist, and Random Forest (RF), were employed. Water erodibility (K) was calculated using three distinct equations: USLE, EPIC, and Sharpley and Smith, while wind erodibility was determined using the Fryrear method.
Methodology: To achieve the research objectives, 1200 surface soil samples were collected from agricultural lands in the study area, and relevant physicochemical properties for the K factor and EF index were measured. Remote sensing variables were extracted from Landsat 8 (OLI sensor) and Sentinel-2 (MSI sensor) imagery. Landsat 8 data with a spatial resolution of 30 m were acquired from the United States Geological Survey (USGS) archive, while Sentinel-2 data (10, 20, and 60 m resolution) were obtained from the Copernicus Open Access Hub to enhance analysis accuracy. The dataset comprised six main spectral bands, four soil-related spectral indices (Clay, Carbonate, Salinity, and Gypsum), five vegetation indices (NDVI, EVI, PVI, RVI, SAVI), and one innovative index (KfactorRS Index). To mitigate the risk of overfitting and reduce dimensionality, the Boruta algorithm was implemented for precise feature selection for each target variable. Subsequently, the dataset was randomly split into training (80%) and testing (20%) subsets. The predictive performance and accuracy of the ML models were evaluated using seven statistical metrics: Mean Error (ME), Root Mean Square Error (RMSE), Normalized RMSE (NRMSE), Pearson correlation coefficient (r), Coefficient of Determination (R²), Ratio of Performance to Deviation (RPD), and Lin’s Concordance Correlation Coefficient (CCC).
Results: The results indicated a significant negative correlation (p < 0.01) between water erodibility coefficients (K) and the wind erodibility index (EF). A justifiable behavioral pattern observed in this study was the emergence of a negative correlation between spectral vegetation indices and clay content; a phenomenon that, contrary to initial assumptions, stems from vegetation degradation driven by salinity accumulation in fine-grained soils. Furthermore, the unexpected direct relationships of the KEPIC and KShSm equations with organic carbon and clay, alongside an inverse relationship with very fine sand, indicate a high sensitivity to textural variables that introduces statistical bias under the region's specific conditions. Consequently, applying these indices is not recommended without recalibration tailored to the local geomorphic setting. Regarding model performance, the Cubist algorithm demonstrated superior predictive capability for most variables (EF, KEPIC, KShSm), while the Random Forest (RF) model outperformed others in predicting KUSLE. Both Cubist and RF, as ensemble-based decision tree methods, showed high competence in modeling soil erosion complexities. Although the SVM model yielded reasonable performance across all variables, it never ranked as the superior model. Conversely, the rpart model exhibited the weakest performance, particularly for KEPIC and KShSm, indicating the inability of simple single-tree models to capture the non-linear complexities of soil erosion in this region.
Conclusion: Although Random Forest is widely recognized as one of the most proven and practical algorithms in DSM, this study demonstrated that rule-based models such as Cubist—which offer greater flexibility in modeling local non-linear relationships—can outperform RF in specific scenarios. Nevertheless, the consistency and stability of RF across all variables (consistently ranking first or second) and its absence of poor performance establish it as a safe and robust choice for predicting soil properties and erosion parameters in digital soil mapping studies. |