Explainable Machine Learning for Soil Health Index Prediction and Degradation-Risk Screening: A Synthetic Multi-Region Proof-of-Concept for Circular-Bio-economy Soil Management

Authors

DOI:

https://doi.org/10.63002/jrecs.404.1645

Keywords:

soil health index, explainable machine learning, permutation importance, circular bio economy, soil degradation, microplastics, synthetic data, precision agriculture

Abstract

Soil health is a multidimensional property shaped by interacting chemical, physical, biological, and management factors that influence agricultural productivity, carbon storage, and ecosystem resilience. Although machine learning (ML) is increasingly applied to soil assessment, much of the existing literature remains prediction-oriented, with comparatively limited emphasis on model interpretability and on integrating circular-bio economy management variables and emerging contaminants. This study presents a proof-of-concept explainable ML framework for predicting a continuous Soil Health Index (SHI) and a binary soil degradation risk label using features spanning soil chemistry, physical condition, biological activity, management, resource-efficiency indicators, and a microplastic risk proxy. The analysis uses a synthetic dataset engineered from an openly available Kaggle soil dataset and is explicitly intended as a methodological demonstration rather than field validation. Three regression models—Linear Regression, Gradient Boosting, and Random Forest—and three corresponding classification approaches—Logistic Regression, Gradient Boosting, and Random Forest—were evaluated under a common pre-processing framework. Predictive performance was assessed using R² for SHI regression and the area under the receiver operating characteristic curve (AUC) for degradation-risk classification, while permutation importance was applied to the best-performing regression model to examine the influence of predictors. Linear Regression produced the strongest regression performance (R² = 0.963), outperforming Gradient Boosting (R² = 0.922) and Random Forest (R² = 0.897). Logistic Regression achieved the highest classification performance (AUC = 0.986), closely followed by Gradient Boosting (AUC = 0.985) and Random Forest (AUC = 0.975). Permutation analysis identified soil organic carbon, microplastic risk, soil moisture, electrical conductivity, compaction, and bulk density among the most influential predictors. The strong performance of comparatively simple linear models suggests that the engineered SHI and degradation-risk relationships in this synthetic setting are predominantly structured and approximately linear. More importantly, the framework demonstrates how conventional soil indicators, circular bioeconomy management dimensions, and an emerging contaminant proxy can be incorporated into a transparent and interpretable predictive workflow. Given the dataset's synthetic nature, the reported performance metrics should not be interpreted as evidence of field-level predictive accuracy. Instead, the study provides a reproducible methodological template for subsequent validation using measured soil and management data.

Downloads

Published

22-08-2026