Abstract:Accurate information on surface soil moisture (SSM) content at a global scale under different climatic conditions is important for hydrological and climatological applications. Machine-learning-based systematic integration of in situ hydrological measurements, complex environmental and climate data, and satellite observation facilitate the generation of reliable data products to monitor and analyse the exchange of water, energy, and carbon in the Earth system at a proper space–time resolution. This study investigates the estimation of daily SSM using 8 optimised machine learning (ML) algorithms and 10 ensemble models (constructed via model bootstrap aggregating techniques and five-fold cross-validation). The algorithmic implementations were trained and tested using International Soil Moisture Network (ISMN) data collected from 1722 stations distributed across the world. The result showed that the K-neighbours Regressor (KNR) had the lowest root-mean-square error (0.0379 cm 3 cm −3 ) on the "test_random" set (for testing the performance of randomly split data during training), the Random Forest Regressor (RFR) had the lowest RMSE (0.0599 cm 3 cm −3 ) on the "test_temporal" set (for testing the performance on the period that was not used in training), and AdaBoost (AB) had the lowest RMSE (0.0786 cm 3 cm −3 ) on the "test_independent-stations" set (for testing the performance on the stations that were not used in training). Independent evaluation on novel stations across different climate zones was conducted. For the optimised ML algorithms, the median RMSE values were below 0.1 cm 3 cm −3 . GradientBoosting (GB), Multi-layer Perceptron Regressor (MLPR), Stochastic Gradient Descent Regressor (SGDR), and RFR achieved a median r score of 0.6 in 12, 11, 9, and 9 climate zones, respectively, out of 15 climate zones. The performance of ensemble models improved significantly, with the median RMSE value below 0.075 cm 3 cm −3 for all climate zones. All voting regressors achieved r scores of above 0.6 in 13 climate zones; BSh (hot semi-arid climate) and BWh (hot desert climate) were the exceptions because of the sparse distribution of training stations. The metric evaluation showed that ensemble models can improve the performance of single ML algorithms and achieve more stable results. Based on the results computed for three different test sets, the ensemble model with KNR, RFR and Extreme Gradient Boosting (XB) performed the best. Overall, our investigation shows that ensemble machine learning algorithms have a greater capability with respect to predicting SSM compared with the optimised or base ML algorithms; this indicates their huge potential applicability in estimating water cycle budgets, managing irrigation, and predicting crop yields.

Model ensembles of artificial neural networks and support vector regression for improved accuracy in the prediction of vegetation conditions

A mixed model approach to drought prediction using artificial neural networks: Case of an operational drought monitoring environment

Support Vector Machine―an Alternative to Artificial Neuron Network for Water Quality Forecasting in an Agricultural Nonpoint Source Polluted River?

Drought risk assessment: integrating meteorological, hydrological, agricultural and socio-economic factors using ensemble models and geospatial techniques

Forecasting actual evapotranspiration without climate data based on stacked integration of DNN and meta-heuristic models across China from 1958 to 2021

Modelling drought in South Africa: meteorological insights and predictive parameters

Forecasting of meteorological drought using ensemble and machine learning models

Evaluating Performance of Multiple Machine Learning Models for Drought Monitoring: A Case Study of Typical Grassland in Inner Mongolia

Robust meteorological drought prediction using antecedent SST fluctuations and machine learning

A stacking ANN ensemble model of ML models for stream water quality prediction of Godavari River Basin, India

A Novel Fusion-Based Methodology for Drought Forecasting

Prediction of meteorological drought by using hybrid support vector regression optimized with HHO versus PSO algorithms

Ensemble of optimised machine learning algorithms for predicting surface soil moisture content at a global scale

How Well Can Machine Learning Models Perform without Hydrologists? Application of Rational Feature Selection to Improve Hydrological Forecasting

Drought forecasting through statistical models using standardised precipitation index: a systematic review and meta-regression analysis

Hybrid Data-Driven Models for Hydrological Simulation and Projection on the Catchment Scale

Multi-models for SPI drought forecasting in the north of Haihe River Basin, China

On the importance of training methods and ensemble aggregation for runoff prediction by means of artificial neural networks

Improving multiple model ensemble predictions of daily precipitation and temperature through machine learning techniques

Combination of data-driven models and best subset regression for predicting the standardized precipitation index (SPI) at the Upper Godavari Basin in India

Bayesian Machine Learning Ensemble Approach to Quantify Model Uncertainty in Predicting Groundwater Storage Change.