Incorporating Subsampling into Bayesian Models for High-Dimensional Spatial Data

Sudipto Saha,Jonathan R. Bradley

DOI: https://doi.org/10.1214/24-BA1426

2024-03-20

Abstract:Additive spatial statistical models with weakly stationary process assumptions have become standard in spatial statistics. However, one disadvantage of such models is the computation time, which rapidly increases with the number of data points. The goal of this article is to apply an existing subsampling strategy to standard spatial additive models and to derive the spatial statistical properties. We call this strategy the ''spatial data subset model'' (SDSM) approach, which can be applied to big datasets in a computationally feasible way. Our approach has the advantage that one does not require any additional restrictive model assumptions. That is, computational gains increase as model assumptions are removed when using our model framework. This provides one solution to the computational bottlenecks that occur when applying methods such as Kriging to ''big data''. We provide several properties of this new spatial data subset model approach in terms of moments, sill, nugget, and range under several sampling designs. An advantage of our approach is that it subsamples without throwing away data, and can be implemented using datasets of any size that can be stored. We present the results of the spatial data subset model approach on simulated datasets, and on a large dataset consists of 150,000 observations of daytime land surface temperatures measured by the MODIS instrument onboard the Terra satellite.

Methodology,Computation

What problem does this paper attempt to address?

The main problem that this paper attempts to solve is that when dealing with high - dimensional spatial data, the computation time of the traditional Bayesian spatial model increases rapidly as the number of data points grows. Specifically, the paper proposes a new "Spatial Data Subset Model (SDSM)" method, which solves this computational bottleneck by applying the existing subsampling strategies in the standard spatial additive model. This method can not only computationally handle large - scale data sets, but also does not require additional restrictive assumptions, thereby improving computational efficiency while removing model assumptions. In addition, the paper also explores the spatial statistical properties of this method under different sampling designs, such as moments, sill, nugget and range, etc., and verifies the effectiveness of this method through simulated data sets and actual data sets (daytime land surface temperature data with 150,000 observations).

Incorporating Subsampling into Bayesian Models for High-Dimensional Spatial Data

Bayesian statistics in spatial epidemiology]

Bayesian Nonstationary Spatial Modeling for Very Large Datasets

High-Dimensional Bayesian Geostatistics

Spatial Factor Modeling: A Bayesian Matrix-Normal Approach for Misaligned Data

A Scalable Partitioned Approach to Model Massive Nonstationary Non-Gaussian Spatial Datasets

Fully Bayesian inference for spatiotemporal data with the multi-resolution approximation

Flexible Basis Representations for Modeling Large Non-Gaussian Spatial Data

A Divide-and-Conquer Bayesian Approach to Large-Scale Kriging

Bayesian Spatial Predictive Synthesis

Additive Models with Spatio-Temporal Data

High-dimensional Multivariate Geostatistics: A Bayesian Matrix-Normal Approach

Adaptive Bayesian nonstationary modeling for large spatial datasets using covariance approximations

Nonstationary Spatial Process Models with Spatially Varying Covariance Kernels

A Variational Approach for Modeling High-dimensional Spatial Generalized Linear Mixed Models

The Spatial Statistic Trinity: A Generic Framework for Spatial Sampling and Inference.

Bayesian Modeling with Spatial Curvature Processes

Bayesian Design with Sampling Windows for Complex Spatial Processes

Adaptive and Stratified Subsampling Techniques for High Dimensional Non-Standard Data Environments

spBayes for large univariate and multivariate point-referenced spatio-temporal data models

Relationship-aware Multivariate Sampling Strategy for Scientific Simulation Data