Divide and Recombine for Large and Complex Data: Model Likelihood Functions using MCMC

Qi Liu,Anindya Bhadra,William S. Cleveland
DOI: https://doi.org/10.48550/arXiv.1801.05007
2018-01-16
Abstract:In Divide & Recombine (D&R), big data are divided into subsets, each analytic method is applied to subsets, and the outputs are recombined. This enables deep analysis and practical computational performance. An innovate D\&R procedure is proposed to compute likelihood functions of data-model (DM) parameters for big data. The likelihood-model (LM) is a parametric probability density function of the DM parameters. The density parameters are estimated by fitting the density to MCMC draws from each subset DM likelihood function, and then the fitted densities are recombined. The procedure is illustrated using normal and skew-normal LMs for the logistic regression DM.
Methodology,Machine Learning
What problem does this paper attempt to address?