MapReduce for Bayesian Network Parameter Learning using the EM Algorithm

Irina Brinster,O. Mengshoel,Aniruddha Basak
Abstract:This work applies the distributed computing framework MapReduce to Bayesian network parameter learning from incomplete data. We formulate the classical Expectation Maximization (EM) algorithm within the MapReduce framework. Analytically and experimentally we analyze the speed-up that can be obtained by means of MapReduce. We present details of the MapReduce formulation of EM, report speed-ups versus the sequential case, and carefully compare various Hadoop cluster configurations in experiments with Bayesian networks of different sizes and structures.
Mathematics,Computer Science
What problem does this paper attempt to address?