Abstract:As a new area of machine learning research, the deep learning algorithm has attracted a lot of attention from the research community. It may bring human beings to a higher cognitive level of data. Its unsupervised pre-training step allows us to find high-dimensional representations or abstract features which work much better than the principal component analysis (PCA) method. However, it will face problems when being applied to deal with large scale data due to its intensive computation from many levels of training process against large scale data. The sequential deep learning algorithms usually can not finish the computation in an acceptable time. In this paper, we propose a many-core algorithm which is based on a parallel method and is used in the Intel Xeon Phi many-core systems to speed up the unsupervised training process of Sparse Autoencoder and Restricted Boltzmann Machine (RBM). Using the sequential training algorithm as a baseline to compare, we adopted several optimization methods to parallelize the algorithm. The experimental results show that our fully-optimized algorithm gains more than 300-fold speedup on parallelized Sparse Autoencoder compared with the original sequential algorithm on the Intel Xeon Phi coprocessor. Also, we ran the fully-optimized code on both the Intel Xeon Phi coprocessor and an expensive Intel Xeon CPU. Our method on the Intel Xeon Phi coprocessor is 7 to 10 times faster than the Intel Xeon CPU for this application. In addition to this, we compared our fully-optimized code on the Intel Xeon Phi with a Matlab code running on single Intel Xeon CPU. Our method on the Intel Xeon Phi runs 16 times faster than the Matlab implementation. The result also suggests that the Intel Xeon Phi can offer an efficient but more general-purposed way to parallelize the deep learning algorithm compared to GPU. It also achieves faster speed with better parallelism than the Intel Xeon CPU.

Optimizing the MapReduce Framework on Intel Xeon Phi Coprocessor

An Efficient MapReduce Framework for Intel MIC Cluster.

Optimizing the MapReduce Framework for CPU-MIC Heterogeneous Cluster

Optimizing and Auto-Tuning Scale-Free Sparse Matrix-Vector Multiplication on Intel Xeon Phi

Optimizing Protein Folding Simulation on Intel Xeon Phi

Towards Modeling Energy Consumption of Xeon Phi

NO2: Speeding Up Parallel Processing of Massive Compute-Intensive Tasks

Cluster-level tuning of a shallow water equation solver on the Intel MIC architecture

FPMR: MapReduce framework on FPGA.

Training Large Scale Deep Neural Networks on the Intel Xeon Phi Many-Core Coprocessor

Deep and Shallow convections in Atmosphere Models on Intel Xeon Phi Coprocessor Systems

Micmr: an Efficient Mapreduce Framework for Cpu-Mic Heterogeneous Architecture

NUMERICAL SIMULATION OF PLANETARY FLUID DYNAMICS ON CPU-MIC HETEROGENEOUS MANY-CORE SYSTEMS

Memory Efficient Two-Pass 3D FFT Algorithm for Intel® Xeon Phi TM Coprocessor

Towards co-designed optimizations in parallel frameworks: A MapReduce case study

Fpmr: Map Reduce Framework on Fpga A Case Study of Rankboost Acceleration

The performance of MapReduce: an in-depth study

The Performance of MapReduce

Optimizing MapReduce for Highly Distributed Environments

LAMMPS' PPPM Long-Range Solver for the Second Generation Xeon Phi

Accelerating Gravitational Microlensing Simulations Using the Xeon Phi Coprocessor