Abstract:As a new area of machine learning research, the deep learning algorithm has attracted a lot of attention from the research community. It may bring human beings to a higher cognitive level of data. Its unsupervised pre-training step allows us to find high-dimensional representations or abstract features which work much better than the principal component analysis (PCA) method. However, it will face problems when being applied to deal with large scale data due to its intensive computation from many levels of training process against large scale data. The sequential deep learning algorithms usually can not finish the computation in an acceptable time. In this paper, we propose a many-core algorithm which is based on a parallel method and is used in the Intel Xeon Phi many-core systems to speed up the unsupervised training process of Sparse Autoencoder and Restricted Boltzmann Machine (RBM). Using the sequential training algorithm as a baseline to compare, we adopted several optimization methods to parallelize the algorithm. The experimental results show that our fully-optimized algorithm gains more than 300-fold speedup on parallelized Sparse Autoencoder compared with the original sequential algorithm on the Intel Xeon Phi coprocessor. Also, we ran the fully-optimized code on both the Intel Xeon Phi coprocessor and an expensive Intel Xeon CPU. Our method on the Intel Xeon Phi coprocessor is 7 to 10 times faster than the Intel Xeon CPU for this application. In addition to this, we compared our fully-optimized code on the Intel Xeon Phi with a Matlab code running on single Intel Xeon CPU. Our method on the Intel Xeon Phi runs 16 times faster than the Matlab implementation. The result also suggests that the Intel Xeon Phi can offer an efficient but more general-purposed way to parallelize the deep learning algorithm compared to GPU. It also achieves faster speed with better parallelism than the Intel Xeon CPU.

ImageNet Training in Minutes

ImageNet Training in Minutes

Highly Scalable Deep Learning Training System with Mixed-Precision: Training ImageNet in Four Minutes

Accurate, Large Minibatch SGD: Training ImageNet in 1 Hour

Extremely Large Minibatch SGD: Training ResNet-50 on ImageNet in 15 Minutes

Training EfficientNets at Supercomputer Scale: 83% ImageNet Top-1 Accuracy in One Hour

Training Multiscale-CNN for Large Microscopy Image Classification in One Hour

Optimizing Network Performance for Distributed DNN Training on GPU Clusters: ImageNet/AlexNet Training in 1.5 Minutes

Accelerating Neural Network Training: A Brief Review

Large Batch Training of Convolutional Networks

AccEPT: an Acceleration Scheme for Speeding Up Edge Pipeline-parallel Training

Speeding Up Image Classifiers with Little Companions

Rapid-INR: Storage Efficient CPU-free DNN Training Using Implicit Neural Representation

A Variable Batch Size Strategy for Large Scale Distributed DNN Training

Distributed Training Large-Scale Deep Architectures

CATERPILLAR: Coarse Grain Reconfigurable Architecture for Accelerating the Training of Deep Neural Networks

Training Large Scale Deep Neural Networks on the Intel Xeon Phi Many-Core Coprocessor

An efficient approach to escalate the speed of training convolution neural networks

PowerAI DDL

Pipelined Backpropagation at Scale: Training Large Models without Batches

FastHebb: Scaling Hebbian Training of Deep Neural Networks to ImageNet Level