Abstract:The ever-increasing number of layers, millions of parameters, and large data volume make deep learning workloads resource-intensive and power-hungry. In this paper, we develop a convolutional neural network (CNN) acceleration framework, named MLCNN, which explores algorithm-hardware co-design to achieve cross-layer cooperative optimization and acceleration. MLCNN dramatically reduces computation and on-off chip communication, improving CNN's performance. To achieve this, MLCNN reorders the position of nonlinear activation layers and pooling layers, which we prove results in a negligible accuracy loss; then the convolutional layer and pooling layer are co-optimized by means of redundant multiplication elimination, local addition reuse, and global addition reuse. To the best of our knowledge, MLCNN is the first of its kind that incorporates cooperative optimization across convolutional, activation, and pooling layers. We further customize the MLCNN accelerator to take full advantage of cross-layer CNN optimization to reduce both computation and on-off chip communication. Our analysis shows that MLCNN can significantly reduce (up to 98%) multiplications and additions. We have implemented a prototype of MLCNN and evaluated its performance on several widely used CNN models using both an accelerator-level cycle and energy model and RTL implementation. Experimental results show that MLCNN achieves 3.2x speedup and 2.9x energy efficiency compared with dense CNNs. MLCNN's optimization methods are orthogonal to other CNN acceleration techniques, such as quantization and pruning. Combined with quantization, our quantized MLCNN gains a 12.8x speedup and 11.3x energy efficiency compared with DCNN.

Reduce Computational Complexity for Convolutional Layers by Skipping Zeros

Deep Neural Network Acceleration with Sparse Prediction Layers

MLCNN: Cross-Layer Cooperative Optimization and Accelerator Architecture for Speeding Up Deep Learning Applications

A Convolutional Neural Network Accelerator Architecture with Fine-Granular Mixed Precision Configurability.

Accelerating Convolutional Neural Networks by Removing Interspatial and Interkernel Redundancies.

Layer-Wise Training To Create Efficient Convolutional Neural Networks

LACS: A High-Computational-Efficiency Accelerator for CNNs

An Efficient Accelerator for Multiple Convolutions From the Sparsity Perspective

Kernel-Segregated Transpose Convolution Operation

A Sparse CNN Accelerator for Eliminating Redundant Computations in Intra- and Inter-Convolutional/Pooling Layers

Comb Convolution for Efficient Convolutional Architecture

Accelerating convolutional neural network by exploiting sparsity on GPUs

Optimizing GPU Memory Transactions for Convolution Operations

Zero and data reuse-aware fast convolution for deep neural networks on GPU

A Pre-defined Sparse Kernel Based Convolution for Deep CNNs

Learning Efficient Convolutional Networks Through Network Slimming.

SlimConv: Reducing Channel Redundancy in Convolutional Neural Networks by Features Recombining

TabConv: Low-Computation CNN Inference via Table Lookups

KCNN: Kernel-wise Quantization to Remarkably Decrease Multiplications in Convolutional Neural Network.

High Performance Zero-Memory Overhead Direct Convolutions

A High Efficient Architecture for Convolution Neural Network Accelerator