Abstract:Binary Neural Networks (BNNs) are showing tremendous success on realistic image classification tasks. Notably, their accuracy is similar to the state-of-the-art accuracy obtained by full-precision models tailored to edge devices. In this regard, BNNs are very amenable to edge devices since they employ 1-bit to store the inputs and weights, and thus, their storage requirements are low. Also, BNNs computations are mainly done using xnor and pop-counts operations which are implemented very efficiently using simple hardware structures. Nonetheless, supporting BNNs efficiently on mobile CPUs is far from trivial since their benefits are hindered by frequent memory accesses to load weights and inputs. In BNNs, a weight or an input is stored using one bit, and aiming to increase storage and computation efficiency, several of them are packed together as a sequence of bits. In this work, we observe that the number of unique sequences representing a set of weights is typically low. Also, we have seen that during the evaluation of a BNN layer, a small group of unique sequences is employed more frequently than others. Accordingly, we propose exploiting this observation by using Huffman Encoding to encode the bit sequences and then using an indirection table to decode them during the BNN evaluation. Also, we propose a clustering scheme to identify the most common sequences of bits and replace the less common ones with some similar common sequences. Hence, we decrease the storage requirements and memory accesses since common sequences are encoded with fewer bits. We extend a mobile CPU by adding a small hardware structure that can efficiently cache and decode the compressed sequence of bits. We evaluate our scheme using the ReAacNet model with the Imagenet dataset. Our experimental results show that our technique can reduce memory requirement by 1.32x and improve performance by 1.35x.

Run-Time Efficient RNN Compression for Inference on Edge Devices

Condense: A Framework for Device and Frequency Adaptive Neural Network Models on the Edge.

Pushing the limits of RNN Compression

MCMC: Multi-Constrained Model Compression Via One-Stage Envelope Reinforcement Learning.

A 3.89-Gops/mw Scalable Recurrent Neural Network Processor with Improved Efficiency on Memory and Computation

Rank and run-time aware compression of NLP Applications

Exploiting Symmetric Temporally Sparse BPTT for Efficient RNN Training

MPDCompress - Matrix Permutation Decomposition Algorithm for Deep Neural Network Compression

Lightweight compression of neural network feature tensors for collaborative intelligence

Deep Compression: Compressing Deep Neural Networks with Pruning, Trained Quantization and Huffman Coding

Lightweight Compression of Intermediate Neural Network Features for Collaborative Intelligence

Compression of Recurrent Neural Networks using Matrix Factorization

To Compress, or Not to Compress: Characterizing Deep Learning Model Compression for Embedded Inference

Edge AI: Evaluation of Model Compression Techniques for Convolutional Neural Networks

Conv-inheritance: A hardware-efficient method to compress convolutional neural networks for edge applications

Tensor train decompositions on recurrent networks

A deep neural network compression algorithm based on knowledge transfer for edge devices

Incremental Training and Group Convolution Pruning for Runtime DNN Performance Scaling on Heterogeneous Embedded Platforms

Multi-Tree Compact Hierarchical Tensor Recurrent Neural Networks for Intelligent Transportation System Edge Devices

RNC: Efficient RRAM-aware NAS and Compilation for DNNs on Resource-Constrained Edge Devices

Exploiting Kernel Compression on BNNs