Abstract:Binary Neural Networks (BNNs) are showing tremendous success on realistic image classification tasks. Notably, their accuracy is similar to the state-of-the-art accuracy obtained by full-precision models tailored to edge devices. In this regard, BNNs are very amenable to edge devices since they employ 1-bit to store the inputs and weights, and thus, their storage requirements are low. Also, BNNs computations are mainly done using xnor and pop-counts operations which are implemented very efficiently using simple hardware structures. Nonetheless, supporting BNNs efficiently on mobile CPUs is far from trivial since their benefits are hindered by frequent memory accesses to load weights and inputs. In BNNs, a weight or an input is stored using one bit, and aiming to increase storage and computation efficiency, several of them are packed together as a sequence of bits. In this work, we observe that the number of unique sequences representing a set of weights is typically low. Also, we have seen that during the evaluation of a BNN layer, a small group of unique sequences is employed more frequently than others. Accordingly, we propose exploiting this observation by using Huffman Encoding to encode the bit sequences and then using an indirection table to decode them during the BNN evaluation. Also, we propose a clustering scheme to identify the most common sequences of bits and replace the less common ones with some similar common sequences. Hence, we decrease the storage requirements and memory accesses since common sequences are encoded with fewer bits. We extend a mobile CPU by adding a small hardware structure that can efficiently cache and decode the compressed sequence of bits. We evaluate our scheme using the ReAacNet model with the Imagenet dataset. Our experimental results show that our technique can reduce memory requirement by 1.32x and improve performance by 1.35x.

O3BNN-R: An Out-of-Order Architecture for High-Performance and Regularized BNN Inference

Class-Aware Pruning for Efficient Neural Networks

All-in-One: A Highly Representative DNN Pruning Framework for Edge Devices with Dynamic Power Management

An Approach of Binary Neural Network Energy-Efficient Implementation

FP-BNN: Binarized neural network on FPGA

TCP-Net: Minimizing Operation Counts of Binarized Neural Network Inference.

Binary Early-Exit Network for Adaptive Inference on Low-Resource Devices

Exploiting Kernel Compression on BNNs

A high-throughput scalable BNN accelerator with fully pipelined architecture

A&B BNN: Add&Bit-Operation-Only Hardware-Friendly Binary Neural Network

FCA-BNN: Flexible and Configurable Accelerator for Binarized Neural Networks on FPGA.

Bisection Neural Network Toward Reconfigurable Hardware Implementation

Cloud–Edge Collaborative Inference with Network Pruning

Reconfigurable Binary Neural Network Accelerator with Adaptive Parallelism Scheme

FOBNN: Fast Oblivious Binarized Neural Network Inference

Binary Neural Networks as a general-propose compute paradigm for on-device computer vision

BitXpro: Regularity-Aware Hardware Runtime Pruning for Deep Neural Networks

FracBNN: Accurate and FPGA-Efficient Binary Neural Networks with Fractional Activations.

Only Train Once: A One-Shot Neural Network Training And Pruning Framework

Accelerating Neural-ODE Inference on FPGAs with Two-Stage Structured Pruning and History-based Stepsize Search.

An Efficient Channel-Aware Sparse Binarized Neural Networks Inference Accelerator