Abstract:In the traditional deep compression framework, iteratively performing network pruning and quantization can reduce the model size and computation cost to meet the deployment requirements. However, such a step-wise application of pruning and quantization may lead to suboptimal solutions and unnecessary time consumption. In this paper, we tackle this issue by integrating network pruning and quantization as a unified joint compression problem and then use AutoML to automatically solve it. We find the pruning process can be regarded as the channel-wise quantization with 0 bit. Thus, the separate two-step pruning and quantization can be simplified as the one-step quantization with mixed precision. This unification not only simplifies the compression pipeline but also avoids the compression divergence. To implement this idea, we propose the automated model compression by jointly applied pruning and quantization (AJPQ). AJPQ is designed with a hierarchical architecture: the layer controller controls the layer sparsity, and the channel controller decides the bit-width for each kernel. Following the same importance criterion, the layer controller and the channel controller collaboratively decide the compression strategy. With the help of reinforcement learning, our one-step compression is automatically achieved. Compared with the state-of-the-art automated compression methods, our method obtains a better accuracy while reducing the storage considerably. For fixed precision quantization, AJPQ can reduce more than five times model size and two times computation with a slight performance increase for Skynet in remote sensing object detection. When mixed-precision is allowed, AJPQ can reduce five times model size with only 1.06% top-5 accuracy decline for MobileNet in the classification task.

AMC: AutoML for Model Compression and Acceleration on Mobile Devices

MCMC: Multi-Constrained Model Compression Via One-Stage Envelope Reinforcement Learning.

Improved Model Compression Method Based on Information Entropy

MetaAMC: Meta Learning and AutoML for Model Compression

AutoMC: Automated Model Compression based on Domain Knowledge and Progressive search strategy

A CNN Compression Method Via Dynamic Channel Ranking Strategy

DeepRebirth: Accelerating Deep Neural Network Execution on Mobile Devices

Compression of Deep Convolutional Neural Networks for Fast and Low Power Mobile Applications

Model Compression for Deep Neural Networks: A Survey

Safety and Performance, Why not Both? Bi-Objective Optimized Model Compression toward AI Software Deployment

Deep Learning Model Compression with Rank Reduction in Tensor Decomposition.

MobilePrune: Neural Network Compression via l(0) Sparse Group Lasso on the Mobile System

Automated Model Compression by Jointly Applied Pruning and Quantization

Gated Compression Layers for Efficient Always-On Models

To Compress, or Not to Compress: Characterizing Deep Learning Model Compression for Embedded Inference

ZipNN: Lossless Compression for AI Models

Compress and Compare: Interactively Evaluating Efficiency and Behavior Across ML Model Compression Experiments

Model Compression in Practice: Lessons Learned from Practitioners Creating On-device Machine Learning Experiences

Analysis of Model Compression Using Knowledge Distillation

Deep Compression: Compressing Deep Neural Networks with Pruning, Trained Quantization and Huffman Coding

Fast Hybrid Search for Automatic Model Compression