Abstract:Abstract This document addresses some inherent problems in Machine Learning (ML), such as the high computational and energy costs associated with their implementation on IoT devices. It aims to study and analyze the performance and efficiency of quantization as an optimization method, as well as the possibility of training ML models directly on an IoT device. Quantization involves reducing the precision of model weights and activations while still maintaining acceptable levels of accuracy. Using representative networks for facial recognition developed with TensorFlow and TensorRT, Post-Training Quantization and Quantization-Aware Training are employed to reduce computational load and improve energy efficiency. The computational experience was conducted on a general-purpose computer featuring an Intel i7-1260P processor and an NVIDIA RTX 3080 graphics card used as an accelerator. Additionally, a NVIDIA Jetson AGX Orin was used as an example of an IoT device. We analyze the feasibility of training on an IoT device, the impact of quantization optimization on knowledge transfer-trained models and evaluate the differences between Post-Training Quantization and Quantization-Aware Training in such networks on different devices. Furthermore, the performance and efficiency of NVIDIA’s inference accelerator (Deep Learning Accelerator - DLA, in its 2.0 version) available at the Jetson Orin architecture are studied. We concluded that the Jetson device is capable of performing training on its own. The IoT device can achieve inference performance similar to that of the more powerful processor, thanks to the optimization process, with better energy efficiency. Post-Training Quantization has shown better performance, while Quantization-Aware Training has demonstrated higher energy efficiency. However, since the accelerator cannot execute certain layers of the models, the use of DLA worsens both the performance and efficiency results.

Research on Model Compression for Embedded Platform through Quantization and Pruning

Single-shot Pruning and Quantization for Hardware-Friendly Neural Network Acceleration

Improved Model Compression Method Based on Information Entropy

MCMC: Multi-Constrained Model Compression Via One-Stage Envelope Reinforcement Learning.

A Compression Pipeline for One-Stage Object Detection Model

Model Compression for Deep Neural Networks: A Survey

To Compress, or Not to Compress: Characterizing Deep Learning Model Compression for Embedded Inference

Automated Model Compression by Jointly Applied Pruning and Quantization

Edge AI: Evaluation of Model Compression Techniques for Convolutional Neural Networks

Deep Compression: Compressing Deep Neural Networks with Pruning, Trained Quantization and Huffman Coding

Downscaling and Overflow-aware Model Compression for Efficient Vision Processors

CAQ: Toward Context-Aware and Self-Adaptive Deep Model Computation for AIoT Applications

Pruning and quantization for deep neural network acceleration: A survey

Comprehensive Study on Performance Evaluation and Optimization of Model Compression: Bridging Traditional Deep Learning and Large Language Models

Quantisation and Pruning for Neural Network Compression and Regularisation

Design of a Novel Neural Network Compression Method for Tiny Machine Learning

OPQ: Compressing Deep Neural Networks with One-shot Pruning-Quantization

Deep Model Compression and Architecture Optimization for Embedded Systems: A Survey

Pruning at a Glance: Global Neural Pruning for Model Compression

Model Compression for Resource-Constrained Mobile Robots

Optimizing convolutional neural networks for IoT devices: performance and energy efficiency of quantization techniques