Abstract:Abstract This document addresses some inherent problems in Machine Learning (ML), such as the high computational and energy costs associated with their implementation on IoT devices. It aims to study and analyze the performance and efficiency of quantization as an optimization method, as well as the possibility of training ML models directly on an IoT device. Quantization involves reducing the precision of model weights and activations while still maintaining acceptable levels of accuracy. Using representative networks for facial recognition developed with TensorFlow and TensorRT, Post-Training Quantization and Quantization-Aware Training are employed to reduce computational load and improve energy efficiency. The computational experience was conducted on a general-purpose computer featuring an Intel i7-1260P processor and an NVIDIA RTX 3080 graphics card used as an accelerator. Additionally, a NVIDIA Jetson AGX Orin was used as an example of an IoT device. We analyze the feasibility of training on an IoT device, the impact of quantization optimization on knowledge transfer-trained models and evaluate the differences between Post-Training Quantization and Quantization-Aware Training in such networks on different devices. Furthermore, the performance and efficiency of NVIDIA’s inference accelerator (Deep Learning Accelerator - DLA, in its 2.0 version) available at the Jetson Orin architecture are studied. We concluded that the Jetson device is capable of performing training on its own. The IoT device can achieve inference performance similar to that of the more powerful processor, thanks to the optimization process, with better energy efficiency. Post-Training Quantization has shown better performance, while Quantization-Aware Training has demonstrated higher energy efficiency. However, since the accelerator cannot execute certain layers of the models, the use of DLA worsens both the performance and efficiency results.

Optimizing convolutional neural networks for IoT devices: performance and energy efficiency of quantization techniques

Hessian-based Mixed-Precision Quantization with Transition Aware Training for Neural Networks

Single-shot Pruning and Quantization for Hardware-Friendly Neural Network Acceleration

Deploy Large-Scale Deep Neural Networks in Resource Constrained IoT Devices with Local Quantization Region

Multi-Component Optimization and Efficient Deployment of Neural-Networks on Resource-Constrained IoT Hardware

On-Device Training Under 256KB Memory

Low Power Inference for On-Device Visual Recognition with a Quantization-Friendly Solution.

Zac: Towards Automatic Optimization and Deployment of Quantized Deep Neural Networks on Embedded Devices

Quantization and Deployment of Deep Neural Networks on Microcontrollers

Performance Characterization of using Quantization for DNN Inference on Edge Devices: Extended Version

Optimized CNN Architectures Benchmarking in Hardware-Constrained Edge Devices in IoT Environments

On-Chip Hardware-Aware Quantization for Mixed Precision Neural Networks

DNN Memory Footprint Reduction via Post-Training Intra-Layer Multi-Precision Quantization

Design of a Novel Neural Network Compression Method for Tiny Machine Learning

DNN Model Compression for IoT Domain-Specific Hardware Accelerators

Towards Efficient Compact Network Training on Edge-Devices

On-Device Training of Fully Quantized Deep Neural Networks on Cortex-M Microcontrollers

To Compress, or Not to Compress: Characterizing Deep Learning Model Compression for Embedded Inference

Post-Training Non-Uniform Quantization for Convolutional Neural Networks

Dataflow-Based Joint Quantization for Deep Neural Networks

Incremental Training and Group Convolution Pruning for Runtime DNN Performance Scaling on Heterogeneous Embedded Platforms