Abstract:Traffic sign recognition is a very important computer vision task for a number of real-world applications such as intelligent transportation surveillance and analysis. While deep neural networks have been demonstrated in recent years to provide state-of-the-art performance traffic sign recognition, a key challenge for enabling the widespread deployment of deep neural networks for embedded traffic sign recognition is the high computational and memory requirements of such networks. As a consequence, there are significant benefits in investigating compact deep neural network architectures for traffic sign recognition that are better suited for embedded devices. In this paper, we introduce MicronNet, a highly compact deep convolutional neural network for real-time embedded traffic sign recognition designed based on macroarchitecture design principles (e.g., spectral macroarchitecture augmentation, parameter precision optimization, etc.) as well as numerical microarchitecture optimization strategies. The resulting overall architecture of MicronNet is thus designed with as few parameters and computations as possible while maintaining recognition performance, leading to optimized information density of the proposed network. The resulting MicronNet possesses a model size of just ~1MB and ~510,000 parameters (~27x fewer parameters than state-of-the-art) while still achieving a human performance level top-1 accuracy of 98.9% on the German traffic sign recognition benchmark. Furthermore, MicronNet requires just ~10 million multiply-accumulate operations to perform inference, and has a time-to-compute of just 32.19 ms on a Cortex-A53 high efficiency processor. These experimental results show that highly compact, optimized deep neural network architectures can be designed for real-time traffic sign recognition that are well-suited for embedded scenarios.

UL-CNN: an Ultra-Lightweight Convolutional Neural Network Aiming at Flash-Based Computing-In-Memory Architecture for Pedestrian Recognition.

DaDianNao: A Machine-Learning Supercomputer

A 40nm 1mb 35.6 TOPS/W MLC NOR-Flash Based Computation-in-Memory Structure for Machine Learning

In-Memory Computing Integrated Structure Circuit Based on Nonvolatile Flash Memory Unit

MicronNet: A Highly Compact Deep Convolutional Neural Network Architecture for Real-time Embedded Traffic Sign Classification

Optimized operation scheme of flash-memory-based neural network online training with ultra-high endurance

Flash-Based Content Addressable Memory with L2 Distance for Memory-Augmented Neural Network

Designing Extremely Memory-Efficient CNNs for On-device Vision Tasks

LR-CNN: Lightweight Row-centric Convolutional Neural Network Training for Memory Reduction

Flash: A Hybrid Private Inference Protocol for Deep CNNs with High Accuracy and Low Latency on CPU

Hardware Implementation of Energy Efficient Deep Learning Neural Network Based on Nanoscale Flash Computing Array

CINT -- an Energy-efficient Mixed-signal In-Memory CNN Accelerator Based on NOR Flash Memory (Poster)

Design of Computing-in-Memory (CIM) with Vertical Split-Gate Flash Memory for Deep Neural Network (DNN) Inference Accelerator

SRAM-Based CIM Architecture Design for Event Detection

MCUNetV2: Memory-Efficient Patch-based Inference for Tiny Deep Learning

A Memory-Optimized and Energy-Efficient CNN Acceleration Architecture Based on FPGA.

A Weight-Reload-Eliminated Compute-in-Memory Accelerator for 60 fps 4K Super-Resolution

Improved JPEG Lossless Compression for Compression of Intermediate Layers in Neural Networks Based on Compute-In-Memory

Face Recognition Based on Lightweight Convolutional Neural Networks

A Vision Chip Architecture for Image Recognition Based on Improved Convolutional Neural Network

Memory-efficient Deep Learning Inference with Incremental Weight Loading and Data Layout Reorganization on Edge Systems.