Abstract:Exploding development of convolutional neural network (CNN) benefits greatly from the hardware-based acceleration to maintain low latency and high utilization of resources. To enhance the processing efficiency of CNN algorithms, Field Programming Gate Array (FPGA)-based accelerators are designed with increased hardware resources to achieve high parallelism and throughput. However, there exist bottlenecks when more processing elements (PEs) in the form of PE clusters are introduced, including 1) the under-utilization of FPGA's fixed hardware resources, which leads to the effective and peak performance mismatch; and 2) the limited clock frequency caused by the sophisticated routing and complex placement. In this paper, a 2-level hierarchical Network-on-Chip (NoC)-based CNN accelerator is proposed. In the upper level, a mesh-based NoC that interconnects multiple PE clusters is introduced. Such a design not only provides increased flexibility to balance different data communication models for better PE utilization and energy efficiency but also enables globally asynchronous, locally synchronous (GALS) architecture for better timing closure. At the lower level, local PEs are organized into a 3D-tiled PE cluster aiming to maximize the data reuse exploiting inherent dataflow of the convolution networks. Implementation and experiments on Xilinx ZU9EG FPGA for 4 benchmark CNN models: ResNet50, ResNet34, VGG16, and Darknet19 show that our work operates at a frequency of 300 MHz and delivers an effective throughput of 0.998 TOPS, 1.022 TOPS, 1.024 TOPS, and 1.026 TOPS. This result corresponds to 92.85%, 95.1%, 95.25%, and 95.46% PE utilization. Compared with the related FPGA-based de

Power Efficient Tiny Yolo CNN Using Reduced Hardware Resources Based on Booth Multiplier and WALLACE Tree Adders

A Convolutional Neural Network Accelerator Architecture with Fine-Granular Mixed Precision Configurability.

Design of Efficient CNN Accelerator Based on Zynq Platform.

A high-speed reusable quantized hardware accelerator design for CNN on constrained edge device

Optimization for Efficient Hardware Implementation of CNN on FPGA

An Efficient Streaming Accelerator for Low Bit-Width Convolutional Neural Networks

Smilodon: an Efficient Accelerator for Low Bit-Width CNNs with Task Partitioning

A Solution To Optimize Multi-Operand Adders In Cnn Architecture On Fpga

Design and Implementation of YOLOv3-Tiny Accelerator Based on PYNQ-Z2 Heterogeneous Platform

Hardware-Efficient Template-Based Deep CNNs Accelerator Design

High-Speed CNN Accelerator SoC Design Based on a Flexible Diagonal Cyclic Array

A CNN Hardware Accelerator Designed for YOLO Algorithm Based on RISC-V SoC

CNN hardware acceleration on a low-power and low-cost APSoC

Optimizing FPGA-based Accelerator Design for Deep Convolutional Neural Networks

Efficient Hardware Architectures for Deep Convolutional Neural Network

An Efficient Sparse CNNs Accelerator on FPGA

Communication-Aware and Resource-Efficient NoC-Based Architecture for CNN Acceleration

A FPGA-based Hardware Accelerator for Multiple Convolutional Neural Networks

YodaNN: An Ultra-Low Power Convolutional Neural Network Accelerator Based on Binary Weights

A High-Efficient and Configurable Hardware Accelerator for Convolutional Neural Network