Abstract:Exploding development of convolutional neural network (CNN) benefits greatly from the hardware-based acceleration to maintain low latency and high utilization of resources. To enhance the processing efficiency of CNN algorithms, Field Programming Gate Array (FPGA)-based accelerators are designed with increased hardware resources to achieve high parallelism and throughput. However, there exist bottlenecks when more processing elements (PEs) in the form of PE clusters are introduced, including 1) the under-utilization of FPGA's fixed hardware resources, which leads to the effective and peak performance mismatch; and 2) the limited clock frequency caused by the sophisticated routing and complex placement. In this paper, a 2-level hierarchical Network-on-Chip (NoC)-based CNN accelerator is proposed. In the upper level, a mesh-based NoC that interconnects multiple PE clusters is introduced. Such a design not only provides increased flexibility to balance different data communication models for better PE utilization and energy efficiency but also enables globally asynchronous, locally synchronous (GALS) architecture for better timing closure. At the lower level, local PEs are organized into a 3D-tiled PE cluster aiming to maximize the data reuse exploiting inherent dataflow of the convolution networks. Implementation and experiments on Xilinx ZU9EG FPGA for 4 benchmark CNN models: ResNet50, ResNet34, VGG16, and Darknet19 show that our work operates at a frequency of 300 MHz and delivers an effective throughput of 0.998 TOPS, 1.022 TOPS, 1.024 TOPS, and 1.026 TOPS. This result corresponds to 92.85%, 95.1%, 95.25%, and 95.46% PE utilization. Compared with the related FPGA-based de

Efficient fast convolution architectures for convolutional neural network

A Convolutional Neural Network Accelerator Architecture with Fine-Granular Mixed Precision Configurability.

Efficient Hardware Architectures for Deep Convolutional Neural Network

A High Efficient Architecture for Convolution Neural Network Accelerator

A Scalable 3D Array Architecture for Accelerating Convolutional Neural Networks

A High-Efficient and Configurable Hardware Accelerator for Convolutional Neural Network

A Parallel Loading Based Accelerator for Convolution Neural Network

An Efficient Streaming Accelerator for Low Bit-Width Convolutional Neural Networks

A Reconfigurable Spatial Architecture for Energy-Efficient Inception Neural Networks

A High Performance Reconfigurable Hardware Architecture for Lightweight Convolutional Neural Network

A Reconfigurable Accelerator for Sparse Convolutional Neural Networks.

A High Performance FPGA-based Accelerator for Large-Scale Convolutional Neural Networks

A FPGA-based Hardware Accelerator for Multiple Convolutional Neural Networks

Efficient Inference of Large-Scale and Lightweight Convolutional Neural Networks on FPGA

An Efficient Accelerator for Multiple Convolutions From the Sparsity Perspective

A High-Performance Systolic Array Accelerator Dedicated for CNN.

A Novel Low-Communication Energy-Efficient Reconfigurable CNN Acceleration Architecture

A high-speed reusable quantized hardware accelerator design for CNN on constrained edge device

Energy-Efficient Architecture for FPGA-based Deep Convolutional Neural Networks with Binary Weights

Communication-Aware and Resource-Efficient NoC-Based Architecture for CNN Acceleration