Abstract:Exploding development of convolutional neural network (CNN) benefits greatly from the hardware-based acceleration to maintain low latency and high utilization of resources. To enhance the processing efficiency of CNN algorithms, Field Programming Gate Array (FPGA)-based accelerators are designed with increased hardware resources to achieve high parallelism and throughput. However, there exist bottlenecks when more processing elements (PEs) in the form of PE clusters are introduced, including 1) the under-utilization of FPGA's fixed hardware resources, which leads to the effective and peak performance mismatch; and 2) the limited clock frequency caused by the sophisticated routing and complex placement. In this paper, a 2-level hierarchical Network-on-Chip (NoC)-based CNN accelerator is proposed. In the upper level, a mesh-based NoC that interconnects multiple PE clusters is introduced. Such a design not only provides increased flexibility to balance different data communication models for better PE utilization and energy efficiency but also enables globally asynchronous, locally synchronous (GALS) architecture for better timing closure. At the lower level, local PEs are organized into a 3D-tiled PE cluster aiming to maximize the data reuse exploiting inherent dataflow of the convolution networks. Implementation and experiments on Xilinx ZU9EG FPGA for 4 benchmark CNN models: ResNet50, ResNet34, VGG16, and Darknet19 show that our work operates at a frequency of 300 MHz and delivers an effective throughput of 0.998 TOPS, 1.022 TOPS, 1.024 TOPS, and 1.026 TOPS. This result corresponds to 92.85%, 95.1%, 95.25%, and 95.46% PE utilization. Compared with the related FPGA-based de

Energy-Efficient And High-Throughput Fpga-Based Accelerator For Convolutional Neural Networks

A Convolutional Neural Network Accelerator Architecture with Fine-Granular Mixed Precision Configurability.

Throughput-Optimized FPGA Accelerator for Deep Convolutional Neural Networks

Optimizing FPGA-based Accelerator Design for Deep Convolutional Neural Networks

A Power-Efficient and High Performance FPGA Accelerator for Convolutional Neural Networks: Work-in-progress.

FPGA-based Accelerator for Convolutional Neural Network

A FPGA-based Hardware Accelerator for Multiple Convolutional Neural Networks

Design and Implementation of a FPGA-based Accelerator for Convolutional Neural Networks

A Memory-Optimized and Energy-Efficient CNN Acceleration Architecture Based on FPGA.

Throughput-Optimized Opencl-Based Fpga Accelerator For Large-Scale Convolutional Neural Networks

An Efficient Sparse CNNs Accelerator on FPGA

A High Performance FPGA-based Accelerator for Large-Scale Convolutional Neural Networks

A High-Performance Accelerator for Large-Scale Convolutional Neural Networks

An FPGA-Based CNN Accelerator Integrating Depthwise Separable Convolution

Scalable FPGA-Based Convolutional Neural Network Accelerator for Embedded Systems

An Efficient CNN Accelerator for Pattern-Compressed Sparse Neural Networks on FPGA

High Throughput CNN Accelerator Design Based on FPGA

An Efficient Convolutional Neural Network Accelerator on FPGA

FPGA-Based High-Throughput CNN Hardware Accelerator With High Computing Resource Utilization Ratio

Communication-Aware and Resource-Efficient NoC-Based Architecture for CNN Acceleration

Improving the computational efficiency and flexibility of FPGA-based CNN accelerator through loop optimization