Abstract:Exploding development of convolutional neural network (CNN) benefits greatly from the hardware-based acceleration to maintain low latency and high utilization of resources. To enhance the processing efficiency of CNN algorithms, Field Programming Gate Array (FPGA)-based accelerators are designed with increased hardware resources to achieve high parallelism and throughput. However, there exist bottlenecks when more processing elements (PEs) in the form of PE clusters are introduced, including 1) the under-utilization of FPGA's fixed hardware resources, which leads to the effective and peak performance mismatch; and 2) the limited clock frequency caused by the sophisticated routing and complex placement. In this paper, a 2-level hierarchical Network-on-Chip (NoC)-based CNN accelerator is proposed. In the upper level, a mesh-based NoC that interconnects multiple PE clusters is introduced. Such a design not only provides increased flexibility to balance different data communication models for better PE utilization and energy efficiency but also enables globally asynchronous, locally synchronous (GALS) architecture for better timing closure. At the lower level, local PEs are organized into a 3D-tiled PE cluster aiming to maximize the data reuse exploiting inherent dataflow of the convolution networks. Implementation and experiments on Xilinx ZU9EG FPGA for 4 benchmark CNN models: ResNet50, ResNet34, VGG16, and Darknet19 show that our work operates at a frequency of 300 MHz and delivers an effective throughput of 0.998 TOPS, 1.022 TOPS, 1.024 TOPS, and 1.026 TOPS. This result corresponds to 92.85%, 95.1%, 95.25%, and 95.46% PE utilization. Compared with the related FPGA-based de

A FPGA-based Hardware Accelerator for Multiple Convolutional Neural Networks

A Convolutional Neural Network Accelerator Architecture with Fine-Granular Mixed Precision Configurability.

FPGA-based Accelerator for Convolutional Neural Network

Design and Implementation of a FPGA-based Accelerator for Convolutional Neural Networks

A High Performance FPGA-based Accelerator for Large-Scale Convolutional Neural Networks

A High-Performance Accelerator for Large-Scale Convolutional Neural Networks

An Efficient Sparse CNNs Accelerator on FPGA

Communication-Aware and Resource-Efficient NoC-Based Architecture for CNN Acceleration

A High-Efficient and Configurable Hardware Accelerator for Convolutional Neural Network

A High Utilization FPGA-Based Accelerator for Variable-Scale Convolutional Neural Network

An FPGA-Based Accelerator Enabling Efficient Support for CNNs with Arbitrary Kernel Sizes

A Block-Floating-Point Arithmetic Based FPGA Accelerator for Convolutional Neural Networks

An Efficient Hardware Accelerator for Structured Sparse Convolutional Neural Networks on FPGAs

A FPGA-based Accelerator of Convolutional Neural Network for Face Feature Extraction

Toward Full-Stack Acceleration of Deep Convolutional Neural Networks on FPGAs

A Power-Efficient and High Performance FPGA Accelerator for Convolutional Neural Networks: Work-in-progress.

High-Performance FPGA-Based CNN Accelerator with Block-Floating-Point Arithmetic.

A Flexible and Efficient FPGA Accelerator for Various Large-Scale and Lightweight CNNs

An Efficient Hardware Accelerator for Sparse Convolutional Neural Networks on FPGAs

A Scalable FPGA Accelerator for Convolutional Neural Networks.

WPU: A FPGA-based Scalable, Efficient and Software/Hardware Co-design Deep Neural Network Inference Acceleration Processor