Abstract:In the modern-day era of technology, a paradigm shift has been witnessed in the areas involving applications of Artificial Intelligence (AI), Machine Learning (ML), and Deep Learning (DL). Specifically, Deep Neural Networks (DNNs) have emerged as a popular field of interest in most AI applications such as computer vision, image and video processing, robotics, etc. In the context of developed digital technologies and the availability of authentic data and data handling infrastructure, DNNs have been a credible choice for solving more complex real-life problems. The performance and accuracy of a DNN is a way better than human intelligence in certain situations. However, it is noteworthy that the DNN is computationally too cumbersome in terms of the resources and time to handle these computations. Furthermore, general-purpose architectures like CPUs have issues in handling such computationally intensive algorithms. Therefore, a lot of interest and efforts have been invested by the research fraternity in specialized hardware architectures such as Graphics Processing Unit (GPU), Field Programmable Gate Array (FPGA), Application Specific Integrated Circuit (ASIC), and Coarse Grained Reconfigurable Array (CGRA) in the context of effective implementation of computationally intensive algorithms. This paper brings forward the various research works on the development and deployment of DNNs using the aforementioned specialized hardware architectures and embedded AI accelerators. The review discusses the detailed description of the specialized hardware-based accelerators used in the training and/or inference of DNN. A comparative study based on factors like power, area, and throughput, is also made on the various accelerators discussed. Finally, future research and development directions, such as future trends in DNN implementation on specialized hardware accelerators, are discussed. This review article is intended to guide hardware architects to accelerate and improve the eff- ctiveness of deep learning research.

New paradigm of FPGA-based computational intelligence from surveying the implementation of DNN accelerators

A Near Memory Computing FPGA Architecture for Neural Network Acceleration

DaDianNao: A Machine-Learning Supercomputer

[DL] A Survey of FPGA-based Neural Network Inference Accelerators

Recent Advances in Efficient Computation of Deep Convolutional Neural Networks

Designing Deep Learning Hardware Accelerator and Efficiency Evaluation

A Survey of FPGA-Based Neural Network Accelerator

A survey of FPGA design for AI era

A Survey of FPGA Based Deep Learning Accelerators: Challenges and Opportunities

Efficient Hardware Architectures for Accelerating Deep Neural Networks: Survey

An All-Digital Compute-In-Memory FPGA Architecture for Deep Learning Acceleration

Research on Convolutional Neural Network Inference Acceleration and Performance Optimization for Edge Intelligence

Challenges in Energy-Efficient Deep Neural Network Training with FPGA.

Power-Driven DNN Dataflow Optimization on FPGA

FPGA/DNN Co-Design: An Efficient Design Methodology for IoT Intelligence on the Edge

Latency optimized Deep Neural Networks (DNNs): An Artificial Intelligence approach at the Edge using Multiprocessor System on Chip (MPSoC)

An Overview of FPGA Based Deep Learning Accelerators: Challenges and Opportunities.

Field-Programmable Gate Array Architecture for Deep Learning: Survey & Future Directions

A Survey on Hardware Accelerator Design of Deep Learning for Edge Devices

Towards Ultra-High Performance and Energy Efficiency of Deep Learning Systems: An Algorithm-Hardware Co-Optimization Framework

Mobile or FPGA? A Comprehensive Evaluation on Energy Efficiency and a Unified Optimization Framework