Abstract:In the modern-day era of technology, a paradigm shift has been witnessed in the areas involving applications of Artificial Intelligence (AI), Machine Learning (ML), and Deep Learning (DL). Specifically, Deep Neural Networks (DNNs) have emerged as a popular field of interest in most AI applications such as computer vision, image and video processing, robotics, etc. In the context of developed digital technologies and the availability of authentic data and data handling infrastructure, DNNs have been a credible choice for solving more complex real-life problems. The performance and accuracy of a DNN is a way better than human intelligence in certain situations. However, it is noteworthy that the DNN is computationally too cumbersome in terms of the resources and time to handle these computations. Furthermore, general-purpose architectures like CPUs have issues in handling such computationally intensive algorithms. Therefore, a lot of interest and efforts have been invested by the research fraternity in specialized hardware architectures such as Graphics Processing Unit (GPU), Field Programmable Gate Array (FPGA), Application Specific Integrated Circuit (ASIC), and Coarse Grained Reconfigurable Array (CGRA) in the context of effective implementation of computationally intensive algorithms. This paper brings forward the various research works on the development and deployment of DNNs using the aforementioned specialized hardware architectures and embedded AI accelerators. The review discusses the detailed description of the specialized hardware-based accelerators used in the training and/or inference of DNN. A comparative study based on factors like power, area, and throughput, is also made on the various accelerators discussed. Finally, future research and development directions, such as future trends in DNN implementation on specialized hardware accelerators, are discussed. This review article is intended to guide hardware architects to accelerate and improve the eff- ctiveness of deep learning research.

Implementation and evaluation of deep neural networks (DNN) on mainstream heterogeneous systems

DaDianNao: A Machine-Learning Supercomputer

A Practical Implementation of GPU based Accelerator for Deep Neural Networks

A Power Efficient Neural Network Implementation on Heterogeneous FPGA and GPU Devices

Performance Modeling and Evaluation of Distributed Deep Learning Frameworks on GPUs

Why is FPGA-GPU Heterogeneity the Best Option for Embedded Deep Neural Networks?

A Metaprogramming and Autotuning Framework for Deploying Deep Learning Applications

Acceleration of Deep Neural Network Training with Resistive Cross-Point Devices

A Performance Analysis Framework for Exploiting GPU Microarchitectural Capability.

A Hybrid Parallelization Approach for Distributed and Scalable Deep Learning

EmBench: Quantifying Performance Variations of Deep Neural Networks across Modern Commodity Devices

HPH: Hybrid Parallelism on Heterogeneous Clusters for Accelerating Large-scale DNNs Training.

Efficient and Robust Parallel DNN Training through Model Parallelism on Multi-GPU Platform

Accelerating DNN Inference with Heterogeneous Multi-DPU Engines

Evaluating Emerging AI/ML Accelerators: IPU, RDU, and NVIDIA/AMD GPUs

Benchmarking TPU, GPU, and CPU Platforms for Deep Learning

HP-GNN: Generating High Throughput GNN Training Implementation on CPU-FPGA Heterogeneous Platform

Artificial Neural Network Computation On Graphic Process Unit

A Unified Architecture for Accelerating Distributed DNN Training in Heterogeneous GPU/CPU Clusters

A Heterogeneous Full-stack AI Platform for Performance Monitoring and Hardware-specific Optimizations

Efficient Hardware Architectures for Accelerating Deep Neural Networks: Survey