Abstract:In the modern-day era of technology, a paradigm shift has been witnessed in the areas involving applications of Artificial Intelligence (AI), Machine Learning (ML), and Deep Learning (DL). Specifically, Deep Neural Networks (DNNs) have emerged as a popular field of interest in most AI applications such as computer vision, image and video processing, robotics, etc. In the context of developed digital technologies and the availability of authentic data and data handling infrastructure, DNNs have been a credible choice for solving more complex real-life problems. The performance and accuracy of a DNN is a way better than human intelligence in certain situations. However, it is noteworthy that the DNN is computationally too cumbersome in terms of the resources and time to handle these computations. Furthermore, general-purpose architectures like CPUs have issues in handling such computationally intensive algorithms. Therefore, a lot of interest and efforts have been invested by the research fraternity in specialized hardware architectures such as Graphics Processing Unit (GPU), Field Programmable Gate Array (FPGA), Application Specific Integrated Circuit (ASIC), and Coarse Grained Reconfigurable Array (CGRA) in the context of effective implementation of computationally intensive algorithms. This paper brings forward the various research works on the development and deployment of DNNs using the aforementioned specialized hardware architectures and embedded AI accelerators. The review discusses the detailed description of the specialized hardware-based accelerators used in the training and/or inference of DNN. A comparative study based on factors like power, area, and throughput, is also made on the various accelerators discussed. Finally, future research and development directions, such as future trends in DNN implementation on specialized hardware accelerators, are discussed. This review article is intended to guide hardware architects to accelerate and improve the eff- ctiveness of deep learning research.

Emerging Neural Workloads and Their Impact on Hardware.

Towards Efficient Neural Networks On-a-chip: Joint Hardware-Algorithm Approaches

A Co-design view of Compute in-Memory with Non-Volatile Elements for Neural Networks

Efficient Hardware Architectures for Accelerating Deep Neural Networks: Survey

Roadmap on Emerging Hardware and Technology for Machine Learning

A Survey on Efficient Convolutional Neural Networks and Hardware Acceleration

Special Topic on Nonvolatile Memory for Efficient Implementation of Neural/Neuromorphic Computing

Analog architectures for neural network acceleration based on non-volatile memory

A Survey of Neural Network Hardware Accelerators in Machine Learning

Hardware-Aware Neural Architecture Search: Survey and Taxonomy

Hardware and Software Optimizations for Accelerating Deep Neural Networks: Survey of Current Trends, Challenges, and the Road Ahead

Evaluating Emerging AI/ML Accelerators: IPU, RDU, and NVIDIA/AMD GPUs

Towards Efficient Neuro-Symbolic AI: From Workload Characterization to Hardware Architecture

Learning on Hardware: A Tutorial on Neural Network Accelerators and Co-Processors

Hardware-friendly Neural Network Architecture for Neuromorphic Computing

A Survey on Neural Network Hardware Accelerators

Hardware-aware training for large-scale and diverse deep learning inference workloads using in-memory computing-based accelerators

A Heterogeneous Full-stack AI Platform for Performance Monitoring and Hardware-specific Optimizations

Energy Scaling Advantages of Resistive Memory Crossbar Based Computation and Its Application to Sparse Coding

A Survey of Near-Data Processing Architectures for Neural Networks

An Automated Design Flow for Adaptive Neural Network Hardware Accelerators