Abstract:With the rapid development of deep learning technology, the demand for computing resources is increasing, and the accelerated optimization of hardware on artificial intelligence (AI) chip has become one of the key ways to solve this challenge. This paper aims to explore the hardware acceleration optimization strategy of deep learning model on AI chip to improve the training and inference performance of the model. In this paper, the method and practice of optimizing deep learning model on AI chip are deeply analyzed by comprehensively considering the hardware characteristics such as parallel processing ability, energy-efficient computing, neural network accelerator, flexibility and programmability, high integration and heterogeneous computing structure. By designing and implementing an efficient convolution accelerator, the computational efficiency of the model is improved. The introduction of energy-efficient computing effectively reduces energy consumption, which provides feasibility for the practical application of mobile devices and embedded systems. At the same time, the optimization design of neural network accelerator becomes the core of hardware acceleration, and deep learning calculation such as convolution and matrix operation are accelerated through special hardware structure, which provides strong support for the real-time performance of the model. By analyzing the actual application cases of hardware accelerated optimization in different application scenarios, this paper highlights the key role of hardware accelerated optimization in improving the performance of deep learning model. Hardware accelerated optimization not only improves the computing efficiency, but also provides efficient and intelligent computing support for AI applications in different fields.

Hardware-Aware Softmax Approximation for Deep Neural Networks

Efficient Hardware Architecture of Softmax Layer in Deep Neural Network

Deep Neural Network Approximation for Custom Hardware

Hardware-Efficient SoftMax Architecture With Bit-Wise Exponentiation and Reciprocal Calculation

Towards Ultra-High Performance and Energy Efficiency of Deep Learning Systems: An Algorithm-Hardware Co-Optimization Framework

Hardware Platform-Aware Binarized Neural Network Model Optimization

HAO: Hardware-aware neural Architecture Optimization for Efficient Inference

Woodpecker-DL: Accelerating Deep Neural Networks via Hardware-Aware Multifaceted Optimizations

Towards Efficient Neural Networks On-a-chip: Joint Hardware-Algorithm Approaches

Learned Hardware/Software Co-Design of Neural Accelerators

Learning on Hardware: A Tutorial on Neural Network Accelerators and Co-Processors

Hardware Approximate Techniques for Deep Neural Network Accelerators: A Survey

A Survey on Efficient Convolutional Neural Networks and Hardware Acceleration

Optimization of Softmax Layer in Deep Neural Network Using Integral Stochastic Computation

Hardware Accelerated Optimization of Deep Learning Model on Artificial Intelligence Chip

EH-DNAS: End-to-End Hardware-aware Differentiable Neural Architecture Search

Hardware-aware training for large-scale and diverse deep learning inference workloads using in-memory computing-based accelerators

A DNN Optimization Framework with Unlabeled Data for Efficient and Accurate Reconfigurable Hardware Inference

Hardware and Software Optimizations for Accelerating Deep Neural Networks: Survey of Current Trends, Challenges, and the Road Ahead

Towards Budget-Driven Hardware Optimization for Deep Convolutional Neural Networks Using Stochastic Computing

Software-defined Design Space Exploration for an Efficient DNN Accelerator Architecture