Abstract:With the rapid development of deep learning technology, the demand for computing resources is increasing, and the accelerated optimization of hardware on artificial intelligence (AI) chip has become one of the key ways to solve this challenge. This paper aims to explore the hardware acceleration optimization strategy of deep learning model on AI chip to improve the training and inference performance of the model. In this paper, the method and practice of optimizing deep learning model on AI chip are deeply analyzed by comprehensively considering the hardware characteristics such as parallel processing ability, energy-efficient computing, neural network accelerator, flexibility and programmability, high integration and heterogeneous computing structure. By designing and implementing an efficient convolution accelerator, the computational efficiency of the model is improved. The introduction of energy-efficient computing effectively reduces energy consumption, which provides feasibility for the practical application of mobile devices and embedded systems. At the same time, the optimization design of neural network accelerator becomes the core of hardware acceleration, and deep learning calculation such as convolution and matrix operation are accelerated through special hardware structure, which provides strong support for the real-time performance of the model. By analyzing the actual application cases of hardware accelerated optimization in different application scenarios, this paper highlights the key role of hardware accelerated optimization in improving the performance of deep learning model. Hardware accelerated optimization not only improves the computing efficiency, but also provides efficient and intelligent computing support for AI applications in different fields.

Deep Learning Compiler Load Balancing Optimization Method for Model Training

ALT: Breaking the Wall Between Data Layout and Loop Optimizations for Deep Learning Compilation

ALT: Boosting Deep Learning Performance by Breaking the Wall between Graph and Operator Level Optimizations

RAF: Holistic Compilation for Deep Learning Model Training

An Optimization Toolchain Design Of Deep Learning Deployment Based On Heterogeneous Computing Platform

AI Powered Compiler Techniques for DL Code Optimization

oneDNN Graph Compiler: A Hybrid Approach for High-Performance Deep Learning Compilation

Pipeline-based Optimization Method for Large-Scale End-to-End Inference.

Woodpecker-DL: Accelerating Deep Neural Networks via Hardware-Aware Multifaceted Optimizations

The Deep Learning Compiler: A Comprehensive Survey

Performance Modeling and Evaluation of Distributed Deep Learning Frameworks on GPUs

Compiler-Level Matrix Multiplication Optimization for Deep Learning

Bring Your Own Codegen to Deep Learning Compiler

Scheduling Computation Graphs of Deep Learning Models on Manycore CPUs

ATFormer: A Learned Performance Model with Transfer Learning Across Devices for Deep Learning Tensor Programs

Performance and Power Evaluation of AI Accelerators for Training Deep Learning Models

Towards Ultra-High Performance and Energy Efficiency of Deep Learning Systems: An Algorithm-Hardware Co-Optimization Framework

Hardware Accelerated Optimization of Deep Learning Model on Artificial Intelligence Chip

Benchmarking State-of-the-Art Deep Learning Software Tools

Deep Optimizer States: Towards Scalable Training of Transformer Models Using Interleaved Offloading

An In-depth Comparison of Compilers for Deep Neural Networks on Hardware