Abstract:Sparse triangular solve（SpTRSV） is a vital operation in preconditioners. In particular, in scientific computing program that solves partial differential equation systems iteratively, structured SpTRSV is a common type of issue and often a performance bottleneck that needs to be addressed by the scientific computing program. The commercial mathematical libraries tailored to the graphics processing unit（GPU） platform, represented by CUSPARSE, parallelize SpTRSV operations by level-scheduling methods. However, this method is weakened by time-consuming preprocessing and serious GPU thread idle when it is employed to deal with structured SpTRSV issues. This study proposes a parallel algorithm tailored to structured SpTRSV issues. The proposed algorithm leverages the special non-zero element distribution pattern of structured SpTRSV issues during task allocation to skip the preprocessing and analysis of the non-zero element structure of the input issue. Furthermore, the element-wise operation strategy used in the existing level-scheduling methods is modified. As a result, the problem of GPU thread idle is effectively alleviated, and the memory access latency of some non-zero elements in the matrix is concealed. This study also adopts a state variable compression technique according to the task allocation characteristics of the proposed algorithm, significantly improving the cache hit rate of the algorithm in state variable operations. Additionally, several hardware features of the GPU, including predicated execution, are investigated to comprehensively optimize algorithm implementation. The proposed algorithm is tested on NVIDIA V100 GPU, achieving an average 2.71× acceleration over CUSPARSE and a peak effective memory-access bandwidth of 225.2 GB/s. The modified element-wise operation strategy, combined with a series of other optimization measures for GPU hardware, attains a prominent optimization effect by yielding a nearly115% increase in the effective memory-access bandwidth of the proposed algorithm.

Atomic Reduction Based Sparse Matrix-Transpose Vector Multiplication on GPUs

Optimizing sparse matrix-vector multiplication based on gpu

Parallel optimization for sparse matrix-vector on GPU

Efficient Algorithm Design of Optimizing SpMV on GPU.

Sgap: Towards Efficient Sparse Tensor Algebra Compilation for GPU

Efficient sparse-matrix multi-vector product on GPUs

TileSpMSpV: A Tiled Algorithm for Sparse Matrix-Sparse Vector Multiplication on GPUs

FastLoad: Speeding Up Data Loading of Both Sparse Matrix and Vector for SpMV on GPUs

TaiChi: A Hybrid Compression Format for Binary Sparse Matrix-Vector Multiplication on GPU

TileSpMV: A Tiled Algorithm for Sparse Matrix-Vector Multiplication on GPUs

Improvement of Sparse Matrix-Vector Multiplication on GPU

A Novel Parallel Algorithm for Sparse Tensor Matrix Chain Multiplication via TCU-Acceleration

Parallel Structured Sparse Triangular Solver for GPU Platform

Performance Optimization for Sparse A(T)Ax in Parallel on Multicore Cpu

An efficient sparse stiffness matrix vector multiplication using compressed sparse row storage format on AMD GPU

An Integral-equation-oriented Vectorized SpMV Algorithm and Its Application on CT Imaging Reconstruction

A New Sparse Matrix Vector Multiplication GPU Algorithm Designed for Finite Element Problems

Block-wise dynamic mixed-precision for sparse matrix-vector multiplication on GPUs

Accelerating Sparse Approximate Matrix Multiplication on GPUs

Smat: An Input Adaptive Auto-Tuner For Sparse Matrix-Vector Multiplication

Sparse matrix partitioning for optimizing SpMV on CPU-GPU heterogeneous platforms