Abstract:Sparse matrix-sparse vector (SpMSpV) multiplication is one of the fundamental and important operations in many high-performance scientific and engineering applications. The inherent irregularity and poor data locality lead to two main challenges to scaling SpMSpV over high-performance computing (HPC) systems: (i) a large amount of redundant data limits the utilization of bandwidth and parallel resources; (ii) the irregular access pattern limits the exploitation of computing resources. This paper proposes a fine-grained parallel SpMSpV ( fgSpMSpV ) framework on Sunway TaihuLight supercomputer to alleviate the challenges for large-scale real-world applications. First, fgSpMSpV adopts an MPI \( + \) OpenMP \( +X \) parallelization model to exploit the multi-stage and hybrid parallelism of heterogeneous HPC architectures and accelerate both pre-/post-processing and main SpMSpV computation. Second, fgSpMSpV utilizes an adaptive parallel execution to reduce the pre-processing, adapt to the parallelism and memory hierarchy of the Sunway system, while still tame redundant and random memory accesses in SpMSpV, including a set of techniques like the fine-grained partitioner, re-collection method, and Compressed Sparse Column Vector (CSCV) matrix format. Third, fgSpMSpV uses several optimization techniques to further utilize the computing resources. fgSpMSpV on the Sunway TaihuLight gains a noticeable performance improvement from the key optimization techniques with various sparsity of the input. Additionally, fgSpMSpV is implemented on an NVIDIA Tesal P100 GPU and applied to the breath-first-search (BFS) application. fgSpMSpV on a P100 GPU obtains the speedup of up to \( 134.38\times \) over the state-of-the-art SpMSpV algorithms, and the BFS application using fgSpMSpV achieves the speedup of up to \( 21.68\times \) over the state-of-the-arts.

fgSpMSpV: A Fine-grained Parallel SpMSpV Framework on HPC Platforms

Towards Efficient SpMV on Sunway Manycore Architectures.

TileSpMSpV: A Tiled Algorithm for Sparse Matrix-Sparse Vector Multiplication on GPUs

FSpGEMM: An OpenCL-based HPC Framework for Accelerating General Sparse Matrix-Matrix Multiplication on FPGAs

FPGA and GPU Implementation of Large Scale SpMV

Efficient Algorithm Design of Optimizing SpMV on GPU.

A Work-Efficient Parallel Sparse Matrix-Sparse Vector Multiplication Algorithm

Characterizing Scalability of Sparse Matrix-Vector Multiplications on Phytium FT-2000+ Many-cores

A New Sparse Matrix Vector Multiplication GPU Algorithm Designed for Finite Element Problems

Esspmv: an Embedded-FPGA-based Hardware Accelerator for Symmetric Sparse Matrix-Vector Multiplication.

An Optimized GP-GPU Warp Scheduling Algorithm for Sparse Matrix-Vector Multiplication

Design and Implementation of Adaptive SpMV Library for Multicore and Many-Core Architecture

Efficient Sparse Matrix Kernels based on Adaptive Workload-Balancing and Parallel-Reduction

A Fast Sparse Triangular Solver for Structured-grid Problems on Sunway Many-core Processor SW26010

Adaptive SpMV/SpMSpV on GPUs for Input Vectors of Varied Sparsity

Parallel Efficient Sparse Matrix-Matrix Multiplication on Multicore Platforms

Parallel optimization for sparse matrix-vector on GPU

A Data Locality-Aware Design Framework For Reconfigurable Sparse Matrix-Vector Multiplication Kernel

Sparse matrix partitioning for optimizing SpMV on CPU-GPU heterogeneous platforms

Distributed-Memory Parallel Algorithms for Sparse Matrix and Sparse Tall-and-Skinny Matrix Multiplication

FSpGEMM: A Framework for Accelerating Sparse General Matrix–Matrix Multiplication Using Gustavson’s Algorithm on FPGAs