Abstract:Sparse matrix-vector multiplication (SpMV) operations are commonly used in various scientific applications. The performance of the SpMV operation often depends on exploiting regularity patterns in the matrix. Various representations have been proposed to minimize the memory bandwidth bottleneck arising from the irregular memory access pattern involved. Among recent representation techniques, tensor decomposition is a popular one used for very large but sparse matrices. Post sparse-tensor decomposition, the new representation involves indirect accesses, making it challenging to optimize for multi-cores and GPUs. Computational neuroscience algorithms often involve sparse datasets while still performing long-running computations on them. The LiFE application is a popular neuroscience algorithm used for pruning brain connectivity graphs. The datasets employed herein involve the Sparse Tucker Decomposition (STD), a widely used tensor decomposition method. Using this decomposition leads to irregular array references, making it very difficult to optimize for both CPUs and GPUs. Recent codes of the LiFE algorithm show that its SpMV operations are the key bottleneck for performance and scaling. In this work, we first propose target-independent optimizations to optimize these SpMV operations, followed by target-dependent optimizations for CPU and GPU systems. The target-independent techniques include: (1) standard compiler optimizations, (2) data restructuring methods, and (3) methods to partition computations among threads. Then we present the optimizations for CPUs and GPUs to exploit platform-specific speed. Our highly optimized CPU code obtain a speedup of 27.12x over the original sequential CPU code running on 16-core Intel Xeon (Skylake-based) system, and our optimized GPU code achieves a speedup of 5.2x over a reference optimized GPU code version on NVIDIA's GeForce RTX 2080 Ti GPU.

Mapping Parallelism in a Functional IR through Constraint Satisfaction

Strategy Preserving Compilation for Parallel Functional Code

Generating performance portable code using rewrite rules: from high-level functional expressions to high-performance OpenCL code

Automatic source code generation for deterministic global optimization with parallel architectures

MLIR-based code generation for GPU tensor cores

Accelerating RTL Simulation with GPUs

Multilevel Granularity Parallelism Synthesis on FPGAs

Exploiting Parallelism in the Simulation of General Purpose Graphics Processing Unit Program

High-Performance GPU-to-CPU Transpilation and Optimization via High-Level Parallel Constructs

Optimizing Spatial Mapping of Nested Loop for Coarse-Grained Reconfigurable Architectures

Optimizing the Linear Fascicle Evaluation Algorithm for Multi-Core and Many-Core Systems

Exploiting Outer Loop Parallelism of Nested Loop on Coarse-Grained Reconfigurable Architectures

Reconfigurable Hardware Generation of Multigrid Solvers with Conjugate Gradient Coarse-Grid Solution

A Compiler-assisted Locality Aware CTA Mapping Scheme

An Efficient Compiler Framework for Cache Bypassing on GPUs

GPT-Driven Source-to-Source Transformation for Generating Compilable Parallel CUDA Code for Nussinov's Algorithm

Parallelism Optimization with Variables Affinity Graph for Reconfigurable Computing Architecture

A Compilation Flow for the Generation of CNN Inference Accelerators on FPGAs

Optimizing Declarative Parallel Distributed Graph Processing by Using Constraint Solvers.

Breaking the Memory Wall: A Study of I/O Patterns and GPU Memory Utilization for Hybrid CPU-GPU Offloaded Optimizers

Acceleration for Timing-Aware Gate-Level Logic Simulation with One-Pass GPU Parallelism