Abstract:The performance of sparse iterative solvers is typically limited by sparse matrix-vector multiplication, which is itself limited by memory system and network performance. As the gap between computation and communication speed continues to widen, these traditional sparse methods will suffer. In this paper we focus on an alternative building block for sparse iterative solvers, the “matrix powers ker-nel” [x, Ax, $\rm A^{2}x$, …, $\rm A^{k}x$], and show that by organizing computations around this kernel, we can achieve near-minimal communication costs. We consider communication very broadly as both network communication in parallel code and memory hierarchy access in sequential code. In particular, we introduce a parallel algorithm for which the number of messages (total latency cost) is independent of the power k, and a sequential algorithm, that reduces both the number and volume of accesses, so that it is independent of k in both latency and bandwidth costs. This is part of a larger project to develop “communication-avoiding Krylov subspace methods,” which also addresses the numerical issues associated with these methods. Our algorithms work for general sparse matrices that “partition well”. We introduce parallel performance models of matrices arising from 2D and 3D problems and show predicted speedups over a conventional algorithm of up to $7x$ on a Petaflop-scale machine and up to $22x$ on computation across the Grid. Analogous sequential performance models of the same problems predict speedups over a conventional algorithm of up to $10x$ on an out-of-core implemen-tation, and up to $2.5x$ when we use our ideas to reduce off-chip latency and bandwidth to DRAM. Finally, we validate the model on an out-of-core sequential implementation and measured a speedup of over $3x$, which is close to the predicted speedup.

The Improved Krylov Subspace Methods for Large and Sparse Linear Systems on Bulk Synchronous Parallel Architectures

A Parallel Simulator for Massive Reservoir Models Utilizing Distributed-Memory Parallel Systems

The IBiCGStab method on bulk synchronous parallel architectures

The Improved BiCG Method for Large and Sparse Linear Systems on Parallel Distributed Memory Architectures.

Parallel Row Operation Algorithm for Banded Linear Systems

An Efficient Parallel Krylov-Schur Method for Eigen-Analysis of Large-Scale Power Systems

Scalable Linear Solvers for Sparse Linear Systems from Large-Scale Numerical Simulations.

A Survey of Accelerating Parallel Sparse Linear Algebra

Parallel Algorithm for Solving Banded Linear Systems

Implementation of a Parallel Sparse Direct Solver on Vector Architecture

Optimizing Algorithm of Sparse Linear Systems on GPU

On the stability and performance of the solution of sparse linear systems by partitioned procedures

A Method for Solving Large Sparse Unsymmetric Linear Equations

Avoiding communication in sparse matrix computations

The Improved Parallel Icgs Method For Large And Sparse Unsymmetric Linear Systems

A Class of Parallel Algorithms for Solving Large Sparse Linear Systems on Multiprocessors

Scalable Hierarchical Parallel Algorithm for the Solution of Super Large-Scale Sparse Linear Equations

A Non-Krylov subspace Method for Solving Large and Sparse Linear System of Equations

Development of Krylov and AMG linear solvers for large-scale sparse matrices on GPUs

Iteratively Solving Sparse Linear System Based on PaRSEC Task Scheduling

Virtual Transmission Method, A New Distributed Algorithm to Solve Sparse Linear Systems