Abstract:Sorting is a primitive operation that is a building block for countless algorithms. As such, it is important to design sorting algorithms that approach peak performance on a range of hardware architectures. Graphics Processing Units (GPUs) are particularly attractive architectures as they provides massive parallelism and computing power. However, the intricacies of their compute and memory hierarchies make designing GPU-efficient algorithms challenging. In this work we present GPU Multiway Mergesort (MMS), a new GPU-efficient multiway mergesort algorithm. MMS employs a new partitioning technique that exposes the parallelism needed by modern GPU architectures. To the best of our knowledge, MMS is the first sorting algorithm for the GPU that is asymptotically optimal in terms of global memory accesses and that is completely free of shared memory bank conflicts. We realize an initial implementation of MMS, evaluate its performance on three modern GPU architectures, and compare it to competitive implementations available in state-of-the-art GPU libraries. Despite these implementations being highly optimized, MMS compares favorably, achieving performance improvements for most random inputs. Furthermore, unlike MMS, state-of-the-art algorithms are susceptible to bank conflicts. We find that for certain inputs that cause these algorithms to incur large numbers of bank conflicts, MMS can achieve up to a 37.6% speedup over its fastest competitor. Overall, even though its current implementation is not fully optimized, due to its efficient use of the memory hierarchy, MMS outperforms the fastest comparison-based sorting implementations available to date.

Parallel Shellsort Algorithm for Many-Core GPUs with CUDA

Parallel Shellsort Algorithm for Many-Core GPUs with CUDA

Count Sort for GPU Computing

Parallel Sorting by Approximate Splitting for Multi-core Processors

A Hybrid Sorting Algorithm on Heterogeneous Architectures

The implementation and optimization of Bitonic sort algorithm based on CUDA

A Real-Time Spike Sorting Method Based on the Embedded GPU

A Memory Bandwidth-Efficient Hybrid Radix Sort on GPUs

An Efficient Multiway Mergesort for GPU Architectures

A study of integer sorting on multicores

GPU accelerate parallel Odd-Even merge sort: An OpenCL method

A Creativity Survey of Parallel Sorting Algorithm

Simple sorting algorithm test based on CUDA

Efficient parallel merge sort for fixed and variable length keys

A Parallel Implementation of Nearest Neighbor Analysis Based on GPGPU

A Parallel Quicksort Algorithm on Manycore Processors in Sunway TaihuLight

Impact of CUDA and OpenCL on Parallel and Distributed Computing

Engineering Worst-Case Inputs for Pairwise Merge Sort on GPUs

On Parallel Solution of Sparse Triangular Linear Systems in CUDA

CUDA Optimization Strategies for Compute- and Memory-Bound Neuroimaging Algorithms

An Algorithm to Parallelise Parton Showers on a GPU