Abstract:Emerging GPU applications exhibit increasingly high computation demands which has led GPU manufacturers to build GPUs with an increasingly large number of streaming multiprocessors (SMs). Providing data to the SMs at high bandwidth puts significant pressure on the memory hierarchy and the Network-on-Chip (NoC). Current GPUs typically partition the memory-side last-level cache (LLC) in equally-sized slices that are shared by all SMs. Although a shared LLC typically results in a lower miss rate, we find that for workloads with high degrees of data sharing across SMs, a private LLC leads to a significant performance advantage because of increased bandwidth to replicated cache lines across different LLC slices. In this paper, we propose adaptive memory-side last-level GPU caching to boost performance for sharing-intensive workloads that need high bandwidth to read-only shared data. Adaptive caching leverages a lightweight performance model that balances increased LLC bandwidth against increased miss rate under private caching. In addition to improving performance for sharing-intensive workloads, adaptive caching also saves energy in a (co-designed) hierarchical two-stage crossbar NoC by power-gating and bypassing the second stage if the LLC is configured as a private cache. Our experimental results using 17 GPU workloads show that adaptive caching improves performance by 28.1% on average (up to 38.1%) compared to a shared LLC for sharing-intensive workloads. In addition, adaptive caching reduces NoC energy by 26.6% on average (up to 29.7%) and total system energy by 6.1% on average (up to 27.2%) when configured as a private cache. Finally, we demonstrate through a GPU NoC design space exploration that a hierarchical two-stage crossbar is both more power- and area-efficient than full and concentrated crossbars with the same bisection bandwidth, thus providing a low-cost cooperative solution to exploit workload sharing behavior in memory-side last-level caches.

A Graph-based Model for GPU Caching Problems

Using GPU to Accelerate Cache Simulation.

Adaptive Memory-Side Last-level GPU Caching

Coordinated Static and Dynamic Cache Bypassing for GPUs

Adaptive Cache Management for Energy-Efficient GPU Computing.

Locality Protected Dynamic Cache Allocation Scheme on GPUs

Optimizing Cache Bypassing and Warp Scheduling for GPUs

Exploring Cache Bypassing and Partitioning for Multi-Tasking on GPUs

Optimizing Graph Processing on GPUs

Orchestrating Cache Management and Memory Scheduling for GPGPU Applications.

Accelerate Cache Simulation with Generic GPU

Locality-Driven Dynamic Gpu Cache Bypassing

Performance Evaluation and Optimization on Gpu

Intra-Cluster Coalescing and Distributed-Block Scheduling to Reduce GPU NoC Pressure.

Locality-protected Cache Allocation Scheme with Low Overhead on GPUs.

An Efficient Compiler Framework for Cache Bypassing on GPUs

A Sharing-Aware L1.5D Cache for Data Reuse in GPGPUs

Cache simulator based on GPU acceleration

Understanding the Tradeoffs Between Software-Managed Vs. Hardware-Managed Caches in GPUs

A BSP Model Graph Processing System on Many Cores

CPU-assisted GPU thread pool model for dynamic task parallelism