Abstract:Emerging GPU applications exhibit increasingly high computation demands which has led GPU manufacturers to build GPUs with an increasingly large number of streaming multiprocessors (SMs). Providing data to the SMs at high bandwidth puts significant pressure on the memory hierarchy and the Network-on-Chip (NoC). Current GPUs typically partition the memory-side last-level cache (LLC) in equally-sized slices that are shared by all SMs. Although a shared LLC typically results in a lower miss rate, we find that for workloads with high degrees of data sharing across SMs, a private LLC leads to a significant performance advantage because of increased bandwidth to replicated cache lines across different LLC slices. In this paper, we propose adaptive memory-side last-level GPU caching to boost performance for sharing-intensive workloads that need high bandwidth to read-only shared data. Adaptive caching leverages a lightweight performance model that balances increased LLC bandwidth against increased miss rate under private caching. In addition to improving performance for sharing-intensive workloads, adaptive caching also saves energy in a (co-designed) hierarchical two-stage crossbar NoC by power-gating and bypassing the second stage if the LLC is configured as a private cache. Our experimental results using 17 GPU workloads show that adaptive caching improves performance by 28.1% on average (up to 38.1%) compared to a shared LLC for sharing-intensive workloads. In addition, adaptive caching reduces NoC energy by 26.6% on average (up to 29.7%) and total system energy by 6.1% on average (up to 27.2%) when configured as a private cache. Finally, we demonstrate through a GPU NoC design space exploration that a hierarchical two-stage crossbar is both more power- and area-efficient than full and concentrated crossbars with the same bisection bandwidth, thus providing a low-cost cooperative solution to exploit workload sharing behavior in memory-side last-level caches.

RACB: Resource Aware Cache Bypass on GPUs

Locality-Driven Dynamic Gpu Cache Bypassing

Coordinated Static and Dynamic Cache Bypassing for GPUs

Selectively GPU Cache Bypassing for Un-Coalesced Loads.

Statistical Cache Bypassing for Non-Volatile Memory

A model-driven approach to warp/thread-block level GPU cache bypassing.

Sbac: A Statistics Based Cache Bypassing Method For Asymmetric-Access Caches

An Efficient Compiler Framework for Cache Bypassing on GPUs

Locality-protected Cache Allocation Scheme with Low Overhead on GPUs.

Adaptive Cache Bypass and Insertion for Many-core Accelerators

DD-L1D: Improving the Decoupled L1D Efficiency for GPU Architecture

Adaptive Cache Management for Energy-Efficient GPU Computing.

Exploring Cache Bypassing and Partitioning for Multi-Tasking on GPUs

Locality Protected Dynamic Cache Allocation Scheme on GPUs

Adaptive Memory-Side Last-level GPU Caching

Optimizing Cache Bypassing and Warp Scheduling for GPUs

Understanding the Tradeoffs Between Software-Managed Vs. Hardware-Managed Caches in GPUs

Improving CPU and GPU Performance Through Sample-Based Dynamic LLC Bypassing

Bandwidth-Effective DRAM Cache for GPUs with Storage-Class Memory

A Sharing-Aware L1.5D Cache for Data Reuse in GPGPUs

BackCache: Mitigating Contention-Based Cache Timing Attacks by Hiding Cache Line Evictions