Abstract:Graphics Processing Units (GPUs) computing has become ubiquitous for embedded system, evidenced by its wide adoption for various general purpose applications. As more and more applications are accelerated by GPUs, multi-tasking scenario starts to emerge. Multi-tasking allows multiple applications to simultaneously execute on the same GPU and share the resource. This brings new challenges due to the contention among the different applications for the shared resources such as caches. However, the caches on GPUs are difficult to use. If used inappropriately, it may hurt the performance instead of improving it. In this paper, we propose to use cache partitioning together with cache bypassing as the shared cache management mechanism for multi-tasking on GPUs. The combined approach aims to reduce the interference among the tasks and preserve the locality for each task. However, the interplay among the cache partitioning and bypassing brings greater challenges. On one hand, the partitioned cache space to each task affects its cache bypassing decision. On the other hand, cache bypassing affects the cache capacity required for each task. To address this, we propose a two-step approach. First, we use cache partitioning to assign dedicated cache space to each task to reduce the interference among the tasks. During this process, we compare cache partitioning with coarse-grained cache bypassing. Then, we use fine-grained cache bypassing to selectively bypass certain data requests and threads for each task. We explore different cache partitioning and bypassing designs and demonstrate the potential benefits of this approach. Experiments using a wide range of applications demonstrate that our technique improves the overall system throughput by 52% on average compared to the default multi-tasking solution on GPUs.

DD-L1D: Improving the Decoupled L1D Efficiency for GPU Architecture

Locality-Driven Dynamic Gpu Cache Bypassing

Selectively GPU Cache Bypassing for Un-Coalesced Loads.

Improving CPU and GPU Performance Through Sample-Based Dynamic LLC Bypassing

Coordinated Static and Dynamic Cache Bypassing for GPUs

DyCache: Dynamic Multi-Grain Cache Management for Irregular Memory Accesses on GPU

Locality Protected Dynamic Cache Allocation Scheme on GPUs

Locality-Based Cache Management and Warp Scheduling for Reducing Cache Contention in GPU

RACB: Resource Aware Cache Bypass on GPUs

A Sharing-Aware L1.5D Cache for Data Reuse in GPGPUs

A model-driven approach to warp/thread-block level GPU cache bypassing.

Adaptive Cache Management for Energy-Efficient GPU Computing.

Adaptive Cache Bypass and Insertion for Many-core Accelerators

An Efficient Compiler Framework for Cache Bypassing on GPUs

Locality-protected Cache Allocation Scheme with Low Overhead on GPUs.

Optimizing Cache Bypassing and Warp Scheduling for GPUs

Improve Llc Bypassing Performance By Memory Controller Improvements In Heterogeneous Multicore System

Exploring Cache Bypassing and Partitioning for Multi-Tasking on GPUs

Orchestrating Cache Management and Memory Scheduling for GPGPU Applications.

Understanding the Tradeoffs Between Software-Managed Vs. Hardware-Managed Caches in GPUs

Incorporating Selective Victim Cache into GPGPU for High-Performance Computing