Abstract:Phase analysis, which classifies the set of execution intervals with similar execution behavior and resource requirements, has been widely used in a variety of systems, including dynamic cache reconfiguration, prefetching, race detection, and sampling simulation. Although phase granularity has been a major factor in the accuracy of phase analysis, it has not been well investigated, and most systems usually adopt a fine-grained scheme. However, such a scheme can only take account of recent local phase information and could be frequently interfered by temporary noise due to instant phase changes, which might notably limit the accuracy. In this article, we make the first investigation on the potential of multilevel phase analysis (MLPA), where different granularity phase analyses are combined together to improve the overall accuracy. The key observation is that the coarse-grained intervals belonging to the same phase usually consist of stably distributed fine-grained phases. Moreover, the phase of a coarse-grained interval can be accurately identified based on the fine-grained intervals at the beginning of its execution. Based on the observation, we design and implement an MLPA scheme. In such a scheme, a coarse-grained phase is first identified based on the fine-grained intervals at the beginning of its execution. The following fine-grained phases in it are then predicted based on the sequence of fine-grained phases in the coarse-grained phase. Experimental results show that such a scheme can notably improve the prediction accuracy. Using a Markov fine-grained phase predictor as the baseline, MLPA can improve prediction accuracy by 20&percnt;, 39&percnt;, and 29&percnt; for next phase, phase change, and phase length prediction for SPEC2000, respectively, yet incur only about 2&percnt; time overhead and 40&percnt; space overhead (about 360 bytes in total). To demonstrate the effectiveness of MLPA, we apply it to a dynamic cache reconfiguration system that dynamically adjusts the cache size to reduce the power consumption and access time of the data cache. Experimental results show that MLPA can further reduce the average cache size by 15&percnt; compared to the fine-grained scheme. Moreover, for MLPA, we also observe that coarse-grained phases can better capture the overall program characteristics with fewer of phases and the last representative phase could be classified in a very early program position, leading to fewer execution internals being functionally simulated. Based on this observation, we also design a multilevel sampling simulation technique that combines both fine- and coarse-grained phase analysis for sampling simulation. Such a scheme uses fine-grained simulation points to represent only the selected coarse-grained simulation points instead of the entire program execution; thus, it could further reduce both the functional and detailed simulation time. Experimental results show that MLPA for sampling simulation can achieve a speedup in simulation time of about 8.3X with similar accuracy compared to 10M SimPoint.

MBAPIS: Multi-Level Behavior Analysis Guided Program Interval Selection for Microarchitecture Studies

WPC: Whole-Picture Workload Characterization Across Intermediate Representation, ISA, and Microarchitecture

Multilevel Phase Analysis

A Hardware-Software Cooperative Interval-Replaying for FPGA-based Architecture Evaluation

A Theoretical Research on Program Instruction Level Parallelism to Guide Microprocessor Design

A Hybrid Modeling Approach to Microarchitecture Design Space Exploring

Improving Dynamic Prediction Accuracy Through Multi-Level Phase Analysis

Measuring Microarchitectural Details of Multi- and Many-Core Memory Systems Through Microbenchmarking

Memory Performance Characterization Of Spec Cpu2006 Benchmarks Using Tsim1

Improving the Representativeness of Simulation Intervals for the Cache Memory System

Multi-level Phase Analysis for Sampling Simulation

MNSIM 2.0: A Behavior-Level Modeling Tool for Processing-In-Memory Architectures.

McPAT-Calib - A Microarchitecture Power Modeling Framework for Modern CPUs.

Benchmarking End-To-End Performance of AI-Based Chip Placement Algorithms

PCAsim: A parallel cycle accurate simulation platform for CMPs

A Novel Memory Subsystem Evaluation Framework for Chip Multiprocessors.

Focusing processor policies via critical-path prediction

Design and implementation of a configurable hardware profiler supporting path profiling and sampling

Hot spots profiling and dataflow analysis in custom dataflow computing SoftProcessors

An Approach to Build Cycle Accurate Full System VLIW Simulation Platform.

Microprocessor architecture-level "Y-behavior" research for branch instructions