Abstract:In the rapidly expanding field of parallel processing, job schedulers are the "operating systems" of modern big data architectures and supercomputing systems. Job schedulers allocate computing resources and control the execution of processes on those resources. Historically, job schedulers were the domain of supercomputers, and job schedulers were designed to run massive, long-running computations over days and weeks. More recently, big data workloads have created a need for a new class of computations consisting of many short computations taking seconds or minutes that process enormous quantities of data. For both supercomputers and big data systems, the efficiency of the job scheduler represents a fundamental limit on the efficiency of the system. Detailed measurement and modeling of the performance of schedulers are critical for maximizing the performance of a large-scale computing system. This paper presents a detailed feature analysis of 15 supercomputing and big data schedulers. For big data workloads, the scheduler latency is the most important performance characteristic of the scheduler. A theoretical model of the latency of these schedulers is developed and used to design experiments targeted at measuring scheduler latency. Detailed benchmarking of four of the most popular schedulers (Slurm, Son of Grid Engine, Mesos, and Hadoop YARN) are conducted. The theoretical model is compared with data and demonstrates that scheduler performance can be characterized by two key parameters: the marginal latency of the scheduler $t_s$ and a nonlinear exponent $\alpha_s$. For all four schedulers, the utilization of the computing system decreases to < 10\% for computations lasting only a few seconds. Multilevel schedulers that transparently aggregate short computations can improve utilization for these short computations to > 90\% for all four of the schedulers that were tested.

Towards Scalable Resource Management for Supercomputers

End-to-End Predictions-Based Resource Management Framework for Supercomputer Jobs

Exploring Plan-Based Scheduling for Large-Scale Computing Systems

Docker-Enabled Scalable Parallel MLFMA System for RCS Evaluation

Scalable System Scheduling for HPC and Big Data

Job Scheduling in High Performance Computing

Improving HPC System Performance by Predicting Job Resources via Supervised Machine Learning

Comprehensive Analysis of Performance, Fault-Tolerance and Scalability in Grid Resource Management System

Full Lifecycle Data Analysis on a Large-scale and Leadership Supercomputer: What Can We Learn from It?

A HPC Co-Scheduler with Reinforcement Learning

Multiple Resources Scheduling for Diverse Workloads in Heterogeneous Datacenter

Energy hardware and workload aware job scheduling towards interconnected HPC environments

ROME: A Multi-Resource Job Scheduling Framework for Exascale HPC Systems

Scalable Hierarchical Parallel Algorithm for the Solution of Super Large-Scale Sparse Linear Equations

Node-Based Job Scheduling for Large Scale Simulations of Short Running Jobs

Online Resource Management in Thermal and Energy Constrained Heterogeneous High Performance Computing

Toward a Dynamic Allocation Strategy for Deadline‐Oriented Resource and Job Management in HPC Systems

A Dynamic, Hierarchical Resource Model for Converged Computing

Performability Evaluation of Resource Scheduling Algorithms for Computational Grids

DRESS: Dynamic RESource-Reservation Scheme for Congested Data-Intensive Computing Platforms

MRSch: Multi-Resource Scheduling for HPC