Abstract:In the rapidly expanding field of parallel processing, job schedulers are the "operating systems" of modern big data architectures and supercomputing systems. Job schedulers allocate computing resources and control the execution of processes on those resources. Historically, job schedulers were the domain of supercomputers, and job schedulers were designed to run massive, long-running computations over days and weeks. More recently, big data workloads have created a need for a new class of computations consisting of many short computations taking seconds or minutes that process enormous quantities of data. For both supercomputers and big data systems, the efficiency of the job scheduler represents a fundamental limit on the efficiency of the system. Detailed measurement and modeling of the performance of schedulers are critical for maximizing the performance of a large-scale computing system. This paper presents a detailed feature analysis of 15 supercomputing and big data schedulers. For big data workloads, the scheduler latency is the most important performance characteristic of the scheduler. A theoretical model of the latency of these schedulers is developed and used to design experiments targeted at measuring scheduler latency. Detailed benchmarking of four of the most popular schedulers (Slurm, Son of Grid Engine, Mesos, and Hadoop YARN) are conducted. The theoretical model is compared with data and demonstrates that scheduler performance can be characterized by two key parameters: the marginal latency of the scheduler $t_s$ and a nonlinear exponent $\alpha_s$. For all four schedulers, the utilization of the computing system decreases to < 10\% for computations lasting only a few seconds. Multilevel schedulers that transparently aggregate short computations can improve utilization for these short computations to > 90\% for all four of the schedulers that were tested.

Low Latency Big Data Processing Without Prior Information

Job Scheduling Without Prior Information in Big Data Processing Systems

PAS: Performance-Aware Job Scheduling for Big Data Processing Systems

Scalable System Scheduling for HPC and Big Data

Preemptive and Low Latency Datacenter Scheduling via Lightweight Containers

A Real-Time Scheduling Strategy Based on Processing Framework of Hadoop

Job-Aware Scheduling for Big Data Processing

Job schedulers for Big data processing in Hadoop environment: testing real-life schedulers using benchmark programs

Job-aware Network Scheduling for Hadoop Cluster.

Deadline-aware Preemptive Job Scheduling in Hadoop YARN Clusters

An Adaptive Scheduling Algorithm for Heterogeneous Hadoop Systems

Performance-Aware Fair Scheduling: Exploiting Demand Elasticity of Data Analytics Jobs

New Scheduling Algorithms for Improving Performance and Resource Utilization in Hadoop YARN Clusters

Minimizing Interference and Maximizing Progress for Hadoop Virtual Machines

A Near Optimal Multi-Faced Job Scheduler For Datacenter Workloads

An Efficiency-Aware Scheduling for Data-Intensive Computations on Mapreduce Clusters

K%-Fair Scheduling: A Flexible Task Scheduling Strategy for Balancing Fairness and Efficiency in MapReduce Systems

Hadoop Scheduling Base On Data Locality

A New Schedule Strategy for Heterogenous Workload-aware in Hadoop

Performance optimization of computing task scheduling based on the Hadoop big data platform

Matchmaking: A New MapReduce Scheduling Technique