Abstract:Most distributed stream processing engines (DSPEs) do not support online task management and cannot adapt to time-varying data flows. Recently, some studies have proposed online task deployment algorithms to solve this problem. However, these approaches do not guarantee the Quality of Service (QoS) when the task deployment changes at runtime, because the task migrations caused by the change of task deployments will impose an exorbitant cost. We study one of the most popular DSPEs, Apache Storm, and find out that when a task needs to be migrated, Storm has to stop the resource (implemented as a process of Worker in Storm) where the task is deployed. This will lead to the stop and restart of all tasks in the resource, resulting in the poor performance of task migrations. Aiming to solve this problem, in this paper, we propose N-Storm (Nonstop Storm), which is a task-resource decoupling DSPE. N-Storm allows tasks allocated to resources to be changed at runtime, which is implemented by a thread-level scheme for task migrations. Particularly, we add a local shared key/value store on each node to make resources aware of the changes in the allocation plan. Thus, each resource can manage its tasks at runtime. Based on N-Storm, we further propose Online Task Deployment (OTD). Differing from traditional task deployment algorithms that deploy all tasks at once without considering the cost of task migrations caused by a task re-deployment, OTD can gradually adjust the current task deployment to an optimized one based on the communication cost and the runtime states of resources. We demonstrate that OTD can adapt to different kinds of applications including computation- and communication-intensive applications. The experimental results on a real DSPE cluster show that N-Storm can avoid the system stop and save up to 87% of the performance degradation time, compared with Apache Storm and other state-of-the-art approaches. In addition, OTD can increase the average CPU usage by 51% for computation-intensive applications and reduce network communication costs by 88% for communication-intensive applications.

Dynamic Resource Management In a Massively Parallel Stream Processing Engine

Using Paralleled-PEs Method to Resolve the Bursting Data in Distributed Stream Processing System

Fault Tolerance for Stream Processing Engines

Online Nonstop Task Management for Storm-Based Distributed Stream Processing Engines

Demeter: Resource-Efficient Distributed Stream Processing under Dynamic Loads with Multi-Configuration Optimization

DRS: Dynamic Resource Scheduling for Real-Time Analytics over Fast Streams

Stromax: Partitioning-Based Scheduler For Real-Time Stream Processing System

Distributed and Parallel Processing for Real-Time and Dynamic Spatio-Temporal Graph

System-aware dynamic partitioning for batch and streaming workloads

An Adaptive Online Scheme for Scheduling and Resource Enforcement in Storm.

Optimal Operator State Migration for Elastic Data Stream Processing

Optimal Multi-Level Interval-based Checkpointing for Exascale Stream Processing Systems

Performance Evaluation of an Algorithm-based Asynchronous Checkpoint-Restart Fault Tolerant Application Using Mixed MPI/GPI-2

An On-the-Fly Scheduling Strategy for Distributed Stream Processing Platform.

Model-driven Scheduling for Distributed Stream Processing Systems

Benchmarking Distributed Stream Data Processing Systems

Adaptive Scheduling for Efficient Execution of Dynamic Stream Workflows

A Utilization Model for Optimization of Checkpoint Intervals in Distributed Stream Processing Systems

Building a Fault Tolerant Application Using the GASPI Communication Layer

Application-Level Differential Checkpointing for HPC Applications with Dynamic Datasets

FastFlow: Efficient Scalable Model-Driven Framework for Processing Massive Mobile Stream Data.