Abstract:Efficiently analyzing geo-distributed datasets is emerging as a major demand in a cloud-edge system. Since the datasets are often generated in closer proximity to end users, traditional works mainly focus on offloading proper tasks from those hotspot edges to the datacenter to decrease the overall completion time of submitted jobs in a one-shot manner. However, optimizing the completion time of <italic>current job</italic> alone is insufficient in a long-term scope since some datasets would be used multiple times. Instead, optimizing the data distribution is much more efficient and could directly benefit forthcoming jobs, although it may postpone the execution of current one. Unfortunately, due to the throwaway feature of data fetcher, existing data analytics systems fail to re-distribute corresponding data out of hotspot edges after the execution of data analytics. In order to minimize the overall completion time for a <italic>sequence</italic> of jobs as well as to guarantee the performance of current one, we propose to re-distribute the data along with task offloading, and formulate corresponding <inline-formula><tex-math notation="LaTeX">$\varepsilon$</tex-math> <alternatives><mml:math><mml:mi>ɛ</mml:mi></mml:math><inline-graphic xlink:href="qian-ieq2-3086274.gif"/></alternatives></inline-formula>-bounded data-driven task scheduling problem over wide area network under the consideration of edge heterogeneity. We design an online schema <inline-formula><tex-math notation="LaTeX">$run$</tex-math> <alternatives><mml:math><mml:mrow><mml:mi>r</mml:mi><mml:mi>u</mml:mi><mml:mi>n</mml:mi></mml:mrow></mml:math><inline-graphic xlink:href="qian-ieq3-3086274.gif"/></alternatives></inline-formula>Data, which offloads proper tasks and related data via piggybacking to the datacenter based on delicately calculated probabilities. Through rigorous theoretical analysis, <inline-formula><tex-math notation="LaTeX">$run$</tex-math> <alternatives><mml:math><mml:mrow><mml:mi>r</mml:mi><mml:mi>u</mml:mi><mml:mi>n</mml:mi></mml:mrow></mml:math><inline-graphic xlink:href="qian-ieq4-3086274.gif"/></alternatives></inline-formula>Data is proved concentrated on its optimum with high probability. We implement <inline-formula><tex-math notation="LaTeX">$run$</tex-math> <alternatives><mml:math><mml:mrow><mml:mi>r</mml:mi><mml:mi>u</mml:mi><mml:mi>n</mml:mi></mml:mrow></mml:math><inline-graphic xlink:href="qian-ieq5-3086274.gif"/></alternatives></inline-formula>Data based on Spark and HDFS. Both testbed results and trace-driven simulations show that <inline-formula><tex-math notation="LaTeX">$run$</tex-math> <alternatives><mml:math><mml:mrow><mml:mi>r</mml:mi><mml:mi>u</mml:mi><mml:mi>n</mml:mi></mml:mrow></mml:math><inline-graphic xlink:href="qian-ieq6-3086274.gif"/></alternatives></inline-formula>Data re-distributes proper data via piggybacking and achieves up to 37 percent reduction on average response time compared with state-of-the-art schemas.

Shadow: Exploiting the Power of Choice for Efficient Shuffling in MapReduce.

Efficient. Scalable and Robust Data Shuffle Service for Distributed MapReduce Computing on Cloud

Join Query Optimization Based on MapReduce under Skewed Data

Templating Shuffles

Exoshuffle: An Extensible Shuffle Architecture

SHadoop: Improving MapReduce Performance by Optimizing Job Execution Mechanism in Hadoop Clusters

Optimization and reconstruction shuffle in MapReduce

Hybrid Coded MapReduce for Latency-Constrained Tasks with Straggling Servers

Performance Analysis of Randomized Data Fetching in Cluster Computing.

Joint Design of Shuffling and Function Assignment in Heterogeneous Coded Distributed Computing

Load scheduling for distributed edge computing: A communication-computation tradeoff

LIBRA: Lightweight Data Skew Mitigation in MapReduce

Performance Optimization for Short MapReduce Job Execution in Hadoop

Moving Hadoop into the Cloud with Flexible Slot Management and Speculative Execution

On Implementation of Parallel Map and Shuffle Phases for Coded Distributed Computing

Improving MapReduce Performance Using Smart Speculative Execution Strategy

Run Data Run! Re-Distributing Data via Piggybacking for Geo-Distributed Data Analytics

Optimizing MapReduce for Highly Distributed Environments

A Real-Time Partition Generation Mechanism for Data Skew Mitigation in Spark Computing Environment

Privacy Enhancement Via Dummy Points in the Shuffle Model