Abstract:Efficiently analyzing geo-distributed datasets is emerging as a major demand in a cloud-edge system. Since the datasets are often generated in closer proximity to end users, traditional works mainly focus on offloading proper tasks from those hotspot edges to the datacenter to decrease the overall completion time of submitted jobs in a one-shot manner. However, optimizing the completion time of <italic>current job</italic> alone is insufficient in a long-term scope since some datasets would be used multiple times. Instead, optimizing the data distribution is much more efficient and could directly benefit forthcoming jobs, although it may postpone the execution of current one. Unfortunately, due to the throwaway feature of data fetcher, existing data analytics systems fail to re-distribute corresponding data out of hotspot edges after the execution of data analytics. In order to minimize the overall completion time for a <italic>sequence</italic> of jobs as well as to guarantee the performance of current one, we propose to re-distribute the data along with task offloading, and formulate corresponding <inline-formula><tex-math notation="LaTeX">$\varepsilon$</tex-math> <alternatives><mml:math><mml:mi>ɛ</mml:mi></mml:math><inline-graphic xlink:href="qian-ieq2-3086274.gif"/></alternatives></inline-formula>-bounded data-driven task scheduling problem over wide area network under the consideration of edge heterogeneity. We design an online schema <inline-formula><tex-math notation="LaTeX">$run$</tex-math> <alternatives><mml:math><mml:mrow><mml:mi>r</mml:mi><mml:mi>u</mml:mi><mml:mi>n</mml:mi></mml:mrow></mml:math><inline-graphic xlink:href="qian-ieq3-3086274.gif"/></alternatives></inline-formula>Data, which offloads proper tasks and related data via piggybacking to the datacenter based on delicately calculated probabilities. Through rigorous theoretical analysis, <inline-formula><tex-math notation="LaTeX">$run$</tex-math> <alternatives><mml:math><mml:mrow><mml:mi>r</mml:mi><mml:mi>u</mml:mi><mml:mi>n</mml:mi></mml:mrow></mml:math><inline-graphic xlink:href="qian-ieq4-3086274.gif"/></alternatives></inline-formula>Data is proved concentrated on its optimum with high probability. We implement <inline-formula><tex-math notation="LaTeX">$run$</tex-math> <alternatives><mml:math><mml:mrow><mml:mi>r</mml:mi><mml:mi>u</mml:mi><mml:mi>n</mml:mi></mml:mrow></mml:math><inline-graphic xlink:href="qian-ieq5-3086274.gif"/></alternatives></inline-formula>Data based on Spark and HDFS. Both testbed results and trace-driven simulations show that <inline-formula><tex-math notation="LaTeX">$run$</tex-math> <alternatives><mml:math><mml:mrow><mml:mi>r</mml:mi><mml:mi>u</mml:mi><mml:mi>n</mml:mi></mml:mrow></mml:math><inline-graphic xlink:href="qian-ieq6-3086274.gif"/></alternatives></inline-formula>Data re-distributes proper data via piggybacking and achieves up to 37 percent reduction on average response time compared with state-of-the-art schemas.

Efficient Data Aggregation Transfers in Data Center Networks

On Traffic-Aware Partition and Aggregation in MapReduce for Big Data Applications

Accelerating Distributed Training Through In-network Aggregation with Idle Resources in Data Centers

Resource Allocation Considering Impact of Network on Performance in a Disaggregated Data Center

Improving performance by network-aware virtual machine clustering and consolidation

On Efficient Data Transfers Across Geographically Dispersed Datacenters

Flexible and Efficient Multicast Transfers in Inter-Datacenter Networks

Constrained In-network Computing with Low Congestion in Datacenter Networks

Exploring Efficient and Scalable Multicast Routing in Future Data Center Networks

A MapReduce-supported Network Structure for Data Centers

MiniForest: Distributed and Dynamic Multicasting in Datacenter Networks

Energy-aware Routing in Data Center Network

QoS-Aware Data Placement for MapReduce Applications in Geo-Distributed Data Centers

Distributed Algorithm for Tree-Structured Data Aggregation Service Placement in Smart Grid

In-Network Aggregation with Transport Transparency for Distributed Training

Multicast Routing and Recovery Based in Data Center Networks.

MIP: Minimizing the idle period of data transmission in data center networks

Opportunistic Data Aggregation Algorithm Using Any-cast

Moving big data to the cloud

Efficient Inter-Datacenter AllReduce With Multiple Trees