Abstract:Recently, MapReduce based spatial query systems have emerged as a cost effective and scalable solution to large scale spatial data processing and analytics. MapReduce based systems achieve massive scalability by partitioning the data and running query tasks on those partitions in parallel. Therefore, effective data partitioning is critical for task parallelization, load balancing, and directly affects system performance. However, several pitfalls of spatial data partitioning make this task particularly challenging. First, data skew is very common in spatial applications. To achieve best query performance, data skew need to be reduced. Second, spatial partitioning approaches generate boundary objects that cross multiple partitions, and add extra query processing overhead. Consequently, boundary objects need to be minimized. Third, the high computational complexity of spatial partitioning algorithms combined with massive amounts of data require an efficient approach for partitioning to achieve overall fast query response. In this paper, we provide a systematic evaluation of multiple spatial partitioning methods with a set of different partitioning strategies, and study their implications on the performance of MapReduce based spatial queries. We also study sampling based partitioning methods and their impact on queries, and propose several MapReduce based high performance spatial partitioning methods. The main objective of our work is to provide a comprehensive guidance for optimal spatial data partitioning to support scalable and fast spatial data processing in massively parallel data processing frameworks such as MapReduce. The algorithms developed in this work are open source and can be easily integrated into different high performance spatial data processing systems.

Handling Partitioning Skew in MapReduce Using LEEN

Join Query Optimization Based on MapReduce under Skewed Data

Performance Evaluation for Distributed Join Based on MapReduce.

LIBRA: Lightweight Data Skew Mitigation in MapReduce

A Comparative Study of Data Skew in Hadoop

Handling Data Skew at Reduce Stage in Spark by ReducePartition

Load Balancing In Heterogeneous Mapreduce Environments

A Real-Time Partition Generation Mechanism for Data Skew Mitigation in Spark Computing Environment

Efficient. Scalable and Robust Data Shuffle Service for Distributed MapReduce Computing on Cloud

Effective Spatial Data Partitioning for Scalable Query Processing

SP-Partitioner: A novel partition method to handle intermediate data skew in spark streaming

Learning-based distributed locality sensitive hashing.

The performance of MapReduce: an in-depth study

A load balance algorithm based on nodes performance in Hadoop cluster

On Traffic-Aware Partition and Aggregation in MapReduce for Big Data Applications

A Heterogeneity-aware Data Distribution and Rebalance Method in Hadoop Cluster

The Performance of MapReduce

Dependency-Aware Data Locality For Mapreduce

TS-Hadoop: Handling Access Skew in MapReduce by Using Tiered Storage Infrastructure

Matchmaking: A New MapReduce Scheduling Technique

Improving MapReduce Performance in a Heterogeneous Cloud: A Measurement Study