Abstract:Recently, MapReduce based spatial query systems have emerged as a cost effective and scalable solution to large scale spatial data processing and analytics. MapReduce based systems achieve massive scalability by partitioning the data and running query tasks on those partitions in parallel. Therefore, effective data partitioning is critical for task parallelization, load balancing, and directly affects system performance. However, several pitfalls of spatial data partitioning make this task particularly challenging. First, data skew is very common in spatial applications. To achieve best query performance, data skew need to be reduced. Second, spatial partitioning approaches generate boundary objects that cross multiple partitions, and add extra query processing overhead. Consequently, boundary objects need to be minimized. Third, the high computational complexity of spatial partitioning algorithms combined with massive amounts of data require an efficient approach for partitioning to achieve overall fast query response. In this paper, we provide a systematic evaluation of multiple spatial partitioning methods with a set of different partitioning strategies, and study their implications on the performance of MapReduce based spatial queries. We also study sampling based partitioning methods and their impact on queries, and propose several MapReduce based high performance spatial partitioning methods. The main objective of our work is to provide a comprehensive guidance for optimal spatial data partitioning to support scalable and fast spatial data processing in massively parallel data processing frameworks such as MapReduce. The algorithms developed in this work are open source and can be easily integrated into different high performance spatial data processing systems.

Optimizing Data Partition for Scaling out Nosql Cluster

Optimizing Data Partition For Nosql Cluster

A Request Skew Aware Heterogeneous Distributed Storage System Based on Cassandra

Data Based Application Partitioning and Workload Balance in Distributed Environment

Distributed Model Based on Data Partition and Load Balance Algorithm

Design of A More Scalable Database System

New Balanced Data Allocating and Online Migrating Algorithms in Database Cluster

Heterogeneous Replicas for Multi-dimensional Data Management

Cost-Based Optimization Of Logical Partitions For A Query Workload In A Hadoop Data Warehouse

An Adaptive Model For Building Service-Partition System

Optimizing Hadoop Block Placement Policy and Cluster Blocks Distribution

Achieving Load-Balanced, Redundancy-Free Cluster Caching with Selective Partition

A spatial data partition algorithm based on statistical cluster

Effective Spatial Data Partitioning for Scalable Query Processing

DynaHash: Efficient Data Rebalancing in Apache AsterixDB (Extended Version)

A Parallel Varied Density-Based Clustering Algorithm with Optimized Data Partition

Optimizing Data Center Traffic of Online Social Networks

Enhancing Storage Efficiency and Performance: A Survey of Data Partitioning Techniques

A Real-Time Partition Generation Mechanism for Data Skew Mitigation in Spark Computing Environment

A survey of data partitioning and sampling methods to support big data analysis

Location-Aware Data Block Allocation Strategy for HDFS-Based Applications in the Cloud