Abstract:With the continued digitization of societal processes, we are seeing an explosion in available data. This is referred to as big data. In a research setting, three aspects of the data are often viewed as the main sources of challenges when attempting to enable value creation from big data: volume, velocity, and variety. Many studies address volume or velocity, while fewer studies concern the variety. Metric spaces are ideal for addressing variety because they can accommodate any data as long as it can be equipped with a distance notion that satisfies the triangle inequality. To accelerate search in metric spaces, a collection of indexing techniques for metric data have been proposed. However, existing surveys offer limited coverage, and a comprehensive empirical study exists has yet to be reported. We offer a comprehensive survey of existing metric indexes that support exact similarity search: we summarize existing partitioning, pruning, and validation techniques used by metric indexes to support exact similarity search; we provide the time and space complexity analyses of index construction; and we offer an empirical comparison of their query processing performance. Empirical studies are important when evaluating metric indexing performance, because performance can depend highly on the effectiveness of available pruning and validation as well as on the data distribution, which means that complexity analyses often offer limited insights. This article aims at revealing strengths and weaknesses of different indexing techniques to offer guidance on selecting an appropriate indexing technique for a given setting, and to provide directions for future research on metric indexing.

Influence Of Data Set Splitting Methods On Similarity Indexing Performance

Speed Partitioning for Indexing Moving Objects.

VPIndexer: velocity-based partitioning for indexing moving objects.

Performance evaluation of air indexing schemes for multi-attribute data broadcast

Indexing Metric Spaces for Exact Similarity Search

Efficient Metric Indexing for Similarity Search

An Efficient Framework for Exact Set Similarity Search Using Tree Structure Indexes.

Pivot-based Metric Indexing

DIMS: Distributed Index for Similarity Search in Metric Spaces

Evaluating instructors' performance based on Set Pair Analysis

A Study of Performance Optimization Method for Massive Spaito-temporal Data Based on Spatio-temporal Partition Clustering

DIDS: Double Indices and Double Summarizations for Fast Similarity Search

A Learned Index for Exact Similarity Search in Metric Spaces

Envelope parameter calculation of similarity indexing structure

Pivot Selection Algorithms in Metric Spaces: a Survey and Experimental Study

Efficient Multimedia Similarity Measurement Using Similar Elements

Indexing Dataspaces with Partitions.

A Bi-metric Framework for Fast Similarity Search

DESIRE: An Efficient Dynamic Cluster-based Forest Indexing for Similarity Search in Multi-Metric Spaces

iDistance: An adaptive B+-tree based indexing method for nearest neighbor search

Efficient Indexing for Large Scale Visual Search