Abstract:As an essential operation in data cleaning, the similarity join has attracted considerable attention from the database community. In this article, we study string similarity joins with edit-distance constraints, which find similar string pairs from two large sets of strings whose edit distance is within a given threshold. Existing algorithms are efficient either for short strings or for long strings, and there is no algorithm that can efficiently and adaptively support both short strings and long strings. To address this problem, we propose a new filter, called the segment filter . We partition a string into a set of segments and use the segments as a filter to find similar string pairs. We first create inverted indices for the segments. Then for each string, we select some of its substrings, identify the selected substrings from the inverted indices, and take strings on the inverted lists of the found substrings as candidates of this string. Finally, we verify the candidates to generate the final answer. We devise efficient techniques to select substrings and prove that our method can minimize the number of selected substrings. We develop novel pruning techniques to efficiently verify the candidates. We also extend our techniques to support normalized edit distance. Experimental results show that our algorithms are efficient for both short strings and long strings, and outperform state-of-the-art methods on real-world datasets.

Hash(Ed)-Join: Approximate String Similarity Join With Hashing

Improved LSH-driven String Similarity Join Filtering-Verification Framework

Trie-join: a Trie-Based Method for Efficient String Similarity Joins

Set Similarity Join Using Partition Index

A Partition-Based Method for String Similarity Joins with Edit-Distance Constraints

Effective Indices for Efficient Approximate String Search and Similarity Join

Massjoin: A Mapreduce-Based Method for Scalable String Similarity Joins

String Similarity Joins

PASS-JOIN: A Partition-based Method for Similarity Joins

An Efficient Framework for Exact Set Similarity Search Using Tree Structure Indexes.

An Efficient MapReduce Algorithm for Similarity Join in Metric Spaces

Efficient Similarity Join and Search on Multi-Attribute Data

An Efficient Partition Based Method for Exact Set Similarity Joins

Min-Max Hash for Jaccard Similarity

Fast and Accurate Hashing Via Iterative Nearest Neighbors Expansion.

Fast-join: an Efficient Method for Fuzzy Token Matching Based String Similarity Join

A Unified Approximate Nearest Neighbor Search Scheme by Combining Data Structure and Hashing.

A unified framework for string similarity search with edit-distance constraint

Efficient String Similarity Join in Multi-Core and Distributed Systems.

Reinforcing Short-Length Hashing

Complementary Hashing for Approximate Nearest Neighbor Search