Abstract:Fuzzy similarity join is an important database operator widely used in practice. So far the research community has focused exclusively on optimizing fuzzy joinscalability. However, practitioners today also struggle to optimize fuzzy-joinquality, because they face a daunting space of parameters (e.g., distance-functions, distance-thresholds, tokenization-options, etc.), and often have to resort to a manual trial-and-error approach to program these parameters in order to optimize fuzzy-join quality. This key challenge of automatically generating high-quality fuzzy-join programs has received surprisingly little attention thus far. In this work, we study the problem of "auto-program'' fuzzy-joins. Leveraging a geometric interpretation of distance-functions, we develop an unsupervised Auto-FuzzyJoin framework that can infer suitable fuzzy-join programs on given input tables, without requiring explicit human input such as labelled training data. Using Auto-FuzzyJoin, users only need to provide two input tables L and R, and a desired precision target τ (say 0.9). Auto-FuzzyJoin leverages the fact that one of the input is a reference table to automatically program fuzzy-joins that meet the precision target τ in expectation, while maximizing fuzzy-join recall (defined as the number of correctly joined records). Experiments on both existing benchmarks and a new benchmark with 50 fuzzy-join tasks created from Wikipedia data suggest that the proposed Auto-FuzzyJoin significantly outperforms existing unsupervised approaches, and is surprisingly competitive even against supervised approaches (e.g., Magellan and DeepMatcher) when 50% of ground-truth labels are used as training data. We have released our code and benchmark on GitHub\footnote\urlhttps://github.com/chu-data-lab/AutomaticFuzzyJoin to facilitate future research.

Fast-join: an Efficient Method for Fuzzy Token Matching Based String Similarity Join

Extending String Similarity Join to Tolerant Fuzzy Token Matching

Improved LSH-driven String Similarity Join Filtering-Verification Framework

Massjoin: A Mapreduce-Based Method for Scalable String Similarity Joins

Improve Semantic Web Services Discovery Through Similarity Search in Metric Space

Trie-join: a Trie-Based Method for Efficient String Similarity Joins

PASS-JOIN: A Partition-based Method for Similarity Joins

String Similarity Joins

Scalable Similarity Joins of Tokenized Strings

Effective Indices for Efficient Approximate String Search and Similarity Join

A Partition-Based Method for String Similarity Joins with Edit-Distance Constraints

An Efficient MapReduce Algorithm for Similarity Join in Metric Spaces

Auto-FuzzyJoin: Auto-Program Fuzzy Similarity Joins Without Labeled Examples

Star-Join: spatio-textual similarity join.

Similarity Joins Of Text With Incomplete Information Formats

An Efficient Partition Based Method for Exact Set Similarity Joins

Can we beat the prefix filtering?: an adaptive framework for similarity join and search.

Auto-FuzzyJoin

String similarity search and join: a survey

Efficient String Similarity Join in Multi-Core and Distributed Systems.

Dynamic Set Similarity Join: an Update Log Based Approach