Abstract:Strings form a fundamental data type in computer systems. String searching has been extensively studied since the inception of computer science. Increasingly many applications have to deal with imprecise strings or strings with fuzzy information in them. String matching becomes a probabilistic event when a string contains uncertainty, i.e. each position of the string can have different probable characters with associated probability of occurrence for each character. Such uncertain strings are prevalent in various applications such as biological sequence data, event monitoring and automatic ECG annotations. We explore the problem of indexing uncertain strings to support efficient string searching. In this paper we consider two basic problems of string searching, namely substring searching and string listing. In substring searching, the task is to find the occurrences of a deterministic string in an uncertain string. We formulate the string listing problem for uncertain strings, where the objective is to output all the strings from a collection of strings, that contain probable occurrence of a deterministic query string. Indexing solution for both these problems are significantly more challenging for uncertain strings than for deterministic strings. Given a construction time probability value $\tau$, our indexes can be constructed in linear space and supports queries in near optimal time for arbitrary values of probability threshold parameter greater than $\tau$. To the best of our knowledge, this is the first indexing solution for searching in uncertain strings that achieves strong theoretical bound and supports arbitrary values of probability threshold parameter. We also propose an approximate substring search index that can answer substring search queries with an additive error in optimal time. We conduct experiments to evaluate the performance of our indexes.

A Probabilistic Approach to String Transformation

A Probabilistic Method for Tag Ranking in Tagging System

A Fast and Accurate Method for Approximate String Search

A Neural Probabilistic Structured-Prediction Method for Transition-Based Natural Language Processing.

The role of grammar in transition-probabilities of subsequent words in English text

Building Probabilistic Models for Natural Language

Robustness to Programmable String Transformations via Augmented Abstract Training

Probabilistic Automata for Computing with Words

Predicting from Strings: Language Model Embeddings for Bayesian Optimization

String Re-Writing Kernel

Probabilistic Threshold Indexing for Uncertain Strings

An Introduction to String Re-Writing Kernel

Keyword Query Reformulation on Structured Data

Probabilistic Transformer: A Probabilistic Dependency Model for Contextual Word Representation

Probabilistic Linguistic Knowledge and Token-level Text Augmentation

A Constraint-Based Probabilistic Framework for Name Disambiguation

A neural probabilistic structured-prediction method for transition-based natural language processing

Estimating Translation Probabilities for Social Tag Suggestion

A Probabilistic Approach to Knowledge Translation

An Optimized String Transformation Algorithm for Real-Time Group Editors

Learning Alternative Name Spellings