Data Dependencies Extended for Variety and Veracity: A Family Tree (extended Abstract)

Shaoxu Song,Fei Gao,Ruihong Huang,Chaokun Wang
DOI: https://doi.org/10.1109/icde55515.2023.00336
2023-01-01
Abstract:To address the variety and veracity issues of big data, data dependencies have been extended as data quality rules to adapt to various data types, ranging from (1) categorical data with equality relationships to (2) heterogeneous data with similarity relationships, and (3) numerical data with order relationships. In this survey, we briefly review the recent proposals on data dependencies categorized into the aforesaid types of data. In addition to (a) the concepts of these data dependency notations, we investigate (b) the extension relationships between data dependencies. It forms a family tree of extensions, mostly rooted in FDs. Moreover, we summarize (c) the discovery of dependencies from data, and (d) the applications of the extended data dependencies. Finally, we conclude with several directions of future studies on the emerging data.
What problem does this paper attempt to address?