Discovery of genuine functional dependencies from relational data with missing values
- Functional dependencies (FDs) play an important role in maintaining data quality. They can be used to enforce data consistency and to guide repairs over a database. In this work, we investigate the problem of missing values and its impact on FD discovery. When using existing FD discovery algorithms, some genuine FDs could not be detected precisely due to missing values or some non-genuine FDs can be discovered even though they are caused by missing values with a certain NULL semantics. We define a notion of genuineness and propose algorithms to compute the genuineness score of a discovered FD. This can be used to identify the genuine FDs among the set of all valid dependencies that hold on the data. We evaluate the quality of our method over various real-world and semi-synthetic datasets with extensive experiments. The results show that our method performs well for relatively large FD sets and is able to accurately capture genuine FDs.
Verfasserangaben: | Laure Berti-EquilleORCiD, Nazar HarmouchORCiD, Felix NaumannORCiDGND, Noel Novelli, Thirumuruganathan Saravanan |
---|---|
DOI: | https://doi.org/10.14778/3204028.3204032 |
ISSN: | 2150-8097 |
Titel des übergeordneten Werks (Englisch): | Proceedings of the VLDB Endowment |
Verlag: | Association for Computing Machinery |
Verlagsort: | New York |
Publikationstyp: | Wissenschaftlicher Artikel |
Sprache: | Englisch |
Datum der Erstveröffentlichung: | 01.04.2018 |
Erscheinungsjahr: | 2018 |
Datum der Freischaltung: | 15.12.2021 |
Band: | 11 |
Ausgabe: | 8 |
Seitenanzahl: | 13 |
Erste Seite: | 880 |
Letzte Seite: | 892 |
Organisationseinheiten: | Digital Engineering Fakultät / Hasso-Plattner-Institut für Digital Engineering GmbH |
DDC-Klassifikation: | 0 Informatik, Informationswissenschaft, allgemeine Werke / 00 Informatik, Wissen, Systeme |
Peer Review: | Referiert |
Publikationsweg: | Open Access / Green Open-Access |