Holistic primary key and foreign key detection

Jiang, Lan; Naumann, Felix

doi:10.1007/s10844-019-00562-z

Treffer 4 von 9

Zurück zur Trefferliste

Holistic primary key and foreign key detection

Lan Jiang, Felix Naumann

Primary keys (PKs) and foreign keys (FKs) are important elements of relational schemata in various applications, such as query optimization and data integration. However, in many cases, these constraints are unknown or not documented. Detecting them manually is time-consuming and even infeasible in large-scale datasets. We study the problem of discovering primary keys and foreign keys automatically and propose an algorithm to detect both, namely Holistic Primary Key and Foreign Key Detection (HoPF). PKs and FKs are subsets of the sets of unique column combinations (UCCs) and inclusion dependencies (INDs), respectively, for which efficient discovery algorithms are known. Using score functions, our approach is able to effectively extract the true PKs and FKs from the vast sets of valid UCCs and INDs. Several pruning rules are employed to speed up the procedure. We evaluate precision and recall on three benchmarks and two real-world datasets. The results show that our method is able to retrieve on average 88% of all primary keys, and 91%Primary keys (PKs) and foreign keys (FKs) are important elements of relational schemata in various applications, such as query optimization and data integration. However, in many cases, these constraints are unknown or not documented. Detecting them manually is time-consuming and even infeasible in large-scale datasets. We study the problem of discovering primary keys and foreign keys automatically and propose an algorithm to detect both, namely Holistic Primary Key and Foreign Key Detection (HoPF). PKs and FKs are subsets of the sets of unique column combinations (UCCs) and inclusion dependencies (INDs), respectively, for which efficient discovery algorithms are known. Using score functions, our approach is able to effectively extract the true PKs and FKs from the vast sets of valid UCCs and INDs. Several pruning rules are employed to speed up the procedure. We evaluate precision and recall on three benchmarks and two real-world datasets. The results show that our method is able to retrieve on average 88% of all primary keys, and 91% of all foreign keys. We compare the performance of HoPF with two baseline approaches that both assume the existence of primary keys.…

Metadaten
Verfasserangaben:	Lan Jiang ORCiD GND, Felix Naumann ORCiD GND
DOI:	https://doi.org/10.1007/s10844-019-00562-z
ISSN:	0925-9902
ISSN:	1573-7675
Titel des übergeordneten Werks (Englisch):	Journal of intelligent information systems : JIIS
Verlag:	Springer
Verlagsort:	Dordrecht
Publikationstyp:	Wissenschaftlicher Artikel
Sprache:	Englisch
Datum der Erstveröffentlichung:	10.06.2019
Erscheinungsjahr:	2020
Datum der Freischaltung:	02.01.2023
Freies Schlagwort / Tag:	Data profiling application; Database; Foreign key; Primary key; management
Band:	54
Ausgabe:	3
Seitenanzahl:	23
Erste Seite:	439
Letzte Seite:	461
Organisationseinheiten:	Digital Engineering Fakultät / Hasso-Plattner-Institut für Digital Engineering GmbH
DDC-Klassifikation:	0 Informatik, Informationswissenschaft, allgemeine Werke / 00 Informatik, Wissen, Systeme / 000 Informatik, Informationswissenschaft, allgemeine Werke
Peer Review:	Referiert

Holistic primary key and foreign key detection

Metadaten exportieren

Weitere Dienste