{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,4,10]],"date-time":"2026-04-10T15:09:59Z","timestamp":1775833799735,"version":"3.50.1"},"reference-count":76,"publisher":"Association for Computing Machinery (ACM)","issue":"4","funder":[{"DOI":"10.13039\/100021130","name":"Bundesministerium f\u00fcr Wirtschaft und Klimaschutz","doi-asserted-by":"crossref","award":["BMWK.IIB5"],"award-info":[{"award-number":["BMWK.IIB5"]}],"id":[{"id":"10.13039\/100021130","id-type":"DOI","asserted-by":"crossref"}]},{"name":"Research Fund of Praezisions-LDS","award":["03EE3061F"],"award-info":[{"award-number":["03EE3061F"]}]},{"name":"LOEWE initiative (Hesse, Germany) within the emergenCITY center","award":["LOEWE\/1\/12\/519\/03\/05.001(0016)\/72"],"award-info":[{"award-number":["LOEWE\/1\/12\/519\/03\/05.001(0016)\/72"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Knowl. Discov. Data"],"published-print":{"date-parts":[[2026,5,31]]},"abstract":"<jats:p>\n                    Data profiling describes the activity of inferring structural metadata, such as functional dependencies, inclusion dependencies, and unique column combinations, from (relational) datasets. Because structural metadata is often not stored explicitly, data profiling plays a crucial role in various data management tasks, including data discovery, cleaning, integration, normalization, and querying. Due to the importance of structural metadata and, in particular, data dependencies, researchers have been actively exploring new types of metadata and efficient algorithms for their automatic discovery. In the past, however, each type of metadata has been considered mostly in isolation. This poses a serious challenge to many use cases that actually require specific\n                    <jats:italic toggle=\"yes\">combinations of data dependencies<\/jats:italic>\n                    because deriving these combinations from individually profiled metadata is as difficult as the initial metadata discovery.\n                  <\/jats:p>\n                  <jats:p>\n                    In this article, we investigate the interaction of data dependencies in (complex) combinations and define\n                    <jats:italic toggle=\"yes\">minimality<\/jats:italic>\n                    and\n                    <jats:italic toggle=\"yes\">completeness<\/jats:italic>\n                    as two essential properties that enable the automatic profiling of data dependency combinations. A notion of minimal dependency combinations and complete dependency combination result sets is a prerequisite for the (automatic) discovery of dependency combinations, because these properties enable search space pruning and effectively restrict the profiling to manageable and meaningful result sizes. Due to the enormous search space of dependency combinations, we also propose a\n                    <jats:italic toggle=\"yes\">minimality constraint<\/jats:italic>\n                    formalism as a novel search space pruning technique. This technique expresses the minimality of any data dependency combination in terms of already well-known, type-specific minimality constraints. Furthermore, we apply a practical, graph-based\n                    <jats:italic toggle=\"yes\">constraint inference<\/jats:italic>\n                    algorithm to automatically derive query-specific minimality constraints for any given metadata query. In an experimental evaluation, we assess the effectiveness of the derived minimality constraints and provide a first impression of the\n                    <jats:italic toggle=\"yes\">possibilities and challenges<\/jats:italic>\n                    that arise when profiling (complex) metadata patterns. Our study covers both theoretical and practical aspects for the profiling of data dependency combinations and is a necessary step towards the development of a holistic data profiling system that efficiently answers metadata pattern queries.\n                  <\/jats:p>","DOI":"10.1145\/3799992","type":"journal-article","created":{"date-parts":[[2026,3,4]],"date-time":"2026-03-04T12:38:11Z","timestamp":1772627891000},"page":"1-49","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":0,"title":["Profiling Minimal Data Dependency Combinations"],"prefix":"10.1145","volume":"20","author":[{"ORCID":"https:\/\/orcid.org\/0009-0001-8483-2212","authenticated-orcid":false,"given":"Marcian","family":"Seeger","sequence":"first","affiliation":[{"name":"University of Marburg, Marburg, Germany"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-6597-9809","authenticated-orcid":false,"given":"Sebastian","family":"Schmidl","sequence":"additional","affiliation":[{"name":"Hasso-Plattner-Institut fur Digital Engineering gGmbH, Potsdam, Germany"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-4019-8221","authenticated-orcid":false,"given":"Thorsten","family":"Papenbrock","sequence":"additional","affiliation":[{"name":"University of Marburg, Marburg, Germany"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2026,4,10]]},"reference":[{"key":"e_1_3_2_2_2","doi-asserted-by":"publisher","DOI":"10.1007\/s00778-015-0389-y"},{"key":"e_1_3_2_3_2","doi-asserted-by":"publisher","DOI":"10.2200\/S00878ED1V01Y201810DTM052"},{"key":"e_1_3_2_4_2","first-page":"580","volume-title":"Proceedings of the IFIP Congress","volume":"74","author":"Armstrong William Ward","year":"1974","unstructured":"William Ward Armstrong. 1974. Dependency structures of data base relationships. In Proceedings of the IFIP Congress, Vol. 74, 580\u2013583."},{"key":"e_1_3_2_5_2","doi-asserted-by":"publisher","DOI":"10.1145\/320064.320066"},{"key":"e_1_3_2_6_2","doi-asserted-by":"publisher","DOI":"10.1145\/509404.509414"},{"key":"e_1_3_2_7_2","doi-asserted-by":"crossref","unstructured":"Catriel Beeri and Peter Honeyman. 1981. Preserving functional dependencies. SIAM Journal on Computing 10 3 (1981) 647\u2013656.","DOI":"10.1137\/0210048"},{"key":"e_1_3_2_8_2","doi-asserted-by":"publisher","DOI":"10.5555\/646235.682563"},{"key":"e_1_3_2_9_2","doi-asserted-by":"publisher","DOI":"10.14778\/3407790.3407824"},{"key":"e_1_3_2_10_2","doi-asserted-by":"publisher","DOI":"10.14778\/3157794.3157800"},{"key":"e_1_3_2_11_2","doi-asserted-by":"publisher","DOI":"10.1145\/3639298"},{"key":"e_1_3_2_12_2","doi-asserted-by":"publisher","DOI":"10.1007\/s10618-011-0229-7"},{"key":"e_1_3_2_13_2","first-page":"243","volume-title":"Proceedings of the 33rd International Conference on Very Large Data Bases (VLDB)","volume":"7","author":"Bravo Loreto","year":"2007","unstructured":"Loreto Bravo, Wenfei, and Fan, Shuai Ma. 2007. Extending dependencies with conditions. In Proceedings of the 33rd International Conference on Very Large Data Bases (VLDB), Vol. 7. VLDB Endowment, 243\u2013254."},{"key":"e_1_3_2_14_2","doi-asserted-by":"publisher","DOI":"10.1145\/588111.588141"},{"key":"e_1_3_2_15_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.tcs.2016.11.004"},{"key":"e_1_3_2_16_2","unstructured":"E. F. Codd. 1971. Further Normalization of the Data Base Relational Model. IBM Research Report RJ909."},{"key":"e_1_3_2_17_2","first-page":"33","article-title":"Further normalization of the data base relational model","volume":"6","author":"Codd Edgar F.","year":"1972","unstructured":"Edgar F. Codd. 1972. Further normalization of the data base relational model. Data Base Systems 6 (1972), 33\u201364.","journal-title":"Data Base Systems"},{"key":"e_1_3_2_18_2","first-page":"1017","volume-title":"Information Processing: Proceedings of the IFIP Congress","author":"Codd E. F.","year":"1974","unstructured":"E. F. Codd. 1974. Recent investigations in relational data base systems. In Information Processing: Proceedings of the IFIP Congress, 1017\u20131021."},{"key":"e_1_3_2_19_2","doi-asserted-by":"publisher","DOI":"10.1109\/tpami.2004.75"},{"key":"e_1_3_2_20_2","doi-asserted-by":"publisher","DOI":"10.1145\/588011.588016"},{"key":"e_1_3_2_21_2","doi-asserted-by":"publisher","DOI":"10.1145\/78935.78937"},{"key":"e_1_3_2_22_2","doi-asserted-by":"publisher","DOI":"10.1016\/B978-0-444-88074-1.50011-1"},{"key":"e_1_3_2_23_2","unstructured":"Desbordante. 2022. Open-Source Data Profiling Tool. Retrieved from https:\/\/desbordante.unidata-platform.ru\/"},{"key":"e_1_3_2_24_2","doi-asserted-by":"publisher","DOI":"10.1145\/3357384.3357916"},{"key":"e_1_3_2_25_2","doi-asserted-by":"publisher","DOI":"10.5441\/002\/EDBT.2016.29"},{"key":"e_1_3_2_26_2","doi-asserted-by":"publisher","DOI":"10.1145\/320557.320571"},{"key":"e_1_3_2_27_2","doi-asserted-by":"publisher","DOI":"10.1145\/1739041.1739076"},{"key":"e_1_3_2_28_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-91458-9_21"},{"key":"e_1_3_2_29_2","doi-asserted-by":"publisher","DOI":"10.1145\/3186728.3164145"},{"key":"e_1_3_2_30_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICDE51399.2021.00045"},{"issue":"2","key":"e_1_3_2_31_2","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/2188349.2188355","article-title":"The implication problem of data dependencies over SQL table definitions: Axiomatic, algorithmic and logical characterizations","volume":"37","author":"Hartmann Sven","year":"2012","unstructured":"Sven Hartmann and Sebastian Link. 2012. The implication problem of data dependencies over SQL table definitions: Axiomatic, algorithmic and logical characterizations. ACM Transactions on Database Systems 37, 2 (2012), 1\u201340.","journal-title":"ACM Transactions on Database Systems"},{"key":"e_1_3_2_32_2","doi-asserted-by":"publisher","DOI":"10.14778\/2732240.2732248"},{"key":"e_1_3_2_33_2","volume-title":"Proceedings of the International Conference on Extending Database Technology (EDBT)","author":"Hellenberg Jan-Eric","year":"2025","unstructured":"Jan-Eric Hellenberg, Fabian Mahling, Lukas Laskowski, Felix Naumann, Matteo Paganelli, and Fabian Panse. 2025. PRISMA: A privacy-preserving schema matcher using functional dependencies. In Proceedings of the International Conference on Extending Database Technology (EDBT)."},{"key":"e_1_3_2_34_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.patcog.2008.08.001"},{"key":"e_1_3_2_35_2","doi-asserted-by":"publisher","DOI":"10.1093\/comjnl\/42.2.100"},{"key":"e_1_3_2_36_2","doi-asserted-by":"publisher","DOI":"10.1007\/s10844-019-00562-z"},{"key":"e_1_3_2_37_2","doi-asserted-by":"publisher","DOI":"10.1145\/3588929"},{"key":"e_1_3_2_38_2","unstructured":"Reza Karegar Parke Godfrey Lukasz Golab Mehdi Kargar Divesh Srivastava and Jaroslaw Szlichta. 2021. Efficient discovery of approximate order dependencies. arXiv:2101.02174. Retrieved from https:\/\/arxiv.org\/abs\/2101.02174"},{"key":"e_1_3_2_39_2","first-page":"1","volume-title":"Proceedings of the Conference on Innovative Data Systems Research (CIDR)","volume":"12","author":"Kossmann Jan","year":"2022","unstructured":"Jan Kossmann, Daniel Lindner, Felix Naumann, and Thorsten Papenbrock. 2022. Workload-driven, lazy discovery of data dependencies for query optimization. In Proceedings of the Conference on Innovative Data Systems Research (CIDR), Vol. 12. CIDR Conference, 1\u20137."},{"key":"e_1_3_2_40_2","doi-asserted-by":"publisher","DOI":"10.1007\/s00778-021-00676-3"},{"key":"e_1_3_2_41_2","doi-asserted-by":"publisher","DOI":"10.1145\/3132847.3133180"},{"key":"e_1_3_2_42_2","doi-asserted-by":"publisher","DOI":"10.5555\/645922.673490"},{"issue":"3","key":"e_1_3_2_43_2","doi-asserted-by":"crossref","first-page":"115","DOI":"10.1016\/S0020-0190(99)00095-2","article-title":"How to prevent interaction of functional and inclusion dependencies","volume":"71","author":"Levene Mark","year":"1999","unstructured":"Mark Levene and George Loizou. 1999. How to prevent interaction of functional and inclusion dependencies. Information Processing Letters 71, 3\u20134 (1999), 115\u2013125.","journal-title":"Information Processing Letters"},{"key":"e_1_3_2_44_2","doi-asserted-by":"publisher","DOI":"10.1109\/69.842267"},{"issue":"1","key":"e_1_3_2_45_2","first-page":"68","article-title":"Mining conditional functional dependency rules on big data","volume":"3","author":"Li Mingda","year":"2019","unstructured":"Mingda Li, Hongzhi Wang, and Jianzhong Li. 2019. Mining conditional functional dependency rules on big data. Big Data Mining and Analytics 3, 1 (2019), 68\u201384.","journal-title":"Big Data Mining and Analytics"},{"key":"e_1_3_2_46_2","unstructured":"Daniel Lindner Daniel Ritter and Felix Naumann. 2024. Enabling data dependency-based query optimization. arXiv:2406.06886. Retrieved from https:\/\/arxiv.org\/abs\/2406.06886"},{"key":"e_1_3_2_47_2","doi-asserted-by":"publisher","DOI":"10.1016\/0022-0000(78)90009-0"},{"key":"e_1_3_2_48_2","doi-asserted-by":"crossref","first-page":"53","DOI":"10.1007\/s10844-007-0048-x","article-title":"Unary and n-ary inclusion dependency discovery in relational databases","volume":"32","author":"De Marchi Fabien","year":"2009","unstructured":"Fabien De Marchi, St\u00e9phane Lopes, and Jean-Marc Petit. 2009. Unary and n-ary inclusion dependency discovery in relational databases. Journal of Intelligent Information Systems 32 (2009), 53\u201373.","journal-title":"Journal of Intelligent Information Systems"},{"key":"e_1_3_2_49_2","doi-asserted-by":"publisher","DOI":"10.1006\/jagm.1997.0898"},{"key":"e_1_3_2_50_2","doi-asserted-by":"publisher","DOI":"10.1023\/A:1009774406717"},{"key":"e_1_3_2_51_2","doi-asserted-by":"publisher","DOI":"10.5555\/645922.673500"},{"key":"e_1_3_2_52_2","unstructured":"MetaBrainz. 2022. MusicBrainz Database. Retrieved November 30 2022 from https:\/\/data.metabrainz.org\/pub\/musicbrainz\/data\/fullexport\/"},{"key":"e_1_3_2_53_2","unstructured":"Microsoft. 2023. AdventureWorks Sample Databases. Retrieved June 18 2023 from https:\/\/learn.microsoft.com\/en-us\/sql\/samples\/adventureworks-install-configure?view=sql-server-ver16"},{"key":"e_1_3_2_54_2","doi-asserted-by":"publisher","DOI":"10.1016\/S0019-9958(83)80002-3"},{"key":"e_1_3_2_55_2","doi-asserted-by":"publisher","DOI":"10.1145\/588058.588067"},{"key":"e_1_3_2_56_2","doi-asserted-by":"publisher","DOI":"10.1145\/1989323.1989458"},{"key":"e_1_3_2_57_2","doi-asserted-by":"publisher","DOI":"10.14778\/2824032.2824086"},{"key":"e_1_3_2_58_2","doi-asserted-by":"publisher","DOI":"10.14778\/2794367.2794377"},{"key":"e_1_3_2_59_2","doi-asserted-by":"publisher","DOI":"10.14778\/2752939.2752946"},{"key":"e_1_3_2_60_2","doi-asserted-by":"publisher","DOI":"10.1145\/2882903.2915203"},{"key":"e_1_3_2_61_2","first-page":"342","volume-title":"Proceedings of the International Conference on Extending Database Technology (EDBT)","volume":"17","author":"Papenbrock Thorsten","year":"2017","unstructured":"Thorsten Papenbrock and Felix Naumann. 2017. Data-driven schema normalization. In Proceedings of the International Conference on Extending Database Technology (EDBT), Vol. 17, 342\u2013353."},{"key":"e_1_3_2_62_2","first-page":"195","volume-title":"Proceedings of the Conference Database Systems for Business, Technology and Web (BTW)","author":"Papenbrock Thorsten","year":"2017","unstructured":"Thorsten Papenbrock and Felix Naumann. 2017. A hybrid approach for efficient unique column combination discovery. In Proceedings of the Conference Database Systems for Business, Technology and Web (BTW), 195\u2013204."},{"key":"e_1_3_2_63_2","doi-asserted-by":"publisher","DOI":"10.14778\/3503585.3503595"},{"key":"e_1_3_2_64_2","doi-asserted-by":"publisher","DOI":"10.1007\/s007780100057"},{"key":"e_1_3_2_65_2","first-page":"1","volume-title":"Proceedings of the ACM Workshop on the Web and Databases (WebDB)","volume":"12","author":"Rostin Alexandra","year":"2009","unstructured":"Alexandra Rostin, Oliver Albrecht, Jana Bauckmann, Felix Naumann, and Ulf Leser. 2009. A machine learning approach to foreign key discovery. In Proceedings of the ACM Workshop on the Web and Databases (WebDB), Vol. 12. ACM New York, NY, 1\u20136."},{"key":"e_1_3_2_66_2","doi-asserted-by":"publisher","DOI":"10.1145\/3392778"},{"key":"e_1_3_2_67_2","doi-asserted-by":"publisher","DOI":"10.1007\/s00778-021-00683-4"},{"key":"e_1_3_2_68_2","volume-title":"Proceedings of the Conference Database Systems for Business, Technology and Web (BTW)","author":"Seeger Marcian","year":"2025","unstructured":"Marcian Seeger and Thorsten Papenbrock. 2025. DPQL: Applications for holistic data profiling. In Proceedings of the Conference Database Systems for Business, Technology and Web (BTW)."},{"key":"e_1_3_2_69_2","doi-asserted-by":"publisher","DOI":"10.18420\/BTW2023-19"},{"key":"e_1_3_2_70_2","doi-asserted-by":"publisher","DOI":"10.14778\/3067421.3067422"},{"key":"e_1_3_2_71_2","doi-asserted-by":"publisher","DOI":"10.14778\/2350229.2350241"},{"key":"e_1_3_2_72_2","unstructured":"TPC. 2022. TPC-H Benchmark. Retrieved November 30 2022 from http:\/\/www.tpc.org\/tpch\/"},{"key":"e_1_3_2_73_2","unstructured":"TPC. 2023. TPC-E Benchmark. Retrieved June 1 2023 from http:\/\/www.tpc.org\/tpce\/"},{"key":"e_1_3_2_74_2","unstructured":"Viadotto. 2022. Make Your Data Profitable with Our Next-Gen Data Profiling Tools. Retrieved from https:\/\/www.viadotto.tech\/"},{"key":"e_1_3_2_75_2","doi-asserted-by":"publisher","DOI":"10.1145\/3070647"},{"key":"e_1_3_2_76_2","doi-asserted-by":"publisher","DOI":"10.14778\/3358701.3358703"},{"issue":"3","key":"e_1_3_2_77_2","doi-asserted-by":"crossref","first-page":"29","DOI":"10.1007\/s00778-025-00910-2","article-title":"Mixed covers: Optimizing updates and queries using minimal keys and functional dependencies","volume":"34","author":"Zhang Zhuoxing","year":"2025","unstructured":"Zhuoxing Zhang and Sebastian Link. 2025. Mixed covers: Optimizing updates and queries using minimal keys and functional dependencies. The VLDB Journal 34, 3 (2025), 29.","journal-title":"The VLDB Journal"}],"container-title":["ACM Transactions on Knowledge Discovery from Data"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3799992","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,4,10]],"date-time":"2026-04-10T14:41:35Z","timestamp":1775832095000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3799992"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,4,10]]},"references-count":76,"journal-issue":{"issue":"4","published-print":{"date-parts":[[2026,5,31]]}},"alternative-id":["10.1145\/3799992"],"URL":"https:\/\/doi.org\/10.1145\/3799992","relation":{},"ISSN":["1556-4681","1556-472X"],"issn-type":[{"value":"1556-4681","type":"print"},{"value":"1556-472X","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,4,10]]},"assertion":[{"value":"2025-09-05","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2026-02-23","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2026-04-10","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}