{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,2,13]],"date-time":"2026-02-13T23:17:50Z","timestamp":1771024670547,"version":"3.50.1"},"reference-count":35,"publisher":"Association for Computing Machinery (ACM)","issue":"5","content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Proc. VLDB Endow."],"published-print":{"date-parts":[[2020,1]]},"abstract":"<jats:p>Duplicate detection is an integral part of data cleaning and serves to identify multiple representations of same real-world entities in (relational) datasets. Existing duplicate detection approaches are effective, but they are also hard to parameterize or require a lot of pre-labeled training data. Both parameterization and pre-labeling are at least domain-specific if not dataset-specific, which is a problem if a new dataset needs to be cleaned.<\/jats:p>\n          <jats:p>\n            For this reason, we propose a novel, rule-based and fully automatic duplicate detection approach that is based on\n            <jats:italic>matching dependencies (MDs).<\/jats:italic>\n            Our system uses automatically discovered MDs, various dataset features, and known gold standards to train a model that selects MDs as duplicate detection rules. Once trained, the model can select useful MDs for duplicate detection on any new dataset. To increase the generally low recall of MD-based data cleaning approaches, we propose an additional boosting step. Our experiments show that this approach reaches up to 94% F-measure and 100% precision on our evaluation datasets, which are good numbers considering that the system does not require domain or target data-specific configuration.\n          <\/jats:p>","DOI":"10.14778\/3377369.3377379","type":"journal-article","created":{"date-parts":[[2020,2,19]],"date-time":"2020-02-19T18:58:53Z","timestamp":1582138733000},"page":"712-725","source":"Crossref","is-referenced-by-count":23,"title":["MDedup"],"prefix":"10.14778","volume":"13","author":[{"given":"loannis","family":"Koumarelas","sequence":"first","affiliation":[{"name":"University of Potsdam, Germany"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Thorsten","family":"Papenbrock","sequence":"additional","affiliation":[{"name":"University of Potsdam, Germany"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Felix","family":"Naumann","sequence":"additional","affiliation":[{"name":"University of Potsdam, Germany"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2020,2,19]]},"reference":[{"key":"e_1_2_1_1_1","volume-title":"Entity Resolution, and Duplicate Detection","author":"Christen Peter","year":"2012","unstructured":"Peter Christen . Data Matching: Concepts and Techniques for Record Linkage , Entity Resolution, and Duplicate Detection . Springer-Verlag Berlin Heidelberg , 2012 . Peter Christen. Data Matching: Concepts and Techniques for Record Linkage, Entity Resolution, and Duplicate Detection. Springer-Verlag Berlin Heidelberg, 2012."},{"key":"e_1_2_1_2_1","doi-asserted-by":"publisher","DOI":"10.5555\/1191547.1191739"},{"key":"e_1_2_1_3_1","doi-asserted-by":"publisher","DOI":"10.1145\/956750.956759"},{"issue":"11","key":"e_1_2_1_4_1","first-page":"1454","article-title":"Distributed representations of tuples for entity resolution","volume":"11","author":"Ebraheem Muhammad","year":"2018","unstructured":"Muhammad Ebraheem , Saravanan Thirumuruganathan , Shafiq Joty , Mourad Ouzzani , and Nan Tang . Distributed representations of tuples for entity resolution . PVLDB , 11 ( 11 ): 1454 -- 1467 , 2018 . Muhammad Ebraheem, Saravanan Thirumuruganathan, Shafiq Joty, Mourad Ouzzani, and Nan Tang. Distributed representations of tuples for entity resolution. PVLDB, 11(11):1454--1467, 2018.","journal-title":"PVLDB"},{"key":"e_1_2_1_5_1","doi-asserted-by":"publisher","DOI":"10.1145\/3183713.3196926"},{"key":"e_1_2_1_6_1","doi-asserted-by":"publisher","DOI":"10.14778\/2021017.2021020"},{"key":"e_1_2_1_7_1","doi-asserted-by":"publisher","DOI":"10.14778\/3149193.3149199"},{"issue":"3","key":"e_1_2_1_8_1","first-page":"278","article-title":"semantics and query answering","volume":"6","author":"Gardezi Jaffer","year":"2012","unstructured":"Jaffer Gardezi , Leopoldo Bertossi , and Iluju Kiringa . Matching dependencies : semantics and query answering . Frontiers of Computer Science , 6 ( 3 ): 278 -- 292 , 2012 . Jaffer Gardezi, Leopoldo Bertossi, and Iluju Kiringa. Matching dependencies: semantics and query answering. Frontiers of Computer Science, 6(3):278--292, 2012.","journal-title":"Frontiers of Computer Science"},{"key":"e_1_2_1_9_1","first-page":"380","volume-title":"Proceedings of the International Conference on Principles of Knowledge Representation and Reasoning (KR)","author":"Bahmani Zeinab","year":"2012","unstructured":"Zeinab Bahmani , Leopoldo E Bertossi , Solmaz Kolahi , and Laks VS Lakshmanan . Declarative entity resolution via matching dependencies and answer set programs . In Proceedings of the International Conference on Principles of Knowledge Representation and Reasoning (KR) , pages 380 -- 390 , 2012 . Zeinab Bahmani, Leopoldo E Bertossi, Solmaz Kolahi, and Laks VS Lakshmanan. Declarative entity resolution via matching dependencies and answer set programs. In Proceedings of the International Conference on Principles of Knowledge Representation and Reasoning (KR), pages 380--390, 2012."},{"key":"e_1_2_1_10_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.datak.2009.10.003"},{"key":"e_1_2_1_11_1","first-page":"17","volume-title":"Proceedings of the Australasian Workshop on Health Data and Knowledge Management (HDKM)","author":"Christen Peter","year":"2008","unstructured":"Peter Christen . Febrl : a freely available record linkage system with a graphical user interface . In Proceedings of the Australasian Workshop on Health Data and Knowledge Management (HDKM) , pages 17 -- 25 , 2008 . Peter Christen. Febrl: a freely available record linkage system with a graphical user interface. In Proceedings of the Australasian Workshop on Health Data and Knowledge Management (HDKM), pages 17--25, 2008."},{"key":"e_1_2_1_12_1","doi-asserted-by":"publisher","DOI":"10.14778\/3229863.3236232"},{"key":"e_1_2_1_13_1","first-page":"590","volume-title":"Proceedings of the International Conference on Management of Data (SIGMOD)","author":"Galhardas Helena","year":"2000","unstructured":"Helena Galhardas , Daniela Florescu , Dennis Shasha , and Eric Simon . Ajax : An extensible data cleaning tool . In Proceedings of the International Conference on Management of Data (SIGMOD) , page 590 , 2000 . Helena Galhardas, Daniela Florescu, Dennis Shasha, and Eric Simon. Ajax: An extensible data cleaning tool. In Proceedings of the International Conference on Management of Data (SIGMOD), page 590, 2000."},{"key":"e_1_2_1_14_1","volume-title":"Steven Euijong Whang, and Jennifer Widom. Swoosh: a generic approach to entity resolution. VLDB Journal (VLDBJ), 18(1):255--276","author":"Benjelloun Omar","year":"2009","unstructured":"Omar Benjelloun , Hector Garcia-Molina , David Menestrina , Qi Su , Steven Euijong Whang, and Jennifer Widom. Swoosh: a generic approach to entity resolution. VLDB Journal (VLDBJ), 18(1):255--276 , 2009 . Omar Benjelloun, Hector Garcia-Molina, David Menestrina, Qi Su, Steven Euijong Whang, and Jennifer Widom. Swoosh: a generic approach to entity resolution. VLDB Journal (VLDBJ), 18(1):255--276, 2009."},{"key":"e_1_2_1_15_1","doi-asserted-by":"publisher","DOI":"10.1109\/TKDE.2013.54"},{"key":"e_1_2_1_16_1","first-page":"136","volume-title":"Journal of Data Semantics (JoDS)","author":"Lehti Patrick","year":"2006","unstructured":"Patrick Lehti and Peter Fankhauser . Unsupervised duplicate detection using sample non-duplicates . In Journal of Data Semantics (JoDS) , pages 136 -- 164 . Springer , 2006 . Patrick Lehti and Peter Fankhauser. Unsupervised duplicate detection using sample non-duplicates. In Journal of Data Semantics (JoDS), pages 136--164. Springer, 2006."},{"key":"e_1_2_1_17_1","doi-asserted-by":"publisher","DOI":"10.1145\/1401890.1401913"},{"key":"e_1_2_1_18_1","doi-asserted-by":"publisher","DOI":"10.1016\/S0020-0255(00)00070-0"},{"key":"e_1_2_1_19_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-23540-0_27"},{"key":"e_1_2_1_20_1","doi-asserted-by":"publisher","DOI":"10.14778\/3157794.3157797"},{"key":"e_1_2_1_21_1","doi-asserted-by":"publisher","DOI":"10.1145\/2396761.2398606"},{"key":"e_1_2_1_22_1","doi-asserted-by":"publisher","DOI":"10.1109\/TKDE.2015.2472010"},{"key":"e_1_2_1_23_1","doi-asserted-by":"publisher","DOI":"10.1145\/1376916.1376940"},{"key":"e_1_2_1_24_1","doi-asserted-by":"publisher","DOI":"10.1145\/1645953.1646135"},{"key":"e_1_2_1_25_1","volume-title":"Efficient discovery of similarity constraints for matching dependencies. Data and Knowledge Engineering (DKE), 87:146--166","author":"Song Shaoxu","year":"2013","unstructured":"Shaoxu Song and Lei Chen . Efficient discovery of similarity constraints for matching dependencies. Data and Knowledge Engineering (DKE), 87:146--166 , 2013 . Shaoxu Song and Lei Chen. Efficient discovery of similarity constraints for matching dependencies. Data and Knowledge Engineering (DKE), 87:146--166, 2013."},{"key":"e_1_2_1_26_1","volume-title":"https:\/\/github.com\/HPI-Information-Systems\/metanome-algorithms. [Online","author":"Metanome","year":"2020","unstructured":"Metanome algorithm repository. https:\/\/github.com\/HPI-Information-Systems\/metanome-algorithms. [Online ; accessed 2- January - 2020 ]. Metanome algorithm repository. https:\/\/github.com\/HPI-Information-Systems\/metanome-algorithms. [Online; accessed 2-January-2020]."},{"issue":"8","key":"e_1_2_1_27_1","first-page":"707","article-title":"Binary codes capable of correcting deletions, insertions, and reversals","volume":"10","author":"Levenshtein Vladimir I.","year":"1966","unstructured":"Vladimir I. Levenshtein . Binary codes capable of correcting deletions, insertions, and reversals . Soviet Physics Doklady , 10 ( 8 ): 707 -- 710 , 1966 . Vladimir I. Levenshtein. Binary codes capable of correcting deletions, insertions, and reversals. Soviet Physics Doklady, 10(8):707--710, 1966.","journal-title":"Soviet Physics Doklady"},{"key":"e_1_2_1_28_1","first-page":"241","article-title":"Distribution de la flore alpine dans le bassin des Dranses et dans quelques r\u00e9gions voisines","volume":"37","author":"Jaccard Paul","year":"1901","unstructured":"Paul Jaccard . Distribution de la flore alpine dans le bassin des Dranses et dans quelques r\u00e9gions voisines . Bulletin de la Soci\u00e9t\u00e9 Vaudoise des Sciences Naturelles , 37 : 241 -- 272 , 1901 . Paul Jaccard. Distribution de la flore alpine dans le bassin des Dranses et dans quelques r\u00e9gions voisines. Bulletin de la Soci\u00e9t\u00e9 Vaudoise des Sciences Naturelles, 37:241--272, 1901.","journal-title":"Bulletin de la Soci\u00e9t\u00e9 Vaudoise des Sciences Naturelles"},{"key":"e_1_2_1_29_1","first-page":"289","volume-title":"International Conference on Machine Learning (ICML)","author":"Nan Ye","year":"2012","unstructured":"Ye Nan , Kian M. Chai , Wee S. Lee , and Hai L. Chieu . Optimizing f-measure: A tale of two approaches. In John Langford and Joelle Pineau, editors , International Conference on Machine Learning (ICML) , pages 289 -- 296 , 2012 . Ye Nan, Kian M. Chai, Wee S. Lee, and Hai L. Chieu. Optimizing f-measure: A tale of two approaches. In John Langford and Joelle Pineau, editors, International Conference on Machine Learning (ICML), pages 289--296, 2012."},{"key":"e_1_2_1_30_1","doi-asserted-by":"publisher","DOI":"10.3115\/1073012.1073017"},{"key":"e_1_2_1_31_1","volume-title":"The elements of statistical learning","author":"Friedman Jerome","year":"2001","unstructured":"Jerome Friedman , Trevor Hastie , and Robert Tibshirani . The elements of statistical learning . Springer-Verlag Berlin Heidelberg , 1 st edition, 2001 . Jerome Friedman, Trevor Hastie, and Robert Tibshirani. The elements of statistical learning. Springer-Verlag Berlin Heidelberg, 1st edition, 2001.","edition":"1"},{"key":"e_1_2_1_32_1","volume-title":"Logistic regression, adaboost and bregman distances. Machine Learning, 48(1--3):253--285","author":"Collins Michael","year":"2002","unstructured":"Michael Collins , Robert E Schapire , and Yoram Singer . Logistic regression, adaboost and bregman distances. Machine Learning, 48(1--3):253--285 , 2002 . Michael Collins, Robert E Schapire, and Yoram Singer. Logistic regression, adaboost and bregman distances. Machine Learning, 48(1--3):253--285, 2002."},{"key":"e_1_2_1_33_1","volume-title":"https:\/\/haifengl.github.io\/smile\/quickstart.html. [Online","author":"Smile","year":"2019","unstructured":"Smile - statistical machine intelligence and learning engine. https:\/\/haifengl.github.io\/smile\/quickstart.html. [Online ; accessed 1- July - 2019 ]. Smile - statistical machine intelligence and learning engine. https:\/\/haifengl.github.io\/smile\/quickstart.html. [Online; accessed 1-July-2019]."},{"key":"e_1_2_1_34_1","volume-title":"A practical guide to support vector classification","author":"Hsu Chih-Wei","year":"2003","unstructured":"Chih-Wei Hsu , Chih-Chung Chang , and Chih-Jen Lin . A practical guide to support vector classification . National Taiwan University , Taipei , 2003 . Chih-Wei Hsu, Chih-Chung Chang, and Chih-Jen Lin. A practical guide to support vector classification. National Taiwan University, Taipei, 2003."},{"key":"e_1_2_1_35_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-02300-7_3"}],"container-title":["Proceedings of the VLDB Endowment"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.14778\/3377369.3377379","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2022,12,28]],"date-time":"2022-12-28T09:33:50Z","timestamp":1672220030000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.14778\/3377369.3377379"}},"subtitle":["duplicate detection with matching dependencies"],"short-title":[],"issued":{"date-parts":[[2020,1]]},"references-count":35,"journal-issue":{"issue":"5","published-print":{"date-parts":[[2020,1]]}},"alternative-id":["10.14778\/3377369.3377379"],"URL":"https:\/\/doi.org\/10.14778\/3377369.3377379","relation":{},"ISSN":["2150-8097"],"issn-type":[{"value":"2150-8097","type":"print"}],"subject":[],"published":{"date-parts":[[2020,1]]}}}