{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,8,20]],"date-time":"2026-08-20T14:53:11Z","timestamp":1787237591028,"version":"build-2736575974"},"reference-count":13,"publisher":"Association for Computing Machinery (ACM)","issue":"9","content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Proc. VLDB Endow."],"published-print":{"date-parts":[[2014,5]]},"abstract":"<jats:p>\n                    Record linkage clusters records such that each cluster corresponds to a single distinct real-world entity. It is a crucial step in data cleaning and data integration. In the big data era, the\n                    <jats:italic>velocity<\/jats:italic>\n                    of data updates is often high, quickly making previous linkage results obsolete. This paper presents an end-to-end framework that can incrementally and efficiently update linkage results when data updates arrive. Our algorithms not only allow merging records in the updates with existing clusters, but also allow leveraging new evidence from the updates to fix previous linkage errors. Experimental results on three real and synthetic data sets show that our algorithms can significantly reduce linkage time without sacrificing linkage quality.\n                  <\/jats:p>","DOI":"10.14778\/2732939.2732943","type":"journal-article","created":{"date-parts":[[2015,5,12]],"date-time":"2015-05-12T11:37:52Z","timestamp":1431430672000},"page":"697-708","source":"Crossref","is-referenced-by-count":80,"title":["Incremental record linkage"],"prefix":"10.14778","volume":"7","author":[{"given":"Anja","family":"Gruenheid","sequence":"first","affiliation":[{"name":"ETH Zurich"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Xin Luna","family":"Dong","sequence":"additional","affiliation":[{"name":"Google Inc."}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Divesh","family":"Srivastava","sequence":"additional","affiliation":[{"name":"AT&amp;T Labs-Research"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2014,5]]},"reference":[{"key":"e_1_2_1_1_1","first-page":"238","volume-title":"Machine Learning","author":"Bansal N.","year":"2002","unstructured":"N. Bansal , A. Blum , and S. Chawla . Correlation clustering . In Machine Learning , pages 238 -- 247 , 2002 . N. Bansal, A. Blum, and S. Chawla. Correlation clustering. In Machine Learning, pages 238--247, 2002."},{"key":"e_1_2_1_2_1","volume-title":"ACM SIGKDD Workshop on Data Cleaning, Record Linkage, and Object Consolidation","author":"Baxter R.","year":"2003","unstructured":"R. Baxter , P. Christen , and T. Churches . A comparison of fast blocking methods for record linkage . In ACM SIGKDD Workshop on Data Cleaning, Record Linkage, and Object Consolidation , 2003 . R. Baxter, P. Christen, and T. Churches. A comparison of fast blocking methods for record linkage. In ACM SIGKDD Workshop on Data Cleaning, Record Linkage, and Object Consolidation, 2003."},{"key":"e_1_2_1_3_1","doi-asserted-by":"publisher","DOI":"10.1007\/s00778-008-0098-x"},{"key":"e_1_2_1_4_1","doi-asserted-by":"publisher","DOI":"10.1145\/956750.956759"},{"key":"e_1_2_1_5_1","doi-asserted-by":"publisher","DOI":"10.1137\/S0097539702418498"},{"key":"e_1_2_1_6_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.1979.4766909"},{"key":"e_1_2_1_7_1","volume-title":"Proceedings of the 18th International Conference on Information Quality (ICIQ)","author":"Draisbach U.","year":"2013","unstructured":"U. Draisbach and F. Naumann . On choosing thresholds for duplicate detection . In Proceedings of the 18th International Conference on Information Quality (ICIQ) , 2013 . U. Draisbach and F. Naumann. On choosing thresholds for duplicate detection. In Proceedings of the 18th International Conference on Information Quality (ICIQ), 2013."},{"key":"e_1_2_1_8_1","doi-asserted-by":"publisher","DOI":"10.14778\/2367502.2367564"},{"key":"e_1_2_1_9_1","doi-asserted-by":"publisher","DOI":"10.14778\/1920841.1920897"},{"key":"e_1_2_1_10_1","doi-asserted-by":"publisher","DOI":"10.14778\/1687627.1687771"},{"key":"e_1_2_1_11_1","first-page":"573","volume-title":"STACS","author":"Mathieu C.","year":"2010","unstructured":"C. Mathieu , O. Sankur , and W. Schudy . Online correlation clustering . In STACS , pages 573 -- 584 , 2010 . C. Mathieu, O. Sankur, and W. Schudy. Online correlation clustering. In STACS, pages 573--584, 2010."},{"key":"e_1_2_1_12_1","doi-asserted-by":"publisher","DOI":"10.14778\/1920841.1921004"},{"key":"e_1_2_1_13_1","doi-asserted-by":"publisher","DOI":"10.1007\/s00778-013-0315-0"}],"container-title":["Proceedings of the VLDB Endowment"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.14778\/2732939.2732943","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2022,12,28]],"date-time":"2022-12-28T05:21:30Z","timestamp":1672204890000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.14778\/2732939.2732943"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2014,5]]},"references-count":13,"journal-issue":{"issue":"9","published-print":{"date-parts":[[2014,5]]}},"alternative-id":["10.14778\/2732939.2732943"],"URL":"https:\/\/doi.org\/10.14778\/2732939.2732943","relation":{},"ISSN":["2150-8097"],"issn-type":[{"value":"2150-8097","type":"print"}],"subject":[],"published":{"date-parts":[[2014,5]]}}}