{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,2,22]],"date-time":"2026-02-22T07:34:54Z","timestamp":1771745694820,"version":"3.50.1"},"reference-count":43,"publisher":"Association for Computing Machinery (ACM)","issue":"2","license":[{"start":{"date-parts":[[2013,6,1]],"date-time":"2013-06-01T00:00:00Z","timestamp":1370044800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["61003004 and 61272090"],"award-info":[{"award-number":["61003004 and 61272090"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100001457","name":"Media Development Authority - Singapore","doi-asserted-by":"publisher","award":["WBS:R-252-300-001-490"],"award-info":[{"award-number":["WBS:R-252-300-001-490"]}],"id":[{"id":"10.13039\/501100001457","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100004147","name":"Tsinghua University","doi-asserted-by":"publisher","award":["20111081073"],"award-info":[{"award-number":["20111081073"]}],"id":[{"id":"10.13039\/501100004147","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100002855","name":"Ministry of Science and Technology of the People's Republic of China","doi-asserted-by":"publisher","award":["2011CB302206"],"award-info":[{"award-number":["2011CB302206"]}],"id":[{"id":"10.13039\/501100002855","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Database Syst."],"published-print":{"date-parts":[[2013,6]]},"abstract":"<jats:p>\n            As an essential operation in data cleaning, the similarity join has attracted considerable attention from the database community. In this article, we study string similarity joins with edit-distance constraints, which find similar string pairs from two large sets of strings whose edit distance is within a given threshold. Existing algorithms are efficient either for short strings or for long strings, and there is no algorithm that can efficiently and adaptively support both short strings and long strings. To address this problem, we propose a new filter, called\n            <jats:italic>the segment filter<\/jats:italic>\n            . We partition a string into a set of segments and use the segments as a filter to find similar string pairs. We first create inverted indices for the segments. Then for each string, we select some of its substrings, identify the selected substrings from the inverted indices, and take strings on the inverted lists of the found substrings as candidates of this string. Finally, we verify the candidates to generate the final answer. We devise efficient techniques to select substrings and prove that our method can minimize the number of selected substrings. We develop novel pruning techniques to efficiently verify the candidates. We also extend our techniques to support normalized edit distance. Experimental results show that our algorithms are efficient for both short strings and long strings, and outperform state-of-the-art methods on real-world datasets.\n          <\/jats:p>","DOI":"10.1145\/2487259.2487261","type":"journal-article","created":{"date-parts":[[2013,7,2]],"date-time":"2013-07-02T14:32:49Z","timestamp":1372775569000},"page":"1-33","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":29,"title":["A partition-based method for string similarity joins with edit-distance constraints"],"prefix":"10.1145","volume":"38","author":[{"given":"Guoliang","family":"Li","sequence":"first","affiliation":[{"name":"Tsinghua University, China"}]},{"given":"Dong","family":"Deng","sequence":"additional","affiliation":[{"name":"Tsinghua University, China"}]},{"given":"Jianhua","family":"Feng","sequence":"additional","affiliation":[{"name":"Tsinghua University, China"}]}],"member":"320","published-online":{"date-parts":[[2013,7,4]]},"reference":[{"key":"e_1_2_2_1_1","doi-asserted-by":"publisher","DOI":"10.14778\/1453856.1453958"},{"key":"e_1_2_2_2_1","volume-title":"Proceedings of the International Conference on Very Large Databases. 918--929","author":"Arasu A.","unstructured":"Arasu , A. , Ganti , V. , and Kaushik , R . 2006. Efficient exact set-similarity joins . In Proceedings of the International Conference on Very Large Databases. 918--929 . Arasu, A., Ganti, V., and Kaushik, R. 2006. Efficient exact set-similarity joins. In Proceedings of the International Conference on Very Large Databases. 918--929."},{"key":"e_1_2_2_3_1","doi-asserted-by":"publisher","DOI":"10.1145\/1242572.1242591"},{"key":"e_1_2_2_4_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICDE.2009.32"},{"key":"e_1_2_2_5_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICDE.2011.5767856"},{"key":"e_1_2_2_6_1","doi-asserted-by":"publisher","DOI":"10.1145\/1376616.1376697"},{"key":"e_1_2_2_7_1","doi-asserted-by":"publisher","DOI":"10.1145\/872757.872796"},{"key":"e_1_2_2_8_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICDE.2006.9"},{"key":"e_1_2_2_9_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICDE.2012.29"},{"key":"e_1_2_2_10_1","volume-title":"Proceedings of the International Conference on Data Engineering.","author":"Deng D.","unstructured":"Deng , D. , Li , G. , Feng , J. , and Li , W . -S. 2013. Top-k string similarity search with edit-distance constraints . In Proceedings of the International Conference on Data Engineering. Deng, D., Li, G., Feng, J., and Li, W.-S. 2013. Top-k string similarity search with edit-distance constraints. In Proceedings of the International Conference on Data Engineering."},{"key":"e_1_2_2_11_1","doi-asserted-by":"publisher","DOI":"10.1007\/s00778-011-0252-8"},{"key":"e_1_2_2_12_1","volume-title":"Proceedings of the International Conference on Very Large Databases. 491--500","author":"Gravano L.","unstructured":"Gravano , L. , Ipeirotis , P. G. , Jagadish , H. V. , Koudas , N. , Muthukrishnan , S. , and Srivastava , D . 2001. Approximate string joins in a database (almost) for free . In Proceedings of the International Conference on Very Large Databases. 491--500 . Gravano, L., Ipeirotis, P. G., Jagadish, H. V., Koudas, N., Muthukrishnan, S., and Srivastava, D. 2001. Approximate string joins in a database (almost) for free. In Proceedings of the International Conference on Very Large Databases. 491--500."},{"key":"e_1_2_2_13_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICDE.2008.4497435"},{"key":"e_1_2_2_14_1","doi-asserted-by":"publisher","DOI":"10.1145\/1559845.1559891"},{"key":"e_1_2_2_15_1","doi-asserted-by":"publisher","DOI":"10.14778\/1687553.1687623"},{"key":"e_1_2_2_16_1","doi-asserted-by":"publisher","DOI":"10.14778\/1453856.1453883"},{"key":"e_1_2_2_17_1","doi-asserted-by":"publisher","DOI":"10.1145\/1366102.1366104"},{"key":"e_1_2_2_18_1","doi-asserted-by":"publisher","DOI":"10.1145\/1807167.1807204"},{"key":"e_1_2_2_19_1","doi-asserted-by":"publisher","DOI":"10.1007\/s00778-007-0061-2"},{"key":"e_1_2_2_20_1","volume-title":"Proceedings of the International Conference on Very Large Databases. 195--206","author":"Lee H.","unstructured":"Lee , H. , Ng , R. T. , and Shim , K . 2007. Extending q-grams to estimate selectivity of string matching with low edit distance . In Proceedings of the International Conference on Very Large Databases. 195--206 . Lee, H., Ng, R. T., and Shim, K. 2007. Extending q-grams to estimate selectivity of string matching with low edit distance. In Proceedings of the International Conference on Very Large Databases. 195--206."},{"key":"e_1_2_2_21_1","doi-asserted-by":"publisher","DOI":"10.14778\/1687627.1687702"},{"key":"e_1_2_2_22_1","doi-asserted-by":"publisher","DOI":"10.14778\/1978665.1978666"},{"key":"e_1_2_2_23_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICDE.2008.4497434"},{"key":"e_1_2_2_24_1","doi-asserted-by":"publisher","DOI":"10.1145\/1989323.1989379"},{"key":"e_1_2_2_25_1","doi-asserted-by":"publisher","DOI":"10.14778\/2078331.2078340"},{"key":"e_1_2_2_26_1","doi-asserted-by":"publisher","DOI":"10.1109\/TKDE.2011.148"},{"key":"e_1_2_2_27_1","doi-asserted-by":"publisher","DOI":"10.1007\/s00778-011-0218-x"},{"key":"e_1_2_2_28_1","doi-asserted-by":"publisher","DOI":"10.1145\/375360.375365"},{"key":"e_1_2_2_29_1","doi-asserted-by":"publisher","DOI":"10.1145\/1989323.1989431"},{"key":"e_1_2_2_30_1","doi-asserted-by":"publisher","DOI":"10.1145\/1007568.1007652"},{"key":"e_1_2_2_31_1","volume-title":"Proceedings of the International Conference on Data Engineering. 892--903","author":"Silva Y. N.","unstructured":"Silva , Y. N. , Aref , W. G. , and Ali , M. H . 2010. The similarity join database operator . In Proceedings of the International Conference on Data Engineering. 892--903 . Silva, Y. N., Aref, W. G., and Ali, M. H. 2010. The similarity join database operator. In Proceedings of the International Conference on Data Engineering. 892--903."},{"key":"e_1_2_2_32_1","volume-title":"Proceedings of the International Workshop on Web and Databases.","author":"Sun C.","unstructured":"Sun , C. and Naughton , J. F . 2011. The token distribution filter for approximate string membership . In Proceedings of the International Workshop on Web and Databases. Sun, C. and Naughton, J. F. 2011. The token distribution filter for approximate string membership. In Proceedings of the International Workshop on Web and Databases."},{"key":"e_1_2_2_33_1","doi-asserted-by":"publisher","DOI":"10.1016\/S0019-9958(85)80046-2"},{"key":"e_1_2_2_34_1","doi-asserted-by":"publisher","DOI":"10.1145\/1807167.1807222"},{"key":"e_1_2_2_35_1","doi-asserted-by":"publisher","DOI":"10.14778\/1920841.1920992"},{"key":"e_1_2_2_36_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICDE.2011.5767865"},{"key":"e_1_2_2_37_1","doi-asserted-by":"publisher","DOI":"10.1145\/2213836.2213847"},{"key":"e_1_2_2_38_1","doi-asserted-by":"publisher","DOI":"10.1145\/1559845.1559925"},{"key":"e_1_2_2_39_1","doi-asserted-by":"publisher","DOI":"10.14778\/1453856.1453957"},{"key":"e_1_2_2_40_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICDE.2009.111"},{"key":"e_1_2_2_41_1","doi-asserted-by":"publisher","DOI":"10.1145\/1367497.1367516"},{"key":"e_1_2_2_42_1","doi-asserted-by":"publisher","DOI":"10.1145\/1376616.1376655"},{"key":"e_1_2_2_43_1","doi-asserted-by":"publisher","DOI":"10.1145\/1807167.1807266"}],"container-title":["ACM Transactions on Database Systems"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/2487259.2487261","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/2487259.2487261","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T08:35:54Z","timestamp":1750235754000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/2487259.2487261"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2013,6]]},"references-count":43,"journal-issue":{"issue":"2","published-print":{"date-parts":[[2013,6]]}},"alternative-id":["10.1145\/2487259.2487261"],"URL":"https:\/\/doi.org\/10.1145\/2487259.2487261","relation":{},"ISSN":["0362-5915","1557-4644"],"issn-type":[{"value":"0362-5915","type":"print"},{"value":"1557-4644","type":"electronic"}],"subject":[],"published":{"date-parts":[[2013,6]]},"assertion":[{"value":"2012-06-01","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2013-02-01","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2013-07-04","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}