{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,27]],"date-time":"2026-06-27T22:51:02Z","timestamp":1782600662821,"version":"3.54.5"},"reference-count":29,"publisher":"Association for Computing Machinery (ACM)","issue":"4","content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Proc. VLDB Endow."],"published-print":{"date-parts":[[2015,12]]},"abstract":"<jats:p>Entity Resolution constitutes a core task for data integration that, due to its quadratic complexity, typically scales to large datasets through blocking methods. These can be configured in two ways. The schema-based configuration relies on schema information in order to select signatures of high distinctiveness and low noise, while the schema-agnostic one treats every token from all attribute values as a signature. The latter approach has significant potential, as it requires no fine-tuning by human experts and it applies to heterogeneous data. Yet, there is no systematic study on its relative performance with respect to the schema-based configuration. This work covers this gap by comparing analytically the two configurations in terms of effectiveness, time efficiency and scalability. We apply them to 9 established blocking methods and to 11 benchmarks of structured data. We provide valuable insights into the internal functionality of the blocking methods with the help of a novel taxonomy. Our studies reveal that the schema-agnostic configuration offers unsupervised and robust definition of blocking keys under versatile settings, trading a higher computational cost for a consistently higher recall than the schema-based one. It also enables the use of state-of-the-art blocking methods without schema knowledge.<\/jats:p>","DOI":"10.14778\/2856318.2856326","type":"journal-article","created":{"date-parts":[[2016,2,1]],"date-time":"2016-02-01T14:10:31Z","timestamp":1454335831000},"page":"312-323","source":"Crossref","is-referenced-by-count":50,"title":["Schema-agnostic vs schema-based configurations for blocking methods on homogeneous data"],"prefix":"10.14778","volume":"9","author":[{"given":"George","family":"Papadakis","sequence":"first","affiliation":[{"name":"University of Athens, Greece"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"George","family":"Alexiou","sequence":"additional","affiliation":[{"name":"IMIS, Research Center \"Athena\", Greece"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"George","family":"Papastefanatos","sequence":"additional","affiliation":[{"name":"IMIS, Research Center \"Athena\", Greece"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Georgia","family":"Koutrika","sequence":"additional","affiliation":[{"name":"HP Labs"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2015,12]]},"reference":[{"key":"e_1_2_1_1_1","unstructured":"Blocking framework: http:\/\/sourceforge.net\/projects\/erframework.  Blocking framework: http:\/\/sourceforge.net\/projects\/erframework."},{"key":"e_1_2_1_2_1","doi-asserted-by":"publisher","DOI":"10.5555\/1105926.1106227"},{"key":"e_1_2_1_3_1","first-page":"25","volume-title":"Workshop on Data Cleaning, Record Linkage and Object Consolidation","author":"Baxter R.","year":"2003","unstructured":"R. Baxter , P. Christen , and T. Churches . A comparison of fast blocking methods for record linkage . In Workshop on Data Cleaning, Record Linkage and Object Consolidation , pages 25 -- 27 , 2003 . R. Baxter, P. Christen, and T. Churches. A comparison of fast blocking methods for record linkage. In Workshop on Data Cleaning, Record Linkage and Object Consolidation, pages 25--27, 2003."},{"key":"e_1_2_1_4_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICDM.2006.13"},{"key":"e_1_2_1_5_1","doi-asserted-by":"publisher","DOI":"10.1145\/1401890.1402020"},{"key":"e_1_2_1_6_1","volume-title":"Data Matching. Data-centric systems and applications","author":"Christen P.","year":"2012","unstructured":"P. Christen . Data Matching. Data-centric systems and applications . Springer , 2012 . P. Christen. Data Matching. Data-centric systems and applications. Springer, 2012."},{"key":"e_1_2_1_7_1","doi-asserted-by":"publisher","DOI":"10.1109\/TKDE.2011.127"},{"key":"e_1_2_1_8_1","first-page":"51","volume-title":"Proceedings of the International Workshop on Quality in Databases (QDB)","author":"Draisbach U.","year":"2009","unstructured":"U. Draisbach and F. Naumann . A comparison and generalization of blocking and windowing algorithms for duplicate detection . In Proceedings of the International Workshop on Quality in Databases (QDB) , pages 51 -- 56 , 2009 . U. Draisbach and F. Naumann. A comparison and generalization of blocking and windowing algorithms for duplicate detection. In Proceedings of the International Workshop on Quality in Databases (QDB), pages 51--56, 2009."},{"key":"e_1_2_1_9_1","volume-title":"Proceedings of the International Workshop on Quality in Databases (QDB)","author":"Draisbach U.","year":"2010","unstructured":"U. Draisbach and F. Naumann . Dude: The duplicate detection toolkit . In Proceedings of the International Workshop on Quality in Databases (QDB) , 2010 . U. Draisbach and F. Naumann. Dude: The duplicate detection toolkit. In Proceedings of the International Workshop on Quality in Databases (QDB), 2010."},{"key":"e_1_2_1_10_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICDE.2012.20"},{"key":"e_1_2_1_11_1","doi-asserted-by":"publisher","DOI":"10.1109\/TKDE.2007.9"},{"key":"e_1_2_1_12_1","doi-asserted-by":"publisher","DOI":"10.1080\/01621459.1969.10501049"},{"key":"e_1_2_1_13_1","doi-asserted-by":"publisher","DOI":"10.14778\/2367502.2367564"},{"key":"e_1_2_1_14_1","doi-asserted-by":"publisher","DOI":"10.1145\/2487575.2506179"},{"key":"e_1_2_1_15_1","first-page":"491","volume-title":"VLDB","author":"Gravano L.","year":"2001","unstructured":"L. Gravano , P. Ipeirotis , H. Jagadish , N. Koudas , S. Muthukrishnan , and D. Srivastava . Approximate string joins in a database (almost) for free . In VLDB , pages 491 -- 500 , 2001 . L. Gravano, P. Ipeirotis, H. Jagadish, N. Koudas, S. Muthukrishnan, and D. Srivastava. Approximate string joins in a database (almost) for free. In VLDB, pages 491--500, 2001."},{"key":"e_1_2_1_16_1","doi-asserted-by":"publisher","DOI":"10.1145\/223784.223807"},{"key":"e_1_2_1_17_1","volume-title":"WebDB","author":"Isele R.","year":"2011","unstructured":"R. Isele , A. Jentzsch , and C. Bizer . Efficient multidimensional blocking for link discovery without losing recall . In WebDB , 2011 . R. Isele, A. Jentzsch, and C. Bizer. Efficient multidimensional blocking for link discovery without losing recall. In WebDB, 2011."},{"key":"e_1_2_1_18_1","doi-asserted-by":"publisher","DOI":"10.14778\/2732296.2732299"},{"key":"e_1_2_1_19_1","first-page":"137","volume-title":"DASFAA","author":"Jin L.","year":"2003","unstructured":"L. Jin , C. Li , and S. Mehrotra . Efficient record linkage in large data sets . In DASFAA , pages 137 -- 146 , 2003 . L. Jin, C. Li, and S. Mehrotra. Efficient record linkage in large data sets. In DASFAA, pages 137--146, 2003."},{"key":"e_1_2_1_20_1","doi-asserted-by":"publisher","DOI":"10.1145\/1142473.1142599"},{"key":"e_1_2_1_21_1","first-page":"342","volume-title":"CIDR","author":"Madhavan J.","year":"2007","unstructured":"J. Madhavan , S. Cohen , X. L. Dong , A. Y. Halevy , S. R. Jeffery , D. Ko , and C. Yu . Web-scale data integration: You can afford to pay as you go . In CIDR , pages 342 -- 350 , 2007 . J. Madhavan, S. Cohen, X. L. Dong, A. Y. Halevy, S. R. Jeffery, D. Ko, and C. Yu. Web-scale data integration: You can afford to pay as you go. In CIDR, pages 342--350, 2007."},{"key":"e_1_2_1_22_1","doi-asserted-by":"publisher","DOI":"10.1145\/347090.347123"},{"key":"e_1_2_1_23_1","first-page":"440","volume-title":"AAAI","author":"Michelson M.","year":"2006","unstructured":"M. Michelson and C. A. Knoblock . Learning blocking schemes for record linkage . In AAAI , pages 440 -- 445 , 2006 . M. Michelson and C. A. Knoblock. Learning blocking schemes for record linkage. In AAAI, pages 440--445, 2006."},{"key":"e_1_2_1_24_1","doi-asserted-by":"publisher","DOI":"10.5555\/2283696.2283783"},{"key":"e_1_2_1_25_1","doi-asserted-by":"publisher","DOI":"10.1145\/2124295.2124305"},{"key":"e_1_2_1_26_1","doi-asserted-by":"publisher","DOI":"10.1109\/TKDE.2012.150"},{"key":"e_1_2_1_27_1","doi-asserted-by":"publisher","DOI":"10.1109\/TKDE.2013.54"},{"key":"e_1_2_1_29_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-642-16518-4_1"},{"key":"e_1_2_1_30_1","doi-asserted-by":"publisher","DOI":"10.1145\/1559845.1559870"}],"container-title":["Proceedings of the VLDB Endowment"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.14778\/2856318.2856326","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2022,12,28]],"date-time":"2022-12-28T10:23:24Z","timestamp":1672223004000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.14778\/2856318.2856326"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2015,12]]},"references-count":29,"journal-issue":{"issue":"4","published-print":{"date-parts":[[2015,12]]}},"alternative-id":["10.14778\/2856318.2856326"],"URL":"https:\/\/doi.org\/10.14778\/2856318.2856326","relation":{},"ISSN":["2150-8097"],"issn-type":[{"value":"2150-8097","type":"print"}],"subject":[],"published":{"date-parts":[[2015,12]]}}}