{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,5]],"date-time":"2026-06-05T14:02:06Z","timestamp":1780668126111,"version":"3.54.1"},"reference-count":33,"publisher":"Frontiers Media SA","license":[{"start":{"date-parts":[[2024,9,9]],"date-time":"2024-09-09T00:00:00Z","timestamp":1725840000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":["frontiersin.org"],"crossmark-restriction":true},"short-container-title":["Front. Big Data"],"abstract":"<jats:p>Data volume has been one of the fast-growing assets of most real-world applications. This increases the rate of human errors such as duplication of records, misspellings, and erroneous transpositions, among other data quality issues. Entity Resolution is an ETL process that aims to resolve data inconsistencies by ensuring entities are referring to the same real-world objects. One of the main challenges of most traditional Entity Resolution systems is ensuring their scalability to meet the rising data needs. This research aims to refactor a working proof-of-concept entity resolution system called the Data Washing Machine to be highly scalable using Apache Spark distributed data processing framework. We solve the single-threaded design problem of the legacy Data Washing Machine by using PySpark's Resilient Distributed Dataset and improve the Data Washing Machine design to use intrinsic metadata information from references. We prove that our systems achieve the same results as the legacy Data Washing Machine using 18 synthetically generated datasets. We also test the scalability of our system using a variety of real-world benchmark ER datasets from a few thousand to millions. Our experimental results show that our proposed system performs better than a MapReduce-based Data Washing Machine. We also compared our system with Famer and concluded that our system can find more clusters when given optimal starting parameters for clustering.<\/jats:p>","DOI":"10.3389\/fdata.2024.1446071","type":"journal-article","created":{"date-parts":[[2024,9,9]],"date-time":"2024-09-09T04:32:15Z","timestamp":1725856335000},"update-policy":"https:\/\/doi.org\/10.3389\/crossmark-policy","source":"Crossref","is-referenced-by-count":2,"title":["SparkDWM: a scalable design of a Data Washing Machine using Apache Spark"],"prefix":"10.3389","volume":"7","author":[{"given":"Nicholas Kofi Akortia","family":"Hagan","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"John R.","family":"Talburt","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1965","published-online":{"date-parts":[[2024,9,9]]},"reference":[{"key":"B1","first-page":"19","article-title":"A scalable, hybrid entity resolution process for unstandardized entity references","volume":"35","author":"Al Sarkhi","year":"2020","journal-title":"J. Comp. Sci. Colleg."},{"key":"B2","first-page":"64","article-title":"An analysis of the effect of stop words on the performance of the matrix comparator for entity resolution","volume":"34","author":"Al Sarkhi","year":"","journal-title":"J. Comp. Sci. Colleg."},{"key":"B3","first-page":"12","article-title":"Estimating the parameters for linking unstandardized references with the matrix comparator","volume":"10","author":"Al Sarkhi","year":"","journal-title":"J. Inform. Technol. Manag."},{"key":"B4","doi-asserted-by":"crossref","first-page":"106","DOI":"10.1007\/978-3-031-47451-4_8","article-title":"\u201cOptimal starting parameters for unsupervised data clustering and cleaning in the data washing machine,\u201d","author":"Anderson","year":"2023","journal-title":"Proceedings of the Future Technologies Conference (FTC) 2023, Volume 2"},{"key":"B5","first-page":"17","article-title":"\u201cA spark-based workflow for probabilistic record linkage of healthcare data,\u201d","author":"Pita","year":"2015","journal-title":"Edbt\/Icdt Workshops"},{"key":"B6","doi-asserted-by":"publisher","first-page":"1537","DOI":"10.1109\/TKDE.2011.127","article-title":"A survey of indexing techniques for scalable record linkage and deduplication","volume":"24","author":"Christen","year":"2012","journal-title":"IEEE Trans. Knowl. Data Eng."},{"key":"B7","doi-asserted-by":"publisher","first-page":"107","DOI":"10.1145\/1327452.1327492","article-title":"MapReduce: simplified data processing on large clusters","volume":"51","author":"Dean","year":"2008","journal-title":"Commun. ACM."},{"key":"B8","first-page":"602","article-title":"\u201cSparker: scaling entity resolution in spark,\u201d","author":"Gagliardelli","year":"2019","journal-title":"Advances in Database Technology-EDBT 2019, 22nd International Conference on Extending Database Technology, Lisbon, Portugal, March 26-29, Proceedings, Vol. 2019"},{"key":"B9","doi-asserted-by":"publisher","first-page":"1296552","DOI":"10.3389\/fdata.2024.1296552","article-title":"A scalable MapReduce-based design of an unsupervised entity resolution system","volume":"7","author":"Hagan","year":"2024","journal-title":"Front. Big Data"},{"key":"B10","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/2064085.2064087","article-title":"\u201cLearning-based entity resolution with MapReduce,\u201d","volume-title":"Proceedings of the Third International Workshop on Cloud Data Management","author":"Kolb","year":"2011"},{"key":"B11","doi-asserted-by":"publisher","first-page":"107","DOI":"10.1007\/s13222-014-0154-1","article-title":"Iterative computation of connected graph components with MapReduce","volume":"14","author":"Kolb","year":"2014","journal-title":"Datenbank-Spektrum"},{"key":"B12","doi-asserted-by":"publisher","first-page":"1878","DOI":"10.14778\/2367502.2367527","article-title":"Dedoop: efficient deduplication with Hadoop","volume":"5","author":"Kolb","year":"2012","journal-title":"Proc. VLDB Endowm."},{"key":"B13","doi-asserted-by":"publisher","first-page":"484","DOI":"10.14778\/1920841.1920904","article-title":"Evaluation of entity resolution approaches on real-world match problems","volume":"3","author":"K\u00f6pcke","year":"2010","journal-title":"Proc. VLDB Endowm."},{"key":"B14","doi-asserted-by":"crossref","first-page":"1087","DOI":"10.1109\/CSCI46756.2018.00211","article-title":"\u201cScoring matrix for unstandardized data in entity resolution,\u201d","volume-title":"2018 International Conference on Computational Science and Computational Intelligence (CSCI)","author":"Li","year":"2018"},{"key":"B15","unstructured":"Distributed holistic clustering on linked data'\n            NentwigM.\n            Gro\u00dfA.\n            M\u00f6llerM.\n            RahmE.\n          arXiv2017"},{"key":"B16","article-title":"\u201cKnowledge graph completion with FAMER,\u201d","author":"Obraczka","year":"2019","journal-title":"Proc. DI2KG"},{"key":"B17","doi-asserted-by":"publisher","first-page":"312","DOI":"10.14778\/2856318.2856326","article-title":"Schema-agnostic vs schema-based configurations for blocking methods on homogeneous data","volume":"9","author":"Papadakis","year":"2015","journal-title":"Proc. VLDB Endowm."},{"key":"B18","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1145\/3377455","article-title":"Blocking and filtering techniques for entity resolution: a survey","volume":"53","author":"Papadakis","year":"2020","journal-title":"ACM Comp. Surv."},{"key":"B19","doi-asserted-by":"crossref","first-page":"278","DOI":"10.1007\/978-3-319-66917-5_19","article-title":"\u201cComparative evaluation of distributed clustering schemes for multi-source entity resolution,\u201d","volume-title":"Advances in Databases and Information Systems: 21st European Conference, ADBIS 2017","author":"Saeedi","year":"2017"},{"key":"B20","doi-asserted-by":"crossref","DOI":"10.1007\/978-3-642-33460-3_35","article-title":"\u201cCC-MR - finding connected components in huge graphs with MapReduce,\u201d","volume-title":"Proceedings of the 2012th European Conference on Machine Learning and Knowledge Discovery in Databases - Volume Part I","author":"Seidl","year":"2012"},{"key":"B21","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1109\/MSST.2010.5496972","article-title":"\u201cThe hadoop distributed file system,\u201d","volume-title":"2010 IEEE 26th Symposium on Mass Storage Systems and Technologies (MSST)","author":"Shvachko","year":"2010"},{"key":"B22","doi-asserted-by":"publisher","first-page":"1173","DOI":"10.14778\/2994509.2994533","article-title":"BLAST: a loosely schema-aware meta-blocking approach for entity resolution","volume":"9","author":"Simonini","year":"2016","journal-title":"Proc. VLDB Endowm."},{"key":"B23","doi-asserted-by":"crossref","first-page":"860","DOI":"10.1109\/HPCS.2018.00138","article-title":"\u201cEnhancing loosely schema-aware entity resolution with user interaction,\u201d","volume-title":"2018 International Conference on High Performance Computing & Simulation (HPCS)","author":"Simonini","year":"2018"},{"key":"B24","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-642-36257-6_11","article-title":"\u201cA practical guide to entity resolution with OYSTER,\u201d","author":"Talburt","year":"2013","journal-title":"Handbook of Data Quality: Research and Practice"},{"key":"B25","doi-asserted-by":"publisher","first-page":"12","DOI":"10.14569\/IJACSA.2020.0111279","article-title":"An iterative, self-assessing entity resolution system: first steps toward a data washing machine","volume":"11","author":"Talburt","year":"2020","journal-title":"Int. J. Adv. Comp. Sci. Appl."},{"key":"B26","doi-asserted-by":"publisher","first-page":"1148331","DOI":"10.3389\/fdata.2023.1148331","article-title":"Editorial: automated data curation and data governance automation","volume":"6","author":"Talburt","year":"2023","journal-title":"Front. Big Data"},{"key":"B27","doi-asserted-by":"publisher","first-page":"295","DOI":"10.1007\/978-3-030-03643-0_14","article-title":"\u201cEvaluating and improving data fusion accuracy,\u201d","author":"Talburt","year":"2019","journal-title":"Information Quality in Information Fusion and Decision Making"},{"key":"B28","volume-title":"Entity Information Life Cycle for Big Data: Master Data Management and Information Integration. 1st edn","author":"Talburt","year":"2015"},{"key":"B29","first-page":"91","article-title":"\u201cSOG: a synthetic occupancy generator to support entity resolution instruction and research,\u201d","author":"Talburt","year":"2009","journal-title":"Proceedings of 14th International Conference on Information Quality (ICIQ 2009)"},{"key":"B30","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/2523616.2523633","article-title":"\u201cApache Hadoop YARN: yet another resource negotiator,\u201d","volume-title":"Proceedings of the 4th Annual Symposium on Cloud Computing. SOCC '13: ACM Symposium on Cloud Computing","author":"Vavilapalli","year":"2013"},{"key":"B31","first-page":"551","article-title":"\u201cParallel duplicate detection in adverse drug reaction databases with spark,\u201d","author":"Wang","year":"2016","journal-title":"EDBT"},{"key":"B32","first-page":"15","article-title":"\u201cResilient distributed datasets: a fault-tolerant abstraction for in-memory cluster computing,\u201d","author":"Zaharia","year":"2012","journal-title":"9th USENIX Symposium on Networked Systems Design and Implementation (NSDI 12)"},{"key":"B33","doi-asserted-by":"publisher","first-page":"56","DOI":"10.1145\/2934664","article-title":"Apache Spark: a unified engine for big data processing","volume":"59","author":"Zaharia","year":"2016","journal-title":"Commun. ACM"}],"container-title":["Frontiers in Big Data"],"original-title":[],"link":[{"URL":"https:\/\/www.frontiersin.org\/articles\/10.3389\/fdata.2024.1446071\/full","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2024,9,9]],"date-time":"2024-09-09T04:32:33Z","timestamp":1725856353000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.frontiersin.org\/articles\/10.3389\/fdata.2024.1446071\/full"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,9,9]]},"references-count":33,"alternative-id":["10.3389\/fdata.2024.1446071"],"URL":"https:\/\/doi.org\/10.3389\/fdata.2024.1446071","relation":{},"ISSN":["2624-909X"],"issn-type":[{"value":"2624-909X","type":"electronic"}],"subject":[],"published":{"date-parts":[[2024,9,9]]},"article-number":"1446071"}}