{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,10]],"date-time":"2026-06-10T14:45:15Z","timestamp":1781102715368,"version":"3.54.1"},"reference-count":25,"publisher":"IGI Global Scientific Publishing","issue":"4","content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":[],"published-print":{"date-parts":[[2011,10,1]]},"abstract":"<p>The quality of real world data that is being fed into a data warehouse is a major concern of today. As the data comes from a variety of sources before loading the data in the data warehouse, it must be checked for errors and anomalies. There may be exact duplicate records or approximate duplicate records in the source data. The presence of incorrect or inconsistent data can significantly distort the results of analyses, often negating the potential benefits of information-driven approaches. This paper addresses issues related to detection and correction of such duplicate records. Also, it analyzes data quality and various factors that degrade it. A brief analysis of existing work is discussed, pointing out its major limitations. Thus, a new framework is proposed that is an improvement over the existing technique.<\/p>","DOI":"10.4018\/ijkbo.2011100104","type":"journal-article","created":{"date-parts":[[2011,10,19]],"date-time":"2011-10-19T12:38:00Z","timestamp":1319027880000},"page":"56-71","source":"Crossref","is-referenced-by-count":8,"title":["An Efficient Algorithm for Data Cleaning"],"prefix":"10.4018","volume":"1","author":[{"given":"Payal","family":"Pahwa","sequence":"first","affiliation":[{"name":"Guru Gobind Singh IndraPrastha University, India"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Rajiv","family":"Arora","sequence":"additional","affiliation":[{"name":"Guru Gobind Singh IndraPrastha University, India"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Garima","family":"Thakur","sequence":"additional","affiliation":[{"name":"Guru Gobind Singh IndraPrastha University, India"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"2432","reference":[{"key":"ijkbo.2011100104-0","doi-asserted-by":"publisher","DOI":"10.1145\/27633.27634"},{"key":"ijkbo.2011100104-1","unstructured":"Business Dictionary. (n. d.). Data warehouse. Retrieved from http:\/\/www.businessdictionary.com"},{"issue":"1","key":"ijkbo.2011100104-2","article-title":"An overview of data warehousing and OLAP technology.","volume":"26","author":"S.Chaudhari","year":"1997","journal-title":"SIGMOD Record"},{"key":"ijkbo.2011100104-3","author":"R.Elmasri","year":"2000","journal-title":"Fundamentals of database systems"},{"key":"ijkbo.2011100104-4","doi-asserted-by":"publisher","DOI":"10.4018\/jdwm.2005040101"},{"key":"ijkbo.2011100104-5","first-page":"4","article-title":"Data cleaning: Problems & current approaches.","volume":"24","author":"D.Hang-Hai","year":"2000","journal-title":"IEEE Bulletin of the Technical Committee on Data Engineering"},{"key":"ijkbo.2011100104-6","doi-asserted-by":"crossref","unstructured":"Hernandez, A. M., & Stolfo, S. J. (1995).The merge\/purge problem for large databases. In Proceedings of the ACM SIGMOD Conference on Management of Data (pp. 127-138).","DOI":"10.1145\/568271.223807"},{"key":"ijkbo.2011100104-7","doi-asserted-by":"publisher","DOI":"10.1023\/A:1009761603038"},{"key":"ijkbo.2011100104-8","first-page":"5","author":"F.Humboldt","year":"2003","journal-title":"Problems, methods, and challenges in comprehensive data cleansing"},{"key":"ijkbo.2011100104-9","doi-asserted-by":"crossref","DOI":"10.1007\/978-3-662-04138-3","author":"M.Jarke","year":"2000","journal-title":"Fundamentals of data warehouses"},{"key":"ijkbo.2011100104-10","doi-asserted-by":"crossref","unstructured":"Lee, M., Lu, H., Ling, T., & Ko, Y. (1999). Cleansing data for mining and warehousing. In Proceedings of the 10th International Conference on Database and Expert Systems Applications (pp. 751-760).","DOI":"10.1007\/3-540-48309-8_70"},{"key":"ijkbo.2011100104-11","unstructured":"Marcus, A., & Maletic, J. (2000, October). Data cleaning: Problems & current approaches. In Proceedings of the Conference on Information Quality."},{"key":"ijkbo.2011100104-12","unstructured":"Monge, A. (1997). Adaptive detection of approximately duplicate database records and the database integration approach to information discovery. In Proceedings of the SIGMOD Workshop on Data Mining & Knowledge Discovery."},{"key":"ijkbo.2011100104-13","unstructured":"Ohanekwu, T., & Ezeife, C. (2003). A token-based data cleaning technique for data warehouse systems. In Proceedings of the IEEE Workshop on Data Quality in Cooperative Information Systems."},{"key":"ijkbo.2011100104-14","doi-asserted-by":"publisher","DOI":"10.1145\/269012.269023"},{"key":"ijkbo.2011100104-15","first-page":"402","author":"P.Ponniah","year":"2001","journal-title":"Data warehousing fundamentals: A comprehensive guide for IT professionals"},{"key":"ijkbo.2011100104-16","author":"T.Redman","year":"1996","journal-title":"Data quality for the information age"},{"key":"ijkbo.2011100104-17","doi-asserted-by":"publisher","DOI":"10.1145\/269012.269025"},{"key":"ijkbo.2011100104-18","unstructured":"Shahri, H. H., Shahri, S. H., Hellerstein, J., & Raman, V. (2001). Potter\u2019s wheel: An interactive data cleaning system. In Proceedings of International Conference on Very Large Databases."},{"issue":"5","key":"ijkbo.2011100104-19","first-page":"117","article-title":"A unified framework and sequential data cleaning approach for a data warehouse.","volume":"8","author":"J.Tamilselvi","year":"2008","journal-title":"International Journal on Computer Science and Network Security"},{"issue":"2","key":"ijkbo.2011100104-20","article-title":"Detection and elimination of duplicate data using token-based method for a data warehouse: A clustering based approach.","volume":"5","author":"J.Tamilselvi","year":"2009","journal-title":"International Journal of Computational Intelligence Research"},{"key":"ijkbo.2011100104-21","unstructured":"Wikipedia. (n. d.). Dirty data. Retrieved from http:\/\/en.wikipedia.org"},{"key":"ijkbo.2011100104-22","author":"W. E.Winkler","year":"1995","journal-title":"Matching and record linkage in business survey methods"},{"key":"ijkbo.2011100104-23","unstructured":"Winkler, W. E. (1999). State of statistical data editing and current research problems. In Proceedings of the UN\/ECE Work Session on Statistical Data Editing."},{"key":"ijkbo.2011100104-24","doi-asserted-by":"publisher","DOI":"10.1023\/A:1009783824328"}],"container-title":["International Journal of Knowledge-Based Organizations"],"original-title":[],"language":"ng","link":[{"URL":"https:\/\/www.igi-global.com\/viewtitle.aspx?TitleId=58919","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2022,6,1]],"date-time":"2022-06-01T16:36:44Z","timestamp":1654101404000},"score":1,"resource":{"primary":{"URL":"https:\/\/services.igi-global.com\/resolvedoi\/resolve.aspx?doi=10.4018\/ijkbo.2011100104"}},"subtitle":[""],"short-title":[],"issued":{"date-parts":[[2011,10,1]]},"references-count":25,"journal-issue":{"issue":"4","published-print":{"date-parts":[[2011,10]]}},"URL":"https:\/\/doi.org\/10.4018\/ijkbo.2011100104","relation":{},"ISSN":["2155-6393","2155-6407"],"issn-type":[{"value":"2155-6393","type":"print"},{"value":"2155-6407","type":"electronic"}],"subject":[],"published":{"date-parts":[[2011,10,1]]}}}