{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,4]],"date-time":"2026-05-04T00:27:40Z","timestamp":1777854460960,"version":"3.51.4"},"reference-count":27,"publisher":"SAGE Publications","issue":"2","license":[{"start":{"date-parts":[[2012,10,15]],"date-time":"2012-10-15T00:00:00Z","timestamp":1350259200000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/journals.sagepub.com\/page\/policies\/text-and-data-mining-license"}],"content-domain":{"domain":["journals.sagepub.com"],"crossmark-restriction":true},"short-container-title":["Journal of Information Science"],"published-print":{"date-parts":[[2013,4]]},"abstract":"<jats:p>Duplicate record detection has been an important issue in the fields of data and records management and various detection methods have been proposed. A new method, which uses an optical character recognition (OCR)-converted source of information for record matching to detect duplicates, is proposed and examined in this paper. First, the design of an experiment for examining the performance of such a duplicate detection method is discussed. The base record set with an OCR-converted title page and its verso were prepared along with two test record sets from different union catalogues, and duplicate records between the base record set and the test sets were manually identified. A duplicate detection system was developed to execute matching (1) between records, (2) between a record and an OCR-converted source of information and (3) using a combination of these. Second, matching performance at the individual data element level is examined. Third, the performance of duplicate record detection based on matching at the element level is examined through rule-based detection and machine learning-based detection. The results of the experiment show the usefulness of incorporating source of information into duplicate detection to a certain extent.<\/jats:p>","DOI":"10.1177\/0165551512459923","type":"journal-article","created":{"date-parts":[[2012,10,15]],"date-time":"2012-10-15T20:36:22Z","timestamp":1350333382000},"page":"153-168","update-policy":"https:\/\/doi.org\/10.1177\/sage-journals-update-policy","source":"Crossref","is-referenced-by-count":4,"title":["Duplicate bibliographic record detection with an OCR-converted source of information"],"prefix":"10.1177","volume":"39","author":[{"given":"Shoichi","family":"Taniguchi","sequence":"first","affiliation":[{"name":"School of Library and Information Science, Keio University, Japan"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"179","published-online":{"date-parts":[[2012,10,15]]},"reference":[{"key":"bibr1-0165551512459923","doi-asserted-by":"publisher","DOI":"10.1177\/016555159802400402"},{"key":"bibr2-0165551512459923","first-page":"213","volume":"13","author":"Lazinger SS","year":"1994","journal-title":"Inform Technol Libraries"},{"key":"bibr3-0165551512459923","doi-asserted-by":"publisher","DOI":"10.1108\/eb047100"},{"key":"bibr4-0165551512459923","doi-asserted-by":"publisher","DOI":"10.1108\/07378830810880379"},{"key":"bibr5-0165551512459923","first-page":"19","volume":"11","author":"Toney SR","year":"1992","journal-title":"Inform Technol Libraries"},{"key":"bibr6-0165551512459923","first-page":"59","volume":"37","author":"O\u2019Neill ET","year":"1993","journal-title":"Library Resourc Techn Serv"},{"key":"bibr7-0165551512459923","doi-asserted-by":"publisher","DOI":"10.28945\/852"},{"key":"bibr8-0165551512459923","doi-asserted-by":"publisher","DOI":"10.6017\/ital.v26i2.3278"},{"key":"bibr9-0165551512459923","first-page":"125","volume":"12","author":"Hickey TB","year":"1979","journal-title":"J Library Autom"},{"key":"bibr10-0165551512459923","first-page":"143","volume":"12","author":"MacLaury KD","year":"1979","journal-title":"J Library Autom"},{"key":"bibr11-0165551512459923","doi-asserted-by":"publisher","DOI":"10.1109\/TKDE.2007.250581"},{"key":"bibr12-0165551512459923","doi-asserted-by":"crossref","unstructured":"Winkler WE. Overview of record linkage and current research direction, http:\/\/www.census.gov\/srd\/papers\/pdf\/rrs2006-02.pdf (2006, accessed April 2012).","DOI":"10.1002\/9780470057339.var022.pub2"},{"key":"bibr13-0165551512459923","doi-asserted-by":"publisher","DOI":"10.1016\/0306-4573(87)90051-3"},{"key":"bibr14-0165551512459923","doi-asserted-by":"publisher","DOI":"10.1300\/J104v07n02_07"},{"key":"bibr15-0165551512459923","doi-asserted-by":"publisher","DOI":"10.1002\/(SICI)1097-4571(199103)42:2<99::AID-ASI3>3.0.CO;2-V"},{"key":"bibr16-0165551512459923","doi-asserted-by":"publisher","DOI":"10.1016\/0306-4573(91)90032-H"},{"key":"bibr17-0165551512459923","first-page":"123","volume-title":"Expert systems in libraries","author":"Vizine-Goetz D","year":"1990"},{"key":"bibr18-0165551512459923","doi-asserted-by":"publisher","DOI":"10.1016\/0306-4573(89)90006-X"},{"key":"bibr19-0165551512459923","doi-asserted-by":"publisher","DOI":"10.1108\/07378830410570494"},{"key":"bibr20-0165551512459923","doi-asserted-by":"publisher","DOI":"10.1108\/07378830410570502"},{"key":"bibr21-0165551512459923","doi-asserted-by":"publisher","DOI":"10.1002\/asi.20414"},{"key":"bibr22-0165551512459923","doi-asserted-by":"publisher","DOI":"10.1002\/asi.20551"},{"key":"bibr23-0165551512459923","unstructured":"ALCTS. Differences between, changes within: Guidelines on when to create a new record, revised edn, http:\/\/www.ala.org\/ala\/mgrps\/divs\/alcts\/resources\/org\/cat\/differences07.pdf (2007, accessed April 2012)."},{"key":"bibr24-0165551512459923","unstructured":"OCLC. WorldCat quality: An OCLC report, http:\/\/www.oclc.org\/reports\/worldcatquality\/214660usb_WorldCat_Quality.pdf (2011, accessed April 2012)."},{"key":"bibr25-0165551512459923","first-page":"124","volume":"57","author":"Taniguchi S","year":"2011","journal-title":"J Jap Soc Library Inform Sci"},{"key":"bibr26-0165551512459923","volume-title":"Data mining: Practical machine learning tools and techniques","author":"Witten IH","year":"2011","edition":"3"},{"key":"bibr27-0165551512459923","unstructured":"Coyle K. Record merging algorithm: for MARC records, http:\/\/kcoyle.net\/merge.html (2007, accessed April 2012)."}],"container-title":["Journal of Information Science"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/journals.sagepub.com\/doi\/pdf\/10.1177\/0165551512459923","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/journals.sagepub.com\/doi\/full-xml\/10.1177\/0165551512459923","content-type":"application\/xml","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/journals.sagepub.com\/doi\/pdf\/10.1177\/0165551512459923","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,4,29]],"date-time":"2026-04-29T23:08:27Z","timestamp":1777504107000},"score":1,"resource":{"primary":{"URL":"https:\/\/journals.sagepub.com\/doi\/10.1177\/0165551512459923"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2012,10,15]]},"references-count":27,"journal-issue":{"issue":"2","published-print":{"date-parts":[[2013,4]]}},"alternative-id":["10.1177\/0165551512459923"],"URL":"https:\/\/doi.org\/10.1177\/0165551512459923","relation":{},"ISSN":["0165-5515","1741-6485"],"issn-type":[{"value":"0165-5515","type":"print"},{"value":"1741-6485","type":"electronic"}],"subject":[],"published":{"date-parts":[[2012,10,15]]}}}