{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2023,7,27]],"date-time":"2023-07-27T00:12:05Z","timestamp":1690416725023},"reference-count":23,"publisher":"Association for Computing Machinery (ACM)","issue":"12","content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Proc. VLDB Endow."],"published-print":{"date-parts":[[2012,8]]},"abstract":"<jats:p>Master data management (MDM) integrates data from multiple structured data sources and builds a consolidated 360-degree view of business entities such as customers and products. Today's MDM systems are not prepared to integrate information from unstructured data sources, such as news reports, emails, call-center transcripts, and chat logs. However, those unstructured data sources may contain valuable information about the same entities known to MDM from the structured data sources. Integrating information from unstructured data into MDM is challenging as textual references to existing MDM entities are often incomplete and imprecise and the additional entity information extracted from text should not impact the trustworthiness of MDM data.<\/jats:p><jats:p>In this paper, we present an architecture for making MDM text-aware and showcase its implementation as IBM Info-Sphere MDM Extension for Unstructured Text Correlation, an add-on to IBM InfoSphere Master Data Management Standard Edition. We highlight how MDM benefits from additional evidence found in documents when doing entity resolution and relationship discovery. We experimentally demonstrate the feasibility of integrating information from unstructured data sources into MDM.<\/jats:p>","DOI":"10.14778\/2367502.2367524","type":"journal-article","created":{"date-parts":[[2014,6,24]],"date-time":"2014-06-24T12:17:57Z","timestamp":1403612277000},"page":"1862-1873","source":"Crossref","is-referenced-by-count":9,"title":["Exploiting evidence from unstructured data to enhance master data management"],"prefix":"10.14778","volume":"5","author":[{"given":"Karin","family":"Murthy","sequence":"first","affiliation":[{"name":"IBM Research - India"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Prasad M.","family":"Deshpande","sequence":"additional","affiliation":[{"name":"IBM Research - India"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Atreyee","family":"Dey","sequence":"additional","affiliation":[{"name":"IBM Research - India"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Ramanujam","family":"Halasipuram","sequence":"additional","affiliation":[{"name":"IBM Research - India"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Mukesh","family":"Mohania","sequence":"additional","affiliation":[{"name":"IBM Research - India"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"P.","family":"Deepak","sequence":"additional","affiliation":[{"name":"IBM Research - India"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Jennifer","family":"Reed","sequence":"additional","affiliation":[{"name":"IBM Software Group"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Scott","family":"Schumacher","sequence":"additional","affiliation":[{"name":"IBM Software Group"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2012,8]]},"reference":[{"key":"e_1_2_1_1_1","doi-asserted-by":"publisher","DOI":"10.1007\/s00778-008-0098-x"},{"key":"e_1_2_1_2_1","first-page":"667","volume-title":"VLDB","author":"Chakaravarthy V. T.","year":"2006","unstructured":"V. T. Chakaravarthy , H. Gupta , P. Roy , and M. Mohania . Efficiently linking text documents with relevant structured information . In VLDB , pages 667 -- 678 , 2006 . V. T. Chakaravarthy, H. Gupta, P. Roy, and M. Mohania. Efficiently linking text documents with relevant structured information. In VLDB, pages 667--678, 2006."},{"key":"e_1_2_1_3_1","first-page":"128","volume-title":"ACL","author":"Chiticariu L.","year":"2010","unstructured":"L. Chiticariu , R. Krishnamurthy , Y. Li , S. Raghavan , F. R. Reiss , and S. Vaithyanathan . SystemT: an algebraic approach to declarative information extraction . In ACL , pages 128 -- 137 , 2010 . L. Chiticariu, R. Krishnamurthy, Y. Li, S. Raghavan, F. R. Reiss, and S. Vaithyanathan. SystemT: an algebraic approach to declarative information extraction. In ACL, pages 128--137, 2010."},{"key":"e_1_2_1_4_1","doi-asserted-by":"crossref","first-page":"89","DOI":"10.1145\/1014052.1014065","volume-title":"KDD","author":"Cohen W. W.","year":"2004","unstructured":"W. W. Cohen and S. Sarawagi . Exploiting dictionaries in named entity extraction: combining semi-markov extraction processes and data integration methods . In KDD , pages 89 -- 98 , 2004 . 10.1145\/1014052.1014065 W. W. Cohen and S. Sarawagi. Exploiting dictionaries in named entity extraction: combining semi-markov extraction processes and data integration methods. In KDD, pages 89--98, 2004. 10.1145\/1014052.1014065"},{"key":"e_1_2_1_5_1","doi-asserted-by":"publisher","DOI":"10.1109\/TKDE.2007.9"},{"key":"e_1_2_1_6_1","doi-asserted-by":"publisher","DOI":"10.1007\/s00778-010-0206-6"},{"key":"e_1_2_1_7_1","doi-asserted-by":"crossref","first-page":"1183","DOI":"10.1080\/01621459.1969.10501049","volume":"64","author":"Fellegi I. P.","year":"1969","unstructured":"I. P. Fellegi and A. B. Sunter . A Theory for Record Linkage. Journal of the American Statistical Association , 64 : 1183 -- 1210 , 1969 . I. P. Fellegi and A. B. Sunter. A Theory for Record Linkage. Journal of the American Statistical Association, 64:1183--1210, 1969.","journal-title":"Journal of the American Statistical Association"},{"key":"e_1_2_1_8_1","unstructured":"IBM Corporation. Master data management and accurate data matching. http:\/\/www-01.ibm.com\/software\/data\/master-data-management\/library.html 2010. IBM Corporation. Master data management and accurate data matching. http:\/\/www-01.ibm.com\/software\/data\/master-data-management\/library.html 2010."},{"key":"e_1_2_1_9_1","first-page":"718","volume-title":"SIGMOD","author":"Jonas J.","year":"2006","unstructured":"J. Jonas . Identity resolution : 23 years of practical experience and observations at scale . In SIGMOD , pages 718 -- 718 , 2006 . 10.1145\/1142473.1142556 J. Jonas. Identity resolution: 23 years of practical experience and observations at scale. In SIGMOD, pages 718--718, 2006. 10.1145\/1142473.1142556"},{"issue":"8","key":"e_1_2_1_10_1","first-page":"707","article-title":"Binary codes capable of correcting deletions, insertions, and reversals","volume":"10","author":"Levenshtein V. I.","year":"1966","unstructured":"V. I. Levenshtein . Binary codes capable of correcting deletions, insertions, and reversals . Soviet Physics Doklady , 10 ( 8 ): 707 -- 710 , 1966 . V. I. Levenshtein. Binary codes capable of correcting deletions, insertions, and reversals. Soviet Physics Doklady, 10(8):707--710, 1966.","journal-title":"Soviet Physics Doklady"},{"key":"e_1_2_1_11_1","first-page":"1341","volume-title":"CIKM","author":"Li G.","year":"2010","unstructured":"G. Li , D. Deng , and J. Feng . Extending dictionary-based entity extraction to tolerate errors . In CIKM , pages 1341 -- 1344 , 2010 . 10.1145\/1871437.1871616 G. Li, D. Deng, and J. Feng. Extending dictionary-based entity extraction to tolerate errors. In CIKM, pages 1341--1344, 2010. 10.1145\/1871437.1871616"},{"key":"e_1_2_1_12_1","first-page":"529","volume-title":"SIGMOD","author":"Li G.","year":"2011","unstructured":"G. Li , D. Deng , and J. Feng . Faerie: efficient filtering algorithms for approximate dictionary-based entity extraction . In SIGMOD , pages 529 -- 540 , 2011 . 10.1145\/1989323.1989379 G. Li, D. Deng, and J. Feng. Faerie: efficient filtering algorithms for approximate dictionary-based entity extraction. In SIGMOD, pages 529--540, 2011. 10.1145\/1989323.1989379"},{"key":"e_1_2_1_13_1","first-page":"1","volume-title":"ICISA","author":"Lim S.","year":"2010","unstructured":"S. Lim . Cleansing noisy city names in spatial data mining . In ICISA , pages 1 -- 8 , 2010 . S. Lim. Cleansing noisy city names in spatial data mining. In ICISA, pages 1--8, 2010."},{"key":"e_1_2_1_14_1","first-page":"29","volume-title":"ICDE","author":"Mansuri I. R.","year":"2006","unstructured":"I. R. Mansuri and S. Sarawagi . Integrating unstructured data into relational databases . In ICDE , pages 29 -- 29 , 2006 . 10.1109\/ICDE.2006.83 I. R. Mansuri and S. Sarawagi. Integrating unstructured data into relational databases. In ICDE, pages 29--29, 2006. 10.1109\/ICDE.2006.83"},{"key":"e_1_2_1_15_1","first-page":"491","volume-title":"ACL","author":"McDonald R.","year":"2005","unstructured":"R. McDonald , F. Pereira , S. Kulick , S. Winters , Y. Jin , and P. White . Simple algorithms for complex relation extraction with applications to biomedical IE . In ACL , pages 491 -- 498 , 2005 . 10.3115\/1219840.1219901 R. McDonald, F. Pereira, S. Kulick, S. Winters, Y. Jin, and P. White. Simple algorithms for complex relation extraction with applications to biomedical IE. In ACL, pages 491--498, 2005. 10.3115\/1219840.1219901"},{"key":"e_1_2_1_16_1","volume-title":"Hidden Treasure","author":"Messerschmidt M.","year":"2011","unstructured":"M. Messerschmidt and J. St\u00fcben . Hidden Treasure . PricewaterhouseCooper AG Wirtschaftspr\u00fcfungsgesellschaft , 2011 . M. Messerschmidt and J. St\u00fcben. Hidden Treasure. PricewaterhouseCooper AG Wirtschaftspr\u00fcfungsgesellschaft, 2011."},{"key":"e_1_2_1_17_1","first-page":"206","volume-title":"COMAD","author":"Murthy K.","year":"2010","unstructured":"K. Murthy , D. P, P. M. Deshpande , S. Kakaraparthy , V. S. Sandeep , V. Shyamsundar , and S. Singh . Content-aware master data management . In COMAD , pages 206 -- 209 , 2010 . K. Murthy, D. P, P. M. Deshpande, S. Kakaraparthy, V. S. Sandeep, V. Shyamsundar, and S. Singh. Content-aware master data management. In COMAD, pages 206--209, 2010."},{"key":"e_1_2_1_18_1","first-page":"851","volume-title":"COLING","author":"Okazaki N.","year":"2010","unstructured":"N. Okazaki and J. Tsujii . Simple and efficient algorithm for approximate dictionary matching . In COLING , pages 851 -- 859 , 2010 . N. Okazaki and J. Tsujii. Simple and efficient algorithm for approximate dictionary matching. In COLING, pages 851--859, 2010."},{"key":"e_1_2_1_19_1","doi-asserted-by":"publisher","DOI":"10.1561\/1900000003"},{"key":"e_1_2_1_20_1","first-page":"1185","volume-title":"NIPS","author":"Sarawagi S.","year":"2004","unstructured":"S. Sarawagi and W. W. Cohen . Semi-markov conditional random fields for information extraction . In NIPS , pages 1185 -- 1192 , 2004 . S. Sarawagi and W. W. Cohen. Semi-markov conditional random fields for information extraction. In NIPS, pages 1185--1192, 2004."},{"key":"e_1_2_1_21_1","volume-title":"IDC","author":"Villars R. L.","year":"2011","unstructured":"R. L. Villars . Scale-Out Storage in the Content-Driven Enterprise: Unleashing the Value of Information Assets . IDC , 2011 . R. L. Villars. Scale-Out Storage in the Content-Driven Enterprise: Unleashing the Value of Information Assets. IDC, 2011."},{"key":"e_1_2_1_22_1","first-page":"759","volume-title":"SIGMOD","author":"Wang W.","year":"2009","unstructured":"W. Wang , C. Xiao , X. Lin , and C. Zhang . Efficient approximate entity extraction with edit distance constraints . In SIGMOD , pages 759 -- 770 , 2009 . 10.1145\/1559845.1559925 W. Wang, C. Xiao, X. Lin, and C. Zhang. Efficient approximate entity extraction with edit distance constraints. In SIGMOD, pages 759--770, 2009. 10.1145\/1559845.1559925"},{"key":"e_1_2_1_23_1","doi-asserted-by":"crossref","first-page":"603","DOI":"10.3115\/1610075.1610160","volume-title":"EMNLP","author":"Wick M.","year":"2006","unstructured":"M. Wick , A. Culotta , and A. McCallum . Learning field compatibilities to extract database records from unstructured text . In EMNLP , pages 603 -- 611 , 2006 . M. Wick, A. Culotta, and A. McCallum. Learning field compatibilities to extract database records from unstructured text. In EMNLP, pages 603--611, 2006."}],"container-title":["Proceedings of the VLDB Endowment"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.14778\/2367502.2367524","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2023,7,26]],"date-time":"2023-07-26T23:56:16Z","timestamp":1690415776000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.14778\/2367502.2367524"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2012,8]]},"references-count":23,"journal-issue":{"issue":"12","published-print":{"date-parts":[[2012,8]]}},"alternative-id":["10.14778\/2367502.2367524"],"URL":"https:\/\/doi.org\/10.14778\/2367502.2367524","relation":{},"ISSN":["2150-8097"],"issn-type":[{"value":"2150-8097","type":"print"}],"subject":[],"published":{"date-parts":[[2012,8]]}}}