{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,4,13]],"date-time":"2026-04-13T21:53:06Z","timestamp":1776117186913,"version":"3.50.1"},"reference-count":18,"publisher":"Association for Computing Machinery (ACM)","issue":"4","license":[{"start":{"date-parts":[[2009,3,20]],"date-time":"2009-03-20T00:00:00Z","timestamp":1237507200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["SIGMOD Rec."],"published-print":{"date-parts":[[2009,3,20]]},"abstract":"<jats:p>Over the past few years, we have been trying to build an end-to-end system at Wisconsin to manage unstructured data, using extraction, integration, and user interaction. This paper describes the key information extraction (IE) challenges that we have run into, and sketches our solutions. We discuss in particular developing a declarative IE language, optimizing for this language, generating IE provenance, incorporating user feedback into the IE process, developing a novel wiki-based user interface for feedback, best-effort IE, pushing IE into RDBMSs, and more. Our work suggests that IE in managing unstructured data can open up many interesting research challenges, and that these challenges can greatly benefit from the wealth of work on managing structured data that has been carried out by the database community.<\/jats:p>","DOI":"10.1145\/1519103.1519106","type":"journal-article","created":{"date-parts":[[2009,4,6]],"date-time":"2009-04-06T16:34:22Z","timestamp":1239035662000},"page":"14-20","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":48,"title":["Information extraction challenges in managing unstructured data"],"prefix":"10.1145","volume":"37","author":[{"given":"AnHai","family":"Doan","sequence":"first","affiliation":[{"name":"University of Wisconsin-Madison"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Jeffrey F.","family":"Naughton","sequence":"additional","affiliation":[{"name":"University of Wisconsin-Madison"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Raghu","family":"Ramakrishnan","sequence":"additional","affiliation":[{"name":"University of Wisconsin-Madison"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Akanksha","family":"Baid","sequence":"additional","affiliation":[{"name":"University of Wisconsin-Madison"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Xiaoyong","family":"Chai","sequence":"additional","affiliation":[{"name":"University of Wisconsin-Madison"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Fei","family":"Chen","sequence":"additional","affiliation":[{"name":"University of Wisconsin-Madison"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Ting","family":"Chen","sequence":"additional","affiliation":[{"name":"University of Wisconsin-Madison"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Eric","family":"Chu","sequence":"additional","affiliation":[{"name":"University of Wisconsin-Madison"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Pedro","family":"DeRose","sequence":"additional","affiliation":[{"name":"University of Wisconsin-Madison"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Byron","family":"Gao","sequence":"additional","affiliation":[{"name":"University of Wisconsin-Madison"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Chaitanya","family":"Gokhale","sequence":"additional","affiliation":[{"name":"University of Wisconsin-Madison"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Jiansheng","family":"Huang","sequence":"additional","affiliation":[{"name":"University of Wisconsin-Madison"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Warren","family":"Shen","sequence":"additional","affiliation":[{"name":"University of Wisconsin-Madison"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Ba-Quy","family":"Vuong","sequence":"additional","affiliation":[{"name":"University of Wisconsin-Madison"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2009,3,20]]},"reference":[{"key":"e_1_2_1_1_1","doi-asserted-by":"publisher","DOI":"10.1145\/375663.375774"},{"key":"e_1_2_1_2_1","doi-asserted-by":"publisher","DOI":"10.5555\/646543.696220"},{"key":"e_1_2_1_3_1","doi-asserted-by":"publisher","DOI":"10.1145\/1066157.1066289"},{"key":"e_1_2_1_5_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICDE.2008.4497503"},{"key":"e_1_2_1_6_1","volume-title":"VLDB","author":"Chu E.","year":"2007","unstructured":"E. Chu , A. Baid , T. Chen , A. Doan , and J. F. Naughton . A relational approach to incrementally extracting and querying structure in unstructured data . In VLDB , 2007 . E. Chu, A. Baid, T. Chen, A. Doan, and J. F. Naughton. A relational approach to incrementally extracting and querying structure in unstructured data. In VLDB, 2007."},{"key":"e_1_2_1_7_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICDE.2008.4497473"},{"key":"e_1_2_1_8_1","volume-title":"VLDB","author":"DeRose P.","year":"2007","unstructured":"P. DeRose , W. Shen , F. Chen , A. Doan , and R. Ramakrishnan . Building structured web community portals: A top-down, compositional, and incremental approach . In VLDB , 2007 . P. DeRose, W. Shen, F. Chen, A. Doan, and R. Ramakrishnan. Building structured web community portals: A top-down, compositional, and incremental approach. In VLDB, 2007."},{"key":"e_1_2_1_9_1","volume-title":"CIDR","author":"DeRose P.","year":"2007","unstructured":"P. DeRose , W. Shen , F. Chen , Y. Lee , D. Burdick , A. Doan , and R. Ramakrishnan . Dblife: A community information management platform for the database research community (demo) . In CIDR , 2007 . P. DeRose, W. Shen, F. Chen, Y. Lee, D. Burdick, A. Doan, and R. Ramakrishnan. Dblife: A community information management platform for the database research community (demo). In CIDR, 2007."},{"key":"e_1_2_1_10_1","volume-title":"Workshop on Information Integration Methods, Architectures, and Systems (IIMAS) at ICDE-08","author":"Doan A.","unstructured":"A. Doan . Data integration research challenges in community information management systems, 2008. Keynote talk , Workshop on Information Integration Methods, Architectures, and Systems (IIMAS) at ICDE-08 . A. Doan. Data integration research challenges in community information management systems, 2008. Keynote talk, Workshop on Information Integration Methods, Architectures, and Systems (IIMAS) at ICDE-08."},{"issue":"2","key":"e_1_2_1_11_1","first-page":"32","article-title":"User-centric research challenges in community information management systems","volume":"30","author":"Doan A.","year":"2007","unstructured":"A. Doan , P. Bohannon , R. Ramakrishnan , X. Chai , P. DeRose , B. Gao , and W. Shen . User-centric research challenges in community information management systems . IEEE Data Engineering Bulletin , 30 ( 2 ): 32 -- 40 , 2007 . A. Doan, P. Bohannon, R. Ramakrishnan, X. Chai, P. DeRose, B. Gao, and W. Shen. User-centric research challenges in community information management systems. IEEE Data Engineering Bulletin, 30(2):32--40, 2007.","journal-title":"IEEE Data Engineering Bulletin"},{"key":"e_1_2_1_12_1","volume-title":"CIDR","author":"Doan A.","year":"2009","unstructured":"A. Doan , J. F. Naughton , A. Baid , X. Chai , F. Chen , T. Chen , E. Chu , P. DeRose , B. Gao , C. Gokhale , J. Huang , W. Shen , and B. Vuong . The case for a structured approach to managing unstructured data . In CIDR , 2009 . A. Doan, J. F. Naughton, A. Baid, X. Chai, F. Chen, T. Chen, E. Chu, P. DeRose, B. Gao, C. Gokhale, J. Huang, W. Shen, and B. Vuong. The case for a structured approach to managing unstructured data. In CIDR, 2009."},{"issue":"1","key":"e_1_2_1_13_1","first-page":"64","article-title":"Community information management","volume":"29","author":"Doan A.","year":"2006","unstructured":"A. Doan , R. Ramakrishnan , F. Chen , P. DeRose , Y. Lee , R. McCann , M. Sayyadian , and W. Shen . Community information management . IEEE Data Engineering Bulletin , 29 ( 1 ): 64 -- 72 , 2006 . A. Doan, R. Ramakrishnan, F. Chen, P. DeRose, Y. Lee, R. McCann, M. Sayyadian, and W. Shen. Community information management. IEEE Data Engineering Bulletin, 29(1):64--72, 2006.","journal-title":"IEEE Data Engineering Bulletin"},{"key":"e_1_2_1_14_1","doi-asserted-by":"publisher","DOI":"10.1145\/1142351.1142352"},{"key":"e_1_2_1_15_1","doi-asserted-by":"publisher","DOI":"10.14778\/1453856.1453936"},{"key":"e_1_2_1_16_1","doi-asserted-by":"publisher","DOI":"10.1145\/1519103.1519105"},{"key":"e_1_2_1_17_1","doi-asserted-by":"publisher","DOI":"10.1145\/1376616.1376718"},{"key":"e_1_2_1_18_1","volume-title":"VLDB","author":"Shen W.","year":"2007","unstructured":"W. Shen , A. Doan , J. F. Naughton , and R. Ramakrishnan . Declarative information extraction using datalog with embedded extraction predicates . In VLDB , 2007 . W. Shen, A. Doan, J. F. Naughton, and R. Ramakrishnan. Declarative information extraction using datalog with embedded extraction predicates. In VLDB, 2007."},{"issue":"4","key":"e_1_2_1_20_1","first-page":"3","article-title":"Provenance in databases: Past, current, and future","volume":"30","author":"Tan W. C.","year":"2007","unstructured":"W. C. Tan . Provenance in databases: Past, current, and future . IEEE Data Eng. Bull. , 30 ( 4 ): 3 -- 12 , 2007 . W. C. Tan. Provenance in databases: Past, current, and future. IEEE Data Eng. Bull., 30(4):3--12, 2007.","journal-title":"IEEE Data Eng. Bull."}],"container-title":["ACM SIGMOD Record"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/1519103.1519106","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/1519103.1519106","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T13:29:53Z","timestamp":1750253393000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/1519103.1519106"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2009,3,20]]},"references-count":18,"journal-issue":{"issue":"4","published-print":{"date-parts":[[2009,3,20]]}},"alternative-id":["10.1145\/1519103.1519106"],"URL":"https:\/\/doi.org\/10.1145\/1519103.1519106","relation":{},"ISSN":["0163-5808"],"issn-type":[{"value":"0163-5808","type":"print"}],"subject":[],"published":{"date-parts":[[2009,3,20]]},"assertion":[{"value":"2009-03-20","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}