{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T04:32:06Z","timestamp":1750221126000,"version":"3.41.0"},"publisher-location":"New York, NY, USA","reference-count":49,"publisher":"ACM","license":[{"start":{"date-parts":[[2019,6,25]],"date-time":"2019-06-25T00:00:00Z","timestamp":1561420800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2019,6,25]]},"DOI":"10.1145\/3299869.3319867","type":"proceedings-article","created":{"date-parts":[[2019,6,18]],"date-time":"2019-06-18T17:41:43Z","timestamp":1560879703000},"page":"247-262","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":14,"title":["Visual Segmentation for Information Extraction from Heterogeneous Visually Rich Documents"],"prefix":"10.1145","author":[{"given":"Ritesh","family":"Sarkhel","sequence":"first","affiliation":[{"name":"Ohio State University, Columbus, OH, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Arnab","family":"Nandi","sequence":"additional","affiliation":[{"name":"Ohio State University, Columbus, OH, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2019,6,25]]},"reference":[{"key":"e_1_3_2_1_1_1","doi-asserted-by":"publisher","DOI":"10.3115\/v1\/D14-1164"},{"key":"e_1_3_2_1_2_1","doi-asserted-by":"crossref","unstructured":"Emilia Apostolova and Noriko Tomuro. 2014. Combining Visual and Textual Features for Information Extraction from Online Flyers.. In EMNLP. 1924--1929.  Emilia Apostolova and Noriko Tomuro. 2014. Combining Visual and Textual Features for Information Extraction from Online Flyers.. In EMNLP. 1924--1929.","DOI":"10.3115\/v1\/D14-1206"},{"key":"e_1_3_2_1_3_1","doi-asserted-by":"publisher","DOI":"10.5555\/647344.724142"},{"key":"e_1_3_2_1_4_1","unstructured":"Deng Cai Shipeng Yu Ji-Rong Wen and Wei-Ying Ma. 2003. Vips: a vision-based page segmentation algorithm. (2003).  Deng Cai Shipeng Yu Ji-Rong Wen and Wei-Ying Ma. 2003. Vips: a vision-based page segmentation algorithm. (2003)."},{"key":"e_1_3_2_1_5_1","volume-title":"LREC","volume":"2012","author":"Chang Angel X","year":"2012","unstructured":"Angel X Chang and Christopher D Manning . 2012 . Sutime: A library for recognizing and normalizing time expressions .. In LREC , Vol. 2012 . 3735--3740. Angel X Chang and Christopher D Manning. 2012. Sutime: A library for recognizing and normalizing time expressions.. In LREC, Vol. 2012. 3735--3740."},{"key":"e_1_3_2_1_6_1","doi-asserted-by":"publisher","DOI":"10.1145\/2160601.2160605"},{"key":"e_1_3_2_1_7_1","doi-asserted-by":"publisher","DOI":"10.1145\/1807167.1807339"},{"volume-title":"Introduction to algorithms","author":"Cormen Thomas H","key":"e_1_3_2_1_8_1","unstructured":"Thomas H Cormen , Charles E Leiserson , Ronald L Rivest , and Clifford Stein . 2009. Introduction to algorithms . MIT press . Thomas H Cormen, Charles E Leiserson, Ronald L Rivest, and Clifford Stein. 2009. Introduction to algorithms .MIT press."},{"key":"e_1_3_2_1_9_1","first-page":"109","article-title":"Roadrunner: Towards automatic data extraction from large web sites","volume":"1","author":"Crescenzi Valter","year":"2001","unstructured":"Valter Crescenzi , Giansalvatore Mecca , Paolo Merialdo , 2001 . Roadrunner: Towards automatic data extraction from large web sites . In VLDB , Vol. 1. 109 -- 118 . Valter Crescenzi, Giansalvatore Mecca, Paolo Merialdo, et almbox. 2001. Roadrunner: Towards automatic data extraction from large web sites. In VLDB, Vol. 1. 109--118.","journal-title":"VLDB"},{"key":"e_1_3_2_1_10_1","doi-asserted-by":"publisher","DOI":"10.1145\/2488388.2488420"},{"key":"e_1_3_2_1_11_1","doi-asserted-by":"publisher","DOI":"10.1145\/1519103.1519106"},{"key":"e_1_3_2_1_12_1","doi-asserted-by":"publisher","DOI":"10.1007\/s11263-009-0275-4"},{"key":"e_1_3_2_1_13_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.knosys.2014.07.007"},{"key":"e_1_3_2_1_14_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-23192-1_27"},{"key":"e_1_3_2_1_15_1","doi-asserted-by":"publisher","DOI":"10.1145\/1242572.1242583"},{"volume-title":"Information extraction from text. Mining text data","author":"Jiang Jing","key":"e_1_3_2_1_16_1","unstructured":"Jing Jiang . 2012. Information extraction from text. Mining text data . Springer , 11--41. Jing Jiang. 2012. Information extraction from text. Mining text data . Springer, 11--41."},{"key":"e_1_3_2_1_17_1","unstructured":"Dan Jurafsky. 2000. Speech & language processing .Pearson Education India.  Dan Jurafsky. 2000. Speech & language processing .Pearson Education India."},{"key":"e_1_3_2_1_18_1","doi-asserted-by":"publisher","DOI":"10.1109\/34.221173"},{"key":"e_1_3_2_1_19_1","doi-asserted-by":"publisher","DOI":"10.1016\/S0004-3702(99)00100-9"},{"key":"e_1_3_2_1_20_1","doi-asserted-by":"publisher","DOI":"10.1145\/565117.565137"},{"key":"e_1_3_2_1_21_1","doi-asserted-by":"publisher","DOI":"10.1109\/TKDE.2009.109"},{"key":"e_1_3_2_1_22_1","unstructured":"Astera LLC. 2018. ReportMiner: A Data Extraction Solution. https:\/\/www.astera.com\/products\/report-miner . (2018). Accessed: 2018-09--30.  Astera LLC. 2018. ReportMiner: A Data Extraction Solution. https:\/\/www.astera.com\/products\/report-miner . (2018). Accessed: 2018-09--30."},{"key":"e_1_3_2_1_23_1","doi-asserted-by":"publisher","DOI":"10.14778\/2824032.2824058"},{"key":"e_1_3_2_1_24_1","unstructured":"Google Maps. 2018. Google Maps Api. https:\/\/developers.google.com\/maps . (2018). Accessed: 2018-09--30.  Google Maps. 2018. Google Maps Api. https:\/\/developers.google.com\/maps . (2018). Accessed: 2018-09--30."},{"key":"e_1_3_2_1_25_1","volume-title":"Survey of multi-objective optimization methods for engineering. Structural and multidisciplinary optimization","author":"Timothy Marler R","year":"2004","unstructured":"R Timothy Marler and Jasbir S Arora . 2004. Survey of multi-objective optimization methods for engineering. Structural and multidisciplinary optimization , Vol. 26 , 6 ( 2004 ), 369--395. R Timothy Marler and Jasbir S Arora. 2004. Survey of multi-objective optimization methods for engineering. Structural and multidisciplinary optimization, Vol. 26, 6 (2004), 369--395."},{"key":"e_1_3_2_1_26_1","unstructured":"Tomas Mikolov Ilya Sutskever Kai Chen Greg S Corrado and Jeff Dean. 2013. Distributed representations of words and phrases and their compositionality. In Advances in neural information processing systems. 3111--3119.   Tomas Mikolov Ilya Sutskever Kai Chen Greg S Corrado and Jeff Dean. 2013. Distributed representations of words and phrases and their compositionality. In Advances in neural information processing systems. 3111--3119."},{"key":"e_1_3_2_1_27_1","doi-asserted-by":"publisher","DOI":"10.5555\/1690219.1690287"},{"key":"e_1_3_2_1_28_1","doi-asserted-by":"publisher","DOI":"10.1007\/s10115-017-1097-2"},{"key":"e_1_3_2_1_29_1","doi-asserted-by":"publisher","DOI":"10.1075\/li.30.1.03nad"},{"key":"e_1_3_2_1_30_1","doi-asserted-by":"publisher","DOI":"10.1145\/1459352.1459355"},{"key":"e_1_3_2_1_31_1","volume-title":"Proceedings of the National Conference on Artificial Intelligence","volume":"22","author":"Nguyen Dat PT","year":"2007","unstructured":"Dat PT Nguyen , Yutaka Matsuo , and Mitsuru Ishizuka . 2007 . Relation extraction from wikipedia using subtree mining . In Proceedings of the National Conference on Artificial Intelligence , Vol. 22 . Menlo Park, CA; Cambridge, MA; London; AAAI Press; MIT Press; 1999, 1414. Dat PT Nguyen, Yutaka Matsuo, and Mitsuru Ishizuka. 2007. Relation extraction from wikipedia using subtree mining. In Proceedings of the National Conference on Artificial Intelligence, Vol. 22. Menlo Park, CA; Cambridge, MA; London; AAAI Press; MIT Press; 1999, 1414."},{"key":"e_1_3_2_1_32_1","first-page":"25","article-title":"DeepDive: Web-scale Knowledge-base Construction using Statistical Learning and Inference","volume":"12","author":"Niu Feng","year":"2012","unstructured":"Feng Niu , Ce Zhang , Christopher R\u00e9 , and Jude W Shavlik . 2012 . DeepDive: Web-scale Knowledge-base Construction using Statistical Learning and Inference . VLDS , Vol. 12 (2012), 25 -- 28 . Feng Niu, Ce Zhang, Christopher R\u00e9, and Jude W Shavlik. 2012. DeepDive: Web-scale Knowledge-base Construction using Statistical Learning and Inference. VLDS, Vol. 12 (2012), 25--28.","journal-title":"VLDS"},{"key":"e_1_3_2_1_33_1","unstructured":"National Institute of Standards and Technology. 2018. NIST Special Database 6. https:\/\/www.nist.gov\/srd\/nist-special-database-6 . (2018). Accessed: 2018-09--30.  National Institute of Standards and Technology. 2018. NIST Special Database 6. https:\/\/www.nist.gov\/srd\/nist-special-database-6 . (2018). Accessed: 2018-09--30."},{"key":"e_1_3_2_1_34_1","volume-title":"Effective slot filling based on shallow distant supervision methods. arXiv preprint arXiv:1401.1158","author":"Roth Benjamin","year":"2014","unstructured":"Benjamin Roth , Tassilo Barth , Michael Wiegand , Mittul Singh , and Dietrich Klakow . 2014. Effective slot filling based on shallow distant supervision methods. arXiv preprint arXiv:1401.1158 ( 2014 ). Benjamin Roth, Tassilo Barth, Michael Wiegand, Mittul Singh, and Dietrich Klakow. 2014. Effective slot filling based on shallow distant supervision methods. arXiv preprint arXiv:1401.1158 (2014)."},{"key":"e_1_3_2_1_35_1","doi-asserted-by":"publisher","DOI":"10.1561\/1900000003"},{"key":"e_1_3_2_1_36_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.patcog.2017.05.022"},{"key":"e_1_3_2_1_37_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.patcog.2016.04.010"},{"key":"e_1_3_2_1_38_1","unstructured":"Karin Kipper Schuler. 2005. VerbNet: A broad-coverage comprehensive verb lexicon. (2005).  Karin Kipper Schuler. 2005. VerbNet: A broad-coverage comprehensive verb lexicon. (2005)."},{"key":"e_1_3_2_1_39_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICDAR.2011.296"},{"key":"e_1_3_2_1_40_1","doi-asserted-by":"publisher","DOI":"10.2307\/2333709"},{"key":"e_1_3_2_1_41_1","doi-asserted-by":"publisher","DOI":"10.5555\/1304596.1304846"},{"key":"e_1_3_2_1_42_1","unstructured":"Rion Snow Daniel Jurafsky and Andrew Y Ng. 2005. Learning syntactic patterns for automatic hypernym discovery. In Advances in neural information processing systems. 1297--1304.   Rion Snow Daniel Jurafsky and Andrew Y Ng. 2005. Learning syntactic patterns for automatic hypernym discovery. In Advances in neural information processing systems. 1297--1304."},{"key":"e_1_3_2_1_43_1","doi-asserted-by":"publisher","DOI":"10.1145\/2009916.2009952"},{"key":"e_1_3_2_1_44_1","doi-asserted-by":"publisher","DOI":"10.1561\/0600000017"},{"key":"e_1_3_2_1_45_1","doi-asserted-by":"publisher","DOI":"10.1145\/3183713.3183729"},{"key":"e_1_3_2_1_46_1","doi-asserted-by":"crossref","unstructured":"Yudong Yang Yu Chen and HongJiang Zhang. 2003. HTML page analysis based on visual cues. Web Document Analysis: Challenges and Opportunities. World Scientific 113--131.  Yudong Yang Yu Chen and HongJiang Zhang. 2003. HTML page analysis based on visual cues. Web Document Analysis: Challenges and Opportunities. World Scientific 113--131.","DOI":"10.1142\/9789812775375_0007"},{"key":"e_1_3_2_1_47_1","doi-asserted-by":"publisher","DOI":"10.1145\/775047.775058"},{"key":"e_1_3_2_1_48_1","volume-title":"AIP Conference Proceedings","volume":"1864","author":"Zhang Xiaowen","year":"2017","unstructured":"Xiaowen Zhang and Bingfeng Chen . 2017 . A construction scheme of web page comment information extraction system based on frequent subtree mining . In AIP Conference Proceedings , Vol. 1864 . AIP Publishing, 0 20059. Xiaowen Zhang and Bingfeng Chen. 2017. A construction scheme of web page comment information extraction system based on frequent subtree mining. In AIP Conference Proceedings, Vol. 1864. AIP Publishing, 020059."},{"key":"e_1_3_2_1_49_1","unstructured":"Ziyan Zhou and Muntasir Mashuq. 2014. Web Content Extraction Through Machine Learning. (2014).  Ziyan Zhou and Muntasir Mashuq. 2014. Web Content Extraction Through Machine Learning. (2014)."}],"event":{"name":"SIGMOD\/PODS '19: International Conference on Management of Data","sponsor":["SIGMOD ACM Special Interest Group on Management of Data"],"location":"Amsterdam Netherlands","acronym":"SIGMOD\/PODS '19"},"container-title":["Proceedings of the 2019 International Conference on Management of Data"],"original-title":[],"link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3299869.3319867","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3299869.3319867","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T01:02:16Z","timestamp":1750208536000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3299869.3319867"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2019,6,25]]},"references-count":49,"alternative-id":["10.1145\/3299869.3319867","10.1145\/3299869"],"URL":"https:\/\/doi.org\/10.1145\/3299869.3319867","relation":{},"subject":[],"published":{"date-parts":[[2019,6,25]]},"assertion":[{"value":"2019-06-25","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}