{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,1,11]],"date-time":"2026-01-11T05:17:08Z","timestamp":1768108628129,"version":"3.49.0"},"reference-count":62,"publisher":"Association for Computing Machinery (ACM)","issue":"5","content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Proc. VLDB Endow."],"published-print":{"date-parts":[[2021,1]]},"abstract":"<jats:p>Along with textual content, visual features play an essential role in the semantics of visually rich documents. Information extraction (IE) tasks perform poorly on these documents if these visual cues are not taken into account. In this paper, we present Artemis - a visually aware, machine-learning-based IE method for heterogeneous visually rich documents. Artemis represents a visual span in a document by jointly encoding its visual and textual context for IE tasks. Our main contribution is two-fold. First, we develop a deep-learning model that identifies the local context boundary of a visual span with minimal human-labeling. Second, we describe a deep neural network that encodes the multimodal context of a visual span into a fixed-length vector by taking its textual and layout-specific features into account. It identifies the visual span(s) containing a named entity by leveraging this learned representation followed by an inference task. We evaluate Artemis on four heterogeneous datasets from different domains over a suite of information extraction tasks. Results show that it outperforms state-of-the-art text-based methods by up to 17 points in F1-score.<\/jats:p>","DOI":"10.14778\/3446095.3446104","type":"journal-article","created":{"date-parts":[[2021,3,23]],"date-time":"2021-03-23T16:36:58Z","timestamp":1616517418000},"page":"822-834","source":"Crossref","is-referenced-by-count":7,"title":["Improving information extraction from visually rich documents using visual span representations"],"prefix":"10.14778","volume":"14","author":[{"given":"Ritesh","family":"Sarkhel","sequence":"first","affiliation":[{"name":"The Ohio State Universtiy"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Arnab","family":"Nandi","sequence":"additional","affiliation":[{"name":"The Ohio State University"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2021,3,23]]},"reference":[{"key":"e_1_2_1_1_1","volume-title":"Neural machine translation by jointly learning to align and translate. arXiv preprint arXiv:1409.0473","author":"Bahdanau Dzmitry","year":"2014"},{"key":"e_1_2_1_2_1","doi-asserted-by":"publisher","DOI":"10.5555\/1304596.1304952"},{"key":"e_1_2_1_3_1","unstructured":"Deng Cai Shipeng Yu Ji-Rong Wen and Wei-Ying Ma. 2003. Vips: a vision-based page segmentation algorithm. (2003).  Deng Cai Shipeng Yu Ji-Rong Wen and Wei-Ying Ma. 2003. Vips: a vision-based page segmentation algorithm. (2003)."},{"key":"e_1_2_1_4_1","doi-asserted-by":"publisher","DOI":"10.1145\/2160601.2160605"},{"key":"e_1_2_1_5_1","doi-asserted-by":"publisher","DOI":"10.1145\/1815330.1815333"},{"key":"e_1_2_1_6_1","volume-title":"Thirty-second AAAI conference on artificial intelligence.","author":"Dai Quanyu","year":"2018"},{"key":"e_1_2_1_7_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICDAR.2019.00030"},{"key":"e_1_2_1_8_1","volume-title":"Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805","author":"Devlin Jacob","year":"2018"},{"key":"e_1_2_1_9_1","doi-asserted-by":"publisher","DOI":"10.1145\/1519103.1519106"},{"key":"e_1_2_1_10_1","doi-asserted-by":"publisher","DOI":"10.1007\/s11263-009-0275-4"},{"key":"e_1_2_1_11_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-23192-1_27"},{"key":"e_1_2_1_12_1","doi-asserted-by":"publisher","DOI":"10.5555\/2969033.2969125"},{"key":"e_1_2_1_13_1","unstructured":"The Stanford NLP Group. 2020. Stanford Part-Of-Speech Tagger. Accessed: 2020-01-31.  The Stanford NLP Group. 2020. Stanford Part-Of-Speech Tagger. Accessed: 2020-01-31."},{"key":"e_1_2_1_14_1","unstructured":"The Stanford NLP Group. 2020. Stanford Word Tokenizer. Accessed: 2020-01-31.  The Stanford NLP Group. 2020. Stanford Word Tokenizer. Accessed: 2020-01-31."},{"key":"e_1_2_1_15_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICDAR.2015.7333910"},{"key":"e_1_2_1_16_1","unstructured":"Nurse Tech Inc. 2018. NurseBrains. Accessed: 2019-01-25.  Nurse Tech Inc. 2018. NurseBrains. Accessed: 2019-01-25."},{"key":"e_1_2_1_17_1","volume-title":"Image-to-image translation with conditional adversarial networks. arXiv preprint","author":"Isola Phillip","year":"2017"},{"key":"e_1_2_1_18_1","volume-title":"Chargrid: Towards understanding 2d documents. arXiv preprint arXiv:1809.08799","author":"Katti Anoop Raveendra","year":"2018"},{"key":"e_1_2_1_19_1","volume-title":"Keras: Deep Learning for Humans. Accessed: 2018-09-30.","year":"2018"},{"key":"e_1_2_1_20_1","volume-title":"International Conference on Learning Representations (ICLR)","volume":"5","author":"Kinga D","year":"2015"},{"key":"e_1_2_1_21_1","doi-asserted-by":"publisher","DOI":"10.1016\/S0004-3702(99)00100-9"},{"key":"e_1_2_1_22_1","unstructured":"Matthew Lamm. 2020. Natural Language Processing with Deep Learning. Accessed: 2020-01-31.  Matthew Lamm. 2020. Natural Language Processing with Deep Learning. Accessed: 2020-01-31."},{"key":"e_1_2_1_23_1","volume-title":"Deep learning. nature 521, 7553","author":"LeCun Yann","year":"2015"},{"key":"e_1_2_1_24_1","doi-asserted-by":"publisher","DOI":"10.1145\/1148170.1148307"},{"key":"e_1_2_1_25_1","volume-title":"A hierarchical neural autoencoder for paragraphs and documents. arXiv preprint arXiv:1506.01057","author":"Li Jiwei","year":"2015"},{"key":"e_1_2_1_26_1","volume-title":"Graph convolution for multimodal information extraction from visually rich documents. arXiv preprint arXiv:1903.11279","author":"Liu Xiaojing","year":"2019"},{"key":"e_1_2_1_27_1","unstructured":"Astera LLC. 2018. ReportMiner: A Data Extraction Solution. Accessed: 2018-09-30.  Astera LLC. 2018. ReportMiner: A Data Extraction Solution. Accessed: 2018-09-30."},{"key":"e_1_2_1_28_1","volume-title":"Arash Einolghozati, and Prashant Shiralkar.","author":"Lockard Colin","year":"2018"},{"key":"e_1_2_1_29_1","volume-title":"A general framework for information extraction using dynamic span graphs. arXiv preprint arXiv:1904.03296","author":"Luan Yi","year":"2019"},{"key":"e_1_2_1_30_1","volume-title":"End-to-end sequence labeling via bidirectional lstm-cnns-crf. arXiv preprint arXiv:1603.01354","author":"Ma Xuezhe","year":"2016"},{"key":"e_1_2_1_31_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2020.acl-main.580"},{"key":"e_1_2_1_32_1","doi-asserted-by":"publisher","DOI":"10.14778\/2824032.2824058"},{"key":"e_1_2_1_33_1","unstructured":"Christopher Manning. 2017. Representations for language: From word embeddings to sentence meanings. Accessed: 2020-01-31.  Christopher Manning. 2017. Representations for language: From word embeddings to sentence meanings. Accessed: 2020-01-31."},{"key":"e_1_2_1_34_1","doi-asserted-by":"publisher","DOI":"10.3115\/v1\/P14-5010"},{"key":"e_1_2_1_35_1","doi-asserted-by":"publisher","DOI":"10.1007\/s10115-017-1097-2"},{"key":"e_1_2_1_36_1","doi-asserted-by":"publisher","DOI":"10.1177\/1555343418825429"},{"key":"e_1_2_1_37_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICDAR.2015.7333803"},{"key":"e_1_2_1_38_1","unstructured":"NIST. 2018. NIST Special Database 6. Accessed: 2018-09-30.  NIST. 2018. NIST Special Database 6. Accessed: 2018-09-30."},{"key":"e_1_2_1_39_1","first-page":"25","article-title":"DeepDive: Webscale Knowledge-base Construction using Statistical Learning and Inference","volume":"12","author":"Niu Feng","year":"2012","journal-title":"VLDS"},{"key":"e_1_2_1_40_1","doi-asserted-by":"publisher","DOI":"10.1109\/TKDE.2009.191"},{"key":"e_1_2_1_41_1","volume-title":"SEGAN: Speech enhancement generative adversarial network. arXiv preprint arXiv:1703.09452","author":"Pascual Santiago","year":"2017"},{"key":"e_1_2_1_42_1","volume-title":"An introduction to digital image processing. online]: http:\/\/www.programmersheaven.com\/articles\/patin\/ImageProc.pdf","author":"Patin Fr\u00e9d\u00e9ric","year":"2003"},{"key":"e_1_2_1_43_1","unstructured":"P David Pearson Michael L Kamil Peter B Mosenthal Rebecca Barr etal 2016. Handbook of reading research. Routledge.  P David Pearson Michael L Kamil Peter B Mosenthal Rebecca Barr et al. 2016. Handbook of reading research. Routledge."},{"key":"e_1_2_1_44_1","doi-asserted-by":"publisher","DOI":"10.14778\/3157794.3157797"},{"key":"e_1_2_1_45_1","doi-asserted-by":"publisher","DOI":"10.1145\/3035918.3056442"},{"key":"e_1_2_1_46_1","doi-asserted-by":"publisher","DOI":"10.5555\/3157382.3157497"},{"key":"e_1_2_1_47_1","doi-asserted-by":"publisher","DOI":"10.5555\/2969239.2969250"},{"key":"e_1_2_1_48_1","doi-asserted-by":"publisher","DOI":"10.1561\/1900000003"},{"key":"e_1_2_1_49_1","volume-title":"Proceedings of the 28th International Conference on Computational Linguistics. 6871--6882","author":"Sarkhel Ritesh","year":"2020"},{"key":"e_1_2_1_50_1","doi-asserted-by":"publisher","DOI":"10.5555\/3367471.3367508"},{"key":"e_1_2_1_51_1","doi-asserted-by":"publisher","DOI":"10.1145\/3299869.3319867"},{"key":"e_1_2_1_52_1","doi-asserted-by":"publisher","DOI":"10.1177\/2327857918071045"},{"key":"e_1_2_1_53_1","doi-asserted-by":"publisher","DOI":"10.1145\/3068335"},{"key":"e_1_2_1_54_1","doi-asserted-by":"publisher","DOI":"10.5555\/1304596.1304846"},{"key":"e_1_2_1_55_1","doi-asserted-by":"publisher","DOI":"10.1145\/2009916.2009952"},{"key":"e_1_2_1_56_1","volume-title":"Symposium on document image understanding and technology (SDIUT). 199--205","author":"Thoma GFG","year":"2003"},{"key":"e_1_2_1_57_1","doi-asserted-by":"publisher","DOI":"10.5555\/1858681.1858721"},{"key":"e_1_2_1_58_1","volume-title":"relation, and event extraction with contextualized span representations. arXiv preprint arXiv:1909.03546","author":"Wadden David","year":"2019"},{"key":"e_1_2_1_59_1","doi-asserted-by":"publisher","DOI":"10.1145\/3183713.3183729"},{"key":"e_1_2_1_60_1","volume-title":"Joint extraction of events and entities within a document context. arXiv preprint arXiv:1609.03632","author":"Yang Bishan","year":"2016"},{"key":"e_1_2_1_61_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.462"},{"key":"e_1_2_1_62_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.media.2019.101552"}],"container-title":["Proceedings of the VLDB Endowment"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.14778\/3446095.3446104","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2022,12,28]],"date-time":"2022-12-28T11:23:13Z","timestamp":1672226593000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.14778\/3446095.3446104"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2021,1]]},"references-count":62,"journal-issue":{"issue":"5","published-print":{"date-parts":[[2021,1]]}},"alternative-id":["10.14778\/3446095.3446104"],"URL":"https:\/\/doi.org\/10.14778\/3446095.3446104","relation":{},"ISSN":["2150-8097"],"issn-type":[{"value":"2150-8097","type":"print"}],"subject":[],"published":{"date-parts":[[2021,1]]}}}