{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,6]],"date-time":"2026-07-06T12:42:20Z","timestamp":1783341740687,"version":"3.54.6"},"reference-count":27,"publisher":"Association for Computing Machinery (ACM)","issue":"6","content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Proc. VLDB Endow."],"published-print":{"date-parts":[[2021,2]]},"abstract":"<jats:p>\n            Extracting structured information from templatic documents is an important problem with the potential to automate many real-world business workflows such as payment, procurement, and payroll. The core challenge is that such documents can be laid out in virtually infinitely different ways. A good solution to this problem is one that generalizes well not only to\n            <jats:italic>known<\/jats:italic>\n            templates such as invoices from a known vendor, but also to\n            <jats:italic>unseen<\/jats:italic>\n            ones.\n          <\/jats:p>\n          <jats:p>We developed a system called Glean to tackle this problem. Given a target schema for a document type and some labeled documents of that type, Glean uses machine learning to automatically extract structured information from other documents of that type. In this paper, we describe the overall architecture of Glean, and discuss three key data management challenges : 1) managing the quality of ground truth data, 2) generating training data for the machine learning model using labeled documents, and 3) building tools that help a developer rapidly build and improve a model for a given document type. Through empirical studies on a real-world dataset, we show that these data management techniques allow us to train a model that is over 5 F1 points better than the exact same model architecture without the techniques we describe. We argue that for such information-extraction problems, designing abstractions that carefully manage the training data is at least as important as choosing a good model architecture.<\/jats:p>","DOI":"10.14778\/3447689.3447703","type":"journal-article","created":{"date-parts":[[2021,4,12]],"date-time":"2021-04-12T16:20:06Z","timestamp":1618244406000},"page":"997-1005","source":"Crossref","is-referenced-by-count":9,"title":["Glean"],"prefix":"10.14778","volume":"14","author":[{"given":"Sandeep","family":"Tata","sequence":"first","affiliation":[{"name":"Google"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Navneet","family":"Potti","sequence":"additional","affiliation":[{"name":"Google"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"James B.","family":"Wendt","sequence":"additional","affiliation":[{"name":"Google"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Lauro Beltr\u00e3o","family":"Costa","sequence":"additional","affiliation":[{"name":"Google"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Marc","family":"Najork","sequence":"additional","affiliation":[{"name":"Google"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Beliz","family":"Gunel","sequence":"additional","affiliation":[{"name":"Stanford University"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2021,4,12]]},"reference":[{"key":"e_1_2_1_1_1","doi-asserted-by":"publisher","DOI":"10.1145\/3132847.3133083"},{"key":"e_1_2_1_2_1","doi-asserted-by":"publisher","DOI":"10.5555\/1766091.1766143"},{"key":"e_1_2_1_3_1","doi-asserted-by":"publisher","DOI":"10.1145\/1008992.1009070"},{"key":"e_1_2_1_4_1","first-page":"182","article-title":"Template Mining for Information Extraction from Digital Documents","volume":"48","author":"Chowdhury Gobinda G.","year":"1999","journal-title":"Library Trends"},{"key":"e_1_2_1_5_1","doi-asserted-by":"publisher","DOI":"10.1145\/3148148"},{"key":"e_1_2_1_6_1","volume-title":"Proceedings of the 8th Biennial Conference on Innovative Data Systems Research.","author":"Deng Dong","year":"2017"},{"key":"e_1_2_1_7_1","volume-title":"Denk and Christian Reisswig","author":"Timo","year":"2019"},{"key":"e_1_2_1_8_1","volume-title":"Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 4171--4186","author":"Devlin Jacob","year":"2019"},{"key":"e_1_2_1_9_1","doi-asserted-by":"publisher","DOI":"10.14778\/1687627.1687749"},{"key":"e_1_2_1_10_1","volume-title":"LAMBERT: Layout-Aware language Modeling using BERT for information extraction. arXiv:2002.08087 [cs.CL]","author":"Garncarek \u0141ukasz","year":"2020"},{"key":"e_1_2_1_11_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/D18-1476"},{"key":"e_1_2_1_12_1","volume-title":"Learning to Rank for Information Retrieval","author":"Liu Tie-Yan"},{"key":"e_1_2_1_13_1","unstructured":"Yinhan Liu Myle Ott Naman Goyal Jingfei Du Mandar Joshi Danqi Chen Omer Levy Mike Lewis Luke Zettlemoyer and Veselin Stoyanov. 2019. RoBERTa: A Robustly Optimized BERT Pretraining Approach. arXiv:1907.11692 [cs.CL]  Yinhan Liu Myle Ott Naman Goyal Jingfei Du Mandar Joshi Danqi Chen Omer Levy Mike Lewis Luke Zettlemoyer and Veselin Stoyanov. 2019. RoBERTa: A Robustly Optimized BERT Pretraining Approach. arXiv:1907.11692 [cs.CL]"},{"key":"e_1_2_1_14_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2020.acl-main.580"},{"key":"e_1_2_1_15_1","unstructured":"Rachel Millner. 2008. Four regular expressions to check email addresses. https:\/\/www.wired.com\/2008\/08\/four-regular-expressions-to-check-email-addresses\/  Rachel Millner. 2008. Four regular expressions to check email addresses. https:\/\/www.wired.com\/2008\/08\/four-regular-expressions-to-check-email-addresses\/"},{"key":"e_1_2_1_16_1","doi-asserted-by":"publisher","DOI":"10.14778\/3157794.3157797"},{"key":"e_1_2_1_17_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2016.2577031"},{"key":"e_1_2_1_18_1","volume-title":"Proceedings of the 10th Conference on Innovative Data Systems Research.","author":"Rezig El Kindi","year":"2020"},{"key":"e_1_2_1_19_1","doi-asserted-by":"publisher","DOI":"10.1145\/3299869.3319867"},{"key":"e_1_2_1_20_1","doi-asserted-by":"publisher","DOI":"10.14778\/1929861.1929862"},{"key":"e_1_2_1_21_1","volume-title":"13th IAPR International Workshop on Document Analysis Systems - Short Papers Booklet. 21--22","author":"Walker Jake","year":"2018"},{"key":"e_1_2_1_22_1","doi-asserted-by":"publisher","DOI":"10.1145\/3183713.3183729"},{"key":"e_1_2_1_23_1","unstructured":"Yiheng Xu Minghao Li Lei Cui Shaohan Huang Furu Wei and Ming Zhou. 2019. LayoutLM: Pre-training of Text and Layout for Document Image Understanding. arXiv:1912.13318 [cs.CL]  Yiheng Xu Minghao Li Lei Cui Shaohan Huang Furu Wei and Ming Zhou. 2019. LayoutLM: Pre-training of Text and Layout for Document Image Understanding. arXiv:1912.13318 [cs.CL]"},{"key":"e_1_2_1_24_1","volume-title":"Cohen","author":"Yang Zhilin","year":"2017"},{"key":"e_1_2_1_25_1","doi-asserted-by":"publisher","DOI":"10.1145\/775152.775155"},{"key":"e_1_2_1_26_1","doi-asserted-by":"publisher","DOI":"10.1145\/3292500.3330773"},{"key":"e_1_2_1_27_1","doi-asserted-by":"publisher","DOI":"10.1145\/1150402.1150457"}],"container-title":["Proceedings of the VLDB Endowment"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.14778\/3447689.3447703","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2022,12,28]],"date-time":"2022-12-28T11:19:51Z","timestamp":1672226391000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.14778\/3447689.3447703"}},"subtitle":["structured extractions from templatic documents"],"short-title":[],"issued":{"date-parts":[[2021,2]]},"references-count":27,"journal-issue":{"issue":"6","published-print":{"date-parts":[[2021,2]]}},"alternative-id":["10.14778\/3447689.3447703"],"URL":"https:\/\/doi.org\/10.14778\/3447689.3447703","relation":{},"ISSN":["2150-8097"],"issn-type":[{"value":"2150-8097","type":"print"}],"subject":[],"published":{"date-parts":[[2021,2]]}}}