{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,12]],"date-time":"2026-05-12T14:15:53Z","timestamp":1778595353141,"version":"3.51.4"},"reference-count":39,"publisher":"MIT Press - Journals","license":[{"start":{"date-parts":[[2022,3,21]],"date-time":"2022-03-21T00:00:00Z","timestamp":1647820800000},"content-version":"vor","delay-in-days":79,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":["direct.mit.edu"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2022,3,18]]},"abstract":"<jats:title>Abstract<\/jats:title>\n               <jats:p>This paper presents a new task of predicting the coverage of a text document for relation extraction (RE): Does the document contain many relational tuples for a given entity? Coverage predictions are useful in selecting the best documents for knowledge base construction with large input corpora. To study this problem, we present a dataset of 31,366 diverse documents for 520 entities. We analyze the correlation of document coverage with features like length, entity mention frequency, Alexa rank, language complexity, and information retrieval scores. Each of these features has only moderate predictive power. We employ methods combining features with statistical models like TF-IDF and language models like BERT. The model combining features and BERT, HERB, achieves an F1 score of up to 46%. We demonstrate the utility of coverage predictions on two use cases: KB construction and claim refutation.<\/jats:p>","DOI":"10.1162\/tacl_a_00456","type":"journal-article","created":{"date-parts":[[2022,3,21]],"date-time":"2022-03-21T19:09:54Z","timestamp":1647889794000},"page":"207-223","update-policy":"https:\/\/doi.org\/10.1162\/mitpressjournals.corrections.policy","source":"Crossref","is-referenced-by-count":3,"title":["Predicting Document Coverage for Relation Extraction"],"prefix":"10.1162","volume":"10","author":[{"given":"Sneha","family":"Singhania","sequence":"first","affiliation":[{"name":"Max Planck Institute for Informatics, Germany. ssinghan@mpi-inf.mpg.de\""}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Simon","family":"Razniewski","sequence":"additional","affiliation":[{"name":"Max Planck Institute for Informatics, Germany. srazniew@mpi-inf.mpg.de\""}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Gerhard","family":"Weikum","sequence":"additional","affiliation":[{"name":"Max Planck Institute for Informatics, Germany. weikum@mpi-inf.mpg.de\""}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"281","published-online":{"date-parts":[[2022,3,18]]},"reference":[{"key":"2022033118511904300_bib1","doi-asserted-by":"publisher","first-page":"100661","DOI":"10.1016\/j.websem.2021.100661","article-title":"Negative statements considered useful","volume":"71","author":"Arnaout","year":"2021","journal-title":"Journal of Web Semantics"},{"key":"2022033118511904300_bib2","first-page":"993","article-title":"Latent dirichl et allocation","volume":"3","author":"Blei","year":"2003","journal-title":"Journal of Machine Learning Research"},{"key":"2022033118511904300_bib3","first-page":"542","article-title":"Seeing things from a different angle: Discovering diverse perspectives about claims","volume-title":"Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers)","author":"Chen","year":"2019"},{"key":"2022033118511904300_bib4","doi-asserted-by":"publisher","first-page":"985","DOI":"10.1145\/3331184.3331303","article-title":"Deeper text understanding for IR with contextual neural language modeling","volume-title":"Proceedings of the 42nd International ACM SIGIR Conference on Research and Development in Information Retrieval","author":"Dai","year":"2019"},{"key":"2022033118511904300_bib5","doi-asserted-by":"publisher","first-page":"66","DOI":"10.1007\/978-3-642-41335-3_5","article-title":"Completeness statements about RDF data sources and their use for query answering","volume-title":"The Semantic Web \u2013 ISWC 2013","author":"Darari","year":"2013"},{"key":"2022033118511904300_bib6","first-page":"4171","article-title":"BERT: Pre-training of deep bidirectional transformers for language understanding","volume-title":"Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers)","author":"Devlin","year":"2019"},{"issue":"5","key":"2022033118511904300_bib7","doi-asserted-by":"crossref","first-page":"378","DOI":"10.1037\/h0031619","article-title":"Measuring nominal scale agreement among many raters","volume":"76","author":"Fleiss","year":"1971","journal-title":"Psychological Bulletin"},{"key":"2022033118511904300_bib8","doi-asserted-by":"publisher","DOI":"10.1037\/h0031619","volume-title":"The Art of Readable Writing","author":"Flesch","year":"1949"},{"key":"2022033118511904300_bib9","doi-asserted-by":"publisher","first-page":"375","DOI":"10.1145\/3018661.3018739","article-title":"Predicting completeness in knowledge bases","volume-title":"Proceedings of the Tenth ACM International Conference on Web Search and Data Mining","author":"Gal\u00e1rraga","year":"2017"},{"key":"2022033118511904300_bib10","doi-asserted-by":"publisher","first-page":"708","DOI":"10.18653\/v1\/N18-1065","article-title":"Newsroom: A dataset of 1.3 million summaries with diverse extractive strategies","volume-title":"Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers)","author":"Grusky","year":"2018"},{"key":"2022033118511904300_bib11","first-page":"745","article-title":"More data, more relations, more context and more openness: A review and outlook for relation extraction","volume-title":"Proceedings of the 1st Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics and the 10th International Joint Conference on Natural Language Processing","author":"Han","year":"2020"},{"issue":"8","key":"2022033118511904300_bib12","doi-asserted-by":"publisher","first-page":"1735","DOI":"10.1162\/neco.1997.9.8.1735","article-title":"Long short-term memory","volume":"9","author":"Hochreiter","year":"1997","journal-title":"Neural Computation"},{"issue":"4","key":"2022033118511904300_bib13","doi-asserted-by":"publisher","first-page":"96","DOI":"10.1145\/3447772","article-title":"Knowledge graphs","volume":"64","author":"Hogan","year":"2021","journal-title":"ACM Computing Surveys"},{"key":"2022033118511904300_bib14","doi-asserted-by":"publisher","first-page":"200","DOI":"10.18653\/v1\/N18-3025","article-title":"Demand-weighted completeness prediction for a knowledge base","volume-title":"Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 3 (Industry Papers)","author":"Hopkinson","year":"2018"},{"issue":"4","key":"2022033118511904300_bib15","doi-asserted-by":"crossref","first-page":"21\u2013es","DOI":"10.1145\/1292609.1292611","article-title":"Towards a query optimizer for text-centric tasks","volume":"32","author":"Ipeirotis","year":"2007","journal-title":"ACM Transactions on Database Systems"},{"issue":"4","key":"2022033118511904300_bib16","doi-asserted-by":"publisher","first-page":"422","DOI":"10.1145\/582415.582418","article-title":"Cumulated gain-based evaluation of IR techniques","volume":"20","author":"J\u00e4rvelin","year":"2002","journal-title":"ACM Transactions of Information Systems"},{"key":"2022033118511904300_bib17","doi-asserted-by":"crossref","first-page":"2124","DOI":"10.18653\/v1\/P16-1200","article-title":"Neural relation extraction with selective attention over instances","volume-title":"Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)","author":"Lin","year":"2016"},{"key":"2022033118511904300_bib18","doi-asserted-by":"publisher","first-page":"225","DOI":"10.1007\/978-3-662-44851-9_15","article-title":"Optimal thresholding of classifiers to maximize F1 measure","volume-title":"Machine Learning and Knowledge Discovery in Databases","author":"Lipton","year":"2014"},{"key":"2022033118511904300_bib19","doi-asserted-by":"publisher","first-page":"453","DOI":"10.1007\/978-3-030-30793-6_26","article-title":"Non-parametric class completeness estimators for collaborative knowledge graphs\u2014the case of wikidata","volume-title":"The Semantic Web \u2013 ISWC 2019","author":"Luggen","year":"2019"},{"key":"2022033118511904300_bib20","doi-asserted-by":"publisher","first-page":"1003","DOI":"10.3115\/1690219.1690287","article-title":"Distant supervision for relation extraction without labeled data","volume-title":"Proceedings of the Joint Conference of the 47th Annual Meeting of the ACL and the 4th International Joint Conference on Natural Language Processing of the AFNLP","author":"Mintz","year":"2009"},{"issue":"5","key":"2022033118511904300_bib21","doi-asserted-by":"publisher","first-page":"103","DOI":"10.1145\/3191513","article-title":"Never-ending learning","volume":"61","author":"Mitchell","year":"2018","journal-title":"Communications of the ACM"},{"key":"2022033118511904300_bib22","doi-asserted-by":"publisher","first-page":"1009","DOI":"10.3115\/v1\/P14-1095","article-title":"Language-aware truth assessment of fact candidates","volume-title":"Proceedings of the 52nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)","author":"Nakashole","year":"2014"},{"key":"2022033118511904300_bib23","article-title":"Passage re-ranking with BERT","volume":"abs\/1901.04085","author":"Nogueira","year":"2020","journal-title":"ArXiv"},{"key":"2022033118511904300_bib24","doi-asserted-by":"publisher","first-page":"708","DOI":"10.18653\/v1\/2020.findings-emnlp.63","article-title":"Document ranking with a pretrained sequence-to-sequence model","volume-title":"Findings of the Association for Computational Linguistics: EMNLP 2020","author":"Nogueira","year":"2020"},{"key":"2022033118511904300_bib25","doi-asserted-by":"publisher","first-page":"1532","DOI":"10.3115\/v1\/D14-1162","article-title":"GloVe: Global vectors for word representation","volume-title":"Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP)","author":"Pennington","year":"2014"},{"issue":"140","key":"2022033118511904300_bib26","first-page":"1","article-title":"Exploring the limits of transfer learning with a unified text-to-text transformer","volume":"21","author":"Raffel","year":"2020","journal-title":"Journal of Machine Learning Research"},{"key":"2022033118511904300_bib27","doi-asserted-by":"publisher","first-page":"2931","DOI":"10.18653\/v1\/D17-1317","article-title":"Truth of varying shades: Analyzing language in fake news and political fact-checking","volume-title":"Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing","author":"Rashkin","year":"2017"},{"key":"2022033118511904300_bib28","doi-asserted-by":"publisher","first-page":"5771","DOI":"10.18653\/v1\/D19-1583","article-title":"Coverage of information extraction from sentences and paragraphs","volume-title":"Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP)","author":"Razniewski","year":"2019"},{"key":"2022033118511904300_bib29","doi-asserted-by":"publisher","first-page":"148","DOI":"10.1007\/978-3-642-15939-8_10","article-title":"Modeling relations and their mentions without labeled text","volume-title":"Machine Learning and Knowledge Discovery in Databases","author":"Riedel","year":"2010"},{"key":"2022033118511904300_bib30","first-page":"109","article-title":"Okapi at TREC-3","volume-title":"Overview of the Third Text REtrieval Conference (TREC-3)","author":"Robertson","year":"1995"},{"key":"2022033118511904300_bib31","first-page":"2373","article-title":"A topic-aligned multilingual corpus of wikipedia articles for studying information asymmetry in low resource languages","volume-title":"Proceedings of the 12th Language Resources and Evaluation Conference","author":"Roy","year":"2020"},{"issue":"12","key":"2022033118511904300_bib32","first-page":"e26752","article-title":"The New York Times Annotated Corpus","volume":"6","author":"Sandhaus","year":"2008","journal-title":"Linguistic Data Consortium"},{"key":"2022033118511904300_bib33","doi-asserted-by":"crossref","first-page":"2895","DOI":"10.18653\/v1\/P19-1279","article-title":"Matching the blanks: Distributional similarity for relation learning","volume-title":"Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics","author":"Soares","year":"2019"},{"key":"2022033118511904300_bib34","doi-asserted-by":"publisher","first-page":"809","DOI":"10.18653\/v1\/N18-1074","article-title":"FEVER: A large-scale dataset for Fact Extraction and VERification","volume-title":"Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers)","author":"Thorne","year":"2018"},{"key":"2022033118511904300_bib35","doi-asserted-by":"publisher","first-page":"578","DOI":"10.1109\/ICDE.2019.00058","article-title":"MIDAS: Finding the right web sources to fill knowledge gaps","volume-title":"2019 IEEE 35th International Conference on Data Engineering (ICDE)","author":"Wang","year":"2019"},{"issue":"2\u20134","key":"2022033118511904300_bib36","doi-asserted-by":"publisher","first-page":"108","DOI":"10.1561\/1900000064","article-title":"Machine knowledge: Creation and curation of comprehensive knowledge bases","volume":"10","author":"Weikum","year":"2021","journal-title":"Foundations and Trends in Databases"},{"key":"2022033118511904300_bib37","first-page":"14149","article-title":"Entity structure within and throughout: Modeling mention dependencies for document-level relation extraction","volume-title":"Proceedings of the AAAI Conference on Artificial Intelligence","author":"Benfeng","year":"2021"},{"key":"2022033118511904300_bib38","doi-asserted-by":"publisher","first-page":"764","DOI":"10.18653\/v1\/P19-1074","article-title":"DocRED: A large-scale document-level relation extraction dataset","volume-title":"Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics","author":"Yao","year":"2019"},{"key":"2022033118511904300_bib39","doi-asserted-by":"publisher","first-page":"35","DOI":"10.18653\/v1\/D17-1004","article-title":"Position-aware attention and supervised data improve slot filling","volume-title":"Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing","author":"Zhang","year":"2017"}],"container-title":["Transactions of the Association for Computational Linguistics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/direct.mit.edu\/tacl\/article-pdf\/doi\/10.1162\/tacl_a_00456\/2002678\/tacl_a_00456.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/direct.mit.edu\/tacl\/article-pdf\/doi\/10.1162\/tacl_a_00456\/2002678\/tacl_a_00456.pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2022,3,31]],"date-time":"2022-03-31T23:52:40Z","timestamp":1648770760000},"score":1,"resource":{"primary":{"URL":"https:\/\/direct.mit.edu\/tacl\/article\/doi\/10.1162\/tacl_a_00456\/110010\/Predicting-Document-Coverage-for-Relation"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022]]},"references-count":39,"URL":"https:\/\/doi.org\/10.1162\/tacl_a_00456","relation":{},"ISSN":["2307-387X"],"issn-type":[{"value":"2307-387X","type":"electronic"}],"subject":[],"published-other":{"date-parts":[[2022]]},"published":{"date-parts":[[2022]]}}}