{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2022,10,20]],"date-time":"2022-10-20T12:27:15Z","timestamp":1666268835051},"reference-count":47,"publisher":"Cambridge University Press (CUP)","issue":"1","license":[{"start":{"date-parts":[[2011,3,9]],"date-time":"2011-03-09T00:00:00Z","timestamp":1299628800000},"content-version":"unspecified","delay-in-days":0,"URL":"https:\/\/www.cambridge.org\/core\/terms"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Nat. Lang. Eng."],"published-print":{"date-parts":[[2012,1]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:p>A vast amount of usable electronic data is in the form of unstructured text. The relation extraction task aims to identify useful information in text (e.g. <jats:italic>PersonW<\/jats:italic> works for <jats:italic>OrganisationX<\/jats:italic>, <jats:italic>GeneY<\/jats:italic> encodes <jats:italic>ProteinZ<\/jats:italic>) and recode it in a format such as a relational database or RDF triplestore that can be more effectively used for querying and automated reasoning. A number of resources have been developed for training and evaluating automatic systems for relation extraction in different domains. However, comparative evaluation is impeded by the fact that these corpora use different markup formats and notions of what constitutes a relation. We describe the preparation of corpora for comparative evaluation of relation extraction across domains based on the publicly available ACE 2004, ACE 2005 and BioInfer data sets. We present a common document type using token standoff and including detailed linguistic markup, while maintaining all information in the original annotation. The subsequent reannotation process normalises the two data sets so that they comply with a notion of relation that is intuitive, simple and informed by the semantic web. For the ACE data, we describe an automatic process that automatically converts many relations involving nested, nominal entity mentions to relations involving non-nested, named or pronominal entity mentions. For example, the first entity is mapped from \u2018one\u2019 to \u2018Amidu Berry\u2019 in the membership relation described in \u2018Amidu Berry, one half of PBS\u2019. Moreover, we describe a comparably reannotated version of the BioInfer corpus that flattens nested relations, maps part-whole to part-part relations and maps n-ary to binary relations. Finally, we summarise experiments that compare approaches to generic relation extraction, a knowledge discovery task that uses minimally supervised techniques to achieve maximally portable extractors. These experiments illustrate the utility of the corpora.<jats:sup>1<\/jats:sup><\/jats:p>","DOI":"10.1017\/s1351324911000106","type":"journal-article","created":{"date-parts":[[2011,3,9]],"date-time":"2011-03-09T05:44:14Z","timestamp":1299649454000},"page":"21-59","source":"Crossref","is-referenced-by-count":10,"title":["Datasets for generic relation extraction"],"prefix":"10.1017","volume":"18","author":[{"given":"B.","family":"HACHEY","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"C.","family":"GROVER","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"R.","family":"TOBIN","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"56","published-online":{"date-parts":[[2011,3,9]]},"reference":[{"key":"S1351324911000106_ref22","unstructured":"Hachey B. 2009 b. Towards Generic Relation Extraction. Ph.D. thesis, University of Edinburgh."},{"key":"S1351324911000106_ref41","doi-asserted-by":"publisher","DOI":"10.1016\/j.jbi.2003.10.001"},{"key":"S1351324911000106_ref26","doi-asserted-by":"publisher","DOI":"10.1186\/1747-5333-2-4"},{"key":"S1351324911000106_ref20","first-page":"19","volume-title":"Proceedings of the EACL Workshop on Multi-dimensional Markup in Natural Language Processing","author":"Grover","year":"2006"},{"key":"S1351324911000106_ref45","doi-asserted-by":"publisher","DOI":"10.1145\/1132956.1132957"},{"key":"S1351324911000106_ref44","doi-asserted-by":"publisher","DOI":"10.1353\/pbm.1986.0087"},{"key":"S1351324911000106_ref36","first-page":"201","volume-title":"Proceedings of the 1st International Natural Language Generation Conference","author":"Minnen","year":"2000"},{"key":"S1351324911000106_ref28","volume-title":"Annotation Guidelines for Entity Detection and Tracking (EDT)","year":"2004"},{"key":"S1351324911000106_ref14","first-page":"91","volume-title":"Proceedings of the 11th Meeting of the European Chapter of the Association for Computational Linguistics","author":"Curran","year":"2003"},{"key":"S1351324911000106_ref33","doi-asserted-by":"publisher","DOI":"10.1017\/S1351324901002765"},{"key":"S1351324911000106_ref25","first-page":"45","volume-title":"Proceedings of the 3rd International Symposium on Semantic Mining in Biomedicine","author":"Heimonen","year":"2008"},{"key":"S1351324911000106_ref24","unstructured":"Hasegawa T. , Sekine S. and Grishman R. 2005. Unsupervised paraphrase acquisition via relation discovery. Technical Report 05-012, Proteus Project, Computer Science Department, New York University."},{"key":"S1351324911000106_ref11","first-page":"38","volume-title":"Proceedings of the ACL-ISMB Workshop on Linking Biological Literature, Ontologies and Databases: Mining Biological Semantics","author":"Cohen","year":"2005"},{"key":"S1351324911000106_ref23","first-page":"415","volume-title":"Proceedings of the 42nd Annual Meeting of Association of Computational Linguistics","author":"Hasegawa","year":"2004"},{"key":"S1351324911000106_ref15","first-page":"837","volume-title":"Proceedings of the 4th International Conference on Language Resources and Evaluation","author":"Doddington","year":"2004"},{"key":"S1351324911000106_ref47","doi-asserted-by":"publisher","DOI":"10.1186\/1471-2105-5-155"},{"key":"S1351324911000106_ref27","doi-asserted-by":"publisher","DOI":"10.1080\/01638539809545028"},{"key":"S1351324911000106_ref7","doi-asserted-by":"publisher","DOI":"10.1007\/10704656_11"},{"key":"S1351324911000106_ref19","doi-asserted-by":"publisher","DOI":"10.1073\/pnas.0307752101"},{"key":"S1351324911000106_ref43","first-page":"73","volume-title":"Proceedings of the 25th International ACM SIGIR Conference on Research and Development in Information Retrieval","author":"Smith","year":"2002"},{"key":"S1351324911000106_ref39","doi-asserted-by":"publisher","DOI":"10.1186\/1471-2105-9-S3-S6"},{"key":"S1351324911000106_ref5","first-page":"1","article-title":"Linked data \u2013 the story so far","volume":"5","author":"Bizer","year":"2009","journal-title":"International Journal on Semantic Web and Information Systems"},{"key":"S1351324911000106_ref34","first-page":"313","article-title":"Building a large annotated corpus of English: the Penn treebank","volume":"19","author":"Marcus","year":"1993","journal-title":"Computational Linguistics"},{"key":"S1351324911000106_ref1","doi-asserted-by":"publisher","DOI":"10.1145\/336597.336644"},{"key":"S1351324911000106_ref29","volume-title":"Annotation Guidelines for Relation Detection and Characterization (RDC)","year":"2004"},{"key":"S1351324911000106_ref17","first-page":"247","volume-title":"Recent Advances in Natural Language Processing III","author":"Filatova","year":"2003"},{"key":"S1351324911000106_ref38","first-page":"99","volume-title":"New Directions in Question Answering","author":"Pustejovsky","year":"2004"},{"key":"S1351324911000106_ref37","volume-title":"ACE 2004 Multilingual Training Corpus","author":"Mitchell","year":"2005"},{"key":"S1351324911000106_ref30","volume-title":"ACE (Automatic Content Extraction) English Annotation Guidelines for Entities","year":"2005"},{"key":"S1351324911000106_ref3","doi-asserted-by":"publisher","DOI":"10.1145\/1526709.1526793"},{"key":"S1351324911000106_ref4","doi-asserted-by":"publisher","DOI":"10.1137\/1037127"},{"key":"S1351324911000106_ref6","first-page":"993","article-title":"Latent Dirichlet allocation","volume":"3","author":"Blei","year":"2003","journal-title":"Journal of Machine Learning Research"},{"key":"S1351324911000106_ref9","unstructured":"Byrne K. 2009. Populating the Semantic Web \u2013 Combining Text and Relational Databases as RDF Graphs. PhD thesis, University of Edinburgh."},{"key":"S1351324911000106_ref8","doi-asserted-by":"publisher","DOI":"10.1016\/j.artmed.2004.07.016"},{"key":"S1351324911000106_ref10","volume-title":"Proceedings of the 7th Message Understanding Conference","author":"Chinchor","year":"1998"},{"key":"S1351324911000106_ref32","first-page":"317","volume-title":"Proceedings of the LREC Workshop Evaluation of Parsing Systems","author":"Lin","year":"1998"},{"key":"S1351324911000106_ref40","doi-asserted-by":"publisher","DOI":"10.1186\/1471-2105-8-50"},{"key":"S1351324911000106_ref12","doi-asserted-by":"publisher","DOI":"10.1186\/1471-2105-7-S3-S5"},{"key":"S1351324911000106_ref46","volume-title":"ACE 2005 Multilingual Training Corpus","author":"Walker","year":"2006"},{"key":"S1351324911000106_ref21","first-page":"420","volume-title":"Proceedings of the Conference on Empirical Methods in Natural Language Processing","author":"Hachey","year":"2009"},{"key":"S1351324911000106_ref18","unstructured":"Ginter F. , Pyysalo S. , Bj\u00f6rne J. , Heimonen J. , and Salakoski T. 2007. BioInfer relationship annotation manual. Technical Report 806, Turku Centre for Computer Science."},{"key":"S1351324911000106_ref35","first-page":"491","volume-title":"Proceedings of the 43rd Annual Meeting of the Association for Computational Linguistics","author":"McDonald","year":"2005"},{"key":"S1351324911000106_ref16","doi-asserted-by":"publisher","DOI":"10.1007\/BF02288367"},{"key":"S1351324911000106_ref2","volume-title":"Proceedings of the 7th Message Understanding Conference","author":"Aone","year":"1998"},{"key":"S1351324911000106_ref13","first-page":"260","volume-title":"Proceedings of the 17th International ACM SIGIR Conference on Research and Development in Information Retrieval","author":"Conrad","year":"1994"},{"key":"S1351324911000106_ref42","doi-asserted-by":"publisher","DOI":"10.3115\/1273073.1273167"},{"key":"S1351324911000106_ref31","volume-title":"ACE (Automatic Content Extraction) English Annotation Guidelines for Relations","year":"2005"}],"container-title":["Natural Language Engineering"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.cambridge.org\/core\/services\/aop-cambridge-core\/content\/view\/S1351324911000106","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2019,4,25]],"date-time":"2019-04-25T23:20:59Z","timestamp":1556234459000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.cambridge.org\/core\/product\/identifier\/S1351324911000106\/type\/journal_article"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2011,3,9]]},"references-count":47,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2012,1]]}},"alternative-id":["S1351324911000106"],"URL":"https:\/\/doi.org\/10.1017\/s1351324911000106","relation":{},"ISSN":["1351-3249","1469-8110"],"issn-type":[{"value":"1351-3249","type":"print"},{"value":"1469-8110","type":"electronic"}],"subject":[],"published":{"date-parts":[[2011,3,9]]}}}