{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,3,19]],"date-time":"2025-03-19T10:05:00Z","timestamp":1742378700243},"reference-count":23,"publisher":"Oxford University Press (OUP)","issue":"6","content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":[],"published-print":{"date-parts":[[2006,3,15]]},"abstract":"<jats:title>Abstract<\/jats:title>\n               <jats:p>Motivation: Many entity taggers and information extraction systems make use of lists of terms of entities such as people, places, genes or chemicals. These lists have traditionally been constructed manually. We show that distributional clustering methods which group words based on the contexts that they appear in, including neighboring words and syntactic relations extracted using a shallow parser, can be used to aid in the construction of term lists.<\/jats:p>\n               <jats:p>Results: Experiments on learning lists of terms and using them as part of a gene tagger on a corpus of abstracts from the scientific literature show that our automatically generated term lists significantly boost the precision of a state-of-the-art CRF-based gene tagger to a degree that is competitive with using hand curated lists and boosts recall to a degree that surpasses that of the hand-curated lists. Our results also show that these distributional clustering methods do not generate lists as helpful as those generated by supervised techniques, but that they can be used to complement supervised techniques so as to obtain better performance.<\/jats:p>\n               <jats:p>Availability: The code used in this paper is available from<\/jats:p>\n               <jats:p>Contact: \u00a0tsandler@seas.upenn.edu<\/jats:p>","DOI":"10.1093\/bioinformatics\/bti733","type":"journal-article","created":{"date-parts":[[2005,10,27]],"date-time":"2005-10-27T00:12:37Z","timestamp":1130371957000},"page":"651-657","source":"Crossref","is-referenced-by-count":6,"title":["Automatic term list generation for entity tagging"],"prefix":"10.1093","volume":"22","author":[{"given":"Ted","family":"Sandler","sequence":"first","affiliation":[{"name":"Department of Computer and Information Science, University of Pennsylvania \u00a0 3330 Walnut Street, Philadelphia, PA 19104, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Andrew I.","family":"Schein","sequence":"additional","affiliation":[{"name":"Department of Computer and Information Science, University of Pennsylvania \u00a0 3330 Walnut Street, Philadelphia, PA 19104, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Lyle H.","family":"Ungar","sequence":"additional","affiliation":[{"name":"Department of Computer and Information Science, University of Pennsylvania \u00a0 3330 Walnut Street, Philadelphia, PA 19104, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"286","published-online":{"date-parts":[[2005,10,25]]},"reference":[{"key":"2023012408500642100_b1","first-page":"194","article-title":"Nymble: a high-performance learning name-finder","author":"Bikel","year":"1997"},{"key":"2023012408500642100_b2","first-page":"467","article-title":"Class-based n-gram models of natural language","volume":"18","author":"Brown","year":"1992","journal-title":"Comput. Linguist."},{"key":"2023012408500642100_b3","first-page":"22","article-title":"Word association norms, mutual information, and lexicography","volume":"16","author":"Church","year":"1990","journal-title":"Comput. Linguist."},{"key":"2023012408500642100_b4","first-page":"89","article-title":"Information-theoretic co-clustering","author":"Dhillon","year":"2003"},{"key":"2023012408500642100_b5","first-page":"61","article-title":"Accurate methods for the statistics of surprise and coincidence","volume":"19","author":"Dunning","year":"1993","journal-title":"Comput. Linguist."},{"key":"2023012408500642100_b6","first-page":"262","article-title":"Trained named entity recognition using distributional clusters","author":"Freitag","year":"2004"},{"key":"2023012408500642100_b7","first-page":"577","article-title":"Boosted wrapper induction","author":"Freitag","year":"2000"},{"key":"2023012408500642100_b8","first-page":"539","article-title":"Automatic acquisition of hyponyms from large text corpora","author":"Hearst","year":"1992"},{"key":"2023012408500642100_b9","first-page":"268","article-title":"Noun classification from predicate-argument structures","author":"Hindle","year":"1990"},{"key":"2023012408500642100_b10","doi-asserted-by":"crossref","first-page":"247","DOI":"10.1016\/S1532-0464(03)00014-5","article-title":"Rutabaga by any other name: extracting biological names","volume":"35","author":"Hirschman","year":"2002","journal-title":"J. Biomed. Inform."},{"key":"2023012408500642100_b11","doi-asserted-by":"crossref","DOI":"10.1186\/1471-2105-6-S1-S1","article-title":"Overview of BioCreAtIvE: critical assessment of information extraction for biology","volume-title":"BMC Bioinformatics","author":"Hirschman","year":"2005"},{"key":"2023012408500642100_b12","first-page":"768","article-title":"Automatic retrieval and clustering of similar words","author":"Lin","year":"1998"},{"key":"2023012408500642100_b13","article-title":"Dependency-based evaluation of MINIPAR","author":"Lin","year":"1998"},{"key":"2023012408500642100_b14","first-page":"317","article-title":"Induction of semantic classes from natural language text","author":"Lin","year":"2001"},{"key":"2023012408500642100_b15","article-title":"Mallet: A machine learning for language toolkit","author":"McCallum","year":"2002"},{"key":"2023012408500642100_b16","article-title":"Identifying gene and protein mentions in text using conditional random fields","volume-title":"BMC Bioinformatics","author":"McDonald","year":"2004"},{"key":"2023012408500642100_b17","first-page":"337","article-title":"Name tagging with word clusters and discriminative training","author":"Miller","year":"2004"},{"key":"2023012408500642100_b18","first-page":"183","article-title":"Distributional clustering of english words","author":"Pereira","year":"1993"},{"key":"2023012408500642100_b19","first-page":"117","article-title":"A corpus-based approach for building semantic lexicons","author":"Riloff","year":"1997"},{"key":"2023012408500642100_b20","first-page":"1110","article-title":"Noun-phrase co-occurrence statistics for semiautomatic semantic lexicon construction","author":"Roark","year":"1998"},{"key":"2023012408500642100_b21","first-page":"9","article-title":"Tagging gene and protein names in full text articles","author":"Tanabe","year":"2002"},{"key":"2023012408500642100_b22","doi-asserted-by":"crossref","first-page":"611","DOI":"10.1142\/S0219720004000399","article-title":"Generation of a large gene\/protein lexicon by morphological pattern analysis","volume":"1","author":"Tanabe","year":"2004","journal-title":"J. Bioinformatics Comput. Biol."},{"key":"2023012408500642100_b23","doi-asserted-by":"crossref","first-page":"255","DOI":"10.1093\/nar\/gkh072","article-title":"Genew: the human gene nomenclature database, 2004 updates","volume":"32","author":"Wain","year":"2004","journal-title":"Nucleic Acids Res."}],"container-title":["Bioinformatics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article-pdf\/22\/6\/651\/48838571\/bioinformatics_22_6_651.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article-pdf\/22\/6\/651\/48838571\/bioinformatics_22_6_651.pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2023,1,24]],"date-time":"2023-01-24T09:17:33Z","timestamp":1674551853000},"score":1,"resource":{"primary":{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article\/22\/6\/651\/294268"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2005,10,25]]},"references-count":23,"journal-issue":{"issue":"6","published-print":{"date-parts":[[2006,3,15]]}},"URL":"https:\/\/doi.org\/10.1093\/bioinformatics\/bti733","relation":{},"ISSN":["1367-4811","1367-4803"],"issn-type":[{"value":"1367-4811","type":"electronic"},{"value":"1367-4803","type":"print"}],"subject":[],"published-other":{"date-parts":[[2006,3,15]]},"published":{"date-parts":[[2005,10,25]]}}}