{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,10,15]],"date-time":"2025-10-15T17:11:42Z","timestamp":1760548302820},"reference-count":22,"publisher":"Oxford University Press (OUP)","issue":"17","content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":[],"published-print":{"date-parts":[[2006,9,1]]},"abstract":"<jats:title>Abstract<\/jats:title>\n               <jats:p>Motivation: Attribute selection is a critical step in development of document classification systems. As a standard practice, words are stemmed and the most informative ones are used as attributes in classification. Owing to high complexity of biomedical terminology, general-purpose stemming algorithms are often conservative and could also remove informative stems. This can lead to accuracy reduction, especially when the number of labeled documents is small. To address this issue, we propose an algorithm that omits stemming and, instead, uses the most discriminative substrings as attributes.<\/jats:p>\n               <jats:p>Results: The approach was tested on five annotated sets of abstracts from iProLINK that report on the experimental evidence about five types of protein post-translational modifications. The experiments showed that Naive Bayes and support vector machine classifiers perform consistently better [with area under the ROC curve (AUC) accuracy in range 0.92\u20130.97] when using the proposed attribute selection than when using attributes obtained by the Porter stemmer algorithm (AUC in 0.86\u20130.93 range). The proposed approach is particularly useful when labeled datasets are small.<\/jats:p>\n               <jats:p>Contact: \u00a0vucetic@ist.temple.edu<\/jats:p>\n               <jats:p>Supplementary Information: The supplementary data are available from<\/jats:p>","DOI":"10.1093\/bioinformatics\/btl350","type":"journal-article","created":{"date-parts":[[2006,7,13]],"date-time":"2006-07-13T00:39:14Z","timestamp":1152751154000},"page":"2136-2142","source":"Crossref","is-referenced-by-count":15,"title":["Substring selection for biomedical document classification"],"prefix":"10.1093","volume":"22","author":[{"given":"Bo","family":"Han","sequence":"first","affiliation":[{"name":"Center for Information Science and Technology, Temple University 1 \u00a0 1 \u00a0 \u00a0 Philadelphia, PA 19122, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Zoran","family":"Obradovic","sequence":"additional","affiliation":[{"name":"Center for Information Science and Technology, Temple University 1 \u00a0 1 \u00a0 \u00a0 Philadelphia, PA 19122, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Zhang-Zhi","family":"Hu","sequence":"additional","affiliation":[{"name":"Department of Biochemistry and Molecular & Cellular Biology, Georgetown University Medical Center 2 \u00a0 2 \u00a0 \u00a0 Washington DC 20007, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Cathy H.","family":"Wu","sequence":"additional","affiliation":[{"name":"Department of Biochemistry and Molecular & Cellular Biology, Georgetown University Medical Center 2 \u00a0 2 \u00a0 \u00a0 Washington DC 20007, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Slobodan","family":"Vucetic","sequence":"additional","affiliation":[{"name":"Center for Information Science and Technology, Temple University 1 \u00a0 1 \u00a0 \u00a0 Philadelphia, PA 19122, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"286","published-online":{"date-parts":[[2006,6,23]]},"reference":[{"key":"2023012409140701500_b1","doi-asserted-by":"crossref","first-page":"600","DOI":"10.1093\/bioinformatics\/14.7.600","article-title":"Automatic extraction of keywords from scientific text: application to the knowledge domain of protein families","volume":"14","author":"Andrade","year":"1998","journal-title":"Bioinformatics"},{"key":"2023012409140701500_b2","doi-asserted-by":"crossref","first-page":"207","DOI":"10.1197\/jamia.M1641","article-title":"Text categorization models for retrieval of high quality articles in internal medicine","volume":"12","author":"Aphinyanaphongs","year":"2005","journal-title":"J. Am. Med. Inform. Assoc."},{"key":"2023012409140701500_b3","doi-asserted-by":"crossref","first-page":"573","DOI":"10.1137\/1037127","article-title":"Using linear algebra for intelligent information retrieval","volume":"37","author":"Berry","year":"1995","journal-title":"SIAM Rev."},{"key":"2023012409140701500_b4","doi-asserted-by":"crossref","first-page":"631","DOI":"10.1090\/S0025-5718-1969-0247736-4","article-title":"Rational Chebyshev approximations for the error function","volume":"22","author":"Cody","year":"1969","journal-title":"Math. Comp."},{"key":"2023012409140701500_b5","doi-asserted-by":"crossref","first-page":"i91","DOI":"10.1093\/bioinformatics\/btg1011","article-title":"Combining NLP and probabilistic categorization for document and term selection for Swiss-Prot medical annotation","volume":"19","author":"Dobrokhotov","year":"2003","journal-title":"Bioinformatics"},{"key":"2023012409140701500_b6","doi-asserted-by":"crossref","first-page":"95","DOI":"10.1145\/772862.772876","article-title":"Automatic scientific text classification using local patterns: KDD Cup 2002 (task 1)","volume":"4","author":"Ghanem","year":"2003","journal-title":"SIGKDD Explor. Newslett."},{"key":"2023012409140701500_b7","doi-asserted-by":"crossref","first-page":"409","DOI":"10.1016\/j.compbiolchem.2004.09.010","article-title":"iProLINK: an integrated protein resource for literature mining","volume":"28","author":"Hu","year":"2004","journal-title":"Comput. Biol. Chem."},{"key":"2023012409140701500_b8","doi-asserted-by":"crossref","first-page":"2759","DOI":"10.1093\/bioinformatics\/bti390","article-title":"Literature mining and database annotation of protein phosphorylation using a rule-based system","volume":"21","author":"Hu","year":"2005","journal-title":"Bioinformatics"},{"key":"2023012409140701500_b9","first-page":"137","article-title":"Text categorization with support vector machines: learning with many relevant features","author":"Joachims","year":"1998"},{"key":"2023012409140701500_b10","first-page":"41","article-title":"Making large-scale SVM learning practical","volume-title":"In Advances in Kernel Methods\u2014Support Vector Learning.","author":"Joachims","year":"1999"},{"key":"2023012409140701500_b11","doi-asserted-by":"crossref","first-page":"359","DOI":"10.1093\/bioinformatics\/17.4.359","article-title":"Mining literature for protein-protein interactions","volume":"17","author":"Marcotte","year":"2001","journal-title":"Bioinformatics"},{"key":"2023012409140701500_b12","first-page":"41","article-title":"A comparison of event models for Na\u00efve Bayes text classification","author":"McCallum","year":"1998"},{"key":"2023012409140701500_b13","first-page":"121","article-title":"Selecting text features for gene name classification: from documents to terms","author":"Nenadic","year":"2003"},{"key":"2023012409140701500_b14","doi-asserted-by":"crossref","first-page":"130","DOI":"10.1108\/eb046814","article-title":"An algorithm for sux stripping","volume":"14","author":"Porter","year":"1980","journal-title":"Program"},{"key":"2023012409140701500_b15","doi-asserted-by":"crossref","first-page":"90","DOI":"10.1145\/772862.772874","article-title":"Rulebased extraction of experimental evidence in the biomedical domain\u2014the KDD Cup 2002 (task 1)","volume":"4","author":"Regev","year":"2003","journal-title":"SIGKDD Explor. Newslett."},{"key":"2023012409140701500_b16","doi-asserted-by":"crossref","first-page":"S22","DOI":"10.1186\/1471-2105-6-S1-S22","article-title":"Mining protein function from text using term-based support vector machines","volume":"6","author":"Rice","year":"2005","journal-title":"BMC Bioinformatics"},{"key":"2023012409140701500_b17","first-page":"93","article-title":"A machine learning approach for the curation of biomedical literature\u2013KDD Cup 2002 (task 1)","volume":"4","author":"Shi","year":"2003","journal-title":"SIGKDD Explor. Newslett."},{"key":"2023012409140701500_b18","doi-asserted-by":"crossref","DOI":"10.1007\/978-1-4757-2440-0","volume-title":"The Nature of Statistical Learning Theory","author":"Vapnik","year":"1995"},{"key":"2023012409140701500_b19","first-page":"918","article-title":"Boosting Na\u00efve Bayesian learning on a large subset of MEDLINE","author":"Wilbur","year":"2000"},{"key":"2023012409140701500_b20","doi-asserted-by":"crossref","first-page":"345","DOI":"10.1093\/nar\/gkg040","article-title":"The Protein information resource","volume":"31","author":"Wu","year":"2003","journal-title":"Nucleic Acids Res."},{"key":"2023012409140701500_b21","doi-asserted-by":"crossref","first-page":"D187","DOI":"10.1093\/nar\/gkj161","article-title":"The Universal Protein Resource (UniProt): an expanding universe of protein information","volume":"34","author":"Wu","year":"2006","journal-title":"Nucleic Acids Res."},{"key":"2023012409140701500_b22","first-page":"412","article-title":"A comparative study on feature selection in text categorization","author":"Yang","year":"1997"}],"container-title":["Bioinformatics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article-pdf\/22\/17\/2136\/48840315\/bioinformatics_22_17_2136.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article-pdf\/22\/17\/2136\/48840315\/bioinformatics_22_17_2136.pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2023,1,24]],"date-time":"2023-01-24T09:53:14Z","timestamp":1674553994000},"score":1,"resource":{"primary":{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article\/22\/17\/2136\/273783"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2006,6,23]]},"references-count":22,"journal-issue":{"issue":"17","published-print":{"date-parts":[[2006,9,1]]}},"URL":"https:\/\/doi.org\/10.1093\/bioinformatics\/btl350","relation":{},"ISSN":["1367-4811","1367-4803"],"issn-type":[{"value":"1367-4811","type":"electronic"},{"value":"1367-4803","type":"print"}],"subject":[],"published-other":{"date-parts":[[2006,9,1]]},"published":{"date-parts":[[2006,6,23]]}}}