{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,16]],"date-time":"2026-07-16T18:56:59Z","timestamp":1784228219298,"version":"3.55.0"},"reference-count":38,"publisher":"Oxford University Press (OUP)","issue":"1","content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":[],"published-print":{"date-parts":[[2005,1,1]]},"abstract":"<jats:title>Abstract<\/jats:title>\n               <jats:p>Motivation: With protein sequences entering into databanks at an explosive pace, the early determination of the family or subfamily class for a newly found enzyme molecule becomes important because this is directly related to the detailed information about which specific target it acts on, as well as to its catalytic process and biological function. Unfortunately, it is both time-consuming and costly to do so by experiments alone. In a previous study, the covariant-discriminant algorithm was introduced to identify the 16 subfamily classes of oxidoreductases. Although the results were quite encouraging, the entire prediction process was based on the amino acid composition alone without including any sequence-order information. Therefore, it is worthy of further investigation.<\/jats:p>\n               <jats:p>Results: To incorporate the sequence-order effects into the predictor, the \u2018amphiphilic pseudo amino acid composition\u2019 is introduced to represent the statistical sample of a protein. The novel representation contains 20 + 2\u03bb discrete numbers: the first 20 numbers are the components of the conventional amino acid composition; the next 2\u03bb numbers are a set of correlation factors that reflect different hydrophobicity and hydrophilicity distribution patterns along a protein chain. Based on such a concept and formulation scheme, a new predictor is developed. It is shown by the self-consistency test, jackknife test and independent dataset tests that the success rates obtained by the new predictor are all significantly higher than those by the previous predictors. The significant enhancement in success rates also implies that the distribution of hydrophobicity and hydrophilicity of the amino acid residues along a protein chain plays a very important role to its structure and function.<\/jats:p>\n               <jats:p>Contact: \u00a0kchou@san.rr.com<\/jats:p>","DOI":"10.1093\/bioinformatics\/bth466","type":"journal-article","created":{"date-parts":[[2004,8,13]],"date-time":"2004-08-13T00:15:36Z","timestamp":1092356136000},"page":"10-19","source":"Crossref","is-referenced-by-count":837,"title":["Using amphiphilic pseudo amino acid composition to predict enzyme subfamily classes"],"prefix":"10.1093","volume":"21","author":[{"given":"Kuo-Chen","family":"Chou","sequence":"first","affiliation":[{"name":"Gordon Life Science Institute San Diego, CA 92130, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"286","published-online":{"date-parts":[[2004,8,12]]},"reference":[{"key":"2023013107190461100_B1","unstructured":"Bahar, I., Atilgan, A.R., Jernigan, R.L., Erman, B. 1997Understanding the recognition of protein structural classes by amino acid composition. Proteins29172\u2013185"},{"key":"2023013107190461100_B2","unstructured":"Bairoch, A. and Apweiler, R. 2000The SWISS-PROT protein sequence data bank and its supplement TrEMBL. Nucleic Acids Res.2531\u201336"},{"key":"2023013107190461100_B3","unstructured":"Cai, Y.D., Li, Y.X., Chou, K.C. 2000Using neural networks for prediction of domain structural classes. Biochim. Biophys. Acta14761\u20132"},{"key":"2023013107190461100_B4","unstructured":"Cedano, J., Aloy, P., P'erez-Pons, J.A., Querol, E. 1997Relation between amino acid composition and cellular location of proteins. J. Mol. Biol.266594\u2013600"},{"key":"2023013107190461100_B5","unstructured":"Chandonia, J.M. and Karplus, M. 1995Neural networks for secondary structure and structural class prediction. Protein Sci.4275\u2013285"},{"key":"2023013107190461100_B6","doi-asserted-by":"crossref","unstructured":"Chou, J.J. and Zhang, C.T. 1993A joint prediction of the folding types of 1490 human proteins from their genetic codons. J. Theoret. Biol.161251\u2013262","DOI":"10.1006\/jtbi.1993.1053"},{"key":"2023013107190461100_B7","doi-asserted-by":"crossref","unstructured":"Chou, K.C. 1995A novel approach to predicting protein structural classes in a (20\u20131)-D amino acid composition space. Proteins21319\u2013344","DOI":"10.1002\/prot.340210406"},{"key":"2023013107190461100_B8","unstructured":"Chou, K.C. 1999A key driving force in determination of protein structural classes. Biochem. Biophys. Res. Commun.264216\u2013224"},{"key":"2023013107190461100_B9","doi-asserted-by":"crossref","unstructured":"Chou, K.C. 2000Review: prediction of protein structural classes and subcellular locations. Curr. Prot. Peptide Sci.1171\u2013208","DOI":"10.2174\/1389203003381379"},{"key":"2023013107190461100_B10","doi-asserted-by":"crossref","unstructured":"Chou, K.C. 2001Prediction of protein cellular attributes using pseudo-amino-acid-composition. Proteins43246\u2013255 [Erratum (2001) Proteins, 44, 60.]","DOI":"10.1002\/prot.1072"},{"key":"2023013107190461100_B11","doi-asserted-by":"crossref","unstructured":"Chou, K.C. and Cai, Y.D. 2004Predicting enzyme family class in a hybridization space. Protein Sci.13 in press","DOI":"10.1110\/ps.04981104"},{"key":"2023013107190461100_B12","unstructured":"Chou, K.C. and Elrod, D.W. 1999Protein subcellular location prediction. Protein Eng.12107\u2013118"},{"key":"2023013107190461100_B13","unstructured":"Chou, K.C. and Elrod, D.W. 2003Prediction of enzyme family classes. J. Proteome Res.2183\u2013190"},{"key":"2023013107190461100_B14","unstructured":"Chou, K.C. and Maggiora, G.M. 1998Domain structural class prediction. Protein Eng.11523\u2013538"},{"key":"2023013107190461100_B15","doi-asserted-by":"crossref","unstructured":"Chou, K.C. and Zhang, C.T. 1994Predicting protein folding types by distance functions that make allowances for amino acid interactions. J. Biol. Chem.26922014\u201322020","DOI":"10.1016\/S0021-9258(17)31748-9"},{"key":"2023013107190461100_B16","unstructured":"Chou, K.C. and Zhang, C.T. 1995Review: Prediction of protein structural classes. Crit. Rev. Biochem. Mol. Biol.30275\u2013349"},{"key":"2023013107190461100_B17","doi-asserted-by":"crossref","unstructured":"Chou, K.C., Zhang, C.T., Maggiora, M.G. 1997Disposition of amphiphilic helices in heteropolar environments. Proteins2899\u2013108","DOI":"10.1002\/(SICI)1097-0134(199705)28:1<99::AID-PROT10>3.0.CO;2-C"},{"key":"2023013107190461100_B18","doi-asserted-by":"crossref","unstructured":"Chou, P.Y. 1989Prediction of protein structural classes from amino acid composition. In Fasman, G.D. (Ed.). Prediction of Protein Structure and the Principles of Protein Conformation , New York  Plenum Press,  pp. 549\u2013586","DOI":"10.1007\/978-1-4613-1571-1_12"},{"key":"2023013107190461100_B19","unstructured":"Deleage, G. and Roux, B. 1987An algorithm for protein secondary structure prediction based on class prediction. Protein Eng.1289\u2013294"},{"key":"2023013107190461100_B20","doi-asserted-by":"crossref","unstructured":"Hopp, T.P. and Woods, K.R. 1981Prediction of protein antigenic determinants from amino acid sequences. Proc. Natl Acad. Sci. USA783824\u20133828","DOI":"10.1073\/pnas.78.6.3824"},{"key":"2023013107190461100_B21","unstructured":"Hua, S. and Sun, Z. 2001Support vector machine approach for protein subcellular localization prediction. Bioinformatics17721\u2013728"},{"key":"2023013107190461100_B22","unstructured":"Klein, P. 1986Prediction of protein structural class by discriminant analysis. Biochim. Biophys. Acta874205\u2013215"},{"key":"2023013107190461100_B23","unstructured":"Klein, P. and Delisi, C. 1986Prediction of protein structural class from amino acid sequence. Biopolymers251659\u20131672"},{"key":"2023013107190461100_B24","unstructured":"Kneller, D.G., Cohen, F.E., Langridge, R. 1990Improvements in protein secondary structure prediction by an enhanced neural network. J. Mol. Biol.214171\u2013182"},{"key":"2023013107190461100_B25","unstructured":"Liu, W. and Chou, K.C. 1998Prediction of protein structural classes by modified Mahalanobis discriminant algorithm. J. Protein Chem.17209\u2013217"},{"key":"2023013107190461100_B26","unstructured":"Mahalanobis, P.C. 1936On the generalized distance in statistics. Proc. Natl Inst. Sci. India249\u201355"},{"key":"2023013107190461100_B27","unstructured":"Mardia, K.V., Kent, J.T., Bibby, J.M. 1979Multivariate Analysis: Chapter 11, Discriminant analysis; Chapter 12, Multivariate analysis of variance; Chapter 13, Cluster analysis.  , London  Academic Press322\u2013381"},{"key":"2023013107190461100_B28","doi-asserted-by":"crossref","unstructured":"Metfessel, B.A., Saurugger, P.N., Connelly, D.P., Rich, S.T. 1993Cross-validation of protein structural class prediction using statistical clustering and neural networks. Protein Sci.21171\u20131182","DOI":"10.1002\/pro.5560020712"},{"key":"2023013107190461100_B29","doi-asserted-by":"crossref","unstructured":"Nakai, K. 2000Protein sorting signals and prediction of subcellular localization. Adv. Protein Chem.54277\u2013344","DOI":"10.1016\/S0065-3233(00)54009-1"},{"key":"2023013107190461100_B30","doi-asserted-by":"crossref","unstructured":"Nakai, K. and Kanehisa, M. 1991Expert system for predicting protein localization sites in Gram-negative bacteria. Proteins1195\u2013110","DOI":"10.1002\/prot.340110203"},{"key":"2023013107190461100_B31","unstructured":"Nakashima, H. and Nishikawa, K. 1994Discrimination of intracellular and extracellular proteins using amino acid composition and residue-pair frequencies. J. Mol. Biol.23854\u201361"},{"key":"2023013107190461100_B32","unstructured":"Nakashima, H., Nishikawa, K., Ooi, T. 1986The folding type of a protein is relevant to the amino acid composition. J. Biochem.99152\u2013162"},{"key":"2023013107190461100_B33","unstructured":"Pillai, K.C.S. 1985Mahalanobis D2. In Kotz, S. and Johnson, N.L. (Eds.). Encyclopedia of Statistical Sciences , New York  John Wiley and Sons Vol 5,  pp. 176\u2013181"},{"key":"2023013107190461100_B34","unstructured":"Tanford, C. 1962Contribution of hydrophobic interactions to the stability of the globular conformation of proteins. J. Am. Chem. Soc.844240\u20134274"},{"key":"2023013107190461100_B35","unstructured":"Webb, E.C. Enzyme Nomenclature1992, San Diego CA  Academic Press"},{"key":"2023013107190461100_B36","doi-asserted-by":"crossref","unstructured":"Zhou, G.P. 1998An intriguing controversy over protein structural class prediction. J. Protein Chem.17,  pp. 729\u2013738","DOI":"10.1023\/A:1020713915365"},{"key":"2023013107190461100_B37","unstructured":"Zhou, G.P. and Assa-Munt, N. 2001Some insights into protein structural class prediction. Proteins4457\u201359"},{"key":"2023013107190461100_B38","unstructured":"Zhou, G.P. and Doctor, K. 2003Subcellular location prediction of apoptosis proteins. Proteins5044\u201348"}],"container-title":["Bioinformatics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article-pdf\/21\/1\/10\/48961885\/bioinformatics_21_1_10.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article-pdf\/21\/1\/10\/48961885\/bioinformatics_21_1_10.pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2023,1,31]],"date-time":"2023-01-31T09:54:51Z","timestamp":1675158891000},"score":1,"resource":{"primary":{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article\/21\/1\/10\/212492"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2004,8,12]]},"references-count":38,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2005,1,1]]}},"URL":"https:\/\/doi.org\/10.1093\/bioinformatics\/bth466","relation":{},"ISSN":["1367-4811","1367-4803"],"issn-type":[{"value":"1367-4811","type":"electronic"},{"value":"1367-4803","type":"print"}],"subject":[],"published-other":{"date-parts":[[2005,1,1]]},"published":{"date-parts":[[2004,8,12]]}}}