{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,10,1]],"date-time":"2025-10-01T15:43:29Z","timestamp":1759333409836},"reference-count":32,"publisher":"Oxford University Press (OUP)","issue":"7","content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":[],"published-print":{"date-parts":[[2005,4,1]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:p>Motivation: Structural genomics projects aim to solve a large number of protein structures with the ultimate objective of representing the entire protein space. The computational challenge is to identify and prioritize a small set of proteins with new, currently unknown, superfamilies or folds.<\/jats:p><jats:p>Results: We develop a method that assigns each protein a likelihood of it belonging to a new, yet undetermined, structural superfamily. The method relies on a variant of ProtoNet, an automatic hierarchical classification scheme of all protein sequences from SwissProt. Our results show that proteins that are remote from solved structures in the ProtoNet hierarchy are more likely to belong to new superfamilies. The results are validated against SCOP releases from recent years that account for about half of the solved structures known to date. We show that our new method and the representation of ProtoNet are superior in detecting new targets, compared to our previous method using ProtoMap classification. Furthermore, our method outperforms PSI-BLAST search in detecting potential new superfamilies.<\/jats:p><jats:p>Availability: An interactive tool implementing this method, named ProTarget, is available at http:\/\/www.protarget.cs.huji.ac.il. It can be used interactively to retrieve a list of candidate proteins for Structural genomics projects. Supplementary material is available at http:\/\/www.protarget.cs.huji.ac.il\/supplement<\/jats:p><jats:p>Contact: \u00a0michall@cc.huji.ac.il<\/jats:p>","DOI":"10.1093\/bioinformatics\/bti135","type":"journal-article","created":{"date-parts":[[2004,11,12]],"date-time":"2004-11-12T01:14:59Z","timestamp":1100222099000},"page":"1020-1027","source":"Crossref","is-referenced-by-count":8,"title":["Predicting fold novelty based on ProtoNet hierarchical classification"],"prefix":"10.1093","volume":"21","author":[{"given":"Ilona","family":"Kifer","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Ori","family":"Sasson","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Michal","family":"Linial","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"286","published-online":{"date-parts":[[2004,11,11]]},"reference":[{"key":"2023013107271574900_B1","doi-asserted-by":"crossref","unstructured":"Altschul, S.F., Madden, T.L., Schaffer, A.A., Zhang, J., Zhang, Z., Miller, W., Lipman, D.J. 1997Gapped BLAST and PSI-BLAST: a new generation of protein database search programs. Nucleic Acids Res.253389\u20133402","DOI":"10.1093\/nar\/25.17.3389"},{"key":"2023013107271574900_B2","doi-asserted-by":"crossref","unstructured":"Bray, J.E., Marsden, R.L., Rison, S.C., Savchenko, A., Edwards, A.M., Thornton, J.M., Orengo, C.A. 2004A practical and robust sequence search strategy for structural genomics target selection. Bioinformatics202288\u20132295","DOI":"10.1093\/bioinformatics\/bth240"},{"key":"2023013107271574900_B3","doi-asserted-by":"crossref","unstructured":"Brenner, S.E. 2000Target selection for structural genomics. Nat. Struct. Biol.7(Suppl),967\u2013969","DOI":"10.1038\/80747"},{"key":"2023013107271574900_B4","doi-asserted-by":"crossref","unstructured":"Brenner, S.E., Chothia, C., Hubbard, T.J. 1998Assessing sequence comparison methods with reliable structurally identified distant evolutionary relationships. Proc. Natl Acad. Sci. USA956073\u20136078","DOI":"10.1073\/pnas.95.11.6073"},{"key":"2023013107271574900_B5","unstructured":"Brenner, S.E. and Levitt, M. 2000Expectations from structural genomics. Protein Sci.9197\u2013200"},{"key":"2023013107271574900_B6","unstructured":"Burley, S.K. and Bonanno, J.B. 2002Structural genomics of proteins from conserved biochemical pathways and processes. Curr. Opin. Struct. Biol.12383\u2013391"},{"key":"2023013107271574900_B7","unstructured":"Carter, P., Liu, J., Rost, B. 2003PEP: predictions for entire proteomes. Nucleic Acids Res.31410\u2013413"},{"key":"2023013107271574900_B8","doi-asserted-by":"crossref","unstructured":"Chance, M.R., Bresnick, A.R., Burley, S.K., Jiang, J.S., Lima, C.D., Sali, A., Almo, S.C., Bonanno, J.B., Buglino, J.A., Boulton, S., et al. 2002Structural genomics: a pipeline for providing structures for the biologist. Protein Sci.11723\u2013738","DOI":"10.1110\/ps.4570102"},{"key":"2023013107271574900_B9","doi-asserted-by":"crossref","unstructured":"Elofsson, A. and Sonnhammer, E.L. 1999A comparison of sequence and structure protein domain families as a basis for structural genomics. Bioinformatics15480\u2013500","DOI":"10.1093\/bioinformatics\/15.6.480"},{"key":"2023013107271574900_B10","unstructured":"Eswaramoorthy, S., Gerchman, S., Graziano, V., Kycia, H., Studier, F.W., Swaminathan, S. 2003Structure of a yeast hypothetical protein selected by a structural genomics approach. Acta Crystallogr. D Biol. Crystallogr.59127\u2013135"},{"key":"2023013107271574900_B11","doi-asserted-by":"crossref","unstructured":"Goldsmith-Fischman, S. and Honig, B. 2003Structural genomics: computational methods for structure analysis. Protein Sci.121813\u20131821","DOI":"10.1110\/ps.0242903"},{"key":"2023013107271574900_B12","doi-asserted-by":"crossref","unstructured":"Gough, J. and Chothia, C. 2002SUPERFAMILY: HMMs representing all proteins of known structure. SCOP sequence searches, alignments and genome assignments. Nucleic Acids Res.30268\u2013272","DOI":"10.1093\/nar\/30.1.268"},{"key":"2023013107271574900_B13","unstructured":"Joachims, T. 1999Making Large-Scale SVM Learning Practical. , Cambridge, MA, USA MIT Press"},{"key":"2023013107271574900_B14","unstructured":"Jones, D.T. 1999Protein secondary structure prediction based on position-specific scoring matrices. J. Mol. Biol.292195\u2013202"},{"key":"2023013107271574900_B15","doi-asserted-by":"crossref","unstructured":"Karplus, K., Barrett, C., Hughey, R. 1998Hidden Markov models for detecting remote protein homologies. Bioinformatics14846\u2013856","DOI":"10.1093\/bioinformatics\/14.10.846"},{"key":"2023013107271574900_B16","unstructured":"Linial, M. and Yona, G. 2000Methodologies for target selection in structural genomics. Progr. Biophys. Mol. Biol.73297\u2013320"},{"key":"2023013107271574900_B17","unstructured":"Liu, J. and Rost, B. 2002Target space for structural genomics revisited. Bioinformatics18922\u2013933"},{"key":"2023013107271574900_B18","doi-asserted-by":"crossref","unstructured":"Liu, J. and Rost, B. 2003Domains, motifs and clusters in the protein universe. Curr. Opin. Chem. Biol.75\u201311","DOI":"10.1016\/S1367-5931(02)00003-0"},{"key":"2023013107271574900_B19","unstructured":"Lo Conte, L., Ailey, B., Hubbard, T.J., Brenner, S.E., Murzin, A.G., Chothia, C. 2000SCOP: a structural classification of proteins database. Nucleic Acids Res.28257\u2013259"},{"key":"2023013107271574900_B20","doi-asserted-by":"crossref","unstructured":"Portugaly, E., Kifer, I., Linial, M. 2002Selecting targets for structural determination by navigating in a graph of protein families. Bioinformatics18899\u2013907","DOI":"10.1093\/bioinformatics\/18.7.899"},{"key":"2023013107271574900_B21","doi-asserted-by":"crossref","unstructured":"Portugaly, E. and Linial, M. 2000Estimating the probability for a protein to have a new fold: a statistical computational model. Proc. Natl Acad. Sci. USA975161\u20135166","DOI":"10.1073\/pnas.090559497"},{"key":"2023013107271574900_B22","unstructured":"Sali, A. 1998100,000 protein structures for the biologist. Nat. Struct. Biol.51029\u20131032"},{"key":"2023013107271574900_B23","doi-asserted-by":"crossref","unstructured":"Sanchez, R., Pieper, U., Melo, F., Eswar, N., Marti-Renom, M.A., Madhusudhan, M.S., Mirkovic, N., Sali, A. 2000Protein structure modeling for structural genomics. Nat. Struct. Biol.7 Suppl.986\u2013990","DOI":"10.1038\/80776"},{"key":"2023013107271574900_B24","doi-asserted-by":"crossref","unstructured":"Sasson, O., Linial, N., Linial, M. 2002The metric space of proteins-comparative study of clustering algorithms. Bioinformatics18S14\u201321","DOI":"10.1093\/bioinformatics\/18.suppl_1.S14"},{"key":"2023013107271574900_B25","unstructured":"Sasson, O., Vaaknin, A., Fleischer, H., Portugaly, E., Bilu, Y., Linial, N., Linial, M. 2003ProtoNet: hierarchical classification of the protein space. Nucleic Acids Res.31348\u2013352"},{"key":"2023013107271574900_B26","doi-asserted-by":"crossref","unstructured":"Shachar, O. and Linial, M. 2004A robust method to detect structural and functional remote homologues. Proteins57531\u2013538","DOI":"10.1002\/prot.20235"},{"key":"2023013107271574900_B27","unstructured":"Vitkup, D., Melamud, E., Moult, J., Sander, C. 2001Completeness in structural genomics. Nat. Struct. Biol.8559\u2013566"},{"key":"2023013107271574900_B28","unstructured":"Westbrook, J., Feng, Z., Chen, L., Yang, H., Berman, H.M. 2003The Protein Data Bank and structural genomics. Nucleic Acids Res.31489\u2013491"},{"key":"2023013107271574900_B29","doi-asserted-by":"crossref","unstructured":"Yona, G., Linial, N., Linial, M. 2000ProtoMap: automatic classification of protein sequences and hierarchy of protein families. Nucleic Acids Res.2849\u201355","DOI":"10.1093\/nar\/28.1.49"},{"key":"2023013107271574900_B30","doi-asserted-by":"crossref","unstructured":"Zarembinski, T.I., Hung, L.W., Mueller-Dieckmann, H.J., Kim, K.K., Yokota, H., Kim, R., Kim, S.H. 1998Structure-based assignment of the biochemical function of a hypothetical protein: a test case of structural genomics. Proc. Natl Acad. Sci. USA9515189\u201315193","DOI":"10.2210\/pdb1mjh\/pdb"},{"key":"2023013107271574900_B31","doi-asserted-by":"crossref","unstructured":"Zavaljevski, N., Stevens, F.J., Reifman, J. 2002Support vector machines with selective kernel scaling for protein classification and identification of key amino acid positions. Bioinformatics18689\u2013696","DOI":"10.1093\/bioinformatics\/18.5.689"},{"key":"2023013107271574900_B32","unstructured":"Zhang, C. and Kim, S.H. 2003Overview of structural genomics: from structure to function. Curr. Opin. Chem. Biol.,728\u201332"}],"container-title":["Bioinformatics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article-pdf\/21\/7\/1020\/48967016\/bioinformatics_21_7_1020.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article-pdf\/21\/7\/1020\/48967016\/bioinformatics_21_7_1020.pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2024,1,15]],"date-time":"2024-01-15T07:38:20Z","timestamp":1705304300000},"score":1,"resource":{"primary":{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article\/21\/7\/1020\/269027"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2004,11,11]]},"references-count":32,"journal-issue":{"issue":"7","published-print":{"date-parts":[[2005,4,1]]}},"URL":"https:\/\/doi.org\/10.1093\/bioinformatics\/bti135","relation":{},"ISSN":["1367-4811","1367-4803"],"issn-type":[{"value":"1367-4811","type":"electronic"},{"value":"1367-4803","type":"print"}],"subject":[],"published-other":{"date-parts":[[2005,4,1]]},"published":{"date-parts":[[2004,11,11]]}}}