{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,8,21]],"date-time":"2025-08-21T17:06:24Z","timestamp":1755795984288},"reference-count":37,"publisher":"Oxford University Press (OUP)","issue":"22","content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":[],"published-print":{"date-parts":[[2005,11,15]]},"abstract":"<jats:title>Abstract<\/jats:title>\n               <jats:p>Motivation: Hidden Markov models (HMMs) calculate the probability that a sequence was generated by a given model. Log-odds scoring provides a context for evaluating this probability, by considering it in relation to a null hypothesis. We have found that using a reverse-sequence null model effectively removes biases owing to sequence length and composition and reduces the number of false positives in a database search.<\/jats:p>\n               <jats:p>Any scoring system is an arbitrary measure of the quality of database matches. Significance estimates of scores are essential, because they eliminate model- and method-dependent scaling factors, and because they quantify the importance of each match. Accurate computation of the significance of reverse-sequence null model scores presents a problem, because the scores do not fit the extreme-value (Gumbel) distribution commonly used to estimate HMM scores' significance.<\/jats:p>\n               <jats:p>Results: To get a better estimate of the significance of reverse-sequence null model scores, we derive a theoretical distribution based on the assumption of a Gumbel distribution for raw HMM scores and compare estimates based on this and other distribution families. We derive estimation methods for the parameters of the distributions based on maximum likelihood and on moment matching (least-squares fit for Student's t-distribution).<\/jats:p>\n               <jats:p>We evaluate the modeled distributions of scores, based on how well they fit the tail of the observed distribution for data not used in the fitting and on the effects of the improved E-values on our HMM-based fold-recognition methods.<\/jats:p>\n               <jats:p>The theoretical distribution provides some improvement in fitting the tail and in providing fewer false positives in the fold-recognition test. An ad hoc distribution based on assuming a stretched exponential tail does an even better job. The use of Student's t to model the distribution fits well in the middle of the distribution, but provides too heavy a tail. The moment-matching methods fit the tails better than maximum-likelihood methods.<\/jats:p>\n               <jats:p>Availability: Information on obtaining the SAM program suite (free for academic use), as well as a server interface, is available at and the open-source random sequence generator with varying compositional biases is available at<\/jats:p>\n               <jats:p>Contact: \u00a0karplus@soe.ucsc.edu<\/jats:p>","DOI":"10.1093\/bioinformatics\/bti629","type":"journal-article","created":{"date-parts":[[2005,8,26]],"date-time":"2005-08-26T00:24:37Z","timestamp":1125015877000},"page":"4107-4115","source":"Crossref","is-referenced-by-count":29,"title":["Calibrating <i>E<\/i>-values for hidden Markov models using reverse-sequence null models"],"prefix":"10.1093","volume":"21","author":[{"given":"Kevin","family":"Karplus","sequence":"first","affiliation":[{"name":"Department of Biomolecular Engineering, University of California 1 \u00a0 1 \u00a0 \u00a0 Santa Cruz, CA 95064, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Rachel","family":"Karchin","sequence":"additional","affiliation":[{"name":"Department of Biopharmaceutical Sciences, University of California 2 \u00a0 2 \u00a0 \u00a0 San Francisco, CA, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"George","family":"Shackelford","sequence":"additional","affiliation":[{"name":"Department of Biomolecular Engineering, University of California 1 \u00a0 1 \u00a0 \u00a0 Santa Cruz, CA 95064, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Richard","family":"Hughey","sequence":"additional","affiliation":[{"name":"Department of Biomolecular Engineering, University of California 1 \u00a0 1 \u00a0 \u00a0 Santa Cruz, CA 95064, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"286","published-online":{"date-parts":[[2005,8,25]]},"reference":[{"key":"2023061007105598600_b1","doi-asserted-by":"crossref","first-page":"555","DOI":"10.1016\/0022-2836(91)90193-A","article-title":"Amino acid substitution matrices from an information theoretic perspective","volume":"219","author":"Altschul","year":"1991","journal-title":"J. Mol. Biol."},{"key":"2023061007105598600_b2","doi-asserted-by":"crossref","first-page":"403","DOI":"10.1016\/S0022-2836(05)80360-2","article-title":"A basic local alignment search tool","volume":"215","author":"Altschul","year":"1990","journal-title":"J. Mol. Biol."},{"key":"2023061007105598600_b3","doi-asserted-by":"crossref","first-page":"3389","DOI":"10.1093\/nar\/25.17.3389","article-title":"Gapped BLAST and PSI-BLAST: a new generation of protein database search programs","volume":"25","author":"Altschul","year":"1997","journal-title":"Nucleic Acids Res."},{"key":"2023061007105598600_b4","doi-asserted-by":"crossref","first-page":"575","DOI":"10.1089\/106652702760138637","article-title":"Estimating and evaluating the statistics of gapped local-alignment scores","volume":"9","author":"Bailey","year":"2002","journal-title":"J. Comput. Biol."},{"key":"2023061007105598600_b5","doi-asserted-by":"crossref","first-page":"1059","DOI":"10.1073\/pnas.91.3.1059","article-title":"Hidden Markov models of biological primary sequence information","volume":"91","author":"Baldi","year":"1994","journal-title":"Proc. Natl Acad. Sci. USA"},{"key":"2023061007105598600_b6","first-page":"191","article-title":"Scoring hidden Markov models","volume":"13","author":"Barrett","year":"1997","journal-title":"Comput. Appl. Biosci."},{"key":"2023061007105598600_b7","article-title":"DCDFLIB: Library of routines for cumulative distribution functions, inverses, and other parameters (C and Fortran)","author":"Brown","year":"1997"},{"key":"2023061007105598600_b8","first-page":"53","article-title":"A generalized profile syntax for biomolecular sequence motifs and its function in automatic sequence interpretation","author":"Bucher","year":"1994"},{"key":"2023061007105598600_b9","doi-asserted-by":"crossref","first-page":"3","DOI":"10.1016\/S0097-8485(96)80003-9","article-title":"A flexible motif search technique based on generalized profiles","volume":"20","author":"Bucher","year":"1996","journal-title":"Comput. Chem."},{"key":"2023061007105598600_b10","doi-asserted-by":"crossref","first-page":"271","DOI":"10.1002\/1097-0134(20001115)41:3<271::AID-PROT10>3.0.CO;2-Z","article-title":"Bayesian probabilistic approach for predicting backbone structures in terms of protein blocks","volume":"41","author":"de Brevern","year":"2000","journal-title":"Proteins"},{"key":"2023061007105598600_b11","article-title":"Culling the PDB by resolution and sequence identity","author":"Dunbrack","year":"2001"},{"key":"2023061007105598600_b12","first-page":"114","article-title":"Multiple alignment using hidden Markov models","author":"Eddy","year":"1995"},{"key":"2023061007105598600_b13","doi-asserted-by":"crossref","first-page":"9","DOI":"10.1089\/cmb.1995.2.9","article-title":"Maximum discrimination hidden Markov models of sequence consensus","volume":"2","author":"Eddy","year":"1995","journal-title":"J. Comput. Biol."},{"key":"2023061007105598600_b14","doi-asserted-by":"crossref","first-page":"566","DOI":"10.1002\/prot.340230412","article-title":"Knowledge-based protein secondary structure assignment","volume":"23","author":"Frishman","year":"1995","journal-title":"Proteins"},{"key":"2023061007105598600_b15","volume-title":"Table of Integrals, Series, and Products","author":"Gradshteyn","year":"1965","edition":"fourth edn"},{"key":"2023061007105598600_b16","first-page":"397","article-title":"Meta-MEME: motif-based hidden Markov models of protein families","volume":"13","author":"Grundy","year":"1997","journal-title":"Comput. Appl. Biosci."},{"key":"2023061007105598600_b17","first-page":"792","article-title":"Protein modeling using hidden Markov models: analysis of globins","author":"Haussler","year":"1993"},{"key":"2023061007105598600_b18","first-page":"95","article-title":"Hidden Markov models for sequence analysis: extension and analysis of the basic method","volume":"12","author":"Hughey","year":"1996","journal-title":"Comput. Appl. Biosci."},{"key":"2023061007105598600_b19","article-title":"SAM: sequence alignment and modeling software system, version 3","volume-title":"Technical Report UCSC-CRL-99-11","author":"Hughey","year":"1999"},{"key":"2023061007105598600_b20","doi-asserted-by":"crossref","first-page":"2577","DOI":"10.1002\/bip.360221211","article-title":"Dictionary of protein secondary structure: pattern recognition of hydrogen-bonded and geometrical features","volume":"22","author":"Kabsch","year":"1983","journal-title":"Biopolymers"},{"key":"2023061007105598600_b21","doi-asserted-by":"crossref","first-page":"772","DOI":"10.1093\/bioinformatics\/14.9.772","article-title":"Weighting hidden Markov models for maximum discrimination","volume":"14","author":"Karchin","year":"1998","journal-title":"Bioinformatics"},{"key":"2023061007105598600_b22","doi-asserted-by":"crossref","first-page":"504","DOI":"10.1002\/prot.10369","article-title":"Hidden Markov models that use predicted local structure for fold recognition: alphabets of backbone geometry","volume":"51","author":"Karchin","year":"2003","journal-title":"Proteins"},{"key":"2023061007105598600_b23","doi-asserted-by":"crossref","first-page":"508","DOI":"10.1002\/prot.20008","article-title":"Evaluation of local structure alphabets based on residue burial","volume":"55","author":"Karchin","year":"2004","journal-title":"Proteins"},{"key":"2023061007105598600_b24","article-title":"gen_sequence: an open-source library","author":"Karplus","year":"2000"},{"key":"2023061007105598600_b25","doi-asserted-by":"crossref","first-page":"134","DOI":"10.1002\/(SICI)1097-0134(1997)1+<134::AID-PROT18>3.0.CO;2-P","article-title":"Predicting protein structure using hidden Markov models","author":"Karplus","year":"1997","journal-title":"Proteins"},{"key":"2023061007105598600_b26","doi-asserted-by":"crossref","first-page":"846","DOI":"10.1093\/bioinformatics\/14.10.846","article-title":"Hidden Markov models for detecting remote protein homologies","volume":"14","author":"Karplus","year":"1998","journal-title":"Bioinformatics"},{"key":"2023061007105598600_b27","doi-asserted-by":"crossref","first-page":"121","DOI":"10.1002\/(SICI)1097-0134(1999)37:3+<121::AID-PROT16>3.0.CO;2-Q","article-title":"Predicting protein structure using only sequence information","author":"Karplus","year":"1999","journal-title":"Proteins"},{"key":"2023061007105598600_b28","doi-asserted-by":"crossref","first-page":"86","DOI":"10.1002\/prot.10021","article-title":"What is the value added by human intervention in protein structure prediction?","volume":"45","author":"Karplus","year":"2001","journal-title":"Proteins"},{"key":"2023061007105598600_b29","doi-asserted-by":"crossref","first-page":"491","DOI":"10.1002\/prot.10540","article-title":"Combining local-structure, fold-recognition, and new-fold methods for protein structure prediction","volume":"53","author":"Karplus","year":"2003","journal-title":"Proteins"},{"key":"2023061007105598600_b30","doi-asserted-by":"crossref","first-page":"1501","DOI":"10.1006\/jmbi.1994.1104","article-title":"Hidden Markov models in computational biology: applications to protein modeling","volume":"235","author":"Krogh","year":"1994","journal-title":"J. Mol. Biol."},{"key":"2023061007105598600_b31","first-page":"155","article-title":"Parameterization studies for the SAM and HMMER methods of hidden Markov model generation","author":"McClure","year":"1996"},{"key":"2023061007105598600_b32","doi-asserted-by":"crossref","first-page":"536","DOI":"10.1016\/S0022-2836(05)80134-2","article-title":"SCOP: a structural classification of proteins database for the investigation of sequences and structures","volume":"247","author":"Murzin","year":"1995","journal-title":"J. Mol. Biol."},{"key":"2023061007105598600_b33","doi-asserted-by":"crossref","first-page":"2994","DOI":"10.1093\/nar\/29.14.2994","article-title":"Improving the accuracy of PSI-BLAST protein database searches with composition-based statistics and other refinements","volume":"29","author":"Sch\u00e4ffer","year":"2001","journal-title":"Nucleic Acids Res."},{"key":"2023061007105598600_b34","doi-asserted-by":"crossref","first-page":"482","DOI":"10.1016\/0196-8858(81)90046-4","article-title":"Comparison of bio-sequences","volume":"2","author":"Smith","year":"1981","journal-title":"Adv. Appl. Math."},{"key":"2023061007105598600_b35","doi-asserted-by":"crossref","first-page":"233","DOI":"10.1016\/0022-2836(86)90308-6","article-title":"Identification of protein sequence homology by consensus template alignment","volume":"188","author":"Taylor","year":"1986","journal-title":"J. Mol. Biol."},{"key":"2023061007105598600_b36","volume-title":"Numerical Recipes in C","author":"Vetterling","year":"1988"},{"key":"2023061007105598600_b37","doi-asserted-by":"crossref","first-page":"249","DOI":"10.1089\/10665270152530845","article-title":"Statistical significance of probabilistic sequence alignment and related local hidden Markov models","volume":"8","author":"Yu","year":"2001","journal-title":"J. Comput. Biol."}],"container-title":["Bioinformatics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article-pdf\/21\/22\/4107\/50566280\/bioinformatics_21_22_4107.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article-pdf\/21\/22\/4107\/50566280\/bioinformatics_21_22_4107.pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2023,6,10]],"date-time":"2023-06-10T07:11:51Z","timestamp":1686381111000},"score":1,"resource":{"primary":{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article\/21\/22\/4107\/194239"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2005,8,25]]},"references-count":37,"journal-issue":{"issue":"22","published-print":{"date-parts":[[2005,11,15]]}},"URL":"https:\/\/doi.org\/10.1093\/bioinformatics\/bti629","relation":{},"ISSN":["1367-4811","1367-4803"],"issn-type":[{"value":"1367-4811","type":"electronic"},{"value":"1367-4803","type":"print"}],"subject":[],"published-other":{"date-parts":[[2005,11,15]]},"published":{"date-parts":[[2005,8,25]]}}}