{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,11,11]],"date-time":"2025-11-11T12:50:27Z","timestamp":1762865427804},"reference-count":28,"publisher":"Oxford University Press (OUP)","issue":"13","license":[{"start":{"date-parts":[[2016,10,2]],"date-time":"2016-10-02T00:00:00Z","timestamp":1475366400000},"content-version":"vor","delay-in-days":3015,"URL":"http:\/\/creativecommons.org\/licenses\/by-nc\/2.0\/uk\/"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":[],"published-print":{"date-parts":[[2008,7,1]]},"abstract":"<jats:title>Abstract<\/jats:title>\n               <jats:p>Motivation: Small organic molecules, from nucleotides and amino acids to metabolites and drugs, play a fundamental role in chemistry, biology and medicine. As databases of small molecules continue to grow and become more open, it is important to develop the tools to search them efficiently. In order to develop a BLAST-like tool for small molecules, one must first understand the statistical behavior of molecular similarity scores.<\/jats:p>\n               <jats:p>Results: We develop a new detailed theory of molecular similarity scores that can be applied to a variety of molecular representations and similarity measures. For concreteness, we focus on the most widely used measure\u2014the Tanimoto measure applied to chem-ical fingerprints. In both the case of empirical fingerprints and fingerprints generated by several stochastic models, we derive accurate approximations for both the distribution and extreme value distribution of similarity scores. These approximation are derived using a ratio of correlated Gaussians approach. The theory enables the calculation of significance scores, such as Z-scores and P-values, and the estimation of the top hits list size. Empirical results obtained using both the random models and real data from the ChemDB database are given to corroborate the theory and show how it can be applied to mine chemical space.<\/jats:p>\n               <jats:p>Availability: Data and related resources are available through http:\/\/cdb.ics.uci.edu<\/jats:p>\n               <jats:p>Contact: \u00a0pfbaldi@ics.uci.edu<\/jats:p>","DOI":"10.1093\/bioinformatics\/btn187","type":"journal-article","created":{"date-parts":[[2008,6,27]],"date-time":"2008-06-27T07:43:13Z","timestamp":1214552593000},"page":"i357-i365","source":"Crossref","is-referenced-by-count":14,"title":["BLASTing small molecules\u2014statistics and extreme statistics of chemical similarity scores"],"prefix":"10.1093","volume":"24","author":[{"given":"Pierre","family":"Baldi","sequence":"first","affiliation":[{"name":"1 Department of Computer Science, 2Institute for Genomics and Bioinformatics and 3Department of Biological Chemistry, University of California, Irvine, CA 92697-3435, USA"},{"name":"1 Department of Computer Science, 2Institute for Genomics and Bioinformatics and 3Department of Biological Chemistry, University of California, Irvine, CA 92697-3435, USA"},{"name":"1 Department of Computer Science, 2Institute for Genomics and Bioinformatics and 3Department of Biological Chemistry, University of California, Irvine, CA 92697-3435, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Ryan W.","family":"Benz","sequence":"additional","affiliation":[{"name":"1 Department of Computer Science, 2Institute for Genomics and Bioinformatics and 3Department of Biological Chemistry, University of California, Irvine, CA 92697-3435, USA"},{"name":"1 Department of Computer Science, 2Institute for Genomics and Bioinformatics and 3Department of Biological Chemistry, University of California, Irvine, CA 92697-3435, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"286","published-online":{"date-parts":[[2008,7,1]]},"reference":[{"key":"2023020210391195400_B1","doi-asserted-by":"crossref","first-page":"147","DOI":"10.1207\/s15516709cog0901_7","article-title":"A learning algorithm for Boltzmann machines","volume":"9","author":"Ackley","year":"1985","journal-title":"Cogn. Sci"},{"key":"2023020210391195400_B2","doi-asserted-by":"crossref","first-page":"3389","DOI":"10.1093\/nar\/25.17.3389","article-title":"Gapped blast and psiblast: a new generation of protein database search programs","volume":"25","author":"Altschul","year":"1997","journal-title":"Nucleic Acids Res"},{"key":"2023020210391195400_B3","doi-asserted-by":"crossref","first-page":"2098","DOI":"10.1021\/ci700200n","article-title":"Lossless compression of chemical fingerprints using integer entropy codes improves storage and retrieval","volume":"47","author":"Baldi","year":"2007","journal-title":"J. Chem. Inform. Model"},{"key":"2023020210391195400_B4","first-page":"1708","article-title":"Similarity searching of chemical databases using atom environment descriptors (molprint 2d): Evaluation of performance","volume":"44","author":"Bender","year":"2004","journal-title":"J. Chem. Inform. Model"},{"key":"2023020210391195400_B5","doi-asserted-by":"crossref","first-page":"3","DOI":"10.1002\/(SICI)1098-1128(199601)16:1<3::AID-MED1>3.0.CO;2-6","article-title":"The art and practice of tructure-based drug design: a molecular modelling perspective","volume":"16","author":"Bohacek","year":"1996","journal-title":"Med. Res. Rev"},{"key":"2023020210391195400_B6","first-page":"99","article-title":"The distribution of the ratio of jointly normal variables","volume":"1","author":"Cedilnik","year":"2004","journal-title":"Metodoloski Zveki"},{"key":"2023020210391195400_B7","doi-asserted-by":"crossref","first-page":"4133","DOI":"10.1093\/bioinformatics\/bti683","article-title":"ChemDB: a public database of small molecules and related chemoinformatics resources","volume":"21","author":"Chen","year":"2005","journal-title":"Bioinformatics"},{"key":"2023020210391195400_B8","doi-asserted-by":"crossref","first-page":"2348","DOI":"10.1093\/bioinformatics\/btm341","article-title":"ChemDB update-full text search and virtual chemical space","volume":"23","author":"Chen","year":"2007","journal-title":"Bioinformatics"},{"key":"2023020210391195400_B9","doi-asserted-by":"crossref","DOI":"10.1007\/978-1-4471-3675-0","volume-title":"An Introduction to Statistical Modeling of Extreme Values","author":"Coles","year":"2001"},{"key":"2023020210391195400_B10","doi-asserted-by":"crossref","first-page":"110","DOI":"10.1198\/004017002317375064","article-title":"A modification of the Jaccard\/Tanimoto similarity index for diverse selection of chemical compounds using binary strings","volume":"44","author":"Fligner","year":"2002","journal-title":"Technometrics"},{"key":"2023020210391195400_B11","doi-asserted-by":"crossref","first-page":"379","DOI":"10.1021\/ci970437z","article-title":"On the properties of bit string-based measures of chemical similarity","volume":"38","author":"Flower","year":"1998","journal-title":"J. Chem. Inform. Comput. Sci"},{"key":"2023020210391195400_B12","doi-asserted-by":"crossref","DOI":"10.7551\/mitpress\/3348.001.0001","volume-title":"Graphical Models for Machine Learning and Digital Communicaiton","author":"Frey","year":"1998"},{"key":"2023020210391195400_B13","volume-title":"The Asymptotic Theory of Extreme Order Statistics","author":"Galambos","year":"1978"},{"key":"2023020210391195400_B14","doi-asserted-by":"crossref","first-page":"283","DOI":"10.1007\/s11030-006-9041-5","article-title":"Cheminformatics analysis and learning in a data pipelining environment","volume":"10","author":"Hassan","year":"2006","journal-title":"Mol. Divers"},{"key":"2023020210391195400_B15","doi-asserted-by":"crossref","first-page":"3256","DOI":"10.1039\/b409865j","article-title":"Comparison of topological descriptors for similarity-based virtual screening using multiple bioactive reference structures","volume":"2","author":"Hert","year":"2004","journal-title":"Org. Biomol. Chem"},{"key":"2023020210391195400_B16","doi-asserted-by":"crossref","first-page":"635","DOI":"10.1093\/biomet\/56.3.635","article-title":"On the ratio of two correlated normal random variables","volume":"56","author":"Hinkley","year":"1969","journal-title":"Biometrika"},{"key":"2023020210391195400_B17","doi-asserted-by":"crossref","first-page":"155","DOI":"10.2174\/1386207024607338","article-title":"Grouping of coefficients for the calculation of inter-molecular similarity and dissimilarity using 2d fragment bit-strings","volume":"5","author":"Holliday","year":"2002","journal-title":"Comb. Chem. High Throughput Screen"},{"key":"2023020210391195400_B18","doi-asserted-by":"crossref","first-page":"177","DOI":"10.1021\/ci049714+","article-title":"ZINC\u2013a free database of commercially available compounds for virtual screening","volume":"45","author":"Irwin","year":"2005","journal-title":"J. Chem. Inform. Comput. Sci"},{"key":"2023020210391195400_B19","unstructured":"James\n              CA\n            \n            \u00a0et al.\n          Daylight Theory Manual\n          2004\n          Available at http:\/\/www.daylight.com\/dayhtml\/doc\/theory\/theory.toc.html"},{"key":"2023020210391195400_B20","volume-title":"An Introduction to Chemoinformatics","author":"Leach","year":"2005"},{"key":"2023020210391195400_B21","doi-asserted-by":"crossref","DOI":"10.1007\/978-1-4612-5449-2","volume-title":"Extremems and Related Properties of Random Sequences and Series","author":"Leadbetter","year":"1983"},{"key":"2023020210391195400_B22","doi-asserted-by":"crossref","first-page":"193","DOI":"10.1080\/01621459.1965.10480783","article-title":"Ratios of normal variables and rations of sums of uniform variables","volume":"60","author":"Marsaglia","year":"1965","journal-title":"J. Ameri. Stat. Assoc"},{"key":"2023020210391195400_B23","doi-asserted-by":"crossref","first-page":"1569","DOI":"10.1080\/03610920600683689","article-title":"Density of the ratio of two normal random variables and applications","volume":"35","author":"Pham-Gia","year":"2006","journal-title":"Commun. Stat.-Theory Methods"},{"key":"2023020210391195400_B24","doi-asserted-by":"crossref","first-page":"580","DOI":"10.1021\/ci00010a002","article-title":"Definition and role of similarity concepts in the chemical and physical sciences","volume":"32","author":"Rouvray","year":"1992","journal-title":"J. Chem. Inform. Comput. Sci"},{"key":"2023020210391195400_B25","doi-asserted-by":"crossref","first-page":"302","DOI":"10.1021\/ci600358f","article-title":"Bounds and algorithms for exact searches of chemical fingerprints in linear and sub-linear time","volume":"47","author":"Swamidass","year":"2007","journal-title":"J. Chem. Inform. Model"},{"key":"2023020210391195400_B26","doi-asserted-by":"crossref","first-page":"327","DOI":"10.1037\/0033-295X.84.4.327","article-title":"Features of similarity","volume":"84","author":"Tversky","year":"1977","journal-title":"Psychol. Rev"},{"key":"2023020210391195400_B27","doi-asserted-by":"crossref","first-page":"1218","DOI":"10.1021\/ci030287u","article-title":"Profile scaling increases the similarity search performance of molecular fingerprints containing numerical descriptors and structural keys","volume":"43","author":"Xue","year":"2003","journal-title":"J. Chem. Inform. Comput. Sci"},{"key":"2023020210391195400_B28","doi-asserted-by":"crossref","first-page":"2032","DOI":"10.1021\/ci0400819","article-title":"Similarity search profiling reveals effects of fingerprint scaling in virtual screening","volume":"44","author":"Xue","year":"2004","journal-title":"J. Chem. Inform. Comput. Sci"}],"container-title":["Bioinformatics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article-pdf\/24\/13\/i357\/49050631\/bioinformatics_24_13_i357.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article-pdf\/24\/13\/i357\/49050631\/bioinformatics_24_13_i357.pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2023,2,2]],"date-time":"2023-02-02T12:22:22Z","timestamp":1675340542000},"score":1,"resource":{"primary":{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article\/24\/13\/i357\/236514"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2008,7,1]]},"references-count":28,"journal-issue":{"issue":"13","published-print":{"date-parts":[[2008,7,1]]}},"URL":"https:\/\/doi.org\/10.1093\/bioinformatics\/btn187","relation":{},"ISSN":["1367-4811","1367-4803"],"issn-type":[{"value":"1367-4811","type":"electronic"},{"value":"1367-4803","type":"print"}],"subject":[],"published-other":{"date-parts":[[2008,7,1]]},"published":{"date-parts":[[2008,7,1]]}}}