{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,8,18]],"date-time":"2026-08-18T09:31:17Z","timestamp":1787045477549,"version":"3.56.0"},"reference-count":38,"publisher":"American Chemical Society (ACS)","issue":"10","license":[{"start":{"date-parts":[[2020,10,6]],"date-time":"2020-10-06T00:00:00Z","timestamp":1601942400000},"content-version":"stm-asf","delay-in-days":0,"URL":"https:\/\/doi.org\/10.15223\/policy-029"},{"start":{"date-parts":[[2020,10,6]],"date-time":"2020-10-06T00:00:00Z","timestamp":1601942400000},"content-version":"stm-asf","delay-in-days":0,"URL":"https:\/\/doi.org\/10.15223\/policy-037"},{"start":{"date-parts":[[2020,10,6]],"date-time":"2020-10-06T00:00:00Z","timestamp":1601942400000},"content-version":"stm-asf","delay-in-days":0,"URL":"https:\/\/doi.org\/10.15223\/policy-045"}],"funder":[{"name":"MRL Postdoctoral Research Program"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":[],"published-print":{"date-parts":[[2020,10,26]]},"abstract":"<jats:title>Abstract<\/jats:title>\n                  <jats:p>While Gaussian process models are typically restricted to smaller data sets, we propose a variation which extends its applicability to the larger data sets common in the industrial drug discovery space, making it relatively novel in the quantitative structure\u2013activity relationship (QSAR) field. By incorporating locality-sensitive hashing for fast nearest neighbor searches, the nearest neighbor Gaussian process model makes predictions with time complexity that is sub-linear with the sample size. The model can be efficiently built, permitting rapid updates to prevent degradation as new data is collected. Given its small number of hyperparameters, it is robust against overfitting and generalizes about as well as other common QSAR models. Like the usual Gaussian process model, it natively produces principled and well-calibrated uncertainty estimates on its predictions. We compare this new model with implementations of random forest, light gradient boosting, and k-nearest neighbors to highlight these promising advantages. The code for the nearest neighbor Gaussian process is available at https:\/\/github.com\/Merck\/nngp.<\/jats:p>","DOI":"10.1021\/acs.jcim.0c00678","type":"journal-article","created":{"date-parts":[[2020,10,6]],"date-time":"2020-10-06T18:47:32Z","timestamp":1602010052000},"page":"4653-4663","source":"Crossref","is-referenced-by-count":8,"title":["Nearest Neighbor Gaussian Process for Quantitative\nStructure\u2013Activity Relationships"],"prefix":"10.1021","volume":"60","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-3151-9150","authenticated-orcid":true,"given":"Anthony","family":"DiFranzo","sequence":"first","affiliation":[{"name":"Merck & Company, Inc. , , , ,","place":["West Point, Pennsylvania, United States, 19486"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-6549-1635","authenticated-orcid":true,"given":"Robert P.","family":"Sheridan","sequence":"additional","affiliation":[{"name":"Merck & Company, Inc. , , , ,","place":["Kenilworth, New Jersey, United States, 07033"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Andy","family":"Liaw","sequence":"additional","affiliation":[{"name":"Merck & Company, Inc. , , , ,","place":["Rahway, New Jersey, United States, 07065"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-8766-195X","authenticated-orcid":true,"given":"Matthew","family":"Tudor","sequence":"additional","affiliation":[{"name":"Merck & Company, Inc. , , , ,","place":["West Point, Pennsylvania, United States, 19486"]}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"316","published-online":{"date-parts":[[2020,10,6]]},"reference":[{"key":"2026081804394937700_cit1","doi-asserted-by":"publisher","first-page":"256","DOI":"10.1002\/minf.201000160","article-title":"Predictivity\nof Simulated ADME AutoQSAR Models over Time","volume":"30","author":"Rodgers","year":"2011","journal-title":"Mol. Inf."},{"key":"2026081804394937700_cit2","doi-asserted-by":"publisher","first-page":"1183","DOI":"10.1021\/mp300466n","article-title":"Quantitative Structure-Activity\nRelationship Models\nThat Stand the Test of Time","volume":"10","author":"Davis","year":"2013","journal-title":"Mol. Pharm."},{"key":"2026081804394937700_cit3","doi-asserted-by":"publisher","first-page":"321","DOI":"10.1007\/s10822-013-9648-4","article-title":"QSAR workbench:\nautomating QSAR modeling to drive compound design","volume":"27","author":"Cox","year":"2013","journal-title":"J. Comput.-Aided Mol. Des."},{"key":"2026081804394937700_cit4","doi-asserted-by":"publisher","first-page":"1083","DOI":"10.1021\/ci500084w","article-title":"Global\nQuantitative Structure-Activity Relationship Models vs Selected Local\nModels as Predictors of Off-Target Activities for Project Compounds","volume":"54","author":"Sheridan","year":"2014","journal-title":"J. Chem. Inf. Model."},{"key":"2026081804394937700_cit5","doi-asserted-by":"publisher","first-page":"67","DOI":"10.1109\/4235.585893","article-title":"No free\nlunch theorems for optimization","volume":"1","author":"Wolpert","year":"1997","journal-title":"IEEE Trans.\nEvol. Comput."},{"key":"2026081804394937700_cit6","unstructured":"Williams, C. K.; Rasmussen, C. E.\n          Gaussian processes\nfor regression. Advances in Neural Information\nProcessing Systems, 1996; pp 514\u2013520."},{"key":"2026081804394937700_cit7","volume-title":"Gaussian Processes\nfor Machine Learning","author":"Rasmussen","year":"2006"},{"key":"2026081804394937700_cit8","doi-asserted-by":"crossref","first-page":"192","DOI":"10.1007\/978-3-319-62416-7_14","article-title":"Over-Fitting\nin Model Selection with Gaussian Process Regression","volume-title":"Machine Learning and Data Mining in Pattern Recognition","author":"Mohammed","year":"2017"},{"key":"2026081804394937700_cit9","doi-asserted-by":"publisher","first-page":"830","DOI":"10.1021\/ci000459c","article-title":"Quantitative\nStructure\u2013Activity Relationship Studies Using Gaussian Processes","volume":"41","author":"Burden","year":"2001","journal-title":"J. Chem. Inf. Comput. Sci."},{"key":"2026081804394937700_cit10","doi-asserted-by":"publisher","first-page":"461","DOI":"10.1080\/10629360108035385","article-title":"Gaussian Process: An Efficient Technique\nto Solve Quantitative Structure-Property Relationship Problems","volume":"12","author":"Enot","year":"2001","journal-title":"SAR QSAR Environ. Res."},{"key":"2026081804394937700_cit11","doi-asserted-by":"publisher","first-page":"1847","DOI":"10.1021\/ci7000633","article-title":"Gaussian Processes:\nA Method for Automatic QSAR Modeling of ADME Properties","volume":"47","author":"Obrezanova","year":"2007","journal-title":"J. Chem. Inf. Model."},{"key":"2026081804394937700_cit12","volume-title":"Efficient Implementation\nof Gaussian Processes","author":"Gibbs","year":"1997"},{"key":"2026081804394937700_cit13","doi-asserted-by":"publisher","first-page":"1836","DOI":"10.1021\/ci060064e","article-title":"Local Lazy Regression:\nMaking Use of the Neighborhood to Improve QSAR Predictions","volume":"46","author":"Guha","year":"2006","journal-title":"J. Chem. Inf. Model."},{"key":"2026081804394937700_cit14","doi-asserted-by":"publisher","first-page":"1984","DOI":"10.1021\/ci060132x","article-title":"A Novel Automated\nLazy Learning QSAR (ALL-QSAR) Approach: Method Development, Applications,\nand Virtual Screening of Chemical Databases Using Validated ALL-QSAR\nModels","volume":"46","author":"Zhang","year":"2006","journal-title":"J. Chem. Inf. Model."},{"key":"2026081804394937700_cit15","doi-asserted-by":"publisher","first-page":"440","DOI":"10.2174\/138620709788167908","article-title":"Review on Lazy Learning Regressors\nand their Applications in QSAR","volume":"12","author":"Kulkarni","year":"2009","journal-title":"Comb. Chem.\nHigh Throughput Screening"},{"key":"2026081804394937700_cit16","doi-asserted-by":"publisher","first-page":"973","DOI":"10.1002\/jcc.21383","article-title":"A new strategy\nto improve the predictive ability of\nthe local lazy regression and its application to the QSAR study of\nmelanin-concentrating hormone receptor 1 antagonists","volume":"31","author":"Li","year":"2010","journal-title":"J. Comput. Chem."},{"key":"2026081804394937700_cit17","doi-asserted-by":"publisher","first-page":"5","DOI":"10.1023\/a:1010933404324","article-title":"Random forests","volume":"45","author":"Breiman","year":"2001","journal-title":"Mach. Learn."},{"key":"2026081804394937700_cit18","doi-asserted-by":"publisher","first-page":"1947","DOI":"10.1021\/ci034160g","article-title":"Random Forest: A\nClassification and\nRegression Tool for Compound Classification and QSAR Modeling","volume":"43","author":"Svetnik","year":"2003","journal-title":"J. Chem. Inf. Comput. Sci."},{"key":"2026081804394937700_cit19","doi-asserted-by":"publisher","first-page":"2353","DOI":"10.1021\/acs.jcim.6b00591","article-title":"Extreme\nGradient Boosting as a Method for Quantitative Structure-Activity\nRelationships","volume":"56","author":"Sheridan","year":"2016","journal-title":"J. Chem. Inf. Model."},{"key":"2026081804394937700_cit20","article-title":"Benchmarking and\nOptimization of Gradient Boosting Decision Tree Algorithms","author":"Anghel","year":"2018"},{"key":"2026081804394937700_cit21","doi-asserted-by":"publisher","first-page":"4150","DOI":"10.1021\/acs.jcim.9b00633","article-title":"LightGBM: An Effective\nand Scalable Algorithm for Prediction of Chemical Toxicity-Application\nto the Tox21 and Mutagenicity Data Sets","volume":"59","author":"Zhang","year":"2019","journal-title":"J.\nChem. Inf. Model."},{"key":"2026081804394937700_cit22","unstructured":"Williams, C. K.; Seeger, M.\n          Using the Nystro\u0308m\nmethod to speed up kernel machines. Advances\nin Neural Information Processing Systems, 2001; pp 682\u2013688."},{"key":"2026081804394937700_cit23","doi-asserted-by":"crossref","unstructured":"Charikar, M. S.\n          \n          Similarity\nEstimation Techniques from Rounding Algorithms. Proceedings of the Thiry-Fourth Annual ACM Symposium on Theory of\nComputing, New York, NY, USA, 2002; pp 380\u2013388.","DOI":"10.1145\/509907.509965"},{"key":"2026081804394937700_cit24","doi-asserted-by":"publisher","first-page":"117","DOI":"10.1145\/1327452.1327494","article-title":"Near-Optimal Hashing\nAlgorithms for Approximate Nearest Neighbor\nin High Dimensions","volume":"51","author":"Andoni","year":"2008","journal-title":"Commun. ACM"},{"key":"2026081804394937700_cit25","unstructured":"Bernhardsson, E.\n          \n          Annoy: Approximate\nNearest Neighbors in C++\/Python, 2018https:\/\/pypi.org\/project\/annoy\/, Python package version 1.16.3 (accessed March 6, 2020)."},{"key":"2026081804394937700_cit26","doi-asserted-by":"publisher","first-page":"1910","DOI":"10.1021\/acs.jcim.0c00029","article-title":"Correction\nto Extreme Gradient Boosting as a Method for Quantitative Structure-Activity\nRelationships","volume":"60","author":"Sheridan","year":"2020","journal-title":"J. Chem. Inf. Model."},{"key":"2026081804394937700_cit27","doi-asserted-by":"publisher","first-page":"783","DOI":"10.1021\/ci400084k","article-title":"Time-Split\nCross-Validation as a Method for Estimating the Goodness of Prospective\nPrediction","volume":"53","author":"Sheridan","year":"2013","journal-title":"J. Chem. Inf. Model."},{"key":"2026081804394937700_cit28","doi-asserted-by":"publisher","first-page":"64","DOI":"10.1021\/ci00046a002","article-title":"Atom pairs as molecular features\nin structure-activity studies: definition and applications","volume":"25","author":"Carhart","year":"1985","journal-title":"J. Chem. Inf. Comput. Sci."},{"key":"2026081804394937700_cit29","doi-asserted-by":"publisher","first-page":"118","DOI":"10.1021\/ci950274j","article-title":"Chemical Similarity Using Physiochemical Property Descriptors","volume":"36","author":"Kearsley","year":"1996","journal-title":"J. Chem. Inf. Comput. Sci."},{"key":"2026081804394937700_cit30","first-page":"2825","article-title":"Scikit-learn: Machine Learning in Python","volume":"12","author":"Pedregosa","year":"2011","journal-title":"J. Mach. Learn. Res."},{"key":"2026081804394937700_cit31","doi-asserted-by":"publisher","first-page":"814","DOI":"10.1021\/ci300004n","article-title":"Three Useful\nDimensions for Domain Applicability in QSAR Models Using Random Forest","volume":"52","author":"Sheridan","year":"2012","journal-title":"J. Chem. Inf. Model."},{"key":"2026081804394937700_cit32","unstructured":"Ke, G.; Meng, Q.; Finley, T.; Wang, T.; Chen, W.; Ma, W.; Ye, Q.; Liu, T.-Y.\n          LightGBM: A Highly Efficient Gradient Boosting Decision\nTree. Advances in Neural Information Processing\nSystems 30, 2017; pp 3146\u20133154."},{"key":"2026081804394937700_cit33","first-page":"371","article-title":"A tutorial on conformal\nprediction","volume":"9","author":"Shafer","year":"2008","journal-title":"J. Mach.\nLearn. Res."},{"key":"2026081804394937700_cit34","doi-asserted-by":"publisher","first-page":"155","DOI":"10.1007\/s10994-014-5453-0","article-title":"Regression\nconformal prediction with random forests","volume":"97","author":"Johansson","year":"2014","journal-title":"Mach.\nLearn."},{"key":"2026081804394937700_cit35","doi-asserted-by":"publisher","first-page":"3370","DOI":"10.1021\/acs.jcim.9b00237","article-title":"Analyzing\nLearned Molecular Representations for Property\nPrediction","volume":"59","author":"Yang","year":"2019","journal-title":"J. Chem. Inf. Model."},{"key":"2026081804394937700_cit36","unstructured":"Head, T.; Louppe, G.; Shcherbatyi, I.; Vini\u0301cius, Z.; Schro\u0308der, C.; Campos, N.; Young, T.; Cereda, S.; Fan, T.; Shi, K. K.; Schwabedal, J.; Pak, M.; Callaway, F.; Este\u0300ve, L.; Besson, L.; Cherti, M.; Pfannschmidt, K.; Linzberger, F.; Cauet, C.; Gut, A.; Mueller, A.; Fabisch, A.\n          scikit-optimize\/scikit-optimize: v0.5.2, 2018, (accessed Oct 29, 2019)."},{"key":"2026081804394937700_cit37","doi-asserted-by":"publisher","first-page":"529","DOI":"10.1111\/j.2517-6161.1994.tb01998.x","article-title":"Automated\nKernel Smoothing of Dependent Data by Using Time Series Cross-Validation","volume":"56","author":"Hart","year":"1994","journal-title":"J. R. Stat. Soc. Series B"},{"key":"2026081804394937700_cit38","doi-asserted-by":"publisher","first-page":"742","DOI":"10.1021\/ci100050t","article-title":"Extended-Connectivity\nFingerprints","volume":"50","author":"Rogers","year":"2010","journal-title":"J. Chem.\nInf. Model."}],"container-title":["Journal of Chemical Information and Modeling"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/pubs.acs.org\/doi\/pdf\/10.1021\/acs.jcim.0c00678","content-type":"application\/pdf","content-version":"vor","intended-application":"unspecified"},{"URL":"https:\/\/pubs.acs.org\/jcisd8\/article-pdf\/60\/10\/4653\/8898711\/ci0c00678.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/pubs.acs.org\/jcisd8\/article-pdf\/60\/10\/4653\/8898711\/ci0c00678.pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,8,18]],"date-time":"2026-08-18T09:00:30Z","timestamp":1787043630000},"score":1,"resource":{"primary":{"URL":"https:\/\/pubs.acs.org\/jcisd8\/article\/60\/10\/4653\/849712\/Nearest-Neighbor-Gaussian-Process-for-Quantitative"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2020,10,6]]},"references-count":38,"journal-issue":{"issue":"10","published-online":{"date-parts":[[2020,10,6]]},"published-print":{"date-parts":[[2020,10,26]]}},"URL":"https:\/\/doi.org\/10.1021\/acs.jcim.0c00678","relation":{},"ISSN":["1549-9596","1549-960X"],"issn-type":[{"value":"1549-9596","type":"print"},{"value":"1549-960X","type":"electronic"}],"subject":[],"published":{"date-parts":[[2020,10,6]]}}}