{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,12,18]],"date-time":"2025-12-18T14:20:35Z","timestamp":1766067635375,"version":"3.37.3"},"reference-count":69,"publisher":"IOP Publishing","issue":"4","license":[{"start":{"date-parts":[[2022,12,19]],"date-time":"2022-12-19T00:00:00Z","timestamp":1671408000000},"content-version":"vor","delay-in-days":18,"URL":"http:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2022,12,19]],"date-time":"2022-12-19T00:00:00Z","timestamp":1671408000000},"content-version":"tdm","delay-in-days":18,"URL":"https:\/\/iopscience.iop.org\/info\/page\/text-and-data-mining"}],"funder":[{"DOI":"10.13039\/100006235","name":"Lawrence Berkeley National Laboratory","doi-asserted-by":"crossref","id":[{"id":"10.13039\/100006235","id-type":"DOI","asserted-by":"crossref"}]},{"name":"National Energy Research Scienti\ufb01c Computing Center"},{"DOI":"10.13039\/100006151","name":"Basic Energy Sciences","doi-asserted-by":"crossref","id":[{"id":"10.13039\/100006151","id-type":"DOI","asserted-by":"crossref"}]},{"name":"O\ufb03ce of Science User Facility","award":["DE-AC02-05CH11231"],"award-info":[{"award-number":["DE-AC02-05CH11231"]}]},{"name":"FWO"},{"DOI":"10.13039\/100000015","name":"U.S. Department of Energy","doi-asserted-by":"crossref","award":["DE-AC02\u201305CH11231"],"award-info":[{"award-number":["DE-AC02\u201305CH11231"]}],"id":[{"id":"10.13039\/100000015","id-type":"DOI","asserted-by":"crossref"}]},{"name":"National Science and Engineering Council of Canada. C.C."}],"content-domain":{"domain":["iopscience.iop.org"],"crossmark-restriction":false},"short-container-title":["Mach. Learn.: Sci. Technol."],"published-print":{"date-parts":[[2022,12,1]]},"abstract":"<jats:title>Abstract<\/jats:title>\n               <jats:p>We examine the zero-temperature Metropolis Monte Carlo (MC) algorithm as a tool for training a neural network by minimizing a loss function. We find that, as expected on theoretical grounds and shown empirically by other authors, Metropolis MC can train a neural net with an accuracy comparable to that of gradient descent (GD), if not necessarily as quickly. The Metropolis algorithm does not fail automatically when the number of parameters of a neural network is large. It can fail when a neural network\u2019s structure or neuron activations are strongly heterogenous, and we introduce an adaptive Monte Carlo algorithm (aMC) to overcome these limitations. The intrinsic stochasticity and numerical stability of the MC method allow aMC to train deep neural networks and recurrent neural networks in which the gradient is too small or too large to allow training by GD. MC methods offer a complement to gradient-based methods for training neural networks, allowing access to a distinct set of network architectures and principles.<\/jats:p>","DOI":"10.1088\/2632-2153\/aca6cd","type":"journal-article","created":{"date-parts":[[2022,11,29]],"date-time":"2022-11-29T02:25:41Z","timestamp":1669688741000},"page":"045026","update-policy":"https:\/\/doi.org\/10.1088\/crossmark-policy","source":"Crossref","is-referenced-by-count":9,"title":["Training neural networks using Metropolis Monte Carlo and an adaptive variant"],"prefix":"10.1088","volume":"3","author":[{"given":"Stephen","family":"Whitelam","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Viktor","family":"Selin","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Ian","family":"Benlolo","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Corneel","family":"Casert","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-8146-6667","authenticated-orcid":true,"given":"Isaac","family":"Tamblyn","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"266","published-online":{"date-parts":[[2022,12,19]]},"reference":[{"key":"mlstaca6cdbib1","doi-asserted-by":"publisher","first-page":"1087","DOI":"10.1063\/1.1699114","article-title":"Equation of state calculations by fast computing machines","volume":"21","author":"Metropolis","year":"1953","journal-title":"J. Chem. Phys."},{"key":"mlstaca6cdbib2","doi-asserted-by":"publisher","DOI":"10.1063\/1.1887186","article-title":"Marshall rosenbluth and the metropolis algorithm","volume":"12","author":"Gubernatis","year":"2005","journal-title":"Phys. Plasmas"},{"key":"mlstaca6cdbib3","doi-asserted-by":"publisher","first-page":"22","DOI":"10.1063\/1.1632112","article-title":"Genesis of the Monte Carlo algorithm for statistical mechanics","volume":"690","author":"Rosenbluth","year":"2003"},{"key":"mlstaca6cdbib4","doi-asserted-by":"crossref","DOI":"10.2172\/1770095","article-title":"Arianna wright rosenbluth","author":"Whitacre","year":"2021"},{"volume":"vol 1","year":"2001","author":"Frenkel","key":"mlstaca6cdbib5"},{"key":"mlstaca6cdbib6","doi-asserted-by":"publisher","first-page":"3","DOI":"10.4018\/joeuc.1999070101","article-title":"Beyond backpropagation: using simulated annealing for training neural networks","volume":"11","author":"Sexton","year":"1999","journal-title":"J. Organ. End User Comput."},{"key":"mlstaca6cdbib7","doi-asserted-by":"publisher","first-page":"137","DOI":"10.1016\/j.procs.2015.12.114","article-title":"Simulated annealing algorithm for deep learning","volume":"72","author":"Rere","year":"2015","journal-title":"Proc. Comput. Sci."},{"article-title":"Rso: a gradient free sampling based approach for training deep neural networks","year":"2020","author":"Tripathi","key":"mlstaca6cdbib8"},{"key":"mlstaca6cdbib9","doi-asserted-by":"publisher","first-page":"85","DOI":"10.1016\/j.neunet.2014.09.003","article-title":"Deep learning in neural networks: an overview","volume":"61","author":"Schmidhuber","year":"2015","journal-title":"Neural Netw."},{"year":"2016","author":"Goodfellow","key":"mlstaca6cdbib10"},{"key":"mlstaca6cdbib11","doi-asserted-by":"publisher","first-page":"66","DOI":"10.1038\/scientificamerican0792-66","article-title":"Genetic algorithms","volume":"267","author":"Holland","year":"1992","journal-title":"Sci. Am."},{"key":"mlstaca6cdbib12","doi-asserted-by":"publisher","first-page":"171","DOI":"10.1016\/0303-2647(94)90040-X","article-title":"On the effectiveness of crossover in simulated evolutionary optimization","volume":"32","author":"Fogel","year":"1994","journal-title":"Biosystems"},{"key":"mlstaca6cdbib13","first-page":"762","article-title":"Training feedforward neural networks using genetic algorithms","volume":"89","author":"Montana","year":"1989"},{"key":"mlstaca6cdbib14","article-title":"Zero temperature means that moves that increase the loss are not accepted. This choice is motivated by the empirical success in machine learning of gradient-descent methods, and by the intuition, derived from Gaussian random surfaces, that loss surfaces possess more downhill directions at large values of the loss [31 32]"},{"key":"mlstaca6cdbib15","doi-asserted-by":"publisher","first-page":"335","DOI":"10.1016\/S0009-2614(91)85070-D","article-title":"Metropolis Monte Carlo method as a numerical technique to solve the Fokker\u2013Planck equation","volume":"185","author":"Kikuchi","year":"1991","journal-title":"Chem. Phys. Lett."},{"key":"mlstaca6cdbib16","doi-asserted-by":"publisher","first-page":"57","DOI":"10.1016\/0009-2614(92)85928-4","article-title":"Metropolis Monte Carlo method for Brownian dynamics simulation generalized to include hydrodynamic interactions","volume":"196","author":"Kikuchi","year":"1992","journal-title":"Chem. Phys. Lett."},{"key":"mlstaca6cdbib17","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1038\/s41467-021-26568-2","article-title":"Correspondence between neuroevolution and gradient descent","volume":"12","author":"Whitelam","year":"2021","journal-title":"Nat. Commun."},{"key":"mlstaca6cdbib18","article-title":"Note that algorithms of this nature do not constitute random search. The proposal step is random (related conceptually to the idea of weight guessing, a method used in the presence of vanishing gradients [42]) but the acceptance criterion is a form of importance sampling, and leads to a dynamics equivalent to noisy gradient descent"},{"article-title":"Evolution strategies as a scalable alternative to reinforcement learning","year":"2017","author":"Salimans","key":"mlstaca6cdbib19"},{"article-title":"Adam: a method for stochastic optimization","year":"2014","author":"Kingma","key":"mlstaca6cdbib20"},{"key":"mlstaca6cdbib21","doi-asserted-by":"publisher","first-page":"436","DOI":"10.1038\/nature14539","article-title":"Deep learning","volume":"521","author":"LeCun","year":"2015","journal-title":"Nature"},{"article-title":"Gradients are not all you need","year":"2021","author":"Metz","key":"mlstaca6cdbib22"},{"key":"mlstaca6cdbib23","article-title":"When will a genetic algorithm outperform hill climbing?","volume":"6","author":"Mitchell","year":"1993"},{"year":"1998","author":"Mitchell","key":"mlstaca6cdbib24"},{"key":"mlstaca6cdbib25","article-title":"In Metropolis Monte Carlo simulations of molecular systems it is usual to propose moves of one particle at a time. If we consider neural-net parameters to be akin to particle coordinates then the analog would be to make changes to one neural-net parameter at a time; see e.g. [8]. However, there is no formal mapping between particles and a neural network, and we could equally well consider the neural-net parameters to be akin to the coordinates of a single particle, in a high-dimensional space, in an external potential equal to the loss function. In the latter case the analog would be to propose a change of all neural-net parameters simultaneously, as we do here"},{"key":"mlstaca6cdbib26","doi-asserted-by":"publisher","first-page":"2278","DOI":"10.1109\/5.726791","article-title":"Gradient-based learning applied to document recognition","volume":"86","author":"LeCun","year":"1998","journal-title":"Proc. IEEE"},{"key":"mlstaca6cdbib27","doi-asserted-by":"publisher","first-page":"2278","DOI":"10.1109\/5.726791","article-title":"Gradient-based learning applied to document recognition","volume":"86","author":"LeCun","year":"1998"},{"article-title":"But what is a neural network? | Chapter 1, Deep learning","year":"2017","author":"","key":"mlstaca6cdbib28"},{"key":"mlstaca6cdbib29","article-title":"We have also found the GD-MC equivalence to break down in other circumstances: for certain learning rates \u03b1, the discrete-update equation (3) sometimes results in moves uphill in loss, in which case the discrete update is not equivalent to the equation x\u02d9=\u2212(\u03b1\/\u0394t)\u2207U(x) , while the latter is equivalent to the small-step-size limit of the finite-temperature Metropolis algorithm [15\u201317]"},{"key":"mlstaca6cdbib30","article-title":"In figure 8 we show that GD and MC can both train a large modern neural network to a classification accuracy in excess of 99% on the same problem"},{"key":"mlstaca6cdbib31","article-title":"Identifying and attacking the saddle point problem in high-dimensional non-convex optimization","volume":"27","author":"Dauphin","year":"2014"},{"key":"mlstaca6cdbib32","doi-asserted-by":"publisher","first-page":"501","DOI":"10.1146\/annurev-conmatphys-031119-050745","article-title":"Statistical mechanics of deep learning","volume":"11","author":"Bahri","year":"2020","journal-title":"Annu. Rev. Condens. Matter Phys."},{"key":"mlstaca6cdbib33","doi-asserted-by":"publisher","first-page":"159","DOI":"10.1162\/106365601750190398","article-title":"Completely derandomized self-adaptation in evolution strategies","volume":"9","author":"Hansen","year":"2001","journal-title":"Evol. Comput."},{"key":"mlstaca6cdbib34","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1162\/106365603321828970","article-title":"Reducing the time complexity of the derandomized evolution strategy with covariance matrix adaptation (cma-es)","volume":"11","author":"Hansen","year":"2003","journal-title":"Evol. Comput."},{"key":"mlstaca6cdbib35","first-page":"pp 75","article-title":"The CMA evolution strategy: a comparing review","author":"Hansen","year":"2006"},{"key":"mlstaca6cdbib36","doi-asserted-by":"publisher","first-page":"175","DOI":"10.1093\/comjnl\/3.3.175","article-title":"An automatic method for finding the greatest or least value of a function","volume":"3","author":"Rosenbrock","year":"1960","journal-title":"Comput. J."},{"key":"mlstaca6cdbib37","doi-asserted-by":"publisher","first-page":"119","DOI":"10.1162\/evco.2006.14.1.119","article-title":"A note on the extended Rosenbrock function","volume":"14","author":"Shang","year":"2006","journal-title":"Evol. Comput."},{"key":"mlstaca6cdbib38","first-page":"pp 837","article-title":"Comparison of minimization methods for Rosenbrock functions","author":"Emiola","year":"2021"},{"key":"mlstaca6cdbib39","doi-asserted-by":"publisher","first-page":"e6","DOI":"10.23915\/distill.00006","article-title":"Why momentum really works","volume":"2","author":"Goh","year":"2017","journal-title":"Distill"},{"key":"mlstaca6cdbib40","first-page":"pp 1","article-title":"Backpropagation: the basic theory","author":"Rumelhart","year":"1995"},{"article-title":"Closing the generalization gap of adaptive gradient methods in training deep neural networks","year":"2018","author":"Chen","key":"mlstaca6cdbib41"},{"key":"mlstaca6cdbib42","doi-asserted-by":"publisher","first-page":"1735","DOI":"10.1162\/neco.1997.9.8.1735","article-title":"Long short-term memory","volume":"9","author":"Hochreiter","year":"1997","journal-title":"Neural Comput."},{"key":"mlstaca6cdbib43","doi-asserted-by":"publisher","first-page":"107","DOI":"10.1142\/S0218488598000094","article-title":"The vanishing gradient problem during learning recurrent neural nets and problem solutions","volume":"6","author":"Hochreiter","year":"1998","journal-title":"Int. J. Uncertain. Fuzziness Knowl.-Based Syst."},{"key":"mlstaca6cdbib44","first-page":"64","article-title":"Recurrent neural networks","volume":"5","author":"Medsker","year":"2001","journal-title":"Des. Appl."},{"key":"mlstaca6cdbib45","first-page":"pp 6645","article-title":"Speech recognition with deep recurrent neural networks","author":"Graves","year":"2013"},{"key":"mlstaca6cdbib46","article-title":"Sequence to sequence learning with neural networks","volume":"27","author":"Sutskever","year":"2014"},{"key":"mlstaca6cdbib47","article-title":"Offline handwriting recognition with multidimensional recurrent neural networks","volume":"21","author":"Graves","year":"2008"},{"key":"mlstaca6cdbib48","doi-asserted-by":"publisher","first-page":"620","DOI":"10.1093\/jigpal\/jzp049","article-title":"Recurrent policy gradients","volume":"18","author":"Wierstra","year":"2010","journal-title":"Logic J. IGPL"},{"key":"mlstaca6cdbib49","doi-asserted-by":"publisher","first-page":"157","DOI":"10.1109\/72.279181","article-title":"Learning long-term dependencies with gradient descent is difficult","volume":"5","author":"Bengio","year":"1994","journal-title":"IEEE Trans. Neural Netw."},{"key":"mlstaca6cdbib50","article-title":"Learning recurrent neural networks with hessian-free optimization","volume":"vol 28","author":"Martens","year":"2011"},{"key":"mlstaca6cdbib51","first-page":"pp 8624","article-title":"Advances in optimizing recurrent networks","author":"Bengio","year":"2013"},{"key":"mlstaca6cdbib52","doi-asserted-by":"crossref","DOI":"10.3115\/v1\/D14-1179","article-title":"Learning phrase representations using RNN encoder\u2013decoder for statistical machine translation","author":"Cho","year":"2014"},{"key":"mlstaca6cdbib53","article-title":"Preventing gradient explosions in gated recurrent units","volume":"30","author":"Kanai","year":"2017"},{"key":"mlstaca6cdbib54","first-page":"pp 1310","article-title":"On the difficulty of training recurrent neural networks","author":"Pascanu","year":"2013"},{"article-title":"Capacity and trainability in recurrent neural networks","year":"2016","author":"Collins","key":"mlstaca6cdbib55"},{"key":"mlstaca6cdbib56","first-page":"pp 9","article-title":"Efficient backprop","author":"LeCun","year":"1996"},{"article-title":"Layer normalization","year":"2016","author":"Ba","key":"mlstaca6cdbib57"},{"key":"mlstaca6cdbib58","article-title":"Pytorch: an imperative style, high-performance deep learning library","volume":"32","author":"Paszke","year":"2019"},{"key":"mlstaca6cdbib59","first-page":"pp 770","article-title":"Deep residual learning for image recognition","author":"He","year":"2016"},{"key":"mlstaca6cdbib60","article-title":"For finite temperature T the move is accepted if \u03be<e(U(x)\u2212U(x\u2032))\/T , where \u03be is a random number drawn uniformly on (0,1]"},{"key":"mlstaca6cdbib61","article-title":"Optimal proposal distributions and adaptive MCMC","volume":"vol 4","author":"Rosenthal","year":"2011"},{"key":"mlstaca6cdbib62","article-title":"This approximation assumes that the output neurons do not change under the move. This is not true, but the intent here is to set the basic move scale, and absolute precision is not necessary"},{"key":"mlstaca6cdbib63","doi-asserted-by":"publisher","first-page":"1262","DOI":"10.1103\/PhysRevE.56.1262","article-title":"Stochastic manhattan learning: Time-evolution operator for the ensemble dynamics","volume":"56","author":"Leen","year":"1997","journal-title":"Phys. Rev. E"},{"key":"mlstaca6cdbib64","doi-asserted-by":"publisher","first-page":"86","DOI":"10.1103\/PhysRevLett.58.86","article-title":"Nonuniversal critical dynamics in Monte Carlo simulations","volume":"58","author":"Swendsen","year":"1987","journal-title":"Phys. Rev. Lett."},{"key":"mlstaca6cdbib65","doi-asserted-by":"publisher","first-page":"361","DOI":"10.1103\/PhysRevLett.62.361","article-title":"Collective Monte Carlo updating for spin systems","volume":"62","author":"Wolff","year":"1989","journal-title":"Phys. Rev. Lett."},{"key":"mlstaca6cdbib66","doi-asserted-by":"publisher","first-page":"11275","DOI":"10.1021\/jp012209k","article-title":"Improving the efficiency of the aggregation-volume-bias Monte Carlo algorithm","volume":"105","author":"Chen","year":"2001","journal-title":"J. Phys. Chem. B"},{"key":"mlstaca6cdbib67","doi-asserted-by":"publisher","DOI":"10.1103\/PhysRevLett.92.035504","article-title":"Rejection-free geometric cluster algorithm for complex fluids","volume":"92","author":"Liu","year":"2004","journal-title":"Phys. Rev. Lett."},{"key":"mlstaca6cdbib68","doi-asserted-by":"publisher","DOI":"10.1063\/1.2790421","article-title":"Avoiding unphysical kinetic traps in Monte Carlo simulations of strongly attractive particles","volume":"127","author":"Whitelam","year":"2007","journal-title":"J. Chem. Phys."},{"year":"2022","author":"Whitelam","key":"mlstaca6cdbib69"}],"container-title":["Machine Learning: Science and Technology"],"original-title":[],"link":[{"URL":"https:\/\/iopscience.iop.org\/article\/10.1088\/2632-2153\/aca6cd","content-type":"text\/html","content-version":"am","intended-application":"text-mining"},{"URL":"https:\/\/iopscience.iop.org\/article\/10.1088\/2632-2153\/aca6cd\/pdf","content-type":"application\/pdf","content-version":"am","intended-application":"text-mining"},{"URL":"https:\/\/iopscience.iop.org\/article\/10.1088\/2632-2153\/aca6cd","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/iopscience.iop.org\/article\/10.1088\/2632-2153\/aca6cd\/pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/iopscience.iop.org\/article\/10.1088\/2632-2153\/aca6cd\/pdf","content-type":"application\/pdf","content-version":"am","intended-application":"syndication"},{"URL":"https:\/\/iopscience.iop.org\/article\/10.1088\/2632-2153\/aca6cd\/pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/iopscience.iop.org\/article\/10.1088\/2632-2153\/aca6cd\/pdf","content-type":"application\/pdf","content-version":"am","intended-application":"similarity-checking"},{"URL":"https:\/\/iopscience.iop.org\/article\/10.1088\/2632-2153\/aca6cd\/pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2022,12,19]],"date-time":"2022-12-19T11:09:52Z","timestamp":1671448192000},"score":1,"resource":{"primary":{"URL":"https:\/\/iopscience.iop.org\/article\/10.1088\/2632-2153\/aca6cd"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,12,1]]},"references-count":69,"journal-issue":{"issue":"4","published-online":{"date-parts":[[2022,12,19]]},"published-print":{"date-parts":[[2022,12,1]]}},"URL":"https:\/\/doi.org\/10.1088\/2632-2153\/aca6cd","relation":{},"ISSN":["2632-2153"],"issn-type":[{"type":"electronic","value":"2632-2153"}],"subject":[],"published":{"date-parts":[[2022,12,1]]},"assertion":[{"value":"Training neural networks using Metropolis Monte Carlo and an adaptive variant","name":"article_title","label":"Article Title"},{"value":"Machine Learning: Science and Technology","name":"journal_title","label":"Journal Title"},{"value":"paper","name":"article_type","label":"Article Type"},{"value":"\u00a9 2022 The Author(s). Published by IOP Publishing Ltd","name":"copyright_information","label":"Copyright Information"},{"value":"2022-08-11","name":"date_received","label":"Date Received","group":{"name":"publication_dates","label":"Publication dates"}},{"value":"2022-11-28","name":"date_accepted","label":"Date Accepted","group":{"name":"publication_dates","label":"Publication dates"}},{"value":"2022-12-19","name":"date_epub","label":"Online publication date","group":{"name":"publication_dates","label":"Publication dates"}}]}}