{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,1,9]],"date-time":"2026-01-09T10:53:48Z","timestamp":1767956028682,"version":"3.49.0"},"reference-count":55,"publisher":"IOP Publishing","issue":"1","license":[{"start":{"date-parts":[[2026,1,9]],"date-time":"2026-01-09T00:00:00Z","timestamp":1767916800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"},{"start":{"date-parts":[[2026,1,9]],"date-time":"2026-01-09T00:00:00Z","timestamp":1767916800000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/iopscience.iop.org\/info\/page\/text-and-data-mining"},{"start":{"date-parts":[[2026,1,9]],"date-time":"2026-01-09T00:00:00Z","timestamp":1767916800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/publishingsupport.iopscience.iop.org\/questions\/alternative-author-rights-policies\/"},{"start":{"date-parts":[[2026,1,9]],"date-time":"2026-01-09T00:00:00Z","timestamp":1767916800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/publishingsupport.iopscience.iop.org\/questions\/alternative-author-rights-policies\/"}],"funder":[{"name":"Research Computing Clusters at Old Dominion University"},{"name":"Hampton Roads Biomedical Research Consortium"},{"DOI":"10.13039\/100000015","name":"U.S. Department of Energy","doi-asserted-by":"crossref","award":["DE-AC05-06OR23177"],"award-info":[{"award-number":["DE-AC05-06OR23177"]}],"id":[{"id":"10.13039\/100000015","id-type":"DOI","asserted-by":"crossref"}]},{"name":"Jefferson Science Associates"}],"content-domain":{"domain":["iopscience.iop.org"],"crossmark-restriction":false},"short-container-title":["Mach. Learn.: Sci. Technol."],"published-print":{"date-parts":[[2026,2,1]]},"abstract":"<jats:title>Abstract<\/jats:title>\n                  <jats:p>We present a reinforcement learning (RL) framework for optimizing particle accelerator experiments that builds explainable physics-based constraints on agent behavior. The goal is to increase transparency and trust by letting users verify that the agent\u2019s decision-making process incorporates suitable physics. Our algorithm uses a learnable surrogate function for physical observables, such as energy, and uses them to fine-tune how actions are chosen. This surrogate can be represented by a neural network or by an interpretable sparse dictionary model. We test our algorithm on a range of particle accelerator optimization environments designed to emulate the Continuous Electron Beam Accelerator Facility at Jefferson Lab. By examining the mathematical form of the learned constraint function, we are able to confirm the agent has learned to use the established physics of each environment. In addition, we find that the introduction of a physics-based surrogate enables our RL algorithms to reliably converge for difficult high-dimensional accelerator optimization environments.<\/jats:p>","DOI":"10.1088\/2632-2153\/ae2fa8","type":"journal-article","created":{"date-parts":[[2025,12,19]],"date-time":"2025-12-19T22:53:48Z","timestamp":1766184828000},"page":"015005","update-policy":"https:\/\/doi.org\/10.1088\/crossmark-policy","source":"Crossref","is-referenced-by-count":0,"title":["Explainable physics-based constraints on reinforcement learning for accelerator optimization"],"prefix":"10.1088","volume":"7","author":[{"ORCID":"https:\/\/orcid.org\/0000-0003-4162-0276","authenticated-orcid":true,"given":"Jonathan","family":"Colen","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-3475-2871","authenticated-orcid":true,"given":"Malachi","family":"Schram","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Kishansingh","family":"Rajput","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Armen","family":"Kasparian","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"266","published-online":{"date-parts":[[2026,1,9]]},"reference":[{"key":"mlstae2fa8bib1","author":"Sutton","year":"2020","edition":"2nd edn","type":"book"},{"key":"mlstae2fa8bib2","doi-asserted-by":"publisher","first-page":"1140","DOI":"10.1126\/science.aar6404","type":"journal-article","article-title":"A general reinforcement learning algorithm that masters chess, shogi and Go through self-play","volume":"362","author":"Silver","year":"2018","journal-title":"Science"},{"key":"mlstae2fa8bib3","doi-asserted-by":"publisher","first-page":"982","DOI":"10.1038\/s41586-023-06419-4","type":"journal-article","article-title":"Champion-level drone racing using deep reinforcement learning","volume":"620","author":"Kaufmann","year":"2023","journal-title":"Nature"},{"key":"mlstae2fa8bib4","doi-asserted-by":"publisher","DOI":"10.1016\/j.aei.2022.101787","type":"journal-article","article-title":"Robotics in construction: a critical review of the reinforcement learning and imitation learning paradigms","volume":"54","author":"Manuel Davila Delgado","year":"2022","journal-title":"Adv. Eng. Inform."},{"key":"mlstae2fa8bib5","doi-asserted-by":"publisher","first-page":"1238","DOI":"10.1177\/0278364913495721","type":"journal-article","article-title":"Reinforcement learning in robotics: a survey","volume":"32","author":"Kober","year":"2013","journal-title":"Int. J. Robot. Res."},{"key":"mlstae2fa8bib6","doi-asserted-by":"publisher","DOI":"10.1103\/PhysRevAccelBeams.23.124801","type":"journal-article","article-title":"Sample-efficient reinforcement learning for CERN accelerator control","volume":"23","author":"Kain","year":"2020","journal-title":"Phys. Rev. Accel. Beams"},{"key":"mlstae2fa8bib7","doi-asserted-by":"publisher","DOI":"10.1038\/s41598-024-66263-y","type":"journal-article","article-title":"Reinforcement learning-trained optimisers and Bayesian optimisation for online particle accelerator tuning","volume":"14","author":"Kaiser","year":"2024","journal-title":"Sci. Rep."},{"key":"mlstae2fa8bib8","doi-asserted-by":"publisher","DOI":"10.1103\/PhysRevResearch.3.033291","type":"journal-article","article-title":"Learning to control active matter","volume":"3","author":"Falk","year":"2021","journal-title":"Phys. Rev. Res."},{"key":"mlstae2fa8bib9","doi-asserted-by":"crossref","DOI":"10.2139\/ssrn.4597487","type":"preprint","article-title":"A survey on physics informed reinforcement learning: review and open problems","author":"Banerjee","year":"2023"},{"key":"mlstae2fa8bib10","doi-asserted-by":"publisher","first-page":"686","DOI":"10.1016\/j.jcp.2018.10.045","type":"journal-article","article-title":"Physics-informed neural networks: a deep learning framework for solving forward and inverse problems involving nonlinear partial differential equations","volume":"378","author":"Raissi","year":"2019","journal-title":"J. Comput. Phys."},{"key":"mlstae2fa8bib11","doi-asserted-by":"publisher","first-page":"422","DOI":"10.1038\/s42254-021-00314-5","type":"journal-article","article-title":"Physics-informed machine learning","volume":"3","author":"Karniadakis","year":"2021","journal-title":"Nat. Rev. Phys."},{"key":"mlstae2fa8bib12","doi-asserted-by":"publisher","first-page":"481","DOI":"10.1016\/j.cell.2023.11.041","type":"journal-article","article-title":"Machine learning interpretable models of cell mechanics from protein images","volume":"187","author":"Schmitt","year":"2024","journal-title":"Cell"},{"key":"mlstae2fa8bib13","doi-asserted-by":"publisher","DOI":"10.1073\/pnas.2508692122","type":"journal-article","article-title":"Sociohydrodynamics: Data-driven modeling of social behavior","volume":"122","author":"Seara","year":"2025","journal-title":"Proc. Natl. Acad. Sci."},{"key":"mlstae2fa8bib14","doi-asserted-by":"publisher","first-page":"81","DOI":"10.1126\/science.1165893","type":"journal-article","article-title":"Distilling free-form natural laws from experimental data","volume":"324","author":"Schmidt","year":"2009","journal-title":"Science"},{"key":"mlstae2fa8bib15","doi-asserted-by":"publisher","first-page":"3932","DOI":"10.1073\/pnas.1517384113","type":"journal-article","article-title":"Discovering governing equations from data by sparse identification of nonlinear dynamical systems","volume":"113","author":"Brunton","year":"2016","journal-title":"Proc. Natl Acad. Sci."},{"key":"mlstae2fa8bib16","doi-asserted-by":"publisher","first-page":"477","DOI":"10.1146\/annurev-fluid-010719-060214","type":"journal-article","article-title":"Machine learning for fluid mechanics","volume":"52","author":"Brunton","year":"2020","journal-title":"Annu. Rev. Fluid Mech."},{"key":"mlstae2fa8bib17","doi-asserted-by":"publisher","DOI":"10.1073\/pnas.1906995116","type":"journal-article","article-title":"Data-driven discovery of coordinates and governing equations","volume":"116","author":"Champion","year":"2019","journal-title":"Proc. Natl Acad. Sci."},{"key":"mlstae2fa8bib18","doi-asserted-by":"publisher","DOI":"10.1109\/ACCESS.2020.3023625","type":"journal-article","article-title":"A unified sparse optimization framework to learn parsimonious physics-informed models from data","volume":"8","author":"Champion","year":"2020","journal-title":"IEEE Access"},{"key":"mlstae2fa8bib19","doi-asserted-by":"publisher","DOI":"10.1073\/pnas.2206994120","type":"journal-article","article-title":"Learning hydrodynamic equations for active matter from particle simulations and experiments","volume":"120","author":"Supekar","year":"2023","journal-title":"Proc. Natl Acad. Sci."},{"key":"mlstae2fa8bib20","doi-asserted-by":"publisher","DOI":"10.1103\/PhysRevLett.129.258001","type":"journal-article","article-title":"Data-driven discovery of active nematic hydrodynamics","volume":"129","author":"Joshi","year":"2022","journal-title":"Phys. Rev. Lett."},{"key":"mlstae2fa8bib21","doi-asserted-by":"publisher","first-page":"eabq6120","DOI":"10.1126\/sciadv.abq6120","type":"journal-article","article-title":"Physically informed data-driven modeling of active nematics","volume":"9","author":"Golden","year":"2023","journal-title":"Sci. Adv."},{"key":"mlstae2fa8bib22","doi-asserted-by":"publisher","DOI":"10.1103\/PhysRevLett.133.107301","type":"journal-article","article-title":"Interpreting neural operators: how nonlinear waves propagate in nonreciprocal solids","volume":"133","author":"Colen","year":"2024","journal-title":"Phys. Rev. Lett."},{"key":"mlstae2fa8bib23","doi-asserted-by":"crossref","DOI":"10.52843\/cassyni.hwpvkj","type":"preprint","article-title":"SINDy-RL: interpretable and efficient model-based reinforcement learning","author":"Zolman","year":"2024"},{"key":"mlstae2fa8bib24","doi-asserted-by":"publisher","DOI":"10.1103\/PhysRevAccelBeams.27.084802","type":"journal-article","article-title":"The continuous electron beam accelerator facility at 12 GeV","volume":"27","author":"Adderley","year":"2024","journal-title":"Phys. Rev. Accel. Beams"},{"key":"mlstae2fa8bib25","doi-asserted-by":"publisher","first-page":"240","DOI":"10.1109\/CCTA48906.2021.9658806)","type":"conference-proceedings","article-title":"Machine learning-based anomaly detection for particle accelerators","author":"Marcato","year":"2021"},{"key":"mlstae2fa8bib26","doi-asserted-by":"publisher","DOI":"10.1016\/j.mlwa.2023.100484","type":"journal-article","article-title":"Multi-module-based CVAE to predict HVCM faults in the SNS accelerator","volume":"13","author":"Alanazi","year":"2023","journal-title":"Mach. Learn. Appl."},{"key":"mlstae2fa8bib27","article-title":"Anomaly detection in particle accelerators using autoencoders","author":"Edelen","year":"2021","type":"preprint"},{"key":"mlstae2fa8bib28","doi-asserted-by":"publisher","DOI":"10.1103\/PhysRevAccelBeams.25.122802","type":"journal-article","article-title":"Uncertainty aware anomaly detection to predict errant beam pulses in the Oak ridge spallation neutron source accelerator","volume":"25","author":"Blokland","year":"2022","journal-title":"Phys. Rev. Accel. Beams"},{"key":"mlstae2fa8bib29","doi-asserted-by":"publisher","DOI":"10.1103\/PhysRevAccelBeams.24.062801","type":"journal-article","article-title":"Multiobjective Bayesian optimization for online accelerator tuning","volume":"24","author":"Roussel","year":"2021","journal-title":"Phys. Rev. Accel. Beams"},{"key":"mlstae2fa8bib30","doi-asserted-by":"publisher","DOI":"10.1103\/PhysRevSTAB.16.102803","type":"journal-article","article-title":"Model-independent particle accelerator tuning","volume":"16","author":"Scheinker","year":"2013","journal-title":"Phys. Rev. ST Accel. Beams"},{"key":"mlstae2fa8bib31","doi-asserted-by":"publisher","DOI":"10.1063\/5.0003423","type":"journal-article","article-title":"Online multi-objective particle accelerator optimization of the AWAKE electron beam line for simultaneous emittance and orbit control","volume":"10","author":"Scheinker","year":"2020","journal-title":"AIP Adv."},{"key":"mlstae2fa8bib32","doi-asserted-by":"publisher","DOI":"10.1088\/2632-2153\/adc221","type":"journal-article","article-title":"Harnessing the power of gradient-based simulations for multi-objective optimization in particle accelerators","volume":"6","author":"Rajput","year":"2025","journal-title":"Mach. learn.: sci. technol."},{"key":"mlstae2fa8bib33","doi-asserted-by":"publisher","DOI":"10.1103\/PhysRevSTAB.17.101003","type":"journal-article","article-title":"Simultaneous optimization of the cavity heat load and trip rates in linacs using a genetic algorithm","volume":"17","author":"Terzi\u0107","year":"2014","journal-title":"Phys. Rev. ST Accel. Beams"},{"key":"mlstae2fa8bib34","doi-asserted-by":"crossref","DOI":"10.1080\/14697688.2022.2062431","type":"preprint","article-title":"Deep differentiable reinforcement learning and optimal trading","author":"Jaisson","year":"2022"},{"key":"mlstae2fa8bib35","article-title":"The information bottleneck method","author":"Tishby","year":"2000","type":"preprint"},{"key":"mlstae2fa8bib36","doi-asserted-by":"publisher","DOI":"10.1103\/PhysRevLett.132.197201","type":"journal-article","article-title":"Machine-learning optimized measurements of chaotic dynamical systems via the information bottleneck","volume":"132","author":"Murphy","year":"2024","journal-title":"Phys. Rev. Lett."},{"key":"mlstae2fa8bib37","doi-asserted-by":"publisher","DOI":"10.1073\/pnas.2312988121","type":"journal-article","article-title":"Information decomposition in complex systems via machine learning","volume":"121","author":"Murphy","year":"2024","journal-title":"Proc. Natl Acad. Sci."},{"key":"mlstae2fa8bib38","doi-asserted-by":"publisher","DOI":"10.1101\/2024.04.19.590281)","type":"other","article-title":"Information theory for data-driven model reduction in physics and biology","author":"Schmitt","year":"2024"},{"key":"mlstae2fa8bib39","doi-asserted-by":"publisher","first-page":"578","DOI":"10.1038\/s41567-018-0081-4","type":"journal-article","article-title":"Mutual information, neural networks and the renormalization group","volume":"14","author":"Koch-Janusz","year":"2018","journal-title":"Nat. Phys."},{"key":"mlstae2fa8bib40","doi-asserted-by":"publisher","DOI":"10.1103\/PhysRevLett.127.240603","type":"journal-article","article-title":"Statistical physics through the lens of real-space mutual information","volume":"127","author":"G\u00f6kmen","year":"2021","journal-title":"Phys. Rev. Lett."},{"key":"mlstae2fa8bib41","doi-asserted-by":"publisher","first-page":"469","DOI":"10.1038\/nphys3644","type":"journal-article","article-title":"A structural approach to relaxation in glassy liquids","volume":"12","author":"Schoenholz","year":"2016","journal-title":"Nat. Phys."},{"key":"mlstae2fa8bib42","doi-asserted-by":"publisher","first-page":"eabk0644","DOI":"10.1126\/sciadv.abk0644","type":"journal-article","article-title":"Analyses of internal structures and defects in materials using physics-informed neural networks","volume":"8","author":"Zhang","year":"2022","journal-title":"Sci. Adv."},{"key":"mlstae2fa8bib43","doi-asserted-by":"publisher","first-page":"52","DOI":"10.1109\/TMBMC.2016.2633265","type":"journal-article","article-title":"Inferring biological networks by sparse identification of nonlinear dynamics","volume":"2","author":"Mangan","year":"2016","journal-title":"IEEE Trans. Mol. Biol. Multi-Scale Commun."},{"key":"mlstae2fa8bib44","doi-asserted-by":"publisher","first-page":"636","DOI":"10.1038\/s42256-022-00503-6","type":"journal-article","article-title":"Learning biophysical determinants of cell fate with deep neural networks","volume":"4","author":"Soelistyo","year":"2022","journal-title":"Nat. Mach. Intell."},{"key":"mlstae2fa8bib45","doi-asserted-by":"publisher","DOI":"10.7554\/eLife.68679","type":"journal-article","article-title":"Learning developmental mode dynamics from single-cell trajectories","volume":"10","author":"Romeo","year":"2021","journal-title":"eLife"},{"key":"mlstae2fa8bib46","article-title":"A longitudinal study of field emission in CEBAF\u2019s SRF cavities 1995-2015","author":"Benesch","year":"2015","type":"preprint"},{"key":"mlstae2fa8bib47","article-title":"Addressing function approximation error in actor-critic methods","author":"Fujimoto","year":"2018","type":"preprint"},{"key":"mlstae2fa8bib48","doi-asserted-by":"publisher","first-page":"119","DOI":"10.1145\/3453474","type":"journal-article","article-title":"The hypervolume indicator: computational problems and algorithms","volume":"54","author":"Guerreiro","year":"2021","journal-title":"ACM Comput. Surv."},{"key":"mlstae2fa8bib49","article-title":"Offline reinforcement learning: tutorial, review, and perspectives on open problems","author":"Levine","year":"2020","type":"preprint"},{"key":"mlstae2fa8bib50","article-title":"Hybrid RL: using both offline and online data can make RL efficient","author":"Song","year":"2023","type":"preprint"},{"key":"mlstae2fa8bib51","first-page":"1702","type":"conference-proceedings","article-title":"Offline-to-online reinforcement learning via balanced replay and pessimistic Q-ensemble","author":"Lee","year":"2022"},{"key":"mlstae2fa8bib52","article-title":"Learning a conserved mechanism for early neuroectoderm morphogenesis","author":"Lefebvre","year":"2024","type":"preprint"},{"key":"mlstae2fa8bib53","article-title":"Cornerstones are the key stones: using interpretable machine learning to probe the clogging process in 2D granular hoppers","author":"Hanlan","year":"2024","type":"preprint"},{"key":"mlstae2fa8bib54","article-title":"Tailoring interactions between active nematic defects with reinforcement learning","author":"Floyd","year":"2024","type":"preprint"},{"key":"mlstae2fa8bib55","article-title":"Code supporting \u201cExplainable physics-based constraints on reinforcement learning for accelerator optimization\u201d","author":"Colen","year":"2025","type":"web-resource"}],"container-title":["Machine Learning: Science and Technology"],"original-title":[],"link":[{"URL":"https:\/\/iopscience.iop.org\/article\/10.1088\/2632-2153\/ae2fa8","content-type":"text\/html","content-version":"am","intended-application":"text-mining"},{"URL":"https:\/\/iopscience.iop.org\/article\/10.1088\/2632-2153\/ae2fa8\/pdf","content-type":"application\/pdf","content-version":"am","intended-application":"text-mining"},{"URL":"https:\/\/iopscience.iop.org\/article\/10.1088\/2632-2153\/ae2fa8","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/iopscience.iop.org\/article\/10.1088\/2632-2153\/ae2fa8\/pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/iopscience.iop.org\/article\/10.1088\/2632-2153\/ae2fa8\/pdf","content-type":"application\/pdf","content-version":"am","intended-application":"syndication"},{"URL":"https:\/\/iopscience.iop.org\/article\/10.1088\/2632-2153\/ae2fa8\/pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/iopscience.iop.org\/article\/10.1088\/2632-2153\/ae2fa8\/pdf","content-type":"application\/pdf","content-version":"am","intended-application":"similarity-checking"},{"URL":"https:\/\/iopscience.iop.org\/article\/10.1088\/2632-2153\/ae2fa8\/pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,1,9]],"date-time":"2026-01-09T09:38:00Z","timestamp":1767951480000},"score":1,"resource":{"primary":{"URL":"https:\/\/iopscience.iop.org\/article\/10.1088\/2632-2153\/ae2fa8"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,1,9]]},"references-count":55,"journal-issue":{"issue":"1","published-online":{"date-parts":[[2026,1,9]]},"published-print":{"date-parts":[[2026,2,1]]}},"URL":"https:\/\/doi.org\/10.1088\/2632-2153\/ae2fa8","relation":{},"ISSN":["2632-2153"],"issn-type":[{"value":"2632-2153","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,1,9]]},"assertion":[{"value":"Explainable physics-based constraints on reinforcement learning for accelerator optimization","name":"article_title","label":"Article Title"},{"value":"Machine Learning: Science and Technology","name":"journal_title","label":"Journal Title"},{"value":"paper","name":"article_type","label":"Article Type"},{"value":"\u00a9 2026 The Author(s). Published by IOP Publishing Ltd","name":"copyright_information","label":"Copyright Information"},{"value":"2025-09-04","name":"date_received","label":"Date Received","group":{"name":"publication_dates","label":"Publication dates"}},{"value":"2025-12-19","name":"date_accepted","label":"Date Accepted","group":{"name":"publication_dates","label":"Publication dates"}},{"value":"2026-01-09","name":"date_epub","label":"Online publication date","group":{"name":"publication_dates","label":"Publication dates"}}]}}