{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,1,21]],"date-time":"2026-01-21T07:26:40Z","timestamp":1768980400160,"version":"3.49.0"},"reference-count":18,"publisher":"Fuji Technology Press Ltd.","issue":"1","funder":[{"DOI":"10.13039\/501100001691","name":"Japan Society for the Promotion of Science","doi-asserted-by":"publisher","award":["25K07685"],"award-info":[{"award-number":["25K07685"]}],"id":[{"id":"10.13039\/501100001691","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["JACIII","J. Adv. Comput. Intell. Intell. Inform."],"published-print":{"date-parts":[[2026,1,20]]},"abstract":"<jats:p>Recently, deep reinforcement learning has demonstrated notable success across various tasks. This success is largely attributable to its high expressive power. However, operating large-scale neural networks requires substantial power consumption. This is particularly a challenge for applications with limited power budgets, such as robotic control, in which energy efficiency is paramount. By contrast, spiking neural networks (SNNs) have garnered considerable attention owing to their high energy efficiency, particularly when implemented on dedicated neuromorphic hardware. Despite these advantages, conventional methods for integrating SNNs into deep reinforcement learning frameworks frequently struggle with training stability. To address this issue, this study introduces a novel algorithm that incorporates SNNs within the actor networks of a twin-delayed deep deterministic policy gradient architecture. To further enhance the performance, a burn-in strategy inspired by recurrent experience replay in distributed reinforcement learning was implemented. This study introduces a burn-in strategy that stabilizes learning and reduces variance by addressing the issue of stale membrane potentials stored in the replay buffer. In this strategy, membrane potentials computed using outdated network parameters are passed through the current network to align them with the updated parameters, to improve the accuracy of action value estimations and strengthen training stability. Furthermore, loss-adjusted prioritization was incorporated to improve the learning efficiency and stability. Experimental evaluations conducted in OpenAI Gym environments demonstrated that the proposed method yielded superior rewards compared with conventional approaches.<\/jats:p>","DOI":"10.20965\/jaciii.2026.p0301","type":"journal-article","created":{"date-parts":[[2026,1,19]],"date-time":"2026-01-19T15:02:06Z","timestamp":1768834926000},"page":"301-310","source":"Crossref","is-referenced-by-count":0,"title":["Enhancing Deep Reinforcement Learning in Spiking Neural Networks via a Burn-In Strategy"],"prefix":"10.20965","volume":"30","author":[{"given":"Takahiro","family":"Iwata","sequence":"first","affiliation":[{"name":"Graduate School of Computer Science, Chiba Institute of Technology, 2-17-1 Tsudanuma, Narashino, Chiba 275-0016, Japan"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Seankein","family":"Yoshioka","sequence":"additional","affiliation":[{"name":"Graduate School of Computer Science, Chiba Institute of Technology, 2-17-1 Tsudanuma, Narashino, Chiba 275-0016, Japan"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Hiroto","family":"Takigasaki","sequence":"additional","affiliation":[{"name":"Graduate School of Computer Science, Chiba Institute of Technology, 2-17-1 Tsudanuma, Narashino, Chiba 275-0016, Japan"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-1408-2927","authenticated-orcid":true,"given":"Daisuke","family":"Miki","sequence":"additional","affiliation":[{"name":"Department of Computer Science, Chiba Institute of Technology, 2-17-1 Tsudanuma, Narashino, Chiba 275-0016, Japan"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"8550","published-online":{"date-parts":[[2026,1,20]]},"reference":[{"key":"key-10.20965\/jaciii.2026.p0301-1","unstructured":"V. Mnih, K. Kavukcuoglu, D. Silver, A. Graves, I. Antonoglou, D. Wierstra, and M. A. Riedmiller, \u201cPlaying atari with deep reinforcement learning,\u201d arXiv:1312.5602, 2013. https:\/\/doi.org\/10.48550\/arXiv.1312.5602"},{"key":"key-10.20965\/jaciii.2026.p0301-2","doi-asserted-by":"crossref","unstructured":"D. Silver, J. Schrittwieser, K. Simonyan, I. Antonoglou, A. Huang, A. Guez, T. Hubert, L. Baker, M. Lai, A. Bolton et al., \u201cMastering the game of go without human knowledge,\u201d Nature, Vol.550, No.7676, pp. 354-359, 2017. https:\/\/doi.org\/10.1038\/nature24270","DOI":"10.1038\/nature24270"},{"key":"key-10.20965\/jaciii.2026.p0301-3","doi-asserted-by":"crossref","unstructured":"T. Haarnoja, S. Ha, A. Zhou, J. Tan, G. Tucker, and S. Levine, \u201cLearning to walk via deep reinforcement learning,\u201d Robotics: Science and Systems XV, 2019. https:\/\/doi.org\/10.15607\/RSS.2019.XV.011","DOI":"10.15607\/RSS.2019.XV.011"},{"key":"key-10.20965\/jaciii.2026.p0301-4","unstructured":"S. Fujimoto, H. van Hoof, and D. Meger, \u201cAddressing function approximation error in actor-critic methods,\u201d Proc. of the 35th Int. Conf. on Machine Learning, 2018."},{"key":"key-10.20965\/jaciii.2026.p0301-5","doi-asserted-by":"crossref","unstructured":"M. Horowitz, \u201c1.1 computing\u2019 energy problem (and what we can do about it),\u201d 2014 IEEE Int. Solid-State Circuits Conf. Digest of Technical Papers (ISSCC), pp. 10-14, 2014. https:\/\/doi.org\/10.1109\/ISSCC.2014.6757323","DOI":"10.1109\/ISSCC.2014.6757323"},{"key":"key-10.20965\/jaciii.2026.p0301-6","unstructured":"Intel Corporation, \u201cNeuromorphic Computing and Engineering, Next Wave of AI Capabilities.\u201d https:\/\/www.intel.com\/content\/www\/us\/en\/research\/neuromorphic-computing.html [Accessed April 1, 2025]"},{"key":"key-10.20965\/jaciii.2026.p0301-7","unstructured":"G. Tang, N. Kumar, R. Yoo, and K. Michmizos, \u201cDeep reinforcement learning with population-coded spiking neural network for continuous control,\u201d Proc. of the 2020 Conf. on Robot Learning, Vol.155, pp. 2016-2029, 2021."},{"key":"key-10.20965\/jaciii.2026.p0301-8","doi-asserted-by":"crossref","unstructured":"K. Naya, K. Kutsuzawa, D. Owaki, and M. Hayashibe, \u201cSpiking neural network discovers energy-efficient hexapod motion in deep reinforcement learning,\u201d IEEE Access, Vol.9, pp. 150345-150354, 2021. https:\/\/doi.org\/10.1109\/ACCESS.2021.3126311","DOI":"10.1109\/ACCESS.2021.3126311"},{"key":"key-10.20965\/jaciii.2026.p0301-9","doi-asserted-by":"crossref","unstructured":"T. Iwata, S. Yoshioka, and D. Miki, \u201cApplying burn-in strategy to deep reinforcement learning with spiking neural networks,\u201d 2024 Joint 13th Int. Conf. on Soft Computing and Intelligent Systems and 25th Int. Symp. on Advanced Intelligent Systems (SCIS&ISIS), 2024. https:\/\/doi.org\/10.1109\/SCISISIS61014.2024.10760159","DOI":"10.1109\/SCISISIS61014.2024.10760159"},{"key":"key-10.20965\/jaciii.2026.p0301-10","unstructured":"T. P. Lillicrap, J. J. Hunt, A. Pritzel, N. Heess, T. Erez, Y. Tassa, D. Silver, and D. Wierstra, \u201cContinuous control with deep reinforcement learning,\u201d arXiv:1509.02971, 2019. https:\/\/doi.org\/10.48550\/arXiv.1509.02971"},{"key":"key-10.20965\/jaciii.2026.p0301-11","unstructured":"S. Fujimoto, D. Meger, and D. Precup, \u201cAn equivalence between loss functions and non-uniform sampling in experience replay,\u201d Proc. of the 34th Int. Conf. on Neural Information Processing Systems (NIPS \u201920), 2020."},{"key":"key-10.20965\/jaciii.2026.p0301-12","unstructured":"S. Kapturowski, G. Ostrovski, W. Dabney, J. Quan, and R. Munos, \u201cRecurrent experience replay in distributed reinforcement learning,\u201d Int. Conf. on Learning Representations, 2019."},{"key":"key-10.20965\/jaciii.2026.p0301-13","doi-asserted-by":"crossref","unstructured":"S. Hochreiter and J. Schmidhuber, \u201cLong short-term memory,\u201d Neural Computation, Vol.9, No.8, pp. 1735-1780, 1997. https:\/\/doi.org\/10.1162\/neco.1997.9.8.1735","DOI":"10.1162\/neco.1997.9.8.1735"},{"key":"key-10.20965\/jaciii.2026.p0301-14","unstructured":"D. Horgan, J. Quan, D. Budden, G. Barth-Maron, M. Hessel, H. van Hasselt, and D. Silver, \u201cDistributed prioritized experience replay,\u201d Int. Conf. on Learning Representations, 2018."},{"key":"key-10.20965\/jaciii.2026.p0301-15","unstructured":"\u201cOpenAI Gym.\u201d https:\/\/gymnasium.farama.org\/environments\/mujoco\/ [Accessed April 1, 2025]"},{"key":"key-10.20965\/jaciii.2026.p0301-16","unstructured":"PyTorch Foundation, \u201cPytorch.\u201d https:\/\/pytorch.org\/ [Accessed April 1, 2025]"},{"key":"key-10.20965\/jaciii.2026.p0301-17","unstructured":"J. K. Eshraghian, \u201csnntorch.\u201d https:\/\/snntorch.readthedocs.io\/en\/latest\/snntorch.html [Accessed April 1, 2025]"},{"key":"key-10.20965\/jaciii.2026.p0301-18","unstructured":"PyTorch Foundation, \u201cAdam.\u201d https:\/\/pytorch.org\/docs\/stable\/generated\/torch.optim.Adam.html [Accessed April 1, 2025]"}],"container-title":["Journal of Advanced Computational Intelligence and Intelligent Informatics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.fujipress.jp\/main\/wp-content\/themes\/Fujipress\/hyosetsu.php?ppno=jacii003000010027","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,1,19]],"date-time":"2026-01-19T15:04:03Z","timestamp":1768835043000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.fujipress.jp\/jaciii\/jc\/jacii003000010301"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,1,20]]},"references-count":18,"journal-issue":{"issue":"1","published-online":{"date-parts":[[2026,1,20]]},"published-print":{"date-parts":[[2026,1,20]]}},"URL":"https:\/\/doi.org\/10.20965\/jaciii.2026.p0301","relation":{},"ISSN":["1883-8014","1343-0130"],"issn-type":[{"value":"1883-8014","type":"electronic"},{"value":"1343-0130","type":"print"}],"subject":[],"published":{"date-parts":[[2026,1,20]]}}}