{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,7]],"date-time":"2026-07-07T12:24:53Z","timestamp":1783427093236,"version":"3.54.6"},"reference-count":71,"publisher":"Verein zur Forderung des Open Access Publizierens in den Quantenwissenschaften","license":[{"start":{"date-parts":[[2025,3,12]],"date-time":"2025-03-12T00:00:00Z","timestamp":1741737600000},"content-version":"unspecified","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"DOI":"10.13039\/501100001809","name":"the National Natural Science Foundation of China","doi-asserted-by":"crossref","award":["92265208"],"award-info":[{"award-number":["92265208"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]},{"name":"the National Key R&D Program of China","award":["2018YFA0306703"],"award-info":[{"award-number":["2018YFA0306703"]}]}],"content-domain":{"domain":["quantum-journal.org"],"crossmark-restriction":false},"short-container-title":["Quantum"],"abstract":"<jats:p>Quantum reinforcement learning (QRL) is a promising paradigm for near-term quantum devices. While existing QRL methods have shown success in discrete action spaces, extending these techniques to continuous domains is challenging due to the curse of dimensionality introduced by discretization. To overcome this limitation, we introduce a quantum Deep Deterministic Policy Gradient (DDPG) algorithm that efficiently addresses both classical and quantum sequential decision problems in continuous action spaces. Moreover, our approach facilitates single-shot quantum state generation: a one-time optimization produces a model that outputs the control sequence required to drive a fixed initial state to any desired target state. In contrast, conventional quantum control methods demand separate optimization for each target state. We demonstrate the effectiveness of our method through simulations and discuss its potential applications in quantum control.<\/jats:p>","DOI":"10.22331\/q-2025-03-12-1660","type":"journal-article","created":{"date-parts":[[2025,3,12]],"date-time":"2025-03-12T08:43:38Z","timestamp":1741769018000},"page":"1660","update-policy":"https:\/\/doi.org\/10.22331\/q-crossmark-policy-page","source":"Crossref","is-referenced-by-count":19,"title":["Quantum reinforcement learning in continuous action space"],"prefix":"10.22331","volume":"9","author":[{"given":"Shaojun","family":"Wu","sequence":"first","affiliation":[{"name":"Institute of Fundamental and Frontier Sciences, University of Electronic Science and Technology of China, Chengdu, 610051, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Shan","family":"Jin","sequence":"additional","affiliation":[{"name":"Institute of Fundamental and Frontier Sciences, University of Electronic Science and Technology of China, Chengdu, 610051, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Dingding","family":"Wen","sequence":"additional","affiliation":[{"name":"Institute of Fundamental and Frontier Sciences, University of Electronic Science and Technology of China, Chengdu, 610051, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Donghong","family":"Han","sequence":"additional","affiliation":[{"name":"Institute of Fundamental and Frontier Sciences, University of Electronic Science and Technology of China, Chengdu, 610051, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Xiaoting","family":"Wang","sequence":"additional","affiliation":[{"name":"Institute of Fundamental and Frontier Sciences, University of Electronic Science and Technology of China, Chengdu, 610051, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"9598","published-online":{"date-parts":[[2025,3,12]]},"reference":[{"key":"0","unstructured":"Richard S. Sutton and Andrew G. Barto. Reinforcement Learning: An Introduction. The MIT Press, second edition, 2018. URL http:\/\/incompleteideas.net\/book\/the-book-2nd.html."},{"key":"1","doi-asserted-by":"publisher","unstructured":"David Silver, Julian Schrittwieser, Karen Simonyan, Ioannis Antonoglou, Aja Huang, Arthur Guez, Thomas Hubert, Lucas Baker, Matthew Lai, Adrian Bolton, et al. Mastering the game of go without human knowledge. Nature(London), 550 (7676): 354\u2013359, 2017. 10.1038\/nature24270. URL https:\/\/doi.org\/10.1038\/nature24270.","DOI":"10.1038\/nature24270"},{"key":"2","doi-asserted-by":"publisher","unstructured":"Mnih Volodymyr, Kavukcuoglu Koray, Silver David, Graves Alex, Antonoglou Ioannis, Wierstra Daan, and Riedmiller Martin. Playing atari with deep reinforcement learning. 2013. 10.48550\/ARXIV.1312.5602. URL http:\/\/arxiv.org\/abs\/1312.5602.","DOI":"10.48550\/ARXIV.1312.5602"},{"key":"3","doi-asserted-by":"publisher","unstructured":"David Silver, Aja Huang, Chris J Maddison, Arthur Guez, Laurent Sifre, George Van Den Driessche, Julian Schrittwieser, Ioannis Antonoglou, Veda Panneershelvam, Marc Lanctot, Sander Dieleman, Dominik Grewe, John Nham, Nal Kalchbrenner, Ilya Sutskever, Timothy Lillicrap, Madeleine Leach, Koray Kavukcuoglu, Thore Graepel, and Demis Hassabis. Mastering the game of go with deep neural networks and tree search. Nature(London), 529 (7587): 484\u2013489, 2016. 10.1038\/nature16961. URL https:\/\/doi.org\/10.1038\/nature16961.","DOI":"10.1038\/nature16961"},{"key":"4","doi-asserted-by":"publisher","unstructured":"Jan Peters, Sethu Vijayakumar, and Stefan Schaal. Reinforcement learning for humanoid robotics. In Proceedings of the third IEEE-RAS international conference on humanoid robots, pages 1\u201320, 2003. 10.1109\/LARS\/SBR\/WRE51543.2020.9307084. URL https:\/\/ieeexplore.ieee.org\/document\/9307084.","DOI":"10.1109\/LARS\/SBR\/WRE51543.2020.9307084"},{"key":"5","doi-asserted-by":"publisher","unstructured":"Yan Duan, Xi Chen, Rein Houthooft, John Schulman, and Pieter Abbeel. Benchmarking deep reinforcement learning for continuous control. In Proceedings of the 33rd International Conference on International Conference on Machine Learning - Volume 48, ICML&apos;16, page 1329\u20131338, 2016. 10.5555\/3045390.3045531. URL https:\/\/dl.acm.org\/doi\/10.5555\/3045390.3045531.","DOI":"10.5555\/3045390.3045531"},{"key":"6","doi-asserted-by":"publisher","unstructured":"Hendrik Poulsen Nautrup, Nicolas Delfosse, Vedran Dunjko, Hans J. Briegel, and Nicolai Friis. Optimizing Quantum Error Correction Codes with Reinforcement Learning. Quantum, 3: 215, December 2019. ISSN 2521-327X. 10.22331\/q-2019-12-16-215. URL https:\/\/doi.org\/10.22331\/q-2019-12-16-215.","DOI":"10.22331\/q-2019-12-16-215"},{"key":"7","doi-asserted-by":"publisher","unstructured":"Philip Andreasson, Joel Johansson, Simon Liljestrand, and Mats Granath. Quantum error correction for the toric code using deep reinforcement learning. Quantum, 3: 183, September 2019. ISSN 2521-327X. 10.22331\/q-2019-09-02-183. URL https:\/\/doi.org\/10.22331\/q-2019-09-02-183.","DOI":"10.22331\/q-2019-09-02-183"},{"key":"8","doi-asserted-by":"publisher","unstructured":"Pantita Palittapongarnpim, Peter Wittek, Ehsan Zahedinejad, Shakib Vedaie, and Barry C. Sanders. Learning in quantum control: High-dimensional global optimization for noisy quantum dynamics. Neurocomputing, 268: 116 \u2013 126, 2017. ISSN 0925-2312. https:\/\/doi.org\/10.1016\/j.neucom.2016.12.087. URL http:\/\/www.sciencedirect.com\/science\/article\/pii\/S0925231217307531.","DOI":"10.1016\/j.neucom.2016.12.087"},{"key":"9","doi-asserted-by":"publisher","unstructured":"Zheng An and D. L. Zhou. Deep reinforcement learning for quantum gate control. EPL (Europhysics Letters), 126 (6): 60002, jul 2019. 10.1209\/0295-5075\/126\/60002. URL https:\/\/doi.org\/10.1209\/0295-5075\/126\/60002.","DOI":"10.1209\/0295-5075\/126\/60002"},{"key":"10","doi-asserted-by":"publisher","unstructured":"Marin Bukov, Alexandre G. R. Day, Dries Sels, Phillip Weinberg, Anatoli Polkovnikov, and Pankaj Mehta. Reinforcement learning in different phases of quantum control. Phys. Rev. X, 8: 031086, Sep 2018. 10.1103\/PhysRevX.8.031086. URL https:\/\/doi.org\/10.1103\/PhysRevX.8.031086.","DOI":"10.1103\/PhysRevX.8.031086"},{"key":"11","doi-asserted-by":"publisher","unstructured":"Murphy Yuezhen Niu, Sergio Boixo, Vadim N Smelyanskiy, and Hartmut Neven. Universal quantum control through deep reinforcement learning. npj Quantum Information, 5 (1): 1\u20138, 2019. 10.1038\/s41534-019-0141-3. URL https:\/\/doi.org\/10.1038\/s41534-019-0141-3.","DOI":"10.1038\/s41534-019-0141-3"},{"key":"12","doi-asserted-by":"publisher","unstructured":"Han Xu, Junning Li, Liqiang Liu, Yu Wang, Haidong Yuan, and Xin Wang. Generalizable control for quantum parameter estimation through reinforcement learning. npj Quantum Information, 5 (82): 1\u20138, 2019. 10.1038\/s41534-019-0198-z. URL https:\/\/doi.org\/10.1038\/s41534-019-0198-z.","DOI":"10.1038\/s41534-019-0198-z"},{"key":"13","doi-asserted-by":"publisher","unstructured":"Xiao-Ming Zhang, Zezhu Wei, Raza Asad, Xu-Chen Yang, and Xin Wang. When does reinforcement learning stand out in quantum control? a comparative study on state preparation. npj Quantum Information, 5 (85): 1\u20137, 2019. 10.1038\/s41534-019-0201-8. URL https:\/\/doi.org\/10.1038\/s41534-019-0201-8.","DOI":"10.1038\/s41534-019-0201-8"},{"key":"14","doi-asserted-by":"publisher","unstructured":"Matteo M. Wauters, Emanuele Panizon, Glen B. Mbeng, and Giuseppe E. Santoro. Reinforcement-learning-assisted quantum optimization. Phys. Rev. Research, 2: 033446, Sep 2020. 10.1103\/PhysRevResearch.2.033446. URL https:\/\/doi.org\/10.1103\/PhysRevResearch.2.033446.","DOI":"10.1103\/PhysRevResearch.2.033446"},{"key":"15","doi-asserted-by":"publisher","unstructured":"Christopher J. C. H. Watkins. Learning from delayed rewards. PhD thesis, University of Cambridge, 1989. URL https:\/\/doi.org\/10.1016\/0921-8890(95)00026-C.","DOI":"10.1016\/0921-8890(95)00026-C"},{"key":"16","doi-asserted-by":"publisher","unstructured":"Christopher J. C. H. Watkins and Peter Dayan. Q-learning. Machine Learning, 8 (3-4): 279\u2013292, 1992. 10.1007\/BF00992698. URL https:\/\/doi.org\/10.1007\/BF00992698.","DOI":"10.1007\/BF00992698"},{"key":"17","doi-asserted-by":"publisher","unstructured":"Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A Rusu, Joel Veness, Marc G Bellemare, Alex Graves, Martin Riedmiller, Andreas K Fidjeland, Georg Ostrovski, Stig Petersen, Charles Beattie, Amir Sadik, Ioannis Antonoglou, Helen King, Dharshan Kumaran, Daan Wierstra, Shane Legg, and Demis Hassabis. Human-level control through deep reinforcement learning. Nature(London), 518 (7540): 529\u2013533, 2015. 10.1038\/nature14236. URL https:\/\/doi.org\/10.1038\/nature14236.","DOI":"10.1038\/nature14236"},{"key":"18","doi-asserted-by":"publisher","unstructured":"Timothy P. Lillicrap, Jonathan J. Hunt, Alexander Pritzel, Nicolas Heess, Tom Erez, Yuval Tassa, David Silver, and Daan Wierstra. Continuous control with deep reinforcement learning. 2015. 10.48550\/ARXIV.1509.02971. URL https:\/\/arxiv.org\/abs\/1509.02971.","DOI":"10.48550\/ARXIV.1509.02971"},{"key":"19","doi-asserted-by":"publisher","unstructured":"Michael A. Nielsen and Isaac L. Chuang. Quantum Computation and Quantum Information. Cambridge University Press, USA, 10th edition, 2011. ISBN 1107002176. https:\/\/doi.org\/10.1017\/CBO9780511976667.","DOI":"10.1017\/CBO9780511976667"},{"key":"20","doi-asserted-by":"publisher","unstructured":"Jacob Biamonte, Peter Wittek, Nicola Pancotti, Patrick Rebentrost, Nathan Wiebe, and Seth Lloyd. Quantum machine learning. Nature(London), 549 (7671): 195\u2013202, 2017. 10.1038\/nature23474. URL https:\/\/doi.org\/10.1038\/nature23474.","DOI":"10.1038\/nature23474"},{"key":"21","doi-asserted-by":"publisher","unstructured":"Peter W Shor. Algorithms for quantum computation: discrete logarithms and factoring. In Proceedings 35th Annual Symposium on Foundations of Computer Science, pages 124\u2013134, 1994. 10.1109\/SFCS.1994.365700. URL https:\/\/ieeexplore.ieee.org\/document\/365700.","DOI":"10.1109\/SFCS.1994.365700"},{"key":"22","doi-asserted-by":"publisher","unstructured":"Lov Kumar Grover. Quantum mechanics helps in searching for a needle in a haystack. Phys. Rev. Lett., 79: 325\u2013328, Jul 1997. 10.1103\/PhysRevLett.79.325. URL https:\/\/doi.org\/10.1103\/PhysRevLett.79.325.","DOI":"10.1103\/PhysRevLett.79.325"},{"key":"23","doi-asserted-by":"publisher","unstructured":"Nathan Wiebe, Daniel Braun, and Seth Lloyd. Quantum algorithm for data fitting. Phys. Rev. Lett., 109: 050505, Aug 2012. 10.1103\/PhysRevLett.109.050505. URL https:\/\/doi.org\/10.1103\/PhysRevLett.109.050505.","DOI":"10.1103\/PhysRevLett.109.050505"},{"key":"24","doi-asserted-by":"publisher","unstructured":"Patrick Rebentrost, Masoud Mohseni, and Seth Lloyd. Quantum support vector machine for big data classification. Phys. Rev. Lett., 113: 130503, Sep 2014. 10.1103\/PhysRevLett.113.130503. URL https:\/\/doi.org\/10.1103\/PhysRevLett.113.130503.","DOI":"10.1103\/PhysRevLett.113.130503"},{"key":"25","doi-asserted-by":"publisher","unstructured":"Seth Lloyd and Christian Weedbrook. Quantum generative adversarial learning. Phys. Rev. Lett., 121: 040502, Jul 2018. 10.1103\/PhysRevLett.121.040502. URL https:\/\/doi.org\/10.1103\/PhysRevLett.121.040502.","DOI":"10.1103\/PhysRevLett.121.040502"},{"key":"26","doi-asserted-by":"publisher","unstructured":"Sankar Das Sarma, Dong-Ling Deng, and Lu-Ming Duan. Machine learning meets quantum physics. Physics Today, 72 (3): 48\u201354, Mar 2019. ISSN 1945-0699. 10.1063\/pt.3.4164. URL http:\/\/dx.doi.org\/10.1063\/PT.3.4164.","DOI":"10.1063\/pt.3.4164"},{"key":"27","doi-asserted-by":"publisher","unstructured":"Seth Lloyd, Masoud Mohseni, and Patrick Rebentrost. Quantum principal component analysis. Nature Physics, 10 (9): 631\u2013633, 2014. 10.1038\/nphys3029. URL https:\/\/doi.org\/10.1038\/nphys3029.","DOI":"10.1038\/nphys3029"},{"key":"28","unstructured":"Nico Meyer, Christian Ufrecht, Maniraman Periyasamy, Daniel D. Scherer, Axel Plinge, and Christopher Mutschler. A survey on quantum reinforcement learning, 2024. URL https:\/\/arxiv.org\/abs\/2211.03464."},{"key":"29","doi-asserted-by":"publisher","unstructured":"Daoyi Dong, Chunlin Chen, Hanxiong Li, and Tzyh-Jong Tarn. Quantum reinforcement learning. IEEE Transactions on Systems, Man, and Cybernetics, Part B (Cybernetics), 38 (5): 1207\u20131220, 2008. 10.1109\/TSMCB.2008.925743. URL https:\/\/ieeexplore.ieee.org\/document\/4579244.","DOI":"10.1109\/TSMCB.2008.925743"},{"key":"30","doi-asserted-by":"publisher","unstructured":"Vedran Dunjko, Jacob M. Taylor, and Hans J. Briegel. Quantum-enhanced machine learning. Phys. Rev. Lett., 117: 130501, Sep 2016. 10.1103\/PhysRevLett.117.130501. URL https:\/\/doi.org\/10.1103\/PhysRevLett.117.130501.","DOI":"10.1103\/PhysRevLett.117.130501"},{"key":"31","doi-asserted-by":"publisher","unstructured":"Giuseppe Davide Paparo, Vedran Dunjko, Adi Makmal, Miguel Angel Martin-Delgado, and Hans J. Briegel. Quantum speedup for active learning agents. Phys. Rev. X, 4: 031002, Jul 2014a. 10.1103\/PhysRevX.4.031002. URL https:\/\/doi.org\/10.1103\/PhysRevX.4.031002.","DOI":"10.1103\/PhysRevX.4.031002"},{"key":"32","doi-asserted-by":"publisher","unstructured":"Vedran Dunjko, Jacob M Taylor, and Hans J Briegel. Advances in quantum reinforcement learning. In 2017 IEEE International Conference on Systems, Man, and Cybernetics (SMC), pages 282\u2013287, 2017a. 10.1109\/SMC.2017.8122616. URL https:\/\/ieeexplore.ieee.org\/document\/8122616.","DOI":"10.1109\/SMC.2017.8122616"},{"key":"33","doi-asserted-by":"publisher","unstructured":"Vedran Dunjko and Hans J Briegel. Machine learning & artificial intelligence in the quantum domain: a review of recent progress. Reports on Progress in Physics, 81 (7): 074001, jun 2018. 10.1088\/1361-6633\/aab406. URL https:\/\/doi.org\/10.1088\/1361-6633\/aab406.","DOI":"10.1088\/1361-6633\/aab406"},{"key":"34","doi-asserted-by":"publisher","unstructured":"Sofiene Jerbi, Lea M. Trenkwalder, Hendrik Poulsen Nautrup, Hans J. Briegel, and Vedran Dunjko. Quantum enhancements for deep reinforcement learning in large spaces. PRX Quantum, 2: 010328, Feb 2021a. 10.1103\/PRXQuantum.2.010328. URL https:\/\/doi.org\/10.1103\/PRXQuantum.2.010328.","DOI":"10.1103\/PRXQuantum.2.010328"},{"key":"35","doi-asserted-by":"publisher","unstructured":"Samuel Yen-Chi Chen, Chao-Han Huck Yang, Jun Qi, Pin-Yu Chen, Xiaoli Ma, and Hsi-Sheng Goan. Variational quantum circuits for deep reinforcement learning. IEEE Access, 8: 141007\u2013141024, 2020. 10.1109\/ACCESS.2020.3010470. URL https:\/\/ieeexplore.ieee.org\/abstract\/document\/9144562.","DOI":"10.1109\/ACCESS.2020.3010470"},{"key":"36","unstructured":"Owen Lockwood and Mei Si. Reinforcement learning with quantum variational circuits. In Proceedings of the Sixteenth AAAI Conference on Artificial Intelligence and Interactive Digital Entertainment, AIIDE&apos;20. AAAI Press, 2020. ISBN 978-1-57735-849-7. URL https:\/\/dl.acm.org\/doi\/abs\/10.5555\/3505464.3505499."},{"key":"37","doi-asserted-by":"publisher","unstructured":"Andrea Skolik, Sofiene Jerbi, and Vedran Dunjko. Quantum agents in the Gym: a variational quantum algorithm for deep Q-learning. Quantum, 6: 720, May 2022. ISSN 2521-327X. 10.22331\/q-2022-05-24-720. URL https:\/\/doi.org\/10.22331\/q-2022-05-24-720.","DOI":"10.22331\/q-2022-05-24-720"},{"key":"38","unstructured":"Owen Lockwood and Mei Si. Playing atari with hybrid quantum-classical reinforcement learning. In Luca Bertinetto, Jo\u00e3o F. Henriques, Samuel Albanie, Michela Paganini, and G\u00fcl Varol, editors, NeurIPS 2020 Workshop on Pre-registration in Machine Learning, volume 148 of Proceedings of Machine Learning Research, pages 285\u2013301. PMLR, 11 Dec 2021. URL https:\/\/proceedings.mlr.press\/v148\/lockwood21a.html."},{"key":"39","doi-asserted-by":"publisher","unstructured":"Samuel Yen-Chi Chen. Quantum Deep Q-Learning with Distributed Prioritized Experience Replay . In 2023 IEEE International Conference on Quantum Computing and Engineering (QCE), pages 31\u201335, Los Alamitos, CA, USA, September 2023. IEEE Computer Society. 10.1109\/QCE57702.2023.10180. URL https:\/\/doi.ieeecomputersociety.org\/10.1109\/QCE57702.2023.10180.","DOI":"10.1109\/QCE57702.2023.10180"},{"key":"40","unstructured":"Sofiene Jerbi, Casper Gyurik, Simon Marshall, Hans Briegel, and Vedran Dunjko. Parametrized quantum policies for reinforcement learning. In M. Ranzato, A. Beygelzimer, Y. Dauphin, P.S. Liang, and J. Wortman Vaughan, editors, Advances in Neural Information Processing Systems, volume 34, pages 28362\u201328375. Curran Associates, Inc., 2021b. URL https:\/\/proceedings.neurips.cc\/paper\/2021\/file\/eec96a7f788e88184c0e713456026f3f-Paper.pdf."},{"key":"41","unstructured":"Nico Meyer, Daniel Scherer, Axel Plinge, Christopher Mutschler, and Michael Hartmann. Quantum policy gradient algorithm with optimized action decoding. In Andreas Krause, Emma Brunskill, Kyunghyun Cho, Barbara Engelhardt, Sivan Sabato, and Jonathan Scarlett, editors, Proceedings of the 40th International Conference on Machine Learning, volume 202 of Proceedings of Machine Learning Research, pages 24592\u201324613. PMLR, 23\u201329 Jul 2023a. URL https:\/\/proceedings.mlr.press\/v202\/meyer23a.html."},{"key":"42","doi-asserted-by":"publisher","unstructured":"Nico Meyer, Daniel D. Scherer, Axel Plinge, Christopher Mutschler, and Michael J. Hartmann. Quantum Natural Policy Gradients: Towards Sample-Efficient Reinforcement Learning . In 2023 IEEE International Conference on Quantum Computing and Engineering (QCE), pages 36\u201341, Los Alamitos, CA, USA, September 2023b. IEEE Computer Society. 10.1109\/QCE57702.2023.10181. URL https:\/\/doi.ieeecomputersociety.org\/10.1109\/QCE57702.2023.10181.","DOI":"10.1109\/QCE57702.2023.10181"},{"key":"43","doi-asserted-by":"publisher","unstructured":"Andr\u00e9 Sequeira, Luis Paulo Santos, and Luis Soares Barbosa. Policy gradients using variational quantum circuits. Quantum Machine Intelligence, 5 (1): 18, 2023. ISSN 2524-4914. 10.1007\/s42484-023-00101-8. URL https:\/\/doi.org\/10.1007\/s42484-023-00101-8.","DOI":"10.1007\/s42484-023-00101-8"},{"key":"44","doi-asserted-by":"publisher","unstructured":"Valeria Saggio, Beate E Asenbeck, Arne Hamann, Teodor Str\u00f6mberg, Peter Schiansky, Vedran Dunjko, Nicolai Friis, Nicholas C Harris, Michael Hochberg, Dirk Englund, et al. Experimental quantum speed-up in reinforcement learning agents. Nature, 591 (7849): 229\u2013233, 2021. 10.1038\/s41586-021-03242-7. URL https:\/\/doi.org\/10.1038\/s41586-021-03242-7.","DOI":"10.1038\/s41586-021-03242-7"},{"key":"45","doi-asserted-by":"publisher","unstructured":"Vedran Dunjko, Jacob M. Taylor, and Hans J. Briegel. Advances in quantum reinforcement learning. In 2017 IEEE International Conference on Systems, Man, and Cybernetics (SMC), page 282\u2013287. IEEE Press, 2017b. 10.1109\/SMC.2017.8122616. URL https:\/\/doi.org\/10.1109\/SMC.2017.8122616.","DOI":"10.1109\/SMC.2017.8122616"},{"key":"46","doi-asserted-by":"publisher","unstructured":"El Amine Cherrat, Iordanis Kerenidis, and Anupam Prakash. Quantum reinforcement learning via policy iteration. Quantum Machine Intelligence, 5 (2): 30, 2023. ISSN 2524-4914. 10.1007\/s42484-023-00116-1. URL https:\/\/doi.org\/10.1007\/s42484-023-00116-1.","DOI":"10.1007\/s42484-023-00116-1"},{"key":"47","unstructured":"Daochen Wang, Aarthi Sundaram, Robin Kothari, Ashish Kapoor, and Martin Roetteler. Quantum algorithms for reinforcement learning with a generative model. In Marina Meila and Tong Zhang, editors, Proceedings of the 38th International Conference on Machine Learning, volume 139 of Proceedings of Machine Learning Research, pages 10916\u201310926. PMLR, 18\u201324 Jul 2021. URL https:\/\/proceedings.mlr.press\/v139\/wang21w.html."},{"key":"48","doi-asserted-by":"publisher","unstructured":"Hans J. Briegel and Gemma De las Cuevas. Projective simulation for artificial intelligence. Scientific Reports, 2 (1): 400, 2012. ISSN 2045-2322. 10.1038\/srep00400. URL https:\/\/doi.org\/10.1038\/srep00400.","DOI":"10.1038\/srep00400"},{"key":"49","doi-asserted-by":"publisher","unstructured":"Alexey A. Melnikov, Adi Makmal, Vedran Dunjko, and Hans J. Briegel. Projective simulation with generalization. Scientific Reports, 7 (1): 14430, 2017. ISSN 2045-2322. 10.1038\/s41598-017-14740-y. URL https:\/\/doi.org\/10.1038\/s41598-017-14740-y.","DOI":"10.1038\/s41598-017-14740-y"},{"key":"50","doi-asserted-by":"publisher","unstructured":"Giuseppe Davide Paparo, Vedran Dunjko, Adi Makmal, Miguel Angel Martin-Delgado, and Hans J. Briegel. Quantum speedup for active learning agents. Phys. Rev. X, 4: 031002, Jul 2014b. 10.1103\/PhysRevX.4.031002. URL https:\/\/doi.org\/10.1103\/PhysRevX.4.031002.","DOI":"10.1103\/PhysRevX.4.031002"},{"key":"51","doi-asserted-by":"publisher","unstructured":"V Dunjko, N Friis, and H J Briegel. Quantum-enhanced deliberation of learning agents using trapped ions. New Journal of Physics, 17 (2): 023006, jan 2015. 10.1088\/1367-2630\/17\/2\/023006. URL https:\/\/dx.doi.org\/10.1088\/1367-2630\/17\/2\/023006.","DOI":"10.1088\/1367-2630\/17\/2\/023006"},{"key":"52","doi-asserted-by":"publisher","unstructured":"Th Sriarunothai, S W\u00f6lk, G S Giri, N Friis, V Dunjko, H J Briegel, and Ch Wunderlich. Speeding-up the decision making of a learning agent using an ion trap quantum processor. Quantum Science and Technology, 4 (1): 015014, dec 2018. 10.1088\/2058-9565\/aaef5e. URL https:\/\/dx.doi.org\/10.1088\/2058-9565\/aaef5e.","DOI":"10.1088\/2058-9565\/aaef5e"},{"key":"53","doi-asserted-by":"publisher","unstructured":"Martijn Van Otterlo and Marco Wiering. Reinforcement learning and markov decision processes. In Reinforcement Learning, pages 3\u201342. Springer Berlin Heidelberg, 2012. 10.1007\/978-3-642-27645-3_1. URL https:\/\/doi.org\/10.1007\/978-3-642-27645-3_1.","DOI":"10.1007\/978-3-642-27645-3_1"},{"key":"54","unstructured":"Gavin A Rummery and Mahesan Niranjan. On-line q-learning using connectionist systems. Technical report, 1994. URL http:\/\/mi.eng.cam.ac.uk\/reports\/svr-ftp\/auto-pdf\/rummery_tr166.pdf."},{"key":"55","unstructured":"Sham M Kakade. A natural policy gradient. In Advances in Neural Information Processing Systems, volume 14, pages 1531\u20131538. MIT Press, 2002. URL https:\/\/proceedings.neurips.cc\/paper\/2001\/file\/4b86abe48d358ecf194c56c69108433e-Paper.pdf."},{"key":"56","doi-asserted-by":"publisher","unstructured":"Alberto Peruzzo, Jarrod McClean, Peter Shadbolt, Man-Hong Yung, Xiao-Qi Zhou, Peter J. Love, Al\u00e1n Aspuru-Guzik, and Jeremy L. O&apos;Brien. A variational eigenvalue solver on a photonic quantum processor. Nature communications, 5: 4213, 2014. 10.1038\/ncomms5213. URL https:\/\/doi.org\/10.1038\/ncomms5213.","DOI":"10.1038\/ncomms5213"},{"key":"57","doi-asserted-by":"publisher","unstructured":"Edward Farhi, Jeffrey Goldstone, and Sam Gutmann. A quantum approximate optimization algorithm. 10.48550\/ARXIV.1411.4028. URL https:\/\/arxiv.org\/abs\/1411.4028.","DOI":"10.48550\/ARXIV.1411.4028"},{"key":"58","doi-asserted-by":"publisher","unstructured":"Marcello Benedetti, Erika Lloyd, Stefan Sack, and Mattia Fiorentini. Parameterized quantum circuits as machine learning models. Quantum Science and Technology, 4 (4): 043001, nov 2019. 10.1088\/2058-9565\/ab4eb5. URL https:\/\/doi.org\/10.1088\/2058-9565\/ab4eb5.","DOI":"10.1088\/2058-9565\/ab4eb5"},{"key":"59","unstructured":"Long-Ji Lin. Reinforcement Learning for Robots Using Neural Networks. PhD thesis, USA, 1992. URL https:\/\/dl.acm.org\/doi\/10.5555\/168871."},{"key":"60","unstructured":"Diederik P Kingma and Jimmy Ba. Adam: A method for stochastic optimization. In Proceedings of International Conference on Learning Representations, 2015. URL http:\/\/arxiv.org\/abs\/1412.6980."},{"key":"61","unstructured":"Xiaokai Hou, Guanyu Zhou, Qingyu Li, Shan Jin, and Xiaoting Wang. A universal duplication-free quantum neural network. URL https:\/\/arxiv.org\/abs\/2106.13211."},{"key":"62","doi-asserted-by":"publisher","unstructured":"Dorit Aharonov and Tomer Naveh. Quantum np-a survey. arXiv preprint quant-ph\/0210077, 2002. 10.48550\/ARXIV.QUANT-PH\/0210077. URL https:\/\/arxiv.org\/abs\/quant-ph\/0210077.","DOI":"10.48550\/ARXIV.QUANT-PH\/0210077"},{"key":"63","doi-asserted-by":"publisher","unstructured":"John Watrous. Quantum computational complexity. 10.48550\/ARXIV.0804.3401. URL https:\/\/arxiv.org\/abs\/0804.3401.","DOI":"10.48550\/ARXIV.0804.3401"},{"key":"64","doi-asserted-by":"publisher","unstructured":"Sevag Gharibian, Yichen Huang, Zeph Landau, and Seung Woo Shin. Quantum hamiltonian complexity. pages 7174\u20137201, 2009. 10.1007\/978-0-387-30440-3_428. URL https:\/\/doi.org\/10.1007\/978-0-387-30440-3_428.","DOI":"10.1007\/978-0-387-30440-3_428"},{"key":"65","doi-asserted-by":"publisher","unstructured":"Julia Kempe, Alexei Kitaev, and Oded Regev. The complexity of the local hamiltonian problem. In FSTTCS 2004: Foundations of Software Technology and Theoretical Computer Science, volume 35, pages 372\u2013383, 2006. URL https:\/\/doi.org\/10.1007\/978-3-540-30538-5_31.","DOI":"10.1007\/978-3-540-30538-5_31"},{"key":"66","doi-asserted-by":"publisher","unstructured":"Daniel S. Abrams and Seth Lloyd. Quantum algorithm providing exponential speed increase for finding eigenvalues and eigenvectors. Phys. Rev. Lett., 83: 5162\u20135165, Dec 1999. 10.1103\/PhysRevLett.83.5162. URL https:\/\/doi.org\/10.1103\/PhysRevLett.83.5162.","DOI":"10.1103\/PhysRevLett.83.5162"},{"key":"67","doi-asserted-by":"publisher","unstructured":"Navin Khaneja, Timo Reiss, Cindie Kehlet, Thomas Schulte-Herbr\u00fcggen, and Steffen J. Glaser. Optimal control of coupled spin dynamics: design of nmr pulse sequences by gradient ascent algorithms. Journal of Magnetic Resonance, 172 (2): 296\u2013305, 2005. ISSN 1090-7807. https:\/\/doi.org\/10.1016\/j.jmr.2004.11.004. URL https:\/\/www.sciencedirect.com\/science\/article\/pii\/S1090780704003696.","DOI":"10.1016\/j.jmr.2004.11.004"},{"key":"68","doi-asserted-by":"crossref","unstructured":"Warwick Masson, Pravesh Ranchod, and George Konidaris. Reinforcement learning with parameterized actions. In Proceedings of the Thirtieth AAAI Conference on Artificial Intelligence, AAAI&apos;16, page 1934\u20131940. AAAI Press, 2016. URL https:\/\/ojs.aaai.org\/index.php\/AAAI\/article\/view\/10226.","DOI":"10.1609\/aaai.v30i1.10226"},{"key":"69","doi-asserted-by":"publisher","unstructured":"Kenji Doya. Reinforcement learning in continuous time and space. Neural Computation, 12 (1): 219\u2013245, 2000. 10.1162\/089976600300015961. URL https:\/\/ieeexplore.ieee.org\/document\/6789455.","DOI":"10.1162\/089976600300015961"},{"key":"70","doi-asserted-by":"publisher","unstructured":"Greg Brockman, Vicki Cheung, Ludwig Pettersson, Jonas Schneider, John Schulman, Jie Tang, and Wojciech Zaremba. Openai gym. 10.48550\/ARXIV.1606.01540. URL https:\/\/arxiv.org\/abs\/1606.01540.","DOI":"10.48550\/ARXIV.1606.01540"}],"container-title":["Quantum"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/quantum-journal.org\/papers\/q-2025-03-12-1660\/pdf\/","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"}],"deposited":{"date-parts":[[2025,3,12]],"date-time":"2025-03-12T08:43:49Z","timestamp":1741769029000},"score":1,"resource":{"primary":{"URL":"https:\/\/quantum-journal.org\/papers\/q-2025-03-12-1660\/"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,3,12]]},"references-count":71,"URL":"https:\/\/doi.org\/10.22331\/q-2025-03-12-1660","archive":["CLOCKSS"],"relation":{},"ISSN":["2521-327X"],"issn-type":[{"value":"2521-327X","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,3,12]]},"article-number":"1660"}}