{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,11]],"date-time":"2026-06-11T14:03:11Z","timestamp":1781186591954,"version":"3.54.1"},"reference-count":45,"publisher":"Springer Science and Business Media LLC","issue":"2","license":[{"start":{"date-parts":[[2026,6,11]],"date-time":"2026-06-11T00:00:00Z","timestamp":1781136000000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2026,6,11]],"date-time":"2026-06-11T00:00:00Z","timestamp":1781136000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"name":"INL - International Iberian Nanotechnology Laboratory"}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Quantum Mach. Intell."],"published-print":{"date-parts":[[2026,12]]},"abstract":"<jats:title>Abstract<\/jats:title>\n                  <jats:p>Reinforcement learning (RL) provides a principled framework for decision-making in partially observable environments, which can be modeled as Markov decision processes and compactly represented through dynamic decision Bayesian networks. Recent advances demonstrate that inference on sparse Bayesian networks can be accelerated using quantum rejection sampling combined with amplitude amplification, leading to a computational speedup in estimating acceptance probabilities. Building on this result, we introduce Quantum Bayesian Reinforcement Learning (QBRL), a hybrid quantum-classical look-ahead algorithm for model-based RL in partially observable environments. We present a rigorous, oracle-free time complexity analysis under fault-tolerant assumptions for the quantum device. Unlike standard treatments that assume a black-box oracle, we explicitly specify the inference process, allowing our bounds to more accurately reflect the true computational cost. We show that, for environments whose dynamics form a sparse Bayesian network, horizon-based near-optimal planning can be achieved sub-quadratically faster through quantum-enhanced belief updates. On the other hand, we show that there is no quantum speed-up for environments that are either fully observable, or characterized by Bayesian networks whose maximum in-degree is not small. Furthermore, we present numerical experiments benchmarking QBRL against its classical counterpart on simple yet illustrative decision-making tasks. Our results offer a detailed analysis of how the quantum computational advantage translates into decision-making performance, highlighting that the magnitude of the advantage can vary significantly across different deployment settings.<\/jats:p>","DOI":"10.1007\/s42484-026-00401-9","type":"journal-article","created":{"date-parts":[[2026,6,11]],"date-time":"2026-06-11T13:35:34Z","timestamp":1781184934000},"update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":0,"title":["Quantum Bayesian networks can speed up reinforcement learning in partially observable environments"],"prefix":"10.1007","volume":"8","author":[{"given":"Gilberto","family":"Cunha","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Alexandra","family":"Ram\u00f4a","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Andr\u00e9","family":"Sequeira","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Michael","family":"de Oliveira","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Lu\u00eds","family":"Barbosa","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"297","published-online":{"date-parts":[[2026,6,11]]},"reference":[{"key":"401_CR1","doi-asserted-by":"publisher","unstructured":"Bai Y, Gao Y, Wan R, Zhang S, Song R (2024) A review of reinforcement learning in financial applications. Annual Review of Statistics and Its Application. https:\/\/doi.org\/10.1146\/annurev-statistics-112723-034423","DOI":"10.1146\/annurev-statistics-112723-034423"},{"key":"401_CR2","doi-asserted-by":"publisher","unstructured":"Bergholm V, Vartiainen JJ, M\u00f6tt\u00f6nen M, Salomaa MM (2005) Quantum circuits with uniformly controlled one-qubit gates. Phys Rev A 71. https:\/\/doi.org\/10.1103\/physreva.71.052330","DOI":"10.1103\/physreva.71.052330"},{"key":"401_CR3","unstructured":"Borujeni SE, Nannapaneni S (2021) Modeling time-dependent systems using dynamic quantum Bayesian networks. arXiv preprint arXiv:2107.00713"},{"key":"401_CR4","doi-asserted-by":"publisher","DOI":"10.1016\/j.eswa.2021.114768","volume":"176","author":"SE Borujeni","year":"2021","unstructured":"Borujeni SE, Nannapaneni S, Nguyen NH, Behrman EC, Steck JE (2021) Quantum circuit representation of Bayesian networks. Expert Syst Appl 176:114768","journal-title":"Expert Syst Appl"},{"key":"401_CR5","doi-asserted-by":"publisher","first-page":"53","DOI":"10.1090\/conm\/305\/05215","volume":"305","author":"G Brassard","year":"2002","unstructured":"Brassard G, Hoyer P, Mosca M, Tapp A (2002) Quantum amplitude amplification and estimation. Contemp Math 305:53\u201374","journal-title":"Contemp Math"},{"key":"401_CR6","doi-asserted-by":"publisher","unstructured":"Bukov M, Day AG, Sels D, Weinberg P, Polkovnikov A, Mehta P (2018) Reinforcement learning in different phases of quantum control. Phys Rev X 8. https:\/\/doi.org\/10.1103\/physrevx.8.031086","DOI":"10.1103\/physrevx.8.031086"},{"key":"401_CR7","doi-asserted-by":"publisher","unstructured":"Chen C-L, Dong D-Y (2008) Superposition-inspired reinforcement learning and quantum reinforcement learning. I-Tech Education and Publishing, In Reinforcement Learning. https:\/\/doi.org\/10.5772\/5275","DOI":"10.5772\/5275"},{"key":"401_CR8","doi-asserted-by":"publisher","unstructured":"Cherrat EA, Kerenidis I, Prakash A (2023a) Quantum reinforcement learning via policy iteration. Quant Mach Intell 5:30. https:\/\/doi.org\/10.1007\/s42484-023-00116-1","DOI":"10.1007\/s42484-023-00116-1"},{"key":"401_CR9","doi-asserted-by":"publisher","unstructured":"Cherrat EA, Raj S, Kerenidis I, Shekhar A, Wood B, Dee J, Chakrabarti S, Chen R, Herman D, Hu S, Minssen P, Shaydulin R, Sun Y, Yalovetzky R, Pistoia M (2023b) Quantum Deep Hedging. https:\/\/doi.org\/10.48550\/arXiv.2303.16585","DOI":"10.48550\/arXiv.2303.16585"},{"key":"401_CR10","doi-asserted-by":"crossref","unstructured":"de Oliveira M, Barbosa LS (2021) Quantum Bayesian decision-making. Foundations of Science, (pp 1\u201321)","DOI":"10.1007\/s10699-021-09781-6"},{"key":"401_CR11","doi-asserted-by":"publisher","unstructured":"Dunjko V Liu Y-K, Wu X, Taylor JM (2018) Exponential improvements for quantum-accessible reinforcement learning. https:\/\/doi.org\/10.48550\/arXiv.1710.11160","DOI":"10.48550\/arXiv.1710.11160"},{"key":"401_CR12","unstructured":"Ernst JO, Chatterjee A, Franzmeyer T, Kuhn A (2025) Reinforcement learning for quantum control under physical constraints. arxiv.org\/abs\/2501.14372"},{"key":"401_CR13","doi-asserted-by":"publisher","DOI":"10.1103\/PRXQuantum.2.020303","volume":"2","author":"LJ Fiderer","year":"2021","unstructured":"Fiderer LJ, Schuff J, Braun D (2021) Neural-network heuristics for adaptive bayesian quantum estimation. PRX Quantum 2:020303. https:\/\/doi.org\/10.1103\/PRXQuantum.2.020303","journal-title":"PRX Quantum"},{"key":"401_CR14","doi-asserted-by":"publisher","unstructured":"Ganguly B, Xu Y, Aggarwal V (2024) Quantum Speedups in Regret Analysis of Infinite Horizon Average-Reward Markov Decision Processes. https:\/\/doi.org\/10.48550\/arXiv.2310.11684","DOI":"10.48550\/arXiv.2310.11684"},{"key":"401_CR15","volume":"12","author":"X Gao","year":"2022","unstructured":"Gao X, Anschuetz ER, Wang S-T, Cirac JI, Lukin MD (2022) Enhancing generative models via quantum correlations. Phys Rev X 12:021037","journal-title":"Phys Rev X"},{"key":"401_CR16","doi-asserted-by":"publisher","first-page":"359","DOI":"10.1561\/2200000049","volume":"8","author":"M Ghavamzadeh","year":"2015","unstructured":"Ghavamzadeh M, Mannor S, Pineau J, Tamar A (2015) Bayesian reinforcement learning: A survey. Foundations and Trends\u00ae in Machine Learning 8:359\u2013483. https:\/\/doi.org\/10.1561\/2200000049","journal-title":"Foundations and Trends\u00ae in Machine Learning"},{"key":"401_CR17","doi-asserted-by":"publisher","unstructured":"Grover LK (1996) A fast quantum mechanical algorithm for database search. In Proceedings of the 28tj Annual ACM Symposium on Theory of Computing Stoc \u201996 (pp 212\u2013219). addressNew York, NY, USA: Association for Computing Machinery. https:\/\/doi.org\/10.1145\/237814.237866","DOI":"10.1145\/237814.237866"},{"key":"401_CR18","unstructured":"Herman E, Strang G (2016) Calculus Volume 2. OpenStax"},{"key":"401_CR19","doi-asserted-by":"publisher","unstructured":"Jerbi S, Cornelissen A, Ozols M, Dunjko V (2022) Quantum policy gradient algorithms. https:\/\/doi.org\/10.4230\/LIPIcs.TQC.2023.13, arXiv:2212.09328","DOI":"10.4230\/LIPIcs.TQC.2023.13"},{"key":"401_CR20","unstructured":"Jerbi,S, Gyurik C, Marshall S, Briegel H, Dunjko V (2021) Parametrized Quantum Policies for Reinforcement Learning. In: Advances in neural information processing systems (pp 28362\u201328375). Curran Associates, Inc. volume\u00a034"},{"key":"401_CR21","doi-asserted-by":"publisher","first-page":"99","DOI":"10.1016\/S0004-3702(98)00023-X","volume":"101","author":"LP Kaelbling","year":"1998","unstructured":"Kaelbling LP, Littman ML, Cassandra AR (1998) Planning and acting in partially observable stochastic domains. Artif Intell 101:99\u2013134","journal-title":"Artif Intell"},{"key":"401_CR22","unstructured":"Kaldari J, Tariq S, Al-Kuwari S, Chen SY-C, Chatzinotas S, Shin H (2025) Quantum reinforcement learning: Recent advances and future directions. arXiv:2510.14595"},{"key":"401_CR23","doi-asserted-by":"publisher","unstructured":"Liu X-H, Xue Z, Pang J-C, Jiang S, Xu F, Yu Y (2021) Regret Minimization Experience Replay in Off-Policy Reinforcement Learning. Adv Neural Inf Process Syst. https:\/\/doi.org\/10.48550\/ARXIV.2105.07253","DOI":"10.48550\/ARXIV.2105.07253"},{"key":"401_CR24","doi-asserted-by":"crossref","unstructured":"Low GH, Yoder TJ, Chuang IL (2014a) Quantum inference on Bayesian networks. Phys Rev A 89:062315","DOI":"10.1103\/PhysRevA.89.062315"},{"key":"401_CR25","doi-asserted-by":"publisher","unstructured":"Low GH, Yoder TJ, Chuang IL (2014b) Quantum Inference on Bayesian Networks. Phys Rev A 89:062315. https:\/\/doi.org\/10.1103\/PhysRevA.89.062315","DOI":"10.1103\/PhysRevA.89.062315"},{"key":"401_CR26","unstructured":"Meyer N, Ufrecht C, Periyasamy M, Scherer DD, Plinge A, Mutschler C (2024) A survey on quantum reinforcement learning. arXiv:2211.03464"},{"key":"401_CR27","unstructured":"Moss RJ, Corso A, Caers J, Kochenderfer MJ (2023) Betazero: Belief-state planning for long-horizon pomdps using learned approximations. arXiv preprint arXiv:2306.00249"},{"key":"401_CR28","doi-asserted-by":"publisher","unstructured":"Mu\u0161kardin E, Tappler M, Aichernig BK, Pill I (2024) Reinforcement Learning Under Partial Observability Guided by Learned Environment Models. In: Herber P, Wijs A (eds), Integrated Formal Methods (pp 257\u2013276). addressCham: Springer Nature Switzerland. https:\/\/doi.org\/10.1007\/978-3-031-47705-8_14","DOI":"10.1007\/978-3-031-47705-8_14"},{"key":"401_CR29","doi-asserted-by":"publisher","DOI":"10.1016\/j.compchemeng.2020.106886","volume":"139","author":"R Nian","year":"2020","unstructured":"Nian R, Liu J, Huang B (2020) A review On reinforcement learning: Introduction and applications in industrial process control. Comput Chem Eng 139:106886. https:\/\/doi.org\/10.1016\/j.compchemeng.2020.106886","journal-title":"Comput Chem Eng"},{"key":"401_CR30","doi-asserted-by":"publisher","unstructured":"Niu MY, Boixo S, Smelyanskiy VN, Neven H (2019) Universal quantum control through deep reinforcement learning. In: AIAA Scitech 2019 Forum. American Institute of Aeronautics and Astronautics. https:\/\/doi.org\/10.2514\/6.2019-0954","DOI":"10.2514\/6.2019-0954"},{"key":"401_CR31","doi-asserted-by":"crossref","unstructured":"Papadimitriou CH, Tsitsiklis JN (1987) The Complexity of Markov Decision Processes. Mathematics of Operations Research 12:441\u2013450. arXiv:3689975","DOI":"10.1287\/moor.12.3.441"},{"key":"401_CR32","unstructured":"Pineau J, Gordon G, Thrun S (2003) Point-based value iteration: An anytime algorithm for pomdps. In: IJCAI"},{"key":"401_CR33","unstructured":"Ram\u00f4a A (2024) Quantum bayesian reinforcement learning. https:\/\/github.com\/alexandra-frca\/QBRL"},{"key":"401_CR34","doi-asserted-by":"publisher","unstructured":"Sannia A, Giordano A, Gullo NL, Mastroianni C, Plastina F (2023) A hybrid classical-quantum approach to speed-up Q-learning. Sci Rep 13. https:\/\/doi.org\/10.1038\/s41598-023-30990-5","DOI":"10.1038\/s41598-023-30990-5"},{"key":"401_CR35","doi-asserted-by":"publisher","first-page":"24","DOI":"10.1016\/0315-0860(92)90053-E","volume":"19","author":"E Seneta","year":"1992","unstructured":"Seneta E (1992) On the history of the Strong Law of Large Numbers and Boole\u2019s inequality. Hist Math 19:24\u201339. https:\/\/doi.org\/10.1016\/0315-0860(92)90053-E","journal-title":"Hist Math"},{"key":"401_CR36","unstructured":"Sidford A, Wang M, Wu X, Yang LF, Ye Y (2018) Near-optimal time and sample complexities for solving discounted Markov decision process with a generative model. arXiv preprint arXiv:1806.01492"},{"key":"401_CR37","unstructured":"Silver D, Veness J (2010) Partially observable monte-carlo planning (pomcp). In: UAI"},{"key":"401_CR38","doi-asserted-by":"publisher","first-page":"484","DOI":"10.1038\/nature16961","volume":"529","author":"D Silver","year":"2016","unstructured":"Silver D, Huang A, Maddison CJ, Guez A, Sifre L, Van Den Driessche G, Schrittwieser J, Antonoglou I, Panneershelvam V, Lanctot M et al (2016) Mastering the game of Go with deep neural networks and tree search. Nature 529:484\u2013489","journal-title":"Nature"},{"key":"401_CR39","doi-asserted-by":"publisher","unstructured":"Skolik A, Jerbi S, Dunjko V (2022) Quantum agents in the Gym: A variational quantum algorithm for deep Q-learning. Quantum 6:720. https:\/\/doi.org\/10.22331\/q-2022-05-24-720, arXiv:2103.15084","DOI":"10.22331\/q-2022-05-24-720"},{"key":"401_CR40","unstructured":"Smith T, Simmons R (2012) Heuristic search value iteration for pomdps (hsvi). In: UAI"},{"key":"401_CR41","doi-asserted-by":"publisher","first-page":"195","DOI":"10.1613\/jair.1659","volume":"24","author":"MT Spaan","year":"2005","unstructured":"Spaan MT, Vlassis N (2005) Perseus: Randomized point-based value iteration for pomdps. J Artif Intell Res 24:195\u2013220","journal-title":"J Artif Intell Res"},{"key":"401_CR42","unstructured":"Wang D, Sundaram A, Kothari R, Kapoor A, Roetteler M (2021) Quantum algorithms for reinforcement learning with a generative model. In: International Conference on Machine Learning (pp 10916\u201310926). PMLR"},{"key":"401_CR43","doi-asserted-by":"publisher","unstructured":"Wang S, Zhang S, Zhang J, Hu R, Li X, Zhang T, Li J, Wu F, Wang G, Hovy E (2025) Reinforcement Learning Enhanced LLMs: A Survey. https:\/\/doi.org\/10.48550\/arXiv.2412.10400, arXiv:2412.10400","DOI":"10.48550\/arXiv.2412.10400"},{"key":"401_CR44","doi-asserted-by":"publisher","unstructured":"Xu Y, Aggarwal V (2025) Accelerating Quantum Reinforcement Learning with a Quantum Natural Policy Gradient Based Approach. https:\/\/doi.org\/10.48550\/arXiv.2501.16243. arXiv:2501.16243","DOI":"10.48550\/arXiv.2501.16243"},{"key":"401_CR45","doi-asserted-by":"publisher","first-page":"61","DOI":"10.1109\/TCST.2024.3437142","volume":"33","author":"H Yu","year":"2025","unstructured":"Yu H, Zhao X, Chen C (2025) Quantum-inspired reinforcement learning for quantum control. IEEE Trans Control Syst Technol 33:61\u201376. https:\/\/doi.org\/10.1109\/TCST.2024.3437142","journal-title":"IEEE Trans Control Syst Technol"}],"container-title":["Quantum Machine Intelligence"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s42484-026-00401-9.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s42484-026-00401-9","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s42484-026-00401-9.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,6,11]],"date-time":"2026-06-11T13:35:44Z","timestamp":1781184944000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s42484-026-00401-9"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,6,11]]},"references-count":45,"journal-issue":{"issue":"2","published-print":{"date-parts":[[2026,12]]}},"alternative-id":["401"],"URL":"https:\/\/doi.org\/10.1007\/s42484-026-00401-9","relation":{},"ISSN":["2524-4906","2524-4914"],"issn-type":[{"value":"2524-4906","type":"print"},{"value":"2524-4914","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,6,11]]},"assertion":[{"value":"29 July 2025","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"18 May 2026","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"11 June 2026","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"The authors declare no competing interests.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Competing interests"}}],"article-number":"65"}}