{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,8]],"date-time":"2026-05-08T07:12:45Z","timestamp":1778224365926,"version":"3.51.4"},"reference-count":57,"publisher":"Maximum Academic Press","issue":"1","content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["KER"],"published-print":{"date-parts":[[2026]]},"DOI":"10.48130\/ker-0026-0001","type":"journal-article","created":{"date-parts":[[2026,5,8]],"date-time":"2026-05-08T06:48:59Z","timestamp":1778222939000},"page":"0-0","source":"Crossref","is-referenced-by-count":0,"title":["A comparison of state-of-the-art reinforcement learning algorithms applied to the traveling salesman problem"],"prefix":"10.48130","volume":"41","author":[{"given":"Kenneth","family":"Schr\u00f6der","sequence":"first","affiliation":[{}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Alexander","family":"Kastius","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Rainer","family":"Schlosser","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"27968","reference":[{"key":"key-10.48130\/ker-0026-0001-1","unstructured":"<p>Applegate D, Bixby RE, Chv\u00e1tal V, Cook WJ. 2019. <i>Concorde TSP Solver<\/i>. <ext-link ext-link-type=\"uri\" xlink:href=\"https:\/\/www.math.uwaterloo.ca\/tsp\/concorde.html\">www.math.uwaterloo.ca\/tsp\/concorde.html<\/ext-link> (Accessed 16 May 2024)<\/p>"},{"key":"key-10.48130\/ker-0026-0001-2","unstructured":"<p>Gurobi Optimization, LLC. 2022. <i>Gurobi Optimizer Reference Manual<\/i>. <ext-link ext-link-type=\"uri\" xlink:href=\"https:\/\/www.gurobi.com\">www.gurobi.com<\/ext-link><\/p>"},{"key":"key-10.48130\/ker-0026-0001-3","doi-asserted-by":"publisher","unstructured":"<p>Flood MM. 1956. The traveling-salesman problem. <i><ext-link ext-link-type=\"uri\" xmlns:xlink=\"http:\/\/www.w3.org\/1999\/xlink\" xlink:href=\"https:\/\/doi.org\/10.1287\/opre.4.1.61\">Operations Research<\/ext-link><\/i> 4(1):61\u221275<\/p>","DOI":"10.1287\/opre.4.1.61"},{"key":"key-10.48130\/ker-0026-0001-4","doi-asserted-by":"publisher","unstructured":"<p>Rosenkrantz DJ, Stearns RE, Lewis PM. 1977. An analysis of several heuristics for the traveling salesman problem. <i><ext-link ext-link-type=\"uri\" xmlns:xlink=\"http:\/\/www.w3.org\/1999\/xlink\" xlink:href=\"https:\/\/doi.org\/10.1137\/0206041\">SIAM Journal on Computing<\/ext-link><\/i> 6(3):563\u2212581<\/p>","DOI":"10.1137\/0206041"},{"key":"key-10.48130\/ker-0026-0001-5","doi-asserted-by":"publisher","unstructured":"<p>Mazyavkina N, Sviridov S, Ivanov S, Burnaev E. 2021. Reinforcement learning for combinatorial optimization: a survey. <i><ext-link ext-link-type=\"uri\" xmlns:xlink=\"http:\/\/www.w3.org\/1999\/xlink\" xlink:href=\"https:\/\/doi.org\/10.1016\/j.cor.2021.105400\">Computers &amp; Operations Research<\/ext-link><\/i> 134:105400<\/p>","DOI":"10.1016\/j.cor.2021.105400"},{"key":"key-10.48130\/ker-0026-0001-6","doi-asserted-by":"publisher","unstructured":"<p>Chen D, Imdahl C, Lai D, Van Woensel T. 2025. The Dynamic Traveling Salesman Problem with Time-Dependent and Stochastic travel times: a deep reinforcement learning approach. <i><ext-link ext-link-type=\"uri\" xmlns:xlink=\"http:\/\/www.w3.org\/1999\/xlink\" xlink:href=\"https:\/\/doi.org\/10.1016\/j.trc.2025.105022\">Transportation Research Part C: Emerging Technologies<\/ext-link><\/i> 172:105022<\/p>","DOI":"10.1016\/j.trc.2025.105022"},{"key":"key-10.48130\/ker-0026-0001-7","doi-asserted-by":"publisher","unstructured":"<p>L\u00e4hdeaho O, Hilmola OP. 2024. An exploration of quantitative models and algorithms for vehicle routing optimization and traveling salesman problems. <i><ext-link ext-link-type=\"uri\" xmlns:xlink=\"http:\/\/www.w3.org\/1999\/xlink\" xlink:href=\"https:\/\/doi.org\/10.1016\/j.sca.2023.100056\">Supply Chain Analytics<\/ext-link><\/i> 5:100056<\/p>","DOI":"10.1016\/j.sca.2023.100056"},{"key":"key-10.48130\/ker-0026-0001-8","doi-asserted-by":"publisher","unstructured":"<p>Li J, Ma Y, Gao R, Cao Z, Lim A, et al. 2022. Deep reinforcement learning for solving the heterogeneous capacitated vehicle routing problem. <i><ext-link ext-link-type=\"uri\" xmlns:xlink=\"http:\/\/www.w3.org\/1999\/xlink\" xlink:href=\"https:\/\/doi.org\/10.1109\/TCYB.2021.3111082\">IEEE Transactions on Cybernetics<\/ext-link><\/i> 52(12):13572\u221213585<\/p>","DOI":"10.1109\/TCYB.2021.3111082"},{"key":"key-10.48130\/ker-0026-0001-9","doi-asserted-by":"crossref","unstructured":"<p>Zhang R, Prokhorchuk A, Dauwels J. 2020. Deep reinforcement learning for traveling salesman problem with time windows and rejections. <i>2020 International Joint Conference on Neural Networks (IJCNN). July 19\u221224, 2020. Glasgow, United Kingdom<\/i>. USA: IEEE. pp. 1\u22128 doi: <ext-link ext-link-type=\"uri\" xlink:href=\"https:\/\/doi.org\/10.1109\/ijcnn48605.2020.9207026\">10.1109\/ijcnn48605.2020.9207026<\/ext-link><\/p>","DOI":"10.1109\/IJCNN48605.2020.9207026"},{"key":"key-10.48130\/ker-0026-0001-10","doi-asserted-by":"publisher","unstructured":"<p>Zhang R, Zhang C, Cao Z, Song W, Tan PS, et al. 2023. Learning to solve multiple-TSP with time window and rejections <i>via<\/i> deep reinforcement learning. <i><ext-link ext-link-type=\"uri\" xmlns:xlink=\"http:\/\/www.w3.org\/1999\/xlink\" xlink:href=\"https:\/\/doi.org\/10.1109\/TITS.2022.3207011\">IEEE Transactions on Intelligent Transportation Systems<\/ext-link><\/i> 24(1):1325\u22121336<\/p>","DOI":"10.1109\/TITS.2022.3207011"},{"key":"key-10.48130\/ker-0026-0001-11","doi-asserted-by":"publisher","unstructured":"<p>Golden BL, Levy L, Vohra R. 1987. The orienteering problem. <i><ext-link ext-link-type=\"uri\" xmlns:xlink=\"http:\/\/www.w3.org\/1999\/xlink\" xlink:href=\"https:\/\/doi.org\/10.1002\/1520-6750(198706)34:3307::aid-nav3220340302%3E3.0.co;2-d\">Naval Research Logistics<\/ext-link><\/i> 34(3):307\u2212318<\/p>","DOI":"10.1002\/1520-6750(198706)34:3307::aid-nav3220340302>3.0.co;2-d"},{"key":"key-10.48130\/ker-0026-0001-12","doi-asserted-by":"publisher","unstructured":"<p>Tsiligirides T. 1984. Heuristic methods applied to orienteering. <i><ext-link ext-link-type=\"uri\" xmlns:xlink=\"http:\/\/www.w3.org\/1999\/xlink\" xlink:href=\"https:\/\/doi.org\/10.1057\/jors.1984.162\">Journal of the Operational Research Society<\/ext-link><\/i> 35(9):797\u2212809<\/p>","DOI":"10.1057\/jors.1984.162"},{"key":"key-10.48130\/ker-0026-0001-13","doi-asserted-by":"publisher","unstructured":"<p>Kobeaga G, Merino M, Lozano JA. 2020. A revisited branch-and-cut algorithm for large-scale orienteering problems. <i><ext-link ext-link-type=\"uri\" xmlns:xlink=\"http:\/\/www.w3.org\/1999\/xlink\" xlink:href=\"https:\/\/doi.org\/10.48550\/arXiv.2011.02743\">arXiv<\/ext-link><\/i> 2011.02743<\/p>","DOI":"10.48550\/arXiv.2011.02743"},{"key":"key-10.48130\/ker-0026-0001-14","doi-asserted-by":"publisher","unstructured":"<p>Kobeaga G, Merino M, Lozano JA. 2018. An efficient evolutionary algorithm for the orienteering problem. <i><ext-link ext-link-type=\"uri\" xmlns:xlink=\"http:\/\/www.w3.org\/1999\/xlink\" xlink:href=\"https:\/\/doi.org\/10.1016\/j.cor.2017.09.003\">Computers &amp; Operations Research<\/ext-link><\/i> 90:42\u221259<\/p>","DOI":"10.1016\/j.cor.2017.09.003"},{"key":"key-10.48130\/ker-0026-0001-15","doi-asserted-by":"publisher","unstructured":"<p>Bellman R. 1957. A Markovian decision process. <i><ext-link ext-link-type=\"uri\" xmlns:xlink=\"http:\/\/www.w3.org\/1999\/xlink\" xlink:href=\"https:\/\/doi.org\/10.1512\/iumj.1957.6.56038\">Indiana University Mathematics Journal<\/ext-link><\/i> 6(4):679\u2212684<\/p>","DOI":"10.1512\/iumj.1957.6.56038"},{"key":"key-10.48130\/ker-0026-0001-16","doi-asserted-by":"publisher","unstructured":"<p>Karp RM. 1977. Probabilistic analysis of partitioning algorithms for the traveling-salesman problem in the plane. <i><ext-link ext-link-type=\"uri\" xmlns:xlink=\"http:\/\/www.w3.org\/1999\/xlink\" xlink:href=\"https:\/\/doi.org\/10.1287\/moor.2.3.209\">Mathematics of Operations Research<\/ext-link><\/i> 2(3):209\u2212224<\/p>","DOI":"10.1287\/moor.2.3.209"},{"key":"key-10.48130\/ker-0026-0001-17","doi-asserted-by":"crossref","unstructured":"<p>Traub V, Vygen J. 2024. <i>Approximation Algorithms for Traveling Salesman Problems<\/i>. Cambridge, UK: Cambridge University Press. doi: <ext-link ext-link-type=\"uri\" xlink:href=\"https:\/\/doi.org\/10.1017\/9781009445436\">10.1017\/9781009445436<\/ext-link><\/p>","DOI":"10.1017\/9781009445436"},{"key":"key-10.48130\/ker-0026-0001-18","doi-asserted-by":"publisher","unstructured":"<p>Strutz T. 2021. Travelling santa problem: optimization of a million-households tour within one hour. <i><ext-link ext-link-type=\"uri\" xmlns:xlink=\"http:\/\/www.w3.org\/1999\/xlink\" xlink:href=\"https:\/\/doi.org\/10.3389\/frobt.2021.652417\">Frontiers in Robotics and AI<\/ext-link><\/i> 8:652417<\/p>","DOI":"10.3389\/frobt.2021.652417"},{"key":"key-10.48130\/ker-0026-0001-19","doi-asserted-by":"publisher","unstructured":"<p>Valenzuela CL, Jones AJ. 1993. Evolutionary divide and conquer (I): a novel genetic approach to the TSP. <i><ext-link ext-link-type=\"uri\" xmlns:xlink=\"http:\/\/www.w3.org\/1999\/xlink\" xlink:href=\"https:\/\/doi.org\/10.1162\/evco.1993.1.4.313\">Evolutionary Computation<\/ext-link><\/i> 1(4):313\u2212333<\/p>","DOI":"10.1162\/evco.1993.1.4.313"},{"key":"key-10.48130\/ker-0026-0001-20","doi-asserted-by":"publisher","unstructured":"<p>Liao E, Liu C. 2018. A hierarchical algorithm based on density peaks clustering and ant colony optimization for traveling salesman problem. <i><ext-link ext-link-type=\"uri\" xmlns:xlink=\"http:\/\/www.w3.org\/1999\/xlink\" xlink:href=\"https:\/\/doi.org\/10.1109\/ACCESS.2018.2853129\">IEEE Access<\/ext-link><\/i> 6:38921\u221238933<\/p>","DOI":"10.1109\/ACCESS.2018.2853129"},{"key":"key-10.48130\/ker-0026-0001-21","doi-asserted-by":"publisher","unstructured":"<p>Mariescu-Istodor R, Fr\u00e4nti P. 2021. Solving the large-scale TSP problem in 1 h: santa Claus challenge 2020. <i><ext-link ext-link-type=\"uri\" xmlns:xlink=\"http:\/\/www.w3.org\/1999\/xlink\" xlink:href=\"https:\/\/doi.org\/10.3389\/frobt.2021.689908\">Frontiers in Robotics and AI<\/ext-link><\/i> 8:689908<\/p>","DOI":"10.3389\/frobt.2021.689908"},{"key":"key-10.48130\/ker-0026-0001-22","doi-asserted-by":"publisher","unstructured":"<p>Alanzi E, El Bachir Menai M. 2025. Solving the traveling salesman problem with machine learning: a review of recent advances and challenges. <i><ext-link ext-link-type=\"uri\" xmlns:xlink=\"http:\/\/www.w3.org\/1999\/xlink\" xlink:href=\"https:\/\/doi.org\/10.1007\/s10462-025-11267-x\">Artificial Intelligence Review<\/ext-link><\/i> 58(9):267<\/p>","DOI":"10.1007\/s10462-025-11267-x"},{"key":"key-10.48130\/ker-0026-0001-23","doi-asserted-by":"publisher","unstructured":"<p>Bengio Y, Lodi A, Prouvost A. 2021. Machine learning for combinatorial optimization: a methodological tour d'horizon. <i><ext-link ext-link-type=\"uri\" xmlns:xlink=\"http:\/\/www.w3.org\/1999\/xlink\" xlink:href=\"https:\/\/doi.org\/10.1016\/j.ejor.2020.07.063\">European Journal of Operational Research<\/ext-link><\/i> 290(2):405\u2212421<\/p>","DOI":"10.1016\/j.ejor.2020.07.063"},{"key":"key-10.48130\/ker-0026-0001-24","doi-asserted-by":"crossref","unstructured":"<p>Deudon M, Cournut P, Lacoste A, Adulyasak Y, Rousseau LM. 2018. Learning heuristics for the TSP by policy gradient. In <i>Integration of Constraint Programming, Artificial Intelligence, and Operations Research<\/i>, ed. van Hoeve WJ. Cham: Springer. pp. 170\u2212181 doi: <ext-link ext-link-type=\"uri\" xlink:href=\"10.1007\/978-3-319-93031-2_12\">10.1007\/978-3-319-93031-2_12<\/ext-link><\/p>","DOI":"10.1007\/978-3-319-93031-2_12"},{"key":"key-10.48130\/ker-0026-0001-25","unstructured":"<p>Vinyals O, Fortunato M, Jaitly N. 2015. Pointer Networks. <i>Advances in Neural Information Processing Systems 28 (NIPS 2015)<\/i>. pp. 1\u22129 <ext-link ext-link-type=\"uri\" xlink:href=\"https:\/\/proceedings.neurips.cc\/paper_files\/paper\/2015\/hash\/29921001f2f04bd3baee84a12e98098f-Abstract.html\">https:\/\/proceedings.neurips.cc\/paper_files\/paper\/2015\/hash\/29921001f2f04bd3baee84a12e98098f-Abstract.html<\/ext-link><\/p>"},{"key":"key-10.48130\/ker-0026-0001-26","doi-asserted-by":"publisher","unstructured":"<p>Bello I, Pham H, Le QV, Norouzi M, Bengio S. 2016. Neural combinatorial optimization with reinforcement learning. <i><ext-link ext-link-type=\"uri\" xmlns:xlink=\"http:\/\/www.w3.org\/1999\/xlink\" xlink:href=\"https:\/\/doi.org\/10.48550\/arXiv.1611.09940\">arXiv<\/ext-link><\/i> 1611.09940<\/p>","DOI":"10.48550\/arXiv.1611.09940"},{"key":"key-10.48130\/ker-0026-0001-27","doi-asserted-by":"publisher","unstructured":"<p>Kool W, van Hoof H, Welling M. 2018. Attention, learn to solve routing problems! <i><ext-link ext-link-type=\"uri\" xmlns:xlink=\"http:\/\/www.w3.org\/1999\/xlink\" xlink:href=\"https:\/\/doi.org\/10.48550\/arXiv.1803.08475\">arXiv<\/ext-link><\/i> 1803.08475<\/p>","DOI":"10.48550\/arXiv.1803.08475"},{"key":"key-10.48130\/ker-0026-0001-28","doi-asserted-by":"publisher","unstructured":"<p>Wang J, Xiao C, Wang S, Ruan Y. 2023. Reinforcement learning for the traveling salesman problem: Performance comparison of three algorithms. <i><ext-link ext-link-type=\"uri\" xmlns:xlink=\"http:\/\/www.w3.org\/1999\/xlink\" xlink:href=\"https:\/\/doi.org\/10.1049\/tje2.12303\">The Journal of Engineering<\/ext-link><\/i> 2023(9):e12303<\/p>","DOI":"10.1049\/tje2.12303"},{"key":"key-10.48130\/ker-0026-0001-29","doi-asserted-by":"publisher","unstructured":"<p>Bresson X, Laurent T. 2021. The transformer network for the traveling salesman problem. <i><ext-link ext-link-type=\"uri\" xmlns:xlink=\"http:\/\/www.w3.org\/1999\/xlink\" xlink:href=\"https:\/\/doi.org\/10.48550\/arXiv.2103.03012\">arXiv<\/ext-link><\/i> 2103.03012<\/p>","DOI":"10.48550\/arXiv.2103.03012"},{"key":"key-10.48130\/ker-0026-0001-30","doi-asserted-by":"publisher","unstructured":"<p>Dai H, Khalil EB, Zhang Y, Dilkina B, Song L. 2017. Learning combinatorial optimization algorithms over graphs. <i><ext-link ext-link-type=\"uri\" xmlns:xlink=\"http:\/\/www.w3.org\/1999\/xlink\" xlink:href=\"https:\/\/doi.org\/10.48550\/arXiv.1704.01665\">arXiv<\/ext-link><\/i> 1704.01665<\/p>","DOI":"10.48550\/arXiv.1704.01665"},{"key":"key-10.48130\/ker-0026-0001-31","doi-asserted-by":"publisher","unstructured":"<p>Joshi CK, Laurent T, Bresson X. 2019. An efficient graph convolutional network technique for the travelling salesman problem. <i><ext-link ext-link-type=\"uri\" xmlns:xlink=\"http:\/\/www.w3.org\/1999\/xlink\" xlink:href=\"https:\/\/doi.org\/10.48550\/arXiv.1906.01227\">arXiv<\/ext-link><\/i> 1906.01227<\/p>","DOI":"10.48550\/arXiv.1906.01227"},{"key":"key-10.48130\/ker-0026-0001-32","doi-asserted-by":"publisher","unstructured":"<p>BinJubier MB, Ismail MA, Tusher EH, Aljanabi M, University A. 2024. A GPU accelerated parallel genetic algorithm for the traveling salesman problem. <i><ext-link ext-link-type=\"uri\" xmlns:xlink=\"http:\/\/www.w3.org\/1999\/xlink\" xlink:href=\"https:\/\/doi.org\/10.30880\/jscdm.2024.05.02.010\">Journal of Soft Computing and Data Mining<\/ext-link><\/i> 5(2):137\u2212150<\/p>","DOI":"10.30880\/jscdm.2024.05.02.010"},{"key":"key-10.48130\/ker-0026-0001-33","doi-asserted-by":"publisher","unstructured":"<p>Ruan Y, Cai W, Wang J. 2024. Combining reinforcement learning algorithm and genetic algorithm to solve the traveling salesman problem. <i><ext-link ext-link-type=\"uri\" xmlns:xlink=\"http:\/\/www.w3.org\/1999\/xlink\" xlink:href=\"https:\/\/doi.org\/10.1049\/tje2.12393\">The Journal of Engineering<\/ext-link><\/i> 2024(6):e12393<\/p>","DOI":"10.1049\/tje2.12393"},{"key":"key-10.48130\/ker-0026-0001-34","doi-asserted-by":"publisher","unstructured":"<p>Watkins CJCH, Dayan P. 1992. Q-learning. <i><ext-link ext-link-type=\"uri\" xmlns:xlink=\"http:\/\/www.w3.org\/1999\/xlink\" xlink:href=\"https:\/\/doi.org\/10.1007\/BF00992698\">Machine Learning<\/ext-link><\/i> 8(3):279\u2212292<\/p>","DOI":"10.1007\/BF00992698"},{"key":"key-10.48130\/ker-0026-0001-35","doi-asserted-by":"publisher","unstructured":"<p>Mnih V, Kavukcuoglu K, Silver D, Rusu AA, Veness J, et al. 2015. Human-level control through deep reinforcement learning. <i><ext-link ext-link-type=\"uri\" xmlns:xlink=\"http:\/\/www.w3.org\/1999\/xlink\" xlink:href=\"https:\/\/doi.org\/10.1038\/nature14236\">Nature<\/ext-link><\/i> 518:529\u2212533<\/p>","DOI":"10.1038\/nature14236"},{"key":"key-10.48130\/ker-0026-0001-36","doi-asserted-by":"publisher","unstructured":"<p>Mnih V, Kavukcuoglu K, Silver D, Graves A, Antonoglou I, et al. 2013. Playing atari with deep reinforcement learning. <i><ext-link ext-link-type=\"uri\" xmlns:xlink=\"http:\/\/www.w3.org\/1999\/xlink\" xlink:href=\"https:\/\/doi.org\/10.48550\/arXiv.1312.5602\">arXiv<\/ext-link><\/i> 1312.5602<\/p>","DOI":"10.48550\/arXiv.1312.5602"},{"key":"key-10.48130\/ker-0026-0001-37","doi-asserted-by":"publisher","unstructured":"<p>Van Hasselt H, Guez A, Silver D. 2016. Deep reinforcement learning with double Q-learning. <i><ext-link ext-link-type=\"uri\" xmlns:xlink=\"http:\/\/www.w3.org\/1999\/xlink\" xlink:href=\"https:\/\/doi.org\/10.1609\/aaai.v30i1.10295\">Proceedings of the AAAI Conference on Artificial Intelligence<\/ext-link><\/i> 30(1):2094\u22123100<\/p>","DOI":"10.1609\/aaai.v30i1.10295"},{"key":"key-10.48130\/ker-0026-0001-38","unstructured":"<p>Hasselt H. 2010. Double Q-learning. <i>Advances in Neural Information Processing Systems 23 (NIPS 2010)<\/i>. pp. 1\u22129 <ext-link ext-link-type=\"uri\" xlink:href=\"https:\/\/proceedings.neurips.cc\/paper\/2010\/hash\/091d584fced301b442654dd8c23b3fc9-Abstract.html\">https:\/\/proceedings.neurips.cc\/paper\/2010\/hash\/091d584fced301b442654dd8c23b3fc9-Abstract.html<\/ext-link><\/p>"},{"key":"key-10.48130\/ker-0026-0001-39","doi-asserted-by":"publisher","unstructured":"<p>Schulman J, Moritz P, Levine S, Jordan M, Abbeel P. 2015. High-dimensional continuous control using generalized advantage estimation. <i><ext-link ext-link-type=\"uri\" xmlns:xlink=\"http:\/\/www.w3.org\/1999\/xlink\" xlink:href=\"https:\/\/doi.org\/10.48550\/arXiv.1506.02438\">arXiv<\/ext-link><\/i> 1506.02438<\/p>","DOI":"10.48550\/arXiv.1506.02438"},{"key":"key-10.48130\/ker-0026-0001-40","unstructured":"<p>Sutton RS, McAllester D, Singh S, Mansour Y. 1999. Policy gradient methods for reinforcement learning with function approximation. <i>Proceedings of the 13<sup>th<\/sup> International Conference on Neural Information Processing Systems, 29 November 1999, Denver, CO<\/i>. ACM. pp. 1057\u22121063 <ext-link ext-link-type=\"uri\" xlink:href=\"https:\/\/proceedings.neurips.cc\/paper_files\/paper\/1999\/hash\/464d828b85b0bed98e80ade0a5c43b0f-Abstract.html\">https:\/\/proceedings.neurips.cc\/paper_files\/paper\/1999\/hash\/464d828b85b0bed98e80ade0a5c43b0f-Abstract.html<\/ext-link><\/p>"},{"key":"key-10.48130\/ker-0026-0001-41","unstructured":"<p>Weng L. 2018. <i>Exploration strategies in deep reinforcement learning<\/i>. <ext-link ext-link-type=\"uri\" xlink:href=\"https:\/\/lilianweng.github.io\/posts\/2020-06-07-exploration-drl\/\">https:\/\/lilianweng.github.io\/posts\/2020-06-07-exploration-drl\/<\/ext-link> (Accessed 16 May 2024)<\/p>"},{"key":"key-10.48130\/ker-0026-0001-42","unstructured":"<p>Achiam J. 2018. <i>Spinning up in deep reinforcement learning<\/i>. <ext-link ext-link-type=\"uri\" xlink:href=\"https:\/\/spinningup.openai.com\/en\/latest\/algorithms\/sac.html\">https:\/\/spinningup.openai.com\/en\/latest\/algorithms\/sac.html<\/ext-link> (Accessed 16 May 2024)<\/p>"},{"key":"key-10.48130\/ker-0026-0001-43","doi-asserted-by":"publisher","unstructured":"<p>Williams RJ. 1992. Simple statistical gradient-following algorithms for connectionist reinforcement learning. <i><ext-link ext-link-type=\"uri\" xmlns:xlink=\"http:\/\/www.w3.org\/1999\/xlink\" xlink:href=\"https:\/\/doi.org\/10.1007\/BF00992696\">Machine Learning<\/ext-link><\/i> 8(3):229\u2212256<\/p>","DOI":"10.1007\/BF00992696"},{"key":"key-10.48130\/ker-0026-0001-44","unstructured":"<p>Mnih V, Badia AP, Mirza M, Graves A, Lillicrap T, et al. 2016. Asynchronous methods for deep reinforcement learning. <i>International Conference on Machine Learning, 20\u201322 June 2016, New York, USA<\/i>. vol. 48. PMLR. pp. 1928\u20131937 <ext-link ext-link-type=\"uri\" xlink:href=\"https:\/\/proceedings.mlr.press\/v48\/mniha16.html?ref=\">https:\/\/proceedings.mlr.press\/v48\/mniha16.html?ref=<\/ext-link><\/p>"},{"key":"key-10.48130\/ker-0026-0001-45","doi-asserted-by":"publisher","unstructured":"<p>Kirkpatrick J, Pascanu R, Rabinowitz N, Veness J, Desjardins G, et al. 2017. Overcoming catastrophic forgetting in neural networks. <i><ext-link ext-link-type=\"uri\" xmlns:xlink=\"http:\/\/www.w3.org\/1999\/xlink\" xlink:href=\"https:\/\/doi.org\/10.1073\/pnas.1611835114\">Proceedings of the National Academy of Sciences of the United States of America<\/ext-link><\/i> 114(13):3521\u22123526<\/p>","DOI":"10.1073\/pnas.1611835114"},{"key":"key-10.48130\/ker-0026-0001-46","doi-asserted-by":"publisher","unstructured":"<p>Schulman J, Wolski F, Dhariwal P, Radford A, Klimov O. 2017. Proximal policy optimization algorithms. <i><ext-link ext-link-type=\"uri\" xmlns:xlink=\"http:\/\/www.w3.org\/1999\/xlink\" xlink:href=\"https:\/\/doi.org\/10.48550\/arXiv.1707.06347\">arXiv<\/ext-link><\/i> 1707.06347<\/p>","DOI":"10.48550\/arXiv.1707.06347"},{"key":"key-10.48130\/ker-0026-0001-47","unstructured":"<p>Haarnoja T, Zhou A, Abbeel P, Levine S. 2018. Soft Actor-critic: off-policy maximum entropy deep reinforcement learning with a stochastic actor. <i>Proceedings of the 35<sup>th<\/sup> International Conference on Machine Learning, 10\u201315 July 2018, Stockholmsm\u00e4ssan, Stockholm Sweden<\/i>. vol. 80. PMLR. pp. 1861\u20131870 <ext-link ext-link-type=\"uri\" xlink:href=\"https:\/\/proceedings.mlr.press\/v80\/haarnoja18b\">https:\/\/proceedings.mlr.press\/v80\/haarnoja18b<\/ext-link><\/p>"},{"key":"key-10.48130\/ker-0026-0001-48","doi-asserted-by":"publisher","unstructured":"<p>Haarnoja T, Zhou A, Hartikainen K, Tucker G, Ha S, et al. 2018. Soft actor-critic algorithms and applications. <i><ext-link ext-link-type=\"uri\" xmlns:xlink=\"http:\/\/www.w3.org\/1999\/xlink\" xlink:href=\"https:\/\/doi.org\/10.48550\/arXiv.1812.05905\">arXiv<\/ext-link><\/i> 1812.05905<\/p>","DOI":"10.48550\/arXiv.1812.05905"},{"key":"key-10.48130\/ker-0026-0001-49","doi-asserted-by":"publisher","unstructured":"<p>Duan J, Wang W, Xiao L, Gao J, Li SE, et al. 2025. Distributional soft actor-critic with three refinements. <i><ext-link ext-link-type=\"uri\" xmlns:xlink=\"http:\/\/www.w3.org\/1999\/xlink\" xlink:href=\"https:\/\/doi.org\/10.1109\/TPAMI.2025.3537087\">IEEE Transactions on Pattern Analysis and Machine Intelligence<\/ext-link><\/i> 47(5):3935\u22123946<\/p>","DOI":"10.1109\/TPAMI.2025.3537087"},{"key":"key-10.48130\/ker-0026-0001-50","doi-asserted-by":"publisher","unstructured":"<p>Bahdanau D, Cho K, Bengio Y. 2014. Neural machine translation by jointly learning to align and translate. <i><ext-link ext-link-type=\"uri\" xmlns:xlink=\"http:\/\/www.w3.org\/1999\/xlink\" xlink:href=\"https:\/\/doi.org\/10.48550\/arXiv.1409.0473\">arXiv<\/ext-link><\/i> 1409.0473<\/p>","DOI":"10.48550\/arXiv.1409.0473"},{"key":"key-10.48130\/ker-0026-0001-51","unstructured":"<p>Vaswani A, Shazeer N, Parmar N, Uszkoreit J, Jones L, et al. 2017. Attention is all you need. In <i>Advances in Neural Information Processing Systems<\/i>. pp. 5998\u20136008 <ext-link ext-link-type=\"uri\" xlink:href=\"https:\/\/proceedings.neurips.cc\/paper\/2017\/hash\/3f5ee243547dee91fbd053c1c4a845aa-Abstract.html\">https:\/\/proceedings.neurips.cc\/paper\/2017\/hash\/3f5ee243547dee91fbd053c1c4a845aa-Abstract.html<\/ext-link><\/p>"},{"key":"key-10.48130\/ker-0026-0001-52","doi-asserted-by":"publisher","unstructured":"<p>Hochreiter S, Schmidhuber J. 1997. Long short-term memory. <i><ext-link ext-link-type=\"uri\" xmlns:xlink=\"http:\/\/www.w3.org\/1999\/xlink\" xlink:href=\"https:\/\/doi.org\/10.1162\/neco.1997.9.8.1735\">Neural Computation<\/ext-link><\/i> 9(8):1735\u22121780<\/p>","DOI":"10.1162\/neco.1997.9.8.1735"},{"key":"key-10.48130\/ker-0026-0001-53","unstructured":"<p>Alammar J. 2018. <i>The illustrated transformer<\/i>. <ext-link ext-link-type=\"uri\" xlink:href=\"https:\/\/jalammar.github.io\/illustrated-transformer\">https:\/\/jalammar.github.io\/illustrated-transformer<\/ext-link> (Accessed 16 May 2024)<\/p>"},{"key":"key-10.48130\/ker-0026-0001-54","doi-asserted-by":"publisher","unstructured":"<p>Nazari M, Oroojlooy A, Snyder LV, Tak\u00e1\u010d M. 2018. Reinforcement learning for solving the vehicle routing problem. <i><ext-link ext-link-type=\"uri\" xmlns:xlink=\"http:\/\/www.w3.org\/1999\/xlink\" xlink:href=\"https:\/\/doi.org\/10.48550\/arXiv.1802.04240\">arXiv<\/ext-link><\/i> 1802.04240<\/p>","DOI":"10.48550\/arXiv.1802.04240"},{"key":"key-10.48130\/ker-0026-0001-55","doi-asserted-by":"publisher","unstructured":"<p>Weng J, Chen H, Yan D, You K, Duburcq A, et al. 2021. Tianshou: a highly modularized deep reinforcement learning library. <i><ext-link ext-link-type=\"uri\" xmlns:xlink=\"http:\/\/www.w3.org\/1999\/xlink\" xlink:href=\"https:\/\/doi.org\/10.48550\/arXiv.2107.14171\">arXiv<\/ext-link><\/i> 2107.14171<\/p>","DOI":"10.48550\/arXiv.2107.14171"},{"key":"key-10.48130\/ker-0026-0001-56","unstructured":"<p>Pinto L, Davidson J, Sukthankar R, Gupta A. 2017. Robust adversarial reinforcement learning. <i>Proceedings of the 34<sup>th<\/sup> International Conference on Machine Learning. August 6\u221211, 2017, Sydney, NSW, Australia<\/i>. vol. 70. PMLR. pp. 2817\u22122826 <ext-link ext-link-type=\"uri\" xlink:href=\"https:\/\/proceedings.mlr.press\/v70\/pinto17a.html\">https:\/\/proceedings.mlr.press\/v70\/pinto17a.html<\/ext-link><\/p>"},{"key":"key-10.48130\/ker-0026-0001-57","doi-asserted-by":"crossref","unstructured":"<p>Liessner R, Schmitt J, Dietermann A, B\u00e4ker B. 2019. Hyperparameter optimization for deep reinforcement learning in vehicle energy management. <i>Proceedings of the 11<sup>th<\/sup> International Conference on Agents and Artificial Intelligence. February 19\u221221, 2019. Prague, Czech Republic<\/i>. Portugal: SciTePress. pp. 134\u2212144 doi: <ext-link ext-link-type=\"uri\" xlink:href=\"https:\/\/doi.org\/10.5220\/0007364701340144\">10.5220\/0007364701340144<\/ext-link><\/p>","DOI":"10.5220\/0007364701340144"}],"container-title":["The Knowledge Engineering Review"],"original-title":[],"deposited":{"date-parts":[[2026,5,8]],"date-time":"2026-05-08T06:49:26Z","timestamp":1778222966000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.maxapress.com\/article\/doi\/10.48130\/ker-0026-0001"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026]]},"references-count":57,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2026]]}},"URL":"https:\/\/doi.org\/10.48130\/ker-0026-0001","relation":{},"ISSN":["1469-8005"],"issn-type":[{"value":"1469-8005","type":"print"}],"subject":[],"published":{"date-parts":[[2026]]}}}