{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,9,16]],"date-time":"2025-09-16T17:57:50Z","timestamp":1758045470965,"version":"3.44.0"},"reference-count":64,"publisher":"Springer Science and Business Media LLC","issue":"10","license":[{"start":{"date-parts":[[2025,7,19]],"date-time":"2025-07-19T00:00:00Z","timestamp":1752883200000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2025,7,19]],"date-time":"2025-07-19T00:00:00Z","timestamp":1752883200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Artif Intell Rev"],"abstract":"<jats:title>Abstract<\/jats:title>\n          <jats:p>Market-making is an essential activity in every financial market. They provide liquidity to the system by placing buy and sell orders at multiple price levels. While performing this task, they aim to earn profit and manage inventory levels simultaneously. However, financial markets are not stationary environments; they constantly evolve, influenced by changes in participants, the occurrence of economic events, or the market trading hours, among others. This study introduces a novel approach to address the challenge of market-making in non-stationary financial markets with multi-objective Reinforcement Learning (RL). Traditional RL methods often struggle when applied to non-stationary environments, as the learned optimal policy may not be adapted to the new dynamics. We present Policy Weighting through Discounted Thompson Sampling (POW-dTS), a novel dynamic algorithm that adapts to changing market conditions by effectively weighting pre-trained policies across various contexts. Unlike some conventional methods, POW-dTS does not require additional artifacts such as change-point detection or models of transitions, making it robust against the unpredictability inherent in financial markets. Our approach focuses on optimizing trade profitability and managing inventory risk, the dual objectives of market makers. Through a detailed comparative analysis, we highlight the strengths and adaptability of POW-dTS against traditional techniques in non-stationary environments, demonstrating its potential to enhance market liquidity and efficiency.<\/jats:p>","DOI":"10.1007\/s10462-025-11312-9","type":"journal-article","created":{"date-parts":[[2025,7,19]],"date-time":"2025-07-19T07:05:40Z","timestamp":1752908740000},"update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":0,"title":["Policy weighting via discounted Thomson sampling for non-stationary market-making"],"prefix":"10.1007","volume":"58","author":[{"given":"\u00d3scar","family":"Fern\u00e1ndez Vicente","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Javier","family":"Garc\u00eda","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Fernando","family":"Fern\u00e1ndez","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"297","published-online":{"date-parts":[[2025,7,19]]},"reference":[{"issue":"46","key":"11312_CR1","first-page":"1","volume":"17","author":"S Abdallah","year":"2016","unstructured":"Abdallah S, Kaisers M (2016) Addressing environment non-stationarity by repeating Q-learning updates. J Mach Learn Res 17(46):1\u201331","journal-title":"J Mach Learn Res"},{"issue":"6","key":"11312_CR2","doi-asserted-by":"publisher","first-page":"26","DOI":"10.1109\/MSP.2017.2743240","volume":"34","author":"K Arulkumaran","year":"2017","unstructured":"Arulkumaran K, Deisenroth MP, Brundage M, Bharath AA (2017) Deep reinforcement learning: a brief survey. IEEE Signal Process Mag 34(6):26\u201338. https:\/\/doi.org\/10.1109\/MSP.2017.2743240","journal-title":"IEEE Signal Process Mag"},{"issue":"1","key":"11312_CR3","doi-asserted-by":"publisher","DOI":"10.1088\/1742-6596\/1722\/1\/012096","volume":"1722","author":"FC Asyuraa","year":"2021","unstructured":"Asyuraa FC, Abdullah S, Sutanto TE (2021) Empirical evaluation on discounted Thompson sampling for multi-armed bandit problem with piecewise-stationary Bernoulli arms. J Phys Conf Ser 1722(1):012096. https:\/\/doi.org\/10.1088\/1742-6596\/1722\/1\/012096","journal-title":"J Phys Conf Ser"},{"key":"11312_CR4","doi-asserted-by":"publisher","first-page":"291","DOI":"10.1016\/j.neucom.2020.11.050","volume":"428","author":"C Atkinson","year":"2021","unstructured":"Atkinson C, McCane B, Szymanski L, Robins A (2021) Pseudo-rehearsal: achieving deep reinforcement learning without catastrophic forgetting. Neurocomputing 428:291\u2013307. https:\/\/doi.org\/10.1016\/j.neucom.2020.11.050","journal-title":"Neurocomputing"},{"key":"11312_CR5","doi-asserted-by":"publisher","first-page":"217","DOI":"10.1080\/14697680701381228","volume":"8","author":"M Avellaneda","year":"2008","unstructured":"Avellaneda M, Stoikov S (2008) High-frequency trading in a limit order book. Quant Financ 8:217\u2013224. https:\/\/doi.org\/10.1080\/14697680701381228","journal-title":"Quant Financ"},{"key":"11312_CR6","doi-asserted-by":"publisher","DOI":"10.1016\/j.physa.2023.128901","volume":"622","author":"A Brini","year":"2023","unstructured":"Brini A, Tantari D (2023) Deep reinforcement trading with predictable returns. Physica A 622:128901. https:\/\/doi.org\/10.1016\/j.physa.2023.128901","journal-title":"Physica A"},{"key":"11312_CR7","doi-asserted-by":"publisher","unstructured":"Buzzega P, Boschini M, Porrello A, Calderara S (2021) Rethinking experience replay: a bag of tricks for continual learning. In: 2020 25th international conference on pattern recognition (ICPR). IEEE Computer Society, Los Alamitos, CA, pp 2180\u20132187. https:\/\/doi.org\/10.1109\/ICPR48806.2021.9412614","DOI":"10.1109\/ICPR48806.2021.9412614"},{"key":"11312_CR8","doi-asserted-by":"crossref","unstructured":"Byrd D, Hybinette M, Balch TH (2020) ABIDES: towards high-fidelity multi-agent market simulation. In: Proceedings of the 2020 ACM SIGSIM conference on principles of advanced discrete simulation. Association for Computing Machinery, New York, NY, pp 11\u201322","DOI":"10.1145\/3384441.3395986"},{"key":"11312_CR9","doi-asserted-by":"publisher","unstructured":"Chen Z, Liu B (2018) Continual learning and catastrophic forgetting. In: Lifelong machine learning. Springer International Publishing, Cham, pp 55\u201375. https:\/\/doi.org\/10.1007\/978-3-031-01581-6_4","DOI":"10.1007\/978-3-031-01581-6_4"},{"key":"11312_CR10","doi-asserted-by":"publisher","unstructured":"Choi SPM, Yeung D-Y, Zhang NL (2001) Hidden-mode Markov decision processes for nonstationary sequential decision making. In Sun R, Giles CL (eds) Sequence learning: paradigms, algorithms, and applications. Springer, Berlin\/Heidelberg, pp 264\u2013287. https:\/\/doi.org\/10.1007\/3-540-44565-X_12","DOI":"10.1007\/3-540-44565-X_12"},{"key":"11312_CR11","doi-asserted-by":"publisher","unstructured":"Chung G, Chung M, Lee Y, Kim WC (2022) Market making under order stacking framework: a deep reinforcement learning approach. In: Proceedings of the third ACM international conference on AI in finance. Association for Computing Machinery, New York, NY, pp 223\u2013231. https:\/\/doi.org\/10.1145\/3533271.3561789","DOI":"10.1145\/3533271.3561789"},{"issue":"5","key":"11312_CR12","doi-asserted-by":"publisher","first-page":"1457","DOI":"10.1111\/j.1540-6261.1983.tb03834.x","volume":"38","author":"TE Copeland","year":"1983","unstructured":"Copeland TE, Galai D (1983) Information effects on the bid-ask spread. J Financ 38(5):1457\u20131469. https:\/\/doi.org\/10.1111\/j.1540-6261.1983.tb03834.x","journal-title":"J Financ"},{"key":"11312_CR13","doi-asserted-by":"publisher","unstructured":"da Silva BC, Basso EW, Perotto FS, C\u00a0Bazzan AL, Engel PM (2006) Improving reinforcement learning with context detection. In: Proceedings of the fifth international joint conference on autonomous agents and multiagent systems. Association for Computing Machinery, New York, NY, pp 810\u2013812. https:\/\/doi.org\/10.1145\/1160633.1160779","DOI":"10.1145\/1160633.1160779"},{"key":"11312_CR14","unstructured":"Dewey D (2014) Reinforcement learning and the reward engineering principle. In: 2014 AAAI spring symposium series"},{"key":"11312_CR15","doi-asserted-by":"publisher","DOI":"10.2139\/ssrn.2868473","author":"MF Dixon","year":"2017","unstructured":"Dixon MF (2017) High frequency market making with machine learning. SSRN Electron J. https:\/\/doi.org\/10.2139\/ssrn.2868473","journal-title":"SSRN Electron J"},{"key":"11312_CR16","doi-asserted-by":"publisher","unstructured":"Eschmann J (2021) Reward function design in reinforcement learning. In: Reinforcement learning algorithms: analysis and applications. Springer International Publishing, Cham, pp 25\u201333. https:\/\/doi.org\/10.1007\/978-3-030-41188-6_3","DOI":"10.1007\/978-3-030-41188-6_3"},{"key":"11312_CR64","doi-asserted-by":"publisher","DOI":"10.1016\/j.eswa.2025.128867","author":"\u00d3 Fern\u00e1ndez","year":"2025","unstructured":"Fern\u00e1ndez \u00d3, Garc\u00eda J, Fern\u00e1ndez F (2025). Optimizing market-making strategies: A multi-objective reinforcement learning approach with Pareto fronts. Exp Syst Appl 128867. https:\/\/doi.org\/10.1016\/j.eswa.2025.128867","journal-title":"Exp Syst Appl"},{"issue":"4","key":"11312_CR17","doi-asserted-by":"publisher","first-page":"128","DOI":"10.1016\/S1364-6613(99)01294-2","volume":"3","author":"RM French","year":"1999","unstructured":"French RM (1999) Catastrophic forgetting in connectionist networks. Trends Cogn Sci 3(4):128\u2013135","journal-title":"Trends Cogn Sci"},{"key":"11312_CR18","doi-asserted-by":"publisher","DOI":"10.1016\/j.engappai.2020.104021","volume":"96","author":"J Garc\u00eda","year":"2020","unstructured":"Garc\u00eda J, Majadas R, Fern\u00e1ndez F (2020) Learning adversarial attack policies through multi-objective reinforcement learning. Eng Appl Artif Intell 96:104021. https:\/\/doi.org\/10.1016\/j.engappai.2020.104021","journal-title":"Eng Appl Artif Intell"},{"key":"11312_CR19","doi-asserted-by":"publisher","first-page":"167","DOI":"10.1007\/s10994-006-8365-9","volume":"65","author":"AP George","year":"2006","unstructured":"George AP, Powell WB (2006) Adaptive stepsizes for recursive estimation with applications in approximate dynamic programming. Mach Learn 65:167\u2013198","journal-title":"Mach Learn"},{"issue":"5","key":"11312_CR20","doi-asserted-by":"publisher","first-page":"295","DOI":"10.1007\/s42979-020-00312-x","volume":"1","author":"K Goutam","year":"2020","unstructured":"Goutam K, Balasubramanian S, Gera D, Sarma RR (2020) Layerout: freezing layers in deep neural networks. SN Comput Sci 1(5):295. https:\/\/doi.org\/10.1007\/s42979-020-00312-x","journal-title":"SN Comput Sci"},{"key":"11312_CR21","doi-asserted-by":"publisher","DOI":"10.1201\/b21350","volume-title":"The financial mathematics of market liquidity: from optimal execution to market making","author":"O Gueant","year":"2016","unstructured":"Gueant O (2016) The financial mathematics of market liquidity: from optimal execution to market making, vol 33. CRC Press, Boca Raton"},{"issue":"3","key":"11312_CR22","doi-asserted-by":"publisher","first-page":"454","DOI":"10.1016\/j.econlet.2013.09.026","volume":"121","author":"SK Guharay","year":"2013","unstructured":"Guharay SK, Thakur GS, Goodman FJ, Rosen SL, Houser D (2013) Analysis of non-stationary dynamics in the financial system. Econ Lett 121(3):454\u2013457. https:\/\/doi.org\/10.1016\/j.econlet.2013.09.026","journal-title":"Econ Lett"},{"key":"11312_CR23","unstructured":"Hadoux E, Beynier A, Weng P (2014) Sequential decision-making under non-stationary environments via sequential change-point detection. In: Learning over multiple contexts (LMCE), Nancy, France. https:\/\/hal.science\/hal-01200817"},{"issue":"1","key":"11312_CR24","doi-asserted-by":"publisher","first-page":"26","DOI":"10.1007\/s10458-022-09552-y","volume":"36","author":"CF Hayes","year":"2022","unstructured":"Hayes CF, R\u0103dulescu R, Bargiacchi E, K\u00e4llstr\u00f6m J, Macfarlane M, Reymond M, Roijers DM (2022) A practical guide to multi-objective reinforcement learning and planning. Auton Agent Multi-Agent Syst 36(1):26. https:\/\/doi.org\/10.1007\/s10458-022-09552-y","journal-title":"Auton Agent Multi-Agent Syst"},{"issue":"1","key":"11312_CR25","doi-asserted-by":"publisher","first-page":"47","DOI":"10.1016\/0304-405X(81)90020-9","volume":"9","author":"T Ho","year":"1981","unstructured":"Ho T, Stoll HR (1981) Optimal dealer pricing under transactions and return uncertainty. J Financ Econ 9(1):47\u201373. https:\/\/doi.org\/10.1016\/0304-405X(81)90020-9","journal-title":"J Financ Econ"},{"key":"11312_CR26","doi-asserted-by":"publisher","DOI":"10.1016\/j.eswa.2023.119556","volume":"218","author":"J Jang","year":"2023","unstructured":"Jang J, Seong N (2023) Deep reinforcement learning for stock portfolio optimization by connecting with modern portfolio theory. Expert Syst Appl 218:119556. https:\/\/doi.org\/10.1016\/j.eswa.2023.119556","journal-title":"Expert Syst Appl"},{"key":"11312_CR27","unstructured":"Julian R, Swanson B, Sukhatme G, Levine S, Finn C, Hausman K (2021) Never stop learning: the effectiveness of fine-tuning in robotic reinforcement learning. In: Kober J, Ramos F, Tomlin C (eds) Proceedings of the 2020 conference on robot learning, vol\u00a0155. PMLR, pp 2120\u20132136. https:\/\/proceedings.mlr.press\/v155\/julian21a.html"},{"key":"11312_CR28","doi-asserted-by":"publisher","unstructured":"Kemker R, McClure M, Abitino A, Hayes T, Kanan C (2018) Measuring catastrophic forgetting in neural networks. In: Proceedings of the AAAI conference on artificial intelligence, vol 32, no 1. https:\/\/doi.org\/10.1609\/aaai.v32i1.11651","DOI":"10.1609\/aaai.v32i1.11651"},{"issue":"13","key":"11312_CR29","doi-asserted-by":"publisher","first-page":"3521","DOI":"10.1073\/pnas.1611835114","volume":"114","author":"J Kirkpatrick","year":"2017","unstructured":"Kirkpatrick J, Pascanu R, Rabinowitz N, Veness J, Desjardins G, Rusu AA, Milan K, Quan J, Ramalho T, Grabska-Barwinska A, Hassabis D (2017) Overcoming catastrophic forgetting in neural networks. Proc Natl Acad Sci 114(13):3521\u20133526. https:\/\/doi.org\/10.1073\/pnas.1611835114","journal-title":"Proc Natl Acad Sci"},{"key":"11312_CR30","unstructured":"Kumar P (2023) Deep reinforcement learning for high-frequency market making. In: Khan E, Gonen M (eds) Proceedings of the 14th Asian conference on machine learning, vol\u00a0189. PMLR, pp 531\u2013546. https:\/\/proceedings.mlr.press\/v189\/kumar23a.html"},{"key":"11312_CR31","doi-asserted-by":"publisher","first-page":"3366","DOI":"10.1109\/TPAMI.2021.3057446","volume":"44","author":"MD Lange","year":"2022","unstructured":"Lange MD, Aljundi R, Masana M, Parisot S, Jia X, Leonardis A, Tuytelaars T (2022) A continual learning survey: defying forgetting in classification tasks. IEEE Trans Pattern Anal Mach Intell 44:3366\u20133385. https:\/\/doi.org\/10.1109\/TPAMI.2021.3057446","journal-title":"IEEE Trans Pattern Anal Mach Intell"},{"key":"11312_CR32","unstructured":"Lee S, Seo Y, Lee K, Abbeel P, Shin J (2022) Offline-to-online reinforcement learning via balanced replay and pessimistic q-ensemble. In: Faust A, Hsu D, Neumann G (eds) Proceedings of the 5th conference on robot learning, vol\u00a0164. PMLR, pp 1702\u20131712. https:\/\/proceedings.mlr.press\/v164\/lee22d.html"},{"key":"11312_CR33","doi-asserted-by":"publisher","first-page":"399","DOI":"10.1016\/j.neunet.2022.06.023","volume":"153","author":"M-F Leung","year":"2022","unstructured":"Leung M-F, Wang J, Che H (2022) Cardinality-constrained portfolio selection via two-timescale duplex neurodynamic optimization. Neural Netw 153:399\u2013410. https:\/\/doi.org\/10.1016\/j.neunet.2022.06.023","journal-title":"Neural Netw"},{"issue":"3","key":"11312_CR35","doi-asserted-by":"publisher","first-page":"385","DOI":"10.1109\/TSMC.2014.2358639","volume":"45","author":"C Liu","year":"2015","unstructured":"Liu C, Xu X, Hu D (2015) Multiobjective reinforcement learning: a comprehensive overview. IEEE Trans Syst Man Cybern Syst 45(3):385\u2013398. https:\/\/doi.org\/10.1109\/TSMC.2014.2358639","journal-title":"IEEE Trans Syst Man Cybern Syst"},{"issue":"4","key":"11312_CR34","doi-asserted-by":"publisher","first-page":"596","DOI":"10.1007\/s11704-014-3312-6","volume":"8","author":"X Li","year":"2014","unstructured":"Li X, Deng X, Zhu S, Wang F, Xie H (2014) An intelligent market making strategy in algorithmic trading. Front Comput Sci 8(4):596\u2013608. https:\/\/doi.org\/10.1007\/s11704-014-3312-6","journal-title":"Front Comput Sci"},{"issue":"2","key":"11312_CR36","doi-asserted-by":"publisher","first-page":"134","DOI":"10.1016\/S0167-2789(96)00139-X","volume":"99","author":"R Manuca","year":"1996","unstructured":"Manuca R, Savit R (1996) Stationarity and nonstationarity in time series analysis. Physica D 99(2):134\u2013161. https:\/\/doi.org\/10.1016\/S0167-2789(96)00139-X","journal-title":"Physica D"},{"key":"11312_CR37","doi-asserted-by":"publisher","DOI":"10.1017\/CBO9780511805448","volume-title":"Dynamics of markets: the new financial economics","author":"JL McCauley","year":"2009","unstructured":"McCauley JL (2009) Dynamics of markets: the new financial economics. Cambridge University Press, Cambridge"},{"issue":"4","key":"11312_CR38","doi-asserted-by":"publisher","first-page":"287","DOI":"10.1016\/S1058-3300(02)00060-5","volume":"11","author":"TH McInish","year":"2002","unstructured":"McInish TH, Van Ness BF, Van Ness RA (2002) After-hours trading of NYSE stocks on the regional stock exchanges. Rev Financ Econ 11(4):287\u2013297. https:\/\/doi.org\/10.1016\/S1058-3300(02)00060-5","journal-title":"Rev Financ Econ"},{"key":"11312_CR39","doi-asserted-by":"publisher","first-page":"529","DOI":"10.1038\/nature14236","volume":"518","author":"V Mnih","year":"2015","unstructured":"Mnih V, Kavukcuoglu K, Silver D, Rusu AA, Veness J, Bellemare MG, Hassabis D (2015) Human-level control through deep reinforcement learning. Nature 518:529\u2013533. https:\/\/doi.org\/10.1038\/nature14236","journal-title":"Nature"},{"key":"11312_CR40","doi-asserted-by":"crossref","unstructured":"Ngatchou P, Zarei A, El-Sharkawi A (2005) Pareto multi objective optimization. In: Proceedings of the 13th international conference on, intelligent systems application to power systems. pp 84\u201391","DOI":"10.1109\/ISAP.2005.1599245"},{"key":"11312_CR41","doi-asserted-by":"crossref","unstructured":"Noda I (2010) Recursive adaptation of stepsize parameter for non-stationary environments. In: Adaptive and learning agents: second workshop, ala 2009, held as part of the AAMAS 2009 conference in Budapest, Hungary, May 12, 2009. Revised selected papers 9. pp 74\u201390","DOI":"10.1007\/978-3-642-11814-2_5"},{"issue":"3","key":"11312_CR42","doi-asserted-by":"publisher","first-page":"407","DOI":"10.1016\/S0167-4870(02)00189-7","volume":"25","author":"T Oberlechner","year":"2004","unstructured":"Oberlechner T, Hocking S (2004) Information sources, news, and rumors in financial markets: insights into the foreign exchange market. J Econ Psychol 25(3):407\u2013424. https:\/\/doi.org\/10.1016\/S0167-4870(02)00189-7","journal-title":"J Econ Psychol"},{"issue":"4","key":"11312_CR43","doi-asserted-by":"publisher","first-page":"361","DOI":"10.2307\/2330686","volume":"21","author":"M O\u2019Hara","year":"1986","unstructured":"O\u2019Hara M, Oldfield GS (1986) The microeconomics of market making. J Financ Quant Anal 21(4):361\u2013376. https:\/\/doi.org\/10.2307\/2330686","journal-title":"J Financ Quant Anal"},{"key":"11312_CR44","doi-asserted-by":"publisher","DOI":"10.1145\/3459991","author":"S Padakandla","year":"2021","unstructured":"Padakandla S (2021) A survey of reinforcement learning algorithms for dynamically varying environments. ACM Comput Surv. https:\/\/doi.org\/10.1145\/3459991","journal-title":"ACM Comput Surv"},{"key":"11312_CR45","doi-asserted-by":"publisher","unstructured":"Padakandla S, J PK, Bhatnagar S (2019) Reinforcement learning in non-stationary environments. Appl Intell. https:\/\/doi.org\/10.1007\/s10489-020-01758-5","DOI":"10.1007\/s10489-020-01758-5"},{"issue":"2","key":"11312_CR46","doi-asserted-by":"publisher","first-page":"123","DOI":"10.1080\/09540099550039318","volume":"7","author":"A Robins","year":"1995","unstructured":"Robins A (1995) Catastrophic forgetting, rehearsal and pseudorehearsal. Connect Sci 7(2):123\u2013146. https:\/\/doi.org\/10.1080\/09540099550039318","journal-title":"Connect Sci"},{"key":"11312_CR47","unstructured":"Rolnick D, Ahuja A, Schwarz J, Lillicrap T, Wayne G (2019) Experience replay for continual learning. In: Wallach H, Larochelle H, Beygelzimer A, d\u2019Alch\u00e9-Buc F, Fox E, Garnett R (eds) Advances in neural information processing systems, vol\u00a032. Curran Associates, Inc. https:\/\/proceedings.neurips.cc\/paper_files\/paper\/2019\/file\/fa7cdfad1a5aaf8370ebeda47a1ff1c3-Paper.pdf"},{"key":"11312_CR48","unstructured":"Serra J, Suris D, Miron M, Karatzoglou A (2018) Overcoming catastrophic forgetting with hard attention to the task. In: Dy J, Krause A (eds) Proceedings of the 35th international conference on machine learning, vol\u00a080. PMLR, pp 4548\u20134557. https:\/\/proceedings.mlr.press\/v80\/serra18a.html"},{"key":"11312_CR49","doi-asserted-by":"publisher","unstructured":"Shen Z, Liu Z, Qin J, Savvides M, Cheng K-T (2021) Partial is better than all: revisiting fine-tuning strategy for few-shot learning. In: Proceedings of the AAAI conference on artificial intelligence, vol 35, no 11. pp 9594\u20139602. https:\/\/doi.org\/10.1609\/aaai.v35i11.17155","DOI":"10.1609\/aaai.v35i11.17155"},{"key":"11312_CR50","doi-asserted-by":"publisher","unstructured":"Shi J, Tang SH, Zhou C (2024) Market-making and hedging with market impact using deep reinforcement learning. In: Proceedings of the 5th ACM international conference on AI in finance. Association for Computing Machinery, New York, NY, pp 652\u2013659. https:\/\/doi.org\/10.1145\/3677052.3698646","DOI":"10.1145\/3677052.3698646"},{"key":"11312_CR51","doi-asserted-by":"publisher","unstructured":"Sorrenti A, Bellitto G, Salanitri F, Pennisi M, Spampinato C, Palazzo S (2023) Selective freezing for efficient continual learning. In: 2023 IEEE\/CVF international conference on computer vision workshops (ICCVW). IEEE Computer Society, pp 3542\u20133551. https:\/\/doi.org\/10.1109\/ICCVW60793.2023.00381","DOI":"10.1109\/ICCVW60793.2023.00381"},{"issue":"7","key":"11312_CR52","doi-asserted-by":"publisher","first-page":"1054","DOI":"10.1177\/0093650217705528","volume":"45","author":"N Strau\u00df","year":"2018","unstructured":"Strau\u00df N, Vliegenthart R, Verhoeven P (2018) Intraday news trading: the reciprocal relationships between the stock market and economic news. Commun Res 45(7):1054\u20131077. https:\/\/doi.org\/10.1177\/0093650217705528","journal-title":"Commun Res"},{"key":"11312_CR53","volume-title":"Reinforcement learning: an introduction","author":"RS Sutton","year":"2018","unstructured":"Sutton RS, Barto AG (2018) Reinforcement learning: an introduction, 2nd edn. The MIT Press, Cambridge","edition":"2"},{"issue":"3\/4","key":"11312_CR54","doi-asserted-by":"publisher","first-page":"285","DOI":"10.2307\/2332286","volume":"25","author":"WR Thompson","year":"1933","unstructured":"Thompson WR (1933) On the likelihood that one unknown probability exceeds another in view of the evidence of two samples. Biometrika 25(3\/4):285\u2013294","journal-title":"Biometrika"},{"issue":"1","key":"11312_CR56","doi-asserted-by":"publisher","first-page":"51","DOI":"10.1007\/s10994-010-5232-5","volume":"84","author":"P Vamplew","year":"2011","unstructured":"Vamplew P, Dazeley R, Berry A, Issabekov R, Dekker E (2011) Empirical evaluation methods for multiobjective reinforcement learning algorithms. Mach Learn 84(1):51\u201380. https:\/\/doi.org\/10.1007\/s10994-010-5232-5","journal-title":"Mach Learn"},{"key":"11312_CR55","doi-asserted-by":"crossref","unstructured":"Vamplew P, Yearwood J, Dazeley R, Berry A (2008) On the limitations of scalarisation for multi-objective reinforcement learning of pareto fronts. In: Wobcke W, Zhang M (eds) AI 2008: advances in artificial intelligence. Springer, Berlin\/Heidelberg, pp 372\u2013378","DOI":"10.1007\/978-3-540-89378-3_37"},{"key":"11312_CR57","doi-asserted-by":"crossref","unstructured":"Vicente \u00d3F, Fern\u00e1ndez F, Garc\u00eda J (2022) Deep Q-learning market makers in a multi-agent simulated stock market. In: Proceedings of the second ACM international conference on AI in finance. Association for Computing Machinery, New York, NY","DOI":"10.1145\/3490354.3494448"},{"key":"11312_CR58","doi-asserted-by":"publisher","DOI":"10.1007\/s10489-023-04647-9","author":"\u00d3F Vicente","year":"2023","unstructured":"Vicente \u00d3F, Fern\u00e1ndez F, Garc\u00eda J (2023) Automated market maker inventory management with deep reinforcement learning. Appl Intell. https:\/\/doi.org\/10.1007\/s10489-023-04647-9","journal-title":"Appl Intell"},{"key":"11312_CR59","doi-asserted-by":"publisher","unstructured":"Wang Y, Savani R, Gu A, Mascioli C, Turocy T, Wellman M (2024) Market making with learned beta policies. In: Proceedings of the 5th ACM international conference on AI in finance. Association for Computing Machinery, New York, NY, pp 643\u2013651. https:\/\/doi.org\/10.1145\/3677052.3698623","DOI":"10.1145\/3677052.3698623"},{"key":"11312_CR60","doi-asserted-by":"publisher","first-page":"142","DOI":"10.1016\/j.ins.2020.05.066","volume":"538","author":"X Wu","year":"2020","unstructured":"Wu X, Chen H, Wang J, Troiano L, Loia V, Fujita H (2020) Adaptive stock trading strategies with deep reinforcement learning methods. Inf Sci 538:142\u2013158. https:\/\/doi.org\/10.1016\/j.ins.2020.05.066","journal-title":"Inf Sci"},{"key":"11312_CR61","doi-asserted-by":"publisher","unstructured":"Zelman J, Stefanik M, Weiss M, Teichmann J (2024) Adversarial inverse reinforcement learning for market making. In: Proceedings of the 5th ACM international conference on AI in finance. Association for Computing Machinery, New York, NY, pp 81\u201389. https:\/\/doi.org\/10.1145\/3677052.3698641","DOI":"10.1145\/3677052.3698641"},{"key":"11312_CR62","doi-asserted-by":"publisher","unstructured":"Zhao M, Linetsky V (2021) High frequency automated market making algorithms with adverse selection risk control via reinforcement learning. In: Proceedings of the second ACM international conference on AI in finance. Association for Computing Machinery, New York, NY, pp 1\u20139. https:\/\/doi.org\/10.1145\/3490354.3494398","DOI":"10.1145\/3490354.3494398"},{"issue":"12","key":"11312_CR63","doi-asserted-by":"publisher","first-page":"2402","DOI":"10.1109\/LCOMM.2019.2941717","volume":"23","author":"M Zhou","year":"2019","unstructured":"Zhou M, Wang T, Wang S (2019) Spectrum sensing across multiple service providers: a discounted Thompson sampling method. IEEE Commun Lett 23(12):2402\u20132406. https:\/\/doi.org\/10.1109\/LCOMM.2019.2941717","journal-title":"IEEE Commun Lett"}],"container-title":["Artificial Intelligence Review"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10462-025-11312-9.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s10462-025-11312-9\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10462-025-11312-9.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,9,12]],"date-time":"2025-09-12T18:10:35Z","timestamp":1757700635000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s10462-025-11312-9"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,7,19]]},"references-count":64,"journal-issue":{"issue":"10","published-online":{"date-parts":[[2025,10]]}},"alternative-id":["11312"],"URL":"https:\/\/doi.org\/10.1007\/s10462-025-11312-9","relation":{},"ISSN":["1573-7462"],"issn-type":[{"type":"electronic","value":"1573-7462"}],"subject":[],"published":{"date-parts":[[2025,7,19]]},"assertion":[{"value":"25 June 2025","order":1,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"19 July 2025","order":2,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"The authors have no conflict of interest to declare that are relevant to the content of this article.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Conflict of interest"}},{"value":"Not applicable.","order":3,"name":"Ethics","group":{"name":"EthicsHeading","label":"Ethical approval"}},{"value":"Not applicable.","order":4,"name":"Ethics","group":{"name":"EthicsHeading","label":"Consent to participate"}},{"value":"The participant has consented to the submission of the case report to the journal.","order":5,"name":"Ethics","group":{"name":"EthicsHeading","label":"Consent for publication"}}],"article-number":"318"}}