{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,4,1]],"date-time":"2026-04-01T09:51:31Z","timestamp":1775037091667,"version":"3.50.1"},"reference-count":40,"publisher":"Springer Science and Business Media LLC","issue":"1","license":[{"start":{"date-parts":[[2025,10,17]],"date-time":"2025-10-17T00:00:00Z","timestamp":1760659200000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2025,10,17]],"date-time":"2025-10-17T00:00:00Z","timestamp":1760659200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"DOI":"10.13039\/100000086","name":"Directorate for Mathematical and Physical Sciences","doi-asserted-by":"publisher","award":["2204240"],"award-info":[{"award-number":["2204240"]}],"id":[{"id":"10.13039\/100000086","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["12471452"],"award-info":[{"award-number":["12471452"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100001230","name":"Macquarie University","doi-asserted-by":"crossref","id":[{"id":"10.13039\/501100001230","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["J Optim Theory Appl"],"published-print":{"date-parts":[[2026,1]]},"abstract":"<jats:title>Abstract<\/jats:title>\n                  <jats:p>This paper develops a hybrid deep reinforcement learning approach to manage an insurance portfolio for diffusion models. To address the model uncertainty, we adopt the recently developed modelling of exploration and exploitation strategies in a continuous-time decision-making process with reinforcement learning. We consider an insurance portfolio management problem in which an entropy-regularized reward function and corresponding relaxed stochastic controls are formulated. To obtain the optimal relaxed stochastic controls, we develop a Markov chain approximation and stochastic approximation-based iterative deep reinforcement learning algorithm where the probability distribution of the optimal stochastic controls is approximated by neural networks. In our hybrid algorithm, both Markov chain approximation and stochastic approximation are adopted in the learning processes. The idea of using the Markov chain approximation method to find initial guesses is proposed. A stochastic approximation is adopted to estimate the parameters of neural networks. Convergence analysis of the algorithm is presented. Numerical examples are provided to illustrate the performance of the algorithm.<\/jats:p>","DOI":"10.1007\/s10957-025-02858-3","type":"journal-article","created":{"date-parts":[[2025,10,17]],"date-time":"2025-10-17T04:15:53Z","timestamp":1760674553000},"update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":0,"title":["A Hybrid Deep Reinforcement Learning Method for Insurance Portfolio Management"],"prefix":"10.1007","volume":"208","author":[{"given":"Xiang","family":"Cheng","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-9488-2993","authenticated-orcid":false,"given":"Zhuo","family":"Jin","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Hailiang","family":"Yang","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"George","family":"Yin","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"297","published-online":{"date-parts":[[2025,10,17]]},"reference":[{"key":"2858_CR1","doi-asserted-by":"publisher","DOI":"10.1002\/9781119412540","volume-title":"Reinsurance: Actuarial and Statistical Aspects","author":"H Albrecher","year":"2017","unstructured":"Albrecher, H., Beirlant, J., Teugels, J.L.: Reinsurance: Actuarial and Statistical Aspects. Wiley, West Sussex (2017)"},{"key":"2858_CR2","first-page":"941","volume":"53","author":"K Arrow","year":"1963","unstructured":"Arrow, K.: Uncertainty and the welfare economics of medical care. Am. Econ. Rev. 53, 941\u2013973 (1963)","journal-title":"Am. Econ. Rev."},{"issue":"1","key":"2858_CR3","doi-asserted-by":"publisher","first-page":"143","DOI":"10.1007\/s11009-019-09767-9","volume":"24","author":"A Bachouch","year":"2022","unstructured":"Bachouch, A., Hur\u00e9, C., Langren\u00e9, N., Pham, H.: Deep neural networks algorithms for stochastic control problems on finite horizon: numerical applications. Methodol. Comput. Appl. Probab. 24(1), 143\u2013178 (2022)","journal-title":"Methodol. Comput. Appl. Probab."},{"issue":"4","key":"2858_CR4","first-page":"170","volume":"1","author":"K Borch","year":"1960","unstructured":"Borch, K.: Reciprocal reinsurance treaties. astin. Bulletin 1(4), 170\u2013191 (1960)","journal-title":"Bulletin"},{"issue":"6","key":"2858_CR5","doi-asserted-by":"publisher","first-page":"4065","DOI":"10.1214\/21-AAP1715","volume":"32","author":"R Carmona","year":"2022","unstructured":"Carmona, R., Laurir\u0300e, M.: Convergence analysis of machine learning algorithms for the numerical solution of mean field control and games: II\u2013the finite horizon case. Ann. Appl. Probab. 32(6), 4065\u20134105 (2022)","journal-title":"Ann. Appl. Probab."},{"issue":"7","key":"2858_CR6","doi-asserted-by":"publisher","first-page":"4306","DOI":"10.1287\/mnsc.2023.4902","volume":"70","author":"Z Chen","year":"2024","unstructured":"Chen, Z., Lu, Y., Zhang, J., Zhu, W.: Managing weather risk with a neural network-based index insurance. Manage. Sci. 70(7), 4306\u20134327 (2024)","journal-title":"Manage. Sci."},{"issue":"2","key":"2858_CR7","doi-asserted-by":"publisher","first-page":"449","DOI":"10.1017\/asb.2020.9","volume":"50","author":"X Cheng","year":"2020","unstructured":"Cheng, X., Jin, Z., Yang, H.: Optimal insurance strategies: a hybrid deep learning Markov chain approximation approach. ASTIN Bulletin. 50(2), 449\u2013477 (2020)","journal-title":"ASTIN Bulletin."},{"key":"2858_CR8","first-page":"485","volume":"101","author":"M Denuit","year":"2021","unstructured":"Denuit, M., Charpentier, A., Trufin, J.: Autocalibration and Tweedie-dominance for insurance pricing with machine learning. Insur.: Math. Econ. 101, 485\u2013497 (2021)","journal-title":"Insur.: Math. Econ."},{"key":"2858_CR9","doi-asserted-by":"crossref","unstructured":"E, W., Han, J., Jentzen, A.: Deep learning-based numerical methods for high-dimensional parabolic partial differential equations and backward stochastic differential equations. Commun. Math. Stat. 5(4), 349\u2013380 (2017)","DOI":"10.1007\/s40304-017-0117-6"},{"issue":"3","key":"2858_CR10","doi-asserted-by":"publisher","first-page":"1175","DOI":"10.1017\/asb.2018.23","volume":"48","author":"MA Fahrenwaldt","year":"2018","unstructured":"Fahrenwaldt, M.A., Weber, S., Weske, K.: Pricing of cyber insurance contracts in a network model. ASTIN Bulletin 48(3), 1175\u20131218 (2018)","journal-title":"ASTIN Bulletin"},{"key":"2858_CR11","first-page":"433","volume":"2","author":"B Finetti","year":"1957","unstructured":"Finetti, B.: Su un\u2019impostazione alternativa della teoria collettiva del rischio. Trans. XVth Int. Congr. Actuar. 2, 433\u2013443 (1957)","journal-title":"Trans. XVth Int. Congr. Actuar."},{"key":"2858_CR12","doi-asserted-by":"publisher","DOI":"10.1016\/j.automatica.2022.110177","volume":"139","author":"D Firoozi","year":"2022","unstructured":"Firoozi, D., Jaimungal, S.: Exploratory lqg mean field games with entropy regularization. Automatica 139, 110177 (2022)","journal-title":"Automatica"},{"issue":"1","key":"2858_CR13","doi-asserted-by":"publisher","first-page":"37","DOI":"10.1287\/opre.20.1.37","volume":"20","author":"HU Gerber","year":"1972","unstructured":"Gerber, H.U.: Games of economic survival with discrete and continuous income processes. Oper. Res. 20(1), 37\u201345 (1972)","journal-title":"Oper. Res."},{"issue":"3","key":"2858_CR14","doi-asserted-by":"publisher","first-page":"591","DOI":"10.1111\/jori.12359","volume":"88","author":"C Gomes","year":"2021","unstructured":"Gomes, C., Jin, Z., Yang, H.: Insurance fraud detection with unsupervised deep learning. J. Risk Insur. 88(3), 591\u2013624 (2021)","journal-title":"J. Risk Insur."},{"issue":"4","key":"2858_CR15","doi-asserted-by":"publisher","first-page":"1168","DOI":"10.1137\/20M1360700","volume":"3","author":"H Gu","year":"2021","unstructured":"Gu, H., Guo, X., Wei, X., Xu, R.: Mean-field controls with q-learning for cooperative marl: convergence and complexity analysis. SIAM J. Math. Data Sci. 3(4), 1168\u20131196 (2021)","journal-title":"SIAM J. Math. Data Sci."},{"issue":"4","key":"2858_CR16","doi-asserted-by":"publisher","first-page":"3239","DOI":"10.1287\/moor.2021.1238","volume":"47","author":"X Guo","year":"2022","unstructured":"Guo, X., Xu, R., Zariphopoulou, T.: Entropy regularization for mean field games with learning. Math. Oper. Res. 47(4), 3239\u20133260 (2022)","journal-title":"Math. Oper. Res."},{"issue":"2","key":"2858_CR17","doi-asserted-by":"publisher","first-page":"481","DOI":"10.1017\/asb.2017.45","volume":"48","author":"D Hainaut","year":"2018","unstructured":"Hainaut, D.: A neural-network analyzer for mortality forecast. ASTIN Bulletin. 48(2), 481\u2013508 (2018)","journal-title":"ASTIN Bulletin."},{"key":"2858_CR18","unstructured":"Han, J., E, W.: Deep learning approximation for stochastic control problems. arXiv preprint arXiv:1611.07422. (2016)"},{"issue":"37","key":"2858_CR19","doi-asserted-by":"publisher","first-page":"9163","DOI":"10.1073\/pnas.1811243115","volume":"115","author":"LP Hansen","year":"2018","unstructured":"Hansen, L.P., Miao, J.: Aversion to ambiguity and model misspecification in dynamic stochastic environments. Proc. Natl. Acad. Sci. 115(37), 9163\u20139168 (2018)","journal-title":"Proc. Natl. Acad. Sci."},{"issue":"2","key":"2858_CR20","doi-asserted-by":"publisher","first-page":"60","DOI":"10.1257\/aer.91.2.60","volume":"91","author":"LP Hansen","year":"2001","unstructured":"Hansen, L.P., Sargent, T.J.: Robust control and model uncertainty. Am. Econ. Rev. 91(2), 60\u201366 (2001)","journal-title":"Am. Econ. Rev."},{"issue":"2","key":"2858_CR21","doi-asserted-by":"publisher","first-page":"153","DOI":"10.1111\/1467-9965.00066","volume":"9","author":"BH H\u00f8jgaard","year":"1999","unstructured":"H\u00f8jgaard, B.H., Taksar, M.: Controlling risk exposure and dividends payout schemes: insurance company example. Math. Financ. 9(2), 153\u2013182 (1999)","journal-title":"Math. Financ."},{"issue":"11","key":"2858_CR22","first-page":"1039","volume":"4","author":"J Hu","year":"2003","unstructured":"Hu, J., Wellman, M.P.: Nash q-learning for general-sum stochastic games. J. Mach. Learn. Res. 4(11), 1039\u20131069 (2003)","journal-title":"J. Mach. Learn. Res."},{"issue":"1","key":"2858_CR23","doi-asserted-by":"publisher","first-page":"525","DOI":"10.1137\/20M1316640","volume":"59","author":"C Hur\u00e9","year":"2020","unstructured":"Hur\u00e9, C., Pham, H., Bachouch, A., Langren\u00e9, N.: Deep neural networks algorithms for stochastic control problems on finite horizon: convergence analysis. SIAM J. Numer. Anal. 59(1), 525\u2013557 (2020)","journal-title":"SIAM J. Numer. Anal."},{"issue":"1","key":"2858_CR24","doi-asserted-by":"publisher","first-page":"103","DOI":"10.1007\/s00780-021-00467-2","volume":"26","author":"S Jaimungal","year":"2022","unstructured":"Jaimungal, S.: Reinforcement learning and stochastic optimisation. Finance Stochast. 26(1), 103\u2013129 (2022)","journal-title":"Finance Stochast."},{"key":"2858_CR25","doi-asserted-by":"publisher","first-page":"246","DOI":"10.1007\/s10957-012-0263-7","volume":"159","author":"Z Jin","year":"2013","unstructured":"Jin, Z., Yin, G.: Numerical methods for optimal dividend payment and investment strategies of Markov-modulated jump diffusion models with regular and singular controls. J. Optim. Theory Appl. 159, 246\u2013271 (2013)","journal-title":"J. Optim. Theory Appl."},{"key":"2858_CR26","first-page":"262","volume":"96","author":"Z Jin","year":"2021","unstructured":"Jin, Z., Yang, H., Yin, G.: A hybrid deep learning method for optimal insurance strategies: algorithms and convergence analysis. Insur.: Math. Econ. 96, 262\u2013275 (2021)","journal-title":"Insur.: Math. Econ."},{"key":"2858_CR27","volume-title":"Approximation and Weak Convergence Methods for Random Processes with Applications to Stochastic Systems Theory","author":"H Kushner","year":"1984","unstructured":"Kushner, H.: Approximation and Weak Convergence Methods for Random Processes with Applications to Stochastic Systems Theory. MIT Press, Cambridge, Massachusetts (1984)"},{"key":"2858_CR28","doi-asserted-by":"publisher","DOI":"10.1007\/978-1-4613-0007-6","volume-title":"Numerical Methods for Stochstic Control Problems in Continuous Time","author":"H Kushner","year":"2001","unstructured":"Kushner, H., Dupuis, P.: Numerical Methods for Stochstic Control Problems in Continuous Time, 2nd edn. Springer, New York (2001)","edition":"2"},{"key":"2858_CR29","volume-title":"Stochastic Approximation and Recursive Algorithms and Applications","author":"H Kushner","year":"2003","unstructured":"Kushner, H., Yin, G.: Stochastic Approximation and Recursive Algorithms and Applications, 2nd edn. Springer, New York (2003)","edition":"2"},{"issue":"7540","key":"2858_CR30","doi-asserted-by":"publisher","first-page":"529","DOI":"10.1038\/nature14236","volume":"518","author":"V Mnih","year":"2015","unstructured":"Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A.A., Veness, J., Bellemare, M.G., Graves, A., Riedmiller, M., Fidjeland, A.K., Ostrovski, G., Petersen, S.: Human-level control through deep reinforcement learning. Nature 518(7540), 529\u2013533 (2015)","journal-title":"Nature"},{"issue":"4","key":"2858_CR31","doi-asserted-by":"publisher","first-page":"875","DOI":"10.1109\/72.935097","volume":"12","author":"J Moody","year":"2001","unstructured":"Moody, J., Saffell, M.: Learning to trade via direct reinforcement. IEEE Trans. Neural Netw. 12(4), 875\u2013889 (2001)","journal-title":"IEEE Trans. Neural Netw."},{"issue":"2","key":"2858_CR32","doi-asserted-by":"publisher","first-page":"207","DOI":"10.1017\/S1748499520000238","volume":"15","author":"R Richman","year":"2021","unstructured":"Richman, R.: Ai in actuarial science-a review of recent advances-part 1. Ann. Actuar. Sci. 15(2), 207\u2013229 (2021)","journal-title":"Ann. Actuar. Sci."},{"issue":"2","key":"2858_CR33","doi-asserted-by":"publisher","first-page":"230","DOI":"10.1017\/S174849952000024X","volume":"15","author":"R Richman","year":"2021","unstructured":"Richman, R.: Ai in actuarial science-a review of recent advances-part 2. Ann. Actuar. Sci. 15(2), 230\u2013258 (2021)","journal-title":"Ann. Actuar. Sci."},{"key":"2858_CR34","doi-asserted-by":"publisher","first-page":"560","DOI":"10.1109\/WIIAT.2008.88","volume":"2","author":"AA Salkham","year":"2008","unstructured":"Salkham, A.A., Cunningham, R., Garg, A., Cahill, V.: A collaborative reinforcement learning approach to urban traffic control optimization. IEEE\/WIC\/ACM International Conference on Web Intelligence and Intelligent Agent Technology. 2, 560\u2013566 (2008)","journal-title":"IEEE\/WIC\/ACM International Conference on Web Intelligence and Intelligent Agent Technology."},{"key":"2858_CR35","volume-title":"Reinforcement learning: An introduction","author":"RS Sutton","year":"2018","unstructured":"Sutton, R.S., Barto, A.G.: Reinforcement learning: An introduction. MIT Press, Cambridge, MA (2018)"},{"issue":"4","key":"2858_CR36","doi-asserted-by":"publisher","first-page":"1273","DOI":"10.1111\/mafi.12281","volume":"30","author":"H Wang","year":"2020","unstructured":"Wang, H., Zhou, X.Y.: Continuous-time mean-variance portfolio selection: a reinforcement learning framework. Math. Financ. 30(4), 1273\u20131308 (2020)","journal-title":"Math. Financ."},{"issue":"198","key":"2858_CR37","first-page":"1","volume":"21","author":"H Wang","year":"2020","unstructured":"Wang, H., Zariphopoulou, T., Zhou, X.Y.: Reinforcement learning in continuous time and space: a stochastic control approach. J. Mach. Learn. Res. 21(198), 1\u201334 (2020)","journal-title":"J. Mach. Learn. Res."},{"key":"2858_CR38","doi-asserted-by":"crossref","unstructured":"W\u00fcthrich, M.V., Buser, C.: Data analytics for non-life insurance pricing. Swiss Finance Institute Research Paper. 16-68 (2017)","DOI":"10.2139\/ssrn.2870308"},{"key":"2858_CR39","doi-asserted-by":"publisher","first-page":"465","DOI":"10.1080\/03461238.2018.1428681","volume":"6","author":"MV W\u00fcthrich","year":"2018","unstructured":"W\u00fcthrich, M.V.: Machine learning in individual claims reserving. Scand. Actuar. J. 6, 465\u2013480 (2018)","journal-title":"Scand. Actuar. J."},{"issue":"1","key":"2858_CR40","doi-asserted-by":"publisher","first-page":"240","DOI":"10.1137\/S1052623401392901","volume":"13","author":"G Yin","year":"2002","unstructured":"Yin, G., Liu, R., Zhang, Q.: Recursive algorithms for stock liquidation: a stochastic optimization approach. SIAM J. Optim. 13(1), 240\u2013263 (2002)","journal-title":"SIAM J. Optim."}],"container-title":["Journal of Optimization Theory and Applications"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10957-025-02858-3.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s10957-025-02858-3","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10957-025-02858-3.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,4,1]],"date-time":"2026-04-01T07:22:51Z","timestamp":1775028171000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s10957-025-02858-3"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,10,17]]},"references-count":40,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2026,1]]}},"alternative-id":["2858"],"URL":"https:\/\/doi.org\/10.1007\/s10957-025-02858-3","relation":{},"ISSN":["0022-3239","1573-2878"],"issn-type":[{"value":"0022-3239","type":"print"},{"value":"1573-2878","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,10,17]]},"assertion":[{"value":"17 October 2023","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"26 September 2025","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"17 October 2025","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}}],"article-number":"34"}}