{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,26]],"date-time":"2026-07-26T05:02:08Z","timestamp":1785042128538,"version":"3.55.0"},"reference-count":70,"publisher":"Springer Science and Business Media LLC","issue":"4","license":[{"start":{"date-parts":[[2024,3,18]],"date-time":"2024-03-18T00:00:00Z","timestamp":1710720000000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2024,3,18]],"date-time":"2024-03-18T00:00:00Z","timestamp":1710720000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"name":"Philosophy and Social Science Planning Project of Guangdong Province","award":["GD23YGL24"],"award-info":[{"award-number":["GD23YGL24"]}]},{"name":"Excellent Youth Foundation of Sichuan Province","award":["2020JDJQ0021"],"award-info":[{"award-number":["2020JDJQ0021"]}]},{"name":"Tianfu Ten-Thousand Talents Program of Sichuan Province","award":["0082204151153"],"award-info":[{"award-number":["0082204151153"]}]},{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"crossref","award":["72171160"],"award-info":[{"award-number":["72171160"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Artif Intell Rev"],"abstract":"<jats:title>Abstract<\/jats:title>\n                  <jats:p>\n                    In recent years, deep reinforcement learning (DRL) models have been successfully utilised to solve various classification problems. However, these models have never been applied to customer credit scoring in peer-to-peer (P2P) lending. Moreover, the imbalanced class distribution in experience replay, which may affect the performance of DRL models, has rarely been considered. Therefore, this article proposes a novel DRL model, namely a deep Q-network based on a balanced stratified prioritized experience replay (DQN-BSPER) model, for customer credit scoring in P2P lending. Firstly, customer credit scoring is formulated as a discrete-time finite-Markov decision process. Subsequently, a balanced stratified prioritized experience replay technology is presented to optimize the loss function of the deep Q-network model. This technology can not only balance the numbers of minority and majority experience samples in the mini-batch by using stratified sampling technology but also select more important experience samples for replay based on the priority principle. To verify the model performance, four evaluation measures are introduced for the empirical analysis of two real-world customer credit scoring datasets in P2P lending. The experimental results show that the DQN-BSPER model can outperform four benchmark DRL models and seven traditional benchmark classification models. In addition, the DQN-BSPER model with a discount factor\n                    <jats:italic>\u03b3<\/jats:italic>\n                    of 0.1 has excellent credit scoring performance.\n                  <\/jats:p>","DOI":"10.1007\/s10462-023-10697-9","type":"journal-article","created":{"date-parts":[[2024,3,18]],"date-time":"2024-03-18T01:01:52Z","timestamp":1710723712000},"update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":20,"title":["Deep reinforcement learning based on balanced stratified prioritized experience replay for customer credit scoring in peer-to-peer lending"],"prefix":"10.1007","volume":"57","author":[{"given":"Yadong","family":"Wang","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Yanlin","family":"Jia","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Sha","family":"Fan","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Jin","family":"Xiao","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"297","published-online":{"date-parts":[[2024,3,18]]},"reference":[{"issue":"4","key":"10697_CR1","doi-asserted-by":"crossref","first-page":"589","DOI":"10.1111\/j.1540-6261.1968.tb00843.x","volume":"23","author":"EI Altman","year":"1968","unstructured":"Altman EI (1968) Financial ratios, discriminant analysis and the prediction of corporate bankruptcy. J Finance 23(4):589\u2013609","journal-title":"J Finance"},{"issue":"6","key":"10697_CR2","doi-asserted-by":"crossref","first-page":"627","DOI":"10.1057\/palgrave.jors.2601545","volume":"54","author":"B Baesens","year":"2003","unstructured":"Baesens B, Van Gestel T, Viaene S, Stepanova M, Suykens J, Vanthienen J (2003) Benchmarking state-of-the-art classification algorithms for credit scoring. J Oper Res Soc 54(6):627\u2013635","journal-title":"J Oper Res Soc"},{"key":"10697_CR3","doi-asserted-by":"crossref","first-page":"209","DOI":"10.1016\/j.eswa.2019.05.042","volume":"134","author":"K Bastani","year":"2019","unstructured":"Bastani K, Asgari E, Namavari H (2019) Wide and deep learning for peer-to-peer lending. Expert Syst Appl 134:209\u2013224","journal-title":"Expert Syst Appl"},{"issue":"1","key":"10697_CR4","doi-asserted-by":"crossref","first-page":"20","DOI":"10.1145\/1007730.1007735","volume":"6","author":"GE Batista","year":"2004","unstructured":"Batista GE, Prati RC, Monard MC (2004) A study of the behavior of several methods for balancing machine learning training data. ACM SIGKDD Explor Newsl 6(1):20\u201329","journal-title":"ACM SIGKDD Explor Newsl"},{"issue":"1","key":"10697_CR5","doi-asserted-by":"crossref","first-page":"26","DOI":"10.1080\/01605682.2020.1838960","volume":"73","author":"G Blumenstock","year":"2022","unstructured":"Blumenstock G, Lessmann S, Seow H-V (2022) Deep learning for survival and competing risk modelling. J Oper Res Soc 73(1):26\u201338","journal-title":"J Oper Res Soc"},{"issue":"6","key":"10697_CR6","doi-asserted-by":"crossref","first-page":"1461","DOI":"10.1287\/opre.1110.0973","volume":"59","author":"E Borgonovo","year":"2011","unstructured":"Borgonovo E, Smith CL (2011) A study of interactions in the risk assessment of complex engineering systems: an application to space PSA. Oper Res 59(6):1461\u20131476","journal-title":"Oper Res"},{"issue":"7","key":"10697_CR7","doi-asserted-by":"crossref","first-page":"1145","DOI":"10.1016\/S0031-3203(96)00142-2","volume":"30","author":"AP Bradley","year":"1997","unstructured":"Bradley AP (1997) The use of the area under the ROC curve in the evaluation of machine learning algorithms. Pattern Recognit 30(7):1145\u20131159","journal-title":"Pattern Recognit"},{"key":"10697_CR8","doi-asserted-by":"crossref","first-page":"937","DOI":"10.1109\/TIFS.2020.3026553","volume":"16","author":"R Cai","year":"2020","unstructured":"Cai R, Li H, Wang S, Chen C, Kot A (2020) DRL-FAS: a novel framework based on deep reinforcement learning for face anti-spoofing. IEEE Trans Inf Forensics Secur 16:937\u2013951","journal-title":"IEEE Trans Inf Forensics Secur"},{"key":"10697_CR9","doi-asserted-by":"crossref","unstructured":"Chatterjee M, Namin A-S (2019) Detecting phishing websites through deep reinforcement learning. In: Proceedings of the IEEE 43rd annual computer software and applications conference, 2019. IEEE, pp 227\u2013232","DOI":"10.1109\/COMPSAC.2019.10211"},{"key":"10697_CR10","doi-asserted-by":"crossref","unstructured":"Chen S-Y, Yu Y, Da Q, Tan J, Huang H-K, Tang H-H (2018) Stabilizing reinforcement learning in dynamic environment with application to online recommendation. In: Proceedings of the 24th ACM SIGKDD international conference on knowledge discovery and data mining, 2018. ACM, pp 1187\u20131196","DOI":"10.1145\/3219819.3220122"},{"issue":"1","key":"10697_CR11","doi-asserted-by":"crossref","first-page":"224","DOI":"10.1016\/j.ijforecast.2011.07.006","volume":"28","author":"SF Crone","year":"2012","unstructured":"Crone SF, Finlay S (2012) Instance sampling in credit scoring: an empirical study of sample size and balancing. Int J Forecast 28(1):224\u2013238","journal-title":"Int J Forecast"},{"key":"10697_CR12","doi-asserted-by":"crossref","DOI":"10.1016\/j.asoc.2020.106263","volume":"91","author":"X Dastile","year":"2020","unstructured":"Dastile X, Celik T, Potsane M (2020) Statistical and machine learning models in credit scoring: a systematic literature survey. Appl Soft Comput 91:106263","journal-title":"Appl Soft Comput"},{"issue":"2","key":"10697_CR13","doi-asserted-by":"crossref","first-page":"535","DOI":"10.1016\/j.ejor.2021.10.045","volume":"301","author":"BJ De Moor","year":"2022","unstructured":"De Moor BJ, Gijsbrechts J, Boute RN (2022) Reward shaping to improve the performance of deep reinforcement learning in perishable inventory management. Eur J Oper Res 301(2):535\u2013545","journal-title":"Eur J Oper Res"},{"issue":"Jan","key":"10697_CR14","first-page":"1","volume":"7","author":"J Dem\u0161ar","year":"2006","unstructured":"Dem\u0161ar J (2006) Statistical comparisons of classifiers over multiple data sets. J Mach Learn Res 7(Jan):1\u201330","journal-title":"J Mach Learn Res"},{"key":"10697_CR15","doi-asserted-by":"crossref","DOI":"10.1016\/j.aei.2019.100977","volume":"42","author":"Y Ding","year":"2019","unstructured":"Ding Y, Ma L, Ma J, Suo M, Tao L, Cheng Y, Lu C (2019) Intelligent fault diagnosis for rotating machinery using deep Q-network based health state classification: a deep reinforcement learning approach. Adv Eng Inform 42:100977","journal-title":"Adv Eng Inform"},{"issue":"1","key":"10697_CR16","doi-asserted-by":"crossref","first-page":"315","DOI":"10.1287\/mnsc.2018.3216","volume":"66","author":"N Du","year":"2020","unstructured":"Du N, Li L, Lu T, Lu X (2020) Prosocial compliance in P2P lending: a natural field experiment. Manag Sci 66(1):315\u2013333","journal-title":"Manag Sci"},{"issue":"3","key":"10697_CR17","doi-asserted-by":"crossref","first-page":"1178","DOI":"10.1016\/j.ejor.2021.06.053","volume":"297","author":"E Dumitrescu","year":"2022","unstructured":"Dumitrescu E, Hue S, Hurlin C, Tokpavi S (2022) Machine learning for credit scoring: improving logistic regression with non-linear decision-tree effects. Eur J Oper Res 297(3):1178\u20131192","journal-title":"Eur J Oper Res"},{"issue":"6","key":"10697_CR18","doi-asserted-by":"crossref","first-page":"317","DOI":"10.1038\/s42256-020-0177-2","volume":"2","author":"C Fan","year":"2020","unstructured":"Fan C, Zeng L, Sun Y, Liu Y-Y (2020) Finding key players in complex networks through deep reinforcement learning. Nat Mach Intell 2(6):317\u2013324","journal-title":"Nat Mach Intell"},{"issue":"2","key":"10697_CR19","doi-asserted-by":"crossref","first-page":"517","DOI":"10.1016\/j.ejor.2015.07.013","volume":"249","author":"GB Fernandes","year":"2016","unstructured":"Fernandes GB, Artes R (2016) Spatial dependence in credit risk and its improvement in credit scoring. Eur J Oper Res 249(2):517\u2013524","journal-title":"Eur J Oper Res"},{"issue":"1","key":"10697_CR20","doi-asserted-by":"crossref","first-page":"86","DOI":"10.1214\/aoms\/1177731944","volume":"11","author":"M Friedman","year":"1940","unstructured":"Friedman M (1940) A comparison of alternative tests of significance for the problem of m rankings. Ann Math Stat 11(1):86\u201392","journal-title":"Ann Math Stat"},{"issue":"2","key":"10697_CR21","doi-asserted-by":"crossref","first-page":"178","DOI":"10.1287\/ijoc.1080.0305","volume":"21","author":"A Gosavi","year":"2009","unstructured":"Gosavi A (2009) Reinforcement learning: a tutorial survey and recent advances. INFORMS J Comput 21(2):178\u2013192","journal-title":"INFORMS J Comput"},{"issue":"1","key":"10697_CR22","doi-asserted-by":"crossref","first-page":"292","DOI":"10.1016\/j.ejor.2021.03.006","volume":"295","author":"BR Gunnarsson","year":"2021","unstructured":"Gunnarsson BR, Vanden Broucke S, Baesens B, \u00d3skarsd\u00f3ttir M, Lemahieu W (2021) Deep learning for credit scoring: do or don\u2019t? Eur J Oper Res 295(1):292\u2013305","journal-title":"Eur J Oper Res"},{"issue":"2","key":"10697_CR23","doi-asserted-by":"crossref","first-page":"417","DOI":"10.1016\/j.ejor.2015.05.050","volume":"249","author":"Y Guo","year":"2016","unstructured":"Guo Y, Zhou W, Luo C, Liu C, Xiong H (2016) Instance-based credit risk assessment for investment decisions in P2P lending. Eur J Oper Res 249(2):417\u2013426","journal-title":"Eur J Oper Res"},{"key":"10697_CR24","doi-asserted-by":"crossref","DOI":"10.1002\/9781118548387","volume-title":"Applied logistic regression","author":"DW Hosmer Jr","year":"2013","unstructured":"Hosmer DW Jr, Lemeshow S, Sturdivant RX (2013) Applied logistic regression, vol 398. Wiley, Hoboken"},{"issue":"6","key":"10697_CR25","doi-asserted-by":"crossref","first-page":"571","DOI":"10.1080\/03610928008827904","volume":"9","author":"RL Iman","year":"1980","unstructured":"Iman RL, Davenport JM (1980) Approximations of the critical region of the Fbietkan statistic. Commun Stat Theory Methods 9(6):571\u2013595","journal-title":"Commun Stat Theory Methods"},{"key":"10697_CR27","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1016\/j.ins.2014.02.137","volume":"275","author":"G Kou","year":"2014","unstructured":"Kou G, Peng Y, Wang G (2014) Evaluation of clustering algorithms for financial risk analysis using MCDM methods. Inf Sci 275:1\u201312","journal-title":"Inf Sci"},{"issue":"1","key":"10697_CR26","doi-asserted-by":"crossref","first-page":"39","DOI":"10.1186\/s40854-021-00256-y","volume":"7","author":"G Kou","year":"2021","unstructured":"Kou G, Olgu Akdeniz \u00d6, Din\u00e7er H, Y\u00fcksel S (2021a) Fintech investments in European banks: a hybrid IT2 fuzzy multidimensional decision-making approach. Financ Innov 7(1):39","journal-title":"Financ Innov"},{"key":"10697_CR28","volume":"140","author":"G Kou","year":"2021","unstructured":"Kou G, Xu Y, Peng Y, Shen F, Chen Y, Chang K, Kou S (2021b) Bankruptcy prediction for SMEs using transactional data and two-stage multiobjective feature selection. Decis Support Syst 140:113429","journal-title":"Decis Support Syst"},{"key":"10697_CR29","volume":"140","author":"K Lei","year":"2020","unstructured":"Lei K, Zhang B, Li Y, Yang M, Shen Y (2020) Time-driven feature-aware jointly deep reinforcement learning for financial signal representation and algorithmic trading. Expert Syst Appl 140:112872","journal-title":"Expert Syst Appl"},{"issue":"1","key":"10697_CR30","doi-asserted-by":"crossref","first-page":"124","DOI":"10.1016\/j.ejor.2015.05.030","volume":"247","author":"S Lessmann","year":"2015","unstructured":"Lessmann S, Baesens B, Seow H-V, Thomas LC (2015) Benchmarking state-of-the-art classification algorithms for credit scoring: an update of research. Eur J Oper Res 247(1):124\u2013136","journal-title":"Eur J Oper Res"},{"key":"10697_CR31","volume":"204","author":"H Li","year":"2020","unstructured":"Li H, Xu H (2020) Deep reinforcement learning for robust emotional classification in facial expression recognition. Knowl Based Syst 204:106172","journal-title":"Knowl Based Syst"},{"issue":"3","key":"10697_CR32","doi-asserted-by":"crossref","first-page":"1103","DOI":"10.1016\/j.ejor.2020.03.078","volume":"286","author":"Y Li","year":"2020","unstructured":"Li Y, Wang X, Djehiche B, Hu X (2020) Credit scoring by incorporating dynamic networked information. Eur J Oper Res 286(3):1103\u20131112","journal-title":"Eur J Oper Res"},{"issue":"10","key":"10697_CR33","first-page":"1202","volume":"33","author":"M Lim","year":"2021","unstructured":"Lim M, Abdullah A, Jhanjhi N (2021) Performance optimization of criminal network hidden link prediction model with deep reinforcement learning. J King Saud Univ Comput Inf Sci 33(10):1202\u20131210","journal-title":"J King Saud Univ Comput Inf Sci"},{"key":"10697_CR34","first-page":"1","volume":"5","author":"E Lin","year":"2020","unstructured":"Lin E, Chen Q, Qi X (2020) Deep reinforcement learning for imbalanced classification. Appl Intell 5:1\u201315","journal-title":"Appl Intell"},{"issue":"1","key":"10697_CR35","doi-asserted-by":"crossref","first-page":"166","DOI":"10.1016\/j.ejor.2019.10.049","volume":"283","author":"Y Liu","year":"2020","unstructured":"Liu Y, Chen Y, Jiang T (2020) Dynamic selective maintenance optimization for multi-state systems over a finite horizon: a deep reinforcement learning approach. Eur J Oper Res 283(1):166\u2013181","journal-title":"Eur J Oper Res"},{"key":"10697_CR36","doi-asserted-by":"crossref","DOI":"10.1016\/j.eswa.2019.112963","volume":"141","author":"M Lopez-Martin","year":"2020","unstructured":"Lopez-Martin M, Carro B, Sanchez-Esguevillas A (2020) Application of deep reinforcement learning to intrusion detection for supervised problems. Expert Syst Appl 141:112963","journal-title":"Expert Syst Appl"},{"key":"10697_CR37","doi-asserted-by":"crossref","first-page":"935","DOI":"10.1016\/j.neucom.2015.04.120","volume":"175","author":"O Loyola-Gonz\u00e1lez","year":"2016","unstructured":"Loyola-Gonz\u00e1lez O, Mart\u00ednez-Trinidad JF, Carrasco-Ochoa JA, Garc\u00eda-Borroto M (2016) Study of the impact of resampling methods for contrast pattern based classifiers in imbalanced databases. Neurocomputing 175:935\u2013947","journal-title":"Neurocomputing"},{"issue":"12","key":"10697_CR38","doi-asserted-by":"crossref","first-page":"3337","DOI":"10.1109\/TCYB.2018.2821369","volume":"48","author":"B Luo","year":"2018","unstructured":"Luo B, Yang Y, Liu D (2018) Adaptive Q-Learning for data-based optimal output regulation with experience replay. IEEE Trans Cybern 48(12):3337\u20133348","journal-title":"IEEE Trans Cybern"},{"issue":"7","key":"10697_CR39","doi-asserted-by":"crossref","first-page":"1060","DOI":"10.1057\/jors.2012.120","volume":"64","author":"AI Marqu\u00e9s","year":"2013","unstructured":"Marqu\u00e9s AI, Garc\u00eda V, S\u00e1nchez JS (2013) On the suitability of resampling techniques for the class imbalance problem in credit scoring. J Oper Res Soc 64(7):1060\u20131070","journal-title":"J Oper Res Soc"},{"key":"10697_CR40","doi-asserted-by":"crossref","DOI":"10.1016\/j.knosys.2019.105290","volume":"190","author":"C Martinez","year":"2020","unstructured":"Martinez C, Ramasso E, Perrin G, Rombaut M (2020) Adaptive early classification of temporal sequences using deep reinforcement learning. Knowl Based Syst 190:105290","journal-title":"Knowl Based Syst"},{"key":"10697_CR41","unstructured":"Mnih V, Kavukcuoglu K, Silver D, Graves A, Antonoglou I, Wierstra D, Riedmiller M (2013) Playing Atari with deep reinforcement learning. ArXiv preprint arXiv:1312.5602"},{"issue":"7540","key":"10697_CR42","doi-asserted-by":"crossref","first-page":"529","DOI":"10.1038\/nature14236","volume":"518","author":"V Mnih","year":"2015","unstructured":"Mnih V, Kavukcuoglu K, Silver D, Rusu AA, Veness J, Bellemare MG et al (2015) Human-level control through deep reinforcement learning. Nature 518(7540):529\u2013533","journal-title":"Nature"},{"key":"10697_CR43","doi-asserted-by":"crossref","first-page":"26","DOI":"10.1016\/j.asoc.2018.10.004","volume":"74","author":"M \u00d3skarsd\u00f3ttir","year":"2019","unstructured":"\u00d3skarsd\u00f3ttir M, Bravo C, Sarraute C, Vanthienen J, Baesens B (2019) The value of big data for credit scoring: enhancing financial inclusion using mobile phone data and social network analytics. Appl Soft Comput 74:26\u201339","journal-title":"Appl Soft Comput"},{"key":"10697_CR44","doi-asserted-by":"crossref","first-page":"108","DOI":"10.1016\/j.neunet.2019.08.009","volume":"120","author":"D Patel","year":"2019","unstructured":"Patel D, Hazan H, Saunders DJ, Siegelmann HT, Kozma R (2019) Improved robustness of reinforcement learning policies upon conversion to spiking neuronal network platforms applied to Atari Breakout game. Neural Netw 120:108\u2013115","journal-title":"Neural Netw"},{"issue":"2","key":"10697_CR45","first-page":"1","volume":"73","author":"G Petrides","year":"2020","unstructured":"Petrides G, Moldovan D, Coenen L, Guns T, Verbeke W (2020) Cost-sensitive learning for profit-driven credit scoring. J Oper Res Soc 73(2):1\u201313","journal-title":"J Oper Res Soc"},{"issue":"2","key":"10697_CR46","first-page":"103","volume":"11","author":"E Protopapadakis","year":"2019","unstructured":"Protopapadakis E, Niklis D, Doumpos M, Doulamis A, Zopounidis C (2019) Sample selection algorithms for credit risk modelling through data mining techniques. Int J Data Min Model Manag 11(2):103\u2013128","journal-title":"Int J Data Min Model Manag"},{"issue":"22","key":"10697_CR47","first-page":"41","volume":"3","author":"I Rish","year":"2001","unstructured":"Rish I (2001) An empirical study of the naive Bayes classifier. Workshop Empir Methods Artif Intell 3(22):41\u201346","journal-title":"Workshop Empir Methods Artif Intell"},{"key":"10697_CR48","first-page":"05952","volume":"1511","author":"T Schaul","year":"2015","unstructured":"Schaul T, Quan J, Antonoglou I, Silver D (2015) Prioritized experience replay. ArXiv Preprint arXiv 1511:05952","journal-title":"ArXiv Preprint arXiv"},{"issue":"3","key":"10697_CR49","doi-asserted-by":"crossref","first-page":"993","DOI":"10.1016\/j.ejor.2021.04.050","volume":"296","author":"M Schnaubelt","year":"2022","unstructured":"Schnaubelt M (2022) Deep reinforcement learning for the optimal placement of cryptocurrency limit orders. Eur J Oper Res 296(3):993\u20131006","journal-title":"Eur J Oper Res"},{"key":"10697_CR50","doi-asserted-by":"crossref","first-page":"113","DOI":"10.1016\/j.dss.2016.06.014","volume":"89","author":"C Serrano-Cinca","year":"2016","unstructured":"Serrano-Cinca C, Guti\u00e9rrez-Nieto B (2016) The use of profit scoring as an alternative to credit scoring systems in peer-to-peer (P2P) lending. Decis Support Syst 89:113\u2013122","journal-title":"Decis Support Syst"},{"issue":"6419","key":"10697_CR51","doi-asserted-by":"crossref","first-page":"1140","DOI":"10.1126\/science.aar6404","volume":"362","author":"D Silver","year":"2018","unstructured":"Silver D, Hubert T, Schrittwieser J, Antonoglou I, Lai M, Guez A et al (2018) A general reinforcement learning algorithm that masters chess, Shogi, and Go through self-play. Science 362(6419):1140\u20131144","journal-title":"Science"},{"issue":"1","key":"10697_CR52","doi-asserted-by":"crossref","first-page":"123","DOI":"10.1016\/j.ejor.2011.01.023","volume":"212","author":"MM So","year":"2011","unstructured":"So MM, Thomas LC (2011) Modelling the profitability of credit cards by Markov decision processes. Eur J Oper Res 212(1):123\u2013130","journal-title":"Eur J Oper Res"},{"key":"10697_CR53","volume":"278","author":"AY Sun","year":"2020","unstructured":"Sun AY (2020) Optimal carbon storage reservoir management through deep reinforcement learning. Appl Energy 278:115660","journal-title":"Appl Energy"},{"key":"10697_CR54","volume-title":"Reinforcement learning: an introduction","author":"R Sutton","year":"1998","unstructured":"Sutton R, Barto A (1998) Reinforcement learning: an introduction. MIT Press, Cambridge"},{"issue":"1","key":"10697_CR55","doi-asserted-by":"crossref","first-page":"281","DOI":"10.1109\/TSMCB.2008.2002909","volume":"39","author":"Y Tang","year":"2008","unstructured":"Tang Y, Zhang Y-Q, Chawla NV, Krasser S (2008) SVMs modeling for highly imbalanced classification. IEEE Trans Syst Man Cybern B 39(1):281\u2013288","journal-title":"IEEE Trans Syst Man Cybern B"},{"issue":"3","key":"10697_CR56","doi-asserted-by":"crossref","first-page":"893","DOI":"10.1016\/j.ejor.2005.07.024","volume":"173","author":"TB Trafalis","year":"2006","unstructured":"Trafalis TB, Gilbert RC (2006) Robust classification and regression using support vector machines. Eur J Oper Res 173(3):893\u2013909","journal-title":"Eur J Oper Res"},{"key":"10697_CR57","doi-asserted-by":"publisher","DOI":"10.1007\/s10479-022-04572-z","author":"W van Heeswijk","year":"2022","unstructured":"van Heeswijk W (2022) Strategic bidding in freight transport using deep reinforcement learning. Ann Oper Res. https:\/\/doi.org\/10.1007\/s10479-022-04572-z","journal-title":"Ann Oper Res"},{"key":"10697_CR58","doi-asserted-by":"crossref","first-page":"111","DOI":"10.1016\/j.dss.2018.06.011","volume":"112","author":"D Veganzones","year":"2018","unstructured":"Veganzones D, S\u00e9verin E (2018) An investigation of bankruptcy prediction in imbalanced datasets. Decis Support Syst 112:111\u2013124","journal-title":"Decis Support Syst"},{"issue":"4","key":"10697_CR59","doi-asserted-by":"crossref","first-page":"923","DOI":"10.1080\/01605682.2019.1705193","volume":"72","author":"H Wang","year":"2021","unstructured":"Wang H, Kou G, Peng Y (2021) Multi-class misclassification cost matrix for credit ratings in peer-to-peer lending. J Oper Res Soc 72(4):923\u2013934","journal-title":"J Oper Res Soc"},{"key":"10697_CR60","volume":"200","author":"Y Wang","year":"2022","unstructured":"Wang Y, Jia Y, Tian Y, Xiao J (2022) Deep reinforcement learning with the confusion-matrix-based dynamic reward function for customer credit scoring. Expert Syst Appl 200:117013","journal-title":"Expert Syst Appl"},{"issue":"3\u20134","key":"10697_CR61","first-page":"279","volume":"8","author":"CJ Watkins","year":"1992","unstructured":"Watkins CJ, Dayan P (1992) Q-learning. Mach Learn 8(3\u20134):279\u2013292","journal-title":"Mach Learn"},{"issue":"3","key":"10697_CR62","doi-asserted-by":"crossref","first-page":"1097","DOI":"10.1016\/j.ejor.2016.11.018","volume":"259","author":"M Wauters","year":"2017","unstructured":"Wauters M, Vanhoucke M (2017) A nearest neighbour extension to project duration forecasting with artificial intelligence. Eur J Oper Res 259(3):1097\u20131111","journal-title":"Eur J Oper Res"},{"key":"10697_CR63","doi-asserted-by":"crossref","unstructured":"Wilcoxon F (1992) Individual comparisons by ranking methods. In: Breakthroughs in statistics. Springer, Berlin, pp 196\u2013202","DOI":"10.1007\/978-1-4612-4380-9_16"},{"issue":"7896","key":"10697_CR64","doi-asserted-by":"crossref","first-page":"223","DOI":"10.1038\/s41586-021-04357-7","volume":"602","author":"PR Wurman","year":"2022","unstructured":"Wurman PR, Barrett S, Kawamoto K, MacGlashan J, Subramanian K, Walsh TJ et al (2022) Outracing champion Gran Turismo drivers with deep reinforcement learning. Nature 602(7896):223\u2013228","journal-title":"Nature"},{"key":"10697_CR65","volume":"159","author":"Y Xia","year":"2020","unstructured":"Xia Y, Zhao J, He L, Li Y, Niu M (2020) A novel tree-based dynamic heterogeneous ensemble method for credit scoring. Expert Syst Appl 159:113615","journal-title":"Expert Syst Appl"},{"key":"10697_CR67","doi-asserted-by":"crossref","DOI":"10.1016\/j.knosys.2019.105118","volume":"189","author":"J Xiao","year":"2020","unstructured":"Xiao J, Zhou X, Zhong Y, Xie L, Gu X, Liu D (2020) Cost-sensitive semi-supervised selective ensemble model for customer credit scoring. Knowl Based Syst 189:105118","journal-title":"Knowl Based Syst"},{"key":"10697_CR66","doi-asserted-by":"crossref","first-page":"508","DOI":"10.1016\/j.ins.2021.05.029","volume":"569","author":"J Xiao","year":"2021","unstructured":"Xiao J, Wang Y, Chen J, Xie L, Huang J (2021) Impact of resampling methods and classification models on the imbalanced credit scoring problems. Inf Sci 569:508\u2013526","journal-title":"Inf Sci"},{"issue":"1","key":"10697_CR68","doi-asserted-by":"crossref","first-page":"288","DOI":"10.1016\/j.ijinfomgt.2017.10.002","volume":"38","author":"B Yeo","year":"2018","unstructured":"Yeo B, Grant D (2018) Predicting service industry performance using decision tree analysis. Int J Inf Manag 38(1):288\u2013300","journal-title":"Int J Inf Manag"},{"key":"10697_CR69","volume":"227","author":"G Zhang","year":"2021","unstructured":"Zhang G, Hu W, Cao D, Liu W, Huang R, Huang Q et al (2021) Data-driven optimal energy management for a wind\u2013solar\u2013diesel-battery-reverse osmosis hybrid energy system using a deep reinforcement learning approach. Energy Convers Manag 227:113608","journal-title":"Energy Convers Manag"},{"issue":"4","key":"10697_CR70","doi-asserted-by":"crossref","first-page":"356","DOI":"10.1109\/TCDS.2016.2614675","volume":"9","author":"D Zhao","year":"2016","unstructured":"Zhao D, Chen Y, Lv L (2016) Deep reinforcement learning with visual attention for vehicle classification. IEEE Trans Cogn Dev Syst 9(4):356\u2013367","journal-title":"IEEE Trans Cogn Dev Syst"}],"container-title":["Artificial Intelligence Review"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10462-023-10697-9.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s10462-023-10697-9\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10462-023-10697-9.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2024,4,13]],"date-time":"2024-04-13T03:11:47Z","timestamp":1712977907000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s10462-023-10697-9"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,3,18]]},"references-count":70,"journal-issue":{"issue":"4","published-online":{"date-parts":[[2024,4]]}},"alternative-id":["10697"],"URL":"https:\/\/doi.org\/10.1007\/s10462-023-10697-9","relation":{"has-preprint":[{"id-type":"doi","id":"10.21203\/rs.3.rs-2422835\/v1","asserted-by":"object"}]},"ISSN":["1573-7462"],"issn-type":[{"value":"1573-7462","type":"electronic"}],"subject":[],"published":{"date-parts":[[2024,3,18]]},"assertion":[{"value":"18 March 2024","order":1,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"The authors declare no competing interests.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Conflict of interest"}}],"article-number":"93"}}