{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,3,4]],"date-time":"2026-03-04T17:17:04Z","timestamp":1772644624226,"version":"3.50.1"},"reference-count":35,"publisher":"Springer Science and Business Media LLC","issue":"7","license":[{"start":{"date-parts":[[2024,4,1]],"date-time":"2024-04-01T00:00:00Z","timestamp":1711929600000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2024,4,22]],"date-time":"2024-04-22T00:00:00Z","timestamp":1713744000000},"content-version":"vor","delay-in-days":21,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["62071240"],"award-info":[{"award-number":["62071240"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100012246","name":"Priority Academic Program Development of Jiangsu Higher Education Institutions","doi-asserted-by":"publisher","id":[{"id":"10.13039\/501100012246","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100004608","name":"Natural Science Foundation of Jiangsu Province","doi-asserted-by":"publisher","award":["BK20220804, BK20231142"],"award-info":[{"award-number":["BK20220804, BK20231142"]}],"id":[{"id":"10.13039\/501100004608","id-type":"DOI","asserted-by":"publisher"}]},{"name":"Innovation Program for Quantum Science and Technology","award":["2021ZD0302901"],"award-info":[{"award-number":["2021ZD0302901"]}]}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Appl Intell"],"published-print":{"date-parts":[[2024,4]]},"abstract":"<jats:sec>\n                <jats:title>Abstract<\/jats:title>\n                <jats:p>Reinforcement learning is widely used in financial markets to assist investors in developing trading strategies. However, most existing models primarily focus on simple volume-price factors, and there is a need for further improvement in the returns of stock trading. To address these challenges, a multi-factor stock trading strategy based on Deep Q-Network (DQN) with Multi-layer Bidirectional Gated Recurrent Unit (Multi-BiGRU) and multi-head ProbSparse self-attention is proposed. Our strategy comprehensively characterizes the determinants of stock prices by considering various factors such as financial quality, valuation, and sentiment factors. We first use Light Gradient Boosting Machine (LightGBM) to classify turning points for stock data. Then, in the reinforcement learning strategy, Multi-BiGRU, which holds the bidirectional learning of historical data, is integrated into DQN, aiming to enhance the model\u2019s ability to understand the dynamics of the stock market. Moreover, the multi-head ProbSparse self-attention mechanism effectively captures interactions between different factors, providing the model with deeper market insights. We validate our strategy\u2019s effectiveness through extensive experimental research on stocks from Chinese and US markets. The results show that our method outperforms both temporal and non-temporal models in terms of stock trading returns. Ablation studies confirm the critical role of LightGBM and multi-head ProbSparse self-attention mechanism. The experiment section also demonstrates the significant advantages of our model through the presentation of box plots and statistical tests. Overall, by fully considering the multi-factor data and the model\u2019s feature extraction capabilities, our work is expected to provide investors with more precise trading decision support.<\/jats:p>\n              <\/jats:sec><jats:sec>\n                <jats:title>Graphical abstract<\/jats:title>\n                \n              <\/jats:sec>","DOI":"10.1007\/s10489-024-05463-5","type":"journal-article","created":{"date-parts":[[2024,4,22]],"date-time":"2024-04-22T12:02:13Z","timestamp":1713787333000},"page":"5417-5440","update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":10,"title":["Multi-factor stock trading strategy based on DQN with multi-BiGRU and multi-head ProbSparse self-attention"],"prefix":"10.1007","volume":"54","author":[{"given":"Wenjie","family":"Liu","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-7267-3541","authenticated-orcid":false,"given":"Yuchen","family":"Gu","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yebo","family":"Ge","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"297","published-online":{"date-parts":[[2024,4,22]]},"reference":[{"key":"5463_CR1","doi-asserted-by":"publisher","unstructured":"Almahdi S, Yang SY (2017) An adaptive portfolio trading system: A risk-return portfolio optimization using recurrent reinforcement learning with expected maximum drawdown. Expert Syst Appl 87:267\u2013279. https:\/\/doi.org\/10.1016\/j.eswa.2017.06.023","DOI":"10.1016\/j.eswa.2017.06.023"},{"key":"5463_CR2","doi-asserted-by":"publisher","unstructured":"Aseeri AO (2023) Effective short-term forecasts of saudi stock price trends using technical indicators and large-scale multivariate time series. Peerj Comput Sci 9:e1205. https:\/\/doi.org\/10.7717\/peerj-cs.1205","DOI":"10.7717\/peerj-cs.1205"},{"issue":"1","key":"5463_CR3","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1016\/j.jocs.2010.12.007","volume":"2","author":"J Bollen","year":"2011","unstructured":"Bollen J, Mao H, Zeng X (2011) Twitter mood predicts the stock market. J Comput Sci 2(1):1\u20138. https:\/\/doi.org\/10.1016\/j.jocs.2010.12.007","journal-title":"J Comput Sci"},{"key":"5463_CR4","doi-asserted-by":"publisher","unstructured":"Chakole JB, Kolhe MS, Mahapurush GD et al (2021) A q-learning agent for automated trading in equity stock markets. Expert Syst Appl 163:113761. https:\/\/doi.org\/10.1016\/j.eswa.2020.113761","DOI":"10.1016\/j.eswa.2020.113761"},{"key":"5463_CR5","doi-asserted-by":"publisher","unstructured":"Cho K, van Merri\u00ebnboer B, Gulcehre C et\u00a0al (2014) Learning phrase representations using RNN encoder\u2013decoder for statistical machine translation. In: Proceedings of the 2014 conference on empirical methods in natural language processing (EMNLP), pp 1724\u20131734. https:\/\/doi.org\/10.3115\/v1\/D14-1179","DOI":"10.3115\/v1\/D14-1179"},{"key":"5463_CR6","doi-asserted-by":"publisher","unstructured":"Cui C, Wang P, Li Y et al (2023) Mcvcsb: A new hybrid deep learning network for stock index prediction. Expert Syst Appl 232. https:\/\/doi.org\/10.1016\/j.eswa.2023.120902","DOI":"10.1016\/j.eswa.2023.120902"},{"key":"5463_CR7","doi-asserted-by":"publisher","unstructured":"Deng C, Huang Y, Hasan N et al (2022) Multi-step-ahead stock price index forecasting using long short-term memory model with multivariate empirical mode decomposition. Inf Sci 607:297\u2013321. https:\/\/doi.org\/10.1016\/j.ins.2022.05.088","DOI":"10.1016\/j.ins.2022.05.088"},{"issue":"10","key":"5463_CR8","doi-asserted-by":"publisher","first-page":"7177","DOI":"10.1007\/s10489-021-02249-x","volume":"51","author":"D Fister","year":"2021","unstructured":"Fister D, Perc M, Jagric T (2021) Two robust long short-term memory frameworks for trading stocks. Appl Intell 51(10):7177\u20137195. https:\/\/doi.org\/10.1007\/s10489-021-02249-x","journal-title":"Appl Intell"},{"issue":"3","key":"5463_CR9","doi-asserted-by":"publisher","first-page":"685","DOI":"10.1016\/j.dss.2013.02.006","volume":"55","author":"M Hagenau","year":"2013","unstructured":"Hagenau M, Liebmann M, Neumann D (2013) Automated news reading: Stock price prediction based on financial news using context-capturing features. Decision Support Syst 55(3):685\u2013697. https:\/\/doi.org\/10.1016\/j.dss.2013.02.006","journal-title":"Decision Support Syst"},{"key":"5463_CR10","doi-asserted-by":"publisher","unstructured":"Han H, Xie L, Chen S et al (2023) Stock trend prediction based on industry relationships driven hypergraph attention networks. Appl Intell. https:\/\/doi.org\/10.1007\/s10489-023-05035-z","DOI":"10.1007\/s10489-023-05035-z"},{"key":"5463_CR11","doi-asserted-by":"publisher","unstructured":"Huang Z, Gong W, Duan J (2023) Tbdqn: A novel two-branch deep q-network for crude oil and natural gas futures trading. Appl Energy 347. https:\/\/doi.org\/10.1016\/j.apenergy.2023.121321","DOI":"10.1016\/j.apenergy.2023.121321"},{"key":"5463_CR12","doi-asserted-by":"publisher","unstructured":"Huang Z, Li N, Mei W et al (2023) Algorithmic trading using combinational rule vector and deep reinforcement learning. Appl Soft Comput 147. https:\/\/doi.org\/10.1016\/j.asoc.2023.110802","DOI":"10.1016\/j.asoc.2023.110802"},{"key":"5463_CR13","doi-asserted-by":"publisher","unstructured":"Lei K, Zhang B, Li Y et al (2020) Time-driven feature-aware jointly deep reinforcement learning for financial signal representation and algorithmic trading. Expert Syst Appl 140:112872. https:\/\/doi.org\/10.1016\/j.eswa.2019.112872","DOI":"10.1016\/j.eswa.2019.112872"},{"key":"5463_CR14","doi-asserted-by":"publisher","unstructured":"Li Y, Ni P, Chang V (2020) Application of deep reinforcement learning in stock trading strategies and stock forecasting. Computing 102(6, SI):1305\u20131322. https:\/\/doi.org\/10.1007\/s00607-019-00773-w","DOI":"10.1007\/s00607-019-00773-w"},{"issue":"4","key":"5463_CR15","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1145\/3488378","volume":"16","author":"C Liu","year":"2022","unstructured":"Liu C, Yan J, Guo F et al (2022) Forecasting the market with machine learning algorithms: An application of NMC-BERT-LSTM-DQN-X algorithm in quantitative trading. ACM Trans Knowl Disc Data 16(4):1\u201322. https:\/\/doi.org\/10.1145\/3488378","journal-title":"ACM Trans Knowl Disc Data"},{"issue":"2","key":"5463_CR16","doi-asserted-by":"publisher","first-page":"1683","DOI":"10.1007\/s10489-022-03321-w","volume":"53","author":"P Liu","year":"2023","unstructured":"Liu P, Zhang Y, Bao F et al (2023) Multi-type data fusion framework based on deep reinforcement learning for algorithmic trading. Appl Intell 53(2):1683\u20131706. https:\/\/doi.org\/10.1007\/s10489-022-03321-w","journal-title":"Appl Intell"},{"key":"5463_CR17","doi-asserted-by":"publisher","unstructured":"Liu W, Ge Y, Gu Y (2024) Multi-factor stock price prediction based on gan-trellisnet. Knowl Inf Syst. https:\/\/doi.org\/10.1007\/s10115-024-02085-8","DOI":"10.1007\/s10115-024-02085-8"},{"key":"5463_CR18","doi-asserted-by":"publisher","unstructured":"Ma C, Zhang J, Liu J et al (2021) A parallel multi-module deep reinforcement learning algorithm for stock trading. Neurocomputing 449:290\u2013302. https:\/\/doi.org\/10.1016\/j.neucom.2021.04.005","DOI":"10.1016\/j.neucom.2021.04.005"},{"key":"5463_CR19","doi-asserted-by":"publisher","unstructured":"Ma G, Chen P, Liu Z et al (2022) The prediction of enterprise stock change trend by deep neural network model. Comput Intell Neurosci 2022:9. https:\/\/doi.org\/10.1155\/2022\/9193055","DOI":"10.1155\/2022\/9193055"},{"key":"5463_CR20","unstructured":"Meng Q (2017) Lightgbm: A highly efficient gradient boosting decision tree. In: Neural information processing systems, pp 3149\u20133157"},{"key":"5463_CR21","unstructured":"Mnih V, Kavukcuoglu K, Silver D et\u00a0al (2013) Playing atari with deep reinforcement learning. arXiv:1312.5602"},{"issue":"5\u20136","key":"5463_CR22","doi-asserted-by":"publisher","first-page":"1550","DOI":"10.1016\/j.engappai.2013.01.009","volume":"26","author":"K Park","year":"2013","unstructured":"Park K, Shin H (2013) Stock price prediction based on a complex interrelation network of economic factors. Eng Appl Artif Intell 26(5\u20136):1550\u20131561. https:\/\/doi.org\/10.1016\/j.engappai.2013.01.009","journal-title":"Eng Appl Artif Intell"},{"key":"5463_CR23","doi-asserted-by":"publisher","unstructured":"Shi Y, Li W, Zhu L et al (2021) Stock trading rule discovery with double deep q-network. Appl Soft Comput 107:107320. https:\/\/doi.org\/10.1016\/j.asoc.2021.107320","DOI":"10.1016\/j.asoc.2021.107320"},{"key":"5463_CR24","doi-asserted-by":"publisher","unstructured":"Soleymani F, Paquet E (2020) Financial portfolio optimization with online deep reinforcement learning and restricted stacked autoencoder-deepbreath. Expert Syst Appl 156. https:\/\/doi.org\/10.1016\/j.eswa.2020.113456","DOI":"10.1016\/j.eswa.2020.113456"},{"key":"5463_CR25","doi-asserted-by":"publisher","unstructured":"Staffini A (2022) Stock price forecasting by a deep convolutional generative adversarial network. Front Artif Intell 5. https:\/\/doi.org\/10.3389\/frai.2022.837596","DOI":"10.3389\/frai.2022.837596"},{"key":"5463_CR26","doi-asserted-by":"publisher","unstructured":"Taghian M, Asadi A, Safabakhsh R (2022) Learning financial asset-specific trading rules via deep reinforcement learning. Expert Syst Appl 195. https:\/\/doi.org\/10.1016\/j.eswa.2022.116523","DOI":"10.1016\/j.eswa.2022.116523"},{"key":"5463_CR27","doi-asserted-by":"publisher","unstructured":"Takara LdA, Santos AAP, Mariani VC et\u00a0al (2024) Deep reinforcement learning applied to a sparse-reward trading environment with intraday data. Expert Syst Appl 238(C). https:\/\/doi.org\/10.1016\/j.eswa.2023.121897","DOI":"10.1016\/j.eswa.2023.121897"},{"issue":"1","key":"5463_CR28","doi-asserted-by":"publisher","first-page":"126","DOI":"10.1186\/s40537-021-00512-z","volume":"8","author":"Y Touzani","year":"2021","unstructured":"Touzani Y, Douzi K (2021) An LSTM and GRU based trading strategy adapted to the Moroccan market. J Big Data 8(1):126. https:\/\/doi.org\/10.1186\/s40537-021-00512-z","journal-title":"J Big Data"},{"key":"5463_CR29","doi-asserted-by":"publisher","unstructured":"Wang J, Jing F, He M (2023) Stock trading strategy of reinforcement learning driven by turning point classification. Neural Process Lett 55(3, SI):3489\u20133508. https:\/\/doi.org\/10.1007\/s11063-022-11019-w","DOI":"10.1007\/s11063-022-11019-w"},{"key":"5463_CR30","unstructured":"Watkins CJCH (1989) Learning from delayed rewards. PhD thesis, Cambridge University"},{"issue":"11","key":"5463_CR31","doi-asserted-by":"publisher","first-page":"8119","DOI":"10.1007\/s10489-021-02262-0","volume":"51","author":"ME Wu","year":"2021","unstructured":"Wu ME, Syu JH, Lin JCW et al (2021) Portfolio management system in equity market neutral using reinforcement learning. Appl Intell 51(11):8119\u20138131. https:\/\/doi.org\/10.1007\/s10489-021-02262-0","journal-title":"Appl Intell"},{"key":"5463_CR32","doi-asserted-by":"publisher","unstructured":"Wu X, Chen H, Wang J et al (2020) Adaptive stock trading strategies with deep reinforcement learning methods. Inf Sci 538:142\u2013158. https:\/\/doi.org\/10.1016\/j.ins.2020.05.066","DOI":"10.1016\/j.ins.2020.05.066"},{"key":"5463_CR33","doi-asserted-by":"publisher","unstructured":"Yang Z, Zhao T, Wang S et\u00a0al (2024) Mdf-dmc: A stock prediction model combining multi-view stock data features with dynamic market correlation information. Expert Syst Appl 238(E). https:\/\/doi.org\/10.1016\/j.eswa.2023.122134","DOI":"10.1016\/j.eswa.2023.122134"},{"key":"5463_CR34","doi-asserted-by":"publisher","unstructured":"Yu X, Li D (2021) Important trading point prediction using a hybrid convolutional recurrent neural network. Appl Sci-Basel 11(9):3984. https:\/\/doi.org\/10.3390\/app11093984","DOI":"10.3390\/app11093984"},{"key":"5463_CR35","doi-asserted-by":"crossref","unstructured":"Zhou H, Zhang S, Peng J et\u00a0al (2021) Informer: Beyond efficient transformer for long sequence time-series forecasting. In: Proceedings of AAAI, pp 11106\u201311115","DOI":"10.1609\/aaai.v35i12.17325"}],"container-title":["Applied Intelligence"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10489-024-05463-5.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s10489-024-05463-5\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10489-024-05463-5.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2024,5,13]],"date-time":"2024-05-13T14:13:47Z","timestamp":1715609627000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s10489-024-05463-5"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,4]]},"references-count":35,"journal-issue":{"issue":"7","published-print":{"date-parts":[[2024,4]]}},"alternative-id":["5463"],"URL":"https:\/\/doi.org\/10.1007\/s10489-024-05463-5","relation":{},"ISSN":["0924-669X","1573-7497"],"issn-type":[{"value":"0924-669X","type":"print"},{"value":"1573-7497","type":"electronic"}],"subject":[],"published":{"date-parts":[[2024,4]]},"assertion":[{"value":"11 April 2024","order":1,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"22 April 2024","order":2,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"The authors declare that they have no known competing financial interests or personal relationships that could have appeared to influence the work reported in this paper.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Competing interest"}},{"value":"This article does not contain any studies with human participants or animals performed by any of the authors.","order":3,"name":"Ethics","group":{"name":"EthicsHeading","label":"Ethical and informed consent for data used"}}]}}