{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,1,24]],"date-time":"2026-01-24T22:00:27Z","timestamp":1769292027815,"version":"3.49.0"},"reference-count":33,"publisher":"Wiley","issue":"2","license":[{"start":{"date-parts":[[2026,1,19]],"date-time":"2026-01-19T00:00:00Z","timestamp":1768780800000},"content-version":"vor","delay-in-days":18,"URL":"http:\/\/onlinelibrary.wiley.com\/termsAndConditions#vor"},{"start":{"date-parts":[[2026,1,1]],"date-time":"2026-01-01T00:00:00Z","timestamp":1767225600000},"content-version":"tdm","delay-in-days":0,"URL":"http:\/\/doi.wiley.com\/10.1002\/tdm_license_1.1"}],"funder":[{"DOI":"10.13039\/100007540","name":"Jiangsu Agricultural Science and Technology Innovation Fund","doi-asserted-by":"publisher","award":["CX(22)3104"],"award-info":[{"award-number":["CX(22)3104"]}],"id":[{"id":"10.13039\/100007540","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["61906021"],"award-info":[{"award-number":["61906021"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["onlinelibrary.wiley.com"],"crossmark-restriction":true},"short-container-title":["Concurrency and Computation"],"published-print":{"date-parts":[[2026,1]]},"abstract":"<jats:title>ABSTRACT<\/jats:title>\n                  <jats:p>\n                    In reinforcement learning (RL), the assumption of fixed control frequency often leads to computational resource wastage and degraded policy performance, while traditional single\u2010step temporal difference (TD) learning suffers from accumulated state\u2010value estimation bias. This paper proposes the multi\u2010state soft elastic actor\u2010critic (MSSEAC) algorithm to address these issues: First, the paper introduces a temporal consumption penalty mechanism and reconstructs the actor network's dual\u2010branch output structure to simultaneously generate control actions and time consumption estimates, enabling autonomous control frequency adjustment. Second, the multi\u2010state temporal difference (MSTD) framework is developed to address the limitations of conventional single\u2010step TD learning. Specifically, an innovative experience replay buffer management strategy is proposed, where historical actions are utilized to stabilize the learning process during initial training phases, with a gradual transition to policy\u2010generated actions in later stages to enhance estimation accuracy. The multi\u2010state\u2010value estimation effectively mitigates the bias accumulation problem inherent in single\u2010step TD methods through weighted fusion of return distributions from multiple future states. Code is available at:\n                    <jats:ext-link xmlns:xlink=\"http:\/\/www.w3.org\/1999\/xlink\" xlink:href=\"https:\/\/github.com\/asdwqqqq\/MSSEAC.git\">https:\/\/github.com\/asdwqqqq\/MSSEAC.git<\/jats:ext-link>\n                    .\n                  <\/jats:p>","DOI":"10.1002\/cpe.70555","type":"journal-article","created":{"date-parts":[[2026,1,20]],"date-time":"2026-01-20T02:56:37Z","timestamp":1768877797000},"update-policy":"https:\/\/doi.org\/10.1002\/crossmark_policy","source":"Crossref","is-referenced-by-count":0,"title":["MSSEAC: Multi\u2010State Soft Elastic Actor\u2010Critic"],"prefix":"10.1002","volume":"38","author":[{"given":"Yuwan","family":"Gu","sequence":"first","affiliation":[{"name":"Department of Computer Science and Artificial Intelligence Changzhou University  Changzhou China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Jie","family":"Hao","sequence":"additional","affiliation":[{"name":"Department of Computer Science and Artificial Intelligence Changzhou University  Changzhou China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Fang","family":"Meng","sequence":"additional","affiliation":[{"name":"Department of Computer Science and Artificial Intelligence Changzhou University  Changzhou China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yan","family":"Chen","sequence":"additional","affiliation":[{"name":"Department of Computer Science and Artificial Intelligence Changzhou University  Changzhou China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Ronghai","family":"Miao","sequence":"additional","affiliation":[{"name":"Department of Computer Science and Artificial Intelligence Changzhou University  Changzhou China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Jidong","family":"Lv","sequence":"additional","affiliation":[{"name":"Department of Computer Science and Artificial Intelligence Changzhou University  Changzhou China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"311","published-online":{"date-parts":[[2026,1,19]]},"reference":[{"key":"e_1_2_12_2_1","first-page":"28430","article-title":"RL\u2010GPT: Integrating Reinforcement Learning and Code\u2010As\u2010Policy","volume":"37","author":"Liu S.","year":"2024","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_2_12_3_1","unstructured":"Y.Wang Y.\u2010C.Zhao Z.Wang J.Zhu andZ.Liu \u201cRL\u2010VLM\u2010F: Reinforcement Learning From Vision Language Foundation Model Feedback \u201d(2024). arXiv preprint arXiv:2402.03681."},{"key":"e_1_2_12_4_1","volume-title":"Proceedings of the Twelfth International Conference on Learning Representations (ICLR)","author":"Yuan H.","year":"2024"},{"key":"e_1_2_12_5_1","doi-asserted-by":"publisher","DOI":"10.1038\/s41586-024-07024-9"},{"key":"e_1_2_12_6_1","unstructured":"S.Mahankali L.Wang Y.Reddy andA.Rajeswaran \u201cRandom Latent Exploration for Deep Reinforcement Learning \u201d(2024). arXiv preprint arXiv:2407.13755."},{"key":"e_1_2_12_7_1","first-page":"9976","volume-title":"International Conference on Machine Learning","author":"Feng S.","year":"2023"},{"key":"e_1_2_12_8_1","doi-asserted-by":"publisher","DOI":"10.1038\/s41586-023-06004-9"},{"key":"e_1_2_12_9_1","unstructured":"N.Rudin \u201cGeneralizing Agile Legged Locomotion With Deep Reinforcement Learning \u201d(Doctoral Dissertation ETH Zurich) (2025)."},{"key":"e_1_2_12_10_1","doi-asserted-by":"publisher","DOI":"10.52202\/079017-4228"},{"key":"e_1_2_12_11_1","doi-asserted-by":"publisher","DOI":"10.52202\/079017-3718"},{"key":"e_1_2_12_12_1","first-page":"14386","article-title":"Recurrent Reinforcement Learning With Memoroids","volume":"37","author":"Morad S.","year":"2024","journal-title":"Advances in Neural Information Processing Systems"},{"issue":"200","key":"e_1_2_12_13_1","first-page":"1","article-title":"Distributionally Robust Model\u2010Based Offline Reinforcement Learning With Near\u2010Optimal Sample Complexity","volume":"25","author":"Shi L.","year":"2024","journal-title":"Journal of Machine Learning Research"},{"key":"e_1_2_12_14_1","doi-asserted-by":"publisher","DOI":"10.1038\/s41586-021-04301-9"},{"key":"e_1_2_12_15_1","unstructured":"H.Fu H.Tang andJ.Hao \u201cEfficient Meta Reinforcement Learning via Meta Goal Generation \u201d(2019). arXiv preprint arXiv:1905.07624."},{"key":"e_1_2_12_16_1","doi-asserted-by":"crossref","unstructured":"A.Kumar Z.Fu D.Pathak andJ.Malik \u201cRMA: Rapid Motor Adaptation for Legged Robots \u201d(2021). arXiv preprint arXiv:2107.04034.","DOI":"10.15607\/RSS.2021.XVII.011"},{"key":"e_1_2_12_17_1","unstructured":"P.Brunzema A.vonRohr S.Trimpe andF.Solowjow \u201cEvent\u2010Triggered Time\u2010Varying Bayesian Optimization \u201d(2022). arXiv preprint arXiv:2208.10790."},{"key":"e_1_2_12_18_1","doi-asserted-by":"publisher","DOI":"10.1109\/TNNLS.2021.3121546"},{"key":"e_1_2_12_19_1","doi-asserted-by":"publisher","DOI":"10.1109\/SMC.2017.8122622"},{"key":"e_1_2_12_20_1","doi-asserted-by":"publisher","DOI":"10.1109\/CDC.2012.6425820"},{"key":"e_1_2_12_21_1","first-page":"118953","article-title":"Multi\u2010Turn Reinforcement Learning With Preference Human Feedback","volume":"37","author":"Shani L.","year":"2024","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_2_12_22_1","unstructured":"D.QiaoandY.\u2010X.Wang \u201cNear\u2010Optimal Reinforcement Learning With Self\u2010Play Under Adaptivity Constraints \u201d(2024). arXiv preprint arXiv:2402.01111."},{"key":"e_1_2_12_23_1","unstructured":"B. S.Pavse J. P.Hanna andC.Dann \u201cLearning to Stabilize Online Reinforcement Learning in Unbounded State Spaces \u201d(2023). arXiv preprint arXiv:2306.01896."},{"key":"e_1_2_12_24_1","first-page":"25226","volume-title":"Proceedings of the 41st International Conference on Machine Learning","author":"Koppel A.","year":"2024"},{"key":"e_1_2_12_25_1","first-page":"471","volume-title":"Proceedings of the 41st International Conference on Machine Learning","author":"Agrawal S.","year":"2024"},{"key":"e_1_2_12_26_1","doi-asserted-by":"publisher","DOI":"10.52202\/079017-1610"},{"key":"e_1_2_12_27_1","first-page":"106440","article-title":"Normalization and Effective Learning Rates in Reinforcement Learning","volume":"37","author":"Lyle C.","year":"2024","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_2_12_28_1","unstructured":"J.Obando\u2010Ceron A.Courville andP. S.Castro \u201cIn Value\u2010Based Deep Reinforcement Learning a Pruned Network Is a Good Network \u201d(2024). arXiv preprint arXiv:2402.12479."},{"key":"e_1_2_12_29_1","unstructured":"E.Korkmaz \u201cUnderstanding and Diagnosing Deep Reinforcement Learning \u201d(2024). arXiv preprint arXiv:2406.16979."},{"key":"e_1_2_12_30_1","volume-title":"Proceedings of the Thirteenth International Conference on Learning Representations (ICLR)","author":"Liu S.","year":"2025"},{"key":"e_1_2_12_31_1","unstructured":"J.Liu Y.Chai andP.Li \u201cNeuroplastic Expansion in Deep Reinforcement Learning \u201d(2024). arXiv preprint arXiv:2410.07994."},{"key":"e_1_2_12_32_1","first-page":"2528","article-title":"Reinforcement Learning With Adaptive Regularization for Safe Control of Critical Systems","volume":"37","author":"Tian H.","year":"2024","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_2_12_33_1","doi-asserted-by":"publisher","DOI":"10.1145\/3624720"},{"key":"e_1_2_12_34_1","unstructured":"R. C.Castanyer K.Tuyls andJ.Foerster \u201cStable Gradients for Stable Learning at Scale in Deep Reinforcement Learning \u201d(2025). arXiv preprint arXiv:2506.15544."}],"container-title":["Concurrency and Computation: Practice and Experience"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/onlinelibrary.wiley.com\/doi\/pdf\/10.1002\/cpe.70555","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/onlinelibrary.wiley.com\/doi\/full-xml\/10.1002\/cpe.70555","content-type":"application\/xml","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/onlinelibrary.wiley.com\/doi\/pdf\/10.1002\/cpe.70555","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,1,23]],"date-time":"2026-01-23T12:30:40Z","timestamp":1769171440000},"score":1,"resource":{"primary":{"URL":"https:\/\/onlinelibrary.wiley.com\/doi\/10.1002\/cpe.70555"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,1]]},"references-count":33,"journal-issue":{"issue":"2","published-print":{"date-parts":[[2026,1]]}},"alternative-id":["10.1002\/cpe.70555"],"URL":"https:\/\/doi.org\/10.1002\/cpe.70555","archive":["Portico"],"relation":{},"ISSN":["1532-0626","1532-0634"],"issn-type":[{"value":"1532-0626","type":"print"},{"value":"1532-0634","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,1]]},"assertion":[{"value":"2025-10-16","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-12-29","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2026-01-19","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}],"article-number":"e70555"}}