{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,4,7]],"date-time":"2026-04-07T09:47:10Z","timestamp":1775555230193,"version":"3.50.1"},"reference-count":41,"publisher":"Springer Science and Business Media LLC","issue":"1","license":[{"start":{"date-parts":[[2026,2,2]],"date-time":"2026-02-02T00:00:00Z","timestamp":1769990400000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2026,2,17]],"date-time":"2026-02-17T00:00:00Z","timestamp":1771286400000},"content-version":"vor","delay-in-days":15,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["62462031"],"award-info":[{"award-number":["62462031"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100004479","name":"Natural Science Foundation of Jiangxi Province","doi-asserted-by":"publisher","award":["20242BAB26023"],"award-info":[{"award-number":["20242BAB26023"]}],"id":[{"id":"10.13039\/501100004479","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Hum-Cent Intell Syst"],"abstract":"<jats:title>Abstract<\/jats:title>\n                  <jats:p>In offline reinforcement learning, model-based approaches have demonstrated superior data efficiency by leveraging learned dynamics models to generate additional training samples. However, due to inevitable model inaccuracies, directly deriving policies from such models often leads to suboptimal performance under the constraints of the offline setting. Prior work has attempted to mitigate this issue by adopting conservative strategies that avoid reliance on out-of-distribution transitions. Nevertheless, these methods still face notable challenges, as dynamics models trained solely on historical data typically struggle to generalize to unseen state-action pairs. In this paper, we propose a novel offline reinforcement learning method Dynamic Reward-Guided Multi-Head Attention for Actor-Critic Policy Learning Optimization (DRMAAC). DRMAAC introduces a dynamic-aware paradigm that focuses on capturing the intrinsic characteristics of the behavior policy. It leverages inverse reinforcement learning to recover a reward-consistent dynamics model and identify high-return states. Meanwhile, an Actor-Critic architecture enhanced with multi-head attention makes decisions guided by these high-value states. This integration enables the model to better capture long-term dependencies and prioritize informative features in complex state spaces. Empirical evaluations on the D4RL benchmark show that DRMAAC consistently outperforms previous state-of-the-art methods across a variety of tasks. These results highlight not only improved data efficiency but also strong generalization capabilities under diverse environmental conditions. Overall, DRMAAC presents a promising direction for advancing model-based offline reinforcement learning by combining attention mechanisms with reward-consistent dynamics modeling.<\/jats:p>","DOI":"10.1007\/s44230-026-00135-8","type":"journal-article","created":{"date-parts":[[2026,2,2]],"date-time":"2026-02-02T15:43:29Z","timestamp":1770047009000},"page":"77-92","update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":0,"title":["Dynamic Reward-Guided with Multi-Head Attention for Actor-Critic Policy Learning Optimization"],"prefix":"10.1007","volume":"6","author":[{"given":"Xiaohui","family":"Huang","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Junhang","family":"Zong","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Xiaofei","family":"Yang","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Ying","family":"Yu","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"297","published-online":{"date-parts":[[2026,2,2]]},"reference":[{"issue":"2","key":"135_CR1","doi-asserted-by":"publisher","DOI":"10.1007\/s11432-022-3696-5","volume":"67","author":"F-M Luo","year":"2024","unstructured":"Luo F-M, Xu T, Lai H, Chen X-H, Zhang W, Yu Y. A survey on model-based reinforcement learning. Sci Chin Inf Sci. 2024;67(2):121101.","journal-title":"Sci Chin Inf Sci"},{"issue":"1","key":"135_CR2","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1561\/2200000086","volume":"16","author":"TM Moerland","year":"2023","unstructured":"Moerland TM, Broekens J, Plaat A, Jonker CM. Model-based reinforcement learning: A survey. Found Trends Mach Learn. 2023;16(1):1\u2013118.","journal-title":"Found Trends Mach Learn"},{"key":"135_CR3","doi-asserted-by":"publisher","DOI":"10.1016\/j.eswa.2019.112963","volume":"141","author":"M Lopez-Martin","year":"2020","unstructured":"Lopez-Martin M, Carro B, Sanchez-Esguevillas A. Application of deep reinforcement learning to intrusion detection for supervised problems. Expert Syst Appl. 2020;141:112963.","journal-title":"Expert Syst Appl"},{"key":"135_CR4","unstructured":"Kumar A, Fu J, Soh M, Tucker G, Levine S. Stabilizing off-policy q-learning via bootstrapping error reduction. Adv Neural Inf Process Syst 2019;32"},{"key":"135_CR5","unstructured":"Levine S, Kumar A, Tucker G, Fu J. Offline reinforcement learning: Tutorial, review, and perspectives on open problems. arXiv preprint arXiv:2005.01643 2020."},{"key":"135_CR6","doi-asserted-by":"crossref","unstructured":"Torabi F, Warnell G, Stone P. Behavioral cloning from observation. In: Proceedings of the Twenty-Seventh International Joint Conference on Artificial Intelligence, 2018;4950\u20134957. IJCAI","DOI":"10.24963\/ijcai.2018\/687"},{"key":"135_CR7","first-page":"1179","volume":"33","author":"A Kumar","year":"2020","unstructured":"Kumar A, Zhou A, Tucker G, Levine S. Conservative q-learning for offline reinforcement learning. Adv Neural Inf Process Syst. 2020;33:1179\u201391.","journal-title":"Adv Neural Inf Process Syst"},{"key":"135_CR8","unstructured":"Kostrikov I, Agrawal KK, Dwibedi D, Levine S, Tompson J. Discriminator-actor-critic: Addressing sample inefficiency and reward bias in adversarial imitation learning. In: 7th International Conference on Learning Representations 2019."},{"key":"135_CR9","first-page":"7436","volume":"34","author":"G An","year":"2021","unstructured":"An G, Moon S, Kim J-H, Song HO. Uncertainty-based offline reinforcement learning with diversified q-ensemble. Adv Neural Inf Process Syst. 2021;34:7436\u201347.","journal-title":"Adv Neural Inf Process Syst"},{"key":"135_CR10","unstructured":"Luo F-M, Cao X, Qin R-J, Yu Y. Transferable reward learning by dynamics-agnostic discriminator ensemble"},{"key":"135_CR11","first-page":"1273","volume":"34","author":"M Janner","year":"2021","unstructured":"Janner M, Li Q, Levine S. Offline reinforcement learning as one big sequence modeling problem. Adv Neural Inf Process Syst. 2021;34:1273\u201386.","journal-title":"Adv Neural Inf Process Syst"},{"key":"135_CR12","unstructured":"Schulman J, Wolski F, Dhariwal P, Radford A, Klimov O. Proximal policy optimization algorithms. arXiv preprint arXiv:1707.06347 2017."},{"key":"135_CR13","unstructured":"Sutton RS, McAllester D, Singh S, Mansour Y. Policy gradient methods for reinforcement learning with function approximation. Advances in neural information processing systems 1999;12"},{"key":"135_CR14","unstructured":"Finn C, Levine S, Abbeel P. Guided cost learning: Deep inverse optimal control via policy optimization. In: International Conference on Machine Learning, 2016;49\u201358. PMLR"},{"key":"135_CR15","first-page":"28954","volume":"34","author":"T Yu","year":"2021","unstructured":"Yu T, Kumar A, Rafailov R, Rajeswaran A, Levine S, Finn C. Combo: Conservative offline model-based policy optimization. Adv Neural Inf Process Syst. 2021;34:28954\u201367.","journal-title":"Adv Neural Inf Process Syst"},{"key":"135_CR16","first-page":"16082","volume":"35","author":"M Rigter","year":"2022","unstructured":"Rigter M, Lacerda B, Hawes N. Rambo-rl: Robust adversarial model-based offline reinforcement learning. Adv Neural Inf Process Syst. 2022;35:16082\u201397.","journal-title":"Adv Neural Inf Process Syst"},{"key":"135_CR17","first-page":"14129","volume":"33","author":"T Yu","year":"2020","unstructured":"Yu T, Thomas G, Yu L, Ermon S, Zou JY, Levine S, et al. Mopo: Model-based offline policy optimization. Adv Neural Inf Process Syst. 2020;33:14129\u201342.","journal-title":"Adv Neural Inf Process Syst"},{"key":"135_CR18","unstructured":"Vaswani A, Shazeer N, Parmar N, Uszkoreit J, Jones L, Gomez AN, Kaiser \u0141, Polosukhin I. Attention is all you need. Adv Neural Inf Process Syst 2017;30"},{"key":"135_CR19","doi-asserted-by":"crossref","unstructured":"Voita E, Talbot D, Moiseev F, Sennrich R, Titov I. Analyzing multi-head self-attention: Specialized heads do the heavy lifting, the rest can be pruned. In: Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, 2019;5797\u20135808. ACL","DOI":"10.18653\/v1\/P19-1580"},{"key":"135_CR20","unstructured":"Cordonnier J-B, Loukas A, Jaggi M. Multi-head attention: Collaborate instead of concatenate. In: International Conference on Learning Representations 2021."},{"key":"135_CR21","unstructured":"Konda V, Tsitsiklis J. Actor-critic algorithms. Adv Neural Inf Process Syst 1999;12"},{"key":"135_CR22","unstructured":"Sun Y, Zhang J, Jia C, Lin H, Ye J, Yu Y. Model-bellman inconsistency for model-based offline reinforcement learning. In: International Conference on Machine Learning, 2023;33177\u201333194. PMLR"},{"key":"135_CR23","first-page":"2914","volume":"33","author":"N Rajaraman","year":"2020","unstructured":"Rajaraman N, Yang L, Jiao J, Ramchandran K. Toward the fundamental limits of imitation learning. Adv Neural Inf Process Syst. 2020;33:2914\u201324.","journal-title":"Adv Neural Inf Process Syst"},{"key":"135_CR24","unstructured":"Fu J, Kumar A, Nachum O, Tucker G, Levine S. D4rl: Datasets for deep data-driven reinforcement learning. In: International Conference on Learning Representations 2021."},{"key":"135_CR25","doi-asserted-by":"publisher","DOI":"10.1016\/j.eswa.2025.126842","volume":"274","author":"X Li","year":"2025","unstructured":"Li X, Ling X. Mild evaluation policy via dataset constraint for offline reinforcement learning. Expert Syst Appl. 2025;274:126842.","journal-title":"Expert Syst Appl"},{"key":"135_CR26","doi-asserted-by":"crossref","unstructured":"Lange S, Gabel T, Riedmiller M. Batch reinforcement learning. In: Reinforcement Learning: State-of-the-art, pp. 45\u201373. Springer, Berlin, Heidelberg 2012.","DOI":"10.1007\/978-3-642-27645-3_2"},{"key":"135_CR27","unstructured":"Xu T, Li Z, Yu Y, Luo Z-Q. Provably efficient adversarial imitation learning with unknown transitions. In: Uncertainty in Artificial Intelligence, 2023;2367\u20132378. PMLR"},{"key":"135_CR28","first-page":"20132","volume":"34","author":"S Fujimoto","year":"2021","unstructured":"Fujimoto S, Gu SS. A minimalist approach to offline reinforcement learning. Adv Neural Inf Process Syst. 2021;34:20132\u201345.","journal-title":"Adv Neural Inf Process Syst"},{"key":"135_CR29","unstructured":"Kostrikov I, Nachum O, Tompson J. Imitation learning via off-policy distribution matching. In: 8th International Conference on Learning Representations 2020."},{"key":"135_CR30","unstructured":"Ghasemipour SKS, Zemel R, Gu S. A divergence minimization perspective on imitation learning methods. In: Conference on Robot Learning, 2020;1259\u20131277. PMLR"},{"key":"135_CR31","unstructured":"Bai C, Lingxiao W, Zhuoran Y, Zhihong D, Animesh G, Peng L, Zhaoran W. Pessimistic bootstrapping for uncertainty-driven offline reinforcement learning. In: International Conference on Learning Representations 2022."},{"key":"135_CR32","unstructured":"Janner M, Fu J, Zhang M, Levine S. When to trust your model: Model-based policy optimization. Adv Neural Inf Process Syst, 2019;32"},{"key":"135_CR33","first-page":"21810","volume":"33","author":"R Kidambi","year":"2020","unstructured":"Kidambi R, Rajeswaran A, Netrapalli P, Joachims T. Morel: Model-based offline reinforcement learning. Adv Neural Inf Process Syst. 2020;33:21810\u201323.","journal-title":"Adv Neural Inf Process Syst"},{"key":"135_CR34","first-page":"8432","volume":"34","author":"X-H Chen","year":"2021","unstructured":"Chen X-H, Yu Y, Li Q, Luo F-M, Qin Z, Shang W, et al. Offline model-based adaptable policy learning. Adv Neural Inf Process Syst. 2021;34:8432\u201343.","journal-title":"Adv Neural Inf Process Syst"},{"key":"135_CR35","unstructured":"Ho J, Ermon S. Generative adversarial imitation learning. Adv Neural Inf Process Syst 2016;29"},{"key":"135_CR36","unstructured":"Finn C, Christiano P, Abbeel P, Levine S. A connection between generative adversarial networks, inverse reinforcement learning, and energy-based models. arXiv preprint arXiv:1611.03852 2016."},{"key":"135_CR37","doi-asserted-by":"publisher","first-page":"70654","DOI":"10.52202\/075280-3097","volume":"36","author":"X-H Chen","year":"2023","unstructured":"Chen X-H, Yu Y, Zhu Z, Yu Z, Zhenjun C, Wang C, et al. Adversarial counterfactual environment model learning. Adv Neural Inf Process Syst. 2023;36:70654\u2013706.","journal-title":"Adv Neural Inf Process Syst"},{"key":"135_CR38","doi-asserted-by":"crossref","unstructured":"Mnih V, Kavukcuoglu K, Silver D, Rusu AA, Veness J, Bellemare MG, Graves A, Riedmiller M, Fidjeland AK, Ostrovski G. Human-level control through deep reinforcement learning. nature 2015;518(7540):529\u2013533.","DOI":"10.1038\/nature14236"},{"issue":"10","key":"135_CR39","doi-asserted-by":"publisher","first-page":"6968","DOI":"10.1109\/TPAMI.2021.3096966","volume":"44","author":"T Xu","year":"2021","unstructured":"Xu T, Li Z, Yu Y. Error bounds of imitating policies and environments for reinforcement learning. IEEE Trans Pattern Anal Mach Intell. 2021;44(10):6968\u201380.","journal-title":"IEEE Trans Pattern Anal Mach Intell"},{"key":"135_CR40","unstructured":"Ziebart BD, Maas AL, Bagnell JA, Dey AK. Maximum entropy inverse reinforcement learning. In: Proceedings of the AAAI Conference on Artificial Intelligence, 2008;8:1433\u20131438. AAAI"},{"key":"135_CR41","unstructured":"Lin T, Jin C, Jordan M. On gradient descent ascent for nonconvex-concave minimax problems. In: International Conference on Machine Learning, 2020;6083\u20136093. PMLR"}],"container-title":["Human-Centric Intelligent Systems"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s44230-026-00135-8","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s44230-026-00135-8.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s44230-026-00135-8.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,4,7]],"date-time":"2026-04-07T08:50:18Z","timestamp":1775551818000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s44230-026-00135-8"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,2,2]]},"references-count":41,"journal-issue":{"issue":"1","published-online":{"date-parts":[[2026,3]]}},"alternative-id":["135"],"URL":"https:\/\/doi.org\/10.1007\/s44230-026-00135-8","relation":{},"ISSN":["2667-1336"],"issn-type":[{"value":"2667-1336","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,2,2]]},"assertion":[{"value":"3 September 2025","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"13 January 2026","order":2,"name":"revised","label":"Revised","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"15 January 2026","order":3,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"2 February 2026","order":4,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"The authors of this work have no known competing financial interests or personal relationships that could have appeared to influence the work reported in this paper.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Conflict of interest"}},{"value":"Not applicable.","order":3,"name":"Ethics","group":{"name":"EthicsHeading","label":"Ethics approval and consent to participate"}},{"value":"Not applicable.","order":4,"name":"Ethics","group":{"name":"EthicsHeading","label":"Consent for publication"}}]}}