{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,3]],"date-time":"2026-06-03T12:26:24Z","timestamp":1780489584396,"version":"3.54.1"},"reference-count":30,"publisher":"Springer Science and Business Media LLC","issue":"1","license":[{"start":{"date-parts":[[2022,5,27]],"date-time":"2022-05-27T00:00:00Z","timestamp":1653609600000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2022,5,27]],"date-time":"2022-05-27T00:00:00Z","timestamp":1653609600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"DOI":"10.13039\/501100001691","name":"Japan Society for the Promotion of Science","doi-asserted-by":"publisher","award":["20H04245"],"award-info":[{"award-number":["20H04245"]}],"id":[{"id":"10.13039\/501100001691","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100001691","name":"Japan Society for the Promotion of Science","doi-asserted-by":"publisher","award":["17KT0044"],"award-info":[{"award-number":["17KT0044"]}],"id":[{"id":"10.13039\/501100001691","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Auton. Intell. Syst."],"published-print":{"date-parts":[[2022,12]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:p>We propose a two-stage reward allocation method with decay using an extension of replay memory to adapt this rewarding method for deep reinforcement learning (DRL), to generate coordinated behaviors for tasks that can be completed by executing a few subtasks sequentially by heterogeneous agents. An independent learner in cooperative multi-agent systems needs to learn its policies for effective execution of its own responsible subtask, as well as for coordinated behaviors under a certain coordination structure. Although the reward scheme is an issue for DRL, it is difficult to design it to learn both policies. Our proposed method attempts to generate these different behaviors in multi-agent DRL by dividing the timing of rewards into two stages and varying the ratio between them over time. By introducing the coordinated delivery and execution problem with an expiration time, where a task can be executed sequentially by two heterogeneous agents, we experimentally analyze the effect of using various ratios of the reward division in the two-stage allocations on the generated behaviors. The results demonstrate that the proposed method could improve the overall performance relative to those with the conventional one-time or fixed reward and can establish robust coordinated behavior.<\/jats:p>","DOI":"10.1007\/s43684-022-00029-z","type":"journal-article","created":{"date-parts":[[2022,5,27]],"date-time":"2022-05-27T09:03:26Z","timestamp":1653642206000},"update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":1,"title":["Two-stage reward allocation with decay for multi-agent coordinated behavior for sequential cooperative task by using deep reinforcement learning"],"prefix":"10.1007","volume":"2","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-1676-9346","authenticated-orcid":false,"given":"Yuki","family":"Miyashita","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Toshiharu","family":"Sugawara","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"297","published-online":{"date-parts":[[2022,5,27]]},"reference":[{"issue":"2","key":"29_CR1","doi-asserted-by":"publisher","first-page":"257","DOI":"10.1162\/evco.2008.16.2.257","volume":"16","author":"A. Agogino","year":"2008","unstructured":"A. Agogino, K. Tumer, Efficient evaluation functions for evolving coordination. Evol. Comput. 16(2), 257\u2013288 (2008)","journal-title":"Evol. Comput."},{"key":"29_CR2","series-title":"AAMAS \u201904","first-page":"980","volume-title":"Proceedings of the Third International Joint Conference on Autonomous Agents and Multiagent Systems \u2013 Volume 2","author":"A.K. Agogino","year":"2004","unstructured":"A.K. Agogino, K. Tumer, Unifying temporal and structural credit assignment problems, in Proceedings of the Third International Joint Conference on Autonomous Agents and Multiagent Systems \u2013 Volume 2. AAMAS \u201904 (IEEE Comput. Soc., Los Alamitos, 2004), pp. 980\u2013987"},{"key":"29_CR3","first-page":"2669","volume-title":"AAAI-18: 32nd AAAI Conference on Artificial Intelligence","author":"M. Alshiekh","year":"2018","unstructured":"M. Alshiekh, R. Bloem, R. Ehlers et al., Safe reinforcement learning via shielding, in AAAI-18: 32nd AAAI Conference on Artificial Intelligence (2018), pp. 2669\u20132678"},{"key":"29_CR4","first-page":"196","volume-title":"Proceedings of the 20th International Conference on Autonomous Agents and MultiAgent Systems","author":"R. Beal","year":"2021","unstructured":"R. Beal, G. Chalkiadakis, T.J. Norman et al., Optimising long-term outcomes using real-world fluent objectives: an application to football, in Proceedings of the 20th International Conference on Autonomous Agents and MultiAgent Systems (2021), pp. 196\u2013204"},{"key":"29_CR5","doi-asserted-by":"publisher","first-page":"423","DOI":"10.1613\/jair.1497","volume":"22","author":"R. Becker","year":"2004","unstructured":"R. Becker, S. Zilberstein, V. Lesser et al., Solving transition independent decentralized Markov decision processes. J. Artif. Intell. Res. 22, 423\u2013455 (2004). https:\/\/doi.org\/10.1613\/jair.1497","journal-title":"J. Artif. Intell. Res."},{"issue":"2","key":"29_CR6","doi-asserted-by":"publisher","first-page":"156","DOI":"10.1109\/TSMCC.2007.913919","volume":"38","author":"L. Busoniu","year":"2008","unstructured":"L. Busoniu, R. Babuska, B. De Schutter, A comprehensive survey of multiagent reinforcement learning. IEEE Trans. Syst. Man Cybern., Part C 38(2), 156\u2013172 (2008)","journal-title":"IEEE Trans. Syst. Man Cybern., Part C"},{"key":"29_CR7","first-page":"807","volume-title":"Proceedings of the 16th International Conference on Neural Information Processing Systems","author":"Y.H. Chang","year":"2003","unstructured":"Y.H. Chang, T. Ho, L.P. Kaelbling, All learning is local: multi-agent learning in global reward games, in Proceedings of the 16th International Conference on Neural Information Processing Systems, vol. NIPS\u201903 (MIT Press, Cambridge, 2003), pp. 807\u2013814"},{"key":"29_CR8","first-page":"165","volume-title":"Proceedings of the 2014 International Conference on Autonomous Agents and Multi-Agent Systems","author":"S. Devlin","year":"2014","unstructured":"S. Devlin, L. Yliniemi, D. Kudenko et al., Potential-based difference rewards for multiagent reinforcement learning, in Proceedings of the 2014 International Conference on Autonomous Agents and Multi-Agent Systems (2014), pp. 165\u2013172"},{"key":"29_CR9","first-page":"483","volume-title":"Proceedings of the 20th International Conference on Autonomous Agents and MultiAgent Systems","author":"I. ElSayed-Aly","year":"2021","unstructured":"I. ElSayed-Aly, S. Bharadwaj, C. Amato et al., Safe multi-agent reinforcement learning via shielding, in Proceedings of the 20th International Conference on Autonomous Agents and MultiAgent Systems (2021), pp. 483\u2013491"},{"key":"29_CR10","first-page":"1146","volume-title":"Proceedings of the 34th International Conference on Machine Learning","author":"J. Foerster","year":"2017","unstructured":"J. Foerster, N. Nardelli, G. Farquhar et al., Stabilising experience replay for deep multi-agent reinforcement learning, in Proceedings of the 34th International Conference on Machine Learning, vol. 70 (2017), pp. 1146\u20131155"},{"key":"29_CR11","volume-title":"AAAI 2018: Proceedings of the Thirty-Second AAAI Conference on Artificial Intelligence","author":"J.N. Foerster","year":"2018","unstructured":"J.N. Foerster, G. Farquhar, T. Afouras et al., Counterfactual multi-agent policy gradients, in AAAI 2018: Proceedings of the Thirty-Second AAAI Conference on Artificial Intelligence (2018)"},{"key":"29_CR12","doi-asserted-by":"publisher","first-page":"3389","DOI":"10.1109\/ICRA.2017.7989385","volume-title":"2017 IEEE International Conference on Robotics and Automation (ICRA)","author":"S. Gu","year":"2017","unstructured":"S. Gu, E. Holly, T. Lillicrap et al., Deep reinforcement learning for robotic manipulation with asynchronous off-policy updates, in 2017 IEEE International Conference on Robotics and Automation (ICRA) (IEEE Comput. Soc., Los Alamitos, 2017), pp. 3389\u20133396"},{"key":"29_CR13","doi-asserted-by":"publisher","first-page":"66","DOI":"10.1007\/978-3-319-71682-4_5","volume-title":"International Conference on Autonomous Agents and Multiagent Systems","author":"J.K. Gupta","year":"2017","unstructured":"J.K. Gupta, M. Egorov, M. Kochenderfer, Cooperative multi-agent control using deep reinforcement learning, in International Conference on Autonomous Agents and Multiagent Systems (Springer, Berlin, 2017), pp. 66\u201383"},{"key":"29_CR14","first-page":"602","volume-title":"Proceedings of the 20th International Conference on Autonomous Agents and MultiAgent Systems","author":"K. He","year":"2021","unstructured":"K. He, B. Banerjee, P. Doshi, Cooperative-competitive reinforcement learning with history-dependent rewards, in Proceedings of the 20th International Conference on Autonomous Agents and MultiAgent Systems (2021), pp. 602\u2013610"},{"key":"29_CR15","first-page":"7254","volume-title":"Advances in Neural Information Processing Systems","author":"J. Jiang","year":"2018","unstructured":"J. Jiang, Z. Lu, Learning attentional communication for multi-agent cooperation, in Advances in Neural Information Processing Systems (2018), pp. 7254\u20137264"},{"key":"29_CR16","first-page":"2140","volume-title":"AAAI","author":"G. Lample","year":"2017","unstructured":"G. Lample, D.S. Chaplot, Playing fps games with deep reinforcement learning, in AAAI (2017), pp. 2140\u20132146"},{"key":"29_CR17","doi-asserted-by":"crossref","first-page":"257","DOI":"10.1007\/978-3-030-63833-7_22","volume-title":"Proceedings of 27th International Conference Neural Information Processing (ICONIP 2020)","author":"Y. Miyashita","year":"2020","unstructured":"Y. Miyashita, T. Sugawara, Coordinated behavior for sequential cooperative task using two-stage reward assignment with decay, in Proceedings of 27th International Conference Neural Information Processing (ICONIP 2020) (Springer, Berlin, 2020), pp. 257\u2013269"},{"key":"29_CR18","doi-asserted-by":"publisher","first-page":"550","DOI":"10.1007\/978-3-030-33792-6_40","volume-title":"PRIMA 2019: Principles and Practice of Multi-Agent Systems","author":"Y. Miyashita","year":"2019","unstructured":"Y. Miyashita, T. Sugawara, Coordination in collaborative work by deep reinforcement learning with various state descriptions, in PRIMA 2019: Principles and Practice of Multi-Agent Systems, ed. by M. Baldoni, M. Dastani, B. Liao et al. (Springer, Cham, 2019), pp. 550\u2013558"},{"key":"29_CR19","unstructured":"V. Mnih, K. Kavukcuoglu, D. Silver et al., Playing atari with deep reinforcement learning (2013). arXiv preprint. arXiv:1312.5602"},{"key":"29_CR20","first-page":"278","volume-title":"Proceedings of the Sixteenth International Conference on Machine Learning","author":"A.Y. Ng","year":"1999","unstructured":"A.Y. Ng, D. Harada, S. Russell, Policy invariance under reward transformations: theory and application to reward shaping, in Proceedings of the Sixteenth International Conference on Machine Learning (Morgan Kaufmann, San Mateo, 1999), pp. 278\u2013287"},{"key":"29_CR21","first-page":"8113","volume-title":"Proceedings of the 32nd International Conference on Neural Information Processing Systems","author":"D.T. Nguyen","year":"2018","unstructured":"D.T. Nguyen, A. Kumar, H.C. Lau, Credit assignment for collective multiagent rl with global rewards, in Proceedings of the 32nd International Conference on Neural Information Processing Systems, vol. NIPS\u201918 (Curran Associates, Red Hook, 2018), pp. 8113\u20138124"},{"key":"29_CR22","first-page":"443","volume-title":"Proceedings of the 17th International Conference on Autonomous Agents and MultiAgent Systems, International Foundation for Autonomous Agents and Multiagent Systems","author":"G. Palmer","year":"2018","unstructured":"G. Palmer, K. Tuyls, D. Bloembergen et al., Lenient multi-agent deep reinforcement learning, in Proceedings of the 17th International Conference on Autonomous Agents and MultiAgent Systems, International Foundation for Autonomous Agents and Multiagent Systems (2018), pp. 443\u2013451"},{"key":"29_CR23","first-page":"1","volume-title":"2018 IEEE International Conference on Robotics and Automation (ICRA)","author":"X.B. Peng","year":"2018","unstructured":"X.B. Peng, M. Andrychowicz, W. Zaremba et al., Sim-to-real transfer of robotic control with dynamics randomization, in 2018 IEEE International Conference on Robotics and Automation (ICRA) (IEEE Comput. Soc., Los Alamitos, 2018), pp. 1\u20138"},{"issue":"1","key":"29_CR24","doi-asserted-by":"publisher","first-page":"73","DOI":"10.1109\/TETCI.2018.2823329","volume":"3","author":"K. Shao","year":"2018","unstructured":"K. Shao, Y. Zhu, D. Zhao, Starcraft micromanagement with reinforcement learning and curriculum transfer learning. IEEE Trans. Emerg. Top. Comput. Intell. 3(1), 73\u201384 (2018)","journal-title":"IEEE Trans. Emerg. Top. Comput. Intell."},{"key":"29_CR25","unstructured":"Y. Shoham, R. Powers, T. Grenager, Multi-agent reinforcement learning: a critical survey. Technical report, Stanford University (2003)"},{"key":"29_CR26","doi-asserted-by":"publisher","first-page":"1321","DOI":"10.1145\/1143997.1144202","volume-title":"Proceedings of the 8th Annual Conference on Genetic and Evolutionary Computation","author":"M.E. Taylor","year":"2006","unstructured":"M.E. Taylor, S. Whiteson, P. Stone, Comparing evolutionary and temporal difference methods in a reinforcement learning domain, in Proceedings of the 8th Annual Conference on Genetic and Evolutionary Computation (2006), pp. 1321\u20131328"},{"issue":"2","key":"29_CR27","first-page":"26","volume":"4","author":"T. Tieleman","year":"2012","unstructured":"T. Tieleman, G. Hinton, Lecture 6.5-rmsprop: divide the gradient by a running average of its recent magnitude. COURSERA: Neural Netw. Mach. Learn. 4(2), 26\u201331 (2012)","journal-title":"COURSERA: Neural Netw. Mach. Learn."},{"issue":"4\u20135","key":"29_CR28","doi-asserted-by":"publisher","first-page":"475","DOI":"10.1142\/S0219525909002295","volume":"12","author":"K. Tumer","year":"2009","unstructured":"K. Tumer, A. Agogino, Multiagent learning for black box system reward functions. Adv. Complex Syst. 12(4\u20135), 475\u2013492 (2009)","journal-title":"Adv. Complex Syst."},{"key":"29_CR29","first-page":"5","volume-title":"AAAI","author":"H. Van Hasselt","year":"2016","unstructured":"H. Van Hasselt, A. Guez, D. Silver, Deep reinforcement learning with double q-learning, in AAAI, Phoenix (2016), p. 5"},{"key":"29_CR30","doi-asserted-by":"publisher","first-page":"355","DOI":"10.1142\/9789812777263_0020","volume-title":"Modeling Complexity in Economic and Social Systems","author":"D.H. Wolpert","year":"2002","unstructured":"D.H. Wolpert, K. Tumer, Optimal payoff functions for members of collectives, in Modeling Complexity in Economic and Social Systems (World Scientific, Singapore, 2002), pp. 355\u2013369"}],"container-title":["Autonomous Intelligent Systems"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s43684-022-00029-z.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s43684-022-00029-z\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s43684-022-00029-z.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2023,2,6]],"date-time":"2023-02-06T13:07:13Z","timestamp":1675688833000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s43684-022-00029-z"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,5,27]]},"references-count":30,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2022,12]]}},"alternative-id":["29"],"URL":"https:\/\/doi.org\/10.1007\/s43684-022-00029-z","relation":{},"ISSN":["2730-616X"],"issn-type":[{"value":"2730-616X","type":"electronic"}],"subject":[],"published":{"date-parts":[[2022,5,27]]},"assertion":[{"value":"27 October 2021","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"26 February 2022","order":2,"name":"revised","label":"Revised","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"13 May 2022","order":3,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"27 May 2022","order":4,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"The authors declare that they have no competing interests.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Competing interests"}}],"article-number":"10"}}