{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,11,15]],"date-time":"2025-11-15T10:26:55Z","timestamp":1763202415182,"version":"3.37.3"},"reference-count":29,"publisher":"Springer Science and Business Media LLC","issue":"2","license":[{"start":{"date-parts":[[2020,9,16]],"date-time":"2020-09-16T00:00:00Z","timestamp":1600214400000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2020,9,16]],"date-time":"2020-09-16T00:00:00Z","timestamp":1600214400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"DOI":"10.13039\/501100001691","name":"KAKENHI","doi-asserted-by":"crossref","award":["17KT0044"],"award-info":[{"award-number":["17KT0044"]}],"id":[{"id":"10.13039\/501100001691","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Appl Intell"],"published-print":{"date-parts":[[2021,2]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:p>Cooperation and coordination are major issues in studies on multi-agent systems because the entire performance of such systems is greatly affected by these activities. The issues are challenging however, because appropriate coordinated behaviors depend on not only environmental characteristics but also other agents\u2019 strategies. On the other hand, advances in multi-agent deep reinforcement learning (MADRL) have recently attracted attention, because MADRL can considerably improve the entire performance of multi-agent systems in certain domains. The characteristics of learned coordination structures and agent\u2019s resulting behaviors, however, have not been clarified sufficiently. Therefore, we focus here on MADRL in which agents have their own deep Q-networks (DQNs), and we analyze their coordinated behaviors and structures for the<jats:italic>pickup and floor laying problem<\/jats:italic>, which is an abstraction of our target application. In particular, we analyze the behaviors around scarce resources and long narrow passages in which conflicts such as collisions are likely to occur. We then indicated that different types of inputs to the networks exhibit similar performance but generate various coordination structures with associated behaviors, such as division of labor and a shared social norm, with no direct communication.<\/jats:p>","DOI":"10.1007\/s10489-020-01832-y","type":"journal-article","created":{"date-parts":[[2020,9,16]],"date-time":"2020-09-16T05:02:30Z","timestamp":1600232550000},"page":"1069-1085","update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":10,"title":["Analysis of coordinated behavior structures with multi-agent deep reinforcement learning"],"prefix":"10.1007","volume":"51","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-1676-9346","authenticated-orcid":false,"given":"Yuki","family":"Miyashita","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Toshiharu","family":"Sugawara","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"297","published-online":{"date-parts":[[2020,9,16]]},"reference":[{"key":"1832_CR1","doi-asserted-by":"crossref","unstructured":"Agmon N, Kraus S, Kaminka GA (2008) Multi-robot perimeter patrol in adversarial settings. In: 2008 IEEE International conference on robotics and automation. IEEE, pp 2339\u20132345","DOI":"10.1109\/ROBOT.2008.4543563"},{"key":"1832_CR2","doi-asserted-by":"publisher","first-page":"659","DOI":"10.1613\/jair.4818","volume":"53","author":"D Bloembergen","year":"2015","unstructured":"Bloembergen D, Tuyls K, Hennes D, Kaisers M (2015) Evolutionary dynamics of multi-agent learning: a survey. J Artif Intell Res 53:659\u2013697","journal-title":"J Artif Intell Res"},{"issue":"2","key":"1832_CR3","doi-asserted-by":"publisher","first-page":"156","DOI":"10.1109\/TSMCC.2007.913919","volume":"38","author":"L Busoniu","year":"2008","unstructured":"Busoniu L, Babuska R, De Schutter B (2008) A comprehensive survey of multiagent reinforcement learning. IEEE Trans Systems, Man, and Cybernetics Part C 38(2):156\u2013172","journal-title":"IEEE Trans Systems, Man, and Cybernetics Part C"},{"key":"1832_CR4","doi-asserted-by":"crossref","unstructured":"Diallo EAO, Sugawara T (2018) Learning strategic group formation for coordinated behavior in adversarial multi-agent with double DQN. In: International Conference on Principles and Practice of Multi-Agent Systems. Springer, pp 458\u2013466","DOI":"10.1007\/978-3-030-03098-8_30"},{"key":"1832_CR5","unstructured":"Foerster J, Nardelli N, Farquhar G, Torr P, Kohli P, Whiteson S, et al. (2017) Stabilising experience replay for deep multi-agent reinforcement learning. arXiv:1702.08887"},{"key":"1832_CR6","doi-asserted-by":"crossref","unstructured":"Giuggioli L, Arye I, Robles AH, Kaminka GA (2018) From ants to birds: A novel bio-inspired approach to online area coverage. In: Distributed autonomous robotic systems. Springer, pp 31\u201343","DOI":"10.1007\/978-3-319-73008-0_3"},{"key":"1832_CR7","doi-asserted-by":"crossref","unstructured":"Gu S, Holly E, Lillicrap T, Levine S (2017) Deep reinforcement learning for robotic manipulation with asynchronous off-policy updates. In: 2017 IEEE International Conference on Robotics and Automation (ICRA). IEEE, pp 3389\u20133396","DOI":"10.1109\/ICRA.2017.7989385"},{"key":"1832_CR8","doi-asserted-by":"crossref","unstructured":"Lample G, Chaplot DS (2017) Playing FPS games with deep reinforcement learning. In: AAAI, pp 2140\u20132146","DOI":"10.1609\/aaai.v31i1.10827"},{"issue":"1","key":"1832_CR9","first-page":"1334","volume":"17","author":"S Levine","year":"2016","unstructured":"Levine S, Finn C, Darrell T, Abbeel P (2016) End-to-end training of deep visuomotor policies. J Mach Learning Res 17(1):1334\u20131373","journal-title":"J Mach Learning Res"},{"key":"1832_CR10","unstructured":"Liu M, Ma H, Li J, Koenig S (2019) Task and path planning for multi-agent pickup and delivery. In: Proceedings of the 18th international conference on autonomous agents and multiagent systems. IFAAMAS, pp 1152\u20131160"},{"key":"1832_CR11","unstructured":"Lowe R, Wu Y, Tamar A, Harb J, Abbeel OP, Mordatch I (2017) Multi-agent actor-critic for mixed cooperative-competitive environments. In: Advances in neural information processing systems, pp 6382\u20136393"},{"key":"1832_CR12","unstructured":"Luo W, Nam C, Kantor G, Sycara K (2019) Distributed environmental modeling and adaptive sampling for multi-robot sensor coverage. In: Proceedings of the 18th international conference on autonomous agents and multiagent systems. IFAAMAS, pp 1488\u20131496"},{"key":"1832_CR13","unstructured":"Mao H, Zhang Z, Xiao Z, Gong Z (2019) Modelling the dynamic joint policy of teammates with attention multi-agent DDPG. In: Proceedings of the 18th international conference on autonomous agents and multiagent systems. IFAAMAS, pp 1108\u20131116"},{"key":"1832_CR14","unstructured":"Mnih V, Kavukcuoglu K, Silver D, Graves A, Antonoglou I, Wierstra D, Riedmiller M (2013) Playing atari with deep reinforcement learning. arXiv:1312.5602"},{"issue":"7540","key":"1832_CR15","doi-asserted-by":"publisher","first-page":"529","DOI":"10.1038\/nature14236","volume":"518","author":"V Mnih","year":"2015","unstructured":"Mnih V, Kavukcuoglu K, Silver D, Rusu AA, Veness J, Bellemare MG, Graves A, Riedmiller M, Fidjeland AK, Ostrovski G, et al. (2015) Human-level control through deep reinforcement learning. Nature 518(7540):529","journal-title":"Nature"},{"key":"1832_CR16","unstructured":"Palmer G, Tuyls K, Bloembergen D, Savani R (2018) Lenient multi-agent deep reinforcement learning. In: Proceedings of the 17th international conference on autonomous agents and multiagent systems. IFAAMAS, pp 443\u2013451"},{"key":"1832_CR17","doi-asserted-by":"crossref","unstructured":"Peng XB, Andrychowicz M, Zaremba W, Abbeel P (2018) Sim-to-real transfer of robotic control with dynamics randomization. In: 2018 IEEE International Conference on Robotics and Automation (ICRA). IEEE, pp 1\u20138","DOI":"10.1109\/ICRA.2018.8460528"},{"key":"1832_CR18","doi-asserted-by":"crossref","unstructured":"Sartoretti G, Wu Y, Paivine W, Kumar TS, Koenig S, Choset H (2019) Distributed reinforcement learning for multi-robot decentralized collective construction. In: Distributed autonomous robotic systems. DARS, Springer, pp 35\u201349","DOI":"10.1007\/978-3-030-05816-6_3"},{"key":"1832_CR19","unstructured":"Schaul T, Quan J, Antonoglou I, Silver D (2015) Prioritized experience replay. arXiv:1511.05952"},{"key":"1832_CR20","doi-asserted-by":"crossref","unstructured":"Shao K, Zhu Y, Zhao D (2018) StarCraft micromanagement with reinforcement learning and curriculum transfer learning. IEEE Transactions on emerging topics in computational intelligence (99):1\u201312","DOI":"10.1109\/TETCI.2018.2823329"},{"key":"1832_CR21","unstructured":"Shoham Y, Powers R, Grenager T (2003) Multi-agent reinforcement learning: a critical survey. Tech. rep. Technical report, Stanford University"},{"key":"1832_CR22","unstructured":"Silver D, Lever G, Heess N, Degris T, Wierstra D, Riedmiller M (2014) Deterministic policy gradient algorithms. In: Proceedings of the 31st international conference on international conference on machine learning - volume 32: I\u2013387\u2013I\u2013395. ICML\u201914, JMLR.org"},{"key":"1832_CR23","doi-asserted-by":"crossref","unstructured":"Sugiyama A, Sea V, Sugawara T (2018) Emergence of divisional cooperation with negotiation and re-learning and evaluation of flexibility in continuous cooperative patrol problem. Knowledge and Information Systems","DOI":"10.1007\/s10115-018-1285-8"},{"key":"1832_CR24","unstructured":"Sunehag P, Lever G, Gruslys A, Czarnecki WM, Zambaldi V, Jaderberg M, Lanctot M, Sonnerat N, Leibo JZ, Tuyls K, et al. (2017) Value-decomposition networks for cooperative multi-agent learning. arXiv:1706.05296"},{"issue":"2","key":"1832_CR25","first-page":"26","volume":"4","author":"T Tieleman","year":"2012","unstructured":"Tieleman T, Hinton G (2012) Lecture 6.5-rmsprop: Divide the gradient by a running average of its recent magnitude. COURSERA: Neural networks for machine learning 4(2):26\u201331","journal-title":"COURSERA: Neural networks for machine learning"},{"key":"1832_CR26","doi-asserted-by":"crossref","unstructured":"Van Hasselt H, Guez A, Silver D (2016) Deep reinforcement learning with double Q-learning","DOI":"10.1609\/aaai.v30i1.10295"},{"issue":"3-4","key":"1832_CR27","doi-asserted-by":"publisher","first-page":"279","DOI":"10.1007\/BF00992698","volume":"8","author":"CJ Watkins","year":"1992","unstructured":"Watkins CJ, Dayan P (1992) Q-learning. Machine learning 8(3-4):279\u2013292","journal-title":"Machine learning"},{"key":"1832_CR28","doi-asserted-by":"crossref","unstructured":"Xie M, Tachibana A (2007) Cooperative behavior acquisition for multi-agent systems by Q-learning. In: Foundations of computational intelligence, 2007. FOCI 2007. IEEE Symposium on. pp 424\u2013428. IEEE","DOI":"10.1109\/FOCI.2007.371506"},{"key":"1832_CR29","doi-asserted-by":"crossref","unstructured":"Miyashita Y, Sugawara T (2019) Coordination in collaborative work by deep reinforcement learning with various state descriptions. In: International conference on principles and practice of multi-agent systems, pp 550\u2013558","DOI":"10.1007\/978-3-030-33792-6_40"}],"container-title":["Applied Intelligence"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10489-020-01832-y.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s10489-020-01832-y\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10489-020-01832-y.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2022,11,18]],"date-time":"2022-11-18T14:57:12Z","timestamp":1668783432000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s10489-020-01832-y"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2020,9,16]]},"references-count":29,"journal-issue":{"issue":"2","published-print":{"date-parts":[[2021,2]]}},"alternative-id":["1832"],"URL":"https:\/\/doi.org\/10.1007\/s10489-020-01832-y","relation":{},"ISSN":["0924-669X","1573-7497"],"issn-type":[{"type":"print","value":"0924-669X"},{"type":"electronic","value":"1573-7497"}],"subject":[],"published":{"date-parts":[[2020,9,16]]},"assertion":[{"value":"16 September 2020","order":1,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Compliance with Ethical Standards"}},{"value":"The authors declare that they have no conflict of interest.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"<!--Emphasis Type='Bold' removed-->Conflict of interests"}}]}}