{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,6,19]],"date-time":"2025-06-19T04:16:05Z","timestamp":1750306565544,"version":"3.41.0"},"reference-count":28,"publisher":"Association for Computing Machinery (ACM)","issue":"4","license":[{"start":{"date-parts":[[2014,12,19]],"date-time":"2014-12-19T00:00:00Z","timestamp":1418947200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"DOI":"10.13039\/100000001","name":"National Science Foundation","doi-asserted-by":"publisher","id":[{"id":"10.13039\/100000001","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100004853","name":"Chinese University of Hong Kong","doi-asserted-by":"publisher","id":[{"id":"10.13039\/501100004853","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Auton. Adapt. Syst."],"published-print":{"date-parts":[[2015,1,14]]},"abstract":"<jats:p>\n            Most previous works on coordination in cooperative multiagent systems study the problem of how two (or more) players can coordinate on Pareto-optimal Nash equilibrium(s) through fixed and repeated interactions in the context of cooperative games. However, in practical complex environments, the interactions between agents can be sparse, and each agent's interacting partners may change frequently and randomly. To this end, we investigate the multiagent coordination problems in cooperative environments under a social learning framework. We consider a large population of agents where each agent interacts with another agent randomly chosen from the population in each round. Each agent learns its policy through repeated interactions with the rest of the agents via social learning. It is not clear a priori if all agents can learn a consistent optimal coordination policy in such a situation. We distinguish two different types of learners depending on the amount of information each agent can perceive:\n            <jats:italic>individual action learner and joint action learner<\/jats:italic>\n            . The learning performance of both types of learners is evaluated under a number of challenging deterministic and stochastic cooperative games, and the influence of the information sharing degree on the learning performance also is investigated\u2014a key difference from the learning framework involving repeated interactions among fixed agents.\n          <\/jats:p>","DOI":"10.1145\/2644819","type":"journal-article","created":{"date-parts":[[2014,12,22]],"date-time":"2014-12-22T13:53:23Z","timestamp":1419256403000},"page":"1-20","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":13,"title":["Multiagent Reinforcement Social Learning toward Coordination in Cooperative Multiagent Systems"],"prefix":"10.1145","volume":"9","author":[{"given":"Jianye","family":"Hao","sequence":"first","affiliation":[{"name":"Massachusetts Institute of Technology Shenzhen University, Cambridge, MA, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Ho-Fung","family":"Leung","sequence":"additional","affiliation":[{"name":"The Chinese University of Hong Kong, Shatin, Hong Kong"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Zhong","family":"Ming","sequence":"additional","affiliation":[{"name":"Shenzhen University, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2014,12,19]]},"reference":[{"key":"e_1_2_1_1_1","first-page":"555","article-title":"Game artificial intelligence based using reinforcement learning","volume":"50","author":"Agung Albertus","year":"2012","unstructured":"Albertus Agung and Ford Lumban Gaol . 2012 . Game artificial intelligence based using reinforcement learning . Procedia Engineering 50 (2012), 555 -- 565 . Albertus Agung and Ford Lumban Gaol. 2012. Game artificial intelligence based using reinforcement learning. Procedia Engineering 50 (2012), 555--565.","journal-title":"Procedia Engineering"},{"key":"e_1_2_1_2_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.artint.2004.04.013"},{"key":"e_1_2_1_3_1","doi-asserted-by":"publisher","DOI":"10.1109\/CDC.2011.6161294"},{"volume-title":"Proceedings of AAAI'98","author":"Claus C.","key":"e_1_2_1_4_1","unstructured":"C. Claus and C. Boutilier . 1998. The dynamics of reinforcement learning in cooperative multiagent systems . In Proceedings of AAAI'98 . 746--752. C. Claus and C. Boutilier. 1998. The dynamics of reinforcement learning in cooperative multiagent systems. In Proceedings of AAAI'98. 746--752."},{"key":"e_1_2_1_5_1","unstructured":"R. C. Eberhart Y. H. Shi and J. Kennedy. 2001. Swarm Intelligence. Elsevier.  R. C. Eberhart Y. H. Shi and J. Kennedy. 2001. Swarm Intelligence. Elsevier."},{"key":"e_1_2_1_6_1","unstructured":"I. Foster and C. Kesselman. 2003. The Grid 2: Blueprint for a New Computing Infrastructure. Morgan Kaufmann.   I. Foster and C. Kesselman. 2003. The Grid 2: Blueprint for a New Computing Infrastructure. Morgan Kaufmann."},{"volume-title":"Proceedings of IJCAI'07","author":"Fulda N.","key":"e_1_2_1_7_1","unstructured":"N. Fulda and D. Ventura . 2007. Predicting and preventing coordination problems in cooperative learning systems . In Proceedings of IJCAI'07 . 780--785. N. Fulda and D. Ventura. 2007. Predicting and preventing coordination problems in cooperative learning systems. In Proceedings of IJCAI'07. 780--785."},{"key":"e_1_2_1_8_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.geb.2012.06.001"},{"key":"e_1_2_1_9_1","volume-title":"Moore","author":"Kaelbling Leslie Pack","year":"1996","unstructured":"Leslie Pack Kaelbling , Michael L. Littman , and Andrew W . Moore . 1996 . Reinforcement learning: A survey. Artificial Intelligence Research ( 1996), 237--285. Leslie Pack Kaelbling, Michael L. Littman, and Andrew W. Moore. 1996. Reinforcement learning: A survey. Artificial Intelligence Research (1996), 237--285."},{"volume-title":"Proceedings of AAAI'02","author":"Kapetanakis S.","key":"e_1_2_1_10_1","unstructured":"S. Kapetanakis and D. Kudenko . 2002. Reinforcement learning of coordination in cooperative multiagent systems . In Proceedings of AAAI'02 . 326--331. S. Kapetanakis and D. Kudenko. 2002. Reinforcement learning of coordination in cooperative multiagent systems. In Proceedings of AAAI'02. 326--331."},{"key":"e_1_2_1_11_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.geb.2013.02.006"},{"volume-title":"Proceedings of ICML'00","author":"Lauer M.","key":"e_1_2_1_12_1","unstructured":"M. Lauer and M. Riedmiller . 2000. An algorithm for distributed reinforcement learning in cooperative multi-agent systems . In Proceedings of ICML'00 . 535--542. M. Lauer and M. Riedmiller. 2000. An algorithm for distributed reinforcement learning in cooperative multi-agent systems. In Proceedings of ICML'00. 535--542."},{"key":"e_1_2_1_13_1","doi-asserted-by":"crossref","unstructured":"N. E. Leonard T. Shen B. Nabet L. Scardovi I. D. Couzin and S. A. Levin. 2012. Decision versus compromise for animal groups in motion. In Proceedings of the National Academy of Sciences. 227--232.  N. E. Leonard T. Shen B. Nabet L. Scardovi I. D. Couzin and S. A. Levin. 2012. Decision versus compromise for animal groups in motion. In Proceedings of the National Academy of Sciences. 227--232.","DOI":"10.1073\/pnas.1118318108"},{"key":"e_1_2_1_14_1","volume-title":"AAMAS'08 Workshop: MSDM. 77--91","author":"Matignon L.","year":"2008","unstructured":"L. Matignon , G. J. Laurent , and N. Le For-Piat . 2008 . A study of FMQ heuristic in cooperative multi-agent games . In AAMAS'08 Workshop: MSDM. 77--91 . L. Matignon, G. J. Laurent, and N. Le For-Piat. 2008. A study of FMQ heuristic in cooperative multi-agent games. In AAMAS'08 Workshop: MSDM. 77--91."},{"key":"e_1_2_1_15_1","doi-asserted-by":"publisher","DOI":"10.1017\/S0269888912000057"},{"key":"e_1_2_1_16_1","volume-title":"Proceedings of IROS'07","author":"Matignon L.","year":"2007","unstructured":"L. Matignon , G. J. Laurent , and N. Le Fort-Piat . 2007 . Hysteretic q-learning: An algorithm for dynamic reinforcement learning in cooperative multiagent teams . In Proceedings of IROS'07 . 64--69. L. Matignon, G. J. Laurent, and N. Le Fort-Piat. 2007. Hysteretic q-learning: An algorithm for dynamic reinforcement learning in cooperative multiagent teams. In Proceedings of IROS'07. 64--69."},{"key":"e_1_2_1_17_1","doi-asserted-by":"publisher","DOI":"10.1109\/65.967595"},{"key":"e_1_2_1_18_1","doi-asserted-by":"publisher","DOI":"10.1007\/s10458-005-2631-2"},{"key":"e_1_2_1_19_1","doi-asserted-by":"publisher","DOI":"10.1145\/1160633.1160776"},{"key":"#cr-split#-e_1_2_1_20_1.1","unstructured":"L. Rendell R. Boyd D. Cownden M. Enquist K. Eriksson M. W. Feldman L. Fogarty S. Ghirlanda T. Lillicrap and K. N. Laland. 2010. Why copy others&quest"},{"key":"#cr-split#-e_1_2_1_20_1.2","doi-asserted-by":"crossref","unstructured":"insights from the social learning strategies tournament. Science 328(5975) (2010) 208 213. L. Rendell R. Boyd D. Cownden M. Enquist K. Eriksson M. W. Feldman L. Fogarty S. Ghirlanda T. Lillicrap and K. N. Laland. 2010. Why copy others&quest","DOI":"10.1126\/science.1184719"},{"key":"#cr-split#-e_1_2_1_20_1.3","doi-asserted-by":"crossref","unstructured":"insights from the social learning strategies tournament. Science 328(5975) (2010) 208 213.","DOI":"10.1126\/science.1184719"},{"volume-title":"Proceedings of IJCAI'07","author":"Sen S.","key":"e_1_2_1_21_1","unstructured":"S. Sen and S. Airiau . 2007. Emergence of norms through social learning . In Proceedings of IJCAI'07 . 1507--1512. S. Sen and S. Airiau. 2007. Emergence of norms through social learning. In Proceedings of IJCAI'07. 1507--1512."},{"key":"e_1_2_1_22_1","doi-asserted-by":"publisher","DOI":"10.1038\/nature09203"},{"key":"e_1_2_1_23_1","doi-asserted-by":"publisher","DOI":"10.5555\/2283396.2283465"},{"key":"e_1_2_1_24_1","doi-asserted-by":"publisher","DOI":"10.1145\/2451248.2451250"},{"volume-title":"Proceedings of NIPS'02","author":"Wang X.","key":"e_1_2_1_25_1","unstructured":"X. Wang and T. Sandholm . 2002. Reinforcement learning to play an optimal Nash equilibrium in team Markov games . In Proceedings of NIPS'02 . 1571--1578. X. Wang and T. Sandholm. 2002. Reinforcement learning to play an optimal Nash equilibrium in team Markov games. In Proceedings of NIPS'02. 1571--1578."},{"key":"e_1_2_1_26_1","doi-asserted-by":"publisher","DOI":"10.1007\/BF00992698"}],"container-title":["ACM Transactions on Autonomous and Adaptive Systems"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/2644819","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/2644819","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T06:13:12Z","timestamp":1750227192000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/2644819"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2014,12,19]]},"references-count":28,"journal-issue":{"issue":"4","published-print":{"date-parts":[[2015,1,14]]}},"alternative-id":["10.1145\/2644819"],"URL":"https:\/\/doi.org\/10.1145\/2644819","relation":{},"ISSN":["1556-4665","1556-4703"],"issn-type":[{"type":"print","value":"1556-4665"},{"type":"electronic","value":"1556-4703"}],"subject":[],"published":{"date-parts":[[2014,12,19]]},"assertion":[{"value":"2013-10-01","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2014-07-01","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2014-12-19","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}