{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,4]],"date-time":"2026-06-04T16:11:01Z","timestamp":1780589461630,"version":"3.54.1"},"publisher-location":"California","reference-count":0,"publisher":"International Joint Conferences on Artificial Intelligence Organization","content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":[],"published-print":{"date-parts":[[2022,7]]},"abstract":"<jats:p>The theoretical underpinnings of multiagent reinforcement learning has recently attracted much attention.\n\nIn this work, we focus on the generalized social learning (GSL) protocol --- an agent interaction protocol that is widely adopted in the literature, and aim to develop an accurate theoretical model for the Q-learning dynamics under this protocol.\n\nNoting that previous models fail to characterize the effects of local interactions and incomplete information that arise from GSL, we model the Q-values dynamics of each individual agent as a system of stochastic differential equations (SDE). \n\nBased on the SDE, we express the time evolution of the probability density function of Q-values in the population with a Fokker-Planck equation.\n\nWe validate the correctness of our model through extensive comparisons with agent-based simulation results across different types of symmetric games. \n\nIn addition, we show that as the interactions between agents are more limited and information is less complete, the population can converge to a outcome that is qualitatively different than that with global interactions and complete information.<\/jats:p>","DOI":"10.24963\/ijcai.2022\/55","type":"proceedings-article","created":{"date-parts":[[2022,7,16]],"date-time":"2022-07-16T02:55:56Z","timestamp":1657940156000},"page":"384-390","source":"Crossref","is-referenced-by-count":4,"title":["Modelling the Dynamics of Multi-Agent Q-learning: The Stochastic Effects of Local Interaction and Incomplete Information"],"prefix":"10.24963","author":[{"given":"Chin-wing","family":"Leung","sequence":"first","affiliation":[{"name":"The Chinese University of Hong Kong"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Shuyue","family":"Hu","sequence":"additional","affiliation":[{"name":"Shanghai Artificial Intelligence Laboratory"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Ho-fung","family":"Leung","sequence":"additional","affiliation":[{"name":"The Chinese University of Hong Kong"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"10584","event":{"name":"Thirty-First International Joint Conference on Artificial Intelligence {IJCAI-22}","theme":"Artificial Intelligence","location":"Vienna, Austria","acronym":"IJCAI-2022","number":"31","sponsor":["International Joint Conferences on Artificial Intelligence Organization (IJCAI)"],"start":{"date-parts":[[2022,7,23]]},"end":{"date-parts":[[2022,7,29]]}},"container-title":["Proceedings of the Thirty-First International Joint Conference on Artificial Intelligence"],"original-title":[],"deposited":{"date-parts":[[2022,7,18]],"date-time":"2022-07-18T11:07:26Z","timestamp":1658142446000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.ijcai.org\/proceedings\/2022\/55"}},"subtitle":[],"proceedings-subject":"Artificial Intelligence Research Articles","short-title":[],"issued":{"date-parts":[[2022,7]]},"references-count":0,"URL":"https:\/\/doi.org\/10.24963\/ijcai.2022\/55","relation":{},"subject":[],"published":{"date-parts":[[2022,7]]}}}