{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,19]],"date-time":"2026-05-19T04:45:15Z","timestamp":1779165915740,"version":"3.51.4"},"reference-count":26,"publisher":"MDPI AG","issue":"3","license":[{"start":{"date-parts":[[2022,2,22]],"date-time":"2022-02-22T00:00:00Z","timestamp":1645488000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"DOI":"10.13039\/501100008982","name":"National Science Foundation","doi-asserted-by":"publisher","award":["CCF-1527486 and CNS-1618335."],"award-info":[{"award-number":["CCF-1527486 and CNS-1618335."]}],"id":[{"id":"10.13039\/501100008982","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Algorithms"],"abstract":"<jats:p>Stochastic games provide a framework for interactions among multiple agents and enable a myriad of applications. In these games, agents decide on actions simultaneously. After taking an action, the state of every agent updates to the next state, and each agent receives a reward. However, finding an equilibrium (if exists) in this game is often difficult when the number of agents becomes large. This paper focuses on finding a mean-field equilibrium (MFE) in an action-coupled stochastic game setting in an episodic framework. It is assumed that an agent can approximate the impact of the other agents\u2019 by the empirical distribution of the mean of the actions. All agents know the action distribution and employ lower-myopic best response dynamics to choose the optimal oblivious strategy. This paper proposes a posterior sampling-based approach for reinforcement learning in the mean-field game, where each agent samples a transition probability from the previous transitions. We show that the policy and action distributions converge to the optimal oblivious strategy and the limiting distribution, respectively, which constitute an MFE.<\/jats:p>","DOI":"10.3390\/a15030073","type":"journal-article","created":{"date-parts":[[2022,2,22]],"date-time":"2022-02-22T22:35:09Z","timestamp":1645569309000},"page":"73","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":4,"title":["Reinforcement Learning for Mean-Field Game"],"prefix":"10.3390","volume":"15","author":[{"given":"Mridul","family":"Agarwal","sequence":"first","affiliation":[{"name":"School of Electrical and Computer Engineering, Purdue University, West Lafayette, IN 47907, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Vaneet","family":"Aggarwal","sequence":"additional","affiliation":[{"name":"School of Electrical and Computer Engineering, Purdue University, West Lafayette, IN 47907, USA"},{"name":"School of Industrial Engineering, Purdue University, West Lafayette, IN 47907, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Arnob","family":"Ghosh","sequence":"additional","affiliation":[{"name":"Department of Electrical and Computer Engineering, Ohio State University, Columbus, OH 43210, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Nilay","family":"Tiwari","sequence":"additional","affiliation":[{"name":"Department of Electrical Engineering, I.I.T. Kanpur, Kanpur 208016, UP, India"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2022,2,22]]},"reference":[{"key":"ref_1","unstructured":"Tan, M. (1993, January 27\u201329). Multi-agent reinforcement learning: Independent versus cooperative agents. Proceedings of the 10th International Conference on International Conference on Machine Learning, Amherst, MA, USA."},{"key":"ref_2","doi-asserted-by":"crossref","first-page":"387","DOI":"10.1007\/s10458-005-2631-2","article-title":"Cooperative multi-agent learning: The state of the art","volume":"11","author":"Panait","year":"2005","journal-title":"Auton. Agents Multi-Agent Syst."},{"key":"ref_3","unstructured":"Littman, M.L. (July, January 28). Friend-or-foe Q-learning in general-sum games. Proceedings of the ICML, Williamstown, MA, USA."},{"key":"ref_4","first-page":"1039","article-title":"Nash Q-learning for general-sum stochastic games","volume":"4","author":"Hu","year":"2003","journal-title":"J. Mach. Learn. Res."},{"key":"ref_5","unstructured":"Yang, Y., Luo, R., Li, M., Zhou, M., Zhang, W., and Wang, J. (2018). Mean Field Multi-Agent Reinforcement Learning. arXiv."},{"key":"ref_6","doi-asserted-by":"crossref","first-page":"1550147719831180","DOI":"10.1177\/1550147719831180","article-title":"Optimal defense strategy based on the mean field game model for cyber security","volume":"15","author":"Miao","year":"2019","journal-title":"Int. J. Distrib. Sens. Netw."},{"key":"ref_7","doi-asserted-by":"crossref","first-page":"669","DOI":"10.1007\/s00245-016-9389-6","article-title":"Mean-field-game model for botnet defense in cyber-security","volume":"74","author":"Kolokoltsov","year":"2016","journal-title":"Appl. Math. Optim."},{"key":"ref_8","doi-asserted-by":"crossref","unstructured":"Li, J., Xia, B., Geng, X., Ming, H., Shakkottai, S., Subramanian, V., and Xie, L. (2015, January 15\u201319). Energy Coupon: A Mean Field Game Perspective on Demand Response in Smart Grids. Proceedings of the 2015 ACM SIGMETRICS International Conference on Measurement and Modeling of Computer Systems, Portland, OR, USA.","DOI":"10.1145\/2745844.2745890"},{"key":"ref_9","doi-asserted-by":"crossref","unstructured":"Farzaneh, H., Kebriaei, H., and Aminifar, F. (2018, January 28\u201329). Deterministic Mean Field Game for Energy Management in a Utility with Many Users. Proceedings of the 2018 Smart Grid Conference (SGC), Sanandaj, Iran.","DOI":"10.1109\/SGC.2018.8777759"},{"key":"ref_10","doi-asserted-by":"crossref","first-page":"971","DOI":"10.1287\/opre.2013.1192","article-title":"Mean field equilibrium in dynamic games with strategic complementarities","volume":"61","author":"Adlakha","year":"2013","journal-title":"Oper. Res."},{"key":"ref_11","doi-asserted-by":"crossref","unstructured":"Adlakha, S., and Johari, R. (2010). Mean Field Equilibrium in Dynamic Games with Complementarities. arXiv.","DOI":"10.2139\/ssrn.1583456"},{"key":"ref_12","doi-asserted-by":"crossref","first-page":"1375","DOI":"10.3982\/ECTA6158","article-title":"Markov perfect industry dynamics with many firms","volume":"76","author":"Weintraub","year":"2008","journal-title":"Econometrica"},{"key":"ref_13","doi-asserted-by":"crossref","first-page":"269","DOI":"10.1016\/j.jet.2013.07.002","article-title":"Equilibria of dynamic games with many players: Existence, approximation, and market structure","volume":"156","author":"Adlakha","year":"2015","journal-title":"J. Econ. Theory"},{"key":"ref_14","unstructured":"Agrawal, S., and Jia, R. (2017). Optimistic posterior sampling for reinforcement learning: Worst-case regret bounds. arXiv."},{"key":"ref_15","unstructured":"Osband, I., Russo, D., and Van Roy, B. (2013). (More) efficient reinforcement learning via posterior sampling. arXiv."},{"key":"ref_16","unstructured":"Subramanian, J., and Mahajan, A. (2019, January 13\u201317). Reinforcement learning in stationary mean-field games. Proceedings of the 18th International Conference on Autonomous Agents and MultiAgent Systems, Montreal, QC, Canada."},{"key":"ref_17","unstructured":"Guo, X., Hu, A., Xu, R., and Zhang, J. (2019). Learning mean-field games. arXiv."},{"key":"ref_18","unstructured":"Anahtarc\u0131, B., Kar\u0131ks\u0131z, C.D., and Saldi, N. (2019). Fitted Q-Learning in Mean-field Games. arXiv."},{"key":"ref_19","unstructured":"Fu, Z., Yang, Z., Chen, Y., and Wang, Z. (2019). Actor-critic provably finds Nash equilibria of linear-quadratic mean-field games. arXiv."},{"key":"ref_20","unstructured":"Yang, J., Ye, X., Trivedi, R., Xu, H., and Zha, H. (2017). Learning deep mean field games for modeling large population behavior. arXiv."},{"key":"ref_21","unstructured":"Pong, V., Gu, S., Dalal, M., and Levine, S. (2018). Temporal difference models: Model-free deep rl for model-based control. arXiv."},{"key":"ref_22","doi-asserted-by":"crossref","unstructured":"Mguni, D., Jennings, J., and de Cote, E.M. (2018, January 2\u20137). Decentralised learning in systems with many, many strategic agents. Proceedings of the 32nd AAAI Conference on Artificial Intelligence, New Orleans, LA, USA.","DOI":"10.1609\/aaai.v32i1.11586"},{"key":"ref_23","unstructured":"Elie, R., P\u00e9rolat, J., Lauri\u00e8re, M., Geist, M., and Pietquin, O. (2019). Approximate fictitious play for mean field games. arXiv."},{"key":"ref_24","doi-asserted-by":"crossref","unstructured":"Light, B., and Weintraub, G.Y. (2018). Mean field equilibrium: Uniqueness, existence, and comparative statics. Columbia Business School Research Paper, Columbia Business School Publishing. Paper No. 19-3.","DOI":"10.2139\/ssrn.3265048"},{"key":"ref_25","unstructured":"Puterman, M.L. (2014). Markov Decision Processes: Discrete Stochastic Dynamic Programming, John Wiley & Sons."},{"key":"ref_26","unstructured":"Mondal, W.U., Agarwal, M., Aggarwal, V., and Ukkusuri, S.V. (2021). On the approximation of cooperative heterogeneous multi-agent reinforcement learning (marl) using mean field control (mfc). arXiv."}],"container-title":["Algorithms"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1999-4893\/15\/3\/73\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,10]],"date-time":"2025-10-10T22:25:06Z","timestamp":1760135106000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1999-4893\/15\/3\/73"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,2,22]]},"references-count":26,"journal-issue":{"issue":"3","published-online":{"date-parts":[[2022,3]]}},"alternative-id":["a15030073"],"URL":"https:\/\/doi.org\/10.3390\/a15030073","relation":{},"ISSN":["1999-4893"],"issn-type":[{"value":"1999-4893","type":"electronic"}],"subject":[],"published":{"date-parts":[[2022,2,22]]}}}