{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,25]],"date-time":"2026-06-25T16:26:21Z","timestamp":1782404781312,"version":"3.54.5"},"reference-count":50,"publisher":"Springer Science and Business Media LLC","issue":"13","license":[{"start":{"date-parts":[[2022,12,23]],"date-time":"2022-12-23T00:00:00Z","timestamp":1671753600000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2022,12,23]],"date-time":"2022-12-23T00:00:00Z","timestamp":1671753600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"name":"Strategic Priority Research Program of Chinese Academy of Science","award":["XDA27000000"],"award-info":[{"award-number":["XDA27000000"]}]},{"DOI":"10.13039\/501100012165","name":"Key Technologies Research and Development Program","doi-asserted-by":"publisher","award":["2021YFA1000403"],"award-info":[{"award-number":["2021YFA1000403"]}],"id":[{"id":"10.13039\/501100012165","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["11991022"],"award-info":[{"award-number":["11991022"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Appl Intell"],"published-print":{"date-parts":[[2023,7]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:p>Multiagent reinforcement learning (MARL) has been used extensively in the game environment. One of the main challenges in MARL is that the environment of the agent system is dynamic, and the other agents are also updating their strategies. Therefore, modeling the opponents\u2019 learning process and adopting specific strategies to shape learning is an effective way to obtain better training results. Previous studies such as DRON, LOLA and SOS approximated the opponent\u2019s learning process and gave effective applications. However, these studies modeled only transient changes in opponent strategies and lacked stability in the improvement of equilibrium efficiency. In this article, we design the MOL (modeling opponent learning) method based on the Stackelberg game. We use best response theory to approximate the opponents\u2019 preferences for different actions and explore stable equilibrium with higher rewards. We find that MOL achieves better results in several games with classical structures (the Prisoner\u2019s Dilemma, Stackelberg Leader game and Stag Hunt with 3 players), and in randomly generated bimatrix games. MOL performs well in competitive games played against different opponents and converges to stable points that score above the Nash equilibrium in repeated game environments. The results may provide a reference for the definition of equilibrium in multiagent reinforcement learning systems, and contribute to the design of learning objectives in MARL to avoid local disadvantageous equilibrium and improve general efficiency.<\/jats:p>","DOI":"10.1007\/s10489-022-04249-x","type":"journal-article","created":{"date-parts":[[2022,12,23]],"date-time":"2022-12-23T18:11:30Z","timestamp":1671819090000},"page":"17194-17210","update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":7,"title":["Modeling opponent learning in multiagent repeated games"],"prefix":"10.1007","volume":"53","author":[{"given":"Yudong","family":"Hu","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-3445-4620","authenticated-orcid":false,"given":"Congying","family":"Han","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Haoran","family":"Li","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Tiande","family":"Guo","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"297","published-online":{"date-parts":[[2022,12,23]]},"reference":[{"key":"4249_CR1","unstructured":"Fudenberg D, Levine DK (1998) The theory of learning in games. vol 1. MIT Press Books"},{"issue":"1","key":"4249_CR2","doi-asserted-by":"publisher","first-page":"82","DOI":"10.1016\/0899-8256(91)90006-Z","volume":"3","author":"P Milgrom","year":"1991","unstructured":"Milgrom P, Roberts J (1991) Adaptive and sophisticated learning in normal form games. Games Econom Behav 3(1):82\u2013100","journal-title":"Games Econom Behav"},{"issue":"6","key":"4249_CR3","doi-asserted-by":"publisher","first-page":"1255","DOI":"10.2307\/2938316","volume":"58","author":"P Milgrom","year":"1990","unstructured":"Milgrom P, Roberts J (1990) Rationalizability, learning, and equilibrium in games with strategic complementarities. Econometrica 58(6):1255\u20131277","journal-title":"Econometrica"},{"issue":"2","key":"4249_CR4","doi-asserted-by":"publisher","first-page":"165","DOI":"10.1006\/jeth.1999.2576","volume":"89","author":"E Dekel","year":"1999","unstructured":"Dekel E, Fudenberg D, Levine D (1999) Payoff information and self-confirming equilibrium. J Econ Theory 89(2):165\u2013185","journal-title":"J Econ Theory"},{"issue":"3","key":"4249_CR5","doi-asserted-by":"publisher","first-page":"523","DOI":"10.2307\/2951716","volume":"61","author":"D Fudenberg","year":"1993","unstructured":"Fudenberg D, Levine D (1993) Self-confirming equilibrium. Econometrica 61(3):523\u2013545","journal-title":"Econometrica"},{"issue":"2","key":"4249_CR6","doi-asserted-by":"publisher","first-page":"363","DOI":"10.1111\/1467-937X.00091","volume":"66","author":"K Binmore","year":"1999","unstructured":"Binmore K, Samuelson L (1999) Evolutionary drift and equilibrium selection. Rev Econ Stud 66(2):363\u2013393","journal-title":"Rev Econ Stud"},{"issue":"5","key":"4249_CR7","doi-asserted-by":"publisher","first-page":"3215","DOI":"10.1007\/s10462-020-09938-y","volume":"54","author":"W Du","year":"2021","unstructured":"Du W, Ding S (2021) A survey on multi-agent deep reinforcement learning: from the perspective of challenges and applications. Artif Intell Rev 54(5):3215\u20133238","journal-title":"Artif Intell Rev"},{"key":"4249_CR8","doi-asserted-by":"crossref","unstructured":"Gupta JK, Egorov M, Kochenderfer M (2017) Cooperative multi-agent control using deep reinforcement learning. In: International conference on autonomous agents and multiagent systems. Springer, Cham, pp 66\u201383","DOI":"10.1007\/978-3-319-71682-4_5"},{"key":"4249_CR9","unstructured":"Jiang J, Lu Z (2018) Learning attentional communication for multi-agent cooperation. In: Advances in neural information processing systems 31, pp 7265\u20137275"},{"key":"4249_CR10","doi-asserted-by":"crossref","unstructured":"Ge H, Ge Z, Sun L, et al. (2022) Enhancing cooperation by cognition differences and consistent representation in multi-agent reinforcement learning. Applied Intelligence","DOI":"10.1007\/s10489-021-02873-7"},{"key":"4249_CR11","doi-asserted-by":"crossref","unstructured":"Deng C, Wen C, Wang W, et al. (2022) Distributed adaptive tracking control for high-order nonlinear multi-agent systems over event-triggered communication. IEEE Transactions on Automatic Control","DOI":"10.1109\/TAC.2022.3148384"},{"key":"4249_CR12","unstructured":"Sunehag P, Lever G, Gruslys A, Czarnecki W, Zambaldi V, Jaderberg M, Lanctot M, Sonnerat N, Leibo J, Tuyls K, Graepe T (2017) Value-decomposition networks for cooperative multi-agent learning based on team reward. In: Proceedings of the 17th international conference on autonomous agents and multiagent systems"},{"key":"4249_CR13","unstructured":"Rashid T, Samvelyan M, Witt CD, Farquhar G, Foerster J, Whiteson S (2018) QMIX: monotonic value function factorisation for deep multi-agent reinforcement learning. In: Proceedings of the 35th international conference on machine learning, pp 4295\u20134304"},{"key":"4249_CR14","first-page":"10199","volume":"33","author":"T Rashid","year":"2020","unstructured":"Rashid T, Farquhar G, Peng B, et al. (2020) Weighted qmix: expanding monotonic value function factorisation for deep multi-agent reinforcement learning. Adv Neural Inform Process Syst 33:10199\u201310210","journal-title":"Adv Neural Inform Process Syst"},{"issue":"1","key":"4249_CR15","first-page":"7234","volume":"21","author":"T Rashid","year":"2020","unstructured":"Rashid T, Samvelyan M, De Witt CS, et al. (2020) Monotonic value function factorisation for deep multi-agent reinforcement learning. J Mach Learn Res 21(1):7234\u20137284","journal-title":"J Mach Learn Res"},{"key":"4249_CR16","doi-asserted-by":"publisher","first-page":"82","DOI":"10.1016\/j.neucom.2016.01.031","volume":"190","author":"L Kraemer","year":"2016","unstructured":"Kraemer L, Banerjee B (2016) Multi-agent reinforcement learning as a rehearsal for decentralized planning. Neurocomputing 190:82\u201394","journal-title":"Neurocomputing"},{"issue":"1","key":"4249_CR17","first-page":"7234","volume":"21","author":"T Rashid","year":"2020","unstructured":"Rashid T, Samvelyan M, De Witt CS, et al. (2020) Monotonic value function factorisation for deep multi-agent reinforcement learning. J Mach Learn Res 21(1):7234\u20137284","journal-title":"J Mach Learn Res"},{"issue":"6419","key":"4249_CR18","doi-asserted-by":"publisher","first-page":"1140","DOI":"10.1126\/science.aar6404","volume":"362","author":"D Silver","year":"2018","unstructured":"Silver D, Hubert T, Schrittwieser J, et al. (2018) A general reinforcement learning algorithm that masters chess, shogi, and Go through self-play. Science 362(6419):1140\u20131144","journal-title":"Science"},{"key":"4249_CR19","doi-asserted-by":"crossref","unstructured":"Brown N, Kroer C, Sandholm T (2017) Dynamic thresholding and pruning for regret minimization. In: Proceedings of the AAAI conference on artificial intelligence, vol 31(1)","DOI":"10.1609\/aaai.v31i1.10603"},{"issue":"6374","key":"4249_CR20","doi-asserted-by":"publisher","first-page":"418","DOI":"10.1126\/science.aao1733","volume":"359","author":"N Brown","year":"2018","unstructured":"Brown N, Sandholm T (2018) Superhuman AI for heads-up no-limit poker Libratus beats top professionals. Science 359(6374):418\u2013424","journal-title":"Science"},{"key":"4249_CR21","doi-asserted-by":"crossref","unstructured":"Jiang Q, Li K, Du B, Chen H, Fang H (2019) DeltaDou: expert-level Doudizhu AI through self-play. In: Proceedings of the twenty-eighth international joint conference on artificial intelligence, pp 1265\u20131271","DOI":"10.24963\/ijcai.2019\/176"},{"key":"4249_CR22","unstructured":"Zha D, Xie J, Ma W, Zhang S, Lian X, Hu X, Liu J (2021) DouZero: mastering DouDizhu with self-play deep reinforcement learning. In: Proceedings of the 38th international conference on machine learning, vol 139, pp 12333\u201312344"},{"issue":"1","key":"4249_CR23","first-page":"1582","volume":"17","author":"S Abdallah","year":"2016","unstructured":"Abdallah S, Kaisers M (2016) Addressing environment non-stationarity by repeating Q-learning updates. J Mach Learn Res 17(1):1582\u20131612","journal-title":"J Mach Learn Res"},{"key":"4249_CR24","unstructured":"Tang Z, Yu C, Chen B, Xu H, Wang X, Fang F, Du S, Wang Y, Wu Y (2021) Discovering diverse multi-agent strategic behavior via reward randomization. In: International conference on learning representations, pp 1\u201326"},{"key":"4249_CR25","unstructured":"He H, Boyd-Graber J, Kwok K, Daume H (2016) Opponent modeling in deep reinforcement learning. In: Proceedings of the 33rd international conference on machine learning, pp 1804\u20131813"},{"key":"4249_CR26","unstructured":"Foerster J, Chen R, Al-Shedivat M, Whiteson S, Abbeel P, Mordatch I (2018) Learning with opponent-learning awareness. In: Proceedings of the 17th international conference on autonomous agents and multiagent systems, pp 122\u2013130"},{"key":"4249_CR27","unstructured":"Sch\u00e4fer F, Anandkumar A (2019) Competitive gradient descent. Adv Neural Inf Process Syst, 32"},{"key":"4249_CR28","unstructured":"Willi T, Letcher A, Treutlein J, et al. (2022) COLA: consistent learning with opponent-learning awareness. In: International conference on machine learning. PMLR, pp 23804\u201323831"},{"issue":"1","key":"4249_CR29","doi-asserted-by":"publisher","first-page":"105","DOI":"10.1006\/game.1998.0687","volume":"28","author":"EV Damme","year":"1999","unstructured":"Damme EV, Hurkens S (1999) Endogenous Stackelberg leadership. Games Econ Behav 28 (1):105\u2013129","journal-title":"Games Econ Behav"},{"key":"4249_CR30","doi-asserted-by":"publisher","first-page":"510","DOI":"10.1016\/j.neucom.2020.06.066","volume":"411","author":"H Liu","year":"2020","unstructured":"Liu H, Wang X, Zhang W, et al. (2020) Infrared head pose estimation with multi-scales feature fusion on the IRHP database for human attention recognition. Neurocomputing 411:510\u2013520","journal-title":"Neurocomputing"},{"issue":"1","key":"4249_CR31","doi-asserted-by":"publisher","first-page":"544","DOI":"10.1109\/TII.2019.2934728","volume":"16","author":"T Liu","year":"2019","unstructured":"Liu T, Liu H, Li YF, et al. (2019) Flexible FTIR spectral imaging enhancement for industrial robot infrared vision sensing. IEEE Trans Industr Inform 16(1):544\u2013554","journal-title":"IEEE Trans Industr Inform"},{"key":"4249_CR32","doi-asserted-by":"publisher","first-page":"310","DOI":"10.1016\/j.neucom.2020.09.068","volume":"433","author":"H Liu","year":"2021","unstructured":"Liu H, Nie H, Zhang Z, et al. (2021) Anisotropic angle distribution learning for head pose estimation and attention understanding in human-computer interaction. Neurocomputing 433:310\u2013322","journal-title":"Neurocomputing"},{"key":"4249_CR33","unstructured":"Bowling M, Veloso M (2001) Rational and convergent learning in stochastic games. In: Proceedings of seventeenth international joint conference on artificial intelligence, pp 1021\u20131026"},{"issue":"1","key":"4249_CR34","first-page":"23","volume":"67","author":"V Conitzer","year":"2003","unstructured":"Conitzer V, Sandholm T (2003) AWESOME: a general multiagent learning algorithm that converges in self-play and learns a best response against stationary opponents. Mach Learn 67(1):23\u201343","journal-title":"Mach Learn"},{"key":"4249_CR35","unstructured":"Osband I, Blundell C, Pritzel A, et al. (2016) Deep exploration via bootstrapped DQN. Advances in Neural Information Processing Systems, 9"},{"key":"4249_CR36","doi-asserted-by":"publisher","first-page":"7107","DOI":"10.1109\/TII.2022.3143605","volume":"18","author":"H Liu","year":"2022","unstructured":"Liu H, Liu T, Zhang Z, et al. (2022) ARHPE: asymmetric relation-aware representation learning for head pose estimation in industrial human-machine interaction. IEEE Trans Industr Inform 18:7107\u20137117","journal-title":"IEEE Trans Industr Inform"},{"issue":"7","key":"4249_CR37","doi-asserted-by":"publisher","first-page":"4361","DOI":"10.1109\/TII.2021.3128240","volume":"18","author":"H Liu","year":"2021","unstructured":"Liu H, Zheng C, Li D, et al. (2021) EDMF: efficient deep matrix factorization with review feature learning for industrial recommender system. IEEE Trans Industr Inform 18(7):4361\u20134371","journal-title":"IEEE Trans Industr Inform"},{"key":"4249_CR38","doi-asserted-by":"publisher","first-page":"4434","DOI":"10.1007\/s10489-020-02034-2","volume":"51","author":"T Aotani","year":"2021","unstructured":"Aotani T, Kobayashi T, Sugimoto K (2021) Bottom-up multi-agent reinforcement learning by reward shaping for cooperative-competitive tasks. Appl Intell 51:4434\u20134452","journal-title":"Appl Intell"},{"key":"4249_CR39","unstructured":"Letcher A, Foerster J, Balduzzi D, Rocktaschel T, Whiteson S (2019) Stable opponent shaping in differentiable games. In: International conference on learning representations, pp 1\u201320"},{"key":"4249_CR40","doi-asserted-by":"crossref","unstructured":"Zhang C, Lesser V (2010) Multi-agent learning with policy prediction. In: Proceedings of the twenty-fourth AAAI conference on artificial intelligence, pp 927\u2013934","DOI":"10.1609\/aaai.v24i1.7639"},{"key":"4249_CR41","unstructured":"Wen Y, Chen H, Yang Y, Tian Z, Li M, Chen X, Wang J (2021) Multi-agent trust region learning. In: Proceedings of the seventh international conference on learning representations, pp 1\u201320"},{"key":"4249_CR42","unstructured":"Kim DK, Liu M, Riemer MD, et al. (2021) A policy gradient algorithm for learning to learn in multiagent reinforcement learning. In: International conference on machine learning PMLR, pp 5541\u20135550"},{"key":"4249_CR43","unstructured":"Raileanu R, Denton E, Szlam A, Fergus R (2018) Modeling others using oneself in multi-agent reinforcement learning. In: Proceedings of the 35th international conference on machine learning, pp 4257\u20134266"},{"issue":"2","key":"4249_CR44","first-page":"1015","volume":"51","author":"Z Zhen","year":"2019","unstructured":"Zhen Z, Yew-Soon D, Xue B (2019) Wang a collaborative multiagent reinforcement learning method based on policy gradient potential. IEEE Trans Cybern 51(2):1015\u20131027","journal-title":"IEEE Trans Cybern"},{"issue":"4","key":"4249_CR45","doi-asserted-by":"publisher","first-page":"647","DOI":"10.1109\/TCYB.2014.2332042","volume":"45","author":"Y Hu","year":"2015","unstructured":"Hu Y, Gao Y, An B (2015) Multiagent reinforcement learning with unshared value functions. IEEE Trans Cybern 45(4):647\u2013 662","journal-title":"IEEE Trans Cybern"},{"issue":"4","key":"4249_CR46","doi-asserted-by":"publisher","first-page":"861","DOI":"10.1111\/1468-0262.00223","volume":"69","author":"S Athey","year":"2001","unstructured":"Athey S (2001) Single crossing properties and the existence of pure strategy equilibria in games of incomplete information. Econometrica 69(4):861\u2013889","journal-title":"Econometrica"},{"key":"4249_CR47","unstructured":"Marris L, Muller P, Lanctot M, Tuyls K, Graepel T (2021) Multi-agent training beyond zero-sum with correlated equilibrium meta-solvers. In: Proceedings of the 38th international conference on machine learning, vol 139, pp 7480\u20137491"},{"key":"4249_CR48","doi-asserted-by":"publisher","first-page":"1002","DOI":"10.1007\/s10489-018-1307-y","volume":"49","author":"B Wang","year":"2019","unstructured":"Wang B, Zhang Y, Zhou ZH, et al. (2019) On repeated stackelberg security game with the cooperative human behavior model for wildlife protection. Appl Intell 49:1002\u20131015","journal-title":"Appl Intell"},{"issue":"2","key":"4249_CR49","doi-asserted-by":"publisher","first-page":"235","DOI":"10.1023\/A:1013689704352","volume":"47","author":"P Auer","year":"2002","unstructured":"Auer P, Cesa-Bianchi N, Fischer P (2002) Finite-time analysis of the multiarmed bandit problem. Mach Learn 47(2):235\u2013 256","journal-title":"Mach Learn"},{"key":"4249_CR50","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1007\/BF01737554","volume":"2","author":"J Harsanyi","year":"1973","unstructured":"Harsanyi J (1973) Games with randomly disturbed payoffs: a new rationale for mixed-strategy equilibrium points. Int J Game Theory 2:1\u201323","journal-title":"Int J Game Theory"}],"container-title":["Applied Intelligence"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10489-022-04249-x.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s10489-022-04249-x\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10489-022-04249-x.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2023,7,1]],"date-time":"2023-07-01T05:16:01Z","timestamp":1688188561000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s10489-022-04249-x"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,12,23]]},"references-count":50,"journal-issue":{"issue":"13","published-print":{"date-parts":[[2023,7]]}},"alternative-id":["4249"],"URL":"https:\/\/doi.org\/10.1007\/s10489-022-04249-x","relation":{},"ISSN":["0924-669X","1573-7497"],"issn-type":[{"value":"0924-669X","type":"print"},{"value":"1573-7497","type":"electronic"}],"subject":[],"published":{"date-parts":[[2022,12,23]]},"assertion":[{"value":"6 October 2022","order":1,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"23 December 2022","order":2,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}}]}}