{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,10,9]],"date-time":"2025-10-09T17:11:02Z","timestamp":1760029862302,"version":"build-2065373602"},"reference-count":44,"publisher":"MDPI AG","issue":"2","license":[{"start":{"date-parts":[[2025,2,7]],"date-time":"2025-02-07T00:00:00Z","timestamp":1738886400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Symmetry"],"abstract":"<jats:p>Self-play methods have achieved remarkable success in two-player zero-sum games, attaining superhuman performance in many complex game domains. Parallelizing learners is a feasible approach to handle complex games. However, parallelizing learners often leads to the suboptimal exploitation of computational resources, resulting in inefficiencies. This paper introduces the Mixed Hierarchical Oracle (MHO), which is designed to enhance training efficiency and performance in complex two-player zero-sum games. MHO efficiently leverages interaction data among parallelized solvers during the Parallelized Oracle (PO) process, while employing Model Soups (MS) to consolidate fragmented computational resources and Hierarchical Exploration (HE) to balance exploration and exploitation. These carefully designed enhancements for parallelized systems significantly improve the training performance of self-play. Additionally, MiniStar is introduced as an open source environment focused on small-scale combat scenarios, developed to facilitate research in self-play algorithms. The MHO is evaluated on both the AlphaStar888 matrix game and MiniStar environment, and ablation studies further demonstrates its effectiveness in improving the agent\u2019s decision-making capabilities. This work highlight the potential of the MHO to optimize compute resource utilization and improve performance in self-play methods.<\/jats:p>","DOI":"10.3390\/sym17020250","type":"journal-article","created":{"date-parts":[[2025,2,7]],"date-time":"2025-02-07T05:04:33Z","timestamp":1738904673000},"page":"250","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":0,"title":["Efficient Parallel Design for Self-Play in Two-Player Zero-Sum Games"],"prefix":"10.3390","volume":"17","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-1366-3749","authenticated-orcid":false,"given":"Hongsong","family":"Tang","sequence":"first","affiliation":[{"name":"School of Science, Beijing University of Posts and Telecommunications, Beijing 100876, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Bo","family":"Chen","sequence":"additional","affiliation":[{"name":"School of Artificial Intelligence, Beijing University of Posts and Telecommunications, Beijing 100876, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yingzhuo","family":"Liu","sequence":"additional","affiliation":[{"name":"School of Artificial Intelligence, Beijing University of Posts and Telecommunications, Beijing 100876, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Kuoye","family":"Han","sequence":"additional","affiliation":[{"name":"Information Science Academy (ISA), China Electronics Technology Group Corporation (CETC), Beijing 100043, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0008-8949-9691","authenticated-orcid":false,"given":"Jingqian","family":"Liu","sequence":"additional","affiliation":[{"name":"Chinatelecom Group Corporation, Beijing 100032, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Zhaowei","family":"Qu","sequence":"additional","affiliation":[{"name":"School of Science, Beijing University of Posts and Telecommunications, Beijing 100876, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2025,2,7]]},"reference":[{"key":"ref_1","unstructured":"Albrecht, S.V., Christianos, F., and Sch\u00e4fer, L. (2024). Multi-Agent Reinforcement Learning: Foundations and Modern Approaches, MIT Press."},{"key":"ref_2","unstructured":"Mahajan, A., Rashid, T., Samvelyan, M., and Whiteson, S. (2019, January 8\u201314). Maven: Multi-agent variational exploration. Proceedings of the 33rd Conference on Neural Information Processing Systems (NeurIPS 2019), Vancouver, BC, Canada."},{"key":"ref_3","first-page":"1","article-title":"Monotonic value function factorisation for deep multi-agent reinforcement learning","volume":"21","author":"Rashid","year":"2020","journal-title":"J. Mach. Learn. Res."},{"key":"ref_4","doi-asserted-by":"crossref","first-page":"210","DOI":"10.1147\/rd.33.0210","article-title":"Some studies in machine learning using the game of checkers","volume":"3","author":"Samuel","year":"1959","journal-title":"IBM J. Res. Dev."},{"key":"ref_5","unstructured":"Bansal, T., Pachocki, J., Sidor, S., Sutskever, I., and Mordatch, I. (2017). Emergent complexity via multi-agent competition. arXiv."},{"key":"ref_6","doi-asserted-by":"crossref","first-page":"484","DOI":"10.1038\/nature16961","article-title":"Mastering the game of Go with deep neural networks and tree search","volume":"529","author":"Silver","year":"2016","journal-title":"Nature"},{"key":"ref_7","doi-asserted-by":"crossref","first-page":"354","DOI":"10.1038\/nature24270","article-title":"Mastering the game of go without human knowledge","volume":"550","author":"Silver","year":"2017","journal-title":"Nature"},{"key":"ref_8","unstructured":"Silver, D., Hubert, T., Schrittwieser, J., Antonoglou, I., Lai, M., Guez, A., Lanctot, M., Sifre, L., Kumaran, D., and Graepel, T. (2017). Mastering chess and shogi by self-play with a general reinforcement learning algorithm. arXiv."},{"key":"ref_9","doi-asserted-by":"crossref","first-page":"604","DOI":"10.1038\/s41586-020-03051-4","article-title":"Mastering atari, go, chess and shogi by planning with a learned model","volume":"588","author":"Schrittwieser","year":"2020","journal-title":"Nature"},{"key":"ref_10","doi-asserted-by":"crossref","first-page":"508","DOI":"10.1126\/science.aam6960","article-title":"Deepstack: Expert-level artificial intelligence in heads-up no-limit poker","volume":"356","author":"Schmid","year":"2017","journal-title":"Science"},{"key":"ref_11","unstructured":"Heinrich, J., Lanctot, M., and Silver, D. (2015, January 6\u201311). Fictitious self-play in extensive-form games. Proceedings of the International Conference on Machine Learning, Lille, France."},{"key":"ref_12","unstructured":"Berner, C., Brockman, G., Chan, B., Cheung, V., Debiak, P., Dennison, C., Farhi, D., Fischer, Q., Hashme, S., and Hesse, C. (2019). Dota 2 with large scale deep reinforcement learning. arXiv."},{"key":"ref_13","doi-asserted-by":"crossref","first-page":"350","DOI":"10.1038\/s41586-019-1724-z","article-title":"Grandmaster level in StarCraft II using multi-agent reinforcement learning","volume":"575","author":"Vinyals","year":"2019","journal-title":"Nature"},{"key":"ref_14","unstructured":"McMahan, H.B., Gordon, G.J., and Blum, A. (2003, January 21\u201324). Planning in the presence of cost functions controlled by an adversary. Proceedings of the 20th International Conference on Machine Learning (ICML-03), Washington, DC, USA."},{"key":"ref_15","unstructured":"Heinrich, J., and Silver, D. (2016). Deep reinforcement learning from self-play in imperfect-information games. arXiv."},{"key":"ref_16","doi-asserted-by":"crossref","unstructured":"Hernandez, D., Denamgana\u00ef, K., Gao, Y., York, P., Devlin, S., Samothrakis, S., and Walker, J.A. (2019, January 20\u201323). A generalized framework for self-play training. Proceedings of the 2019 IEEE Conference on Games (CoG), London, UK.","DOI":"10.1109\/CIG.2019.8848006"},{"key":"ref_17","unstructured":"Yang, Y., Luo, J., Wen, Y., Slumbers, O., Graves, D., Ammar, H.B., Wang, J., and Taylor, M.E. (2021). Diverse auto-curriculum is critical for successful real-world multiagent learning systems. arXiv."},{"key":"ref_18","unstructured":"Lanctot, M., Zambaldi, V., Gruslys, A., Lazaridou, A., Tuyls, K., P\u00e9rolat, J., Silver, D., and Graepel, T. (2017, January 4\u20139). A unified game-theoretic approach to multiagent reinforcement learning. Proceedings of the 31st Conference on Neural Information Processing Systems (NIPS 2017), Long Beach, CA, USA."},{"key":"ref_19","unstructured":"Sutton, R.S. (2018). Reinforcement learning: An introduction. A Bradford Book, MIT Press."},{"key":"ref_20","unstructured":"Wellman, M.P. (2006, January 16\u201320). Methods for empirical game-theoretic analysis. Proceedings of the AAAI, Boston, MA, USA."},{"key":"ref_21","unstructured":"Wellman, M.P., Tuyls, K., and Greenwald, A. (2024). Empirical game-theoretic analysis: A survey. arXiv."},{"key":"ref_22","doi-asserted-by":"crossref","unstructured":"Bighashdel, A., Wang, Y., McAleer, S., Savani, R., and Oliehoek, F.A. (2024). Policy Space Response Oracles: A Survey. arXiv.","DOI":"10.24963\/ijcai.2024\/880"},{"key":"ref_23","first-page":"20238","article-title":"Pipeline psro: A scalable approach for finding approximate nash equilibria in large games","volume":"33","author":"McAleer","year":"2020","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_24","unstructured":"Balduzzi, D., Garnelo, M., Bachrach, Y., Czarnecki, W., Perolat, J., Jaderberg, M., and Graepel, T. (2019, January 9\u201315). Open-ended learning in symmetric zero-sum games. Proceedings of the International Conference on Machine Learning, Long Beach, CA, USA."},{"key":"ref_25","unstructured":"Ellis, B., Cook, J., Moalla, S., Samvelyan, M., Sun, M., Mahajan, A., Foerster, J., and Whiteson, S. (2023, January 10\u201316). Smacv2: An improved benchmark for cooperative multi-agent reinforcement learning. Proceedings of the 37th Conference on Neural Information Processing Systems (NeurIPS 2023) Track on Datasets and Benchmarks, New Orleans, LA, USA."},{"key":"ref_26","unstructured":"Beck, J., Vuorio, R., Liu, E.Z., Xiong, Z., Zintgraf, L., Finn, C., and Whiteson, S. (2023). A survey of meta-reinforcement learning. arXiv."},{"key":"ref_27","unstructured":"Rutherford, A., Ellis, B., Gallici, M., Cook, J., Lupu, A., Ingvarsson, G., Willi, T., Khan, A., de Witt, C.S., and Souly, A. (2023). Jaxmarl: Multi-agent rl environments in jax. arXiv."},{"key":"ref_28","first-page":"1","article-title":"Heterogeneous-agent reinforcement learning","volume":"25","author":"Zhong","year":"2024","journal-title":"J. Mach. Learn. Res."},{"key":"ref_29","unstructured":"Espeholt, L., Soyer, H., Munos, R., Simonyan, K., Mnih, V., Ward, T., Doron, Y., Firoiu, V., Harley, T., and Dunning, I. (2018, January 10\u201315). Impala: Scalable distributed deep-rl with importance weighted actor-learner architectures. Proceedings of the International Conference on Machine Learning, Stockholm, Sweden."},{"key":"ref_30","unstructured":"Horgan, D., Quan, J., Budden, D., Barth-Maron, G., Hessel, M., Van Hasselt, H., and Silver, D. (2018). Distributed prioritized experience replay. arXiv."},{"key":"ref_31","first-page":"1","article-title":"MALib: A parallel framework for population-based multi-agent reinforcement learning","volume":"24","author":"Zhou","year":"2023","journal-title":"J. Mach. Learn. Res."},{"key":"ref_32","unstructured":"Vinyals, O., Ewalds, T., Bartunov, S., Georgiev, P., Vezhnevets, A.S., Yeo, M., Makhzani, A., K\u00fcttler, H., Agapiou, J., and Schrittwieser, J. (2017). Starcraft ii: A new challenge for reinforcement learning. arXiv."},{"key":"ref_33","doi-asserted-by":"crossref","unstructured":"Kurach, K., Raichuk, A., Sta\u0144czyk, P., Zajac, M., Bachem, O., Espeholt, L., Riquelme, C., Vincent, D., Michalski, M., and Bousquet, O. (2020, January 7\u201312). Google research football: A novel reinforcement learning environment. Proceedings of the AAAI Conference on Artificial Intelligence, New York, NY, USA.","DOI":"10.1609\/aaai.v34i04.5878"},{"key":"ref_34","first-page":"621","article-title":"Towards playing full moba games with deep reinforcement learning","volume":"33","author":"Ye","year":"2020","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_35","doi-asserted-by":"crossref","first-page":"908","DOI":"10.1109\/TNNLS.2020.3029475","article-title":"Supervised learning achieves human-level performance in moba games: A case study of honor of kings","volume":"33","author":"Ye","year":"2020","journal-title":"IEEE Trans. Neural Netw. Learn. Syst."},{"key":"ref_36","first-page":"11881","article-title":"Honor of kings arena: An environment for generalization in competitive reinforcement learning","volume":"35","author":"Wei","year":"2022","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_37","first-page":"21","article-title":"Imitation learning: A survey of learning methods","volume":"50","author":"Hussein","year":"2017","journal-title":"ACM Comput. Surv. (CSUR)"},{"key":"ref_38","unstructured":"Lin, F., Huang, S., Pearce, T., Chen, W., and Tu, W.W. (2023). Tizero: Mastering multi-agent football with curriculum learning and self-play. arXiv."},{"key":"ref_39","unstructured":"Huang, S., Chen, W., Zhang, L., Li, Z., Zhu, F., Ye, D., Chen, T., and Zhu, J. (2021). TiKick: Towards Playing Multi-agent Football Full Games from Single-agent Demonstrations. arXiv."},{"key":"ref_40","unstructured":"Samvelyan, M., Rashid, T., De Witt, C.S., Farquhar, G., Nardelli, N., Rudner, T.G., Hung, C.M., Torr, P.H., Foerster, J., and Whiteson, S. (2019). The starcraft multi-agent challenge. arXiv."},{"key":"ref_41","unstructured":"Fudenberg, D., and Tirole, J. (1991). Game Theory, MIT Press."},{"key":"ref_42","first-page":"374","article-title":"Iterative solution of games by fictitious play","volume":"13","author":"Brown","year":"1951","journal-title":"Act. Anal. Prod Alloc."},{"key":"ref_43","first-page":"17443","article-title":"Real world games look like spinning tops","volume":"33","author":"Czarnecki","year":"2020","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_44","unstructured":"Yu, C., Velu, A., Vinitsky, E., Wang, Y., Bayen, A., and Wu, Y. (2021). The surprising effectiveness of ppo in cooperative, multi-agent games. arXiv."}],"container-title":["Symmetry"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/2073-8994\/17\/2\/250\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,9]],"date-time":"2025-10-09T16:28:41Z","timestamp":1760027321000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/2073-8994\/17\/2\/250"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,2,7]]},"references-count":44,"journal-issue":{"issue":"2","published-online":{"date-parts":[[2025,2]]}},"alternative-id":["sym17020250"],"URL":"https:\/\/doi.org\/10.3390\/sym17020250","relation":{},"ISSN":["2073-8994"],"issn-type":[{"type":"electronic","value":"2073-8994"}],"subject":[],"published":{"date-parts":[[2025,2,7]]}}}