{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,10,10]],"date-time":"2025-10-10T01:23:55Z","timestamp":1760059435100,"version":"build-2065373602"},"reference-count":52,"publisher":"MDPI AG","issue":"6","license":[{"start":{"date-parts":[[2025,6,13]],"date-time":"2025-06-13T00:00:00Z","timestamp":1749772800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Algorithms"],"abstract":"<jats:p>Games have long been benchmarks for AI algorithms and, with the boost of computational power and the application of new algorithms, AI systems have achieved superhuman performance in games for which it was once thought that they could only be mastered by humans due to their high complexity [...]<\/jats:p>","DOI":"10.3390\/a18060363","type":"journal-article","created":{"date-parts":[[2025,6,13]],"date-time":"2025-06-13T06:19:28Z","timestamp":1749795568000},"page":"363","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":0,"title":["Algorithms for Game AI"],"prefix":"10.3390","volume":"18","author":[{"given":"Wenxin","family":"Li","sequence":"first","affiliation":[{"name":"School of Computer Science, Peking University, Beijing 100871, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Haifeng","family":"Zhang","sequence":"additional","affiliation":[{"name":"Institute of Automation, Chinese Academy of Sciences, Beijing 100190, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2025,6,13]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","first-page":"57","DOI":"10.1016\/S0004-3702(01)00129-1","article-title":"Deep blue","volume":"134","author":"Campbell","year":"2002","journal-title":"Artif. Intell."},{"key":"ref_2","doi-asserted-by":"crossref","first-page":"484","DOI":"10.1038\/nature16961","article-title":"Mastering the game of Go with deep neural networks and tree search","volume":"529","author":"Silver","year":"2016","journal-title":"Nature"},{"key":"ref_3","doi-asserted-by":"crossref","first-page":"354","DOI":"10.1038\/nature24270","article-title":"Mastering the game of go without human knowledge","volume":"550","author":"Silver","year":"2017","journal-title":"Nature"},{"key":"ref_4","doi-asserted-by":"crossref","first-page":"1140","DOI":"10.1126\/science.aar6404","article-title":"A general reinforcement learning algorithm that masters chess, shogi, and Go through self-play","volume":"362","author":"Silver","year":"2018","journal-title":"Science"},{"key":"ref_5","doi-asserted-by":"crossref","first-page":"508","DOI":"10.1126\/science.aam6960","article-title":"Deepstack: Expert-level artificial intelligence in heads-up no-limit poker","volume":"356","author":"Schmid","year":"2017","journal-title":"Science"},{"key":"ref_6","doi-asserted-by":"crossref","first-page":"418","DOI":"10.1126\/science.aao1733","article-title":"Superhuman AI for heads-up no-limit poker: Libratus beats top professionals","volume":"359","author":"Brown","year":"2018","journal-title":"Science"},{"key":"ref_7","doi-asserted-by":"crossref","first-page":"885","DOI":"10.1126\/science.aay2400","article-title":"Superhuman AI for multiplayer poker","volume":"365","author":"Brown","year":"2019","journal-title":"Science"},{"key":"ref_8","first-page":"20","article-title":"Alphastar: Mastering the real-time strategy game starcraft ii","volume":"2","author":"Vinyals","year":"2019","journal-title":"Deep. Blog"},{"key":"ref_9","unstructured":"Berner, C., Brockman, G., Chan, B., Cheung, V., D\u0119biak, P., Dennison, C., Farhi, D., Fischer, Q., Hashme, S., and Hesse, C. (2019). Dota 2 with large scale deep reinforcement learning. arXiv."},{"key":"ref_10","first-page":"621","article-title":"Towards playing full moba games with deep reinforcement learning","volume":"33","author":"Ye","year":"2020","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_11","doi-asserted-by":"crossref","unstructured":"Xia, B., Ye, X., and Abuassba, A.O. (2020, January 15\u201319). Recent research on ai in games. Proceedings of the 2020 International Wireless Communications and Mobile Computing (IWCMC), Limassol, Cyprus.","DOI":"10.1109\/IWCMC48107.2020.9148327"},{"key":"ref_12","doi-asserted-by":"crossref","first-page":"100","DOI":"10.1109\/TSSC.1968.300136","article-title":"A formal basis for the heuristic determination of minimum cost paths","volume":"4","author":"Hart","year":"1968","journal-title":"IEEE Trans. Syst. Sci. Cybern."},{"key":"ref_13","doi-asserted-by":"crossref","first-page":"179","DOI":"10.1016\/0004-3702(79)90016-X","article-title":"A minimax algorithm better than alpha-beta?","volume":"12","author":"Stockman","year":"1979","journal-title":"Artif. Intell."},{"key":"ref_14","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1109\/TCIAIG.2012.2186810","article-title":"A survey of monte carlo tree search methods","volume":"4","author":"Browne","year":"2012","journal-title":"IEEE Trans. Comput. Intell. AI Games"},{"key":"ref_15","doi-asserted-by":"crossref","unstructured":"Kocsis, L., and Szepesv\u00e1ri, C. (2006, January 18\u201322). Bandit based monte-carlo planning. Proceedings of the European Conference on Machine Learning, Berlin, Germany.","DOI":"10.1007\/11871842_29"},{"key":"ref_16","doi-asserted-by":"crossref","unstructured":"Gelly, S., and Silver, D. (2007, January 20\u201324). Combining online and offline knowledge in UCT. Proceedings of the 24th International Conference on Machine Learning, Corvallis, OR, USA.","DOI":"10.1145\/1273496.1273531"},{"key":"ref_17","doi-asserted-by":"crossref","unstructured":"Chaslot, G.M.B., Winands, M.H., and van Den Herik, H.J. (October, January 29). Parallel monte-carlo tree search. Proceedings of the Computers and Games: 6th International Conference, CG 2008, Proceedings 6, Beijing, China.","DOI":"10.1007\/978-3-540-87608-3_6"},{"key":"ref_18","unstructured":"Rechenberg, I. (October, January 29). Evolutionsstrategien. Proceedings of the Simulationsmethoden in der Medizin und Biologie: Workshop, Hannover, Germany."},{"key":"ref_19","unstructured":"Rusu, A.A., Colmenarejo, S.G., Gulcehre, C., Desjardins, G., Kirkpatrick, J., Pascanu, R., Mnih, V., Kavukcuoglu, K., and Hadsell, R. (2015). Policy distillation. arXiv."},{"key":"ref_20","doi-asserted-by":"crossref","unstructured":"Sutton, R.S., and Barto, A.G. (1998). Reinforcement Learning: An Introduction, MIT Press.","DOI":"10.1109\/TNN.1998.712192"},{"key":"ref_21","unstructured":"Mnih, V., Kavukcuoglu, K., Silver, D., Graves, A., Antonoglou, I., Wierstra, D., and Riedmiller, M. (2013). Playing atari with deep reinforcement learning. arXiv."},{"key":"ref_22","doi-asserted-by":"crossref","first-page":"229","DOI":"10.1007\/BF00992696","article-title":"Simple statistical gradient-following algorithms for connectionist reinforcement learning","volume":"8","author":"Williams","year":"1992","journal-title":"Mach. Learn."},{"key":"ref_23","unstructured":"Konda, V., and Tsitsiklis, J. (December, January 29). Actor-critic algorithms. Proceedings of the Neural Information Processing Systems, Denver, CO, USA."},{"key":"ref_24","unstructured":"Lillicrap, T.P., Hunt, J.J., Pritzel, A., Heess, N., Erez, T., Tassa, Y., Silver, D., and Wierstra, D. (2015). Continuous control with deep reinforcement learning. arXiv."},{"key":"ref_25","unstructured":"Mnih, V., Badia, A.P., Mirza, M., Graves, A., Lillicrap, T.P., Harley, T., Silver, D., and Kavukcuoglu, K. (2006, January 19\u201324). Asynchronous methods for deep reinforcement learning. Proceedings of the International Conference on Machine Learning, PmLR, New York, NY, USA."},{"key":"ref_26","unstructured":"Espeholt, L., Soyer, H., Munos, R., Simonyan, K., Mnih, V., Ward, T., Doron, Y., Firoiu, V., Harley, T., and Dunning, I. (2018, January 10\u201315). Impala: Scalable distributed deep-rl with importance weighted actor-learner architectures. Proceedings of the International Conference on Machine Learning, PMLR, Stockholm, Sweden."},{"key":"ref_27","unstructured":"Schulman, J., Levine, S., Abbeel, P., Jordan, M., and Moritz, P. (2015, January 6\u201311). Trust region policy optimization. Proceedings of the International Conference on Machine Learning, PMLR, Lille, France."},{"key":"ref_28","unstructured":"Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O. (2017). Proximal policy optimization algorithms. arXiv."},{"key":"ref_29","unstructured":"Zinkevich, M., Johanson, M., Bowling, M., and Piccione, C. (2007, January 3\u20136). Regret minimization in games with incomplete information. Proceedings of the Neural Information Processing Systems, Vancouver, BC, Canada."},{"key":"ref_30","unstructured":"Brown, N., and Sandholm, T. (February, January 27). Solving imperfect-information games via discounted regret minimization. Proceedings of the AAAI Conference on Artificial Intelligence, Honolulu, HI, USA."},{"key":"ref_31","unstructured":"Lanctot, M., Waugh, K., Zinkevich, M., and Bowling, M. (2009, January 7\u201310). Monte Carlo sampling for regret minimization in extensive games. Proceedings of the Neural Information Processing Systems, Vancouver, BC, Canada."},{"key":"ref_32","unstructured":"Brown, N., Lerer, A., Gross, S., and Sandholm, T. (2019, January 9\u201315). Deep counterfactual regret minimization. Proceedings of the International Conference on Machine Learning, PMLR, Long Beach, CA, USA."},{"key":"ref_33","unstructured":"Steinberger, E., Lerer, A., and Brown, N. (2020). Dream: Deep regret minimization with advantage baselines and model-free learning. arXiv."},{"key":"ref_34","unstructured":"Heinrich, J., Lanctot, M., and Silver, D. (2015, January 6\u201311). Fictitious self-play in extensive-form games. Proceedings of the International Conference on Machine Learning, PMLR, Lille, France."},{"key":"ref_35","unstructured":"Heinrich, J., and Silver, D. (2016). Deep reinforcement learning from self-play in imperfect-information games. arXiv."},{"key":"ref_36","unstructured":"McMahan, H.B., Gordon, G.J., and Blum, A. (2003, January 21\u201324). Planning in the presence of cost functions controlled by an adversary. Proceedings of the 20th International Conference on Machine Learning (ICML-03), Washington, DC, USA."},{"key":"ref_37","unstructured":"Lanctot, M., Zambaldi, V., Gruslys, A., Lazaridou, A., Tuyls, K., Perolat, J., Silver, D., and Graepel, T. (2017, January 4\u20139). A unified game-theoretic approach to multiagent reinforcement learning. Proceedings of the Neural Information Processing Systems, Long Beach, CA, USA."},{"key":"ref_38","unstructured":"Sunehag, P., Lever, G., Gruslys, A., Czarnecki, W.M., Zambaldi, V., Jaderberg, M., Lanctot, M., Sonnerat, N., Leibo, J.Z., and Tuyls, K. (2017). Value-decomposition networks for cooperative multi-agent learning. arXiv."},{"key":"ref_39","first-page":"1","article-title":"Monotonic value function factorisation for deep multi-agent reinforcement learning","volume":"21","author":"Rashid","year":"2020","journal-title":"J. Mach. Learn. Res."},{"key":"ref_40","unstructured":"Lowe, R., Wu, Y.I., Tamar, A., Harb, J., Pieter Abbeel, O., and Mordatch, I. (2017, January 4\u20139). Multi-agent actor-critic for mixed cooperative-competitive environments. Proceedings of the Neural Information Processing Systems, Long Beach, CA, USA."},{"key":"ref_41","doi-asserted-by":"crossref","unstructured":"Wang, W., Sun, D., Jiang, F., Chen, X., and Zhu, C. (2022). Research and Challenges of Reinforcement Learning in Cyber Defense Decision-Making for Intranet Security. Algorithms, 15.","DOI":"10.3390\/a15040134"},{"key":"ref_42","doi-asserted-by":"crossref","unstructured":"Sanjaya, R., Wang, J., and Yang, Y. (2022). Measuring the Non-Transitivity in Chess. Algorithms, 15.","DOI":"10.3390\/a15050152"},{"key":"ref_43","doi-asserted-by":"crossref","unstructured":"Yang, X., Wang, Z., Zhang, H., Ma, N., Yang, N., Liu, H., Zhang, H., and Yang, L. (2022). A Review: Machine Learning for Combinatorial Optimization Problems in Energy Areas. Algorithms, 15.","DOI":"10.3390\/a15060205"},{"key":"ref_44","doi-asserted-by":"crossref","unstructured":"Lu, Y., and Li, W. (2022). Techniques and Paradigms in Modern Game AI Systems. Algorithms, 15.","DOI":"10.3390\/a15080282"},{"key":"ref_45","doi-asserted-by":"crossref","unstructured":"Lu, Y., Li, W., and Li, W. (2023). Official International Mahjong: A New Playground for AI Research. Algorithms, 16.","DOI":"10.3390\/a16050235"},{"key":"ref_46","doi-asserted-by":"crossref","unstructured":"Li, Z., Chen, X., Fu, J., Xie, N., and Zhao, T. (2024). Reducing Q-Value Estimation Bias via Mutual Estimation and Softmax Operation in MADRL. Algorithms, 17.","DOI":"10.3390\/a17010036"},{"key":"ref_47","doi-asserted-by":"crossref","unstructured":"Li, J., Xie, N., and Zhao, T. (2024). Optimizing Reinforcement Learning Using a Generative Action-Translator Transformer. Algorithms, 17.","DOI":"10.3390\/a17010037"},{"key":"ref_48","doi-asserted-by":"crossref","unstructured":"Schaa, H., and Barriga, N.A. (2024). Evaluating the Expressive Range of Super Mario Bros Level Generators. Algorithms, 17.","DOI":"10.3390\/a17070307"},{"key":"ref_49","doi-asserted-by":"crossref","unstructured":"Zhang, L., Zou, H., and Zhu, Y. (2024). An Efficient Optimization of the Monte Carlo Tree Search Algorithm for Amazons. Algorithms, 17.","DOI":"10.3390\/a17080334"},{"key":"ref_50","doi-asserted-by":"crossref","unstructured":"Hsieh, Y.-H., Kao, C.-C., and Yuan, S.-M. (2025). Imitating Human Go Players via Vision Transformer. Algorithms, 18.","DOI":"10.3390\/a18020061"},{"key":"ref_51","doi-asserted-by":"crossref","unstructured":"Penelas, G., Barbosa, L., Reis, A., Barroso, J., and Pinto, T. (2025). Machine Learning for Decision Support and Automation in Games: A Study on Vehicle Optimal Path. Algorithms, 18.","DOI":"10.3390\/a18020106"},{"key":"ref_52","unstructured":"Haarnoja, T., Zhou, A., Abbeel, P., and Levine, S. (2018, January 10\u201315). Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor. Proceedings of the 35th International Conference on Machine Learning, Stockholmsm\u00e4ssan, Stockholm, Sweden."}],"container-title":["Algorithms"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1999-4893\/18\/6\/363\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,9]],"date-time":"2025-10-09T17:51:17Z","timestamp":1760032277000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1999-4893\/18\/6\/363"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,6,13]]},"references-count":52,"journal-issue":{"issue":"6","published-online":{"date-parts":[[2025,6]]}},"alternative-id":["a18060363"],"URL":"https:\/\/doi.org\/10.3390\/a18060363","relation":{},"ISSN":["1999-4893"],"issn-type":[{"type":"electronic","value":"1999-4893"}],"subject":[],"published":{"date-parts":[[2025,6,13]]}}}