{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,30]],"date-time":"2026-07-30T21:51:15Z","timestamp":1785448275897,"version":"3.56.0"},"reference-count":49,"publisher":"Institution of Engineering and Technology (IET)","issue":"3","license":[{"start":{"date-parts":[[2026,2,27]],"date-time":"2026-02-27T00:00:00Z","timestamp":1772150400000},"content-version":"vor","delay-in-days":0,"URL":"http:\/\/creativecommons.org\/licenses\/by\/4.0\/"},{"start":{"date-parts":[[2026,2,27]],"date-time":"2026-02-27T00:00:00Z","timestamp":1772150400000},"content-version":"tdm","delay-in-days":0,"URL":"http:\/\/doi.wiley.com\/10.1002\/tdm_license_1.1"}],"funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["62236011"],"award-info":[{"award-number":["62236011"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["62276285"],"award-info":[{"award-number":["62276285"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["ietresearch.onlinelibrary.wiley.com"],"crossmark-restriction":true},"short-container-title":["CAAI Trans on Intel Tech"],"published-print":{"date-parts":[[2026,6]]},"abstract":"<jats:title>ABSTRACT<\/jats:title>\n                  <jats:p>Large language models (LLMs) have made remarkable advances in natural language processing, demonstrating great potential in modelling structured sequences. However, adapting these capabilities to machine gaming tasks such as Go remains challenging due to limitations in strategy generalisation and optimisation efficiency. This paper presents multitype game optimisation (MyGO), a two\u2010stage fine\u2010tuning framework tailored for two\u2010player perfect information board games, exploring the applicability of LLMs to nonlinguistic decision\u2010making domains. In the supervised fine\u2010tuning stage, we propose a unified structural encoding method, action semantic unit (ASU), which efficiently converts heterogeneous game records into discrete token sequences compatible with LLMs. In the reinforcement learning stage, we design TA\u2010PPO (token\u2010level adaptive proximal policy optimisation), an enhanced PPO\u2010based algorithm to address the issue of sparse feedback commonly encountered in game reinforcement learning. Experimental results demonstrate that the fine\u2010tuned models achieve superior or comparable performance to traditional game\u2010playing algorithms in terms of strategy quality, rule generalisation and inference efficiency. This work provides a scalable paradigm for fine\u2010tuning LLMs in complex decision\u2010making tasks and lays a foundation for future research in game AI and generalisable strategy optimisation.<\/jats:p>","DOI":"10.1049\/cit2.70108","type":"journal-article","created":{"date-parts":[[2026,3,8]],"date-time":"2026-03-08T10:29:34Z","timestamp":1772965774000},"page":"739-753","update-policy":"https:\/\/doi.org\/10.1002\/crossmark_policy","source":"Crossref","is-referenced-by-count":7,"title":["Multitype Game Optimisation: A Two\u2010Stage Fine\u2010Tuning Framework for Multi\u2011Game Optimisation With Large Language Models"],"prefix":"10.1049","volume":"11","author":[{"ORCID":"https:\/\/orcid.org\/0000-0001-7950-6204","authenticated-orcid":false,"given":"Xiali","family":"Li","sequence":"first","affiliation":[{"name":"School of Information and Engineering Minzu University of China  Beijing China"},{"name":"Key Laboratory of Ethnic Language Intelligent Analysis and Security Governance of MOE Minzu University of China  Beijing China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Jingshi","family":"Gu","sequence":"additional","affiliation":[{"name":"School of Information and Engineering Minzu University of China  Beijing China"},{"name":"Key Laboratory of Ethnic Language Intelligent Analysis and Security Governance of MOE Minzu University of China  Beijing China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Feifan","family":"He","sequence":"additional","affiliation":[{"name":"School of Information and Engineering Minzu University of China  Beijing China"},{"name":"Key Laboratory of Ethnic Language Intelligent Analysis and Security Governance of MOE Minzu University of China  Beijing China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Yang","family":"Xiao","sequence":"additional","affiliation":[{"name":"School of Information and Engineering Minzu University of China  Beijing China"},{"name":"Key Laboratory of Ethnic Language Intelligent Analysis and Security Governance of MOE Minzu University of China  Beijing China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Yuanli","family":"Jia","sequence":"additional","affiliation":[{"name":"School of Information and Engineering Minzu University of China  Beijing China"},{"name":"Key Laboratory of Ethnic Language Intelligent Analysis and Security Governance of MOE Minzu University of China  Beijing China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Ping","family":"Lan","sequence":"additional","affiliation":[{"name":"School of Information Science and Technology Xizang University  Lhasa China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"265","published-online":{"date-parts":[[2026,2,27]]},"reference":[{"key":"e_1_2_9_2_1","article-title":"Gpt\u20104 Technical Report","author":"Achiam J.","year":"2023","journal-title":"arXiv preprint arXiv:2303.08774"},{"key":"e_1_2_9_3_1","article-title":"Deepseek\u2010v3 Technical Report","author":"Liu A.","year":"2024","journal-title":"arXiv preprint arXiv:2412.19437"},{"key":"e_1_2_9_4_1","doi-asserted-by":"publisher","DOI":"10.48550\/ARXIV.2205.15241"},{"key":"e_1_2_9_5_1","article-title":"The Chess Transformer: Mastering Play Using Generative Language Models","author":"Noever D.","year":"2020","journal-title":"arXiv preprint arXiv:2008.04057"},{"key":"e_1_2_9_6_1","doi-asserted-by":"publisher","DOI":"10.1109\/AI4I49448.2020.00012"},{"key":"e_1_2_9_7_1","doi-asserted-by":"publisher","DOI":"10.1126\/science.aar6404"},{"key":"e_1_2_9_8_1","doi-asserted-by":"publisher","DOI":"10.1038\/s41586\u2010020\u201003051\u20104"},{"key":"e_1_2_9_9_1","article-title":"Secrets of Rlhf in Large Language Models Part I: Ppo","author":"Zheng R.","year":"2023","journal-title":"arXiv preprint arXiv:2307.04964"},{"key":"e_1_2_9_10_1","article-title":"Dpo Meets Ppo: Reinforced Token Optimization for Rlhf","author":"Zhong H.","year":"2024","journal-title":"arXiv preprint arXiv:2404.18922"},{"key":"e_1_2_9_11_1","article-title":"Will gpt\u20104 Run Doom?","author":"de Wynter A.","year":"2024","journal-title":"IEEE Transactions on Games"},{"key":"e_1_2_9_12_1","doi-asserted-by":"publisher","DOI":"10.1049\/cit2.12298"},{"key":"e_1_2_9_13_1","article-title":"Word Play for Playing Othello (Reverses)","author":"Noever S. E. M.","year":"2022","journal-title":"arXiv preprint arXiv:2207.08766"},{"key":"e_1_2_9_14_1","unstructured":"B.Albert Gramaje \u201cExploring Gpt\u2019s Capabilities in Chess\u2010Puzzles \u201d2023."},{"key":"e_1_2_9_15_1","doi-asserted-by":"publisher","DOI":"10.48550\/ARXIV.2306.09200"},{"key":"e_1_2_9_16_1","article-title":"\u201cOptimizing Language Models for Chess: The Impact of Custom Notation and Elo\u2010Based Fine\u2010Tuning,\u201d ZHAW Z\u00fcrcher Hochschule F\u00fcr Angewandte Wissenschaften","author":"Schmid L.","year":"2024","journal-title":"Technical Reports Series"},{"key":"e_1_2_9_17_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-031-65572-2_7"},{"key":"e_1_2_9_18_1","article-title":"Complete Chess Games Enable Llm Become a Chess Master","author":"Zhang Y.","year":"2025","journal-title":"arXiv preprint arXiv:2501.17186"},{"key":"e_1_2_9_19_1","article-title":"Can Large Language Models Play Games? a Case Study of a Self\u2010Play Approach","author":"Guo H.","year":"2024","journal-title":"arXiv preprint arXiv:2403.05632"},{"key":"e_1_2_9_20_1","doi-asserted-by":"publisher","DOI":"10.26599\/tst.2020.9010017"},{"key":"e_1_2_9_21_1","article-title":"Playing Atari With Deep Reinforcement Learning","author":"Mnih V.","year":"2013","journal-title":"arXiv preprint arXiv:1312.5602"},{"key":"e_1_2_9_22_1","first-page":"1861","volume-title":"International Conference on Machine Learning","author":"Haarnoja T.","year":"2018"},{"key":"e_1_2_9_23_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICSPCC59353.2023.10400213"},{"key":"e_1_2_9_24_1","doi-asserted-by":"publisher","DOI":"10.1109\/PacificVis60374.2024.00051"},{"key":"e_1_2_9_25_1","doi-asserted-by":"publisher","DOI":"10.1117\/12.3031933"},{"key":"e_1_2_9_26_1","doi-asserted-by":"publisher","DOI":"10.1109\/access.2024.3416179"},{"key":"e_1_2_9_27_1","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v32i1.11796"},{"key":"e_1_2_9_28_1","article-title":"Relational Deep Reinforcement Learning","author":"Zambaldi V.","year":"2018","journal-title":"arXiv preprint arXiv:1806.01830"},{"key":"e_1_2_9_29_1","doi-asserted-by":"publisher","DOI":"10.1126\/science.aao1733"},{"key":"e_1_2_9_30_1","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.1705.02955"},{"key":"e_1_2_9_31_1","doi-asserted-by":"publisher","DOI":"10.1049\/cit2.12031"},{"key":"e_1_2_9_32_1","doi-asserted-by":"publisher","DOI":"10.1038\/s41586\u2010019\u20101724\u2010z"},{"key":"e_1_2_9_33_1","doi-asserted-by":"publisher","DOI":"10.1038\/nature16961"},{"key":"e_1_2_9_34_1","doi-asserted-by":"publisher","DOI":"10.1038\/nature24270"},{"key":"e_1_2_9_35_1","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.1707.01067"},{"key":"e_1_2_9_36_1","doi-asserted-by":"publisher","DOI":"10.1109\/access.2019.2938240"},{"key":"e_1_2_9_37_1","doi-asserted-by":"publisher","DOI":"10.1155\/2020\/4708075"},{"key":"e_1_2_9_38_1","volume-title":"Research and Implementation of Computer Game Algorithms for \u201cJiuqi\u201d","author":"Wang S.","year":"2023"},{"key":"e_1_2_9_39_1","doi-asserted-by":"publisher","DOI":"10.1109\/COMPSAC57700.2023.00059"},{"key":"e_1_2_9_40_1","article-title":"Policy Gradient Methods for Reinforcement Learning With Function Approximation","volume":"12","author":"Sutton R. S.","year":"1999","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_2_9_41_1","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2203.02155"},{"key":"e_1_2_9_42_1","article-title":"Proximal Policy Optimization Algorithms","author":"Schulman J.","year":"2017","journal-title":"arXiv preprint arXiv:1707.06347"},{"key":"e_1_2_9_43_1","unstructured":"J.Liu A.Cohen R.Pasunuru Y.Choi H.Hajishirzi andA.Celikyilmaz \u201cMaking Ppo Even Better: Value\u2010Guided Monte\u2010Carlo Tree Search Decoding \u201d2023."},{"key":"e_1_2_9_44_1","article-title":"Remax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models","author":"Li Z.","year":"2023","journal-title":"arXiv preprint arXiv:2310.10505"},{"key":"e_1_2_9_45_1","article-title":"What\u2019s Behind Ppo\u2019s Collapse in Long\u2010Cot? Value Optimization Holds the Secret","author":"Yuan Y.","year":"2025","journal-title":"arXiv preprint arXiv:2503.01491"},{"key":"e_1_2_9_46_1","doi-asserted-by":"publisher","DOI":"10.1145\/3689031.3696075"},{"key":"e_1_2_9_47_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.eswa.2025.130975"},{"key":"e_1_2_9_48_1","doi-asserted-by":"publisher","DOI":"10.1109\/tnnls.2025.3568019"},{"key":"e_1_2_9_49_1","doi-asserted-by":"publisher","DOI":"10.3969\/j.issn.1674-8425(z).2024.05.015"},{"key":"e_1_2_9_50_1","doi-asserted-by":"publisher","DOI":"10.1109\/COMPSAC61105.2024.00211"}],"container-title":["CAAI Transactions on Intelligence Technology"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/ietresearch.onlinelibrary.wiley.com\/doi\/pdf\/10.1049\/cit2.70108","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/ietresearch.onlinelibrary.wiley.com\/doi\/full-xml\/10.1049\/cit2.70108","content-type":"application\/xml","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/ietresearch.onlinelibrary.wiley.com\/doi\/pdf\/10.1049\/cit2.70108","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,7,13]],"date-time":"2026-07-13T07:25:35Z","timestamp":1783927535000},"score":1,"resource":{"primary":{"URL":"https:\/\/ietresearch.onlinelibrary.wiley.com\/doi\/10.1049\/cit2.70108"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,2,27]]},"references-count":49,"journal-issue":{"issue":"3","published-print":{"date-parts":[[2026,6]]}},"alternative-id":["10.1049\/cit2.70108"],"URL":"https:\/\/doi.org\/10.1049\/cit2.70108","archive":["Portico"],"relation":{},"ISSN":["2468-6557","2468-2322"],"issn-type":[{"value":"2468-6557","type":"print"},{"value":"2468-2322","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,2,27]]},"assertion":[{"value":"2025-07-26","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2026-01-27","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2026-02-27","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}