{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,10,11]],"date-time":"2025-10-11T00:37:18Z","timestamp":1760143038501,"version":"build-2065373602"},"reference-count":27,"publisher":"MDPI AG","issue":"1","license":[{"start":{"date-parts":[[2024,1,16]],"date-time":"2024-01-16T00:00:00Z","timestamp":1705363200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"DOI":"10.13039\/501100012166","name":"National Key R&amp;D Program of China","doi-asserted-by":"publisher","award":["2022YFB3104600","2019-YF08-00285-GX","61976156","SCITLAB-30005"],"award-info":[{"award-number":["2022YFB3104600","2019-YF08-00285-GX","61976156","SCITLAB-30005"]}],"id":[{"id":"10.13039\/501100012166","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100019014","name":"Chengdu Science and Technology Project","doi-asserted-by":"publisher","award":["2022YFB3104600","2019-YF08-00285-GX","61976156","SCITLAB-30005"],"award-info":[{"award-number":["2022YFB3104600","2019-YF08-00285-GX","61976156","SCITLAB-30005"]}],"id":[{"id":"10.13039\/501100019014","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["2022YFB3104600","2019-YF08-00285-GX","61976156","SCITLAB-30005"],"award-info":[{"award-number":["2022YFB3104600","2019-YF08-00285-GX","61976156","SCITLAB-30005"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]},{"name":"Intelligent Terminal Key Laboratory of SiChuan Province","award":["2022YFB3104600","2019-YF08-00285-GX","61976156","SCITLAB-30005"],"award-info":[{"award-number":["2022YFB3104600","2019-YF08-00285-GX","61976156","SCITLAB-30005"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Algorithms"],"abstract":"<jats:p>In recent years, with the rapid advancements in Natural Language Processing (NLP) technologies, large models have become widespread. Traditional reinforcement learning algorithms have also started experimenting with language models to optimize training. However, they still fundamentally rely on the Markov Decision Process (MDP) for reinforcement learning, and do not fully exploit the advantages of language models for dealing with long sequences of problems. The Decision Transformer (DT) introduced in 2021 is the initial effort to completely transform the reinforcement learning problem into a challenge within the NLP domain. It attempts to use text generation techniques to create reinforcement learning trajectories, addressing the issue of finding optimal trajectories. However, the article places the training trajectory data of reinforcement learning directly into a basic language model for training. Its aim is to predict the entire trajectory, encompassing state and reward information. This approach deviates from the reinforcement learning training objective of finding the optimal action. Furthermore, it generates redundant information in the output, impacting the final training effectiveness of the agent. This paper proposes a more reasonable network model structure, the Action-Translator Transformer (ATT), to predict only the next action of the agent. This makes the language model more interpretable for the reinforcement learning problem. We test our model in simulated gaming scenarios and compare it with current mainstream methods in the offline reinforcement learning field. Based on the presented experimental results, our model demonstrates superior performance. We hope that introducing this model will inspire new ideas and solutions for combining language models and reinforcement learning, providing fresh perspectives for offline reinforcement learning research.<\/jats:p>","DOI":"10.3390\/a17010037","type":"journal-article","created":{"date-parts":[[2024,1,16]],"date-time":"2024-01-16T04:03:30Z","timestamp":1705377810000},"page":"37","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":1,"title":["Optimizing Reinforcement Learning Using a Generative Action-Translator Transformer"],"prefix":"10.3390","volume":"17","author":[{"given":"Jiaming","family":"Li","sequence":"first","affiliation":[{"name":"Center for Future Media, School of Computer Science and Engineering, and Yibin Park, University of Electronic Science and Technology of China, Chengdu 611731, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Ning","family":"Xie","sequence":"additional","affiliation":[{"name":"Center for Future Media, School of Computer Science and Engineering, and Yibin Park, University of Electronic Science and Technology of China, Chengdu 611731, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Tingting","family":"Zhao","sequence":"additional","affiliation":[{"name":"College of Artificial Intelligence, Tianjin University of Science and Technology, Tianjin 300457, China"},{"name":"RIKEN Center for Advanced Intelligence Project (AIP), Tokyo 103-0027, Japan"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2024,1,16]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","unstructured":"Lu, Y., and Li, W. (2022). Techniques and Paradigms in Modern Game AI Systems. Algorithms, 15.","DOI":"10.3390\/a15080282"},{"key":"ref_2","doi-asserted-by":"crossref","unstructured":"Prudencio, R.F., Maximo, M.R.O.A., and Colombini, E.L. (2023). A survey on offline reinforcement learning: Taxonomy, review, and open problems. IEEE Trans. Neural Netw. Learn. Syst.","DOI":"10.1109\/TNNLS.2023.3250269"},{"key":"ref_3","unstructured":"Devlin, J., Chang, M.W., Lee, K., and Toutanova, K. (2019, January 2\u20137). Bert: Pre-training of deep bidirectional transformers for language understanding. Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Minneapolis, MN, USA."},{"key":"ref_4","unstructured":"Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J.D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., and Askell, A. (2020, January 6\u201312). Language models are few-shot learners. Proceedings of the Advances in Neural Information Processing Systems 33, Online."},{"key":"ref_5","unstructured":"OpenAI (2023). GPT-4 Technical Report. arXiv."},{"key":"ref_6","unstructured":"Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A.N., Kaiser, \u0141., and Polosukhin, I. (2017, January 4\u20139). Attention is all you need. Proceedings of the Advances in Neural Information Processing Systems 30, Los Angeles, CA, USA."},{"key":"ref_7","doi-asserted-by":"crossref","first-page":"1735","DOI":"10.1162\/neco.1997.9.8.1735","article-title":"Long short-term memory","volume":"9","author":"Hochreiter","year":"1997","journal-title":"Neural Comput."},{"key":"ref_8","doi-asserted-by":"crossref","unstructured":"Dey, R., and Salem, F.M. (2017, January 6\u20139). Gate-variants of gated recurrent unit (GRU) neural networks. Proceedings of the 2017 IEEE 60th International Midwest Symposium on Circuits And Systems (MWSCAS), Boston, MA, USA.","DOI":"10.1109\/MWSCAS.2017.8053243"},{"key":"ref_9","unstructured":"Parisotto, E., Song, F., Rae, J., Pascanu, R., Gulcehre, C., Jayakumar, S., Jaderberg, M., Kaufman, R.L., Clark, A., and Noury, S. (2020, January 13\u201318). Stabilizing transformers for reinforcement learning. Proceedings of the 37th International Conference on Machine Learning, Online."},{"key":"ref_10","unstructured":"Banino, A., Badia, A.P., Walker, J., Scholtes, T., Mitrovic, J., and Blundell, C. (2022, January 25\u201329). Coberl: Contrastive bert for reinforcement learning. Proceedings of the Tenth International Conference on Learning Representations (ICLR), Virtual Event."},{"key":"ref_11","unstructured":"Chen, L., Lu, K., Rajeswaran, A., Lee, K., Grover, A., Laskin, M., Abbeel, P., Srinivas, A., and Mordatch, I. (2021, January 6\u201314). Decision transformer: Reinforcement learning via sequence modeling. Proceedings of the Advances in Neural Information Processing Systems 34, Online."},{"key":"ref_12","first-page":"2493","article-title":"Natural language processing (almost) from scratch","volume":"12","author":"Collobert","year":"2011","journal-title":"J. Mach. Learn. Res."},{"key":"ref_13","unstructured":"Sutton, R.S., and Barto, A.G. (2018). Reinforcement Learning: An Introduction, MIT Press."},{"key":"ref_14","unstructured":"Mnih, V., Kavukcuoglu, K., Silver, D., Graves, A., Antonoglou, I., Wierstra, D., and Riedmiller, M. (2013, January 5\u201310). Playing atari with deep reinforcement learning. Proceedings of the Advances in Neural Information Processing Systems 26, Lake Tahoe, NV, USA."},{"key":"ref_15","unstructured":"Sutton, R.S., McAllester, D., Singh, S., and Mansour, Y. (December, January 29). Policy gradient methods for reinforcement learning with function approximation. Proceedings of the Advances in Neural Information Processing Systems 12, Denver, CO, USA."},{"key":"ref_16","unstructured":"Mnih, V., Badia, A.P., Mirza, M., Graves, A., Lillicrap, T., Harley, T., Silver, D., and Kavukcuoglu, K. (2016, January 20\u201322). Asynchronous methods for deep reinforcement learning. Proceedings of the 33rd International Conference on Machine Learning, New York, NY, USA."},{"key":"ref_17","unstructured":"Mishra, N., Rohaninejad, M., Chen, X., and Abbeel, P. (May, January 30). A simple neural attentive meta-learner. Proceedings of the 6th International Conference on Learning Representations (ICLR), Vancouver, BC, Canada."},{"key":"ref_18","unstructured":"Dai, Z., Yang, Z., Yang, Y., Carbonell, J., Le, Q.V., and Salakhutdinov, R. (August, January 28). Transformer-xl: Attentive language models beyond a fixed-length context. Proceedings of the 57th Conference of the Association for Computational Linguistics (ACL), Florence, Italy."},{"key":"ref_19","unstructured":"Kumar, S., Parker, J., and Naderian, P. (2020). Adaptive transformers in RL. arXiv."},{"key":"ref_20","doi-asserted-by":"crossref","unstructured":"Goodman, S., Ding, N., and Soricut, R. (2020, January 16\u201320). TeaForN: Teacher-forcing with n-grams. Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing, Online.","DOI":"10.18653\/v1\/2020.emnlp-main.702"},{"key":"ref_21","doi-asserted-by":"crossref","unstructured":"Todorov, E., Erez, T., and Tassa, Y. (2012, January 7\u201312). Mujoco: A physics engine for model-based control. Proceedings of the 2012 IEEE\/RSJ International Conference on Intelligent Robots and Systems, Vilamoura-Algarve, Portugal.","DOI":"10.1109\/IROS.2012.6386109"},{"key":"ref_22","unstructured":"Fu, J., Kumar, A., Nachum, O., Tucker, G., and Levine, S. (2020). D4rl: Datasets for deep data-driven reinforcement learning. arXiv."},{"key":"ref_23","unstructured":"Kumar, A., Zhou, A., Tucker, G., and Levine, S. (2020, January 6\u201312). Conservative q-learning for offline reinforcement learning. Proceedings of the Advances in Neural Information Processing Systems 33, Online."},{"key":"ref_24","unstructured":"Kumar, A., Fu, J., Tucker, G., and Levine, S. (2019, January 8\u201314). Stabilizing off-policy q-learning via bootstrapping error reduction. Proceedings of the Advances in Neural Information Processing Systems 32, Vancouver, BC, Canada."},{"key":"ref_25","unstructured":"Wu, Y., Tucker, G., and Nachum, O. (2021, January 17\u201319). Behavior regularized offline reinforcement learning. Proceedings of the Asian Conference on Machine Learning (ACML 2021), Virtual Event."},{"key":"ref_26","unstructured":"Peng, X., Kumar, A., Zhang, G., and Levine, S. (2019). Advantage-weighted regression: Simple and scalable off-policy reinforcement learning. arXiv."},{"key":"ref_27","doi-asserted-by":"crossref","unstructured":"Torabi, F., Warnell, G., and Stone, P. (2018, January 13\u201319). Behavioral cloning from observation. Proceedings of the 27th International Joint Conference on Artificial IntelligenceJuly 2018, Stockholm, Sweden.","DOI":"10.24963\/ijcai.2018\/687"}],"container-title":["Algorithms"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1999-4893\/17\/1\/37\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,10]],"date-time":"2025-10-10T13:47:47Z","timestamp":1760104067000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1999-4893\/17\/1\/37"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,1,16]]},"references-count":27,"journal-issue":{"issue":"1","published-online":{"date-parts":[[2024,1]]}},"alternative-id":["a17010037"],"URL":"https:\/\/doi.org\/10.3390\/a17010037","relation":{},"ISSN":["1999-4893"],"issn-type":[{"type":"electronic","value":"1999-4893"}],"subject":[],"published":{"date-parts":[[2024,1,16]]}}}