{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,2]],"date-time":"2026-05-02T06:55:56Z","timestamp":1777704956460,"version":"3.51.4"},"reference-count":24,"publisher":"SAGE Publications","issue":"1","content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["IFS"],"published-print":{"date-parts":[[2021,8,11]]},"abstract":"<jats:p>A novel actor-critic algorithm is introduced and applied to zero-sum differential game. The proposed novel structure consists of two actors and a critic. Different actors represent the control policies of different players, and the critic is used to approximate the state-action utility function. Instead of neural network, the fuzzy inference system is applied as approximators for the actors and critic so that the specific practical meaning can be represented by the linguistic fuzzy rules. Since the goals of the players in the game are completely opposite, the actors for different players are simultaneously updated in opposite directions during the training. One actor is updated updated toward the direction that can minimize the Q value while the other updated toward the direction that can maximize the Q value. A pursuit-evasion problem with two pursuers and one evader is taken as an example to illustrate the validity of our method. In this problem, the two pursuers the same actor and the symmetry in the problem is used to improve the replay buffer. At the end of this paper, some confrontations between the policies with different training episodes are conducted.<\/jats:p>","DOI":"10.3233\/jifs-210032","type":"journal-article","created":{"date-parts":[[2021,6,1]],"date-time":"2021-06-01T14:41:43Z","timestamp":1622558503000},"page":"1069-1082","source":"Crossref","is-referenced-by-count":2,"title":["Minmax fuzzy deterministic policy gradient for zero-sum differential game: Take pursuit-evasion problem as example"],"prefix":"10.1177","volume":"41","author":[{"given":"Wei","family":"Liao","sequence":"first","affiliation":[{"name":"Key Laboratory of Fundamental Science for National Defense-Advanced Design Technology of Flight Vehicle, Nanjing University of Aeronautics and Astronautics, Nanjing, Jiangsu, China"},{"name":"State Key Laboratory of Mechanics and Control of Mechanical Structures, Nanjing University of Aeronautics and Astronautics, Nanjing, Jiangsu, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Xiaohui","family":"Wei","sequence":"additional","affiliation":[{"name":"Key Laboratory of Fundamental Science for National Defense-Advanced Design Technology of Flight Vehicle, Nanjing University of Aeronautics and Astronautics, Nanjing, Jiangsu, China"},{"name":"State Key Laboratory of Mechanics and Control of Mechanical Structures, Nanjing University of Aeronautics and Astronautics, Nanjing, Jiangsu, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Jizhou","family":"Lai","sequence":"additional","affiliation":[{"name":"College of Automation Engineering, Nanjing University of Aeronautics and Astronautics, Nanjing, Jiangsu, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"179","reference":[{"issue":"4","key":"10.3233\/JIFS-210032_ref2","doi-asserted-by":"crossref","first-page":"385","DOI":"10.1109\/TAC.1965.1098197","article-title":"Differential games and optimal pursuit-evasion strategies","volume":"10","author":"Ho","year":"1965","journal-title":"IEEE Transactions on Automatic Control"},{"key":"10.3233\/JIFS-210032_ref3","doi-asserted-by":"crossref","unstructured":"Liubarshchuk I. and Althoefer I. , The problem of approach in differential\u2013difference games, International Journal of Game Theory 45, 02 2015.","DOI":"10.1007\/s00182-015-0467-9"},{"issue":"4","key":"10.3233\/JIFS-210032_ref4","doi-asserted-by":"crossref","first-page":"851","DOI":"10.2514\/1.G003070","article-title":"Optimal evading strategies for two-pursuer\/one-evader problems","volume":"41","author":"Makkapati","year":"2018","journal-title":"Journal of Guidance, Control, and Dynamics"},{"key":"10.3233\/JIFS-210032_ref5","first-page":"3962","article-title":"A time-optimal control strategy for pursuit-evasion games problems, In","volume":"4","author":"Lim","year":"2004","journal-title":"IEEE International Conference on Robotics and Automation, 2004. Proceedings, ICRA \u201904. 2004"},{"key":"10.3233\/JIFS-210032_ref7","doi-asserted-by":"crossref","first-page":"101","DOI":"10.1016\/j.neucom.2020.06.031","article-title":"Cooperative control for multi-player pursuit-evasion games with reinforcement learning","volume":"412","author":"Wang","year":"2020","journal-title":"Neurocomputing"},{"key":"10.3233\/JIFS-210032_ref8","doi-asserted-by":"crossref","first-page":"105529","DOI":"10.1016\/j.ast.2019.105529","article-title":"Guidance strategies for interceptor against active defense spacecraft in two-on-two engagement","volume":"96","author":"Liang","year":"2020","journal-title":"Aerospace Science and Technology"},{"issue":"2","key":"10.3233\/JIFS-210032_ref9","doi-asserted-by":"crossref","first-page":"485","DOI":"10.1002\/jeab.587","article-title":"The dynamics of behavior: Review of sutton and barto: Reinforcement learning: An introduction (2nd ed.)","volume":"113","author":"Staddon","year":"2020","journal-title":"Journal of the Experimental Analysis of Behavior"},{"key":"10.3233\/JIFS-210032_ref10","doi-asserted-by":"crossref","first-page":"482","DOI":"10.1109\/ICMTMA.2009.213","article-title":"A pursuit-evasion algorithm based on hierarchical reinforcement learning","volume":"2","author":"Liu","year":"2009","journal-title":"Measuring Technology and Mechatronics Automation, International Conference on"},{"key":"10.3233\/JIFS-210032_ref11","first-page":"1","article-title":"Optimal and autonomous control using reinforcement learning: A survey","volume":"PP","author":"Kiumarsi","year":"2017","journal-title":"IEEE Transactions on Neural Networks and Learning Systems"},{"key":"10.3233\/JIFS-210032_ref12","first-page":"1","article-title":"A continuous-time markov decision process-based method with application in a pursuit-evasion example","volume":"46","author":"Jia","year":"2015","journal-title":"IEEE Transactions on Systems, Man, and Cybernetics: Systems"},{"issue":"Oct.14","key":"10.3233\/JIFS-210032_ref16","doi-asserted-by":"crossref","first-page":"106","DOI":"10.1016\/j.neucom.2019.07.038","article-title":"A fuzzy deterministic policy gradient algorithm for pursuit-evasion differential games","volume":"362","author":"Wang","year":"2019","journal-title":"Neurocomputing"},{"issue":"1","key":"10.3233\/JIFS-210032_ref17","doi-asserted-by":"crossref","first-page":"22","DOI":"10.1016\/j.robot.2010.09.006","article-title":"Self-learning fuzzy logic controllers for pursuit\u2013evasion differential games","volume":"59","author":"Desouky","year":"2011","journal-title":"Robotics and Autonomous Systems"},{"key":"10.3233\/JIFS-210032_ref18","doi-asserted-by":"crossref","first-page":"1058","DOI":"10.1007\/s40815-016-0284-8","article-title":"A residual gradient fuzzy reinforcement learning algorithm for differential games","volume":"19","author":"Awheda","year":"2017","journal-title":"International Journal of Fuzzy Systems"},{"key":"10.3233\/JIFS-210032_ref19","doi-asserted-by":"crossref","first-page":"443","DOI":"10.1016\/j.neucom.2018.11.072","article-title":"Hybrid hierarchical reinforcement learning for online guidance and navigation with partial observability","volume":"331","author":"Zhou","year":"2019","journal-title":"Neurocomputing"},{"key":"10.3233\/JIFS-210032_ref20","doi-asserted-by":"crossref","first-page":"1291","DOI":"10.1109\/TSMCC.2012.2218595","article-title":"A survey of actor-critic reinforcement learning: Standard and natural policy gradients","volume":"42","author":"Grondman","year":"2012","journal-title":"IEEE Transactions on Systems Man and Cybernetics Part B-Cybernetics"},{"key":"10.3233\/JIFS-210032_ref21","doi-asserted-by":"crossref","first-page":"878","DOI":"10.1016\/j.automatica.2010.02.018","article-title":"Online actor-critic algorithm to solve the continuous-time infinite horizon optimal control problem","volume":"46","author":"Vamvoudakis","year":"2010","journal-title":"Automatica"},{"key":"10.3233\/JIFS-210032_ref23","doi-asserted-by":"crossref","first-page":"73","DOI":"10.1016\/j.neucom.2017.02.051","article-title":"Neural-network-based synchronous iteration learning method for multi-player zero-sum games","volume":"242","author":"Song","year":"2017","journal-title":"Neurocomputing"},{"issue":"3","key":"10.3233\/JIFS-210032_ref24","doi-asserted-by":"crossref","first-page":"338","DOI":"10.1109\/5326.704563","article-title":"Fuzzy inference system learning by reinforcement methods","volume":"28","author":"Jouffe","year":"1998","journal-title":"Trans Sys Man Cyber Part C"},{"key":"10.3233\/JIFS-210032_ref25","doi-asserted-by":"crossref","first-page":"529","DOI":"10.1038\/nature14236","article-title":"Human-level control through deep reinforcement learning","volume":"518","author":"Mnih","year":"2015","journal-title":"Nature"},{"key":"10.3233\/JIFS-210032_ref27","doi-asserted-by":"crossref","first-page":"279","DOI":"10.1007\/BF00992698","article-title":"Q-learning","volume":"8","author":"Watkins","year":"1992","journal-title":"Mach Learn"},{"key":"10.3233\/JIFS-210032_ref30","first-page":"1057","article-title":"Policy gradient methods for reinforcement learning with function approximation","volume":"12","author":"Sutton","year":"2000","journal-title":"Adv Neural Inf Process Syst"},{"key":"10.3233\/JIFS-210032_ref31","doi-asserted-by":"crossref","first-page":"997","DOI":"10.1109\/72.623201","article-title":"Adaptive critic design","volume":"8","author":"Prokhorov","year":"1997","journal-title":"IEEE transactions on neural networks \/ a publication of the IEEE Neural Networks Council"},{"key":"10.3233\/JIFS-210032_ref32","doi-asserted-by":"crossref","first-page":"116","DOI":"10.1109\/TSMC.1985.6313399","article-title":"Fuzzy identification of systems and its applications to modeling and control","volume":"15","author":"TAKAGI","year":"1985","journal-title":"IEEE Trans. Systems, Man, and Cybernet"},{"issue":"10","key":"10.3233\/JIFS-210032_ref33","doi-asserted-by":"crossref","first-page":"5773","DOI":"10.1016\/j.jfranklin.2020.03.009","article-title":"A pursuit\u2013evasion game between two identical differential drive robots","volume":"357","author":"Bravo","year":"2020","journal-title":"Journal of the Franklin Institute"}],"container-title":["Journal of Intelligent &amp; Fuzzy Systems"],"original-title":[],"link":[{"URL":"https:\/\/content.iospress.com\/download?id=10.3233\/JIFS-210032","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,4,29]],"date-time":"2026-04-29T09:42:31Z","timestamp":1777455751000},"score":1,"resource":{"primary":{"URL":"https:\/\/journals.sagepub.com\/doi\/full\/10.3233\/JIFS-210032"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2021,8,11]]},"references-count":24,"journal-issue":{"issue":"1"},"URL":"https:\/\/doi.org\/10.3233\/jifs-210032","relation":{},"ISSN":["1064-1246","1875-8967"],"issn-type":[{"value":"1064-1246","type":"print"},{"value":"1875-8967","type":"electronic"}],"subject":[],"published":{"date-parts":[[2021,8,11]]}}}