{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,1]],"date-time":"2026-07-01T06:25:21Z","timestamp":1782887121494,"version":"3.54.5"},"reference-count":48,"publisher":"Springer Science and Business Media LLC","issue":"1","license":[{"start":{"date-parts":[[2024,6,24]],"date-time":"2024-06-24T00:00:00Z","timestamp":1719187200000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2024,6,24]],"date-time":"2024-06-24T00:00:00Z","timestamp":1719187200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"DOI":"10.13039\/501100007958","name":"Wenzhou University","doi-asserted-by":"publisher","award":["Young Faculty Grant"],"award-info":[{"award-number":["Young Faculty Grant"]}],"id":[{"id":"10.13039\/501100007958","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/100008582","name":"McGill University","doi-asserted-by":"publisher","id":[{"id":"10.13039\/100008582","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Int J Comput Intell Syst"],"abstract":"<jats:title>Abstract<\/jats:title><jats:p>This paper proposes a gradient-based multi-agent actor-critic algorithm for off-policy reinforcement learning using importance sampling. Our algorithm is incremental with full gradients, and its complexity per iteration scales linearly with the size of approximation features. Previous multi-agent actor-critic algorithms are limited to the on-policy setting or off-policy emphatic temporal difference (TD) learning and they do not take advantage of the advances in off-policy gradient temporal difference learning (GTD). As a theoretical contribution, we establish that the critic step of the proposed algorithm converges to the TD solution of the projected Bellman equation and the actor step converges to the set of asymptotically stable fixed points. Numerical experiments on the multi-agent generalization of the Boyan\u2019s chain problem show that the proposed approach provides improved performances in terms of stability and convergence rate as compared with the state-of-the-art baseline algorithm.<\/jats:p>","DOI":"10.1007\/s44196-024-00560-2","type":"journal-article","created":{"date-parts":[[2024,6,24]],"date-time":"2024-06-24T15:01:42Z","timestamp":1719241302000},"update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":4,"title":["Multi-agent Gradient-Based Off-Policy Actor-Critic Algorithm for Distributed Reinforcement Learning"],"prefix":"10.1007","volume":"17","author":[{"ORCID":"https:\/\/orcid.org\/0000-0003-0174-9366","authenticated-orcid":false,"given":"Jineng","family":"Ren","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"297","published-online":{"date-parts":[[2024,6,24]]},"reference":[{"key":"560_CR1","doi-asserted-by":"crossref","unstructured":"Gupta, J.K., Egorov, M., Kochenderfer, M.: Cooperative multi-agent control using deep reinforcement learning. In: Proc. Int. Conf. Autonomous Agents and Multiagent Systems, pp. 66\u201383 (2017)","DOI":"10.1007\/978-3-319-71682-4_5"},{"issue":"19","key":"560_CR2","doi-asserted-by":"publisher","first-page":"70","DOI":"10.2352\/ISSN.2470-1173.2017.19.AVM-023","volume":"2017","author":"AE Sallab","year":"2017","unstructured":"Sallab, A.E., Abdou, M., Perot, E., Yogamani, S.: Deep reinforcement learning framework for autonomous driving. Electron. Imaging 2017(19), 70\u201376 (2017)","journal-title":"Electron. Imaging"},{"issue":"7676","key":"560_CR3","doi-asserted-by":"publisher","first-page":"354","DOI":"10.1038\/nature24270","volume":"550","author":"D Silver","year":"2017","unstructured":"Silver, D., Schrittwieser, J., Simonyan, K., Antonoglou, I., Huang, A., Guez, A., Hubert, T., Baker, L., Lai, M., Bolton, A., et al.: Mastering the game of go without human knowledge. Nature 550(7676), 354\u2013359 (2017)","journal-title":"Nature"},{"key":"560_CR4","doi-asserted-by":"crossref","unstructured":"Zhang, K., Yang, Z., Basar, T.: Networked multi-agent reinforcement learning in continuous spaces. In: Proc. IEEE Conf. Decision and Control, pp. 2771\u20132776 (2018)","DOI":"10.1109\/CDC.2018.8619581"},{"key":"560_CR5","doi-asserted-by":"crossref","unstructured":"Zhang, K., Yang, Z., Liu, H., Zhang, T., Basar, T.: Fully decentralized multi-agent reinforcement learning with networked agents. In: Proc. Int. Conf. Machine Learning, pp. 5872\u20135881 (2018)","DOI":"10.1109\/CDC.2018.8619581"},{"key":"560_CR6","doi-asserted-by":"crossref","unstructured":"Lin, Y., Zhang, K., Yang, Z., Wang, Z., Ba\u015far, T., Sandhu, R., Liu, J.: A communication-efficient multi-agent actor-critic algorithm for distributed reinforcement learning. Proc. IEEE Conf. Decision and Control, 5562\u20135567 (2019)","DOI":"10.1109\/CDC40024.2019.9029257"},{"key":"560_CR7","doi-asserted-by":"crossref","unstructured":"Lin, Y., Gade, S., Sandhu, R., Liu, J.: Toward resilient multi-agent actor-critic algorithms for distributed reinforcement learning. Proc. Am. Control Conf., 3953\u20133958 (2020)","DOI":"10.23919\/ACC45564.2020.9147381"},{"key":"560_CR8","unstructured":"Fakoor, R., Chaudhari, P., Smola, A.J.: P3O: Policy-on policy-off policy optimization. In: Proc. Uncertainty in Artificial Intelligence, pp. 1017\u20131027 (2020)"},{"key":"560_CR9","doi-asserted-by":"publisher","first-page":"610","DOI":"10.1016\/j.neunet.2023.11.046","volume":"170","author":"D Guo","year":"2024","unstructured":"Guo, D., Tang, L., Zhang, X., Liang, Y.-C.: An off-policy multi-agent stochastic policy gradient algorithm for cooperative continuous control. Neural Netw 170, 610\u2013621 (2024)","journal-title":"Neural Netw"},{"key":"560_CR10","doi-asserted-by":"crossref","unstructured":"Zhang, Y., Zavlanos, M.M.: Distributed off-policy actor-critic reinforcement learning with policy consensus. Proc. IEEE Conf. Decision and Control, 4674\u20134679 (2019)","DOI":"10.1109\/CDC40024.2019.9029969"},{"key":"560_CR11","unstructured":"Wang, Y., Han, B., Wang, T., Dong, H., Zhang, C.: Off-policy multi-agent decomposed policy gradients (2020). arXiv preprint arXiv:2007.12322"},{"key":"560_CR12","doi-asserted-by":"publisher","DOI":"10.1016\/j.automatica.2020.109081","volume":"119","author":"C Chen","year":"2020","unstructured":"Chen, C., Lewis, F.L., Xie, K., Xie, S., Liu, Y.: Off-policy learning for adaptive optimal output synchronization of heterogeneous multi-agent systems. Automatica 119, 109081 (2020)","journal-title":"Automatica"},{"key":"560_CR13","doi-asserted-by":"crossref","unstructured":"Stankovi\u0107, M.S., Stankovi\u0107, S.S.: Multi-agent temporal-difference learning with linear function approximation: Weak convergence under time-varying network topologies. Proc. Amer. Control Conf., 167\u2013172 (2016)","DOI":"10.1109\/ACC.2016.7524910"},{"key":"560_CR14","doi-asserted-by":"crossref","unstructured":"Sutton, R.S., Maei, H.R., Precup, D., Bhatnagar, S., Silver, D., Szepesv\u00e1ri, C., Wiewiora, E.: Fast gradient-descent methods for temporal-difference learning with linear function approximation. Proc. Int. Conf. Machine Learning, pp. 993\u20131000 (2009)","DOI":"10.1145\/1553374.1553501"},{"issue":"178","key":"560_CR15","first-page":"1","volume":"24","author":"W Li","year":"2023","unstructured":"Li, W., Jin, B., Wang, X., Yan, J., Zha, H.: F2a2: flexible fully-decentralized approximate actor-critic for cooperative multi-agent reinforcement learning. J. Mach. Learn. Res. 24(178), 1\u201375 (2023)","journal-title":"J. Mach. Learn. Res."},{"key":"560_CR16","unstructured":"Maei, H.R.: Gradient temporal-difference learning algorithms, PhD thesis. University of Alberta (2011)"},{"key":"560_CR17","unstructured":"Maei, H.R., Szepesvari, C., Bhatnagar, S., Precup, D., Silver, D., Sutton, R.S.: Convergent temporal-difference learning with arbitrary smooth function approximation. Proc. Adv. Neur. Inf. Proc. Systems, pp. 1204\u20131212 (2009)"},{"issue":"1","key":"560_CR18","first-page":"2603","volume":"17","author":"RS Sutton","year":"2016","unstructured":"Sutton, R.S., Mahmood, A.R., White, M.: An emphatic approach to the problem of off-policy temporal-difference learning. J. Mach. Learn. Res. 17(1), 2603\u20132631 (2016)","journal-title":"J. Mach. Learn. Res."},{"key":"560_CR19","unstructured":"Degris, T., White, M., Sutton, R.S.: Off-policy actor-critic. Proc. Int. Conf. Mach. Learn., pp. 179\u2013186 (2012)"},{"key":"560_CR20","unstructured":"Ghiassian, S., Patterson, A., Garg, S., Gupta, D., White, A., White, M.: Gradient temporal-difference learning with regularized corrections. Proc. Int. Conf. Machine Learning, pp. 3524\u20133534 (2020)"},{"key":"560_CR21","unstructured":"Imani, E., Graves, E., White, M.: An off-policy policy gradient theorem using emphatic weightings. Proc. Adv. Neur. Inf. Proc. Syst. 96\u2013106 (2018)"},{"key":"560_CR22","unstructured":"Zhang, S., Liu, B., Yao, H., Whiteson, S.: Provably convergent two-timescale off-policy actor-critic with function approximation. Proc. Int. Conf. Mach. Learn., 11204\u201311213 (2020)"},{"key":"560_CR23","unstructured":"Maei, H.R.: Convergent actor-critic algorithms under off-policy training and function approximation (2018). arXiv preprint arXiv:1802.07842"},{"key":"560_CR24","first-page":"1549","volume":"53","author":"W Suttle","year":"2020","unstructured":"Suttle, W., Yang, Z., Zhang, K., Wang, Z., Ba\u015far, T., Liu, J.: A multi-agent off-policy actor-critic algorithm for distributed reinforcement learning. Proc. Int. Fed. Autom. Control 53, 1549\u20131554 (2020)","journal-title":"Proc. Int. Fed. Autom. Control"},{"key":"560_CR25","doi-asserted-by":"crossref","unstructured":"Stankovi\u0107, M.S., Beko, M., Stankovi\u0107, S.S.: Convergent distributed actor-critic algorithm based on gradient temporal difference. In: Proc. European Signal Proc. Conf., pp. 2066\u20132070 (2022)","DOI":"10.23919\/EUSIPCO55093.2022.9909762"},{"key":"560_CR26","doi-asserted-by":"crossref","unstructured":"Chen, Z., Zhou, Y., Chen, R.: Multi-agent off-policy TD learning: finite-time analysis with near-optimal sample complexity and communication complexity (2021). arXiv preprint arXiv:2103.13147","DOI":"10.1109\/IEEECONF53345.2021.9723200"},{"key":"560_CR27","doi-asserted-by":"crossref","unstructured":"Baird, L.: Residual Algorithms: Reinforcement Learning with Function Approximation. In: Machine Learning Proceedings, pp. 30\u201337 (1995)","DOI":"10.1016\/B978-1-55860-377-6.50013-X"},{"key":"560_CR28","unstructured":"Sutton, R.S., Barto, A.G.: Reinforcement Learning: An Introduction. MIT press (2018)"},{"key":"560_CR29","doi-asserted-by":"publisher","first-page":"101280","DOI":"10.1016\/j.jocs.2020.101280","volume":"49","author":"J Ren","year":"2021","unstructured":"Ren, J., Haupt, J., Guo, Z.: Communication-efficient hierarchical distributed optimization for multi-agent policy evaluation. J. Comput. Sci. 49, 101280 (2021)","journal-title":"J. Comput. Sci."},{"key":"560_CR30","doi-asserted-by":"crossref","unstructured":"Wei, Y., Zheng, R.: Multi-robot path planning for mobile sensing through deep reinforcement learning. Proc. Int. Conf. Comp. Comm., 1\u201310 (2021)","DOI":"10.1109\/INFOCOM42981.2021.9488669"},{"key":"560_CR31","doi-asserted-by":"crossref","unstructured":"Perrusqu\u00eda, A., Yu, W., Li, X.: Multi-agent reinforcement learning for redundant robot control in task-space. Int. J. Mach. Learn. Cybern. 12(1), 231\u2013241 (2021)","DOI":"10.1007\/s13042-020-01167-7"},{"key":"560_CR32","doi-asserted-by":"publisher","first-page":"977","DOI":"10.1016\/j.energy.2018.04.042","volume":"153","author":"L Xi","year":"2018","unstructured":"Xi, L., Chen, J., Huang, Y., Xu, Y., Liu, L., Zhou, Y., Li, Y.: Smart generation control based on multi-agent reinforcement learning with the idea of the time tunnel. Energy 153, 977\u2013987 (2018)","journal-title":"Energy"},{"issue":"3","key":"560_CR33","doi-asserted-by":"publisher","first-page":"1464","DOI":"10.1109\/TSG.2013.2248175","volume":"4","author":"E Dall\u2019Anese","year":"2013","unstructured":"Dall\u2019Anese, E., Zhu, H., Giannakis, G.B.: Distributed optimal power flow for smart microgrids. IEEE Trans. Smart Grid 4(3), 1464\u20131475 (2013)","journal-title":"IEEE Trans. Smart Grid"},{"issue":"5","key":"560_CR34","doi-asserted-by":"publisher","first-page":"674","DOI":"10.1109\/9.580874","volume":"42","author":"JN Tsitsiklis","year":"1997","unstructured":"Tsitsiklis, J.N., Van Roy, B.: An analysis of temporal-difference learning with function approximation. IEEE Trans. Autom. Control 42(5), 674\u2013690 (1997)","journal-title":"IEEE Trans. Autom. Control"},{"issue":"21","key":"560_CR35","first-page":"1609","volume":"21","author":"RS Sutton","year":"2008","unstructured":"Sutton, R.S., Szepesv\u00e1ri, C., Maei, H.R.: A convergent o (n) algorithm for off-policy temporal-difference learning with linear function approximation. Proc. Adv. Neur. Inf. Proc. Syst. 21(21), 1609\u20131616 (2008)","journal-title":"Proc. Adv. Neur. Inf. Proc. Syst."},{"key":"560_CR36","first-page":"1057","volume":"12","author":"RS Sutton","year":"2000","unstructured":"Sutton, R.S., McAllester, D., Singh, S., Mansour, Y.: Policy gradient methods for reinforcement learning with function approximation. Proc. Adv. Neur. Inf. Proc. Syst. 12, 1057\u20131063 (2000)","journal-title":"Proc. Adv. Neur. Inf. Proc. Syst."},{"issue":"4","key":"560_CR37","doi-asserted-by":"publisher","first-page":"1070","DOI":"10.1109\/TAC.2014.2352691","volume":"60","author":"JM Hendrickx","year":"2014","unstructured":"Hendrickx, J.M., Shi, G., Johansson, K.H.: Finite-time consensus using stochastic matrices with positive diagonals. IEEE Trans. Autom. Control 60(4), 1070\u20131073 (2014)","journal-title":"IEEE Trans. Autom. Control"},{"issue":"2","key":"560_CR38","doi-asserted-by":"publisher","first-page":"933","DOI":"10.1007\/s12555-015-0407-2","volume":"15","author":"Y Shang","year":"2017","unstructured":"Shang, Y.: Finite-time cluster average consensus for networks via distributed iterations. Int. J. Control Autom. Syst. 15(2), 933\u2013938 (2017)","journal-title":"Int. J. Control Autom. Syst."},{"issue":"1","key":"560_CR39","doi-asserted-by":"publisher","first-page":"75","DOI":"10.1016\/j.scient.2011.03.010","volume":"18","author":"H Sayyaadi","year":"2011","unstructured":"Sayyaadi, H., Doostmohammadian, M.: Finite-time consensus in directed switching network topologies and time-delayed communications. Sci. Iran. 18(1), 75\u201385 (2011)","journal-title":"Sci. Iran."},{"key":"560_CR40","unstructured":"Borkar, V.S.: Stochastic approximation: a dynamical systems viewpoint. Springer 48 (2009)"},{"key":"560_CR41","unstructured":"Yu, H.: On convergence of some gradient-based temporal-differences algorithms for off-policy learning (2017). arXiv preprint arXiv:1712.09652"},{"issue":"1","key":"560_CR42","doi-asserted-by":"publisher","first-page":"48","DOI":"10.1109\/TAC.2008.2009515","volume":"54","author":"A Nedic","year":"2009","unstructured":"Nedic, A., Ozdaglar, A.: Distributed subgradient methods for multi-agent optimization. IEEE Trans. Autom. Control 54(1), 48\u201361 (2009)","journal-title":"IEEE Trans. Autom. Control"},{"issue":"6","key":"560_CR43","doi-asserted-by":"publisher","first-page":"2508","DOI":"10.1109\/TIT.2006.874516","volume":"52","author":"S Boyd","year":"2006","unstructured":"Boyd, S., Ghosh, A., Prabhakar, B., Shah, D.: Randomized gossip algorithms. IEEE Trans. Inform. Theory 52(6), 2508\u20132530 (2006)","journal-title":"IEEE Trans. Inform. Theory"},{"issue":"7","key":"560_CR44","doi-asserted-by":"publisher","first-page":"2748","DOI":"10.1109\/TSP.2009.2016247","volume":"57","author":"TC Aysal","year":"2009","unstructured":"Aysal, T.C., Yildiz, M.E., Sarwate, A.D., Scaglione, A.: Broadcast gossip algorithms for consensus. IEEE Trans. Signal Process. 57(7), 2748\u20132761 (2009)","journal-title":"IEEE Trans. Signal Process."},{"issue":"11","key":"560_CR45","doi-asserted-by":"publisher","first-page":"2471","DOI":"10.1016\/j.automatica.2009.07.008","volume":"45","author":"S Bhatnagar","year":"2009","unstructured":"Bhatnagar, S., Sutton, R.S., Ghavamzadeh, M., Lee, M.: Natural actor-critic algorithms. Automatica 45(11), 2471\u20132482 (2009)","journal-title":"Automatica"},{"key":"560_CR46","unstructured":"Kushner, H., Yin, G.G.: Stochastic approximation and recursive algorithms and applications. Springer Sci. Bus. Media 35 (2003)"},{"key":"560_CR47","unstructured":"Kushner, H.J., Clark, D.S.: Stochastic approximation methods for constrained and unconstrained systems. Springer Sci. Bus. Media 26 (2012)"},{"issue":"2","key":"560_CR48","doi-asserted-by":"publisher","first-page":"233","DOI":"10.1023\/A:1017936530646","volume":"49","author":"JA Boyan","year":"2002","unstructured":"Boyan, J.A.: Technical update: least-squares temporal difference learning. Mach. Learn. 49(2), 233\u2013246 (2002)","journal-title":"Mach. Learn."}],"container-title":["International Journal of Computational Intelligence Systems"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s44196-024-00560-2.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s44196-024-00560-2\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s44196-024-00560-2.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2024,6,24]],"date-time":"2024-06-24T15:26:08Z","timestamp":1719242768000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s44196-024-00560-2"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,6,24]]},"references-count":48,"journal-issue":{"issue":"1","published-online":{"date-parts":[[2024,12]]}},"alternative-id":["560"],"URL":"https:\/\/doi.org\/10.1007\/s44196-024-00560-2","relation":{},"ISSN":["1875-6883"],"issn-type":[{"value":"1875-6883","type":"electronic"}],"subject":[],"published":{"date-parts":[[2024,6,24]]},"assertion":[{"value":"9 January 2024","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"6 June 2024","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"24 June 2024","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"The author reports no conflict of interest.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Conflict of interest"}}],"article-number":"162"}}