{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,3,25]],"date-time":"2026-03-25T10:53:47Z","timestamp":1774436027488,"version":"3.50.1"},"publisher-location":"New York, NY, USA","reference-count":56,"publisher":"ACM","license":[{"start":{"date-parts":[[2023,8,4]],"date-time":"2023-08-04T00:00:00Z","timestamp":1691107200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["61972223, U1936217, 61971267"],"award-info":[{"award-number":["61972223, U1936217, 61971267"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100017582","name":"Beijing National Research Center For Information Science And Technology","doi-asserted-by":"publisher","id":[{"id":"10.13039\/501100017582","id-type":"DOI","asserted-by":"publisher"}]},{"name":"The National Key Research and Development Program of China","award":["2020AAA0106000"],"award-info":[{"award-number":["2020AAA0106000"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2023,8,6]]},"DOI":"10.1145\/3580305.3599359","type":"proceedings-article","created":{"date-parts":[[2023,8,4]],"date-time":"2023-08-04T18:10:58Z","timestamp":1691172658000},"page":"685-697","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":18,"title":["GAT-MF: Graph Attention Mean Field for Very Large Scale Multi-Agent Reinforcement Learning"],"prefix":"10.1145","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-7109-3588","authenticated-orcid":false,"given":"Qianyue","family":"Hao","sequence":"first","affiliation":[{"name":"Tsinghua University, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-0454-7516","authenticated-orcid":false,"given":"Wenzhen","family":"Huang","sequence":"additional","affiliation":[{"name":"Tsinghua University, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-7341-0225","authenticated-orcid":false,"given":"Tao","family":"Feng","sequence":"additional","affiliation":[{"name":"Tsinghua University, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-9734-6056","authenticated-orcid":false,"given":"Jian","family":"Yuan","sequence":"additional","affiliation":[{"name":"Tsinghua University, Tsinghua University, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-5617-1659","authenticated-orcid":false,"given":"Yong","family":"Li","sequence":"additional","affiliation":[{"name":"Tsinghua University, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2023,8,4]]},"reference":[{"key":"e_1_3_2_2_1_1","first-page":"36","article-title":"Reliability-aware: task scheduling in cloud computing using multi-agent reinforcement learning algorithm and neural fitted Q","volume":"18","author":"Balla Husamelddin AM","year":"2021","unstructured":"Husamelddin AM Balla , Chen Guang Sheng , and Weipeng Jing . 2021 . Reliability-aware: task scheduling in cloud computing using multi-agent reinforcement learning algorithm and neural fitted Q . Int. Arab J. Inf. Technol. , Vol. 18 , 1 (2021), 36 -- 47 . Husamelddin AM Balla, Chen Guang Sheng, and Weipeng Jing. 2021. Reliability-aware: task scheduling in cloud computing using multi-agent reinforcement learning algorithm and neural fitted Q. Int. Arab J. Inf. Technol., Vol. 18, 1 (2021), 36--47.","journal-title":"Int. Arab J. Inf. Technol."},{"key":"e_1_3_2_2_2_1","volume-title":"Nature","volume":"599","author":"Bastani Hamsa","year":"2021","unstructured":"Hamsa Bastani , Kimon Drakopoulos , Vishal Gupta , Ioannis Vlachogiannis , Christos Hadjicristodoulou , Pagona Lagiou , Gkikas Magiorkinis , Dimitrios Paraskevis , and Sotirios Tsiodras . 2021 . Efficient and targeted COVID-19 border testing via reinforcement learning . Nature , Vol. 599 , 7883 (2021), 108--113. Hamsa Bastani, Kimon Drakopoulos, Vishal Gupta, Ioannis Vlachogiannis, Christos Hadjicristodoulou, Pagona Lagiou, Gkikas Magiorkinis, Dimitrios Paraskevis, and Sotirios Tsiodras. 2021. Efficient and targeted COVID-19 border testing via reinforcement learning. Nature, Vol. 599, 7883 (2021), 108--113."},{"key":"e_1_3_2_2_3_1","volume-title":"International Conference on Machine Learning. PMLR, 980--991","author":"B\u00f6hmer Wendelin","year":"2020","unstructured":"Wendelin B\u00f6hmer , Vitaly Kurin , and Shimon Whiteson . 2020 . Deep coordination graphs . In International Conference on Machine Learning. PMLR, 980--991 . Wendelin B\u00f6hmer, Vitaly Kurin, and Shimon Whiteson. 2020. Deep coordination graphs. In International Conference on Machine Learning. PMLR, 980--991."},{"key":"e_1_3_2_2_4_1","doi-asserted-by":"publisher","DOI":"10.1146\/annurev-control-042920-020211"},{"key":"e_1_3_2_2_5_1","volume-title":"Jaline Gerardin, Beth Redbird, David Grusky, and Jure Leskovec.","author":"Chang Serina","year":"2021","unstructured":"Serina Chang , Emma Pierson , Pang Wei Koh , Jaline Gerardin, Beth Redbird, David Grusky, and Jure Leskovec. 2021 a. Mobility network models of COVID-19 explain inequities and inform reopening. Nature , Vol. 589 , 7840 (2021), 82--87. Serina Chang, Emma Pierson, Pang Wei Koh, Jaline Gerardin, Beth Redbird, David Grusky, and Jure Leskovec. 2021a. Mobility network models of COVID-19 explain inequities and inform reopening. Nature, Vol. 589, 7840 (2021), 82--87."},{"key":"e_1_3_2_2_6_1","doi-asserted-by":"publisher","DOI":"10.1145\/3447548.3467182"},{"key":"e_1_3_2_2_7_1","doi-asserted-by":"publisher","DOI":"10.1109\/JIOT.2020.3008299"},{"key":"e_1_3_2_2_8_1","volume-title":"Strategic COVID-19 vaccine distribution can simultaneously elevate social utility and equity. Nature Human Behaviour","author":"Chen Lin","year":"2022","unstructured":"Lin Chen , Fengli Xu , Zhenyu Han , Kun Tang , Pan Hui , James Evans , and Yong Li. 2022b. Strategic COVID-19 vaccine distribution can simultaneously elevate social utility and equity. Nature Human Behaviour ( 2022 ), 1--12. Lin Chen, Fengli Xu, Zhenyu Han, Kun Tang, Pan Hui, James Evans, and Yong Li. 2022b. Strategic COVID-19 vaccine distribution can simultaneously elevate social utility and equity. Nature Human Behaviour (2022), 1--12."},{"key":"e_1_3_2_2_9_1","volume-title":"PTDE: Personalized Training with Distillated Execution for Multi-Agent Reinforcement Learning. arXiv preprint arXiv:2210.08872","author":"Chen Yiqun","year":"2022","unstructured":"Yiqun Chen , Hangyu Mao , Tianle Zhang , Shiguang Wu , Bin Zhang , Jianye Hao , Dong Li , Bin Wang , and Hongxing Chang . 2022 a. PTDE: Personalized Training with Distillated Execution for Multi-Agent Reinforcement Learning. arXiv preprint arXiv:2210.08872 (2022). Yiqun Chen, Hangyu Mao, Tianle Zhang, Shiguang Wu, Bin Zhang, Jianye Hao, Dong Li, Bin Wang, and Hongxing Chang. 2022a. PTDE: Personalized Training with Distillated Execution for Multi-Agent Reinforcement Learning. arXiv preprint arXiv:2210.08872 (2022)."},{"key":"e_1_3_2_2_10_1","volume-title":"Mingfei Sun, and Shimon Whiteson.","author":"de Witt Christian Schroeder","year":"2020","unstructured":"Christian Schroeder de Witt , Tarun Gupta , Denys Makoviichuk , Viktor Makoviychuk , Philip HS Torr , Mingfei Sun, and Shimon Whiteson. 2020 . Is independent learning all you need in the starcraft multi-agent challenge? arXiv preprint arXiv:2011.09533 (2020). Christian Schroeder de Witt, Tarun Gupta, Denys Makoviichuk, Viktor Makoviychuk, Philip HS Torr, Mingfei Sun, and Shimon Whiteson. 2020. Is independent learning all you need in the starcraft multi-agent challenge? arXiv preprint arXiv:2011.09533 (2020)."},{"key":"e_1_3_2_2_11_1","volume-title":"Diego de Las Casas, et al","author":"Degrave Jonas","year":"2022","unstructured":"Jonas Degrave , Federico Felici , Jonas Buchli , Michael Neunert , Brendan Tracey , Francesco Carpanese , Timo Ewalds , Roland Hafner , Abbas Abdolmaleki , Diego de Las Casas, et al . 2022 . Magnetic control of tokamak plasmas through deep reinforcement learning. Nature , Vol. 602 , 7897 (2022), 414--419. Jonas Degrave, Federico Felici, Jonas Buchli, Michael Neunert, Brendan Tracey, Francesco Carpanese, Timo Ewalds, Roland Hafner, Abbas Abdolmaleki, Diego de Las Casas, et al. 2022. Magnetic control of tokamak plasmas through deep reinforcement learning. Nature, Vol. 602, 7897 (2022), 414--419."},{"key":"e_1_3_2_2_12_1","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v32i1.11794"},{"key":"e_1_3_2_2_13_1","volume-title":"Proceedings of the Nineteenth International Conference (ICML 2002","author":"Guestrin Carlos","year":"2002","unstructured":"Carlos Guestrin , Michail G. Lagoudakis , and Ronald Parr . 2002 . Coordinated Reinforcement Learning. In Machine Learning , Proceedings of the Nineteenth International Conference (ICML 2002 ), University of New South Wales, Sydney, Australia , July 8-12, 2002, Claude Sammut and Achim G. Hoffmann (Eds.). Morgan Kaufmann, 227--234. Carlos Guestrin, Michail G. Lagoudakis, and Ronald Parr. 2002. Coordinated Reinforcement Learning. In Machine Learning, Proceedings of the Nineteenth International Conference (ICML 2002), University of New South Wales, Sydney, Australia, July 8-12, 2002, Claude Sammut and Achim G. Hoffmann (Eds.). Morgan Kaufmann, 227--234."},{"key":"e_1_3_2_2_14_1","unstructured":"Tuomas Haarnoja Aurick Zhou Kristian Hartikainen George Tucker Sehoon Ha Jie Tan Vikash Kumar Henry Zhu Abhishek Gupta Pieter Abbeel etal 2018. Soft actor-critic algorithms and applications. arXiv preprint arXiv:1812.05905 (2018).  Tuomas Haarnoja Aurick Zhou Kristian Hartikainen George Tucker Sehoon Ha Jie Tan Vikash Kumar Henry Zhu Abhishek Gupta Pieter Abbeel et al. 2018. Soft actor-critic algorithms and applications. arXiv preprint arXiv:1812.05905 (2018)."},{"key":"e_1_3_2_2_15_1","doi-asserted-by":"publisher","DOI":"10.1145\/3534678.3542679"},{"key":"e_1_3_2_2_16_1","doi-asserted-by":"publisher","DOI":"10.1145\/3447548.3467181"},{"key":"e_1_3_2_2_17_1","volume-title":"Graph convolutional reinforcement learning. arXiv preprint arXiv:1810.09202","author":"Jiang Jiechuan","year":"2018","unstructured":"Jiechuan Jiang , Chen Dun , Tiejun Huang , and Zongqing Lu. 2018. Graph convolutional reinforcement learning. arXiv preprint arXiv:1810.09202 ( 2018 ). Jiechuan Jiang, Chen Dun, Tiejun Huang, and Zongqing Lu. 2018. Graph convolutional reinforcement learning. arXiv preprint arXiv:1810.09202 (2018)."},{"key":"e_1_3_2_2_18_1","volume-title":"Trust region policy optimisation in multi-agent reinforcement learning. arXiv preprint arXiv:2109.11251","author":"Kuba Jakub Grudzien","year":"2021","unstructured":"Jakub Grudzien Kuba , Ruiqing Chen , Muning Wen , Ying Wen , Fanglei Sun , Jun Wang , and Yaodong Yang . 2021. Trust region policy optimisation in multi-agent reinforcement learning. arXiv preprint arXiv:2109.11251 ( 2021 ). Jakub Grudzien Kuba, Ruiqing Chen, Muning Wen, Ying Wen, Fanglei Sun, Jun Wang, and Yaodong Yang. 2021. Trust region policy optimisation in multi-agent reinforcement learning. arXiv preprint arXiv:2109.11251 (2021)."},{"key":"e_1_3_2_2_19_1","volume-title":"GARLSched: Generative adversarial deep reinforcement learning task scheduling optimization for large-scale high performance computing systems. Future Generation Computer Systems","author":"Li Jingbo","year":"2022","unstructured":"Jingbo Li , Xingjun Zhang , Jia Wei , Zeyu Ji , and Zheng Wei . 2022. GARLSched: Generative adversarial deep reinforcement learning task scheduling optimization for large-scale high performance computing systems. Future Generation Computer Systems ( 2022 ). Jingbo Li, Xingjun Zhang, Jia Wei, Zeyu Ji, and Zheng Wei. 2022. GARLSched: Generative adversarial deep reinforcement learning task scheduling optimization for large-scale high performance computing systems. Future Generation Computer Systems (2022)."},{"key":"e_1_3_2_2_20_1","unstructured":"Minne Li Zhiwei Qin Yan Jiao Yaodong Yang Jun Wang Chenxi Wang Guobin Wu and Jieping Ye. 2019. Efficient ridesharing order dispatching with mean field multi-agent reinforcement learning. In The world wide web conference. 983--994.  Minne Li Zhiwei Qin Yan Jiao Yaodong Yang Jun Wang Chenxi Wang Guobin Wu and Jieping Ye. 2019. Efficient ridesharing order dispatching with mean field multi-agent reinforcement learning. In The world wide web conference. 983--994."},{"key":"e_1_3_2_2_21_1","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v34i05.6211"},{"key":"e_1_3_2_2_22_1","volume-title":"OpenAI Pieter Abbeel, and Igor Mordatch","author":"Lowe Ryan","year":"2017","unstructured":"Ryan Lowe , Yi I Wu , Aviv Tamar , Jean Harb , OpenAI Pieter Abbeel, and Igor Mordatch . 2017 . Multi-agent actor-critic for mixed cooperative-competitive environments. Advances in neural information processing systems, Vol. 30 (2017). Ryan Lowe, Yi I Wu, Aviv Tamar, Jean Harb, OpenAI Pieter Abbeel, and Igor Mordatch. 2017. Multi-agent actor-critic for mixed cooperative-competitive environments. Advances in neural information processing systems, Vol. 30 (2017)."},{"key":"e_1_3_2_2_23_1","first-page":"23609","article-title":"A hierarchical reinforcement learning based optimization framework for large-scale dynamic pickup and delivery problems","volume":"34","author":"Ma Yi","year":"2021","unstructured":"Yi Ma , Xiaotian Hao , Jianye Hao , Jiawen Lu , Xing Liu , Tong Xialiang , Mingxuan Yuan , Zhigang Li , Jie Tang , and Zhaopeng Meng . 2021 . A hierarchical reinforcement learning based optimization framework for large-scale dynamic pickup and delivery problems . Advances in Neural Information Processing Systems , Vol. 34 (2021), 23609 -- 23620 . Yi Ma, Xiaotian Hao, Jianye Hao, Jiawen Lu, Xing Liu, Tong Xialiang, Mingxuan Yuan, Zhigang Li, Jie Tang, and Zhaopeng Meng. 2021. A hierarchical reinforcement learning based optimization framework for large-scale dynamic pickup and delivery problems. Advances in Neural Information Processing Systems, Vol. 34 (2021), 23609--23620.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_2_2_24_1","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v34i05.6212"},{"key":"e_1_3_2_2_25_1","volume-title":"Modelling the dynamic joint policy of teammates with attention multi-agent DDPG. arXiv preprint arXiv:1811.07029","author":"Mao Hangyu","year":"2018","unstructured":"Hangyu Mao , Zhengchao Zhang , Zhen Xiao , and Zhibo Gong . 2018. Modelling the dynamic joint policy of teammates with attention multi-agent DDPG. arXiv preprint arXiv:1811.07029 ( 2018 ). Hangyu Mao, Zhengchao Zhang, Zhen Xiao, and Zhibo Gong. 2018. Modelling the dynamic joint policy of teammates with attention multi-agent DDPG. arXiv preprint arXiv:1811.07029 (2018)."},{"key":"e_1_3_2_2_26_1","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v34i04.5957"},{"key":"e_1_3_2_2_27_1","volume-title":"Playing atari with deep reinforcement learning. arXiv preprint arXiv:1312.5602","author":"Mnih Volodymyr","year":"2013","unstructured":"Volodymyr Mnih , Koray Kavukcuoglu , David Silver , Alex Graves , Ioannis Antonoglou , Daan Wierstra , and Martin Riedmiller . 2013. Playing atari with deep reinforcement learning. arXiv preprint arXiv:1312.5602 ( 2013 ). Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Alex Graves, Ioannis Antonoglou, Daan Wierstra, and Martin Riedmiller. 2013. Playing atari with deep reinforcement learning. arXiv preprint arXiv:1312.5602 (2013)."},{"key":"e_1_3_2_2_28_1","volume-title":"Proceedings, The Twentieth National Conference on Artificial Intelligence and the Seventeenth Innovative Applications of Artificial Intelligence Conference","author":"Nair Ranjit","year":"2005","unstructured":"Ranjit Nair , Pradeep Varakantham , Milind Tambe , and Makoto Yokoo . 2005 . Networked Distributed POMDPs: A Synthesis of Distributed Constraint Optimization and POMDPs . In Proceedings, The Twentieth National Conference on Artificial Intelligence and the Seventeenth Innovative Applications of Artificial Intelligence Conference , July 9-13, 2005, Pittsburgh, Pennsylvania, USA, Manuela M. Veloso and Subbarao Kambhampati (Eds.). AAAI Press \/ The MIT Press, 133--139. http:\/\/www.aaai.org\/Library\/AAAI\/ 2005\/aaai05-022.php Ranjit Nair, Pradeep Varakantham, Milind Tambe, and Makoto Yokoo. 2005. Networked Distributed POMDPs: A Synthesis of Distributed Constraint Optimization and POMDPs. In Proceedings, The Twentieth National Conference on Artificial Intelligence and the Seventeenth Innovative Applications of Artificial Intelligence Conference, July 9-13, 2005, Pittsburgh, Pennsylvania, USA, Manuela M. Veloso and Subbarao Kambhampati (Eds.). AAAI Press \/ The MIT Press, 133--139. http:\/\/www.aaai.org\/Library\/AAAI\/2005\/aaai05-022.php"},{"key":"e_1_3_2_2_29_1","unstructured":"Yaru Niu Rohan R Paleja and Matthew C Gombolay. 2021a. Multi-Agent Graph-Attention Communication and Teaming.. In AAMAS. 964--973.  Yaru Niu Rohan R Paleja and Matthew C Gombolay. 2021a. Multi-Agent Graph-Attention Communication and Teaming.. In AAMAS. 964--973."},{"key":"e_1_3_2_2_30_1","volume-title":"Multi-Agent Graph-Attention Communication and Teaming. In AAMAS '21: 20th International Conference on Autonomous Agents and Multiagent Systems","author":"Niu Yaru","year":"2021","unstructured":"Yaru Niu , Rohan R. Paleja , and Matthew C. Gombolay . 2021b . Multi-Agent Graph-Attention Communication and Teaming. In AAMAS '21: 20th International Conference on Autonomous Agents and Multiagent Systems , Virtual Event, United Kingdom , May 3-7, 2021 , Frank Dignum, Alessio Lomuscio, Ulle Endriss, and Ann Now\u00e9 (Eds.). ACM, 964--973. https:\/\/doi.org\/10.5555\/3463952.3464065 10.5555\/3463952.3464065 Yaru Niu, Rohan R. Paleja, and Matthew C. Gombolay. 2021b. Multi-Agent Graph-Attention Communication and Teaming. In AAMAS '21: 20th International Conference on Autonomous Agents and Multiagent Systems, Virtual Event, United Kingdom, May 3-7, 2021, Frank Dignum, Alessio Lomuscio, Ulle Endriss, and Ann Now\u00e9 (Eds.). ACM, 964--973. https:\/\/doi.org\/10.5555\/3463952.3464065"},{"key":"e_1_3_2_2_31_1","doi-asserted-by":"publisher","DOI":"10.1613\/jair.1.12136"},{"key":"e_1_3_2_2_32_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.apenergy.2021.116940"},{"key":"e_1_3_2_2_33_1","doi-asserted-by":"publisher","DOI":"10.1109\/ETFA.2019.8869023"},{"key":"e_1_3_2_2_34_1","volume-title":"International conference on machine learning. PMLR, 4295--4304","author":"Rashid Tabish","year":"2018","unstructured":"Tabish Rashid , Mikayel Samvelyan , Christian Schroeder , Gregory Farquhar , Jakob Foerster , and Shimon Whiteson . 2018 Qmix: Monotonic value function factorisation for deep multi-agent reinforcement learning . In International conference on machine learning. PMLR, 4295--4304 . Tabish Rashid, Mikayel Samvelyan, Christian Schroeder, Gregory Farquhar, Jakob Foerster, and Shimon Whiteson. 2018 Qmix: Monotonic value function factorisation for deep multi-agent reinforcement learning. In International conference on machine learning. PMLR, 4295--4304."},{"key":"e_1_3_2_2_35_1","doi-asserted-by":"publisher","DOI":"10.1109\/JIOT.2021.3071531"},{"key":"e_1_3_2_2_36_1","volume-title":"GCS: Graph-Based Coordination Strategy for Multi-Agent Reinforcement Learning. arXiv preprint arXiv:2201.06257","author":"Ruan Jingqing","year":"2022","unstructured":"Jingqing Ruan , Yali Du , Xuantang Xiong , Dengpeng Xing , Xiyun Li , Linghui Meng , Haifeng Zhang , Jun Wang , and Bo Xu . 2022 . GCS: Graph-Based Coordination Strategy for Multi-Agent Reinforcement Learning. arXiv preprint arXiv:2201.06257 (2022). Jingqing Ruan, Yali Du, Xuantang Xiong, Dengpeng Xing, Xiyun Li, Linghui Meng, Haifeng Zhang, Jun Wang, and Bo Xu. 2022. GCS: Graph-Based Coordination Strategy for Multi-Agent Reinforcement Learning. arXiv preprint arXiv:2201.06257 (2022)."},{"key":"e_1_3_2_2_37_1","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v34i05.6214"},{"key":"e_1_3_2_2_38_1","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v34i05.6214"},{"key":"e_1_3_2_2_39_1","volume-title":"Proximal policy optimization algorithms. arXiv preprint arXiv:1707.06347","author":"Schulman John","year":"2017","unstructured":"John Schulman , Filip Wolski , Prafulla Dhariwal , Alec Radford , and Oleg Klimov . 2017. Proximal policy optimization algorithms. arXiv preprint arXiv:1707.06347 ( 2017 ). John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov. 2017. Proximal policy optimization algorithms. arXiv preprint arXiv:1707.06347 (2017)."},{"key":"e_1_3_2_2_40_1","doi-asserted-by":"publisher","DOI":"10.1109\/JIOT.2020.2983741"},{"key":"e_1_3_2_2_41_1","volume-title":"International conference on machine learning. PMLR, 387--395","author":"Silver David","year":"2014","unstructured":"David Silver , Guy Lever , Nicolas Heess , Thomas Degris , Daan Wierstra , and Martin Riedmiller . 2014 . Deterministic policy gradient algorithms . In International conference on machine learning. PMLR, 387--395 . David Silver, Guy Lever, Nicolas Heess, Thomas Degris, Daan Wierstra, and Martin Riedmiller. 2014. Deterministic policy gradient algorithms. In International conference on machine learning. PMLR, 387--395."},{"key":"e_1_3_2_2_42_1","volume-title":"Nature","volume":"550","author":"Silver David","year":"2017","unstructured":"David Silver , Julian Schrittwieser , Karen Simonyan , Ioannis Antonoglou , Aja Huang , Arthur Guez , Thomas Hubert , Lucas Baker , Matthew Lai , Adrian Bolton , 2017 . Mastering the game of go without human knowledge . Nature , Vol. 550 , 7676 (2017), 354--359. David Silver, Julian Schrittwieser, Karen Simonyan, Ioannis Antonoglou, Aja Huang, Arthur Guez, Thomas Hubert, Lucas Baker, Matthew Lai, Adrian Bolton, et al. 2017. Mastering the game of go without human knowledge. Nature, Vol. 550, 7676 (2017), 354--359."},{"key":"e_1_3_2_2_43_1","volume-title":"Conference on Robot Learning. PMLR, 907--917","author":"Sinha Samarth","year":"2022","unstructured":"Samarth Sinha , Ajay Mandlekar , and Animesh Garg . 2022 . S4RL: Surprisingly simple self-supervision for offline reinforcement learning in robotics . In Conference on Robot Learning. PMLR, 907--917 . Samarth Sinha, Ajay Mandlekar, and Animesh Garg. 2022. S4RL: Surprisingly simple self-supervision for offline reinforcement learning in robotics. In Conference on Robot Learning. PMLR, 907--917."},{"key":"e_1_3_2_2_44_1","volume-title":"International conference on machine learning. PMLR, 5887--5896","author":"Son Kyunghwan","year":"2019","unstructured":"Kyunghwan Son , Daewoo Kim , Wan Ju Kang , David Earl Hostallero , and Yung Yi . 2019 . Qtran: Learning to factorize with transformation for cooperative multi-agent reinforcement learning . In International conference on machine learning. PMLR, 5887--5896 . Kyunghwan Son, Daewoo Kim, Wan Ju Kang, David Earl Hostallero, and Yung Yi. 2019. Qtran: Learning to factorize with transformation for cooperative multi-agent reinforcement learning. In International conference on machine learning. PMLR, 5887--5896."},{"key":"e_1_3_2_2_45_1","volume-title":"COVID-19 attack rate increases with city size. arXiv preprint arXiv:2003.10376","author":"Stier Andrew J","year":"2020","unstructured":"Andrew J Stier , Marc G Berman , and Luis Bettencourt . 2020. COVID-19 attack rate increases with city size. arXiv preprint arXiv:2003.10376 ( 2020 ). Andrew J Stier, Marc G Berman, and Luis Bettencourt. 2020. COVID-19 attack rate increases with city size. arXiv preprint arXiv:2003.10376 (2020)."},{"key":"e_1_3_2_2_46_1","unstructured":"Sainbayar Sukhbaatar Rob Fergus etal 2016. Learning multiagent communication with backpropagation. Advances in neural information processing systems Vol. 29 (2016).  Sainbayar Sukhbaatar Rob Fergus et al. 2016. Learning multiagent communication with backpropagation. Advances in neural information processing systems Vol. 29 (2016)."},{"key":"e_1_3_2_2_47_1","volume-title":"Vinicius Zambaldi, Max Jaderberg, Marc Lanctot, Nicolas Sonnerat, Joel Z Leibo, Karl Tuyls, et al.","author":"Sunehag Peter","year":"2017","unstructured":"Peter Sunehag , Guy Lever , Audrunas Gruslys , Wojciech Marian Czarnecki , Vinicius Zambaldi, Max Jaderberg, Marc Lanctot, Nicolas Sonnerat, Joel Z Leibo, Karl Tuyls, et al. 2017 . Value-decomposition networks for cooperative multi-agent learning. arXiv preprint arXiv:1706.05296 (2017). Peter Sunehag, Guy Lever, Audrunas Gruslys, Wojciech Marian Czarnecki, Vinicius Zambaldi, Max Jaderberg, Marc Lanctot, Nicolas Sonnerat, Joel Z Leibo, Karl Tuyls, et al. 2017. Value-decomposition networks for cooperative multi-agent learning. arXiv preprint arXiv:1706.05296 (2017)."},{"key":"e_1_3_2_2_48_1","volume-title":"Policy gradient methods for reinforcement learning with function approximation. Advances in neural information processing systems","author":"Sutton Richard S","year":"1999","unstructured":"Richard S Sutton , David McAllester , Satinder Singh , and Yishay Mansour . 1999. Policy gradient methods for reinforcement learning with function approximation. Advances in neural information processing systems , Vol. 12 ( 1999 ). Richard S Sutton, David McAllester, Satinder Singh, and Yishay Mansour. 1999. Policy gradient methods for reinforcement learning with function approximation. Advances in neural information processing systems, Vol. 12 (1999)."},{"key":"e_1_3_2_2_49_1","volume-title":"Nature","volume":"575","author":"Vinyals Oriol","year":"2019","unstructured":"Oriol Vinyals , Igor Babuschkin , Wojciech M Czarnecki , Micha\u00ebl Mathieu , Andrew Dudzik , Junyoung Chung , David H Choi , Richard Powell , Timo Ewalds , Petko Georgiev , 2019 . Grandmaster level in StarCraft II using multi-agent reinforcement learning . Nature , Vol. 575 , 7782 (2019), 350--354. Oriol Vinyals, Igor Babuschkin, Wojciech M Czarnecki, Micha\u00ebl Mathieu, Andrew Dudzik, Junyoung Chung, David H Choi, Richard Powell, Timo Ewalds, Petko Georgiev, et al. 2019. Grandmaster level in StarCraft II using multi-agent reinforcement learning. Nature, Vol. 575, 7782 (2019), 350--354."},{"key":"e_1_3_2_2_50_1","first-page":"3271","article-title":"Multi-agent reinforcement learning for active voltage control on power distribution networks","volume":"34","author":"Wang Jianhong","year":"2021","unstructured":"Jianhong Wang , Wangkun Xu , Yunjie Gu , Wenbin Song , and Tim C Green . 2021 b. Multi-agent reinforcement learning for active voltage control on power distribution networks . Advances in Neural Information Processing Systems , Vol. 34 (2021), 3271 -- 3284 . Jianhong Wang, Wangkun Xu, Yunjie Gu, Wenbin Song, and Tim C Green. 2021b. Multi-agent reinforcement learning for active voltage control on power distribution networks. Advances in Neural Information Processing Systems, Vol. 34 (2021), 3271--3284.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_2_2_51_1","volume-title":"Adaptive Traffic Signal Control for large-scale scenario with Cooperative Group-based Multi-agent reinforcement learning. Transportation research part C: emerging technologies","author":"Wang Tong","year":"2021","unstructured":"Tong Wang , Jiahua Cao , and Azhar Hussain . 2021a. Adaptive Traffic Signal Control for large-scale scenario with Cooperative Group-based Multi-agent reinforcement learning. Transportation research part C: emerging technologies , Vol. 125 ( 2021 ), 103046. Tong Wang, Jiahua Cao, and Azhar Hussain. 2021a. Adaptive Traffic Signal Control for large-scale scenario with Cooperative Group-based Multi-agent reinforcement learning. Transportation research part C: emerging technologies, Vol. 125 (2021), 103046."},{"key":"e_1_3_2_2_52_1","volume-title":"Machine learning","author":"Watkins Christopher JCH","year":"1992","unstructured":"Christopher JCH Watkins and Peter Dayan . 1992. Q-learning. Machine learning , Vol. 8 , 3 ( 1992 ), 279--292. Christopher JCH Watkins and Peter Dayan. 1992. Q-learning. Machine learning, Vol. 8, 3 (1992), 279--292."},{"key":"e_1_3_2_2_53_1","volume-title":"International conference on machine learning. PMLR, 5571--5580","author":"Yang Yaodong","year":"2018","unstructured":"Yaodong Yang , Rui Luo , Minne Li , Ming Zhou , Weinan Zhang , and Jun Wang . 2018 . Mean field multi-agent reinforcement learning . In International conference on machine learning. PMLR, 5571--5580 . Yaodong Yang, Rui Luo, Minne Li, Ming Zhou, Weinan Zhang, and Jun Wang. 2018. Mean field multi-agent reinforcement learning. In International conference on machine learning. PMLR, 5571--5580."},{"key":"e_1_3_2_2_54_1","first-page":"25476","article-title":"Mastering atari games with limited data","volume":"34","author":"Ye Weirui","year":"2021","unstructured":"Weirui Ye , Shaohuai Liu , Thanard Kurutach , Pieter Abbeel , and Yang Gao . 2021 . Mastering atari games with limited data . Advances in Neural Information Processing Systems , Vol. 34 (2021), 25476 -- 25488 . Weirui Ye, Shaohuai Liu, Thanard Kurutach, Pieter Abbeel, and Yang Gao. 2021. Mastering atari games with limited data. Advances in Neural Information Processing Systems, Vol. 34 (2021), 25476--25488.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_2_2_55_1","volume-title":"The surprising effectiveness of ppo in cooperative, multi-agent games. arXiv preprint arXiv:2103.01955","author":"Yu Chao","year":"2021","unstructured":"Chao Yu , Akash Velu , Eugene Vinitsky , Yu Wang , Alexandre Bayen , and Yi Wu. 2021. The surprising effectiveness of ppo in cooperative, multi-agent games. arXiv preprint arXiv:2103.01955 ( 2021 ). Chao Yu, Akash Velu, Eugene Vinitsky, Yu Wang, Alexandre Bayen, and Yi Wu. 2021. The surprising effectiveness of ppo in cooperative, multi-agent games. arXiv preprint arXiv:2103.01955 (2021)."},{"key":"e_1_3_2_2_56_1","volume-title":"CTDS: Centralized Teacher with Decentralized Student for Multi-Agent Reinforcement Learning","author":"Zhao Jian","year":"2022","unstructured":"Jian Zhao , Xunhan Hu , Mingyu Yang , Wengang Zhou , Jiangcheng Zhu , and Houqiang Li . 2022 . CTDS: Centralized Teacher with Decentralized Student for Multi-Agent Reinforcement Learning . IEEE Transactions on Games ( 2022). Jian Zhao, Xunhan Hu, Mingyu Yang, Wengang Zhou, Jiangcheng Zhu, and Houqiang Li. 2022. CTDS: Centralized Teacher with Decentralized Student for Multi-Agent Reinforcement Learning. IEEE Transactions on Games (2022)."}],"event":{"name":"KDD '23: The 29th ACM SIGKDD Conference on Knowledge Discovery and Data Mining","location":"Long Beach CA USA","acronym":"KDD '23","sponsor":["SIGMOD ACM Special Interest Group on Management of Data","SIGKDD ACM Special Interest Group on Knowledge Discovery in Data"]},"container-title":["Proceedings of the 29th ACM SIGKDD Conference on Knowledge Discovery and Data Mining"],"original-title":[],"link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3580305.3599359","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3580305.3599359","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T16:37:47Z","timestamp":1750178267000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3580305.3599359"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,8,4]]},"references-count":56,"alternative-id":["10.1145\/3580305.3599359","10.1145\/3580305"],"URL":"https:\/\/doi.org\/10.1145\/3580305.3599359","relation":{},"subject":[],"published":{"date-parts":[[2023,8,4]]},"assertion":[{"value":"2023-08-04","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}