{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,10,8]],"date-time":"2025-10-08T16:31:17Z","timestamp":1759941077874,"version":"3.41.0"},"reference-count":62,"publisher":"Association for Computing Machinery (ACM)","issue":"6","license":[{"start":{"date-parts":[[2023,11,14]],"date-time":"2023-11-14T00:00:00Z","timestamp":1699920000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"name":"Hong Kong Research Grant Council","award":["GRF 11218621"],"award-info":[{"award-number":["GRF 11218621"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Intell. Syst. Technol."],"published-print":{"date-parts":[[2023,12,31]]},"abstract":"<jats:p>In Deep Reinforcement Learning (DRL) domain, a compound learning task is often decomposed into several sub-tasks in a divide-and-conquer manner, each trained separately and then fused concurrently to achieve the original task, referred to as policy fusion. However, the state-of-the-art (SOTA) policy fusion methods treat the importance of sub-tasks equally throughout the task process, eliminating the possibility of the agent relying on different sub-tasks at various stages. To address this limitation, we propose a generic policy fusion approach, referred to as Policy Fusion Learning with Dynamic Weights and Prior Reward (PFLDWPR), to automate the time-varying selection of sub-tasks. Specifically, PFLDWPR produces a time-varying one-hot vector for sub-tasks to dynamically select a suitable sub-task and mask the rest throughout the entire task process, enabling the fused strategy to optimally guide the agent in executing the compound task. The sub-tasks with the dynamic one-hot vector are then aggregated to obtain the action policy for the original task. Moreover, we collect sub-tasks\u2019s rewards at the pre-training stage as a prior reward, which, along with the current reward, is used to train the policy fusion network. Thus, this approach reduces fusion bias by leveraging prior experience. Experimental results under three popular learning tasks demonstrate that the proposed method significantly improves three SOTA policy fusion methods in terms of task duration, episode reward, and score difference.<\/jats:p>","DOI":"10.1145\/3623405","type":"journal-article","created":{"date-parts":[[2023,9,11]],"date-time":"2023-09-11T11:49:36Z","timestamp":1694432976000},"page":"1-28","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":3,"title":["Dynamic Weights and Prior Reward in Policy Fusion for Compound Agent Learning"],"prefix":"10.1145","volume":"14","author":[{"ORCID":"https:\/\/orcid.org\/0000-0003-4857-5439","authenticated-orcid":false,"given":"Meng","family":"Xu","sequence":"first","affiliation":[{"name":"Department of Computer Science, City University of Hong Kong, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-6951-2392","authenticated-orcid":false,"given":"Yechao","family":"She","sequence":"additional","affiliation":[{"name":"Department of Computer Science, City University of Hong Kong, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-5337-487X","authenticated-orcid":false,"given":"Yang","family":"Jin","sequence":"additional","affiliation":[{"name":"Department of Computer Science, City University of Hong Kong, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-9318-1482","authenticated-orcid":false,"given":"Jianping","family":"Wang","sequence":"additional","affiliation":[{"name":"Department of Computer Science, City University of Hong Kong, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2023,11,14]]},"reference":[{"key":"e_1_3_2_2_2","first-page":"11","volume-title":"Proceedings of the International Conference on Machine Learning","author":"Abels Axel","year":"2019","unstructured":"Axel Abels, Diederik Roijers, Tom Lenaerts, Ann Now\u00e9, and Denis Steckelmacher. 2019. Dynamic weights in multi-objective deep reinforcement learning. In Proceedings of the International Conference on Machine Learning. PMLR, 11\u201320."},{"key":"e_1_3_2_3_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.12"},{"key":"e_1_3_2_4_2","unstructured":"Mohammad Babaeizadeh Iuri Frosio Stephen Tyree Jason Clemons and Jan Kautz. 2016. Reinforcement learning through asynchronous advantage actor-critic on a gpu. arXiv preprint arXiv:1611.06256 (2016). Retrieved from https:\/\/arxiv.org\/abs\/1611.06256"},{"key":"e_1_3_2_5_2","unstructured":"Glen Berseth Cheng Xie Paul Cernek and Michiel Van de Panne. 2018. Progressive reinforcement learning with distillation for multi-skilled motion control. arXiv preprint arXiv:1802.04765 (2018). Retrieved from https:\/\/arxiv.org\/abs\/1802.04765"},{"key":"e_1_3_2_6_2","doi-asserted-by":"publisher","DOI":"10.1214\/ss\/1177011077"},{"key":"e_1_3_2_7_2","first-page":"169","volume-title":"Proceedings of the ECAI","author":"Bianchi Reinaldo A. C.","year":"2012","unstructured":"Reinaldo A. C. Bianchi, Carlos H. C. Ribeiro, and Anna Helena Reali Costa. 2012. Heuristically accelerated reinforcement learning: Theoretical and experimental results. In Proceedings of the ECAI. 169\u2013174."},{"key":"e_1_3_2_8_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.ifacol.2020.12.2317"},{"key":"e_1_3_2_9_2","doi-asserted-by":"publisher","DOI":"10.1109\/IROS40897.2019.8968092"},{"key":"e_1_3_2_10_2","first-page":"1087","volume-title":"Proceedings of the International Conference on Machine Learning","author":"Czarnecki Wojciech","year":"2018","unstructured":"Wojciech Czarnecki, Siddhant Jayakumar, Max Jaderberg, Leonard Hasenclever, Yee Whye Teh, Nicolas Heess, Simon Osindero, and Razvan Pascanu. 2018. Mix and match agent curricula for reinforcement learning. In Proceedings of the International Conference on Machine Learning. PMLR, 1087\u20131095."},{"key":"e_1_3_2_11_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCDS.2022.3177691"},{"key":"e_1_3_2_12_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2019.2956703"},{"key":"e_1_3_2_13_2","doi-asserted-by":"publisher","DOI":"10.1109\/IROS40897.2019.8968149"},{"key":"e_1_3_2_14_2","unstructured":"Behzad Ghazanfari and Matthew E. Taylor. 2017. Autonomous extracting a hierarchical structure of tasks in reinforcement learning and multi-task reinforcement learning. arXiv preprint arXiv:1709.04579 (2017). Retrieved from https:\/\/arxiv.org\/abs\/1709.04579"},{"key":"e_1_3_2_15_2","unstructured":"Joshua Hare. 2019. Dealing with sparse rewards in reinforcement learning. arXiv preprint arXiv:1910.09281 (2019)."},{"key":"e_1_3_2_16_2","unstructured":"Zhanpeng He Ryan Julian Eric Heiden Hejia Zhang Stefan Schaal Joseph J. Lim Gaurav Sukhatme and Karol Hausman. 2018. Zero-shot skill composition and simulation-to-real transfer by learning task representations. arXiv preprint arXiv:1810.02422 (2018). Retrieved from https:\/\/arxiv.org\/abs\/1810.02422"},{"key":"e_1_3_2_17_2","doi-asserted-by":"publisher","DOI":"10.1109\/IROS40897.2019.8967913"},{"key":"e_1_3_2_18_2","doi-asserted-by":"publisher","DOI":"10.1023\/A:1008858222277"},{"key":"e_1_3_2_19_2","volume-title":"Proceedings of the International Conference on Learning Representations","author":"Lee Youngwoon","year":"2018","unstructured":"Youngwoon Lee, Shao-Hua Sun, Sriram Somasundaram, Edward S. Hu, and Joseph J. Lim. 2018. Composing complex skills by learning transition policies. In Proceedings of the International Conference on Learning Representations."},{"key":"e_1_3_2_20_2","volume-title":"Proceedings of the International Conference on Learning Representations","author":"Lee Youngwoon","year":"2019","unstructured":"Youngwoon Lee, Jingyun Yang, and Joseph J. Lim. 2019. Learning to coordinate manipulation skills via skill behavior diversification. In Proceedings of the International Conference on Learning Representations."},{"key":"e_1_3_2_21_2","unstructured":"Yuxi Li. 2017. Deep reinforcement learning: An overview. arXiv preprint arXiv:1701.07274 (2017). Retrieved from https:\/\/arxiv.org\/abs\/1701.07274"},{"key":"e_1_3_2_22_2","unstructured":"Jorge A. Mendez Harm van Seijen and Eric Eaton. 2022. Modular lifelong reinforcement learning via neural composition. arXiv preprint arXiv:2207.00429 (2022). Retrieved from https:\/\/arxiv.org\/abs\/2207.00429"},{"key":"e_1_3_2_23_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.neunet.2023.01.046"},{"key":"e_1_3_2_24_2","first-page":"1928","volume-title":"Proceedings of the International Conference on Machine Learning","author":"Mnih Volodymyr","year":"2016","unstructured":"Volodymyr Mnih, Adria Puigdomenech Badia, Mehdi Mirza, Alex Graves, Timothy Lillicrap, Tim Harley, David Silver, and Koray Kavukcuoglu. 2016. Asynchronous methods for deep reinforcement learning. In Proceedings of the International Conference on Machine Learning. PMLR, 1928\u20131937."},{"key":"e_1_3_2_25_2","unstructured":"Advances in Neural Information Processing Systems 2014 27 Recurrent models of visual attention"},{"key":"e_1_3_2_26_2","unstructured":"Volodymyr Mnih Koray Kavukcuoglu David Silver Alex Graves Ioannis Antonoglou Daan Wierstra and Martin Riedmiller. 2013. Playing atari with deep reinforcement learning. arXiv preprint arXiv:1312.5602 (2013). Retrieved from https:\/\/arxiv.org\/abs\/1312.5602"},{"key":"e_1_3_2_27_2","doi-asserted-by":"publisher","DOI":"10.1038\/nature14236"},{"key":"e_1_3_2_28_2","unstructured":"Hossam Mossalam Yannis M. Assael Diederik M. Roijers and Shimon Whiteson. 2016. Multi-objective deep reinforcement learning. arXiv preprint arXiv:1610.02707 (2016). Retrieved from https:\/\/arxiv.org\/abs\/1610.02707"},{"key":"e_1_3_2_29_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICHR.2010.5686298"},{"key":"e_1_3_2_30_2","doi-asserted-by":"publisher","DOI":"10.1177\/0278364912472380"},{"key":"e_1_3_2_31_2","unstructured":"Brendan O\u2019Donoghue Remi Munos Koray Kavukcuoglu and Volodymyr Mnih. 2016. Combining policy gradient and q-learning. arXiv preprint arXiv:1611.01626 (2016). Retrieved from https:\/\/arxiv.org\/abs\/1611.01626"},{"key":"e_1_3_2_32_2","first-page":"2661","volume-title":"Proceedings of the International Conference on Machine Learning","author":"Oh Junhyuk","year":"2017","unstructured":"Junhyuk Oh, Satinder Singh, Honglak Lee, and Pushmeet Kohli. 2017. Zero-shot task generalization with multi-task deep reinforcement learning. In Proceedings of the International Conference on Machine Learning. PMLR, 2661\u20132670."},{"key":"e_1_3_2_33_2","doi-asserted-by":"publisher","DOI":"10.1145\/3072959.3073602"},{"key":"e_1_3_2_34_2","unstructured":"Kai Ploeger Michael Lutter and Jan Peters. 2021. High acceleration reinforcement learning for real-world juggling with binary rewards. Conference on Robot Learning PMLR 642\u2013653. Retrieved from https:\/\/arxiv.org\/abs\/2010.13483"},{"key":"e_1_3_2_35_2","unstructured":"John Schulman Filip Wolski Prafulla Dhariwal Alec Radford and Oleg Klimov. 2017. Proximal policy optimization algorithms. arXiv preprint arXiv:1707.06347 (2017). Retrieved from https:\/\/arxiv.org\/abs\/1707.06347"},{"key":"e_1_3_2_36_2","doi-asserted-by":"publisher","DOI":"10.1109\/CoG52621.2021.9618983"},{"key":"e_1_3_2_37_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCDS.2019.2924724"},{"key":"e_1_3_2_38_2","doi-asserted-by":"publisher","DOI":"10.1109\/EIConRusNW.2016.7448188"},{"key":"e_1_3_2_39_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v33i01.33014975"},{"key":"e_1_3_2_40_2","doi-asserted-by":"publisher","DOI":"10.5555\/1689599.1689651"},{"key":"e_1_3_2_41_2","first-page":"2171","volume-title":"Proceedings of the International Conference on Machine Learning","author":"Snoek Jasper","year":"2015","unstructured":"Jasper Snoek, Oren Rippel, Kevin Swersky, Ryan Kiros, Nadathur Satish, Narayanan Sundaram, Mostofa Patwary, Mr Prabhat, and Ryan Adams. 2015. Scalable bayesian optimization using deep neural networks. In Proceedings of the International Conference on Machine Learning. PMLR, 2171\u20132180."},{"key":"e_1_3_2_42_2","doi-asserted-by":"crossref","unstructured":"Ghada Sokar Elena Mocanu Decebal Constantin Mocanu Mykola Pechenizkiy and Peter Stone. 2021. Dynamic sparse training for deep reinforcement learning. arXiv preprint arXiv:2106.04217 (2021). Retrieved from https:\/\/arxiv.org\/abs\/2106.04217","DOI":"10.24963\/ijcai.2022\/477"},{"key":"e_1_3_2_43_2","doi-asserted-by":"publisher","DOI":"10.1109\/CAC48633.2019.8996860"},{"key":"e_1_3_2_44_2","doi-asserted-by":"publisher","DOI":"10.1109\/TNNLS.2020.2996209"},{"key":"e_1_3_2_45_2","unstructured":"Hartmut Surmann Christian Jestel Robin Marchel Franziska Musberg Houssem Elhadj and Mahbube Ardani. 2020. Deep reinforcement learning for real autonomous mobile robot navigation in indoor environments. arXiv preprint arXiv:2005.13857 (2020). Retrieved from https:\/\/arxiv.org\/abs\/2005.13857"},{"key":"e_1_3_2_46_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v31i1.10744"},{"issue":"2","key":"e_1_3_2_47_2","doi-asserted-by":"crossref","first-page":"292","DOI":"10.1109\/TCDS.2016.2607018","article-title":"A reinforcement learning architecture that transfers knowledge between skills when solving multiple tasks","volume":"11","author":"Tommasino Paolo","year":"2016","unstructured":"Paolo Tommasino, Daniele Caligiore, Marco Mirolli, and Gianluca Baldassarre. 2016. A reinforcement learning architecture that transfers knowledge between skills when solving multiple tasks. IEEE Transactions on Cognitive and Developmental Systems 11, 2 (2016), 292\u2013317.","journal-title":"IEEE Transactions on Cognitive and Developmental Systems"},{"key":"e_1_3_2_48_2","first-page":"6401","volume-title":"Proceedings of the International Conference on Machine Learning","author":"Niekerk Benjamin Van","year":"2019","unstructured":"Benjamin Van Niekerk, Steven James, Adam Earle, and Benjamin Rosman. 2019. Composing value functions in reinforcement learning. In Proceedings of the International Conference on Machine Learning. PMLR, 6401\u20136409."},{"key":"e_1_3_2_49_2","article-title":"Sampling rate decay in hindsight experience replay for robot control","author":"Vecchietti Luiz Felipe","year":"2020","unstructured":"Luiz Felipe Vecchietti, Minah Seo, and Dongsoo Har. 2020. Sampling rate decay in hindsight experience replay for robot control. IEEE Transactions on Cybernetics 52, 3 (2020), 1515\u20131526.","journal-title":"IEEE Transactions on Cybernetics"},{"key":"e_1_3_2_50_2","doi-asserted-by":"publisher","DOI":"10.1109\/JIOT.2020.2973193"},{"key":"e_1_3_2_51_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.neunet.2020.12.001"},{"key":"e_1_3_2_52_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2019.2924417"},{"key":"e_1_3_2_53_2","first-page":"10158","volume-title":"Proceedings of the International Conference on Machine Learning","author":"Wei Kaixuan","year":"2020","unstructured":"Kaixuan Wei, Angelica Aviles-Rivero, Jingwei Liang, Ying Fu, Carola-Bibiane Sch\u00f6nlieb, and Hua Huang. 2020. Tuning-free plug-and-play proximal algorithm for inverse imaging problems. In Proceedings of the International Conference on Machine Learning. PMLR, 10158\u201310169."},{"key":"e_1_3_2_54_2","doi-asserted-by":"publisher","DOI":"10.1007\/s10458-020-09451-0"},{"key":"e_1_3_2_55_2","first-page":"10607","volume-title":"Proceedings of the International Conference on Machine Learning","author":"Xu Jie","year":"2020","unstructured":"Jie Xu, Yunsheng Tian, Pingchuan Ma, Daniela Rus, Shinjiro Sueda, and Wojciech Matusik. 2020. Prediction-guided multi-objective reinforcement learning for continuous robot control. In Proceedings of the International Conference on Machine Learning. PMLR, 10607\u201310616."},{"key":"e_1_3_2_56_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICMA.2019.8816420"},{"key":"e_1_3_2_57_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.knosys.2022.109448"},{"key":"e_1_3_2_58_2","doi-asserted-by":"publisher","DOI":"10.1145\/3627824"},{"key":"e_1_3_2_59_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-72062-9_35"},{"key":"e_1_3_2_60_2","article-title":"A generalized algorithm for multi-objective reinforcement learning and policy adaptation","volume":"32","author":"Yang Runzhe","year":"2019","unstructured":"Runzhe Yang, Xingyuan Sun, and Karthik Narasimhan. 2019. A generalized algorithm for multi-objective reinforcement learning and policy adaptation. Advances in Neural Information Processing Systems 32 (2019), 1\u201312.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_2_61_2","doi-asserted-by":"publisher","DOI":"10.1109\/TNNLS.2018.2805379"},{"key":"e_1_3_2_62_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v35i12.17300"},{"key":"e_1_3_2_63_2","unstructured":"Dongyang Zhao Yue Huang Changnan Xiao Yue Li and Shihong Deng. 2021. Hierarchical meta reinforcement learning for multi-task environments. https:\/\/openreview.net\/forum?id=u9ax42K7ND"}],"container-title":["ACM Transactions on Intelligent Systems and Technology"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3623405","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3623405","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T16:36:26Z","timestamp":1750178186000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3623405"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,11,14]]},"references-count":62,"journal-issue":{"issue":"6","published-print":{"date-parts":[[2023,12,31]]}},"alternative-id":["10.1145\/3623405"],"URL":"https:\/\/doi.org\/10.1145\/3623405","relation":{},"ISSN":["2157-6904","2157-6912"],"issn-type":[{"type":"print","value":"2157-6904"},{"type":"electronic","value":"2157-6912"}],"subject":[],"published":{"date-parts":[[2023,11,14]]},"assertion":[{"value":"2023-05-25","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2023-08-25","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2023-11-14","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}