{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T04:11:04Z","timestamp":1750219864817,"version":"3.41.0"},"publisher-location":"New York, NY, USA","reference-count":26,"publisher":"ACM","license":[{"start":{"date-parts":[[2022,12,9]],"date-time":"2022-12-09T00:00:00Z","timestamp":1670544000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"name":"Beijing University of Posts and Telecommunication-China Mobile Research Institution Joint Innovation Center"},{"name":"Guangdong Province Science and Technology Project","award":["2021A0505080015"],"award-info":[{"award-number":["2021A0505080015"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2022,12,9]]},"DOI":"10.1145\/3577530.3577562","type":"proceedings-article","created":{"date-parts":[[2023,3,30]],"date-time":"2023-03-30T22:13:24Z","timestamp":1680214404000},"page":"195-202","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":0,"title":["Policy Transfer via Skill Adaptation and Composition"],"prefix":"10.1145","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-5062-2460","authenticated-orcid":false,"given":"Benhui","family":"Zhuang","sequence":"first","affiliation":[{"name":"State Key Laboratory of Networking and Switching Technology, Beijing University of Posts and Telecommunications, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-3008-1887","authenticated-orcid":false,"given":"Chunhong","family":"Zhang","sequence":"additional","affiliation":[{"name":"Key Laboratory of Universal Wireless Communications, Beijing University of Posts and Telecommunications, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-8874-5466","authenticated-orcid":false,"given":"Zheng","family":"Hu","sequence":"additional","affiliation":[{"name":"State Key Laboratory of Networking and Switching Technology, Beijing University of Posts and Telecommunications, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2023,3,30]]},"reference":[{"key":"e_1_3_2_1_1_1","volume-title":"International Conference on Learning Representations.","author":"Chevalier-Boisvert Maxime","year":"2019","unstructured":"Maxime Chevalier-Boisvert , Dzmitry Bahdanau , Salem Lahlou , Lucas Willems , Chitwan Saharia , Thien\u00a0Huu Nguyen , and Yoshua Bengio . 2019 . BabyAI: First Steps Towards Grounded Language Learning With a Human In the Loop . In International Conference on Learning Representations. Maxime Chevalier-Boisvert, Dzmitry Bahdanau, Salem Lahlou, Lucas Willems, Chitwan Saharia, Thien\u00a0Huu Nguyen, and Yoshua Bengio. 2019. BabyAI: First Steps Towards Grounded Language Learning With a Human In the Loop. In International Conference on Learning Representations."},{"key":"e_1_3_2_1_2_1","volume-title":"Minimalistic Gridworld Environment for OpenAI Gym. https:\/\/github.com\/maximecb\/gym-minigrid. (accessed","author":"Chevalier-Boisvert Maxime","year":"2022","unstructured":"Maxime Chevalier-Boisvert , Lucas Willems , and Suman Pal . 2018. Minimalistic Gridworld Environment for OpenAI Gym. https:\/\/github.com\/maximecb\/gym-minigrid. (accessed 13 August 2022 ). Maxime Chevalier-Boisvert, Lucas Willems, and Suman Pal. 2018. Minimalistic Gridworld Environment for OpenAI Gym. https:\/\/github.com\/maximecb\/gym-minigrid. (accessed 13 August 2022)."},{"key":"e_1_3_2_1_3_1","unstructured":"Anirudh Goyal Alex Lamb Jordan Hoffmann Shagun Sodhani Sergey Levine Yoshua Bengio and Bernhard Sch\u00f6lkopf. 2019. Recurrent Independent Mechanisms. CoRR abs\/1909.10893(2019). arxiv:1909.10893  Anirudh Goyal Alex Lamb Jordan Hoffmann Shagun Sodhani Sergey Levine Yoshua Bengio and Bernhard Sch\u00f6lkopf. 2019. Recurrent Independent Mechanisms. CoRR abs\/1909.10893(2019). arxiv:1909.10893"},{"key":"e_1_3_2_1_4_1","volume-title":"Annual Conference on Robot Learning(Proceedings of Machine Learning Research, Vol.\u00a0100)","author":"Gupta Abhishek","year":"2019","unstructured":"Abhishek Gupta , Vikash Kumar , Corey Lynch , Sergey Levine , and Karol Hausman . 2019 . Relay Policy Learning: Solving Long-Horizon Tasks via Imitation and Reinforcement Learning . In Annual Conference on Robot Learning(Proceedings of Machine Learning Research, Vol.\u00a0100) , Leslie\u00a0Pack Kaelbling, Danica Kragic, and Komei Sugiura (Eds.). PMLR, 1025\u20131037. Abhishek Gupta, Vikash Kumar, Corey Lynch, Sergey Levine, and Karol Hausman. 2019. Relay Policy Learning: Solving Long-Horizon Tasks via Imitation and Reinforcement Learning. In Annual Conference on Robot Learning(Proceedings of Machine Learning Research, Vol.\u00a0100), Leslie\u00a0Pack Kaelbling, Danica Kragic, and Komei Sugiura (Eds.). PMLR, 1025\u20131037."},{"key":"e_1_3_2_1_5_1","unstructured":"Vaibhav Gupta Daksh Anand Praveen Paruchuri and Akshat Kumar. 2021. Action Selection for Composable Modular Deep Reinforcement Learning. In International Conference on Autonomous Agents and Multiagent Systems Frank Dignum Alessio Lomuscio Ulle Endriss and Ann Now\u00e9 (Eds.). ACM 565\u2013573.  Vaibhav Gupta Daksh Anand Praveen Paruchuri and Akshat Kumar. 2021. Action Selection for Composable Modular Deep Reinforcement Learning. In International Conference on Autonomous Agents and Multiagent Systems Frank Dignum Alessio Lomuscio Ulle Endriss and Ann Now\u00e9 (Eds.). ACM 565\u2013573."},{"key":"e_1_3_2_1_6_1","unstructured":"David\u00a0Yu-Tung Hui Maxime Chevalier-Boisvert Dzmitry Bahdanau and Yoshua Bengio. 2020. BabyAI 1.1. CoRR abs\/2007.12770(2020). arxiv:2007.12770  David\u00a0Yu-Tung Hui Maxime Chevalier-Boisvert Dzmitry Bahdanau and Yoshua Bengio. 2020. BabyAI 1.1. CoRR abs\/2007.12770(2020). arxiv:2007.12770"},{"volume-title":"Adam: A Method for Stochastic Optimization. In International Conference on Learning Representations, Yoshua Bengio and Yann LeCun (Eds.).","author":"P.","key":"e_1_3_2_1_7_1","unstructured":"Diederik\u00a0 P. Kingma and Jimmy Ba. 2015 . Adam: A Method for Stochastic Optimization. In International Conference on Learning Representations, Yoshua Bengio and Yann LeCun (Eds.). Diederik\u00a0P. Kingma and Jimmy Ba. 2015. Adam: A Method for Stochastic Optimization. In International Conference on Learning Representations, Yoshua Bengio and Yann LeCun (Eds.)."},{"key":"e_1_3_2_1_8_1","volume-title":"International Conference on Learning Representations. OpenReview.net.","author":"Lee Youngwoon","year":"2019","unstructured":"Youngwoon Lee , Shao-Hua Sun , Sriram Somasundaram , Edward\u00a0 S. Hu , and Joseph\u00a0 J. Lim . 2019 . Composing Complex Skills by Learning Transition Policies . In International Conference on Learning Representations. OpenReview.net. Youngwoon Lee, Shao-Hua Sun, Sriram Somasundaram, Edward\u00a0S. Hu, and Joseph\u00a0J. Lim. 2019. Composing Complex Skills by Learning Transition Policies. In International Conference on Learning Representations. OpenReview.net."},{"key":"e_1_3_2_1_9_1","unstructured":"Jiasen Lu Jordi Salvador Roozbeh Mottaghi and Aniruddha Kembhavi. 2022. ASC me to Do Anything: Multi-task Training for Embodied AI. CoRR abs\/2202.06987(2022).  Jiasen Lu Jordi Salvador Roozbeh Mottaghi and Aniruddha Kembhavi. 2022. ASC me to Do Anything: Multi-task Training for Embodied AI. CoRR abs\/2202.06987(2022)."},{"key":"e_1_3_2_1_10_1","doi-asserted-by":"publisher","DOI":"10.1038\/nature14236"},{"key":"e_1_3_2_1_11_1","volume-title":"MCP: Learning Composable Hierarchical Control with Multiplicative Compositional Policies. In Advances in Neural Information Processing Systems, Vol.\u00a032. Curran Associates","author":"Peng Xue\u00a0Bin","year":"2019","unstructured":"Xue\u00a0Bin Peng , Michael Chang , Grace Zhang , Pieter Abbeel , and Sergey Levine . 2019 . MCP: Learning Composable Hierarchical Control with Multiplicative Compositional Policies. In Advances in Neural Information Processing Systems, Vol.\u00a032. Curran Associates , Inc ., 3686\u20133697. Xue\u00a0Bin Peng, Michael Chang, Grace Zhang, Pieter Abbeel, and Sergey Levine. 2019. MCP: Learning Composable Hierarchical Control with Multiplicative Compositional Policies. In Advances in Neural Information Processing Systems, Vol.\u00a032. Curran Associates, Inc., 3686\u20133697."},{"key":"e_1_3_2_1_12_1","doi-asserted-by":"publisher","DOI":"10.1145\/3528223.3530110"},{"key":"e_1_3_2_1_13_1","volume-title":"Accelerating Reinforcement Learning with Learned Skill Priors. In Annual Conference on Robot Learning(Proceedings of Machine Learning Research, Vol.\u00a0155)","author":"Pertsch Karl","year":"2020","unstructured":"Karl Pertsch , Youngwoon Lee , and Joseph\u00a0 J. Lim . 2020 . Accelerating Reinforcement Learning with Learned Skill Priors. In Annual Conference on Robot Learning(Proceedings of Machine Learning Research, Vol.\u00a0155) , Jens Kober, Fabio Ramos, and Claire\u00a0J. Tomlin (Eds.). PMLR, 188\u2013204. Karl Pertsch, Youngwoon Lee, and Joseph\u00a0J. Lim. 2020. Accelerating Reinforcement Learning with Learned Skill Priors. In Annual Conference on Robot Learning(Proceedings of Machine Learning Research, Vol.\u00a0155), Jens Kober, Fabio Ramos, and Claire\u00a0J. Tomlin (Eds.). PMLR, 188\u2013204."},{"key":"e_1_3_2_1_14_1","volume-title":"Demonstration-Guided Reinforcement Learning with Learned Skills. In Annual Conference on Robot Learning(Proceedings of Machine Learning Research, Vol.\u00a0164)","author":"Pertsch Karl","year":"2021","unstructured":"Karl Pertsch , Youngwoon Lee , Yue Wu , and Joseph\u00a0 J. Lim . 2021 . Demonstration-Guided Reinforcement Learning with Learned Skills. In Annual Conference on Robot Learning(Proceedings of Machine Learning Research, Vol.\u00a0164) , Aleksandra Faust, David Hsu, and Gerhard Neumann (Eds.). PMLR, 729\u2013739. Karl Pertsch, Youngwoon Lee, Yue Wu, and Joseph\u00a0J. Lim. 2021. Demonstration-Guided Reinforcement Learning with Learned Skills. In Annual Conference on Robot Learning(Proceedings of Machine Learning Research, Vol.\u00a0164), Aleksandra Faust, David Hsu, and Gerhard Neumann (Eds.). PMLR, 729\u2013739."},{"key":"e_1_3_2_1_15_1","doi-asserted-by":"publisher","DOI":"10.1162\/neco.1991.3.1.88"},{"key":"e_1_3_2_1_16_1","volume-title":"Mastering atari, go, chess and shogi by planning with a learned model. Nat. 588, 7839","author":"Schrittwieser Julian","year":"2020","unstructured":"Julian Schrittwieser , Ioannis Antonoglou , Thomas Hubert , Karen Simonyan , Laurent Sifre , Simon Schmitt , Arthur Guez , Edward Lockhart , Demis Hassabis , Thore Graepel , 2020. Mastering atari, go, chess and shogi by planning with a learned model. Nat. 588, 7839 ( 2020 ), 604\u2013609. Julian Schrittwieser, Ioannis Antonoglou, Thomas Hubert, Karen Simonyan, Laurent Sifre, Simon Schmitt, Arthur Guez, Edward Lockhart, Demis Hassabis, Thore Graepel, 2020. Mastering atari, go, chess and shogi by planning with a learned model. Nat. 588, 7839 (2020), 604\u2013609."},{"key":"e_1_3_2_1_17_1","volume-title":"High-Dimensional Continuous Control Using Generalized Advantage Estimation. In International Conference on Learning Representations, Yoshua Bengio and Yann LeCun (Eds.).","author":"Schulman John","year":"2016","unstructured":"John Schulman , Philipp Moritz , Sergey Levine , Michael\u00a0 I. Jordan , and Pieter Abbeel . 2016 . High-Dimensional Continuous Control Using Generalized Advantage Estimation. In International Conference on Learning Representations, Yoshua Bengio and Yann LeCun (Eds.). John Schulman, Philipp Moritz, Sergey Levine, Michael\u00a0I. Jordan, and Pieter Abbeel. 2016. High-Dimensional Continuous Control Using Generalized Advantage Estimation. In International Conference on Learning Representations, Yoshua Bengio and Yann LeCun (Eds.)."},{"key":"e_1_3_2_1_18_1","unstructured":"John Schulman Filip Wolski Prafulla Dhariwal Alec Radford and Oleg Klimov. 2017. Proximal Policy Optimization Algorithms. CoRR abs\/1707.06347(2017). arxiv:1707.06347  John Schulman Filip Wolski Prafulla Dhariwal Alec Radford and Oleg Klimov. 2017. Proximal Policy Optimization Algorithms. CoRR abs\/1707.06347(2017). arxiv:1707.06347"},{"key":"e_1_3_2_1_19_1","doi-asserted-by":"publisher","DOI":"10.24963\/ijcai.2021"},{"key":"e_1_3_2_1_20_1","volume-title":"Parrot: Data-Driven Behavioral Priors for Reinforcement Learning. CoRR abs\/2011.10024(2020). arxiv:2011.10024","author":"Singh Avi","year":"2020","unstructured":"Avi Singh , Huihan Liu , Gaoyue Zhou , Albert Yu , Nicholas Rhinehart , and Sergey Levine . 2020 . Parrot: Data-Driven Behavioral Priors for Reinforcement Learning. CoRR abs\/2011.10024(2020). arxiv:2011.10024 Avi Singh, Huihan Liu, Gaoyue Zhou, Albert Yu, Nicholas Rhinehart, and Sergey Levine. 2020. Parrot: Data-Driven Behavioral Priors for Reinforcement Learning. CoRR abs\/2011.10024(2020). arxiv:2011.10024"},{"key":"e_1_3_2_1_21_1","volume-title":"Independent Skill Transfer for Deep Reinforcement Learning. In International Joint Conference on Artificial Intelligence, Christian Bessiere (Ed.). ijcai.org, 2901\u20132907","author":"Tian Qiangxing","year":"2020","unstructured":"Qiangxing Tian , Guanchu Wang , Jinxin Liu , Donglin Wang , and Yachen Kang . 2020 . Independent Skill Transfer for Deep Reinforcement Learning. In International Joint Conference on Artificial Intelligence, Christian Bessiere (Ed.). ijcai.org, 2901\u20132907 . Qiangxing Tian, Guanchu Wang, Jinxin Liu, Donglin Wang, and Yachen Kang. 2020. Independent Skill Transfer for Deep Reinforcement Learning. In International Joint Conference on Artificial Intelligence, Christian Bessiere (Ed.). ijcai.org, 2901\u20132907."},{"key":"e_1_3_2_1_22_1","volume-title":"Unsupervised Visual Attention and Invariance for Reinforcement Learning. In Conference on Computer Vision and Pattern Recognition. Computer Vision Foundation \/ IEEE, 6677\u20136687","author":"Wang Xudong","year":"2021","unstructured":"Xudong Wang , Long Lian , and Stella\u00a0 X. Yu . 2021 . Unsupervised Visual Attention and Invariance for Reinforcement Learning. In Conference on Computer Vision and Pattern Recognition. Computer Vision Foundation \/ IEEE, 6677\u20136687 . Xudong Wang, Long Lian, and Stella\u00a0X. Yu. 2021. Unsupervised Visual Attention and Invariance for Reinforcement Learning. In Conference on Computer Vision and Pattern Recognition. Computer Vision Foundation \/ IEEE, 6677\u20136687."},{"key":"e_1_3_2_1_23_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2021.findings-emnlp.116"},{"key":"e_1_3_2_1_24_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICRA46639.2022.9811668"},{"key":"e_1_3_2_1_25_1","volume-title":"Matthew","author":"You Heng","year":"2022","unstructured":"Heng You , Tianpei Yang , Yan Zheng , Jianye Hao , and E. Taylor , Matthew . 2022 . Cross-domain adaptive transfer reinforcement learning based on state-action correspondence. In UAI, Vol.\u00a0180. PMLR , 2299\u20132309. Heng You, Tianpei Yang, Yan Zheng, Jianye Hao, and E. Taylor, Matthew. 2022. Cross-domain adaptive transfer reinforcement learning based on state-action correspondence. In UAI, Vol.\u00a0180. PMLR, 2299\u20132309."},{"key":"e_1_3_2_1_26_1","unstructured":"Zhuangdi Zhu Kaixiang Lin and Jiayu Zhou. 2020. Transfer Learning in Deep Reinforcement Learning: A Survey. CoRR abs\/2009.07888(2020). arXiv:2009.07888  Zhuangdi Zhu Kaixiang Lin and Jiayu Zhou. 2020. Transfer Learning in Deep Reinforcement Learning: A Survey. CoRR abs\/2009.07888(2020). arXiv:2009.07888"}],"event":{"name":"CSAI 2022: 2022 6th International Conference on Computer Science and Artificial Intelligence","acronym":"CSAI 2022","location":"Beijing China"},"container-title":["Proceedings of the 2022 6th International Conference on Computer Science and Artificial Intelligence"],"original-title":[],"link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3577530.3577562","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3577530.3577562","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T16:47:35Z","timestamp":1750178855000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3577530.3577562"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,12,9]]},"references-count":26,"alternative-id":["10.1145\/3577530.3577562","10.1145\/3577530"],"URL":"https:\/\/doi.org\/10.1145\/3577530.3577562","relation":{},"subject":[],"published":{"date-parts":[[2022,12,9]]},"assertion":[{"value":"2023-03-30","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}