{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,2,7]],"date-time":"2026-02-07T00:48:51Z","timestamp":1770425331413,"version":"3.49.0"},"reference-count":39,"publisher":"Association for Computing Machinery (ACM)","issue":"4","funder":[{"DOI":"10.13039\/100014440","name":"Ministerio de Ciencia, Innovaci\u00f3n y Universidades","doi-asserted-by":"publisher","award":["PID2021-122136OB-C21"],"award-info":[{"award-number":["PID2021-122136OB-C21"]}],"id":[{"id":"10.13039\/100014440","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["Proc. ACM Comput. Graph. Interact. Tech."],"published-print":{"date-parts":[[2025,8,31]]},"abstract":"<jats:p>\n            Performing everyday tasks requires both large-scale body movements driven by the arms, legs, and torso, and fine motor skills, particularly in the hands. However, existing reinforcement learning approaches often struggle to efficiently acquire these diverse motion skills from scratch, leading to slow convergence or suboptimal policies. We observed that the motor skills of distinct body parts exhibit a certain level of independence. This suggests the potential advantage of independently pretraining specific dexterous limbs prior to their integration in full-body motion tasks. Inspired by this, we present\n            <jats:italic toggle=\"yes\">Part-wise Heterogeneous Agents (PHA)<\/jats:italic>\n            , a cooperative multi-agent reinforcement learning approach where body parts are treated as independent agents, allowing for specialized skill acquisition and cooperative execution of complex full-body tasks. Furthermore, our method enables the pretraining of fine motor skills, such as gripping a bar or grabbing a climbing hold, before integrating them with other body parts for complex whole-body coordination, thus introducing part-wise\n            <jats:italic toggle=\"yes\">Reusable Policy Priors<\/jats:italic>\n            . We tested our technique on challenging tasks such as rope climbing, rock bouldering and traversing a horizontal ladder. Our approach not only accelerates convergence, but also improves overall policy quality, achieving motion tasks that single-agent approaches struggle to solve. Our results also demonstrate adaptability, enabling Reusable Policy Priors to adjust their policies to successfully perform complex tasks in scenarios not seen during training.\n          <\/jats:p>","DOI":"10.1145\/3747870","type":"journal-article","created":{"date-parts":[[2025,8,8]],"date-time":"2025-08-08T15:33:31Z","timestamp":1754667211000},"page":"1-22","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":1,"title":["PHA: Part-wise Heterogeneous Agents with Reusable Policy Priors for Physics-Based Motion Synthesis 63"],"prefix":"10.1145","volume":"8","author":[{"ORCID":"https:\/\/orcid.org\/0009-0008-8545-148X","authenticated-orcid":false,"given":"Luis","family":"Carranza","sequence":"first","affiliation":[{"name":"Universitat Politecnica de Catalunya","place":["Barcelona, Spain"]}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-3943-1839","authenticated-orcid":false,"given":"Oscar","family":"Argudo","sequence":"additional","affiliation":[{"name":"Universitat Politecnica de Catalunya","place":["Barcelona, Spain"]}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-8480-4713","authenticated-orcid":false,"given":"Carlos","family":"Andujar","sequence":"additional","affiliation":[{"name":"Universitat Politecnica de Catalunya","place":["Barcelona, Spain"]}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2025,8,8]]},"reference":[{"key":"e_1_3_3_2_1","unstructured":"Adobe. 2020. Adobe\u2019s Mixamo. (2020). https:\/\/www.mixamo.com"},{"key":"e_1_3_3_3_1","doi-asserted-by":"publisher","DOI":"10.1145\/3588432.3591487"},{"key":"e_1_3_3_4_1","doi-asserted-by":"publisher","DOI":"10.1109\/3DV62453.2024.00109"},{"key":"e_1_3_3_5_1","unstructured":"Zixuan Chen Xialin He Yen-Jen Wang Qiayuan Liao Yanjie Ze Zhongyu Li S.\u00a0Shankar Sastry Jiajun Wu Koushil Sreenath Saurabh Gupta and Xue\u00a0Bin Peng. 2024. Learning Smooth Humanoid Locomotion through Lipschitz-Constrained Policies. arxiv preprint arXiv:https:\/\/arXiv.org\/abs\/2410.11825 (2024)."},{"key":"e_1_3_3_6_1","doi-asserted-by":"crossref","unstructured":"Sammy Christen Muhammed Kocabas Emre Aksan Jemin Hwangbo Jie Song and Otmar Hilliges. 2022. D-Grasp: Physically Plausible Dynamic Grasp Synthesis for Hand-Object Interactions. arxiv:https:\/\/arXiv.org\/abs\/2112.03028\u00a0[cs.CV] https:\/\/arxiv.org\/abs\/2112.03028","DOI":"10.1109\/CVPR52688.2022.01992"},{"key":"e_1_3_3_7_1","doi-asserted-by":"publisher","DOI":"10.1145\/3664647.3681582"},{"key":"e_1_3_3_8_1","doi-asserted-by":"publisher","DOI":"10.1145\/3610548.3618205"},{"key":"e_1_3_3_9_1","doi-asserted-by":"publisher","unstructured":"Levi Fussell Kevin Bergamin and Daniel Holden. 2021. SuperTrack: motion tracking for physically simulated characters using supervised learning. ACM Trans. Graph. 40 6 Article 63 (Dec. 2021) 13\u00a0pages. 10.1145\/3478513.3480527","DOI":"10.1145\/3478513.3480527"},{"key":"e_1_3_3_10_1","doi-asserted-by":"publisher","unstructured":"Anindita Ghosh Rishabh Dabral Vladislav Golyanik Christian Theobalt and Philipp Slusallek. 2023. IMoS: Intent-Driven Full-Body Motion Synthesis for Human-Object Interactions. Computer Graphics Forum 42 2 (2023) 1\u201312. 10.1111\/cgf.14739 arXiv:https:\/\/onlinelibrary.wiley.com\/doi\/pdf\/10.1111\/cgf.14739","DOI":"10.1111\/cgf.14739"},{"key":"e_1_3_3_11_1","doi-asserted-by":"publisher","unstructured":"F\u00e9lix\u00a0G. Harvey Mike Yurick Derek Nowrouzezahrai and Christopher Pal. 2020. Robust motion in-betweening. ACM Transactions on Graphics 39 4 (Aug. 2020). 10.1145\/3386569.3392480","DOI":"10.1145\/3386569.3392480"},{"key":"e_1_3_3_12_1","doi-asserted-by":"publisher","unstructured":"Daniel Holden Taku Komura and Jun Saito. 2017. Phase-functioned neural networks for character control. ACM Trans. Graph. 36 4 Article 63 (July 2017) 13\u00a0pages. 10.1145\/3072959.3073663","DOI":"10.1145\/3072959.3073663"},{"key":"e_1_3_3_13_1","unstructured":"Hanwen Jiang Shaowei Liu Jiashun Wang and Xiaolong Wang. 2021. Hand-Object Contact Consistency Reasoning for Human Grasps Generation. arxiv:https:\/\/arXiv.org\/abs\/2104.03304\u00a0[cs.CV] https:\/\/arxiv.org\/abs\/2104.03304"},{"key":"e_1_3_3_14_1","unstructured":"Korrawe Karunratanakul Jinlong Yang Yan Zhang Michael Black Krikamol Muandet and Siyu Tang. 2020. Grasping Field: Learning Implicit Representations for Human Grasps. arxiv:https:\/\/arXiv.org\/abs\/2008.04451\u00a0[cs.CV] https:\/\/arxiv.org\/abs\/2008.04451"},{"key":"e_1_3_3_15_1","doi-asserted-by":"publisher","DOI":"10.1109\/HUMANOIDS.2015.7363441"},{"key":"e_1_3_3_16_1","doi-asserted-by":"publisher","unstructured":"Ariel Kwiatkowski Eduardo Alvarado Vicky Kalogeiton C.\u00a0Karen Liu Julien Pettr\u00e9 Michiel van\u00a0de Panne and Marie-Paule Cani. 2022. A Survey on Reinforcement Learning Methods in Character Animation. Computer Graphics Forum 41 2 (2022) 613\u2013639. 10.1111\/cgf.14504 arXiv:https:\/\/onlinelibrary.wiley.com\/doi\/pdf\/10.1111\/cgf.14504","DOI":"10.1111\/cgf.14504"},{"key":"e_1_3_3_17_1","doi-asserted-by":"publisher","unstructured":"Kang\u00a0Hoon Lee Myung\u00a0Geol Choi and Jehee Lee. 2006. Motion patches: building blocks for virtual environments annotated with motion data. ACM Trans. Graph. 25 3 (July 2006) 898\u2013906. 10.1145\/1141911.1141972","DOI":"10.1145\/1141911.1141972"},{"key":"e_1_3_3_18_1","doi-asserted-by":"publisher","unstructured":"Cheng Li Levi Fussell and Taku Komura. 2021. Multi-agent reinforcement learning for character control. The Visual Computer 37 12 (01 Dec 2021) 3115\u20133123. 10.1007\/s00371-021-02269-1","DOI":"10.1007\/s00371-021-02269-1"},{"key":"e_1_3_3_19_1","doi-asserted-by":"publisher","unstructured":"Libin Liu and Jessica Hodgins. 2018. Learning basketball dribbling skills using trajectory optimization and deep reinforcement learning. ACM Trans. Graph. 37 4 Article 63 (July 2018) 14\u00a0pages. 10.1145\/3197517.3201315","DOI":"10.1145\/3197517.3201315"},{"key":"e_1_3_3_20_1","doi-asserted-by":"publisher","DOI":"10.5555\/3295222.3295385"},{"key":"e_1_3_3_21_1","volume-title":"The Twelfth International Conference on Learning Representations","author":"Luo Zhengyi","year":"2024","unstructured":"Zhengyi Luo, Jinkun Cao, Josh Merel, Alexander Winkler, Jing Huang, Kris\u00a0M. Kitani, and Weipeng Xu. 2024. Universal Humanoid Motion Representations for Physics-Based Control. In The Twelfth International Conference on Learning Representations. https:\/\/openreview.net\/forum?id=OrOd8PxOO2"},{"key":"e_1_3_3_22_1","unstructured":"Viktor Makoviychuk Lukasz Wawrzyniak Yunrong Guo Michelle Lu Kier Storey Miles Macklin David Hoeller Nikita Rudin Arthur Allshire Ankur Handa and Gavriel State. 2021. Isaac Gym: High Performance GPU-Based Physics Simulation For Robot Learning. CoRR abs\/2108.10470 (2021). arXiv:https:\/\/arXiv.org\/abs\/2108.10470https:\/\/arxiv.org\/abs\/2108.10470"},{"key":"e_1_3_3_23_1","doi-asserted-by":"publisher","unstructured":"Josh Merel Saran Tunyasuvunakool Arun Ahuja Yuval Tassa Leonard Hasenclever Vu Pham Tom Erez Greg Wayne and Nicolas Heess. 2020. Catch & Carry: reusable neural controllers for vision-guided whole-body tasks. ACM Trans. Graph. 39 4 Article 63 (Aug. 2020) 14\u00a0pages. 10.1145\/3386569.3392474","DOI":"10.1145\/3386569.3392474"},{"key":"e_1_3_3_24_1","unstructured":"Liang Pan Zeshi Yang Zhiyang Dou Wenjia Wang Buzhen Huang Bo Dai Taku Komura and Jingbo Wang. 2025. TokenHSI: Unified Synthesis of Physical Human-Scene Interactions through Task Tokenization. arxiv:https:\/\/arXiv.org\/abs\/2503.19901\u00a0[cs.CV] https:\/\/arxiv.org\/abs\/2503.19901"},{"key":"e_1_3_3_25_1","doi-asserted-by":"publisher","unstructured":"Xue\u00a0Bin Peng Pieter Abbeel Sergey Levine and Michiel van\u00a0de Panne. 2018. DeepMimic: Example-guided Deep Reinforcement Learning of Physics-based Character Skills. ACM Trans. Graph. 37 4 Article 63 (July 2018) 14\u00a0pages. 10.1145\/3197517.3201311","DOI":"10.1145\/3197517.3201311"},{"key":"e_1_3_3_26_1","doi-asserted-by":"crossref","unstructured":"Xue\u00a0Bin Peng Yunrong Guo Lina Halper Sergey Levine and Sanja Fidler. 2022. ASE: Large-scale Reusable Adversarial Skill Embeddings for Physically Simulated Characters. ACM Trans. Graph. 41 4 Article 63 (July 2022).","DOI":"10.1145\/3528223.3530110"},{"key":"e_1_3_3_27_1","doi-asserted-by":"publisher","unstructured":"Xue\u00a0Bin Peng Ze Ma Pieter Abbeel Sergey Levine and Angjoo Kanazawa. 2021. AMP: Adversarial Motion Priors for Stylized Physics-Based Character Control. ACM Trans. Graph. 40 4 Article 63 (July 2021) 15\u00a0pages. 10.1145\/3450626.3459670","DOI":"10.1145\/3450626.3459670"},{"key":"e_1_3_3_28_1","unstructured":"Tabish Rashid Mikayel Samvelyan Christian\u00a0Schroeder De\u00a0Witt Gregory Farquhar Jakob Foerster and Shimon Whiteson. 2020. Monotonic value function factorisation for deep multi-agent reinforcement learning. J. Mach. Learn. Res. 21 1 Article 63 (Jan. 2020) 51\u00a0pages."},{"key":"e_1_3_3_29_1","unstructured":"John Schulman Filip Wolski Prafulla Dhariwal Alec Radford and Oleg Klimov. 2017. Proximal Policy Optimization Algorithms. CoRR abs\/1707.06347 (2017). arXiv:https:\/\/arXiv.org\/abs\/1707.06347http:\/\/arxiv.org\/abs\/1707.06347"},{"key":"e_1_3_3_30_1","doi-asserted-by":"publisher","DOI":"10.1145\/1342250.1342271"},{"key":"e_1_3_3_31_1","doi-asserted-by":"publisher","unstructured":"Sebastian Starke He Zhang Taku Komura and Jun Saito. 2019. Neural state machine for character-scene interactions. ACM Trans. Graph. 38 6 Article 63 (Nov. 2019) 14\u00a0pages. 10.1145\/3355089.3356505","DOI":"10.1145\/3355089.3356505"},{"key":"e_1_3_3_32_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52688.2022.01291"},{"key":"e_1_3_3_33_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-58548-8_34"},{"key":"e_1_3_3_34_1","doi-asserted-by":"publisher","unstructured":"Xiangjun Tang He Wang Bo Hu Xu Gong Ruifan Yi Qilong Kou and Xiaogang Jin. 2022. Real-time controllable motion transition for characters. ACM Trans. Graph. 41 4 Article 63 (July 2022) 10\u00a0pages. 10.1145\/3528223.3530090","DOI":"10.1145\/3528223.3530090"},{"key":"e_1_3_3_35_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52729.2023.02029"},{"key":"e_1_3_3_36_1","doi-asserted-by":"crossref","unstructured":"Chen Tessler Yunrong Guo Ofir Nabati Gal Chechik and Xue\u00a0Bin Peng. 2024. MaskedMimic: Unified Physics-Based Character Control Through Masked Motion Inpainting. ACM Transactions on Graphics (TOG) (2024).","DOI":"10.1145\/3687951"},{"key":"e_1_3_3_37_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52688.2022.01981"},{"key":"e_1_3_3_38_1","series-title":"(NIPS \u201922)","volume-title":"Proceedings of the 36th International Conference on Neural Information Processing Systems","author":"Yu Chao","year":"2022","unstructured":"Chao Yu, Akash Velu, Eugene Vinitsky, Jiaxuan Gao, Yu Wang, Alexandre Bayen, and Yi Wu. 2022. The surprising effectiveness of PPO in cooperative multi-agent games. In Proceedings of the 36th International Conference on Neural Information Processing Systems (New Orleans, LA, USA) (NIPS \u201922). Curran Associates Inc., Red Hook, NY, USA, Article 63, 14\u00a0pages."},{"key":"e_1_3_3_39_1","doi-asserted-by":"publisher","unstructured":"He Zhang Yuting Ye Takaaki Shiratori and Taku Komura. 2021. ManipNet: neural manipulation synthesis with a hand-object spatial representation. ACM Trans. Graph. 40 4 Article 63 (July 2021) 14\u00a0pages. 10.1145\/3450626.3459830","DOI":"10.1145\/3450626.3459830"},{"key":"e_1_3_3_40_1","unstructured":"Yifan Zhong Jakub\u00a0Grudzien Kuba Xidong Feng Siyi Hu Jiaming Ji and Yaodong Yang. 2024. Heterogeneous-Agent Reinforcement Learning. Journal of Machine Learning Research 25 32 (2024) 1\u201367. http:\/\/jmlr.org\/papers\/v25\/23-0488.html"}],"container-title":["Proceedings of the ACM on Computer Graphics and Interactive Techniques"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3747870","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,8,8]],"date-time":"2025-08-08T16:24:46Z","timestamp":1754670286000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3747870"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,8,8]]},"references-count":39,"journal-issue":{"issue":"4","published-print":{"date-parts":[[2025,8,31]]}},"alternative-id":["10.1145\/3747870"],"URL":"https:\/\/doi.org\/10.1145\/3747870","relation":{},"ISSN":["2577-6193"],"issn-type":[{"value":"2577-6193","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,8,8]]},"assertion":[{"value":"2025-08-08","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}