{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,11]],"date-time":"2026-07-11T03:25:29Z","timestamp":1783740329148,"version":"3.55.0"},"reference-count":54,"publisher":"Association for Computing Machinery (ACM)","issue":"6","license":[{"start":{"date-parts":[[2023,12,5]],"date-time":"2023-12-05T00:00:00Z","timestamp":1701734400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"name":"the Stanford Institute for Human-Centered AI"},{"name":"ONR MURI","award":["N00014-22-1-2740"],"award-info":[{"award-number":["N00014-22-1-2740"]}]},{"name":"Wu Tsai Human Performance Alliance"},{"DOI":"10.13039\/100019827","name":"Meta","doi-asserted-by":"publisher","id":[{"id":"10.13039\/100019827","id-type":"DOI","asserted-by":"publisher"}]},{"name":"the Toyota Research Institute"},{"name":"NSF CCRI","award":["2120095"],"award-info":[{"award-number":["2120095"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Graph."],"published-print":{"date-parts":[[2023,12,5]]},"abstract":"<jats:p>Modeling human behaviors in contextual environments has a wide range of applications in character animation, embodied AI, VR\/AR, and robotics. In real-world scenarios, humans frequently interact with the environment and manipulate various objects to complete daily tasks. In this work, we study the problem of full-body human motion synthesis for the manipulation of large-sized objects. We propose Object MOtion guided human MOtion synthesis (OMOMO), a conditional diffusion framework that can generate full-body manipulation behaviors from only the object motion. Since naively applying diffusion models fails to precisely enforce contact constraints between the hands and the object, OMOMO learns two separate denoising processes to first predict hand positions from object motion and subsequently synthesize full-body poses based on the predicted hand positions. By employing the hand positions as an intermediate representation between the two denoising processes, we can explicitly enforce contact constraints, resulting in more physically plausible manipulation motions. With the learned model, we develop a novel system that captures full-body human manipulation motions by simply attaching a smartphone to the object being manipulated. Through extensive experiments, we demonstrate the effectiveness of our proposed pipeline and its ability to generalize to unseen objects. Additionally, as high-quality human-object interaction datasets are scarce, we collect a large-scale dataset consisting of 3D object geometry, object motion, and human motion. Our dataset contains human-object interaction motion for 15 objects, with a total duration of approximately 10 hours.<\/jats:p>","DOI":"10.1145\/3618333","type":"journal-article","created":{"date-parts":[[2023,12,5]],"date-time":"2023-12-05T10:20:48Z","timestamp":1701771648000},"page":"1-11","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":84,"title":["Object Motion Guided Human Motion Synthesis"],"prefix":"10.1145","volume":"42","author":[{"given":"Jiaman","family":"Li","sequence":"first","affiliation":[{"name":"Stanford University, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Jiajun","family":"Wu","sequence":"additional","affiliation":[{"name":"Stanford University, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"C. Karen","family":"Liu","sequence":"additional","affiliation":[{"name":"Stanford University, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2023,12,5]]},"reference":[{"key":"e_1_2_2_1_1","unstructured":"Luma AI. 2023. Capture 3D. https:\/\/lumalabs.ai\/"},{"key":"e_1_2_2_2_1","volume-title":"CIRCLE: Capture In Rich Contextual Environments. In Conference on Computer Vision and Pattern Recognition (CVPR).","author":"Araujo Joao Pedro","year":"2023","unstructured":"Joao Pedro Araujo, Jiaman Li, Karthik Vetrivel, Rishi Agarwal, Deepak Gopinath, Jiajun Wu, Alexander Clegg, and C Karen Liu. 2023. CIRCLE: Capture In Rich Contextual Environments. In Conference on Computer Vision and Pattern Recognition (CVPR)."},{"key":"e_1_2_2_3_1","volume-title":"BEHAVE: Dataset and Method for Tracking Human Object Interactions. In Conference on Computer Vision and Pattern Recognition (CVPR). 15935--15946","author":"Bhatnagar Bharat Lal","year":"2022","unstructured":"Bharat Lal Bhatnagar, Xianghui Xie, Ilya A. Petrov, Cristian Sminchisescu, Christian Theobalt, and Gerard Pons-Moll. 2022. BEHAVE: Dataset and Method for Tracking Human Object Interactions. In Conference on Computer Vision and Pattern Recognition (CVPR). 15935--15946."},{"key":"e_1_2_2_4_1","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v35i7.16736"},{"key":"e_1_2_2_5_1","volume-title":"Vladislav Golyanik, and Christian Theobalt.","author":"Dabral Rishabh","year":"2023","unstructured":"Rishabh Dabral, Muhammad Hamza Mughal, Vladislav Golyanik, and Christian Theobalt. 2023. MoFusion: A Framework for Denoising-Diffusion-based Motion Synthesis. In Computer Vision and Pattern Recognition (CVPR)."},{"key":"e_1_2_2_6_1","unstructured":"Angela Dai Angel X Chang Manolis Savva Maciej Halber Thomas Funkhouser and Matthias Nie\u00dfner. 2017. ScanNet: Richly-annotated 3D Reconstructions of Indoor Scenes. In Computer Vision and Pattern Recognition (CVPR). 5828--5839."},{"key":"e_1_2_2_7_1","volume-title":"ARCTIC: A Dataset for Dexterous Bimanual Hand-Object Manipulation. In Conference on Computer Vision and Pattern Recognition (CVPR).","author":"Fan Zicong","year":"2023","unstructured":"Zicong Fan, Omid Taheri, Dimitrios Tzionas, Muhammed Kocabas, Manuel Kaufmann, Michael J. Black, and Otmar Hilliges. 2023. ARCTIC: A Dataset for Dexterous Bimanual Hand-Object Manipulation. In Conference on Computer Vision and Pattern Recognition (CVPR)."},{"key":"e_1_2_2_8_1","doi-asserted-by":"crossref","unstructured":"Anindita Ghosh Rishabh Dabral Vladislav Golyanik Christian Theobalt and Philipp Slusallek. 2023. IMoS: Intent-Driven Full-Body Motion Synthesis for Human-Object Interactions. In Eurographics.","DOI":"10.1111\/cgf.14739"},{"key":"e_1_2_2_9_1","volume-title":"Interaction Replica: Tracking human-object interaction and scene changes from human motion. In arXiv.","author":"Guzov Vladimir","year":"2023","unstructured":"Vladimir Guzov, Julian Chibane, Riccardo Marin, Yannan He, Torsten Sattler, and Gerard Pons-Moll. 2023. Interaction Replica: Tracking human-object interaction and scene changes from human motion. In arXiv."},{"key":"e_1_2_2_10_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR46437.2021.00430"},{"key":"e_1_2_2_11_1","volume-title":"Stochastic Scene-Aware Motion Prediction. In International Conference on Computer Vision (ICCV). 11354--11364","author":"Hassan Mohamed","year":"2021","unstructured":"Mohamed Hassan, Duygu Ceylan, Ruben Villegas, Jun Saito, Jimei Yang, Yi Zhou, and Michael Black. 2021. Stochastic Scene-Aware Motion Prediction. In International Conference on Computer Vision (ICCV). 11354--11364."},{"key":"e_1_2_2_12_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2019.00237"},{"key":"e_1_2_2_13_1","doi-asserted-by":"crossref","unstructured":"Mohamed Hassan Yunrong Guo Tingwu Wang Michael Black Sanja Fidler and Xue Bin Peng. 2023. Synthesizing Physical Character-Scene Interactions. (2023) 1--9.","DOI":"10.1145\/3588432.3591525"},{"key":"e_1_2_2_14_1","volume-title":"Nemf: Neural motion fields for kinematic animation. Advances in Neural Information Processing Systems (NeurIPS)","author":"He Chengan","year":"2022","unstructured":"Chengan He, Jun Saito, James Zachary, Holly Rushmeier, and Yi Zhou. 2022. Nemf: Neural motion fields for kinematic animation. Advances in Neural Information Processing Systems (NeurIPS) (2022)."},{"key":"e_1_2_2_15_1","first-page":"6840","article-title":"Denoising diffusion probabilistic models","volume":"33","author":"Ho Jonathan","year":"2020","unstructured":"Jonathan Ho, Ajay Jain, and Pieter Abbeel. 2020. Denoising diffusion probabilistic models. Advances in Neural Information Processing Systems (NeurIPS) 33 (2020), 6840--6851.","journal-title":"Advances in Neural Information Processing Systems (NeurIPS)"},{"key":"e_1_2_2_16_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52729.2023.01607"},{"key":"e_1_2_2_17_1","volume-title":"International Conference on Learning Representations (ICLR).","author":"Kingma Diederik P","year":"2015","unstructured":"Diederik P Kingma and Jimmy Ba. 2015. Adam: A method for stochastic optimization. In International Conference on Learning Representations (ICLR)."},{"key":"e_1_2_2_18_1","unstructured":"Marian Kleineberg. 2023. Mesh2SDF. https:\/\/github.com\/marian42\/mesh_to_sdf"},{"key":"e_1_2_2_19_1","volume-title":"Locomotion-Action-Manipulation: Synthesizing Human-Scene Interactions in Complex 3D Environments. arXiv preprint arXiv:2301.02667","author":"Lee Jiye","year":"2023","unstructured":"Jiye Lee and Hanbyul Joo. 2023. Locomotion-Action-Manipulation: Synthesizing Human-Scene Interactions in Complex 3D Environments. arXiv preprint arXiv:2301.02667 (2023)."},{"key":"e_1_2_2_20_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52729.2023.01644"},{"key":"e_1_2_2_21_1","doi-asserted-by":"publisher","DOI":"10.1109\/TVCG.2007.1033"},{"key":"e_1_2_2_22_1","doi-asserted-by":"publisher","DOI":"10.1145\/3197517.3201315"},{"key":"e_1_2_2_23_1","doi-asserted-by":"publisher","DOI":"10.1145\/2661229.2661273"},{"key":"e_1_2_2_24_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2019.00554"},{"key":"e_1_2_2_25_1","doi-asserted-by":"publisher","DOI":"10.1145\/3386569.3392474"},{"key":"e_1_2_2_26_1","volume-title":"Generating Continual Human Motion in Diverse 3D Scenes. arXiv preprint arXiv:2304.02061","author":"Mir Aymen","year":"2023","unstructured":"Aymen Mir, Xavier Puig, Angjoo Kanazawa, and Gerard Pons-Moll. 2023. Generating Continual Human Motion in Diverse 3D Scenes. arXiv preprint arXiv:2304.02061 (2023)."},{"key":"e_1_2_2_27_1","volume-title":"PyTorch: An Imperative Style","author":"Paszke Adam","unstructured":"Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, Alban Desmaison, Andreas Kopf, Edward Yang, Zachary DeVito, Martin Raison, Alykhan Tejani, Sasank Chilamkurthy, Benoit Steiner, Lu Fang, Junjie Bai, and Soumith Chintala. 2019. PyTorch: An Imperative Style, High-Performance Deep Learning Library. In Advances in Neural Information Processing Systems (NeurIPS)."},{"key":"e_1_2_2_28_1","volume-title":"Conference on Computer Vision and Pattern Recognition (CVPR). 10975--10985","author":"Pavlakos Georgios","unstructured":"Georgios Pavlakos, Vasileios Choutas, Nima Ghorbani, Timo Bolkart, Ahmed A. A. Osman, Dimitrios Tzionas, and Michael J. Black. 2019. Expressive Body Capture: 3D Hands, Face, and Body From a Single Image. In Conference on Computer Vision and Pattern Recognition (CVPR). 10975--10985."},{"key":"e_1_2_2_29_1","doi-asserted-by":"publisher","DOI":"10.1145\/3197517.3201311"},{"key":"e_1_2_2_30_1","doi-asserted-by":"publisher","DOI":"10.1145\/3450626.3459670"},{"key":"e_1_2_2_31_1","volume-title":"International Conference on Computer Vision (ICCV). 4332--4341","author":"Prokudin Sergey","year":"2019","unstructured":"Sergey Prokudin, Christoph Lassner, and Javier Romero. 2019. Efficient learning on point clouds with basis point sets. In International Conference on Computer Vision (ICCV). 4332--4341."},{"key":"e_1_2_2_32_1","volume-title":"Conference on Computer Vision and Pattern Recognition (CVPR). 652--660","author":"Qi Charles R","year":"2017","unstructured":"Charles R Qi, Hao Su, Kaichun Mo, and Leonidas J Guibas. 2017. Pointnet: Deep learning on point sets for 3d classification and segmentation. In Conference on Computer Vision and Pattern Recognition (CVPR). 652--660."},{"key":"e_1_2_2_33_1","first-page":"7462","article-title":"Implicit neural representations with periodic activation functions","volume":"33","author":"Sitzmann Vincent","year":"2020","unstructured":"Vincent Sitzmann, Julien Martel, Alexander Bergman, David Lindell, and Gordon Wetzstein. 2020. Implicit neural representations with periodic activation functions. Advances in Neural Information Processing Systems (NeurIPS) 33 (2020), 7462--7473.","journal-title":"Advances in Neural Information Processing Systems (NeurIPS)"},{"key":"e_1_2_2_34_1","doi-asserted-by":"publisher","DOI":"10.1145\/3355089.3356505"},{"key":"e_1_2_2_35_1","volume-title":"Michael Goesele, Steven Lovegrove, and Richard Newcombe.","author":"Straub Julian","year":"2019","unstructured":"Julian Straub, Thomas Whelan, Lingni Ma, Yufan Chen, Erik Wijmans, Simon Green, Jakob J. Engel, Raul Mur-Artal, Carl Ren, Shobhit Verma, Anton Clarkson, Mingfei Yan, Brian Budge, Yajie Yan, Xiaqing Pan, June Yon, Yuyang Zou, Kimberly Leon, Nigel Carter, Jesus Briales, Tyler Gillingham, Elias Mueggler, Luis Pesqueira, Manolis Savva, Dhruv Batra, Hauke M. Strasdat, Renzo De Nardi, Michael Goesele, Steven Lovegrove, and Richard Newcombe. 2019. The Replica Dataset: A Digital Replica of Indoor Spaces. arXiv preprint arXiv:1906.05797 (2019)."},{"key":"e_1_2_2_36_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52688.2022.01291"},{"key":"e_1_2_2_37_1","volume-title":"GRAB: A Dataset of Whole-Body Human Grasping of Objects. In European Conference on Computer Vision (ECCV). https:\/\/grab.is.tue.mpg.de","author":"Taheri Omid","year":"2020","unstructured":"Omid Taheri, Nima Ghorbani, Michael J. Black, and Dimitrios Tzionas. 2020. GRAB: A Dataset of Whole-Body Human Grasping of Objects. In European Conference on Computer Vision (ECCV). https:\/\/grab.is.tue.mpg.de"},{"key":"e_1_2_2_38_1","volume-title":"Human Motion Diffusion Model. In International Conference on Learning Representations (ICLR).","author":"Tevet Guy","year":"2023","unstructured":"Guy Tevet, Sigal Raab, Brian Gordon, Yonatan Shafir, Amit H Bermano, and Daniel Cohen-Or. 2023. Human Motion Diffusion Model. In International Conference on Learning Representations (ICLR)."},{"key":"e_1_2_2_39_1","volume-title":"EDGE: Editable Dance Generation From Music. In Computer Vision and Pattern Recognition (CVPR).","author":"Tseng Jonathan","year":"2023","unstructured":"Jonathan Tseng, Rodrigo Castellon, and C Karen Liu. 2023. EDGE: Editable Dance Generation From Music. In Computer Vision and Pattern Recognition (CVPR)."},{"key":"e_1_2_2_40_1","volume-title":"Advances in Neural Information Processing Systems (NIPS)","volume":"30","author":"Vaswani Ashish","year":"2017","unstructured":"Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, \u0141ukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. In Advances in Neural Information Processing Systems (NIPS), Vol. 30."},{"key":"e_1_2_2_41_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR46437.2021.00928"},{"key":"e_1_2_2_42_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR46437.2021.01203"},{"key":"e_1_2_2_43_1","volume-title":"HUMANISE: Language-conditioned Human Motion Generation in 3D Scenes. In Advances in Neural Information Processing Systems (NeurIPS).","author":"Wang Zan","year":"2022","unstructured":"Zan Wang, Yixin Chen, Tengyu Liu, Yixin Zhu, Wei Liang, and Siyuan Huang. 2022. HUMANISE: Language-conditioned Human Motion Generation in 3D Scenes. In Advances in Neural Information Processing Systems (NeurIPS)."},{"key":"e_1_2_2_44_1","volume-title":"Saga: Stochastic Whole-Body Grasping with Contact. In European Conference on Computer Vision (ECCV). 257--274","author":"Wu Yan","year":"2022","unstructured":"Yan Wu, Jiahao Wang, Yan Zhang, Siwei Zhang, Otmar Hilliges, Fisher Yu, and Siyu Tang. 2022. Saga: Stochastic Whole-Body Grasping with Contact. In European Conference on Computer Vision (ECCV). 257--274."},{"key":"e_1_2_2_45_1","volume-title":"ACM SIGGRAPH 2022 Conference Proceedings. 1--9.","author":"Xie Zhaoming","unstructured":"Zhaoming Xie, Sebastian Starke, Hung Yu Ling, and Michiel van de Panne. 2022. Learning soccer juggling skills with layer-wise mixture-of-experts. In ACM SIGGRAPH 2022 Conference Proceedings. 1--9."},{"key":"e_1_2_2_46_1","unstructured":"Zhaoming Xie Jonathan Tseng Sebastian Starke Michiel van de Panne and C Karen Liu. 2023. Hierarchical Planning and Control for Box Loco-Manipulation. (2023)."},{"key":"e_1_2_2_47_1","doi-asserted-by":"publisher","DOI":"10.1145\/2185520.2185537"},{"key":"e_1_2_2_48_1","doi-asserted-by":"publisher","DOI":"10.1145\/3478513.3480500"},{"key":"e_1_2_2_49_1","volume-title":"MotionDiffuse: Text-Driven Human Motion Generation with Diffusion Model. arXiv preprint arXiv:2208.15001","author":"Zhang Mingyuan","year":"2022","unstructured":"Mingyuan Zhang, Zhongang Cai, Liang Pan, Fangzhou Hong, Xinying Guo, Lei Yang, and Ziwei Liu. 2022b. MotionDiffuse: Text-Driven Human Motion Generation with Diffusion Model. arXiv preprint arXiv:2208.15001 (2022)."},{"key":"e_1_2_2_50_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-031-20068-7_11"},{"key":"e_1_2_2_51_1","volume-title":"COUCH: Towards Controllable Human-Chair Interactions. In European Conference on Computer Vision (ECCV). 518--535","author":"Zhang Xiaohan","year":"2022","unstructured":"Xiaohan Zhang, Bharat Lal Bhatnagar, Sebastian Starke, Vladimir Guzov, and Gerard Pons-Moll. 2022a. COUCH: Towards Controllable Human-Chair Interactions. In European Conference on Computer Vision (ECCV). 518--535."},{"key":"e_1_2_2_52_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV51070.2023.01354"},{"key":"e_1_2_2_53_1","volume-title":"GIMO: Gaze-Informed Human Motion Prediction in Context. In European Conference on Computer Vision (ECCV).","author":"Zheng Yang","year":"2022","unstructured":"Yang Zheng, Yanchao Yang, Kaichun Mo, Jiaman Li, Tao Yu, Yebin Liu, Karen Liu, and Leonidas Guibas. 2022. GIMO: Gaze-Informed Human Motion Prediction in Context. In European Conference on Computer Vision (ECCV)."},{"key":"e_1_2_2_54_1","doi-asserted-by":"crossref","unstructured":"Yi Zhou Connelly Barnes Jingwan Lu Jimei Yang and Hao Li. 2019. On the continuity of rotation representations in neural networks. In Computer Vision and Pattern Recognition (CVPR).","DOI":"10.1109\/CVPR.2019.00589"}],"container-title":["ACM Transactions on Graphics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3618333","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3618333","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,8,21]],"date-time":"2025-08-21T10:53:41Z","timestamp":1755773621000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3618333"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,12,5]]},"references-count":54,"journal-issue":{"issue":"6","published-print":{"date-parts":[[2023,12,5]]}},"alternative-id":["10.1145\/3618333"],"URL":"https:\/\/doi.org\/10.1145\/3618333","relation":{},"ISSN":["0730-0301","1557-7368"],"issn-type":[{"value":"0730-0301","type":"print"},{"value":"1557-7368","type":"electronic"}],"subject":[],"published":{"date-parts":[[2023,12,5]]},"assertion":[{"value":"2023-12-05","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}