{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,11]],"date-time":"2026-07-11T17:20:22Z","timestamp":1783790422346,"version":"3.55.0"},"reference-count":82,"publisher":"Association for Computing Machinery (ACM)","issue":"1","license":[{"start":{"date-parts":[[2020,9,4]],"date-time":"2020-09-04T00:00:00Z","timestamp":1599177600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"DOI":"10.13039\/501100003977","name":"Israel Science Foundation","doi-asserted-by":"crossref","award":["2366\/16"],"award-info":[{"award-number":["2366\/16"]}],"id":[{"id":"10.13039\/501100003977","id-type":"DOI","asserted-by":"crossref"}]},{"name":"European Union\u2019s Horizon 2020 Research and Innovation Programme","award":["739578"],"award-info":[{"award-number":["739578"]}]},{"name":"National Key R8D Program of China","award":["2018YFB1403900, 2019YFF0302902"],"award-info":[{"award-number":["2018YFB1403900, 2019YFF0302902"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Graph."],"published-print":{"date-parts":[[2021,2,28]]},"abstract":"<jats:p>\n            We introduce\n            <jats:italic>MotioNet<\/jats:italic>\n            , a deep neural network that directly reconstructs the motion of a 3D human skeleton from a monocular video. While previous methods rely on either rigging or inverse kinematics (IK) to associate a consistent skeleton with temporally coherent joint rotations, our method is the first data-driven approach that directly outputs a kinematic skeleton, which is a complete, commonly used motion representation. At the crux of our approach lies a deep neural network with embedded kinematic priors, which decomposes sequences of 2D joint positions into two separate attributes: a single, symmetric skeleton encoded by bone lengths, and a sequence of 3D joint rotations associated with global root positions and foot contact labels. These attributes are fed into an integrated forward kinematics (FK) layer that outputs 3D positions, which are compared to a ground truth. In addition, an adversarial loss is applied to the velocities of the recovered rotations to ensure that they lie on the manifold of natural joint rotations. The key advantage of our approach is that it learns to infer natural joint rotations directly from the training data rather than assuming an underlying model, or inferring them from joint positions using a data-agnostic IK solver. We show that enforcing a single consistent skeleton along with temporally coherent joint rotations constrains the solution space, leading to a more robust handling of self-occlusions and depth ambiguities.\n          <\/jats:p>","DOI":"10.1145\/3407659","type":"journal-article","created":{"date-parts":[[2020,9,4]],"date-time":"2020-09-04T10:07:55Z","timestamp":1599214075000},"page":"1-15","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":92,"title":["MotioNet"],"prefix":"10.1145","volume":"40","author":[{"given":"Mingyi","family":"Shi","sequence":"first","affiliation":[{"name":"Shandong University, China, and AICFVE, Beijing Film Academy, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Kfir","family":"Aberman","sequence":"additional","affiliation":[{"name":"AICFVE, Beijing Film Academy, China, and Tel-Aviv University, Israel"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Andreas","family":"Aristidou","sequence":"additional","affiliation":[{"name":"University of Cyprus and RISE Research Centre, Cyprus"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Taku","family":"Komura","sequence":"additional","affiliation":[{"name":"Edinburgh University, Japan"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Dani","family":"Lischinski","sequence":"additional","affiliation":[{"name":"Shandong University, China and The Hebrew University of Jerusalem, Israel and AICFVE, Beijing Film Academy, Israel"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Daniel","family":"Cohen-Or","sequence":"additional","affiliation":[{"name":"Tel-Aviv University, Israel, and AICFVE, Beijing Film Academy, Israel"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Baoquan","family":"Chen","sequence":"additional","affiliation":[{"name":"CFCS, Peking University, China, and AICFVE, Beijing Film Academy, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2020,9,4]]},"reference":[{"key":"e_1_2_1_1_1","volume-title":"Proceedings of the 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR\u201915)","author":"Akhter Ijaz","unstructured":"Ijaz Akhter and Michael J. Black . 2015. Pose-conditioned joint angle limits for 3D human pose reconstruction . In Proceedings of the 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR\u201915) . IEEE Computer Society, Washington, DC. Ijaz Akhter and Michael J. Black. 2015. Pose-conditioned joint angle limits for 3D human pose reconstruction. In Proceedings of the 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR\u201915). IEEE Computer Society, Washington, DC."},{"key":"e_1_2_1_2_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00875"},{"key":"e_1_2_1_3_1","doi-asserted-by":"publisher","DOI":"10.1145\/1073204.1073207"},{"key":"e_1_2_1_4_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.00351"},{"key":"e_1_2_1_5_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2011.6126356"},{"key":"e_1_2_1_6_1","doi-asserted-by":"crossref","unstructured":"Didier Bieler Semih G\u00fcnel Pascal Fua and Helge Rhodin. 2019. Gravity as a Reference for Estimating a Person\u2019s Height from Video. arxiv:cs.CV\/1909.02211  Didier Bieler Semih G\u00fcnel Pascal Fua and Helge Rhodin. 2019. Gravity as a Reference for Estimating a Person\u2019s Height from Video. arxiv:cs.CV\/1909.02211","DOI":"10.1109\/ICCV.2019.00866"},{"key":"e_1_2_1_7_1","volume-title":"Proceedings of the European Conference on Computer Vision (ECCV\u201916)","author":"Bogo Federica","unstructured":"Federica Bogo , Angjoo Kanazawa , Christoph Lassner , Peter Gehler , Javier Romero , and Michael J. Black . 2016. Keep it SMPL: Automatic estimation of 3D human pose and shape from a single image . In Proceedings of the European Conference on Computer Vision (ECCV\u201916) . Springer, Berlin, Germany, 561--578. Federica Bogo, Angjoo Kanazawa, Christoph Lassner, Peter Gehler, Javier Romero, and Michael J. Black. 2016. Keep it SMPL: Automatic estimation of 3D human pose and shape from a single image. In Proceedings of the European Conference on Computer Vision (ECCV\u201916). Springer, Berlin, Germany, 561--578."},{"key":"e_1_2_1_8_1","doi-asserted-by":"publisher","DOI":"10.1109\/3DV.2016.84"},{"key":"e_1_2_1_9_1","volume-title":"Proceedings of the 2018 IEEE Conference on Computer Vision and Pattern Recognition (CVPR\u201918)","author":"Cao Zhe","year":"2018","unstructured":"Zhe Cao , Gines Hidalgo , Tomas Simon , Shih-En Wei , and Yaser Sheikh . 2018 . OpenPose: Realtime multi-person 2D pose estimation using part affinity fields . In Proceedings of the 2018 IEEE Conference on Computer Vision and Pattern Recognition (CVPR\u201918) . IEEE Computer Society, Washington, DC. Zhe Cao, Gines Hidalgo, Tomas Simon, Shih-En Wei, and Yaser Sheikh. 2018. OpenPose: Realtime multi-person 2D pose estimation using part affinity fields. In Proceedings of the 2018 IEEE Conference on Computer Vision and Pattern Recognition (CVPR\u201918). IEEE Computer Society, Washington, DC."},{"key":"e_1_2_1_10_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.512"},{"key":"e_1_2_1_11_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.610"},{"key":"e_1_2_1_12_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00742"},{"key":"e_1_2_1_13_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00742"},{"key":"e_1_2_1_14_1","unstructured":"CMU. 2019. CMU Graphics Lab Motion Capture Database. Retrieved from http:\/\/mocap.cs.cmu.edu\/.  CMU. 2019. CMU Graphics Lab Motion Capture Database. Retrieved from http:\/\/mocap.cs.cmu.edu\/."},{"key":"e_1_2_1_15_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-01240-3_41"},{"key":"e_1_2_1_16_1","doi-asserted-by":"publisher","DOI":"10.1145\/3136457.3136460"},{"key":"e_1_2_1_17_1","volume-title":"Proceedings of the AAAI Conference on Artificial Intelligence.","author":"Fang Hao-Shu","year":"2018","unstructured":"Hao-Shu Fang ,* Yuanlu Xu ,* Wenguan Wang , Xiaobai Liu , and Song-Chun Zhu . 2018 . Learning pose grammar to encode human body configuration for 3D pose estimation . In Proceedings of the AAAI Conference on Artificial Intelligence. Hao-Shu Fang,*Yuanlu Xu,*Wenguan Wang, Xiaobai Liu, and Song-Chun Zhu. 2018. Learning pose grammar to encode human body configuration for 3D pose estimation. In Proceedings of the AAAI Conference on Artificial Intelligence."},{"key":"e_1_2_1_18_1","doi-asserted-by":"publisher","DOI":"10.1145\/1186562.1015755"},{"key":"e_1_2_1_19_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.01114"},{"key":"e_1_2_1_20_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00762"},{"key":"e_1_2_1_21_1","doi-asserted-by":"crossref","unstructured":"Semih G\u00fcnel Helge Rhodin and Pascal Fua. 2018. What Face and Body Shapes Can Tell Us About Height. arxiv:cs.CV\/1805.10355  Semih G\u00fcnel Helge Rhodin and Pascal Fua. 2018. What Face and Body Shapes Can Tell Us About Height. arxiv:cs.CV\/1805.10355","DOI":"10.1109\/ICCVW.2019.00226"},{"key":"e_1_2_1_22_1","doi-asserted-by":"crossref","unstructured":"Ikhsanul Habibie Weipeng Xu Dushyant Mehta Gerard Pons-Moll and Christian Theobalt. 2019. In the wild human pose estimation using explicit 2D features and intermediate 3D representations. arXiv preprint arXiv:1904.03289 (2019).  Ikhsanul Habibie Weipeng Xu Dushyant Mehta Gerard Pons-Moll and Christian Theobalt. 2019. In the wild human pose estimation using explicit 2D features and intermediate 3D representations. arXiv preprint arXiv:1904.03289 (2019).","DOI":"10.1109\/CVPR.2019.01116"},{"key":"e_1_2_1_23_1","volume-title":"Mask R-CNN. In Proceedings of the IEEE International Conference on Computer Vision. 2961--2969","author":"He Kaiming","year":"2017","unstructured":"Kaiming He , Georgia Gkioxari , Piotr Doll\u00e1r , and Ross Girshick . 2017 . Mask R-CNN. In Proceedings of the IEEE International Conference on Computer Vision. 2961--2969 . Kaiming He, Georgia Gkioxari, Piotr Doll\u00e1r, and Ross Girshick. 2017. Mask R-CNN. In Proceedings of the IEEE International Conference on Computer Vision. 2961--2969."},{"key":"e_1_2_1_24_1","volume-title":"Little","author":"Imtiaz Hossain Mir Rayat","year":"2018","unstructured":"Mir Rayat Imtiaz Hossain and James J . Little . 2018 . Exploiting temporal information for 3d human pose estimation. In Proceedings of the European Conference on Computer Vision. Springer , 69--86. Mir Rayat Imtiaz Hossain and James J. Little. 2018. Exploiting temporal information for 3d human pose estimation. In Proceedings of the European Conference on Computer Vision. Springer, 69--86."},{"key":"e_1_2_1_25_1","volume-title":"Proceedings of the 2017 International Conference on 3D Vision (3DV). IEEE, 421--430","author":"Huang Yinghao","unstructured":"Yinghao Huang , Federica Bogo , Christoph Lassner , Angjoo Kanazawa , Peter V. Gehler , Javier Romero , Ijaz Akhter , and Michael J. Black . 2017. Towards accurate marker-less human shape and pose estimation over time . In Proceedings of the 2017 International Conference on 3D Vision (3DV). IEEE, 421--430 . Yinghao Huang, Federica Bogo, Christoph Lassner, Angjoo Kanazawa, Peter V. Gehler, Javier Romero, Ijaz Akhter, and Michael J. Black. 2017. Towards accurate marker-less human shape and pose estimation over time. In Proceedings of the 2017 International Conference on 3D Vision (3DV). IEEE, 421--430."},{"key":"e_1_2_1_26_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2013.248"},{"key":"e_1_2_1_27_1","doi-asserted-by":"publisher","DOI":"10.5244\/C.24.12"},{"key":"e_1_2_1_28_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00744"},{"key":"e_1_2_1_29_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.00576"},{"key":"e_1_2_1_30_1","doi-asserted-by":"publisher","DOI":"10.1007\/s11263-018-1066-6"},{"key":"e_1_2_1_31_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2019.00234"},{"key":"e_1_2_1_32_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.00463"},{"key":"e_1_2_1_33_1","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR\u201917)","author":"Lassner Christoph","year":"2017","unstructured":"Christoph Lassner , Javier Romero , Martin Kiefel , Federica Bogo , Michael J. Black , and Peter V. Gehler . 2017. Unite the people: Closing the loop between 3D and 2D human representations . In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR\u201917) . IEEE Computer Society, Washington, DC, 4704--4713. DOI:https:\/\/doi.org\/10.1109\/CVPR. 2017 .500 10.1109\/CVPR.2017.500 Christoph Lassner, Javier Romero, Martin Kiefel, Federica Bogo, Michael J. Black, and Peter V. Gehler. 2017. Unite the people: Closing the loop between 3D and 2D human representations. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR\u201917). IEEE Computer Society, Washington, DC, 4704--4713. DOI:https:\/\/doi.org\/10.1109\/CVPR.2017.500"},{"key":"e_1_2_1_34_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-01234-2_8"},{"key":"e_1_2_1_35_1","doi-asserted-by":"publisher","DOI":"10.1145\/3272127.3275071"},{"key":"e_1_2_1_36_1","volume-title":"Generating multiple hypotheses for 3D human pose estimation with mixture density network. arXiv preprint arXiv:1904.05547","author":"Li Chen","year":"2019","unstructured":"Chen Li and Gim Hee Lee . 2019. Generating multiple hypotheses for 3D human pose estimation with mixture density network. arXiv preprint arXiv:1904.05547 ( 2019 ). Chen Li and Gim Hee Lee. 2019. Generating multiple hypotheses for 3D human pose estimation with mixture density network. arXiv preprint arXiv:1904.05547 (2019)."},{"key":"e_1_2_1_37_1","volume-title":"Compositional human pose regression. Comput. Vision Image Understanding 176-177","author":"Liang Shuang","year":"2018","unstructured":"Shuang Liang , Xiao Sun , and Yichen Wei . 2018. Compositional human pose regression. Comput. Vision Image Understanding 176-177 ( 2018 ), 1--8. Shuang Liang, Xiao Sun, and Yichen Wei. 2018. Compositional human pose regression. Comput. Vision Image Understanding 176-177 (2018), 1--8."},{"key":"e_1_2_1_38_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.588"},{"key":"e_1_2_1_39_1","doi-asserted-by":"crossref","first-page":"11","DOI":"10.1109\/TPAMI.2013.47","article-title":"Markerless motion capture of multiple characters using multiview image segmentation","volume":"35","author":"Liu Yebin","year":"2013","unstructured":"Yebin Liu , Juergen Gall , Carsten Stoll , Qionghai Dai , Hans-Peter Seidel , and Christian Theobalt . 2013 . Markerless motion capture of multiple characters using multiview image segmentation . IEEE Trans. Pattern Anal. Mach. Intell. 35 , 11 (Nov. 2013), 2720--2735. DOI:https:\/\/doi.org\/10.1109\/TPAMI.2013.47 10.1109\/TPAMI.2013.47 Yebin Liu, Juergen Gall, Carsten Stoll, Qionghai Dai, Hans-Peter Seidel, and Christian Theobalt. 2013. Markerless motion capture of multiple characters using multiview image segmentation. IEEE Trans. Pattern Anal. Mach. Intell. 35, 11 (Nov. 2013), 2720--2735. DOI:https:\/\/doi.org\/10.1109\/TPAMI.2013.47","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"e_1_2_1_40_1","doi-asserted-by":"publisher","DOI":"10.1145\/2816795.2818013"},{"key":"e_1_2_1_41_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00539"},{"key":"e_1_2_1_42_1","volume-title":"Proceedings of the IEEE International Conference on Computer Vision (ICCV\u201917)","author":"Martinez Julieta","year":"2017","unstructured":"Julieta Martinez , Rayat Hossain , Javier Romero , and James J. Little . 2017. A simple yet effective baseline for 3D human pose estimation . In Proceedings of the IEEE International Conference on Computer Vision (ICCV\u201917) . 2659--2668. DOI:https:\/\/doi.org\/10.1109\/ICCV. 2017 .288 10.1109\/ICCV.2017.288 Julieta Martinez, Rayat Hossain, Javier Romero, and James J. Little. 2017. A simple yet effective baseline for 3D human pose estimation. In Proceedings of the IEEE International Conference on Computer Vision (ICCV\u201917). 2659--2668. DOI:https:\/\/doi.org\/10.1109\/ICCV.2017.288"},{"key":"e_1_2_1_43_1","doi-asserted-by":"publisher","DOI":"10.1109\/3DV.2017.00064"},{"key":"e_1_2_1_44_1","unstructured":"Dushyant Mehta Oleksandr Sotnychenko Franziska Mueller Weipeng Xu Mohamed Elgharib Pascal Fua Hans-Peter Seidel Helge Rhodin Gerard Pons-Moll and Christian Theobalt. 2019. XNect: Real-time Multi-person 3D Human Pose Estimation with a Single RGB Camera. arxiv:cs.CV\/1907.00837  Dushyant Mehta Oleksandr Sotnychenko Franziska Mueller Weipeng Xu Mohamed Elgharib Pascal Fua Hans-Peter Seidel Helge Rhodin Gerard Pons-Moll and Christian Theobalt. 2019. XNect: Real-time Multi-person 3D Human Pose Estimation with a Single RGB Camera. arxiv:cs.CV\/1907.00837"},{"key":"e_1_2_1_45_1","doi-asserted-by":"publisher","DOI":"10.1145\/3072959.3073596"},{"key":"e_1_2_1_46_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.170"},{"key":"e_1_2_1_47_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-46484-8_29"},{"key":"e_1_2_1_48_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.395"},{"key":"e_1_2_1_49_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00763"},{"key":"e_1_2_1_50_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.139"},{"key":"e_1_2_1_51_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00055"},{"key":"e_1_2_1_52_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.00794"},{"key":"e_1_2_1_53_1","unstructured":"Dario Pavllo David Grangier and Michael Auli. 2018. QuaterNet: A Quaternion-based Recurrent Model for Human Motion. arxiv:cs.CV\/1805.06485  Dario Pavllo David Grangier and Michael Auli. 2018. QuaterNet: A Quaternion-based Recurrent Model for Human Motion. arxiv:cs.CV\/1805.06485"},{"key":"e_1_2_1_54_1","doi-asserted-by":"publisher","DOI":"10.1145\/3272127.3275014"},{"key":"e_1_2_1_55_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-642-33765-9_41"},{"key":"e_1_2_1_56_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-01249-6_46"},{"key":"e_1_2_1_57_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00880"},{"key":"e_1_2_1_58_1","volume-title":"Kakadiaris","author":"Sarafianos Nikolaos","year":"2016","unstructured":"Nikolaos Sarafianos , Bogdan Boteanu , Bogdan Ionescu , and Ioannis A . Kakadiaris . 2016 . 3D human pose estimation: A review of the literature and analysis of covariates. Comput. Vis. Image Underst. 152, C (Nov . 2016), 1--20. DOI:https:\/\/doi.org\/10.1016\/j.cviu.2016.09.002 10.1016\/j.cviu.2016.09.002 Nikolaos Sarafianos, Bogdan Boteanu, Bogdan Ionescu, and Ioannis A. Kakadiaris. 2016. 3D human pose estimation: A review of the literature and analysis of covariates. Comput. Vis. Image Underst. 152, C (Nov. 2016), 1--20. DOI:https:\/\/doi.org\/10.1016\/j.cviu.2016.09.002"},{"key":"e_1_2_1_59_1","volume-title":"Proceedings of the 2012 IEEE Conference on Computer Vision and Pattern Recognition (CVPR) (CVPR\u201912)","author":"Sharp Toby","year":"2012","unstructured":"Toby Sharp . 2012 . The vitruvian manifold: Inferring dense correspondences for one-shot human pose estimation . In Proceedings of the 2012 IEEE Conference on Computer Vision and Pattern Recognition (CVPR) (CVPR\u201912) . IEEE Computer Society, Washington, DC, 103--110. Toby Sharp. 2012. The vitruvian manifold: Inferring dense correspondences for one-shot human pose estimation. In Proceedings of the 2012 IEEE Conference on Computer Vision and Pattern Recognition (CVPR) (CVPR\u201912). IEEE Computer Society, Washington, DC, 103--110."},{"key":"e_1_2_1_60_1","doi-asserted-by":"publisher","DOI":"10.1145\/2398356.2398381"},{"key":"e_1_2_1_61_1","doi-asserted-by":"publisher","DOI":"10.1007\/s11263-009-0293-2"},{"key":"e_1_2_1_62_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2017.425"},{"key":"e_1_2_1_63_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.113"},{"key":"e_1_2_1_64_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.603"},{"key":"e_1_2_1_65_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2014.214"},{"key":"e_1_2_1_66_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00901"},{"key":"e_1_2_1_67_1","doi-asserted-by":"publisher","DOI":"10.1145\/1360612.1360696"},{"key":"e_1_2_1_68_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2019.2892452"},{"key":"e_1_2_1_69_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.511"},{"key":"e_1_2_1_70_1","doi-asserted-by":"publisher","DOI":"10.1145\/2366145.2366207"},{"key":"e_1_2_1_71_1","doi-asserted-by":"publisher","DOI":"10.1145\/3181973"},{"key":"e_1_2_1_72_1","unstructured":"Yuanlu Xu Song-Chun Zhu and Tony Tung. 2019. DenseRaC: Joint 3D Pose and Shape Estimation by Dense Render-and-Compare. arxiv:cs.CV\/1910.00116  Yuanlu Xu Song-Chun Zhu and Tony Tung. 2019. DenseRaC: Joint 3D Pose and Shape Estimation by Dense Render-and-Compare. arxiv:cs.CV\/1910.00116"},{"key":"e_1_2_1_73_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00551"},{"key":"e_1_2_1_74_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2014.301"},{"key":"e_1_2_1_75_1","doi-asserted-by":"publisher","DOI":"10.1002\/cav.1887"},{"key":"e_1_2_1_76_1","doi-asserted-by":"crossref","unstructured":"Yusuke Yoshiyasu Ryusuke Sagawa Ko Ayusawa and Akihiko Murai. 2018. Skeleton Transformer Networks: 3D Human Pose and Skinned Mesh from Single RGB Image. arxiv:cs.CV\/1812.11328  Yusuke Yoshiyasu Ryusuke Sagawa Ko Ayusawa and Akihiko Murai. 2018. Skeleton Transformer Networks: 3D Human Pose and Skinned Mesh from Single RGB Image. arxiv:cs.CV\/1812.11328","DOI":"10.1007\/978-3-030-20870-7_30"},{"key":"e_1_2_1_77_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2019.00721"},{"key":"e_1_2_1_78_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2019.00783"},{"key":"e_1_2_1_79_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2017.51"},{"key":"e_1_2_1_80_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-49409-8_17"},{"key":"e_1_2_1_81_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2018.2816031"},{"key":"e_1_2_1_82_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.00589"}],"container-title":["ACM Transactions on Graphics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3407659","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3407659","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T22:41:23Z","timestamp":1750200083000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3407659"}},"subtitle":["3D Human Motion Reconstruction from Monocular Video with Skeleton Consistency"],"short-title":[],"issued":{"date-parts":[[2020,9,4]]},"references-count":82,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2021,2,28]]}},"alternative-id":["10.1145\/3407659"],"URL":"https:\/\/doi.org\/10.1145\/3407659","relation":{},"ISSN":["0730-0301","1557-7368"],"issn-type":[{"value":"0730-0301","type":"print"},{"value":"1557-7368","type":"electronic"}],"subject":[],"published":{"date-parts":[[2020,9,4]]},"assertion":[{"value":"2019-10-01","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2020-06-01","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2020-09-04","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}