{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,27]],"date-time":"2026-07-27T18:48:34Z","timestamp":1785178114015,"version":"3.55.0"},"reference-count":114,"publisher":"Association for Computing Machinery (ACM)","issue":"4","license":[{"start":{"date-parts":[[2017,7,20]],"date-time":"2017-07-20T00:00:00Z","timestamp":1500508800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"name":"Marie Curie Individual Fellowship","award":["707326"],"award-info":[{"award-number":["707326"]}]},{"DOI":"10.13039\/501100000781","name":"European Research Council","doi-asserted-by":"publisher","award":["335545"],"award-info":[{"award-number":["335545"]}],"id":[{"id":"10.13039\/501100000781","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Graph."],"published-print":{"date-parts":[[2017,8,31]]},"abstract":"<jats:p>We present the first real-time method to capture the full global 3D skeletal pose of a human in a stable, temporally consistent manner using a single RGB camera. Our method combines a new convolutional neural network (CNN) based pose regressor with kinematic skeleton fitting. Our novel fully-convolutional pose formulation regresses 2D and 3D joint positions jointly in real time and does not require tightly cropped input frames. A real-time kinematic skeleton fitting method uses the CNN output to yield temporally stable 3D global pose reconstructions on the basis of a coherent kinematic skeleton. This makes our approach the first monocular RGB method usable in real-time applications such as 3D character control---thus far, the only monocular methods for such applications employed specialized RGB-D cameras. Our method's accuracy is quantitatively on par with the best offline 3D monocular RGB pose estimation methods. Our results are qualitatively comparable to, and sometimes better than, results from monocular RGB-D approaches, such as the Kinect. However, we show that our approach is more broadly applicable than RGB-D solutions, i.e., it works for outdoor scenes, community videos, and low quality commodity RGB cameras.<\/jats:p>","DOI":"10.1145\/3072959.3073596","type":"journal-article","created":{"date-parts":[[2017,7,21]],"date-time":"2017-07-21T12:24:07Z","timestamp":1500639847000},"page":"1-14","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":896,"title":["VNect"],"prefix":"10.1145","volume":"36","author":[{"given":"Dushyant","family":"Mehta","sequence":"first","affiliation":[{"name":"Max Planck Institute for Informatics and Saarland University"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Srinath","family":"Sridhar","sequence":"additional","affiliation":[{"name":"Max Planck Institute for Informatics"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Oleksandr","family":"Sotnychenko","sequence":"additional","affiliation":[{"name":"Max Planck Institute for Informatics"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Helge","family":"Rhodin","sequence":"additional","affiliation":[{"name":"Max Planck Institute for Informatics"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Mohammad","family":"Shafiei","sequence":"additional","affiliation":[{"name":"Max Planck Institute for Informatics and Saarland University"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Hans-Peter","family":"Seidel","sequence":"additional","affiliation":[{"name":"Max Planck Institute for Informatics"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Weipeng","family":"Xu","sequence":"additional","affiliation":[{"name":"Max Planck Institute for Informatics"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Dan","family":"Casas","sequence":"additional","affiliation":[{"name":"Universidad Rey Juan Carlos"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Christian","family":"Theobalt","sequence":"additional","affiliation":[{"name":"Max Planck Institute for Informatics"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2017,7,20]]},"reference":[{"key":"e_1_2_2_1_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2006.21"},{"key":"e_1_2_2_2_1","unstructured":"Sameer Agarwal Keir Mierle and Others. 2017. Ceres Solver. http:\/\/ceres-solver.org. (2017).  Sameer Agarwal Keir Mierle and Others. 2017. Ceres Solver. http:\/\/ceres-solver.org. (2017)."},{"key":"e_1_2_2_3_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2015.7298751"},{"key":"e_1_2_2_4_1","doi-asserted-by":"crossref","unstructured":"Sikandar Amin Mykhaylo Andriluka Marcus Rohrbach and Bernt Schiele. 2013. Multiview Pictorial Structures for 3D Human Pose Estimation. In BMVC.  Sikandar Amin Mykhaylo Andriluka Marcus Rohrbach and Bernt Schiele. 2013. Multiview Pictorial Structures for 3D Human Pose Estimation. In BMVC.","DOI":"10.5244\/C.27.45"},{"key":"e_1_2_2_5_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2014.471"},{"key":"e_1_2_2_6_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2009.5206754"},{"key":"e_1_2_2_7_1","doi-asserted-by":"publisher","DOI":"10.5244\/C.29.32"},{"key":"e_1_2_2_8_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2011.6126356"},{"key":"e_1_2_2_9_1","doi-asserted-by":"publisher","DOI":"10.1109\/VSPETS.2005.1570935"},{"key":"e_1_2_2_10_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2007.383340"},{"key":"e_1_2_2_11_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2014.216"},{"key":"e_1_2_2_12_1","volume-title":"Recurrent Human Pose Estimation. arXiv preprint arXiv:1605.02914","author":"Belagiannis Vasileios","year":"2016","unstructured":"Vasileios Belagiannis and Andrew Zisserman . 2016. Recurrent Human Pose Estimation. arXiv preprint arXiv:1605.02914 ( 2016 ). Vasileios Belagiannis and Andrew Zisserman. 2016. Recurrent Human Pose Estimation. arXiv preprint arXiv:1605.02914 (2016)."},{"key":"e_1_2_2_13_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2007.383129"},{"key":"e_1_2_2_14_1","doi-asserted-by":"publisher","DOI":"10.1007\/s11263-008-0204-y"},{"key":"e_1_2_2_15_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-46454-1_34"},{"key":"e_1_2_2_16_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2009.5459303"},{"key":"e_1_2_2_17_1","doi-asserted-by":"publisher","DOI":"10.1109\/3DV.2016.84"},{"key":"e_1_2_2_18_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.1998.698581"},{"key":"e_1_2_2_19_1","volume-title":"Realtime Multi-Person 2D Pose Estimation using Part Affinity Fields. arXiv preprint arXiv:1611.08050","author":"Cao Zhe","year":"2016","unstructured":"Zhe Cao , Tomas Simon , Shih-En Wei , and Yaser Sheikh . 2016. Realtime Multi-Person 2D Pose Estimation using Part Affinity Fields. arXiv preprint arXiv:1611.08050 ( 2016 ). Zhe Cao, Tomas Simon, Shih-En Wei, and Yaser Sheikh. 2016. Realtime Multi-Person 2D Pose Estimation using Part Affinity Fields. arXiv preprint arXiv:1611.08050 (2016)."},{"key":"e_1_2_2_20_1","doi-asserted-by":"publisher","DOI":"10.1145\/2207676.2208639"},{"key":"e_1_2_2_21_1","doi-asserted-by":"publisher","DOI":"10.1145\/1073204.1073248"},{"key":"e_1_2_2_22_1","doi-asserted-by":"publisher","DOI":"10.1109\/3DV.2016.58"},{"key":"e_1_2_2_23_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2008.4587752"},{"key":"e_1_2_2_24_1","doi-asserted-by":"publisher","DOI":"10.1023\/B:VISI.0000043757.18370.9c"},{"key":"e_1_2_2_25_1","doi-asserted-by":"publisher","DOI":"10.1145\/2897824.2925969"},{"key":"e_1_2_2_26_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2004.1315230"},{"key":"e_1_2_2_27_1","volume-title":"MARCOnI - ConvNet-based MARker-less Motion Capture in Outdoor and Indoor Scenes","author":"Elhayek Ahmed","year":"2016","unstructured":"Ahmed Elhayek , Edilson de Aguiar , Arjun Jain , Jonathan Tompson , Leonid Pishchulin , Mykhaylo Andriluka , Christoph Bregler , Bernt Schiele , and Christian Theobalt . 2016. MARCOnI - ConvNet-based MARker-less Motion Capture in Outdoor and Indoor Scenes . IEEE Transactions on Pattern Analysis and Machine Intelligence (PAMI) ( 2016 ). Ahmed Elhayek, Edilson de Aguiar, Arjun Jain, Jonathan Tompson, Leonid Pishchulin, Mykhaylo Andriluka, Christoph Bregler, Bernt Schiele, and Christian Theobalt. 2016. MARCOnI - ConvNet-based MARker-less Motion Capture in Outdoor and Indoor Scenes. IEEE Transactions on Pattern Analysis and Machine Intelligence (PAMI) (2016)."},{"key":"e_1_2_2_28_1","volume-title":"Object detection with discriminatively trained part-based models","author":"Felzenszwalb Pedro F","unstructured":"Pedro F Felzenszwalb , Ross B Girshick , David McAllester , and Deva Ramanan . 2010. Object detection with discriminatively trained part-based models . In IEEE transactions on pattern analysis and machine intelligence. IEEE , 1627--1645. Pedro F Felzenszwalb, Ross B Girshick, David McAllester, and Deva Ramanan. 2010. Object detection with discriminatively trained part-based models. In IEEE transactions on pattern analysis and machine intelligence. IEEE, 1627--1645."},{"key":"e_1_2_2_29_1","doi-asserted-by":"publisher","DOI":"10.1023\/B:VISI.0000042934.15159.49"},{"key":"e_1_2_2_30_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2009.5206495"},{"key":"e_1_2_2_31_1","doi-asserted-by":"publisher","DOI":"10.1007\/s11263-008-0173-1"},{"key":"e_1_2_2_32_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-642-33783-3_53"},{"key":"e_1_2_2_33_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2013.168"},{"key":"e_1_2_2_34_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2011.6126270"},{"key":"e_1_2_2_35_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2011.50"},{"key":"e_1_2_2_36_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.90"},{"key":"e_1_2_2_37_1","first-page":"820","article-title":"Bayesian Reconstruction of 3D Human Motion from Single-Camera Video","volume":"99","author":"Howe Nicholas R","year":"1999","unstructured":"Nicholas R Howe , Michael E Leventon , and William T Freeman . 1999 . Bayesian Reconstruction of 3D Human Motion from Single-Camera Video .. In NIPS , Vol. 99. 820 -- 826 . Nicholas R Howe, Michael E Leventon, and William T Freeman. 1999. Bayesian Reconstruction of 3D Human Motion from Single-Camera Video.. In NIPS, Vol. 99. 820--6.","journal-title":"NIPS"},{"key":"e_1_2_2_38_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.604"},{"key":"e_1_2_2_39_1","volume-title":"VolumeDeform: Real-time Volumetric Non-rigid Reconstruction. (October","author":"Innmann Matthias","year":"2016","unstructured":"Matthias Innmann , Michael Zollh\u00f6fer , Matthias Nie\u00dfner , Christian Theobalt , and Marc Stamminger . 2016. VolumeDeform: Real-time Volumetric Non-rigid Reconstruction. (October 2016 ), 17. Matthias Innmann, Michael Zollh\u00f6fer, Matthias Nie\u00dfner, Christian Theobalt, and Marc Stamminger. 2016. VolumeDeform: Real-time Volumetric Non-rigid Reconstruction. (October 2016), 17."},{"key":"e_1_2_2_40_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-46466-4_3"},{"key":"e_1_2_2_41_1","volume-title":"Proceedings of The 32nd International Conference on Machine Learning. 448--456","author":"Ioffe Sergey","year":"2015","unstructured":"Sergey Ioffe and Christian Szegedy . 2015 . Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift . In Proceedings of The 32nd International Conference on Machine Learning. 448--456 . Sergey Ioffe and Christian Szegedy. 2015. Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift. In Proceedings of The 32nd International Conference on Machine Learning. 448--456."},{"key":"e_1_2_2_42_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2014.215"},{"key":"e_1_2_2_43_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2013.248"},{"key":"e_1_2_2_44_1","doi-asserted-by":"publisher","DOI":"10.1145\/1866158.1866174"},{"key":"e_1_2_2_45_1","doi-asserted-by":"publisher","DOI":"10.1145\/2647868.2654889"},{"key":"e_1_2_2_46_1","doi-asserted-by":"publisher","DOI":"10.5244\/C.24.12"},{"key":"e_1_2_2_47_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2011.5995318"},{"key":"e_1_2_2_48_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2013.169"},{"key":"e_1_2_2_49_1","volume-title":"Asian Conference on Computer Vision (ACCV). 332--347","author":"Li Sijin","year":"2014","unstructured":"Sijin Li and Antoni B Chan . 2014 . 3d human pose estimation from monocular images with deep convolutional neural network . In Asian Conference on Computer Vision (ACCV). 332--347 . Sijin Li and Antoni B Chan. 2014. 3d human pose estimation from monocular images with deep convolutional neural network. In Asian Conference on Computer Vision (ACCV). 332--347."},{"key":"e_1_2_2_50_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2015.326"},{"key":"e_1_2_2_51_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2015.326"},{"key":"e_1_2_2_52_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-46475-6_16"},{"key":"e_1_2_2_53_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-46448-0_2"},{"key":"e_1_2_2_54_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-10584-0_11"},{"key":"e_1_2_2_55_1","doi-asserted-by":"publisher","DOI":"10.1007\/s00371-013-0894-1"},{"key":"e_1_2_2_56_1","volume-title":"Monocular 3D Human Pose Estimation In The Wild Using Improved CNN Supervision. arXiv preprint arXiv:1611.09813v2","author":"Mehta Dushyant","year":"2016","unstructured":"Dushyant Mehta , Helge Rhodin , Dan Casas , Oleksandr Sotnychenko , Weipeng Xu , and Christian Theobalt . 2016. Monocular 3D Human Pose Estimation In The Wild Using Improved CNN Supervision. arXiv preprint arXiv:1611.09813v2 ( 2016 ). Dushyant Mehta, Helge Rhodin, Dan Casas, Oleksandr Sotnychenko, Weipeng Xu, and Christian Theobalt. 2016. Monocular 3D Human Pose Estimation In The Wild Using Improved CNN Supervision. arXiv preprint arXiv:1611.09813v2 (2016)."},{"key":"e_1_2_2_57_1","unstructured":"Alberto Menache. 2000. Understanding motion capture for computer animation and video games. Morgan kaufmann.  Alberto Menache. 2000. Understanding motion capture for computer animation and video games. Morgan kaufmann."},{"key":"e_1_2_2_58_1","unstructured":"Microsoft Corporation. 2010. Kinect for Xbox 360. http:\/\/www.xbox.com\/en-US\/xbox-360\/accessories\/kinect. (2010).  Microsoft Corporation. 2010. Kinect for Xbox 360. http:\/\/www.xbox.com\/en-US\/xbox-360\/accessories\/kinect. (2010)."},{"key":"e_1_2_2_59_1","unstructured":"Microsoft Corporation. 2013. Kinect for Xbox One. http:\/\/www.xbox.com\/en-US\/xbox-one\/accessories\/kinect. (2013).  Microsoft Corporation. 2013. Kinect for Xbox One. http:\/\/www.xbox.com\/en-US\/xbox-one\/accessories\/kinect. (2013)."},{"key":"e_1_2_2_60_1","unstructured":"Microsoft Corporation. 2015. Kinect SDK. https:\/\/developer.microsoft.com\/en-us\/windows\/kinect. (2015).  Microsoft Corporation. 2015. Kinect SDK. https:\/\/developer.microsoft.com\/en-us\/windows\/kinect. (2015)."},{"key":"e_1_2_2_61_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.cviu.2006.08.002"},{"key":"e_1_2_2_62_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2006.149"},{"key":"e_1_2_2_63_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2015.7298631"},{"key":"e_1_2_2_64_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-46484-8_29"},{"key":"e_1_2_2_65_1","first-page":"3","article-title":"Efficient model-based 3D tracking of hand articulations using Kinect","volume":"1","author":"Oikonomidis Iason","year":"2011","unstructured":"Iason Oikonomidis , Nikolaos Kyriazis , and Antonis A Argyros . 2011 . Efficient model-based 3D tracking of hand articulations using Kinect .. In BmVC , Vol. 1. 3 . Iason Oikonomidis, Nikolaos Kyriazis, and Antonis A Argyros. 2011. Efficient model-based 3D tracking of hand articulations using Kinect.. In BmVC, Vol. 1. 3.","journal-title":"BmVC"},{"key":"e_1_2_2_66_1","doi-asserted-by":"publisher","DOI":"10.1145\/2984511.2984517"},{"key":"e_1_2_2_67_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2011.6126243"},{"key":"e_1_2_2_68_1","volume-title":"Coarse-to-Fine Volumetric Prediction for Single-Image 3D Human Pose. arXiv preprint arXiv:1611.07828","author":"Pavlakos Georgios","year":"2016","unstructured":"Georgios Pavlakos , Xiaowei Zhou , Konstantinos G Derpanis , and Kostas Daniilidis . 2016. Coarse-to-Fine Volumetric Prediction for Single-Image 3D Human Pose. arXiv preprint arXiv:1611.07828 ( 2016 ). Georgios Pavlakos, Xiaowei Zhou, Konstantinos G Derpanis, and Kostas Daniilidis. 2016. Coarse-to-Fine Volumetric Prediction for Single-Image 3D Human Pose. arXiv preprint arXiv:1611.07828 (2016)."},{"key":"e_1_2_2_69_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2013.433"},{"key":"e_1_2_2_70_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.533"},{"key":"e_1_2_2_71_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2014.300"},{"key":"e_1_2_2_72_1","unstructured":"Real Madrid C.F. 2016. Cristiano Ronaldo and Coentrao continue their recovery. https:\/\/www.youtube.com\/watch?v=xqiPuX_buOo. (2016).  Real Madrid C.F. 2016. Cristiano Ronaldo and Coentrao continue their recovery. https:\/\/www.youtube.com\/watch?v=xqiPuX_buOo. (2016)."},{"key":"e_1_2_2_73_1","volume-title":"You only look once: Unified, real-time object detection. arXiv preprint arXiv:1506.02640","author":"Redmon Joseph","year":"2015","unstructured":"Joseph Redmon , Santosh Divvala , Ross Girshick , and Ali Farhadi . 2015. You only look once: Unified, real-time object detection. arXiv preprint arXiv:1506.02640 ( 2015 ). Joseph Redmon, Santosh Divvala, Ross Girshick, and Ali Farhadi. 2015. You only look once: Unified, real-time object detection. arXiv preprint arXiv:1506.02640 (2015)."},{"key":"e_1_2_2_74_1","unstructured":"Shaoqing Ren Kaiming He Ross Girshick and Jian Sun. 2015. Faster R-CNN: Towards real-time object detection with region proposal networks. In Advances in neural information processing systems. 91--99.  Shaoqing Ren Kaiming He Ross Girshick and Jian Sun. 2015. Faster R-CNN: Towards real-time object detection with region proposal networks. In Advances in neural information processing systems. 91--99."},{"key":"e_1_2_2_75_1","volume-title":"EgoCap: Egocentric Marker-less Motion Capture with Two Fisheye Cameras. ACM Trans. Graph. (Proc. SIGGRAPH Asia)","author":"Rhodin Helge","year":"2016","unstructured":"Helge Rhodin , Christian Richardt , Dan Casas , Eldar Insafutdinov , Mohammad Shafiei , Hans-Peter Seidel , Bernt Schiele , and Christian Theobalt . 2016a. EgoCap: Egocentric Marker-less Motion Capture with Two Fisheye Cameras. ACM Trans. Graph. (Proc. SIGGRAPH Asia) ( 2016 ). Helge Rhodin, Christian Richardt, Dan Casas, Eldar Insafutdinov, Mohammad Shafiei, Hans-Peter Seidel, Bernt Schiele, and Christian Theobalt. 2016a. EgoCap: Egocentric Marker-less Motion Capture with Two Fisheye Cameras. ACM Trans. Graph. (Proc. SIGGRAPH Asia) (2016)."},{"key":"e_1_2_2_76_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-46454-1_31"},{"key":"e_1_2_2_77_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2015.94"},{"key":"e_1_2_2_78_1","volume-title":"MoCap-guided Data Augmentation for 3D Pose Estimation in the Wild. arXiv preprint arXiv:1607.02046","author":"Rogez Gr\u00e9gory","year":"2016","unstructured":"Gr\u00e9gory Rogez and Cordelia Schmid . 2016. MoCap-guided Data Augmentation for 3D Pose Estimation in the Wild. arXiv preprint arXiv:1607.02046 ( 2016 ). Gr\u00e9gory Rogez and Cordelia Schmid. 2016. MoCap-guided Data Augmentation for 3D Pose Estimation in the Wild. arXiv preprint arXiv:1607.02046 (2016)."},{"key":"e_1_2_2_79_1","doi-asserted-by":"publisher","DOI":"10.1145\/2634212"},{"key":"e_1_2_2_80_1","doi-asserted-by":"publisher","DOI":"10.1109\/HUMO.2000.897366"},{"key":"e_1_2_2_81_1","doi-asserted-by":"publisher","DOI":"10.1007\/s11263-006-5165-4"},{"key":"e_1_2_2_82_1","unstructured":"RUSFENCING-TV. 2017. The Most Beautiful Strike \/ Saber Woman (Translated from Russian). https:\/\/www.youtube.com\/watch?v=0gOcMsWUkCU. (2017).  RUSFENCING-TV. 2017. The Most Beautiful Strike \/ Saber Woman (Translated from Russian). https:\/\/www.youtube.com\/watch?v=0gOcMsWUkCU. (2017)."},{"key":"e_1_2_2_83_1","doi-asserted-by":"publisher","DOI":"10.1145\/2398356.2398381"},{"key":"e_1_2_2_84_1","doi-asserted-by":"publisher","DOI":"10.1007\/3-540-45053-X_45"},{"key":"e_1_2_2_85_1","doi-asserted-by":"publisher","DOI":"10.1007\/s11263-011-0493-4"},{"key":"e_1_2_2_86_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2006.169"},{"key":"e_1_2_2_87_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2007.1111"},{"key":"e_1_2_2_88_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2001.990509"},{"key":"e_1_2_2_89_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2003.1238446"},{"key":"e_1_2_2_90_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2011.6126338"},{"key":"e_1_2_2_91_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2006.83"},{"key":"e_1_2_2_92_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2000.855885"},{"key":"e_1_2_2_93_1","doi-asserted-by":"publisher","DOI":"10.5244\/C.30.130"},{"key":"e_1_2_2_94_1","volume-title":"Fusing 2D Uncertainty and 3D Cues for Monocular Body Pose Estimation. arXiv preprint arXiv:1611.05708","author":"Tekin Bugra","year":"2016","unstructured":"Bugra Tekin , Pablo M\u00e1rquez-Neila , Mathieu Salzmann , and Pascal Fua . 2016b. Fusing 2D Uncertainty and 3D Cues for Monocular Body Pose Estimation. arXiv preprint arXiv:1611.05708 ( 2016 ). Bugra Tekin, Pablo M\u00e1rquez-Neila, Mathieu Salzmann, and Pascal Fua. 2016b. Fusing 2D Uncertainty and 3D Cues for Monocular Body Pose Estimation. arXiv preprint arXiv:1611.05708 (2016)."},{"key":"e_1_2_2_95_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.113"},{"key":"e_1_2_2_96_1","unstructured":"Jonathan J Tompson Arjun Jain Yann LeCun and Christoph Bregler. 2014. Joint training of a convolutional network and a graphical model for human pose estimation. In Advances in Neural Information Processing Systems (NIPS). 1799--1807.  Jonathan J Tompson Arjun Jain Yann LeCun and Christoph Bregler. 2014. Joint training of a convolutional network and a graphical model for human pose estimation. In Advances in Neural Information Processing Systems (NIPS). 1799--1807."},{"key":"e_1_2_2_97_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2014.214"},{"key":"e_1_2_2_98_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.cviu.2006.08.006"},{"key":"e_1_2_2_99_1","doi-asserted-by":"publisher","DOI":"10.1145\/2185520.2185523"},{"key":"e_1_2_2_100_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2014.303"},{"key":"e_1_2_2_101_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.511"},{"key":"e_1_2_2_102_1","doi-asserted-by":"publisher","DOI":"10.1145\/1833349.1778779"},{"key":"e_1_2_2_103_1","doi-asserted-by":"publisher","DOI":"10.1145\/2366145.2366207"},{"key":"e_1_2_2_104_1","doi-asserted-by":"publisher","DOI":"10.1109\/34.598236"},{"key":"e_1_2_2_105_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.535"},{"key":"e_1_2_2_106_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2014.301"},{"key":"e_1_2_2_107_1","volume-title":"European Conference on Computer Vision (ECCV).","author":"Yu Yongkang","year":"2016","unstructured":"Yongkang Yu , Feilinand Yonghao , Zhen Yilin , and Weidong Mohan . 2016 . Marker-less 3D Human Motion Capture with Monocular Image Sequence and Height-Maps . In European Conference on Computer Vision (ECCV). Yongkang Yu, Feilinand Yonghao, Zhen Yilin, and Weidong Mohan. 2016. Marker-less 3D Human Motion Capture with Monocular Image Sequence and Height-Maps. In European Conference on Computer Vision (ECCV)."},{"key":"e_1_2_2_108_1","volume-title":"ADADELTA: an adaptive learning rate method. arXiv preprint arXiv:1212.5701","author":"Zeiler Matthew D","year":"2012","unstructured":"Matthew D Zeiler . 2012. ADADELTA: an adaptive learning rate method. arXiv preprint arXiv:1212.5701 ( 2012 ). Matthew D Zeiler. 2012. ADADELTA: an adaptive learning rate method. arXiv preprint arXiv:1212.5701 (2012)."},{"key":"e_1_2_2_109_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2015.7299074"},{"key":"e_1_2_2_110_1","doi-asserted-by":"crossref","unstructured":"Xingyi Zhou Xiao Sun Wei Zhang Shuang Liang and Yichen Wei. 2016. Deep Kinematic Pose Regression. ECCV Worktp on Geometry Meets Deep Learning.  Xingyi Zhou Xiao Sun Wei Zhang Shuang Liang and Yichen Wei. 2016. Deep Kinematic Pose Regression. ECCV Worktp on Geometry Meets Deep Learning.","DOI":"10.1007\/978-3-319-49409-8_17"},{"key":"e_1_2_2_111_1","volume-title":"Sparse Representation for 3D Shape Estimation: A Convex Relaxation Approach. arXiv preprint arXiv:1509.04309","author":"Zhou Xiaowei","year":"2015","unstructured":"Xiaowei Zhou , Menglong Zhu , Spyridon Leonardos , and Kostas Daniilidis . 2015a. Sparse Representation for 3D Shape Estimation: A Convex Relaxation Approach. arXiv preprint arXiv:1509.04309 ( 2015 ). Xiaowei Zhou, Menglong Zhu, Spyridon Leonardos, and Kostas Daniilidis. 2015a. Sparse Representation for 3D Shape Estimation: A Convex Relaxation Approach. arXiv preprint arXiv:1509.04309 (2015)."},{"key":"e_1_2_2_112_1","volume-title":"IEEE Conference on Computer Vision and Pattern Recognition (CVPR).","author":"Zhou Xiaowei","year":"2015","unstructured":"Xiaowei Zhou , Menglong Zhu , Spyridon Leonardos , Kosta Derpanis , and Kostas Daniilidis . 2015 b. Sparseness Meets Deepness: 3D Human Pose Estimation from Monocular Video . In IEEE Conference on Computer Vision and Pattern Recognition (CVPR). Xiaowei Zhou, Menglong Zhu, Spyridon Leonardos, Kosta Derpanis, and Kostas Daniilidis. 2015b. Sparseness Meets Deepness: 3D Human Pose Estimation from Monocular Video. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR)."},{"key":"e_1_2_2_113_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2011.5995650"},{"key":"e_1_2_2_114_1","doi-asserted-by":"publisher","DOI":"10.1145\/2601097.2601165"}],"container-title":["ACM Transactions on Graphics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3072959.3073596","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3072959.3073596","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T03:37:23Z","timestamp":1750217843000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3072959.3073596"}},"subtitle":["real-time 3D human pose estimation with a single RGB camera"],"short-title":[],"issued":{"date-parts":[[2017,7,20]]},"references-count":114,"journal-issue":{"issue":"4","published-print":{"date-parts":[[2017,8,31]]}},"alternative-id":["10.1145\/3072959.3073596"],"URL":"https:\/\/doi.org\/10.1145\/3072959.3073596","relation":{},"ISSN":["0730-0301","1557-7368"],"issn-type":[{"value":"0730-0301","type":"print"},{"value":"1557-7368","type":"electronic"}],"subject":[],"published":{"date-parts":[[2017,7,20]]},"assertion":[{"value":"2017-07-20","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}