{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,18]],"date-time":"2026-07-18T05:20:47Z","timestamp":1784352047823,"version":"3.55.0"},"reference-count":115,"publisher":"Association for Computing Machinery (ACM)","issue":"4","license":[{"start":{"date-parts":[[2020,8,12]],"date-time":"2020-08-12T00:00:00Z","timestamp":1597190400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"DOI":"10.13039\/100011199","name":"European Research Council","doi-asserted-by":"publisher","award":["770784"],"award-info":[{"award-number":["770784"]}],"id":[{"id":"10.13039\/100011199","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Graph."],"published-print":{"date-parts":[[2020,8,31]]},"abstract":"<jats:p>\n            We present a real-time approach for multi-person 3D motion capture at over 30 fps using a single RGB camera. It operates successfully in generic scenes which may contain occlusions by objects and by other people. Our method operates in subsequent stages. The first stage is a convolutional neural network (CNN) that estimates 2D and 3D pose features along with identity assignments for all visible joints of all individuals. We contribute a new architecture for this CNN, called\n            <jats:italic toggle=\"yes\">SelecSLS Net<\/jats:italic>\n            , that uses novel selective long and short range skip connections to improve the information flow allowing for a drastically faster network without compromising accuracy. In the second stage, a fullyconnected neural network turns the possibly partial (on account of occlusion) 2D pose and 3D pose features for each subject into a complete 3D pose estimate per individual. The third stage applies space-time skeletal model fitting to the predicted 2D and 3D pose per subject to further reconcile the 2D and 3D pose, and enforce temporal coherence. Our method returns the full skeletal pose in joint angles for each subject. This is a further key distinction from previous work that do not produce joint angle results of a coherent skeleton in real time for multi-person scenes. The proposed system runs on consumer hardware at a previously unseen speed of more than 30 fps given 512x320 images as input while achieving state-of-the-art accuracy, which we will demonstrate on a range of challenging real-world scenes.\n          <\/jats:p>","DOI":"10.1145\/3386569.3392410","type":"journal-article","created":{"date-parts":[[2020,8,12]],"date-time":"2020-08-12T11:44:27Z","timestamp":1597232667000},"update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":264,"title":["XNect"],"prefix":"10.1145","volume":"39","author":[{"given":"Dushyant","family":"Mehta","sequence":"first","affiliation":[{"name":"University of British Columbia"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Oleksandr","family":"Sotnychenko","sequence":"additional","affiliation":[{"name":"University of British Columbia"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Franziska","family":"Mueller","sequence":"additional","affiliation":[{"name":"University of British Columbia"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Weipeng","family":"Xu","sequence":"additional","affiliation":[{"name":"University of British Columbia"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Mohamed","family":"Elgharib","sequence":"additional","affiliation":[{"name":"University of British Columbia"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Pascal","family":"Fua","sequence":"additional","affiliation":[{"name":"University of British Columbia"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Hans-Peter","family":"Seidel","sequence":"additional","affiliation":[{"name":"University of British Columbia"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Helge","family":"Rhodin","sequence":"additional","affiliation":[{"name":"University of British Columbia"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Gerard","family":"Pons-Moll","sequence":"additional","affiliation":[{"name":"University of British Columbia"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Christian","family":"Theobalt","sequence":"additional","affiliation":[{"name":"University of British Columbia"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2020,8,12]]},"reference":[{"key":"e_1_2_2_1_1","unstructured":"1MILLION TV. 2016. GDFR - Flo Rida \/ 1MILLION Dance TUTORIAL (1\/2). https:\/\/www.youtube.com\/watch?v=9HkVnFpmXAw."},{"key":"e_1_2_2_2_1","unstructured":"Adobe. 2020. Mixamo. https:\/\/www.mixamo.com\/."},{"key":"e_1_2_2_3_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.00127"},{"key":"e_1_2_2_4_1","doi-asserted-by":"crossref","unstructured":"Thiemo Alldieck Marcus Magnor Weipeng Xu Christian Theobalt and Gerard Pons-Moll. 2018a. Detailed Human Avatars from Monocular Video. In 3DV.","DOI":"10.1109\/3DV.2018.00022"},{"key":"e_1_2_2_5_1","doi-asserted-by":"crossref","unstructured":"Thiemo Alldieck Marcus Magnor Weipeng Xu Christian Theobalt and Gerard Pons-Moll. 2018b. Video Based Reconstruction of 3D People Models. In CVPR.","DOI":"10.1109\/CVPR.2018.00875"},{"key":"e_1_2_2_6_1","doi-asserted-by":"crossref","unstructured":"Mykhaylo Andriluka Leonid Pishchulin Peter Gehler and Bernt Schiele. 2014. 2D Human Pose Estimation: New Benchmark and State of the Art Analysis. In CVPR.","DOI":"10.1109\/CVPR.2014.471"},{"key":"e_1_2_2_7_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.00351"},{"key":"e_1_2_2_8_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2019.00552"},{"key":"e_1_2_2_9_1","volume-title":"Black","author":"Bogo Federica","year":"2016","unstructured":"Federica Bogo, Angjoo Kanazawa, Christoph Lassner, Peter Gehler, Javier Romero, and Michael J. Black. 2016. Keep it SMPL: Automatic Estimation of 3D Human Pose and Shape from a Single Image. In ECCV."},{"key":"e_1_2_2_10_1","unstructured":"Boxing School Alexei Frolov. 2018. Boxing: bouncing - slip - overhand punch. https:\/\/www.youtube.com\/watch?v=dbuz9Q05bsM."},{"key":"e_1_2_2_11_1","unstructured":"Brave Entertainment. 2017. (Rollin') Dance Practice Video. https:\/\/www.youtube.com\/watch?v=ZhuDSdmby8k."},{"key":"e_1_2_2_12_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2017.400"},{"key":"e_1_2_2_13_1","doi-asserted-by":"crossref","unstructured":"Zhe Cao Tomas Simon Shih-En Wei and Yaser Sheikh. 2017. Realtime Multi-Person 2D Pose Estimation using Part Affinity Fields. In CVPR.","DOI":"10.1109\/CVPR.2017.143"},{"key":"e_1_2_2_14_1","doi-asserted-by":"publisher","DOI":"10.1145\/2207676.2208639"},{"key":"e_1_2_2_15_1","doi-asserted-by":"publisher","DOI":"10.1145\/1073204.1073248"},{"key":"e_1_2_2_16_1","doi-asserted-by":"crossref","unstructured":"Wenzheng Chen Huan Wang Yangyan Li Hao Su Zhenhua Wang Changhe Tu Dani Lischinski Daniel Cohen-Or and Baoquan Chen. 2016. Synthesizing Training Images for Boosting Human 3D Pose Estimation. In 3DV.","DOI":"10.1109\/3DV.2016.58"},{"key":"e_1_2_2_17_1","doi-asserted-by":"publisher","DOI":"10.1109\/3DV.2019.00052"},{"key":"e_1_2_2_18_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-01240-3_41"},{"key":"e_1_2_2_19_1","doi-asserted-by":"crossref","unstructured":"A. Elhayek E. Aguiar A. Jain J. Tompson L. Pishchulin M. Andriluka C. Bregler B. Schiele and C. Theobalt. 2016. MARCOnI - ConvNet-based MARker-less Motion Capture in Outdoor and Indoor Scenes. PAMI (2016).","DOI":"10.1109\/TPAMI.2016.2557779"},{"key":"e_1_2_2_20_1","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v32i1.12270"},{"key":"e_1_2_2_21_1","volume-title":"The lottery ticket hypothesis: Finding sparse, trainable neural networks. arXiv preprint arXiv:1803.03635","author":"Frankle Jonathan","year":"2018","unstructured":"Jonathan Frankle and Michael Carbin. 2018. The lottery ticket hypothesis: Finding sparse, trainable neural networks. arXiv preprint arXiv:1803.03635 (2018)."},{"key":"e_1_2_2_22_1","doi-asserted-by":"crossref","unstructured":"Georgia Gkioxari Bharath Hariharan Ross Girshick and Jitendra Malik. 2014. Using k-poselets for detecting people and localizing their keypoints. In CVPR. 3582--3589.","DOI":"10.1109\/CVPR.2014.458"},{"key":"e_1_2_2_23_1","doi-asserted-by":"publisher","unstructured":"Peng Guan A. Weiss A. O. B\u00c3\u010dlan and M. J. Black. 2009. Estimating human shape and pose from a single image. In CVPR. 1381--1388. 10.1109\/ICCV.2009.5459300","DOI":"10.1109\/ICCV.2009.5459300"},{"key":"e_1_2_2_24_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.01114"},{"key":"e_1_2_2_25_1","doi-asserted-by":"publisher","unstructured":"Riza Alp G\u00fcler Natalia Neverova and Iasonas Kokkinos. 2018. Dense-Pose: Dense Human Pose Estimation in the Wild. In CVPR. 7297--7306. 10.1109\/CVPR.2018.00762","DOI":"10.1109\/CVPR.2018.00762"},{"key":"e_1_2_2_26_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCVW.2019.00226"},{"key":"e_1_2_2_27_1","doi-asserted-by":"publisher","DOI":"10.1145\/3311970"},{"key":"e_1_2_2_28_1","volume-title":"Hao Liu and Yaser Sheikh","author":"Lin Gui Bart Nabbe Lei Tan","year":"2015","unstructured":"Lei Tan Lin Gui Bart Nabbe Iain Matthews Takeo Kanade Shohei Nobuhara Hanbyul Joo, Hao Liu and Yaser Sheikh. 2015. Panoptic Studio: A Massively Multiview System for Social Motion Capture. In ICCV."},{"key":"e_1_2_2_29_1","unstructured":"Kaiming He Xiangyu Zhang Shaoqing Ren and Jian Sun. 2016. Deep Residual Learning for Image Recognition. In CVPR."},{"key":"e_1_2_2_30_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2019.00140"},{"key":"e_1_2_2_31_1","volume-title":"Laurens Van Der Maaten, and Kilian Q Weinberger","author":"Huang Gao","year":"2017","unstructured":"Gao Huang, Zhuang Liu, Laurens Van Der Maaten, and Kilian Q Weinberger. 2017b. Densely connected convolutional networks.. In CVPR."},{"key":"e_1_2_2_32_1","doi-asserted-by":"publisher","DOI":"10.1109\/3DV.2017.00055"},{"key":"e_1_2_2_33_1","volume-title":"SqueezeNet: AlexNet-level accuracy with 50x fewer parameters and &lt;0.5MB model size. arXiv:1602.07360","author":"Iandola Forrest N.","year":"2016","unstructured":"Forrest N. Iandola, Song Han, Matthew W. Moskewicz, Khalid Ashraf, William J. Dally, and Kurt Keutzer. 2016. SqueezeNet: AlexNet-level accuracy with 50x fewer parameters and &lt;0.5MB model size. arXiv:1602.07360 (2016)."},{"key":"e_1_2_2_34_1","doi-asserted-by":"crossref","unstructured":"Eldar Insafutdinov Mykhaylo Andriluka Leonid Pishchulin Siyu Tang Evgeny Levinkov Bjoern Andres Bernt Schiele and Saarland Informatics Campus. 2017. ArtTrack: Articulated multi-person tracking in the wild. In CVPR.","DOI":"10.1109\/CVPR.2017.142"},{"key":"e_1_2_2_35_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2013.248"},{"key":"e_1_2_2_36_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-48881-3_44"},{"key":"e_1_2_2_37_1","doi-asserted-by":"publisher","DOI":"10.1145\/1882261.1866174"},{"key":"e_1_2_2_38_1","doi-asserted-by":"publisher","unstructured":"Sam Johnson and Mark Everingham. 2010. Clustered Pose and Nonlinear Appearance Models for Human Pose Estimation. In BMVC. 10.5244\/C.24.12","DOI":"10.5244\/C.24.12"},{"key":"e_1_2_2_39_1","doi-asserted-by":"crossref","unstructured":"Sam Johnson and Mark Everingham. 2011. Learning Effective Human Pose Estimation from Inaccurate Annotation. In CVPR.","DOI":"10.1109\/CVPR.2011.5995318"},{"key":"e_1_2_2_40_1","volume-title":"Panoptic Studio: A Massively Multiview System for Social Motion Capture. In ICCV. 3334--3342.","author":"Joo Hanbyul","year":"2015","unstructured":"Hanbyul Joo, Hao Liu, Lei Tan, Lin Gui, Bart Nabbe, Iain Matthews, Takeo Kanade, Shohei Nobuhara, and Yaser Sheikh. 2015. Panoptic Studio: A Massively Multiview System for Social Motion Capture. In ICCV. 3334--3342."},{"key":"e_1_2_2_41_1","doi-asserted-by":"crossref","unstructured":"Angjoo Kanazawa Michael J. Black David W. Jacobs and Jitendra Malik. 2018. End-to-end Recovery of Human Shape and Pose. In CVPR.","DOI":"10.1109\/CVPR.2018.00744"},{"key":"e_1_2_2_42_1","doi-asserted-by":"crossref","unstructured":"Angjoo Kanazawa Jason Y. Zhang Panna Felsen and Jitendra Malik. 2019. Learning 3D Human Dynamics from Video. In Computer Vision and Pattern Recognition (CVPR).","DOI":"10.1109\/CVPR.2019.00576"},{"key":"e_1_2_2_43_1","doi-asserted-by":"publisher","DOI":"10.1007\/s11263-018-1066-6"},{"key":"e_1_2_2_44_1","unstructured":"KNG Music. 2019. SE\u00c3\u015aORITA - Shawn Mendes Camila Cabello | Violin Cello and Viola Cover [KNG Music]. https:\/\/www.youtube.com\/watch?v=_xCKmEhKQl4."},{"key":"e_1_2_2_45_1","volume-title":"VIBE: Video Inference for Human Body Pose and Shape Estimation. In The IEEE Conference on Computer Vision and Pattern Recognition (CVPR).","author":"Kocabas Muhammed","unstructured":"Muhammed Kocabas, Nikos Athanasiou, and Michael J. Black. 2020. VIBE: Video Inference for Human Body Pose and Shape Estimation. In The IEEE Conference on Computer Vision and Pattern Recognition (CVPR)."},{"key":"e_1_2_2_46_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2019.00234"},{"key":"e_1_2_2_47_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.00463"},{"key":"e_1_2_2_48_1","volume-title":"Gehler","author":"Lassner Christoph","year":"2017","unstructured":"Christoph Lassner, Javier Romero, Martin Kiefel, Federica Bogo, Michael J. Black, and Peter V. Gehler. 2017. Unite the People: Closing the Loop Between 3D and 2D Human Representations. In CVPR."},{"key":"e_1_2_2_49_1","unstructured":"Sijin Li and Antoni B Chan. 2014. 3d human pose estimation from monocular images with deep convolutional neural network. In ACCV."},{"key":"e_1_2_2_50_1","doi-asserted-by":"crossref","unstructured":"Sijin Li Weichen Zhang and Antoni B Chan. 2015. Maximum-margin structured learning with deep networks for 3d human pose estimation. In ICCV. 2848--2856.","DOI":"10.1109\/ICCV.2015.326"},{"key":"e_1_2_2_51_1","volume-title":"Microsoft coco: Common objects in context","author":"Lin Tsung-Yi","unstructured":"Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Doll\u00e1r, and C Lawrence Zitnick. 2014. Microsoft coco: Common objects in context. In ECCV. Springer, 740--755."},{"key":"e_1_2_2_52_1","doi-asserted-by":"publisher","DOI":"10.1145\/2816795.2818013"},{"key":"e_1_2_2_53_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-01264-9_8"},{"key":"e_1_2_2_54_1","volume-title":"Little","author":"Martinez Julieta","year":"2017","unstructured":"Julieta Martinez, Rayat Hossain, Javier Romero, and James J. Little. 2017. A simple yet effective baseline for 3d human pose estimation. In ICCV."},{"key":"e_1_2_2_55_1","doi-asserted-by":"publisher","DOI":"10.1109\/3dv.2017.00064"},{"key":"e_1_2_2_56_1","volume-title":"Single-Shot Multi-Person 3D Pose Estimation From Monocular RGB. In 3DV","author":"Mehta Dushyant","unstructured":"Dushyant Mehta, Oleksandr Sotnychenko, Franziska Mueller, Weipeng Xu, Srinath Sridhar, Gerard Pons-Moll, and Christian Theobalt. 2018b. Single-Shot Multi-Person 3D Pose Estimation From Monocular RGB. In 3DV. IEEE. http:\/\/gvv.mpi-inf.mpg.de\/projects\/SingleShotMultiPerson"},{"key":"e_1_2_2_57_1","doi-asserted-by":"publisher","DOI":"10.1145\/3072959.3073596"},{"key":"e_1_2_2_58_1","doi-asserted-by":"crossref","unstructured":"Sachin Mehta Mohammad Rastegari Anat Caspi Linda Shapiro and Hannaneh Hajishirzi. 2018a. ESPNet: Efficient Spatial Pyramid of Dilated Convolutions for Semantic Segmentation. In ECCV.","DOI":"10.1007\/978-3-030-01249-6_34"},{"key":"e_1_2_2_59_1","volume-title":"Understanding Motion Capture for Computer Animation","author":"Menache Alberto","unstructured":"Alberto Menache. 2010. Understanding Motion Capture for Computer Animation, Second Edition (2nd ed.). Morgan Kaufmann Publishers Inc., San Francisco, CA, USA.","edition":"2"},{"key":"e_1_2_2_60_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2019.01023"},{"key":"e_1_2_2_61_1","unstructured":"Music Express Magazine. 2013a. Arirang. https:\/\/www.youtube.com\/watch?v=kX6xMYlEwLA."},{"key":"e_1_2_2_62_1","unstructured":"Music Express Magazine. 2013b. Down at the Twist Shout. https:\/\/www.youtube.com\/watch?v=lv-h4WNnw0g."},{"key":"e_1_2_2_63_1","volume-title":"Associative Embedding: End-to-End Learning for Joint Detection and Grouping. In NeurIPS.","author":"Newell Alejandro","year":"2017","unstructured":"Alejandro Newell and Jia Deng. 2017. Associative Embedding: End-to-End Learning for Joint Detection and Grouping. In NeurIPS."},{"key":"e_1_2_2_64_1","doi-asserted-by":"crossref","unstructured":"Aiden Nibali Zhen He Stuart Morgan and Luke Prendergast. 2019. 3D Human Pose Estimation with 2D Marginal Heatmaps. In WACV.","DOI":"10.1109\/WACV.2019.00162"},{"key":"e_1_2_2_65_1","doi-asserted-by":"crossref","unstructured":"Mohamed Omran Christop Lassner Gerard Pons-Moll Peter Gehler and Bernt Schiele. 2018. Neural Body Fitting: Unifying Deep Learning and Model Based Human Pose and Shape Estimation. In 3DV.","DOI":"10.1109\/3DV.2018.00062"},{"key":"e_1_2_2_66_1","doi-asserted-by":"crossref","unstructured":"George Papandreou Tyler Zhu Nori Kanazawa Alexander Toshev Jonathan Tompson Chris Bregler and Kevin Murphy. 2017. Towards Accurate Multi-person Pose Estimation in the Wild. In CVPR.","DOI":"10.1109\/CVPR.2017.395"},{"key":"e_1_2_2_67_1","volume-title":"Black","author":"Pavlakos Georgios","year":"2019","unstructured":"Georgios Pavlakos, Vasileios Choutas, Nima Ghorbani, Timo Bolkart, Ahmed A. A. Osman, Dimitrios Tzionas, and Michael J. Black. 2019. Expressive Body Capture: 3D Hands, Face, and Body from a Single Image. (2019)."},{"key":"e_1_2_2_68_1","doi-asserted-by":"crossref","unstructured":"Georgios Pavlakos Xiaowei Zhou and Kostas Daniilidis. 2018a. Ordinal Depth Supervision for 3D Human Pose Estimation. In CVPR.","DOI":"10.1109\/CVPR.2018.00763"},{"key":"e_1_2_2_69_1","doi-asserted-by":"crossref","unstructured":"Georgios Pavlakos Xiaowei Zhou Konstantinos G Derpanis and Kostas Daniilidis. 2017. Coarse-to-Fine Volumetric Prediction for Single-Image 3D Human Pose. In CVPR.","DOI":"10.1109\/CVPR.2017.139"},{"key":"e_1_2_2_70_1","doi-asserted-by":"crossref","unstructured":"Georgios Pavlakos Luyang Zhu Xiaowei Zhou and Kostas Daniilidis. 2018b. Learning to Estimate 3D Human Pose and Shape from a Single Color Image. In CVPR.","DOI":"10.1109\/CVPR.2018.00055"},{"key":"e_1_2_2_71_1","doi-asserted-by":"crossref","unstructured":"Leonid Pishchulin Eldar Insafutdinov Siyu Tang Bjoern Andres Mykhaylo Andriluka Peter Gehler and Bernt Schiele. 2016. DeepCut: Joint Subset Partition and Labeling for Multi Person Pose Estimation. In CVPR.","DOI":"10.1109\/CVPR.2016.533"},{"key":"e_1_2_2_72_1","volume-title":"Articulated people detection and pose estimation: Reshaping the future","author":"Pishchulin Leonid","unstructured":"Leonid Pishchulin, Arjun Jain, Mykhaylo Andriluka, Thorsten Thorm\u00e4hlen, and Bernt Schiele. 2012. Articulated people detection and pose estimation: Reshaping the future. In CVPR. IEEE, 3178--3185."},{"key":"e_1_2_2_73_1","doi-asserted-by":"crossref","unstructured":"Gerard Pons-Moll David J Fleet and Bodo Rosenhahn. 2014. Posebits for monocular human pose estimation. In CVPR. 2337--2344.","DOI":"10.1109\/CVPR.2014.300"},{"key":"e_1_2_2_74_1","unstructured":"Alin-Ionut Popa Mihai Zanfir and Cristian Sminchisescu. 2017. Deep Multitask Architecture for Integrated 2D and 3D Human Sensing. In CVPR."},{"key":"e_1_2_2_75_1","volume-title":"Reconstructing 3d human pose from 2d image landmarks","author":"Ramakrishna Varun","unstructured":"Varun Ramakrishna, Takeo Kanade, and Yaser Sheikh. 2012. Reconstructing 3d human pose from 2d image landmarks. In ECCV. Springer, 573--586."},{"key":"e_1_2_2_76_1","volume-title":"Proceedings of the European Conference on Computer Vision (ECCV). 68--84","author":"Imtiaz Hossain Mir Rayat","year":"2018","unstructured":"Mir Rayat Imtiaz Hossain and James J Little. 2018. Exploiting temporal information for 3d human pose estimation. In Proceedings of the European Conference on Computer Vision (ECCV). 68--84."},{"key":"e_1_2_2_77_1","volume-title":"Regularized evolution for image classifier architecture search. 33","author":"Real Esteban","year":"2019","unstructured":"Esteban Real, Alok Aggarwal, Yanping Huang, and Quoc V Le. 2019. Regularized evolution for image classifier architecture search. 33 (2019), 4780--4789."},{"key":"e_1_2_2_78_1","unstructured":"Shaoqing Ren Kaiming He Ross Girshick and Jian Sun. 2015. Faster R-CNN: Towards real-time object detection with region proposal networks. In NeurIPS. 91--99."},{"key":"e_1_2_2_79_1","doi-asserted-by":"publisher","DOI":"10.1145\/2980179.2980235"},{"key":"e_1_2_2_80_1","doi-asserted-by":"publisher","unstructured":"Helge Rhodin Nadia Robertini Dan Casas Christian Richardt Hans-Peter Seidel and Christian Theobalt. 2016b. General Automatic Human Shape and Motion Capture Using Volumetric Contour Cues. In ECCV. 509--526. 10.1007\/978-3-319-46454-1_31","DOI":"10.1007\/978-3-319-46454-1_31"},{"key":"e_1_2_2_81_1","doi-asserted-by":"crossref","unstructured":"Gregory Rogez Philippe Weinzaepfel and Cordelia Schmid. 2017. LCR-Net: Localization-Classification-Regression for Human Pose. In CVPR.","DOI":"10.1109\/CVPR.2017.134"},{"key":"e_1_2_2_82_1","volume-title":"Multi-person 2D and 3D Pose Detection in Natural Images. PAMI","author":"Rogez Gr\u00e9gory","year":"2019","unstructured":"Gr\u00e9gory Rogez, Philippe Weinzaepfel, and Cordelia Schmid. 2019. LCR-Net++: Multi-person 2D and 3D Pose Detection in Natural Images. PAMI (2019)."},{"key":"e_1_2_2_83_1","doi-asserted-by":"publisher","DOI":"10.1109\/TITS.2017.2750080"},{"key":"e_1_2_2_84_1","volume-title":"Mobilenetv2: Inverted residuals and linear bottlenecks","author":"Sandler Mark","unstructured":"Mark Sandler, Andrew Howard, Menglong Zhu, Andrey Zhmoginov, and Liang-Chieh Chen. 2018. Mobilenetv2: Inverted residuals and linear bottlenecks. In CVPR. IEEE, 4510--4520."},{"key":"e_1_2_2_85_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.cviu.2016.09.002"},{"key":"e_1_2_2_86_1","doi-asserted-by":"publisher","DOI":"10.1007\/s11263-009-0293-2"},{"key":"e_1_2_2_87_1","doi-asserted-by":"publisher","DOI":"10.1109\/MCG.2007.68"},{"key":"e_1_2_2_88_1","doi-asserted-by":"crossref","unstructured":"Carsten Stoll Nils Hasler Juergen Gall Hans-Peter Seidel and Christian Theobalt. 2011. Fast articulated motion tracking using a sums of Gaussians body model. In ICCV. 951--958.","DOI":"10.1109\/ICCV.2011.6126338"},{"key":"e_1_2_2_89_1","doi-asserted-by":"crossref","unstructured":"Ke Sun Bin Xiao Dong Liu and Jingdong Wang. 2019a. Deep High-Resolution Representation Learning for Human Pose Estimation. In CVPR.","DOI":"10.1109\/CVPR.2019.00584"},{"key":"e_1_2_2_90_1","volume-title":"Articulated part-based model for joint object detection and pose estimation","author":"Sun Min","unstructured":"Min Sun and Silvio Savarese. 2011. Articulated part-based model for joint object detection and pose estimation. In ICCV. IEEE, 723--730."},{"key":"e_1_2_2_91_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2017.284"},{"key":"e_1_2_2_92_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-01231-1_33"},{"key":"e_1_2_2_93_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2019.00545"},{"key":"e_1_2_2_94_1","first-page":"12","article-title":"Inception-v4, inception-resnet and the impact of residual connections on learning","volume":"4","author":"Szegedy Christian","year":"2017","unstructured":"Christian Szegedy, Sergey Ioffe, Vincent Vanhoucke, and Alexander A Alemi. 2017. Inception-v4, inception-resnet and the impact of residual connections on learning.. In AAAI, Vol. 4. 12.","journal-title":"AAAI"},{"key":"e_1_2_2_95_1","volume-title":"Efficientnet: Rethinking model scaling for convolutional neural networks. arXiv preprint arXiv:1905.11946","author":"Tan Mingxing","year":"2019","unstructured":"Mingxing Tan and Quoc V Le. 2019. Efficientnet: Rethinking model scaling for convolutional neural networks. arXiv preprint arXiv:1905.11946 (2019)."},{"key":"e_1_2_2_96_1","doi-asserted-by":"crossref","unstructured":"Bugra Tekin Isinsu Katircioglu Mathieu Salzmann Vincent Lepetit and Pascal Fua. 2016. Structured Prediction of 3D Human Pose with Deep Neural Networks. In BMVC.","DOI":"10.5244\/C.30.130"},{"key":"e_1_2_2_97_1","doi-asserted-by":"crossref","unstructured":"Bugra Tekin Pablo M\u00e1rquez-Neila Mathieu Salzmann and Pascal Fua. 2017. Learning to Fuse 2D and 3D Image Cues for Monocular Body Pose Estimation. In ICCV.","DOI":"10.1109\/ICCV.2017.425"},{"key":"e_1_2_2_98_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2019.00782"},{"key":"e_1_2_2_99_1","unstructured":"TPLA:Terra Prime Light Armory. 2016. Commited Sparring. https:\/\/www.youtube.com\/watch?v=xmFVfUKr1MQ."},{"key":"e_1_2_2_100_1","doi-asserted-by":"crossref","unstructured":"Matthew Trumble Andrew Gilbert Charles Malleson Adrian Hilton and John Collomosse. 2017. Total Capture: 3D Human Pose Estimation Fusing Video and Inertial Sensors. In BMVC. 1--13.","DOI":"10.5244\/C.31.14"},{"key":"e_1_2_2_101_1","unstructured":"Hsiao-Yu Tung Hsiao-Wei Tung Ersin Yumer and Katerina Fragkiadaki. 2017. Self-supervised Learning of Motion Capture. In NeurIPS. 5242--5252."},{"key":"e_1_2_2_102_1","doi-asserted-by":"crossref","unstructured":"Timo von Marcard Roberto Henschel Michael Black Bodo Rosenhahn and Gerard Pons-Moll. 2018. Recovering Accurate 3D Human Pose in The Wild Using IMUs and a Moving Camera. In ECCV.","DOI":"10.1007\/978-3-030-01249-6_37"},{"key":"e_1_2_2_103_1","first-page":"8","article-title":"Human Pose Estimation from Video and IMUs","volume":"38","author":"von Marcard Timo","year":"2016","unstructured":"Timo von Marcard, Gerard Pons-Moll, and Bodo Rosenhahn. 2016. Human Pose Estimation from Video and IMUs. PAMI 38, 8 (Jan. 2016), 1533--1547.","journal-title":"PAMI"},{"key":"e_1_2_2_104_1","volume-title":"Deep High-Resolution Representation Learning for Visual Recognition. TPAMI","author":"Wang Jingdong","year":"2019","unstructured":"Jingdong Wang, Ke Sun, Tianheng Cheng, Borui Jiang, Chaorui Deng, Yang Zhao, Dong Liu, Yadong Mu, Mingkui Tan, Xinggang Wang, Wenyu Liu, and Bin Xiao. 2019. Deep High-Resolution Representation Learning for Visual Recognition. TPAMI (2019)."},{"key":"e_1_2_2_105_1","doi-asserted-by":"publisher","DOI":"10.1145\/1778765.1778779"},{"key":"e_1_2_2_106_1","doi-asserted-by":"publisher","DOI":"10.1007\/s11390-017-1742-y"},{"key":"e_1_2_2_107_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.01122"},{"key":"e_1_2_2_108_1","volume-title":"Aggregated residual transformations for deep neural networks","author":"Xie Saining","unstructured":"Saining Xie, Ross Girshick, Piotr Doll\u00e1r, Zhuowen Tu, and Kaiming He. 2017. Aggregated residual transformations for deep neural networks. In CVPR. IEEE, 5987--5995."},{"key":"e_1_2_2_109_1","volume-title":"Mo 2 Cap 2: Real-time Mobile 3D Motion Capture with a Cap-mounted Fisheye Camera","author":"Xu Weipeng","year":"2019","unstructured":"Weipeng Xu, Avishek Chatterjee, Michael Zollhoefer, Helge Rhodin, Pascal Fua, Hans-Peter Seidel, and Christian Theobalt. 2019. Mo 2 Cap 2: Real-time Mobile 3D Motion Capture with a Cap-mounted Fisheye Camera. IEEE transactions on visualization and computer graphics 25, 5 (2019), 2093--2101."},{"key":"e_1_2_2_110_1","doi-asserted-by":"publisher","DOI":"10.1145\/3181973"},{"key":"e_1_2_2_111_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00551"},{"key":"e_1_2_2_112_1","doi-asserted-by":"crossref","unstructured":"Andrei Zanfir Elisabeta Marinoiu and Cristian Sminchisescu. 2018a. Monocular 3D Pose and Shape Estimation of Multiple People in Natural Scenes-The Importance of Multiple Scene Constraints. In CVPR. 2148--2157.","DOI":"10.1109\/CVPR.2018.00229"},{"key":"e_1_2_2_113_1","unstructured":"Andrei Zanfir Elisabeta Marinoiu Mihai Zanfir Alin-Ionut Popa and Cristian Sminchisescu. 2018b. Deep Network for the Integrated 3D Sensing of Multiple People in Natural Images. In NeurIPS."},{"key":"e_1_2_2_114_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00716"},{"key":"e_1_2_2_115_1","doi-asserted-by":"crossref","unstructured":"Xingyi Zhou Qixing Huang Xiao Sun Xiangyang Xue and Yichen Wei. 2017. Towards 3D Human Pose Estimation in the Wild: A Weakly-Supervised Approach. In CVPR. 398--407.","DOI":"10.1109\/ICCV.2017.51"}],"container-title":["ACM Transactions on Graphics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3386569.3392410","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3386569.3392410","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,25]],"date-time":"2025-06-25T05:35:50Z","timestamp":1750829750000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3386569.3392410"}},"subtitle":["real-time multi-person 3D motion capture with a single RGB camera"],"short-title":[],"issued":{"date-parts":[[2020,8,12]]},"references-count":115,"journal-issue":{"issue":"4","published-print":{"date-parts":[[2020,8,31]]}},"alternative-id":["10.1145\/3386569.3392410"],"URL":"https:\/\/doi.org\/10.1145\/3386569.3392410","relation":{},"ISSN":["0730-0301","1557-7368"],"issn-type":[{"value":"0730-0301","type":"print"},{"value":"1557-7368","type":"electronic"}],"subject":[],"published":{"date-parts":[[2020,8,12]]},"assertion":[{"value":"2020-08-12","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}