{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,3]],"date-time":"2026-07-03T19:02:07Z","timestamp":1783105327957,"version":"3.54.6"},"reference-count":72,"publisher":"Association for Computing Machinery (ACM)","issue":"6","license":[{"start":{"date-parts":[[2016,11,11]],"date-time":"2016-11-11T00:00:00Z","timestamp":1478822400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"name":"ERC","award":["335545"],"award-info":[{"award-number":["335545"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Graph."],"published-print":{"date-parts":[[2016,11,11]]},"abstract":"<jats:p>\n            Marker-based and marker-less optical skeletal motion-capture methods use an\n            <jats:italic>outside-in<\/jats:italic>\n            arrangement of cameras placed around a scene, with viewpoints converging on the center. They often create discomfort with marker suits, and their recording volume is severely restricted and often constrained to indoor scenes with controlled backgrounds. Alternative suit-based systems use several inertial measurement units or an exoskeleton to capture motion with an\n            <jats:italic>inside-in<\/jats:italic>\n            setup, i.e. without external sensors. This makes capture independent of a confined volume, but requires substantial, often constraining, and hard to set up body instrumentation. Therefore, we propose a new method for real-time, marker-less, and egocentric motion capture: estimating the full-body skeleton pose from a lightweight stereo pair of fisheye cameras attached to a helmet or virtual reality headset - an\n            <jats:italic>optical inside-in<\/jats:italic>\n            method, so to speak. This allows full-body motion capture in general indoor and outdoor scenes, including crowded scenes with many people nearby, which enables reconstruction in larger-scale activities. Our approach combines the strength of a new generative pose estimation framework for fisheye views with a ConvNet-based body-part detector trained on a large new dataset. It is particularly useful in virtual reality to freely roam and interact, while seeing the fully motion-captured virtual body.\n          <\/jats:p>","DOI":"10.1145\/2980179.2980235","type":"journal-article","created":{"date-parts":[[2016,11,11]],"date-time":"2016-11-11T17:02:54Z","timestamp":1478883774000},"page":"1-11","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":143,"title":["EgoCap"],"prefix":"10.1145","volume":"35","author":[{"given":"Helge","family":"Rhodin","sequence":"first","affiliation":[{"name":"Max Planck Institute for Informatics"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Christian","family":"Richardt","sequence":"additional","affiliation":[{"name":"Max Planck Institute for Informatics and Intel Visual Computing Institute and University of Bath"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Dan","family":"Casas","sequence":"additional","affiliation":[{"name":"Max Planck Institute for Informatics"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Eldar","family":"Insafutdinov","sequence":"additional","affiliation":[{"name":"Max Planck Institute for Informatics"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Mohammad","family":"Shafiei","sequence":"additional","affiliation":[{"name":"Max Planck Institute for Informatics"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Hans-Peter","family":"Seidel","sequence":"additional","affiliation":[{"name":"Max Planck Institute for Informatics"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Bernt","family":"Schiele","sequence":"additional","affiliation":[{"name":"Max Planck Institute for Informatics"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Christian","family":"Theobalt","sequence":"additional","affiliation":[{"name":"Max Planck Institute for Informatics"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2016,12,5]]},"reference":[{"key":"e_1_2_2_1_1","unstructured":"Amin S. Andriluka M. Rohrbach M. and Schiele B. 2009. Multi-view pictorial structures for 3D human pose estimation. In BMVC.  Amin S. Andriluka M. Rohrbach M. and Schiele B. 2009. Multi-view pictorial structures for 3D human pose estimation. In BMVC."},{"key":"e_1_2_2_2_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2014.471"},{"key":"e_1_2_2_3_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2011.6126356"},{"key":"e_1_2_2_4_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2014.216"},{"key":"e_1_2_2_5_1","doi-asserted-by":"publisher","DOI":"10.1145\/311535.311556"},{"key":"e_1_2_2_6_1","unstructured":"Bregler C. and Malik J. 1998. Tracking people with twists and exponential maps. In CVPR.   Bregler C. and Malik J. 1998. Tracking people with twists and exponential maps. In CVPR."},{"key":"e_1_2_2_7_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2013.464"},{"key":"e_1_2_2_8_1","doi-asserted-by":"publisher","DOI":"10.1007\/s00371-005-0287-1"},{"key":"e_1_2_2_9_1","doi-asserted-by":"publisher","DOI":"10.1145\/1073204.1073248"},{"key":"e_1_2_2_10_1","unstructured":"Chen X. and Yuille A. L. 2014. Articulated pose estimation by a graphical model with image dependent pairwise relations. In NIPS.   Chen X. and Yuille A. L. 2014. Articulated pose estimation by a graphical model with image dependent pairwise relations. In NIPS."},{"key":"e_1_2_2_11_1","unstructured":"EgoCap 2016. EgoCap dataset. http:\/\/gvv.mpi-inf.mpg.de\/projects\/EgoCap\/.  EgoCap 2016. EgoCap dataset. http:\/\/gvv.mpi-inf.mpg.de\/projects\/EgoCap\/."},{"key":"e_1_2_2_12_1","doi-asserted-by":"crossref","unstructured":"Elhayek A. de Aguiar E. Jain A. Tompson J. Pishchulin L. Andriluka M. Bregler C. Schiele B. and Theobalt C. 2015. Efficient ConvNet-based markerless motion capture in general scenes with a low number of cameras. In CVPR.  Elhayek A. de Aguiar E. Jain A. Tompson J. Pishchulin L. Andriluka M. Bregler C. Schiele B. and Theobalt C. 2015. Efficient ConvNet-based markerless motion capture in general scenes with a low number of cameras. In CVPR.","DOI":"10.1109\/CVPR.2015.7299005"},{"key":"e_1_2_2_13_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2011.6126269"},{"key":"e_1_2_2_14_1","doi-asserted-by":"publisher","DOI":"10.1007\/s11263-008-0173-1"},{"key":"e_1_2_2_15_1","doi-asserted-by":"publisher","DOI":"10.1145\/2019406.2019424"},{"key":"e_1_2_2_16_1","doi-asserted-by":"crossref","unstructured":"He K. Zhang X. Ren S. and Sun J. 2016. Deep residual learning for image recognition. In CVPR.  He K. Zhang X. Ren S. and Sun J. 2016. Deep residual learning for image recognition. In CVPR.","DOI":"10.1109\/CVPR.2016.90"},{"key":"e_1_2_2_17_1","doi-asserted-by":"publisher","DOI":"10.1109\/JSTSP.2012.2196975"},{"key":"e_1_2_2_18_1","doi-asserted-by":"crossref","unstructured":"Insafutdinov E. Pishchulin L. Andres B. Andriluka M. and Schiele B. 2016. DeeperCut: A deeper stronger and faster multi-person pose estimation model. In ECCV.  Insafutdinov E. Pishchulin L. Andres B. Andriluka M. and Schiele B. 2016. DeeperCut: A deeper stronger and faster multi-person pose estimation model. In ECCV.","DOI":"10.1007\/978-3-319-46466-4_3"},{"key":"e_1_2_2_19_1","unstructured":"Jain A. Tompson J. Andriluka M. Taylor G. W. and Bregler C. 2014. Learning human pose estimation features with convolutional networks. In ICLR.  Jain A. Tompson J. Andriluka M. Taylor G. W. and Bregler C. 2014. Learning human pose estimation features with convolutional networks. In ICLR."},{"key":"e_1_2_2_20_1","doi-asserted-by":"crossref","unstructured":"Jain A. Tompson J. LeCun Y. and Bregler C. 2015. MoDeep: A deep learning framework using motion features for human pose estimation. In ACCV.  Jain A. Tompson J. LeCun Y. and Bregler C. 2015. MoDeep: A deep learning framework using motion features for human pose estimation. In ACCV.","DOI":"10.1007\/978-3-319-16808-1_21"},{"key":"e_1_2_2_21_1","doi-asserted-by":"crossref","unstructured":"Jiang H. and Grauman K. 2016. Seeing invisible poses: Estimating 3D body pose from egocentric video. arXiv:1603.07763.  Jiang H. and Grauman K. 2016. Seeing invisible poses: Estimating 3D body pose from egocentric video. arXiv:1603.07763.","DOI":"10.1109\/CVPR.2017.373"},{"key":"e_1_2_2_22_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2011.5995318"},{"key":"e_1_2_2_23_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVMP.2011.24"},{"key":"e_1_2_2_24_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2015.381"},{"key":"e_1_2_2_25_1","doi-asserted-by":"publisher","DOI":"10.1145\/2380116.2380139"},{"key":"e_1_2_2_26_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2011.5995406"},{"key":"e_1_2_2_27_1","doi-asserted-by":"publisher","DOI":"10.1145\/2661229.2661273"},{"key":"e_1_2_2_28_1","doi-asserted-by":"crossref","unstructured":"Ma M. Fan H. and Kitani K. M. 2016. Going deeper into first-person activity recognition. In CVPR.  Ma M. Fan H. and Kitani K. M. 2016. Going deeper into first-person activity recognition. In CVPR.","DOI":"10.1109\/CVPR.2016.209"},{"key":"e_1_2_2_29_1","doi-asserted-by":"publisher","DOI":"10.1145\/2897824.2925907"},{"key":"e_1_2_2_30_1","volume-title":"Understanding Motion Capture for Computer Animation","author":"Menache A.","unstructured":"Menache , A. 2010. Understanding Motion Capture for Computer Animation , 2 nd ed. Morgan Kaufmann . Menache, A. 2010. Understanding Motion Capture for Computer Animation, 2nd ed. Morgan Kaufmann.","edition":"2"},{"key":"e_1_2_2_31_1","volume-title":"Eds","author":"Moeslund T. B.","year":"2011","unstructured":"Moeslund , T. B. , Hilton , A. , Kr\u00fcger , V. , and Sigal , L. , Eds . 2011 . Visual Analysis of Humans: Looking at People. Springer . Moeslund, T. B., Hilton, A., Kr\u00fcger, V., and Sigal, L., Eds. 2011. Visual Analysis of Humans: Looking at People. Springer."},{"key":"e_1_2_2_32_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2013.403"},{"key":"e_1_2_2_33_1","unstructured":"Murray R. M. Sastry S. S. and Zexiang L. 1994. A Mathematical Introduction to Robotic Manipulation. CRC Press.   Murray R. M. Sastry S. S. and Zexiang L. 1994. A Mathematical Introduction to Robotic Manipulation. CRC Press."},{"key":"e_1_2_2_34_1","doi-asserted-by":"crossref","unstructured":"Newell A. Yang K. and Deng J. 2016. Stacked hourglass networks for human pose estimation. arXiv:1603.06937.  Newell A. Yang K. and Deng J. 2016. Stacked hourglass networks for human pose estimation. arXiv:1603.06937.","DOI":"10.1007\/978-3-319-46484-8_29"},{"key":"e_1_2_2_35_1","doi-asserted-by":"crossref","unstructured":"Ohnishi K. Kanehira A. Kanezaki A. and Harada T. 2016. Recognizing activities of daily living with a wrist-mounted camera. In CVPR.  Ohnishi K. Kanehira A. Kanezaki A. and Harada T. 2016. Recognizing activities of daily living with a wrist-mounted camera. In CVPR.","DOI":"10.1109\/CVPR.2016.338"},{"key":"e_1_2_2_36_1","doi-asserted-by":"publisher","DOI":"10.1145\/1360612.1360695"},{"key":"e_1_2_2_37_1","unstructured":"Park H. S. Jain E. and Sheikh Y. 2012. 3D social saliency from head-mounted cameras. In NIPS.   Park H. S. Jain E. and Sheikh Y. 2012. 3D social saliency from head-mounted cameras. In NIPS."},{"key":"e_1_2_2_38_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2015.222"},{"key":"e_1_2_2_39_1","doi-asserted-by":"crossref","unstructured":"Pishchulin L. Insafutdinov E. Tang S. Andres B. Andriluka M. Gehler P. and Schiele B. 2016. Deep-Cut: Joint subset partition and labeling for multi person pose estimation. In CVPR.  Pishchulin L. Insafutdinov E. Tang S. Andres B. Andriluka M. Gehler P. and Schiele B. 2016. Deep-Cut: Joint subset partition and labeling for multi person pose estimation. In CVPR.","DOI":"10.1109\/CVPR.2016.533"},{"key":"e_1_2_2_40_1","doi-asserted-by":"crossref","unstructured":"Pons-Moll G. Baak A. Helten T. M\u00fcller M. Seidel H.-P. and Rosenhahn B. 2010. Multisensor-fusion for 3D full-body human motion capture. In CVPR.  Pons-Moll G. Baak A. Helten T. M\u00fcller M. Seidel H.-P. and Rosenhahn B. 2010. Multisensor-fusion for 3D full-body human motion capture. In CVPR.","DOI":"10.1109\/CVPR.2010.5540153"},{"key":"e_1_2_2_41_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2011.6126375"},{"key":"e_1_2_2_42_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2014.300"},{"key":"e_1_2_2_43_1","doi-asserted-by":"crossref","unstructured":"Rhinehart N. and Kitani K. M. 2016. Learning action maps of large environments via first-person vision. In CVPR.  Rhinehart N. and Kitani K. M. 2016. Learning action maps of large environments via first-person vision. In CVPR.","DOI":"10.1109\/CVPR.2016.69"},{"key":"e_1_2_2_44_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2015.94"},{"key":"e_1_2_2_45_1","doi-asserted-by":"crossref","unstructured":"Rhodin H. Robertini N. Casas D. Richardt C. Seidel H.-P. and Theobalt C. 2016. General automatic human shape and motion capture using volumetric contour cues. In ECCV.  Rhodin H. Robertini N. Casas D. Richardt C. Seidel H.-P. and Theobalt C. 2016. General automatic human shape and motion capture using volumetric contour cues. In ECCV.","DOI":"10.1007\/978-3-319-46454-1_31"},{"key":"e_1_2_2_46_1","volume-title":"ECCV Workshops.","author":"Rogez G.","unstructured":"Rogez , G. , Khademi , M. , Supancic , III, J. S. , Montiel , J. M. M. , and Ramanan , D . 2014. 3D hand pose detection in egocentric RGB-D images . In ECCV Workshops. Rogez, G., Khademi, M., Supancic, III, J. S., Montiel, J. M. M., and Ramanan, D. 2014. 3D hand pose detection in egocentric RGB-D images. In ECCV Workshops."},{"key":"e_1_2_2_47_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2013.471"},{"key":"e_1_2_2_48_1","doi-asserted-by":"crossref","unstructured":"Scaramuzza D. Martinelli A. and Siegwart R. 2006. A toolbox for easily calibrating omnidirectional cameras. In IROS.  Scaramuzza D. Martinelli A. and Siegwart R. 2006. A toolbox for easily calibrating omnidirectional cameras. In IROS.","DOI":"10.1109\/IROS.2006.282372"},{"key":"e_1_2_2_49_1","doi-asserted-by":"publisher","DOI":"10.1145\/2010324.1964926"},{"key":"e_1_2_2_50_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2011.5995316"},{"key":"e_1_2_2_51_1","doi-asserted-by":"publisher","DOI":"10.1007\/s11263-009-0273-6"},{"key":"e_1_2_2_52_1","doi-asserted-by":"publisher","DOI":"10.1007\/s11263-011-0493-4"},{"key":"e_1_2_2_53_1","doi-asserted-by":"crossref","unstructured":"Sridhar S. Mueller F. Oulasvirta A. and Theobalt C. 2015. Fast and robust hand tracking using detection-guided optimization. In CVPR.  Sridhar S. Mueller F. Oulasvirta A. and Theobalt C. 2015. Fast and robust hand tracking using detection-guided optimization. In CVPR.","DOI":"10.1109\/CVPR.2015.7298941"},{"key":"e_1_2_2_54_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2011.6126338"},{"key":"e_1_2_2_55_1","doi-asserted-by":"crossref","unstructured":"Su Y.-C. and Grauman K. 2016. Detecting engagement in egocentric video. In ECCV.  Su Y.-C. and Grauman K. 2016. Detecting engagement in egocentric video. In ECCV.","DOI":"10.1007\/978-3-319-46454-1_28"},{"key":"e_1_2_2_56_1","doi-asserted-by":"publisher","DOI":"10.1145\/2807442.2807445"},{"key":"e_1_2_2_57_1","doi-asserted-by":"publisher","DOI":"10.1145\/1966394.1966397"},{"key":"e_1_2_2_58_1","doi-asserted-by":"crossref","unstructured":"Tekin B. Rozantsev A. Lepetit V. and Fua P. 2016. Direct prediction of 3D body poses from motion compensated sequences. In CVPR.  Tekin B. Rozantsev A. Lepetit V. and Fua P. 2016. Direct prediction of 3D body poses from motion compensated sequences. In CVPR.","DOI":"10.1109\/CVPR.2016.113"},{"key":"e_1_2_2_59_1","doi-asserted-by":"crossref","unstructured":"Theobalt C. de Aguiar E. Stoll C. Seidel H.-P. and Thrun S. 2010. Performance capture from multi-view video. In Image and Geometry Processing for 3-D Cinematography R. Ronfard and G. Taubin Eds. Springer 127--149.  Theobalt C. de Aguiar E. Stoll C. Seidel H.-P. and Thrun S. 2010. Performance capture from multi-view video. In Image and Geometry Processing for 3-D Cinematography R. Ronfard and G. Taubin Eds. Springer 127--149.","DOI":"10.1007\/978-3-642-12392-4_6"},{"key":"e_1_2_2_60_1","unstructured":"Tompson J. J. Jain A. LeCun Y. and Bregler C. 2014. Joint training of a convolutional network and a graphical model for human pose estimation. In NIPS.   Tompson J. J. Jain A. LeCun Y. and Bregler C. 2014. Joint training of a convolutional network and a graphical model for human pose estimation. In NIPS."},{"key":"e_1_2_2_61_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2014.214"},{"key":"e_1_2_2_62_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.cviu.2006.08.006"},{"key":"e_1_2_2_63_1","doi-asserted-by":"publisher","DOI":"10.1145\/1276377.1276421"},{"key":"e_1_2_2_64_1","doi-asserted-by":"publisher","DOI":"10.1145\/1531326.1531369"},{"key":"e_1_2_2_65_1","doi-asserted-by":"crossref","unstructured":"Wang J. Cheng Y. and Feris R. S. 2016. Walk and learn: Facial attribute representation learning from egocentric video and contextual data. In CVPR.  Wang J. Cheng Y. and Feris R. S. 2016. Walk and learn: Facial attribute representation learning from egocentric video and contextual data. In CVPR.","DOI":"10.1109\/CVPR.2016.252"},{"key":"e_1_2_2_66_1","doi-asserted-by":"publisher","DOI":"10.1145\/2366145.2366207"},{"key":"e_1_2_2_67_1","doi-asserted-by":"crossref","unstructured":"Wei S.-E. Ramakrishna V. Kanade T. and Sheikh Y. 2016. Convolutional pose machines. In CVPR.  Wei S.-E. Ramakrishna V. Kanade T. and Sheikh Y. 2016. Convolutional pose machines. In CVPR.","DOI":"10.1109\/CVPR.2016.511"},{"key":"e_1_2_2_68_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2012.261"},{"key":"e_1_2_2_69_1","doi-asserted-by":"crossref","unstructured":"Yasin H. Iqbal U. Kr\u00fcger B. Weber A. and Gall J. 2016. A dual-source approach for 3D pose estimation from a single image. In CVPR.  Yasin H. Iqbal U. Kr\u00fcger B. Weber A. and Gall J. 2016. A dual-source approach for 3D pose estimation from a single image. In CVPR.","DOI":"10.1109\/CVPR.2016.535"},{"key":"e_1_2_2_70_1","unstructured":"Yin K. and Pai D. K. 2003. Footsee: an interactive animation system. In SCA.   Yin K. and Pai D. K. 2003. Footsee: an interactive animation system. In SCA."},{"key":"e_1_2_2_71_1","volume-title":"International Conference on Machine Vision Applications (MVA).","author":"Yonemoto H.","unstructured":"Yonemoto , H. , Murasaki , K. , Osawa , T. , Sudo , K. , Shimamura , J. , and Taniguchi , Y . 2015. Egocentric articulated pose tracking for action recognition . In International Conference on Machine Vision Applications (MVA). Yonemoto, H., Murasaki, K., Osawa, T., Sudo, K., Shimamura, J., and Taniguchi, Y. 2015. Egocentric articulated pose tracking for action recognition. In International Conference on Machine Vision Applications (MVA)."},{"key":"e_1_2_2_72_1","doi-asserted-by":"publisher","DOI":"10.1145\/2661229.2661286"}],"container-title":["ACM Transactions on Graphics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/2980179.2980235","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/2980179.2980235","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T03:49:57Z","timestamp":1750218597000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/2980179.2980235"}},"subtitle":["egocentric marker-less motion capture with two fisheye cameras"],"short-title":[],"issued":{"date-parts":[[2016,11,11]]},"references-count":72,"journal-issue":{"issue":"6","published-print":{"date-parts":[[2016,11,11]]}},"alternative-id":["10.1145\/2980179.2980235"],"URL":"https:\/\/doi.org\/10.1145\/2980179.2980235","relation":{},"ISSN":["0730-0301","1557-7368"],"issn-type":[{"value":"0730-0301","type":"print"},{"value":"1557-7368","type":"electronic"}],"subject":[],"published":{"date-parts":[[2016,11,11]]},"assertion":[{"value":"2016-12-05","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}