{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,2]],"date-time":"2026-07-02T03:27:43Z","timestamp":1782962863401,"version":"3.54.5"},"reference-count":42,"publisher":"Springer Science and Business Media LLC","issue":"6","license":[{"start":{"date-parts":[[2019,12,17]],"date-time":"2019-12-17T00:00:00Z","timestamp":1576540800000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2019,12,17]],"date-time":"2019-12-17T00:00:00Z","timestamp":1576540800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"DOI":"10.13039\/501100006041","name":"Innovate UK","doi-asserted-by":"publisher","award":["102685"],"award-info":[{"award-number":["102685"]}],"id":[{"id":"10.13039\/501100006041","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100007601","name":"Horizon 2020","doi-asserted-by":"publisher","award":["687800"],"award-info":[{"award-number":["687800"]}],"id":[{"id":"10.13039\/501100007601","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Int J Comput Vis"],"published-print":{"date-parts":[[2020,6]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:p>A real-time motion capture system is presented which uses input from multiple standard video cameras and inertial measurement units (IMUs). The system is able to track multiple people simultaneously and requires no optical markers, specialized infra-red cameras or foreground\/background segmentation, making it applicable to general indoor and outdoor scenarios with dynamic backgrounds and lighting. To overcome limitations of prior video or IMU-only approaches, we propose to use flexible combinations of multiple-view, calibrated video and IMU input along with a pose prior in an online optimization-based framework, which allows the full 6-DoF motion to be recovered including axial rotation of limbs and drift-free global position. A method for sorting and assigning raw input 2D keypoint detections into corresponding subjects is presented which facilitates multi-person tracking and rejection of any bystanders in the scene. The approach is evaluated on data from several indoor and outdoor capture environments with one or more subjects and the trade-off between input sparsity and tracking performance is discussed. State-of-the-art pose estimation performance is obtained on the Total Capture (mutli-view video and IMU) and Human 3.6M (multi-view video) datasets. Finally, a live demonstrator for the approach is presented showing real-time capture, solving and character animation using a light-weight, commodity hardware setup.<\/jats:p>","DOI":"10.1007\/s11263-019-01270-5","type":"journal-article","created":{"date-parts":[[2019,12,17]],"date-time":"2019-12-17T12:02:43Z","timestamp":1576584163000},"page":"1594-1611","update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":50,"title":["Real-Time Multi-person Motion Capture from Multi-view Video and IMUs"],"prefix":"10.1007","volume":"128","author":[{"ORCID":"https:\/\/orcid.org\/0000-0003-0649-3064","authenticated-orcid":false,"given":"Charles","family":"Malleson","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"John","family":"Collomosse","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Adrian","family":"Hilton","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"297","published-online":{"date-parts":[[2019,12,17]]},"reference":[{"key":"1270_CR1","unstructured":"Agarwal, S., & Mierle, K, et al. (2017). Ceres solver. Retrieved July 20, 2017 from http:\/\/ceres-solver.org."},{"key":"1270_CR2","unstructured":"Alp\u00a0G\u00fcler, R., Neverova, N., Kokkinos, I. (2018). Densepose: Dense human pose estimation in the wild. In Conference on computer vision and pattern recognition (CVPR)."},{"key":"1270_CR3","doi-asserted-by":"publisher","unstructured":"Andrews, S., Huerta, I., Komura, T., Sigal, L., & Mitchell, K. (2016). Real-time physics-based motion capture with sparse sensors. In Proceedings of the 13th European conference on visual media production (CVMP 2016). https:\/\/doi.org\/10.1145\/2998559.2998564.","DOI":"10.1145\/2998559.2998564"},{"key":"1270_CR4","doi-asserted-by":"crossref","unstructured":"Cao, Z., Simon, T., Wei, S. E., & Sheikh, Y. (2017). Realtime multi-person 2D pose estimation using part affinity fields. In Conference on computer vision and pattern recognition (CVPR).","DOI":"10.1109\/CVPR.2017.143"},{"key":"1270_CR5","unstructured":"Captury, T. (2017). The Captury markerless motion capture technology. Retrieved July 20, 2017 from http:\/\/thecaptury.com\/."},{"key":"1270_CR6","doi-asserted-by":"publisher","unstructured":"Elhayek, A., De Aguiar, E., Jain, A., Tompson, J., Pishchulin, L., Andriluka, M., et al. (2015). Efficient ConvNet-based marker-less motion capture in general scenes with a low number of cameras. In Conference on Computer Vision and Pattern Recognition (CVPR) (pp. 3810\u20133818). https:\/\/doi.org\/10.1109\/CVPR.2015.7299005.","DOI":"10.1109\/CVPR.2015.7299005"},{"key":"1270_CR7","doi-asserted-by":"crossref","unstructured":"Helten, T., Muller, M., Seidel, H. P., & Theobalt, C. (2013). Real-time body tracking with one depth camera and inertial sensors. In Proceedings of the IEEE international conference on computer vision (ICCV) (pp. 1105\u20131112).","DOI":"10.1109\/ICCV.2013.141"},{"key":"1270_CR8","doi-asserted-by":"crossref","unstructured":"Hochreiter, S., & Schmidhuber, J. (1997). Long short-term memory. In Neural computation (Vol. 9, pp. 1735\u20131780). MIT Press.","DOI":"10.1162\/neco.1997.9.8.1735"},{"key":"1270_CR9","unstructured":"Huang, Y., Kaufmann, M., Aksan, E., Black, M. J., Hilliges, O., & Pons-Moll, G. (2018). Deep inertial poser: Learning to reconstruct human pose from sparse inertial measurements in real time. ACM Transactions on Graphics, (Proc SIGGRAPH Asia), 37, 185:1\u2013185:15, two first authors contributed equally."},{"key":"1270_CR10","doi-asserted-by":"publisher","first-page":"539","DOI":"10.1016\/j.robot.2015.09.029","volume":"75","author":"AE Ichim","year":"2016","unstructured":"Ichim, A. E., & Tombari, F. (2016). Semantic parametric body shape estimation from noisy depth sequences. Robotics and Autonomous Systems, 75, 539\u2013549. https:\/\/doi.org\/10.1016\/j.robot.2015.09.029.","journal-title":"Robotics and Autonomous Systems"},{"key":"1270_CR11","unstructured":"IKinema. (2017). IKinema Orion. Retrieved July 20, 2017 from https:\/\/ikinema.com\/orion."},{"issue":"7","key":"1270_CR12","doi-asserted-by":"publisher","first-page":"1325","DOI":"10.1109\/TPAMI.2013.248","volume":"36","author":"C Ionescu","year":"2014","unstructured":"Ionescu, C., Papava, D., Olaru, V., & Sminchisescu, C. (2014). Human3.6M: Large scale datasets and predictive methods for 3D human sensing in natural environments. IEEE Transactions on Pattern Analysis and Machine Intelligence, 36(7), 1325\u20131339.","journal-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence"},{"key":"1270_CR13","doi-asserted-by":"crossref","unstructured":"Joo, H., Simon, T., & Sheikh, Y. (2018). Total capture: A 3d deformation model for tracking faces, hands, and bodies. In Conference on computer vision and pattern recognition (CVPR).","DOI":"10.1109\/CVPR.2018.00868"},{"key":"1270_CR14","doi-asserted-by":"crossref","unstructured":"Li, S., Zhang, W., & Chan, A. B. (2017). Maximum-margin structured learning with deep networks for 3D human pose estimation. In International conference on computer vision (ICCV).","DOI":"10.1007\/s11263-016-0962-x"},{"key":"1270_CR15","doi-asserted-by":"crossref","unstructured":"Lin, M., Lin, L., Liang, X., Wang, K., & Cheng, H. (2017). Recurrent 3D pose sequence machines. In Conference on computer vision and pattern recognition (CVPR).","DOI":"10.1109\/CVPR.2017.588"},{"key":"1270_CR16","unstructured":"Loper, M., Mahmood, N., Romero, J., Pons-Moll, G., & Black, M. J. (2015). SMPL: A skinned multi-person linear model. ACM Transactions on Graphics (Proc SIGGRAPH Asia), 34(6), 248:1\u2013248:16."},{"key":"1270_CR17","doi-asserted-by":"crossref","unstructured":"Malleson, C., Volino, M., Gilbert, A., Trumble, M., Collomosse, J., Hilton, A. (2017). Real-time full-body motion capture from video and imus. In 2017 fifth international conference on 3D vision (3DV).","DOI":"10.1109\/3DV.2017.00058"},{"key":"1270_CR18","doi-asserted-by":"crossref","unstructured":"Martinez, J., Hossain, R., Romero, J., & Little, J. J. (2017). A simple yet effective baseline for 3d human pose estimation. In 2017 IEEE international conference on computer vision (ICCV) (pp. 2659\u20132668).","DOI":"10.1109\/ICCV.2017.288"},{"key":"1270_CR19","doi-asserted-by":"crossref","unstructured":"Mehta, D., Sotnychenko, O., Mueller, F., Xu, W., Sridhar, S., Pons-Moll, G., et al. (2018). Single-shot multi-person 3d pose estimation from monocular rgb. In International conference on 3D vision (3DV).","DOI":"10.1109\/3DV.2018.00024"},{"issue":"1145\/3072959","key":"1270_CR20","first-page":"3073596","volume":"10","author":"D Mehta","year":"2017","unstructured":"Mehta, D., Sridhar, S., Sotnychenko, O., Rhodin, H., Shafiei, M., Seidel, H. P., et al. (2017). VNect: Real-time 3D human pose estimation with a single RGB camera. ACM Transactions on Graphics. doi, 10(1145\/3072959), 3073596.","journal-title":"ACM Transactions on Graphics. doi"},{"key":"1270_CR21","unstructured":"OptiTrack. (2017). OptiTrack motive. Retrieved July 20, 2017 from http:\/\/www.optitrack.com."},{"key":"1270_CR22","unstructured":"PerceptionNeuron. (2017). Perception neuron. Retrieved July 20, 2017 from http:\/\/www.neuronmocap.com."},{"issue":"6","key":"1270_CR23","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1145\/2980179.2980235","volume":"35","author":"Helge Rhodin","year":"2016","unstructured":"Rhodin, H., Richardt, C., Casas, D., Insafutdinov, E., Shafiei, M., Seidel, H. P., et al. (2016a). EgoCap: Egocentric marker-less motion capture with two fisheye cameras. ACM Transaction on Graphics (TOG), 35(6), 162:1\u2013162:11.","journal-title":"ACM Transactions on Graphics"},{"key":"1270_CR24","series-title":"Lecture Notes in Computer Science","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-46448-0","volume-title":"Computer Vision \u2013 ECCV 2016","year":"2016","unstructured":"Rhodin, H., Robertini, N., Casas, D., Richardt, C., Seidel, H. P., & Theobalt, C. (2016b). General automatic human shape and motion capture using volumetric contour cues. In European conference on computer vision (ECCV) (pp. 509\u2013526). https:\/\/doi.org\/10.1007\/978-3-319-46448-0."},{"key":"1270_CR25","unstructured":"Roetenberg, D., Luinge, H., & Slycke, P. (2013). Xsens MVN: Full 6DOF human motion tracking using miniature inertial sensors. Technical report, pp. 1\u20137."},{"key":"1270_CR26","doi-asserted-by":"publisher","unstructured":"Rosenhahn, B., Schmaltz, C., Brox, T., Weickert, J., & Seidel, H. P. (2008). Staying well grounded in markerless motion capture. In: Pattern recognition DAGM (pp. 385\u2013395). https:\/\/doi.org\/10.1007\/978-3-540-69321-5_39.","DOI":"10.1007\/978-3-540-69321-5_39"},{"key":"1270_CR27","unstructured":"Tekin, B., M\u00e1rquez-Neila, P., Salzmann, M., & Fua, P. (2016). Fusing 2D uncertainty and 3D cues for monocular body pose estimation. CoRR, arXiv:1611.05708."},{"key":"1270_CR28","doi-asserted-by":"crossref","unstructured":"Tome, D., Russell, C., Agapito, L. (2017). Lifting from the deep: Convolutional 3D pose estimation from a single image. In Conference on computer vision and pattern recognition (CVPR).","DOI":"10.1109\/CVPR.2017.603"},{"key":"1270_CR29","doi-asserted-by":"publisher","unstructured":"Tome, D., Toso, M., Agapito, L., & Russell, C. (2018). Rethinking pose in 3d: Multi-stage refinement and recovery for markerless motion capture. In 2018 international conference on 3D vision (3DV) (pp. 474\u2013483). https:\/\/doi.org\/10.1109\/3DV.2018.00061.","DOI":"10.1109\/3DV.2018.00061"},{"key":"1270_CR30","doi-asserted-by":"crossref","unstructured":"Trumble, M., Gilbert, A., Hilton, A., & Collomosse, J. (2016). Deep convolutional networks for marker-less human pose estimation from multiple views. In Proceedings of the 13th European conference on visual media production (CVMP 2016).","DOI":"10.1145\/2998559.2998565"},{"issue":"1-3","key":"1270_CR31","doi-asserted-by":"publisher","first-page":"41","DOI":"10.1016\/j.scitotenv.2003.11.003","volume":"324","author":"K BERUBE","year":"2004","unstructured":"Trumble, M., Gilbert, A., Hilton, A., & Collomosse, J. (2018). Deep autoencoder for combined human pose estimation and body model upscaling. In European conference on computer vision (ECCV). https:\/\/doi.org\/10.1016\/j.scitotenv.2003.11.003. arXiv:1807.01511.","journal-title":"Science of The Total Environment"},{"key":"1270_CR32","doi-asserted-by":"crossref","unstructured":"Trumble, M., Gilbert, A., Malleson, C., Hilton, A., & Collomosse, J. (2017). Total capture: 3D human pose estimation fusing video and inertial sensors. In British machine vision conference (BMVC).","DOI":"10.5244\/C.31.14"},{"key":"1270_CR33","unstructured":"Vicon. (2017). Vicon blade. Retrieved July 20, 2017 from http:\/\/www.vicon.com."},{"key":"1270_CR34","doi-asserted-by":"crossref","unstructured":"von Marcard, T., Henschel, R., Black, M., Rosenhahn, B., & Pons-Moll, G. (2018). Recovering accurate 3d human pose in the wild using imus and a moving camera. In European conference on computer vision (ECCV).","DOI":"10.1007\/978-3-030-01249-6_37"},{"issue":"8","key":"1270_CR35","doi-asserted-by":"publisher","first-page":"1533","DOI":"10.1109\/TPAMI.2016.2522398","volume":"38","author":"T Von Marcard","year":"2016","unstructured":"Von Marcard, T., Pons-Moll, G., & Rosenhahn, B. (2016). Human pose estimation from video and IMUs. IEEE Transactions on Pattern Analysis and Machine Intelligence, 38(8), 1533\u20131547. https:\/\/doi.org\/10.1109\/TPAMI.2016.2522398.","journal-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence"},{"key":"1270_CR36","doi-asserted-by":"crossref","unstructured":"von Marcard, T., Rosenhahn, B., Black, M., & Pons-Moll, G. (2017). Sparse inertial poser: Automatic 3D human pose estimation from sparse IMUs. In Eurographics 2017 (Vol. 36).","DOI":"10.1111\/cgf.13131"},{"key":"1270_CR37","doi-asserted-by":"publisher","unstructured":"Wei, S. E., Ramakrishna, V., Kanade, T., & Sheikh, Y. (2016). Convolutional pose machines. In IEEE conference on computer vision and pattern recognition (pp. 4724\u20134732). https:\/\/doi.org\/10.1109\/CVPR.2016.511, arXiv:1602.00134.","DOI":"10.1109\/CVPR.2016.511"},{"issue":"6","key":"1270_CR38","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1145\/2366145.2366207","volume":"31","author":"X Wei","year":"2012","unstructured":"Wei, X., Zhang, P., & Chai, J. (2012). Accurate realtime full-body motion capture using a single depth camera. ACM Transactions on Graphics, 31(6), 1. https:\/\/doi.org\/10.1145\/2366145.2366207.","journal-title":"ACM Transactions on Graphics"},{"key":"1270_CR39","doi-asserted-by":"publisher","unstructured":"Zanfir, A., Marinoiu, E., & Sminchisescu, C. (2018). Monocular 3D pose and shape estimation of multiple people in natural scenes: The importance of multiple scene constraints. In Conference on computer vision and pattern recognition (CVPR) (pp. 2148\u20132157). https:\/\/doi.org\/10.1109\/CVPR.2018.00229.","DOI":"10.1109\/CVPR.2018.00229"},{"key":"1270_CR40","doi-asserted-by":"crossref","unstructured":"Zhang, Z. (1999). Flexible camera calibration by viewing a plane from unknown orientations. In International conference on computer vision (ICCV) (Vol. 1, pp. 666\u2013673). https:\/\/doi.org\/10.1109\/ICCV.1999.791289.","DOI":"10.1109\/ICCV.1999.791289"},{"key":"1270_CR41","doi-asserted-by":"publisher","unstructured":"Zhao, M., Li, T., Alsheikh, M. A., Tian, Y., Zhao, H., Torralba, A., et al. (2018). Through-wall human pose estimation using radio signals. In Conference on computer vision and pattern recognition (CVPR) (pp. 7356\u20137365). https:\/\/doi.org\/10.1109\/CVPR.2018.00768, arXiv:1011.1669v3.","DOI":"10.1109\/CVPR.2018.00768"},{"key":"1270_CR42","doi-asserted-by":"crossref","unstructured":"Zhou, X., Zhu, M., Leonardos, S., Derpanis, K. G., & Daniilidis, K. (2016). Sparseness meets deepness: 3D human pose estimation from monocular video. In Conference on computer vision and pattern recognition (CVPR) (pp. 4966\u20134975).","DOI":"10.1109\/CVPR.2016.537"}],"container-title":["International Journal of Computer Vision"],"original-title":[],"language":"en","link":[{"URL":"http:\/\/link.springer.com\/content\/pdf\/10.1007\/s11263-019-01270-5.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"http:\/\/link.springer.com\/article\/10.1007\/s11263-019-01270-5\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"http:\/\/link.springer.com\/content\/pdf\/10.1007\/s11263-019-01270-5.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2020,12,16]],"date-time":"2020-12-16T00:12:36Z","timestamp":1608077556000},"score":1,"resource":{"primary":{"URL":"http:\/\/link.springer.com\/10.1007\/s11263-019-01270-5"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2019,12,17]]},"references-count":42,"journal-issue":{"issue":"6","published-print":{"date-parts":[[2020,6]]}},"alternative-id":["1270"],"URL":"https:\/\/doi.org\/10.1007\/s11263-019-01270-5","relation":{},"ISSN":["0920-5691","1573-1405"],"issn-type":[{"value":"0920-5691","type":"print"},{"value":"1573-1405","type":"electronic"}],"subject":[],"published":{"date-parts":[[2019,12,17]]},"assertion":[{"value":"4 January 2019","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"25 November 2019","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"17 December 2019","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}}]}}