{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,30]],"date-time":"2026-07-30T14:04:22Z","timestamp":1785420262130,"version":"3.56.0"},"reference-count":90,"publisher":"MDPI AG","issue":"3","license":[{"start":{"date-parts":[[2014,3,3]],"date-time":"2014-03-03T00:00:00Z","timestamp":1393804800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/3.0\/"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Sensors"],"abstract":"<jats:p>Human Pose Recovery has been studied in the field of Computer Vision for the last 40 years. Several approaches have been reported, and significant improvements have been obtained in both data representation and model design. However, the problem of Human Pose Recovery in uncontrolled environments is far from being solved. In this paper, we define a general taxonomy to group model based approaches for Human Pose Recovery, which is composed of five main modules: appearance, viewpoint, spatial relations, temporal consistence, and behavior. Subsequently, a methodological comparison is performed following the proposed taxonomy, evaluating current SoA approaches in the aforementioned five group categories. As a result of this comparison, we discuss the main advantages and drawbacks of the reviewed literature.<\/jats:p>","DOI":"10.3390\/s140304189","type":"journal-article","created":{"date-parts":[[2014,3,3]],"date-time":"2014-03-03T11:27:10Z","timestamp":1393846030000},"page":"4189-4210","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":57,"title":["A Survey on Model Based Approaches for 2D and 3D Visual Human Pose Recovery"],"prefix":"10.3390","volume":"14","author":[{"given":"Xavier","family":"Perez-Sala","sequence":"first","affiliation":[{"name":"Fundaci\u00f3 Privada Sant Antoni Abat, Vilanova i la Geltr\u00fa, Universitat Polit\u00e8cnica de Catalunya, Vilanova i la Geltr\u00fa 08800, Catalonia, Spain"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Sergio","family":"Escalera","sequence":"additional","affiliation":[{"name":"Department Mathematics (MAIA), Universitat de Barcelona and Computer Vision Center (CVC), Barcelona 08007, Catalonia, Spain"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Cecilio","family":"Angulo","sequence":"additional","affiliation":[{"name":"Automatic Control Department (ESAII), Universitat Polit\u00e8cnica de Catalunya, Vilanova i la Geltr\u00fa 08800, Catalonia, Spain"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Jordi","family":"Gonz\u00e0lez","sequence":"additional","affiliation":[{"name":"Department Computer Science, Universitat Aut\u00f2noma de Barcelona and Computer Vision Center (CVC), Bellaterra 08193, Catalonia, Spain"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1968","published-online":{"date-parts":[[2014,3,3]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","first-page":"90","DOI":"10.1016\/j.cviu.2006.08.002","article-title":"A survey of advances in vision-based human motion capture and analysis","volume":"104","author":"Moeslund","year":"2006","journal-title":"Comput. Vis. Image Underst."},{"key":"ref_2","first-page":"501","article-title":"Representation and recognition of the movements of shapes","volume":"214","author":"Marr","year":"1982","journal-title":"Proc. R. Soc. Lond. Ser. B. Biol. Sci."},{"key":"ref_3","unstructured":"Eichner, M., Marin-Jimenez, M., Zisserman, A., and Ferrari, V. (2010). Articulated Human Pose Estimation and Search in (Almost) Unconstrained Still Images, ETH Zurich. Technical Report No. 272."},{"key":"ref_4","doi-asserted-by":"crossref","unstructured":"Gowsikhaa, D., Abirami, S., and Baskaran, R. (2012). Automated human behavior analysis from surveillance videos: A survey. Artif. Intell. Rev.","DOI":"10.1007\/s10462-012-9341-3"},{"key":"ref_5","first-page":"743","article-title":"Pedestrian detection: An evaluation of the state of the art","volume":"34","author":"Wojek","year":"2011","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_6","doi-asserted-by":"crossref","unstructured":"Singh, V., and Nevatia, R. (2011, January 6\u201313). Action recognition in cluttered dynamic scenes using Pose-Specific Part Models. Barcelona, Brazil.","DOI":"10.1109\/ICCV.2011.6126232"},{"key":"ref_7","unstructured":"Seemann, E., Nickel, K., and Stiefelhagen, R. (2004, January 17\u201319). Head pose estimation using stereo vision for human-robot interaction. Seoul, Korea."},{"key":"ref_8","doi-asserted-by":"crossref","first-page":"1875","DOI":"10.1016\/j.imavis.2005.12.020","article-title":"Visual recognition of pointing gestures for human-robot interaction","volume":"25","author":"Nickel","year":"2007","journal-title":"Image Vis. Comput."},{"key":"ref_9","unstructured":"Escalera, S. (2012). Articulated Motion and Deformable Objects, Springer."},{"key":"ref_10","doi-asserted-by":"crossref","unstructured":"Andriluka, M., Roth, S., and Schiele, B. (2010, January 13\u201318). Monocular 3D pose estimation and tracking by detection. San Francisco, CA, USA.","DOI":"10.1109\/CVPR.2010.5540156"},{"key":"ref_11","doi-asserted-by":"crossref","first-page":"44","DOI":"10.1109\/TPAMI.2006.21","article-title":"Recovering 3D human pose from monocular images","volume":"28","author":"Agarwal","year":"2006","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_12","doi-asserted-by":"crossref","first-page":"2926","DOI":"10.1016\/j.patcog.2008.02.012","article-title":"A spatio-temporal 2D-models framework for human pose recovery in monocular sequences","volume":"41","author":"Rogez","year":"2008","journal-title":"Pattern Recognit."},{"key":"ref_13","doi-asserted-by":"crossref","first-page":"2179","DOI":"10.1109\/TPAMI.2008.260","article-title":"Monocular pedestrian detection: Survey and experiments","volume":"31","author":"Enzweiler","year":"2009","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_14","first-page":"547","article-title":"Computer vision approaches to pedestrian detection: Visible spectrum survey","volume":"4477","author":"Sappa","year":"2007","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_15","unstructured":"Ramanan, D. (2011). Visual Analysis of Humans, Springer."},{"key":"ref_16","doi-asserted-by":"crossref","first-page":"4","DOI":"10.1016\/j.cviu.2006.10.016","article-title":"Vision-based human motion analysis: An overview","volume":"108","author":"Poppe","year":"2007","journal-title":"Comput. Vis. Image Underst."},{"key":"ref_17","unstructured":"Perez-Sala, X., Escalera, S., and Angulo, C. (2012, January 24\u201326). Survey on spatio-temporal view invariant human pose recovery. Catalonia, Spain."},{"key":"ref_18","doi-asserted-by":"crossref","first-page":"82","DOI":"10.1006\/cviu.1998.0716","article-title":"The visual analysis of human movement: A survey","volume":"73","author":"Gavrila","year":"1999","journal-title":"Comput. Vis. Image Underst."},{"key":"ref_19","first-page":"119","article-title":"Real-time human pose recognition in rarts from single depth images","volume":"411","author":"Shotton","year":"2011","journal-title":"Mach. Learn. Comput. Vis. Stud. Comput. Intell."},{"key":"ref_20","doi-asserted-by":"crossref","unstructured":"Hern\u00e1ndez, A., Reyes, M., Escalera, S., and Radeva, P. (2010, January 13\u201318). Spatio-Temporal GrabCut human segmentation for face and pose recovery. San Francisco, CA, USA.","DOI":"10.1109\/CVPRW.2010.5543824"},{"key":"ref_21","doi-asserted-by":"crossref","unstructured":"Hern\u00e1ndez-Vela, A., Zlateva, N., Marinov, A., Reyes, M., Radeva, P., Dimov, D., and Escalera, S. (2012, January 16\u201321). Graph cuts optimization for multi-limb human segmentation in depth maps. Providence, RI, USA.","DOI":"10.1109\/CVPR.2012.6247742"},{"key":"ref_22","unstructured":"Ramanan, D. (2006, January 4\u20137). Learning to parse images of articulated bodies. Vancouver, BC Canada."},{"key":"ref_23","doi-asserted-by":"crossref","unstructured":"Andriluka, M., Roth, S., and Schiele, B. (2009, January 20\u201325). Pictorial structures revisited: People detection and articulated pose estimation. Miami, FL, USA.","DOI":"10.1109\/CVPR.2009.5206754"},{"key":"ref_24","doi-asserted-by":"crossref","unstructured":"Wang, Y., Tran, D., and Liao, Z. (2011, January 20\u201325). Learning hierarchical poselets for human parsing. Providence, RI, USA.","DOI":"10.1109\/CVPR.2011.5995519"},{"key":"ref_25","doi-asserted-by":"crossref","unstructured":"Pirsiavash, H., and Ramanan, D. (2012, January 16\u201321). Steerable part models. Providence, RI, USA.","DOI":"10.1109\/CVPR.2012.6248058"},{"key":"ref_26","unstructured":"Dalal, N., and Triggs, B. (2005, January 20\u201326). Histograms of oriented gradients for human detection. San Diego, CA, USA."},{"key":"ref_27","unstructured":"Bourdev, L.D., and Malik, J. (October,, January 27). Poselets: Body part detectors trained using 3D human pose annotations. Kyoto, Japan."},{"key":"ref_28","unstructured":"Mittal, A., Zhao, L., and Davis, L. (2003, January 21\u201322). Human body pose estimation using silhouette shape analysis. Miami, FL, USA."},{"key":"ref_29","doi-asserted-by":"crossref","first-page":"1582","DOI":"10.1109\/TPAMI.2009.154","article-title":"Evaluating color descriptors for object and scene recognition","volume":"32","author":"Sande","year":"2010","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_30","doi-asserted-by":"crossref","unstructured":"Navarathna, R., Sridharan, S., and Lucey, S. (2011, January 6\u201313). Fourier active appearance models. Barcelona, Spain.","DOI":"10.1109\/ICCV.2011.6126461"},{"key":"ref_31","doi-asserted-by":"crossref","first-page":"1160","DOI":"10.1364\/JOSAA.2.001160","article-title":"others. Uncertainty relation for resolution in space, spatial frequency, and orientation optimized by two-dimensional visual cortical filters","volume":"2","author":"Daugman","year":"1985","journal-title":"J. Opt. Soc. Am. A"},{"key":"ref_32","doi-asserted-by":"crossref","unstructured":"Pugeault, N., and Bowden, R. (2011, January 6\u201313). Spelling it out: Real-time ASL fingerspelling recognition. Barcelona, Spain.","DOI":"10.1109\/ICCVW.2011.6130290"},{"key":"ref_33","doi-asserted-by":"crossref","unstructured":"Plagemann, C., Ganapathi, V., Koller, D., and Thrun, S. (2011, January 6\u201313). Real-time identification and localization of body parts from depth images. Barcelona, Spain.","DOI":"10.1109\/ROBOT.2010.5509559"},{"key":"ref_34","doi-asserted-by":"crossref","first-page":"43","DOI":"10.1007\/BF01420984","article-title":"Performance of optical flow techniques","volume":"12","author":"Barron","year":"1994","journal-title":"Int. J. Comput. Vis."},{"key":"ref_35","doi-asserted-by":"crossref","unstructured":"Laptev, I., Marszalek, M., Schmid, C., and Rozenfeld, B. (2008, January 24\u201326). Learning realistic human actions from movies. Anchorage, AK, USA.","DOI":"10.1109\/CVPR.2008.4587756"},{"key":"ref_36","doi-asserted-by":"crossref","first-page":"396","DOI":"10.1016\/j.cviu.2011.09.010","article-title":"Selective spatio-temporal interest points","volume":"116","author":"Chakraborty","year":"2012","journal-title":"Comput. Vis. Image Underst."},{"key":"ref_37","doi-asserted-by":"crossref","first-page":"107","DOI":"10.1007\/s11263-005-1838-7","article-title":"On space-time interest points","volume":"64","author":"Laptev","year":"2005","journal-title":"Int. J. Comput. Vis."},{"key":"ref_38","doi-asserted-by":"crossref","unstructured":"Yao, B., and Li, F.-F. (2010, January 13\u201318). Grouplet: A structured image representation for recognizing human and object interactions. San Francisco, CA, USA.","DOI":"10.1109\/CVPR.2010.5540234"},{"key":"ref_39","doi-asserted-by":"crossref","first-page":"1627","DOI":"10.1109\/TPAMI.2009.167","article-title":"Object detection with discriminatively trained part-based models","volume":"32","author":"Felzenszwalb","year":"2010","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_40","doi-asserted-by":"crossref","first-page":"309","DOI":"10.1145\/1015706.1015720","article-title":"GrabCut: Interactive Foreground Extraction Using Iterated Graph Cuts","volume":"23","author":"Rother","year":"2004","journal-title":"ACM Trans. Graph."},{"key":"ref_41","doi-asserted-by":"crossref","first-page":"43","DOI":"10.1007\/s11263-005-3848-x","article-title":"A comparison of affine region detectors","volume":"65","author":"Mikolajczyk","year":"2005","journal-title":"Int. J. Comput. Vis."},{"key":"ref_42","doi-asserted-by":"crossref","unstructured":"Karaulova, I., Hall, P., and Marshall, A. (2000, January 11\u201314). A hierarchical model of dynamics for tracking people with a single video camera. Bristol UK.","DOI":"10.5244\/C.14.36"},{"key":"ref_43","doi-asserted-by":"crossref","unstructured":"Savarese, S., and Li, F.-F. (2007, January 14\u201320). 3D generic object categorization, localization and pose estimation. Rio de Janeiro, Brazil.","DOI":"10.1109\/ICCV.2007.4408987"},{"key":"ref_44","unstructured":"Sun, M., Su, H., Savarese, S., and Li, F.-F. (2009, January 20\u201325). A multi-view probabilistic model for 3D object classes. Miami, FL, USA."},{"key":"ref_45","unstructured":"Su, H., Sun, M., Li, F.-F., and Savarese, S. (October, January 27). Learning a dense multi-view representation for detection, viewpoint classification and synthesis of object categories. Kyoto, Japan."},{"key":"ref_46","doi-asserted-by":"crossref","unstructured":"Moreno-Noguer, F., Lepetit, V., and Fua, P. (2008, January 12\u201318). Pose priors for simultaneously solving alignment and correspondence. Marseille, France.","DOI":"10.1007\/978-3-540-88688-4_30"},{"key":"ref_47","doi-asserted-by":"crossref","unstructured":"Salzmann, M., Moreno-Noguer, F., Lepetit, V., and Fua, P. (2008, January 12\u201318). Closed-form solution to non-rigid 3D surface registration. Marseille, France.","DOI":"10.1007\/978-3-540-88693-8_43"},{"key":"ref_48","doi-asserted-by":"crossref","unstructured":"Simo-Serra, E., Ramisa, A., Alenya, G., Torras, C., and Moreno-Noguer, F. (2012, January 16\u201321). Single Image 3D Human Pose Estimation from Noisy Observations. Providence, RI, USA.","DOI":"10.1109\/CVPR.2012.6247988"},{"key":"ref_49","doi-asserted-by":"crossref","unstructured":"S\u00e1nchez-Riera, J., Ostlund, J., Fua, P., and Moreno-Noguer, F. (2010, January 13\u201318). Simultaneous pose, correspondence and non-rigid shape. San Francisco, CA, USA.","DOI":"10.1109\/CVPR.2010.5539831"},{"key":"ref_50","doi-asserted-by":"crossref","first-page":"190","DOI":"10.1007\/s11263-012-0524-9","article-title":"2d articulated human pose estimation and retrieval in (almost) unconstrained still images","volume":"99","author":"Eichner","year":"2012","journal-title":"Int. J. Comput. Vis."},{"key":"ref_51","doi-asserted-by":"crossref","unstructured":"Sapp, B., Weiss, D., and Taskar, B. (2011, January 20\u201325). Parsing human motion with stretchable models. Colorado Springs, CO, USA.","DOI":"10.1109\/CVPR.2011.5995607"},{"key":"ref_52","unstructured":"Ferrari, V., Eichner, M., Marin-Jimenez, M., and Zisserman, A. Buffy Stickmen Dataset. Available online: http:\/\/www.robots.ox.ac.uk\/\u223cvgg\/data\/stickmen\/."},{"key":"ref_53","doi-asserted-by":"crossref","first-page":"67","DOI":"10.1109\/T-C.1973.223602","article-title":"The representation and matching of pictorial structures","volume":"100","author":"Fischler","year":"1973","journal-title":"Comput. Trans."},{"key":"ref_54","doi-asserted-by":"crossref","first-page":"55","DOI":"10.1023\/B:VISI.0000042934.15159.49","article-title":"Pictorial structures for object recognition","volume":"61","author":"Felzenszwalb","year":"2005","journal-title":"Int. J. Comput. Vis."},{"key":"ref_55","unstructured":"Sigal, L., Bhatia, S., Roth, S., Black, M., and Isard, M. (July, January 27). Tracking loose-limbed people. Washington, DC, USA."},{"key":"ref_56","doi-asserted-by":"crossref","unstructured":"Yang, Y., and Ramanan, D. (2011, January 20\u201325). Articulated pose estimation with flexible mixtures-of-parts. Colorado Springs, CO, USA.","DOI":"10.1109\/CVPR.2011.5995741"},{"key":"ref_57","unstructured":"Sminchisescu, C., and Triggs, B. (2003, January 16\u201322). Kinematic jump processes for monocular 3D human tracking. Madison, WI, USA."},{"key":"ref_58","unstructured":"Felzenszwalb, P., and McAllester, D. (2010). Object Detection Grammars, Computer Science TR; University of Chicago. Technical Report."},{"key":"ref_59","first-page":"6","article-title":"Object detection with grammar models","volume":"33","author":"Girshick","year":"2011","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_60","doi-asserted-by":"crossref","first-page":"355","DOI":"10.1109\/TITS.2013.2281207","article-title":"Toward real-time pedestrian detection based on a deformable template model","volume":"15","author":"Pedersoli","year":"2013","journal-title":"Trans. Intell. Transp. Syst."},{"key":"ref_61","doi-asserted-by":"crossref","first-page":"15","DOI":"10.1007\/s11263-011-0493-4","article-title":"Loose-limbed people: Estimating 3d human pose and motion using non-parametric belief propagation","volume":"98","author":"Sigal","year":"2012","journal-title":"Int. J. Comput. Vis."},{"key":"ref_62","unstructured":"Zhu, L., Chen, Y., Lu, Y., Lin, C., and Yuille, A. (2008, January 24\u201326). Max margin and\/or graph learning for parsing the human body. Anchorage, AK, USA."},{"key":"ref_63","first-page":"289","article-title":"Rapid inference on a novel and\/or graph for object detection, segmentation and parsing","volume":"20","author":"Chen","year":"2007","journal-title":"NIPS"},{"key":"ref_64","unstructured":"Lan, X., and Huttenlocher, D. (2005, January 17\u201320). Beyond trees: Common-factor models for 2d human pose recovery. Beijing, China."},{"key":"ref_65","first-page":"314","article-title":"Efficient inference with multiple heterogeneous part detectors for human pose estimation","volume":"6313","author":"Singh","year":"2010","journal-title":"ECCV"},{"key":"ref_66","unstructured":"Agarwal, A., and Triggs, B. (2004, January 11\u201314). Tracking articulated motion with piecewise learned dynamical models. Prague, Czech Republic."},{"key":"ref_67","unstructured":"Wei, X., and Chai, J. (October, January 27). Modeling 3d human poses from uncalibrated monocular images. Kyoto, Japan."},{"key":"ref_68","doi-asserted-by":"crossref","unstructured":"Valmadre, J., and Lucey, S. (2010, January 5\u201311). Deterministic 3D human pose estimation using rigid structure. Heraklion, Crete, Greece.","DOI":"10.1007\/978-3-642-15558-1_34"},{"key":"ref_69","doi-asserted-by":"crossref","first-page":"179","DOI":"10.1023\/B:VISI.0000011203.00237.9b","article-title":"Twist based acquisition and tracking of animal and human kinematics","volume":"56","author":"Bregler","year":"2004","journal-title":"Int. J. Comput. Vis."},{"key":"ref_70","unstructured":"Howe, N., Leventon, M., and Freeman, W. (1999). Bayesian Reconstruction of 3D Human Motion from Single-Camera Video, NIPS."},{"key":"ref_71","doi-asserted-by":"crossref","unstructured":"Gall, J., Stoll, C., de Aguiar, E., Theobalt, C., Rosenhahn, B., and Seidel, H. (2009, January 20\u201325). Motion capture using joint skeleton tracking and surface estimation. Miami, FL, USA.","DOI":"10.1109\/CVPRW.2009.5206755"},{"key":"ref_72","doi-asserted-by":"crossref","first-page":"2907","DOI":"10.1016\/j.patcog.2009.02.012","article-title":"Action-specific motion prior for efficient bayesian 3D human body tracking","volume":"42","author":"Rius","year":"2009","journal-title":"Pattern Recogn."},{"key":"ref_73","doi-asserted-by":"crossref","first-page":"38","DOI":"10.1006\/cviu.1995.1004","article-title":"others. Active shape models-their training and application","volume":"61","author":"Cootes","year":"1995","journal-title":"Comput. Vis. Image Underst."},{"key":"ref_74","doi-asserted-by":"crossref","first-page":"681","DOI":"10.1109\/34.927467","article-title":"Active appearance models","volume":"23","author":"Cootes","year":"2001","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_75","doi-asserted-by":"crossref","first-page":"607","DOI":"10.1109\/TPAMI.2008.106","article-title":"Head Pose Estimation in Computer Vision: A Survey","volume":"31","author":"Trivedi","year":"2009","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_76","doi-asserted-by":"crossref","first-page":"25","DOI":"10.1049\/iet-cvi.2009.0009","article-title":"Gait recognition using active shape model and motion prediction","volume":"4","author":"Kim","year":"2010","journal-title":"Comput. Vis. IET"},{"key":"ref_77","doi-asserted-by":"crossref","first-page":"157","DOI":"10.1016\/j.cviu.2006.08.006","article-title":"Temporal motion models for monocular and multiview 3D human body tracking","volume":"104","author":"Urtasun","year":"2006","journal-title":"Comput. Vis. Image Underst."},{"key":"ref_78","doi-asserted-by":"crossref","first-page":"1442","DOI":"10.1109\/TPAMI.2010.201","article-title":"Trajectory space: A dual representation for nonrigid structure from motion","volume":"33","author":"Akhter","year":"2011","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_79","doi-asserted-by":"crossref","unstructured":"Moreno-Noguer, F., and Porta, J. (2011, January 20\u201325). Probabilistic simultaneous pose and non-rigid shape recovery. Colorado Springs, CO, USA.","DOI":"10.1109\/CVPR.2011.5995532"},{"key":"ref_80","doi-asserted-by":"crossref","unstructured":"Urtasun, R., and Fua, P. (2004, January 11\u201314). 3D human body tracking using deterministic temporal motion models. Prague, Czech Republic.","DOI":"10.1007\/978-3-540-24672-5_8"},{"key":"ref_81","unstructured":"Urtasun, R., Fleet, D., and Fua, P. (2005, January 20\u201326). Monocular 3D tracking of the golf swing. San Diego, CA, USA."},{"key":"ref_82","doi-asserted-by":"crossref","unstructured":"Urtasun, R., Fleet, D., Hertzmann, A., and Fua, P. (2005, January 17\u201320). Priors for people tracking from small training sets. Beijing, China.","DOI":"10.1109\/ICCV.2005.193"},{"key":"ref_83","doi-asserted-by":"crossref","unstructured":"Fossati, A., Salzmann, M., and Fua, P. (2009, January 20\u201325). Observable subspaces for 3D human motion recovery. Miami, FL, USA.","DOI":"10.1109\/CVPR.2009.5206489"},{"key":"ref_84","unstructured":"Akhter, I., Sheikh, Y., Khan, S., and Kanade, T. (2008, January 8\u201311). Nonrigid structure from motion in trajectory space. Vancouver, BC, Canada."},{"key":"ref_85","doi-asserted-by":"crossref","unstructured":"Park, H., Shiratori, T., Matthews, I., and Sheikh, Y. (2010, January 5\u201311). 3D Reconstruction of a Moving Point from a Series of 2D Projections. Heraklion, Crete, Greece.","DOI":"10.1007\/978-3-642-15558-1_12"},{"key":"ref_86","doi-asserted-by":"crossref","unstructured":"Park, H., and Sheikh, Y. (2011, January 6\u201313). 3D reconstruction of a smooth articulated trajectory from a monocular image sequence. Barcelona, Spain.","DOI":"10.1109\/ICCV.2011.6126243"},{"key":"ref_87","doi-asserted-by":"crossref","unstructured":"Shapovalova, N., Fern\u00e1ndez, C., Roca, F., and Gonz\u00e0lez, J. (2011). Semantics of Human Behavior in Image Sequences. Computer Analysis of Human Behavior, Springer.","DOI":"10.1007\/978-0-85729-994-9_7"},{"key":"ref_88","unstructured":"Sigal, L., and Black, M. (2006). Humaneva: Synchronized Video and Motion Capture Dataset for Evaluation of Articulated Human Motion, Brown Univertsity. Technical Report."},{"key":"ref_89","doi-asserted-by":"crossref","unstructured":"Yao, B., and Fei-Fei, L. (2010, January 13\u201318). Modeling mutual context of object and human pose in human-object interaction activities. San Francisco, CA, USA.","DOI":"10.1109\/CVPR.2010.5540235"},{"key":"ref_90","doi-asserted-by":"crossref","first-page":"260","DOI":"10.1007\/978-3-642-31567-1_26","article-title":"Human Context: Modeling human-human interactions for monocular 3D pose estimation","volume":"7378","author":"Andriluka","year":"2012","journal-title":"Articul. Motion Deform. Objects"}],"container-title":["Sensors"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1424-8220\/14\/3\/4189\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,11]],"date-time":"2025-10-11T21:08:48Z","timestamp":1760216928000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1424-8220\/14\/3\/4189"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2014,3,3]]},"references-count":90,"journal-issue":{"issue":"3","published-online":{"date-parts":[[2014,3]]}},"alternative-id":["s140304189"],"URL":"https:\/\/doi.org\/10.3390\/s140304189","relation":{},"ISSN":["1424-8220"],"issn-type":[{"value":"1424-8220","type":"electronic"}],"subject":[],"published":{"date-parts":[[2014,3,3]]}}}