{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,24]],"date-time":"2026-06-24T16:47:11Z","timestamp":1782319631326,"version":"3.54.5"},"reference-count":45,"publisher":"MDPI AG","issue":"3","license":[{"start":{"date-parts":[[2022,2,25]],"date-time":"2022-02-25T00:00:00Z","timestamp":1645747200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"name":"the Strategic Priority Research Program of the Chinese Academy of Sciences","award":["XDC02070600"],"award-info":[{"award-number":["XDC02070600"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Information"],"abstract":"<jats:p>Existing whole-body human pose estimation methods mostly segment the parts of the body\u2019s hands and feet for specific processing, which not only splits the overall semantics of the body, but also increases the amount of calculation and the complexity of the model. To address these drawbacks, we designed a novel semantic\u2013structural graph convolutional network (SSGCN) for whole-body human pose estimation tasks, which leverages the whole-body graph structure to analyze the semantics of the whole-body keypoints through a graph convolutional network and improves the accuracy of pose estimation. Firstly, we introduced a novel heat-map-based keypoint embedding, which encodes the position information and feature information of the keypoints of the human body. Secondly, we propose a novel semantic\u2013structural graph convolutional network consisting of several sets of cascaded structure-based graph layers and data-dependent whole-body non-local layers. Specifically, the proposed method extracts groups of keypoints and constructs a high-level abstract body graph to process the high-level semantic information of the whole-body keypoints. The experimental results showed that our method achieved very promising results on the challenging COCO whole-body dataset.<\/jats:p>","DOI":"10.3390\/info13030109","type":"journal-article","created":{"date-parts":[[2022,2,25]],"date-time":"2022-02-25T10:00:40Z","timestamp":1645783240000},"page":"109","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":5,"title":["Semantic\u2013Structural Graph Convolutional Networks for Whole-Body Human Pose Estimation"],"prefix":"10.3390","volume":"13","author":[{"ORCID":"https:\/\/orcid.org\/0000-0001-5202-9911","authenticated-orcid":false,"given":"Weiwei","family":"Li","sequence":"first","affiliation":[{"name":"Institute of Microelectronics of the Chinese Academy of Sciences, Beijing 100029, China"},{"name":"University of Chinese Academy of Sciences, Beijing 100864, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Rong","family":"Du","sequence":"additional","affiliation":[{"name":"Institute of Microelectronics of the Chinese Academy of Sciences, Beijing 100029, China"},{"name":"University of Chinese Academy of Sciences, Beijing 100864, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Shudong","family":"Chen","sequence":"additional","affiliation":[{"name":"Institute of Microelectronics of the Chinese Academy of Sciences, Beijing 100029, China"},{"name":"University of Chinese Academy of Sciences, Beijing 100864, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1968","published-online":{"date-parts":[[2022,2,25]]},"reference":[{"key":"ref_1","unstructured":"Cimen, G., Maurhofer, C., Sumner, B., and Guay, M. (2018, January 18\u201320). Ar poser: Automatically augmenting mobile pictures with digital avatars imitating poses. Proceedings of the 12th International Conference on Computer Graphics, Visualization, Computer Vision and Image Processing, Madrid, Spain."},{"key":"ref_2","unstructured":"Elhayek, A., Kovalenko, O., Murthy, P., Malik, J., and Stricker, D. (, January 22\u201323). Fully automatic multi-person human motion capture for vr applications. Proceedings of the International Conference on Virtual Reality and Augmented Reality, London, UK."},{"key":"ref_3","doi-asserted-by":"crossref","first-page":"2093","DOI":"10.1109\/TVCG.2019.2898650","article-title":"Mo2cap2: Real-time mobile 3d motion capture with a cap-mounted fisheye camera","volume":"25","author":"Xu","year":"2019","journal-title":"IEEE Trans. Vis. Comput. Graph."},{"key":"ref_4","doi-asserted-by":"crossref","unstructured":"Choi, H., Moon, G., and Lee, K.M. (2020, January 23\u201328). Pose2Mesh: Graph convolutional network for 3D human pose and mesh recovery from a 2D human pose. Proceedings of the European Conference on Computer Vision, online.","DOI":"10.1007\/978-3-030-58571-6_45"},{"key":"ref_5","doi-asserted-by":"crossref","unstructured":"Kundu, J.N., Rakesh, M., Jampani, V., Venkatesh, R.M., and Babu, R.V. (2020, January 23\u201328). Appearance Consensus Driven Self-supervised Human Mesh Recovery. Proceedings of the European Conference on Computer Vision, online.","DOI":"10.1007\/978-3-030-58452-8_46"},{"key":"ref_6","doi-asserted-by":"crossref","unstructured":"Iqbal, U., Xie, K., Guo, Y., Kautz, J., and Molchanov, P. (2021). KAMA: 3D Keypoint Aware Body Mesh Articulation. arXiv.","DOI":"10.1109\/3DV53792.2021.00078"},{"key":"ref_7","doi-asserted-by":"crossref","unstructured":"Kanazawa, A., Black, M.J., Jacobs, D.W., and Malik, J. (2018, January 18\u201322). End-to-end recovery of human shape and pose. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00744"},{"key":"ref_8","unstructured":"Du, Y., Wang, W., and Wang, L. (2015, January 7\u201312). Hierarchical recurrent neural network for skeleton based action recognition. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Boston, MA, USA."},{"key":"ref_9","doi-asserted-by":"crossref","unstructured":"Li, M., Chen, S., Chen, X., Zhang, Y., Wang, Y., and Tian, Q. (2019, January 15\u201320). Actional-structural graph convolutional networks for skeleton-based action recognition. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.00371"},{"key":"ref_10","doi-asserted-by":"crossref","unstructured":"Yan, A., Wang, Y., Li, Z., and Qiao, Y. (2019, January 15\u201320). PA3D: Pose-action 3D machine for video recognition. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.00811"},{"key":"ref_11","doi-asserted-by":"crossref","first-page":"165","DOI":"10.1016\/j.patcog.2019.03.010","article-title":"Part-aligned pose-guided recurrent network for action recognition","volume":"92","author":"Huang","year":"2019","journal-title":"Pattern Recognit."},{"key":"ref_12","doi-asserted-by":"crossref","unstructured":"Luvizon, D.C., Picard, D., and Tabia, H. (2018, January 18\u201322). 2D\/3d pose estimation and action recognition using multitask deep learning. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00539"},{"key":"ref_13","doi-asserted-by":"crossref","first-page":"231","DOI":"10.1006\/cviu.2000.0897","article-title":"A Survey of Computer Vision-Based Human Motion Capture","volume":"81","author":"Moeslund","year":"2001","journal-title":"Comput. Vis. Image Underst."},{"key":"ref_14","doi-asserted-by":"crossref","unstructured":"Tompson, J., Goroshin, R., Jain, A., LeCun, Y., and Bregler, C. (2015, January 7\u201312). Efficient object localization using Convolutional Networks. Proceedings of the 2015 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Boston, MA, USA.","DOI":"10.1109\/CVPR.2015.7298664"},{"key":"ref_15","first-page":"1799","article-title":"Joint Training of a Convolutional Network and a Graphical Model for Human Pose Estimation","volume":"1","author":"Tompson","year":"2014","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_16","doi-asserted-by":"crossref","unstructured":"Ramakrishna, V., Munoz, D., Hebert, M., Bagnell, J.A., and Sheikh, Y. (2014, January 6\u201312). Pose machines: Articulated pose estimation via inference machines. Proceedings of the European Conference on Computer Vision, Z\u00fcrich, Switzerland.","DOI":"10.1007\/978-3-319-10605-2_3"},{"key":"ref_17","doi-asserted-by":"crossref","unstructured":"Yang, W., Li, S., Ouyang, W., Li, H., and Wang, X. (2017, January 22\u201329). Learning feature pyramids for human pose estimation. Proceedings of the IEEE International Conference on Computer Vision, Venice, Italy.","DOI":"10.1109\/ICCV.2017.144"},{"key":"ref_18","doi-asserted-by":"crossref","unstructured":"Luo, Y., Ren, J., Wang, Z., Sun, W., Pan, J., Liu, J., Pang, J., and Lin, L. (2018, January 18\u201322). Lstm pose machines. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00546"},{"key":"ref_19","doi-asserted-by":"crossref","unstructured":"Artacho, B., and Savakis, A. (2020, January 13\u201319). UniPose: Unified Human Pose Estimation in Single Images and Videos. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA.","DOI":"10.1109\/CVPR42600.2020.00706"},{"key":"ref_20","unstructured":"Athitsos, V., and Sclaroff, S. (2003, January 18\u201320). Estimating 3D hand pose from a cluttered image. Proceedings of the 2003 IEEE Computer Society Conference on Computer Vision and Pattern Recognition, Madison, WI, USA."},{"key":"ref_21","doi-asserted-by":"crossref","first-page":"1793","DOI":"10.1109\/TPAMI.2011.33","article-title":"Model-Based 3D Hand Pose Estimation from Monocular Video","volume":"33","author":"Fleet","year":"2011","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_22","doi-asserted-by":"crossref","first-page":"177","DOI":"10.1007\/s11263-013-0667-3","article-title":"Face alignment by explicit shape regression","volume":"107","author":"Cao","year":"2014","journal-title":"Int. J. Comput. Vis."},{"key":"ref_23","doi-asserted-by":"crossref","unstructured":"Tzimiropoulos, G. (2015, January 7\u201312). Project-out cascaded regression with an application to face alignment. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Boston, MA, USA.","DOI":"10.1109\/CVPR.2015.7298989"},{"key":"ref_24","doi-asserted-by":"crossref","unstructured":"Trigeorgis, G., Snape, P., Nicolaou, M.A., Antonakos, E., and Zafeiriou, S. (2016, January 27\u201330). Mnemonic descent method: A recurrent process applied for end-to-end face alignment. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.453"},{"key":"ref_25","doi-asserted-by":"crossref","first-page":"918","DOI":"10.1109\/TPAMI.2015.2469286","article-title":"Learning deep representation for face alignment with auxiliary attributes","volume":"38","author":"Zhang","year":"2015","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_26","doi-asserted-by":"crossref","unstructured":"Deng, J., Guo, J., Ververas, E., Kotsia, I., and Zafeiriou, S. (2020, January 13\u201319). RetinaFace: Single-Shot Multi-Level Face Localisation in the Wild. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA.","DOI":"10.1109\/CVPR42600.2020.00525"},{"key":"ref_27","doi-asserted-by":"crossref","first-page":"1499","DOI":"10.1109\/LSP.2016.2603342","article-title":"Joint face detection and alignment using multitask cascaded convolutional networks","volume":"23","author":"Zhang","year":"2016","journal-title":"IEEE Signal Process. Lett."},{"key":"ref_28","unstructured":"Oberweger, M., Wohlhart, P., and Lepetit, V. (2015, January 9\u201311). Hands deep in deep learning for hand pose estimation. Proceedings of the 20th Computer Vision Winter Workshop, Seggau, Austria."},{"key":"ref_29","doi-asserted-by":"crossref","unstructured":"Oberweger, M., and Lepetit, V. (2017, January 22\u201329). DeepPrior++: Improving Fast and Accurate 3D Hand Pose Estimation. Proceedings of the IEEE International Conference on Computer Vision Workshops, Venice, Italy.","DOI":"10.1109\/ICCVW.2017.75"},{"key":"ref_30","doi-asserted-by":"crossref","unstructured":"Sharp, T., Keskin, C., Robertson, D., Taylor, J., Shotton, J., Kim, D., Rhemann, C., Leichter, I., Vinnikov, A., and Wei, Y. (2015, January 18\u201323). Accurate, robust, and flexible real-time hand tracking. Proceedings of the 33rd Annual ACM Conference on Human Factors in Computing Systems, Seoul, Korea.","DOI":"10.1145\/2702123.2702179"},{"key":"ref_31","doi-asserted-by":"crossref","unstructured":"Sridhar, S., Mueller, F., Oulasvirta, A., and Theobalt, C. (2015, January 7\u201312). Fast and robust hand tracking using detection-guided optimization. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Boston, MA, USA.","DOI":"10.1109\/CVPR.2015.7298941"},{"key":"ref_32","doi-asserted-by":"crossref","first-page":"172","DOI":"10.1109\/TPAMI.2019.2929257","article-title":"OpenPose: Realtime multi-person 2D pose estimation using Part Affinity Fields","volume":"43","author":"Cao","year":"2021","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_33","doi-asserted-by":"crossref","unstructured":"Jin, S., Xu, L., Xu, J., Wang, C., Liu, W., Qian, C., Ouyang, W., and Luo, P. (2020, January 23\u201328). Whole-body human pose estimation in the wild. Proceedings of the European Conference on Computer Vision, online.","DOI":"10.1007\/978-3-030-58545-7_12"},{"key":"ref_34","unstructured":"Hidalgo, G., Raaj, Y., Idrees, H., Xiang, D., Joo, H., Simon, T., and Sheikh, Y. (2019, January 15\u201320). Single-network whole-body pose estimation. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Long Beach, CA, USA."},{"key":"ref_35","doi-asserted-by":"crossref","unstructured":"Cao, Z., Simon, T., Wei, S.E., and Sheikh, Y. (2017, January 21\u201326). Realtime multi-person 2D pose estimation using part affinity fields. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.143"},{"key":"ref_36","doi-asserted-by":"crossref","unstructured":"Lin, T.Y., Maire, M., Belongie, S., Hays, J., Perona, P., Ramanan, D., Doll\u00e1r, P., and Zitnick, C.L. (2014, January 6\u201312). Microsoft COCO: Common objects in context. Proceedings of the European Conference on Computer Vision, Z\u00fcrich, Switzerland.","DOI":"10.1007\/978-3-319-10602-1_48"},{"key":"ref_37","first-page":"91","article-title":"Faster R-CNN: Towards real-time object detection with region proposal networks","volume":"1","author":"Ren","year":"2015","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_38","doi-asserted-by":"crossref","unstructured":"Sun, K., Xiao, B., Liu, D., and Wang, J. (2019, January 15\u201320). Deep high-resolution representation learning for human pose estimation. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.00584"},{"key":"ref_39","unstructured":"Sun, K., Zhao, Y., Jiang, B., Cheng, T., Xiao, B., Liu, D., Mu, Y., Wang, X., Liu, W., and Wang, J. (2019). High-resolution representations for labeling pixels and regions. arXiv."},{"key":"ref_40","doi-asserted-by":"crossref","unstructured":"Samet, N., and Akbas, E. (2021). HPRNet: Hierarchical Point Regression for Whole-Body Human Pose Estimation. arXiv.","DOI":"10.1016\/j.imavis.2021.104285"},{"key":"ref_41","doi-asserted-by":"crossref","unstructured":"Yan, S., Xiong, Y., and Lin, D. (2018). Spatial Temporal Graph Convolutional Networks for Skeleton-Based Action Recognition. arXiv.","DOI":"10.1609\/aaai.v32i1.12328"},{"key":"ref_42","unstructured":"Buades, A., Coll, B., and Morel, J.M. (2005, January 20\u201326). A non-local algorithm for image denoising. Proceedings of the 2005 IEEE Computer Society Conference on Computer Vision and Pattern Recognition (CVPR\u201905), San Diego, CA, USA."},{"key":"ref_43","doi-asserted-by":"crossref","unstructured":"Wang, X., Girshick, R., Gupta, A., and He, K. (2018, January 18\u201322). Non-local Neural Networks. Proceedings of the 2018 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00813"},{"key":"ref_44","doi-asserted-by":"crossref","unstructured":"He, K., Zhang, X., Ren, S., and Sun, J. (2016, January 27\u201330). Deep Residual Learning for Image Recognition. Proceedings of the 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.90"},{"key":"ref_45","first-page":"2274","article-title":"Associative Embedding: End-to-End Learning for Joint Detection and Grouping","volume":"1","author":"Newell","year":"2017","journal-title":"Adv. Neural Inf. Process. Syst."}],"container-title":["Information"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/2078-2489\/13\/3\/109\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,10]],"date-time":"2025-10-10T22:27:06Z","timestamp":1760135226000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/2078-2489\/13\/3\/109"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,2,25]]},"references-count":45,"journal-issue":{"issue":"3","published-online":{"date-parts":[[2022,3]]}},"alternative-id":["info13030109"],"URL":"https:\/\/doi.org\/10.3390\/info13030109","relation":{},"ISSN":["2078-2489"],"issn-type":[{"value":"2078-2489","type":"electronic"}],"subject":[],"published":{"date-parts":[[2022,2,25]]}}}