{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,10,11]],"date-time":"2025-10-11T02:59:04Z","timestamp":1760151544234,"version":"build-2065373602"},"reference-count":51,"publisher":"MDPI AG","issue":"7","license":[{"start":{"date-parts":[[2022,3,28]],"date-time":"2022-03-28T00:00:00Z","timestamp":1648425600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"name":"the Strategic Priority Research Program of the Chinese Academy of Sciences","award":["XDC02070600"],"award-info":[{"award-number":["XDC02070600"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Sensors"],"abstract":"<jats:p>Despite the great progress in 3D pose estimation from videos, there is still a lack of effective means to extract spatio-temporal features of different granularity from complex dynamic skeleton sequences. To tackle this problem, we propose a novel, skeleton-based spatio-temporal U-Net(STUNet) scheme to deal with spatio-temporal features in multiple scales for 3D human pose estimation in video. The proposed STUNet architecture consists of a cascade structure of semantic graph convolution layers and structural temporal dilated convolution layers, progressively extracting and fusing the spatio-temporal semantic features from fine-grained to coarse-grained. This U-shaped network achieves scale compression and feature squeezing by downscaling and upscaling, while abstracting multi-resolution spatio-temporal dependencies through skip connections. Experiments demonstrate that our model effectively captures comprehensive spatio-temporal features in multiple scales and achieves substantial improvements over mainstream methods on real-world datasets.<\/jats:p>","DOI":"10.3390\/s22072573","type":"journal-article","created":{"date-parts":[[2022,3,29]],"date-time":"2022-03-29T21:45:51Z","timestamp":1648590351000},"page":"2573","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":5,"title":["Skeleton-Based Spatio-Temporal U-Network for 3D Human Pose Estimation in Video"],"prefix":"10.3390","volume":"22","author":[{"ORCID":"https:\/\/orcid.org\/0000-0001-5202-9911","authenticated-orcid":false,"given":"Weiwei","family":"Li","sequence":"first","affiliation":[{"name":"Institute of Microelectronics of the Chinese Academy of Sciences, Beijing 100029, China"},{"name":"University of Chinese Academy of Sciences, Beijing 100029, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Rong","family":"Du","sequence":"additional","affiliation":[{"name":"Institute of Microelectronics of the Chinese Academy of Sciences, Beijing 100029, China"},{"name":"University of Chinese Academy of Sciences, Beijing 100029, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Shudong","family":"Chen","sequence":"additional","affiliation":[{"name":"Institute of Microelectronics of the Chinese Academy of Sciences, Beijing 100029, China"},{"name":"University of Chinese Academy of Sciences, Beijing 100029, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2022,3,28]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","unstructured":"Pavlakos, G., Zhou, X., Derpanis, K.G., and Daniilidis, K. (2016, January 27\u201330). Coarse-to-Fine Volumetric Prediction for Single-Image 3D Human Pose. Proceedings of the 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2017.139"},{"key":"ref_2","doi-asserted-by":"crossref","unstructured":"Tekin, B., M\u00e1rquez-Neila, P., Salzmann, M., and Fua, P. (2017, January 22\u201329). Learning to Fuse 2D and 3D Image Cues for Monocular Body Pose Estimation. Proceedings of the 2017 IEEE International Conference on Computer Vision (ICCV), Venice, Italy.","DOI":"10.1109\/ICCV.2017.425"},{"key":"ref_3","doi-asserted-by":"crossref","unstructured":"Martinez, J., Hossain, R., Romero, J., and Little, J. (2017, January 22\u201329). A Simple Yet Effective Baseline for 3d Human Pose Estimation. Proceedings of the 2017 IEEE International Conference on Computer Vision (ICCV), Venice, Italy.","DOI":"10.1109\/ICCV.2017.288"},{"key":"ref_4","first-page":"1","article-title":"Compositional Human Pose Regression","volume":"176\u2013177","author":"Sun","year":"2018","journal-title":"Comput. Vis. Image Underst."},{"key":"ref_5","doi-asserted-by":"crossref","unstructured":"Fang, H., Xu, Y., Wang, W., Liu, X., and Zhu, S.C. (2018, January 18\u201320). Learning Pose Grammar to Encode Human Body Configuration for 3D Pose Estimation. Proceedings of the AAAI 2018, Arlington, VA, USA.","DOI":"10.1609\/aaai.v32i1.12270"},{"key":"ref_6","doi-asserted-by":"crossref","unstructured":"Pavlakos, G., Zhou, X., and Daniilidis, K. (2018, January 18\u201322). Ordinal Depth Supervision for 3D Human Pose Estimation. Proceedings of the 2018 IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00763"},{"key":"ref_7","doi-asserted-by":"crossref","unstructured":"Yang, W., Ouyang, W., Wang, X., Ren, J.S.J., Li, H., and Wang, X. (2018, January 18\u201322). 3D Human Pose Estimation in the Wild by Adversarial Learning. Proceedings of the 2018 IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00551"},{"key":"ref_8","doi-asserted-by":"crossref","unstructured":"Luvizon, D.C., Picard, D., and Tabia, H. (2018, January 18\u201322). 2D\/3D Pose Estimation and Action Recognition Using Multitask Deep Learning. Proceedings of the 2018 IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00539"},{"key":"ref_9","doi-asserted-by":"crossref","unstructured":"Hossain, M.R.I., and Little, J. (2017). Exploiting temporal information for 3D pose estimation. arXiv.","DOI":"10.1007\/978-3-030-01249-6_5"},{"key":"ref_10","doi-asserted-by":"crossref","unstructured":"Lee, K., Lee, I., and Lee, S. (2018, January 8\u201314). Propagating LSTM: 3D Pose Estimation Based on Joint Interdependency. Proceedings of the European Conference on Computer Vision (ECCV), Munich, Germany.","DOI":"10.1007\/978-3-030-01234-2_8"},{"key":"ref_11","doi-asserted-by":"crossref","unstructured":"Zhao, L., Peng, X., Tian, Y., Kapadia, M., and Metaxas, D.N. (2019, January 15\u201320). Semantic Graph Convolutional Networks for 3D Human Pose Regression. Proceedings of the 2019 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.00354"},{"key":"ref_12","doi-asserted-by":"crossref","unstructured":"Cai, Y., Ge, L., Liu, J., Cai, J., Cham, T., Yuan, J., and Magnenat-Thalmann, N. (2019, January 15\u201320). Exploiting Spatial-Temporal Relationships for 3D Pose Estimation via Graph Convolutional Networks. Proceedings of the 2019 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Long Beach, CA, USA.","DOI":"10.1109\/ICCV.2019.00236"},{"key":"ref_13","doi-asserted-by":"crossref","unstructured":"Yan, S., Xiong, Y., and Lin, D. (2018). Spatial Temporal Graph Convolutional Networks for Skeleton-Based Action Recognition. arXiv.","DOI":"10.1609\/aaai.v32i1.12328"},{"key":"ref_14","unstructured":"Gehring, J., Auli, M., Grangier, D., Yarats, D., and Dauphin, Y. (2017, January 6\u201311). Convolutional Sequence to Sequence Learning. Proceedings of the 34th International Conference on Machine Learning, Sydney, Australia."},{"key":"ref_15","unstructured":"Dauphin, Y., Fan, A., Auli, M., and Grangier, D. (2017, January 6\u201311). Language Modeling with Gated Convolutional NetwoOrdinal depth supervision for 3d humanrks. Proceedings of the 34th International Conference on Machine Learning, Sydney, Australia."},{"key":"ref_16","unstructured":"van den Oord, A., Dieleman, S., Zen, H., Simonyan, K., Vinyals, O., Graves, A., Kalchbrenner, N., Senior, A.W., and Kavukcuoglu, K. (2016). WaveNet: A Generative Model for Raw Audio. arXiv."},{"key":"ref_17","unstructured":"Collobert, R., Puhrsch, C., and Synnaeve, G. (2016). Wav2Letter: An End-to-End ConvNet-based Speech Recognition System. arXiv."},{"key":"ref_18","doi-asserted-by":"crossref","unstructured":"Ronneberger, O., Fischer, P., and Brox, T. (2015, January 5\u20139). U-Net: Convolutional Networks for Biomedical Image Segmentation. Proceedings of the MICCAI, Munich, Germany.","DOI":"10.1007\/978-3-319-24574-4_28"},{"key":"ref_19","doi-asserted-by":"crossref","unstructured":"Sminchisescu, C. (2006, January 22\u201324). 3D Human Motion Analysis in Monocular Video Techniques and Challenges. Proceedings of the AVSS, Sydney, Australia.","DOI":"10.1109\/AVSS.2006.3"},{"key":"ref_20","doi-asserted-by":"crossref","unstructured":"Ramakrishna, V., Kanade, T., and Sheikh, Y. (2012, January 7\u201313). Reconstructing 3D Human Pose from 2D Image Landmarks. Proceedings of the 12th European Conference on Computer Vision, Florence, Italy.","DOI":"10.1007\/978-3-642-33765-9_41"},{"key":"ref_21","doi-asserted-by":"crossref","first-page":"1325","DOI":"10.1109\/TPAMI.2013.248","article-title":"Human3.6M: Large Scale Datasets and Predictive Methods for 3D Human Sensing in Natural Environments","volume":"36","author":"Ionescu","year":"2014","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_22","doi-asserted-by":"crossref","unstructured":"Ionescu, C., Carreira, J., and Sminchisescu, C. (2014, January 23\u201328). Iterated Second-Order Label Sensitive Pooling for 3D Human Pose Estimation. Proceedings of the 2014 IEEE Conference on Computer Vision and Pattern Recognition, Columbus, OH, USA.","DOI":"10.1109\/CVPR.2014.215"},{"key":"ref_23","doi-asserted-by":"crossref","unstructured":"Li, S., and Chan, A.B. (2014, January 1\u20135). 3D Human Pose Estimation from Monocular Images with Deep Convolutional Neural Network. Proceedings of the ACCV 2014, Singapore.","DOI":"10.1007\/978-3-319-16808-1_23"},{"key":"ref_24","doi-asserted-by":"crossref","unstructured":"Tekin, B., Rozantsev, A., Lepetit, V., and Fua, P.V. (2016, January 27\u201330). Direct Prediction of 3D Body Poses from Motion Compensated Sequences. Proceedings of the 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.113"},{"key":"ref_25","doi-asserted-by":"crossref","unstructured":"Tekin, B., Katircioglu, I., Salzmann, M., Lepetit, V., and Fua, P.V. (2016). Structured Prediction of 3D Human Pose with Deep Neural Networks. arXiv.","DOI":"10.5244\/C.30.130"},{"key":"ref_26","doi-asserted-by":"crossref","unstructured":"Jiang, H. (2010, January 23\u201326). 3D Human Pose Reconstruction Using Millions of Exemplars. Proceedings of the 2010 20th International Conference on Pattern Recognition, Istanbul, Turkey.","DOI":"10.1109\/ICPR.2010.414"},{"key":"ref_27","doi-asserted-by":"crossref","unstructured":"Chen, C.H., and Ramanan, D. (2017, January 21\u201326). 3D Human Pose Estimation = 2D Pose Estimation + Matching. Proceedings of the 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.610"},{"key":"ref_28","doi-asserted-by":"crossref","unstructured":"Park, S., Hwang, J., and Kwak, N. (2016, January 11\u201314). 3D Human Pose Estimation Using Convolutional Neural Networks with 2D Pose Information. Proceedings of the ECCV Workshops, Amsterdam, The Netherlands.","DOI":"10.1007\/978-3-319-49409-8_15"},{"key":"ref_29","doi-asserted-by":"crossref","unstructured":"Zhou, X., Huang, Q., Sun, X., Xue, X., and Wei, Y. (2017, January 22\u201329). Towards 3D Human Pose Estimation in the Wild: A Weakly-Supervised Approach. Proceedings of the 2017 IEEE International Conference on Computer Vision (ICCV), Venice, Italy.","DOI":"10.1109\/ICCV.2017.51"},{"key":"ref_30","doi-asserted-by":"crossref","unstructured":"Brau, E., and Jiang, H. (2016, January 25\u201328). 3D Human Pose Estimation via Deep Learning from 2D Annotations. Proceedings of the 2016 Fourth International Conference on 3D Vision (3DV), Stanford, CA, USA.","DOI":"10.1109\/3DV.2016.84"},{"key":"ref_31","doi-asserted-by":"crossref","unstructured":"He, Y., Yan, R., Fragkiadaki, K., and Yu, S.I. (2020, January 13\u201319). Epipolar Transformers. Proceedings of the 2020 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA.","DOI":"10.1109\/CVPR42600.2020.00780"},{"key":"ref_32","unstructured":"Plizzari, C., Cannici, M., and Matteucci, M. (, January 10\u201315). Spatial Temporal Transformer Network for Skeleton-based Action Recognition. Proceedings of the ICPR Workshops, Virtual Event."},{"key":"ref_33","doi-asserted-by":"crossref","unstructured":"Hu, W., Zhang, C., Zhan, F., Zhang, L., and Wong, T.T. (2021, January 20\u201324). Conditional Directed Graph Convolution for 3D Human Pose Estimation. Proceedings of the 29th ACM International Conference on Multimedia, Virtual Event.","DOI":"10.1145\/3474085.3475219"},{"key":"ref_34","doi-asserted-by":"crossref","unstructured":"Nan, M., Trascau, M., Florea, A.M., and Iacob, C.C. (2021). Comparison between Recurrent Networks and Temporal Convolutional Networks Approaches for Skeleton-Based Action Recognition. Sensors, 21.","DOI":"10.3390\/s21062051"},{"key":"ref_35","doi-asserted-by":"crossref","unstructured":"Lin, M., Lin, L., Liang, X., Wang, K., and Cheng, H. (2017, January 21\u201326). Recurrent 3D Pose Sequence Machines. Proceedings of the 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.588"},{"key":"ref_36","doi-asserted-by":"crossref","first-page":"1326","DOI":"10.1007\/s11263-018-1066-6","article-title":"Learning Latent Representations of 3D Human Pose with Deep Neural Networks","volume":"126","author":"Katircioglu","year":"2018","journal-title":"Int. J. Comput. Vis."},{"key":"ref_37","doi-asserted-by":"crossref","unstructured":"Hossain, M.R.I., and Little, J. (2018, January 8\u201314). Exploiting Temporal Information for 3D Human Pose Estimation. Proceedings of the European Conference on Computer Vision (ECCV), Munich, Germany.","DOI":"10.1007\/978-3-030-01249-6_5"},{"key":"ref_38","doi-asserted-by":"crossref","unstructured":"Li, M., Chen, S., Chen, X., Zhang, Y., Wang, Y., and Tian, Q. (2019, January 15\u201320). Actional-Structural Graph Convolutional Networks for Skeleton-Based Action Recognition. Proceedings of the 2019 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.00371"},{"key":"ref_39","doi-asserted-by":"crossref","unstructured":"Shi, L., Zhang, Y., Cheng, J., and Lu, H. (2019, January 15\u201320). Skeleton-Based Action Recognition with Directed Graph Neural Networks. Proceedings of the 2019 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.00810"},{"key":"ref_40","doi-asserted-by":"crossref","unstructured":"Shi, L., Zhang, Y., Cheng, J., and Lu, H. (2019, January 15\u201320). Two-Stream Adaptive Graph Convolutional Networks for Skeleton-Based Action Recognition. Proceedings of the 2019 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.01230"},{"key":"ref_41","doi-asserted-by":"crossref","unstructured":"Pavllo, D., Feichtenhofer, C., Grangier, D., and Auli, M. (2019, January 15\u201320). 3D human pose estimation in video with temporal convolutions and semi-supervised training. Proceedings of the 2019 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.00794"},{"key":"ref_42","doi-asserted-by":"crossref","unstructured":"Wang, X., Girshick, R.B., Gupta, A.K., and He, K. (2018, January 18\u201322). Non-local Neural Networks. Proceedings of the 2018 IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00813"},{"key":"ref_43","doi-asserted-by":"crossref","first-page":"4","DOI":"10.1007\/s11263-009-0273-6","article-title":"HumanEva: Synchronized Video and Motion Capture Dataset and Baseline Algorithm for Evaluation of Articulated Human Motion","volume":"87","author":"Sigal","year":"2009","journal-title":"Int. J. Comput. Vis."},{"key":"ref_44","unstructured":"Yeh, R.A., Hu, Y.T., and Schwing, A.G. (2019, January 8\u201314). Chirality Nets for Human Pose Regression. Proceedings of the NeurIPS 2019, Vancouver, BC, Canada."},{"key":"ref_45","doi-asserted-by":"crossref","unstructured":"Xu, J., Yu, Z., Ni, B., Yang, J., Yang, X., and Zhang, W. (2020, January 13\u201319). Deep Kinematics Analysis for Monocular 3D Human Pose Estimation. Proceedings of the 2020 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA.","DOI":"10.1109\/CVPR42600.2020.00098"},{"key":"ref_46","doi-asserted-by":"crossref","unstructured":"Ci, H., Wang, C., Ma, X., and Wang, Y. (2019, January 15\u201320). Optimizing Network Structure for 3D Human Pose Estimation. Proceedings of the 2019 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Long Beach, CA, USA.","DOI":"10.1109\/ICCV.2019.00235"},{"key":"ref_47","unstructured":"Ioffe, S., and Szegedy, C. (2015). Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift. arXiv."},{"key":"ref_48","unstructured":"Reddi, S.J., Kale, S., and Kumar, S. (2018). On the Convergence of Adam and Beyond. arXiv."},{"key":"ref_49","doi-asserted-by":"crossref","unstructured":"Liu, R., Shen, J., Wang, H., Chen, C., Cheung, S.C.S., and Asari, V.K. (2020, January 13\u201319). Attention Mechanism Exploits Temporal Contexts: Real-Time 3D Human Pose Reconstruction. Proceedings of the 2020 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA.","DOI":"10.1109\/CVPR42600.2020.00511"},{"key":"ref_50","unstructured":"Ying, R., You, J., Morris, C., Ren, X., Hamilton, W.L., and Leskovec, J. (2018). Hierarchical Graph Representation Learning with Differentiable Pooling. arXiv."},{"key":"ref_51","unstructured":"Lee, J., Lee, I., and Kang, J. (2019). Self-Attention Graph Pooling. arXiv."}],"container-title":["Sensors"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1424-8220\/22\/7\/2573\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,10]],"date-time":"2025-10-10T22:44:35Z","timestamp":1760136275000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1424-8220\/22\/7\/2573"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,3,28]]},"references-count":51,"journal-issue":{"issue":"7","published-online":{"date-parts":[[2022,4]]}},"alternative-id":["s22072573"],"URL":"https:\/\/doi.org\/10.3390\/s22072573","relation":{},"ISSN":["1424-8220"],"issn-type":[{"type":"electronic","value":"1424-8220"}],"subject":[],"published":{"date-parts":[[2022,3,28]]}}}