{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,1,22]],"date-time":"2026-01-22T09:06:23Z","timestamp":1769072783126,"version":"3.49.0"},"reference-count":41,"publisher":"MDPI AG","issue":"7","license":[{"start":{"date-parts":[[2020,7,4]],"date-time":"2020-07-04T00:00:00Z","timestamp":1593820800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["61702350"],"award-info":[{"award-number":["61702350"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Symmetry"],"abstract":"<jats:p>In this paper, we study the problem of monocular 3D human pose estimation based on deep learning. Due to single view limitations, the monocular human pose estimation cannot avoid the inherent occlusion problem. The common methods use the multi-view based 3D pose estimation method to solve this problem. However, single-view images cannot be used directly in multi-view methods, which greatly limits practical applications. To address the above-mentioned issues, we propose a novel end-to-end 3D pose estimation network for monocular 3D human pose estimation. First, we propose a multi-view pose generator to predict multi-view 2D poses from the 2D poses in a single view. Secondly, we propose a simple but effective data augmentation method for generating multi-view 2D pose annotations, on account of the existing datasets (e.g., Human3.6M, etc.) not containing a large number of 2D pose annotations in different views. Thirdly, we employ graph convolutional network to infer a 3D pose from multi-view 2D poses. From experiments conducted on public datasets, the results have verified the effectiveness of our method. Furthermore, the ablation studies show that our method improved the performance of existing 3D pose estimation networks.<\/jats:p>","DOI":"10.3390\/sym12071116","type":"journal-article","created":{"date-parts":[[2020,7,6]],"date-time":"2020-07-06T11:07:42Z","timestamp":1594033662000},"page":"1116","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":16,"title":["Multi-View Pose Generator Based on Deep Learning for Monocular 3D Human Pose Estimation"],"prefix":"10.3390","volume":"12","author":[{"given":"Jun","family":"Sun","sequence":"first","affiliation":[{"name":"College of Information and Engineering, Sichuan Agricultural University, Yaan 625014, China"},{"name":"The Lab of Agricultural Information Engineering, Sichuan Key Laboratory, Yaan 625014, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Mantao","family":"Wang","sequence":"additional","affiliation":[{"name":"College of Information and Engineering, Sichuan Agricultural University, Yaan 625014, China"},{"name":"The Lab of Agricultural Information Engineering, Sichuan Key Laboratory, Yaan 625014, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Xin","family":"Zhao","sequence":"additional","affiliation":[{"name":"School of Geography and Information Engineering, China University of Geosciences, Wuhan 430074, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-9129-534X","authenticated-orcid":false,"given":"Dejun","family":"Zhang","sequence":"additional","affiliation":[{"name":"School of Geography and Information Engineering, China University of Geosciences, Wuhan 430074, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2020,7,4]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","first-page":"107312","DOI":"10.1016\/j.patcog.2020.107312","article-title":"Learning motion representation for real-time spatio-temporal action localization","volume":"103","author":"Zhang","year":"2020","journal-title":"Pattern Recognit."},{"key":"ref_2","doi-asserted-by":"crossref","first-page":"39","DOI":"10.1016\/j.jvcir.2013.03.011","article-title":"An adaptable system for RGB-D based human body detection and pose estimation","volume":"25","author":"Buys","year":"2014","journal-title":"J. Vis. Commun. Image Represent."},{"key":"ref_3","doi-asserted-by":"crossref","unstructured":"Ancillao, A. (2018). Stereophotogrammetry in Functional Evaluation: History and Modern Protocols, Springer.","DOI":"10.1007\/978-3-319-67437-7_1"},{"key":"ref_4","doi-asserted-by":"crossref","first-page":"169","DOI":"10.1016\/j.dsp.2015.05.011","article-title":"Bayesian classification and analysis of gait disorders using image and depth sensors of Microsoft Kinect","volume":"47","year":"2015","journal-title":"Digit. Signal Process."},{"key":"ref_5","doi-asserted-by":"crossref","first-page":"2722","DOI":"10.1016\/j.jbiomech.2013.08.011","article-title":"Concurrent validity of the Microsoft Kinect for assessment of spatiotemporal gait variables","volume":"46","author":"Clark","year":"2013","journal-title":"J. Biomech."},{"key":"ref_6","unstructured":"Cortes, C., Lawrence, N.D., Lee, D.D., Sugiyama, M., and Garnett, R. (2015). Learning Structured Output Representation using Deep Conditional Generative Models. Advances in Neural Information Processing Systems 28, Curran Associates, Inc."},{"key":"ref_7","unstructured":"Ghahramani, Z., Welling, M., Cortes, C., Lawrence, N.D., and Weinberger, K.Q. (2014). Generative Adversarial Nets. Advances in Neural Information Processing Systems 27, Curran Associates, Inc."},{"key":"ref_8","doi-asserted-by":"crossref","unstructured":"Yang, W., Ouyang, W., Wang, X., Ren, J., Li, H., and Wang, X. (2018, January 18\u201322). 3D Human Pose Estimation in the Wild by Adversarial Learning. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00551"},{"key":"ref_9","unstructured":"Kipf, T.N., and Welling, M. (2017, January 24\u201326). Semi-Supervised Classification with Graph Convolutional Networks. Proceedings of the 5th International Conference on Learning Representations, ICLR 2017, Toulon, France."},{"key":"ref_10","doi-asserted-by":"crossref","unstructured":"Martinez, J., Hossain, R., Romero, J., and Little, J.J. (2017, January 22\u201329). A Simple Yet Effective Baseline for 3d Human Pose Estimation. Proceedings of the IEEE International Conference on Computer Vision, Venice, Italy.","DOI":"10.1109\/ICCV.2017.288"},{"key":"ref_11","first-page":"1069","article-title":"3D Human Pose Machines with Self-supervised Learning","volume":"42","author":"Wang","year":"2020","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_12","doi-asserted-by":"crossref","unstructured":"Kocabas, M., Karagoz, S., and Akbas, E. (2019, January 16\u201320). Self-Supervised Learning of 3D Human Pose Using Multi-View Geometry. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.00117"},{"key":"ref_13","unstructured":"Fang, H., Xu, Y., Wang, W., Liu, X., and Zhu, S.C. (2017). Learning Knowledge-guided Pose Grammar Machine for 3D Human Pose Estimation. arXiv."},{"key":"ref_14","doi-asserted-by":"crossref","unstructured":"Zhou, X., Huang, Q., Sun, X., Xue, X., and Wei, Y. (2017, January 22\u201329). Towards 3D Human Pose Estimation in the Wild: A Weakly-Supervised Approach. Proceedings of the IEEE International Conference on Computer Vision (ICCV), Venice, Italy.","DOI":"10.1109\/ICCV.2017.51"},{"key":"ref_15","doi-asserted-by":"crossref","unstructured":"Zhao, L., Peng, X., Tian, Y., Kapadia, M., and Metaxas, D.N. (2019, January 16\u201320). Semantic Graph Convolutional Networks for 3D Human Pose Regression. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.00354"},{"key":"ref_16","doi-asserted-by":"crossref","unstructured":"Rhodin, H., Sp\u00f6rri, J., Katircioglu, I., Constantin, V., Meyer, F., M\u00fcller, E., Salzmann, M., and Fua, P. (2018, January 18\u201322). Learning Monocular 3D Human Pose Estimation From Multi-View Images. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00880"},{"key":"ref_17","unstructured":"Qiu, H., Wang, C., Wang, J., Wang, N., and Zeng, W. (November, January 27). Cross View Fusion for 3D Human Pose Estimation. Proceedings of the International Conference on Computer Vision, Seoul, Korea."},{"key":"ref_18","doi-asserted-by":"crossref","unstructured":"Dong, J., Jiang, W., Huang, Q., Bao, H., and Zhou, X. (2019, January 16\u201320). Fast and Robust Multi-Person 3D Pose Estimation from Multiple Views. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.00798"},{"key":"ref_19","unstructured":"Iskakov, K., Burkov, E., Lempitsky, V., and Malkov, Y. (November, January 27). Learnable Triangulation of Human Pose. Proceedings of the International Conference on Computer Vision (ICCV), Seoul, Korea."},{"key":"ref_20","doi-asserted-by":"crossref","first-page":"1325","DOI":"10.1109\/TPAMI.2013.248","article-title":"Human3.6M: Large Scale Datasets and Predictive Methods for 3D Human Sensing in Natural Environments","volume":"36","author":"Ionescu","year":"2014","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_21","doi-asserted-by":"crossref","unstructured":"Huang, F., Zeng, A., Liu, M., Lai, Q., and Xu, Q. (2020, January 1\u20135). DeepFuse: An IMU-Aware Network for Real-Time 3D Human Pose Estimation from Multi-View Image. Proceedings of the IEEE Winter Conference on Applications of Computer Vision (WACV), Snowmass Village, CO, USA.","DOI":"10.1109\/WACV45572.2020.9093526"},{"key":"ref_22","doi-asserted-by":"crossref","unstructured":"Ci, H., Wang, C., Ma, X., and Wang, Y. (2019, January 16\u201320). Optimizing Network Structure for 3D Human Pose Estimation. Proceedings of the IEEE International Conference on Computer Vision, Long Beach, CA, USA.","DOI":"10.1109\/ICCV.2019.00235"},{"key":"ref_23","doi-asserted-by":"crossref","unstructured":"Pavlakos, G., Zhou, X., Derpanis, K.G., and Daniilidis, K. (2017, January 21\u201326). Harvesting multiple views for marker-less 3d human pose annotations. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.138"},{"key":"ref_24","first-page":"3","article-title":"Total Capture: 3D Human Pose Estimation Fusing Video and Inertial Sensors","volume":"2","author":"Trumble","year":"2017","journal-title":"BMVC"},{"key":"ref_25","doi-asserted-by":"crossref","unstructured":"Tome, D., Toso, M., Agapito, L., and Russell, C. (2018, January 5\u20138). Rethinking pose in 3d: Multi-stage refinement and recovery for markerless motion capture. Proceedings of the 2018 International Conference on 3D Vision (3DV), Verona, Italy.","DOI":"10.1109\/3DV.2018.00061"},{"key":"ref_26","doi-asserted-by":"crossref","unstructured":"Yan, S., Xiong, Y., and Lin, D. (2018, January 2\u20137). Spatial temporal graph convolutional networks for skeleton-based action recognition. Proceedings of the Thirty-Second AAAI Conference on Artificial Intelligence, New Orleans, LA, USA.","DOI":"10.1609\/aaai.v32i1.12328"},{"key":"ref_27","doi-asserted-by":"crossref","unstructured":"Pham, H.H., Salmane, H., Khoudour, L., Crouzil, A., Velastin, S.A., and Zegers, P. (2020). A unified deep framework for joint 3d pose estimation and action recognition from a single rgb camera. Sensors, 20.","DOI":"10.3390\/s20071825"},{"key":"ref_28","unstructured":"Lin, J., and Lee, G.H. (2019). Trajectory Space Factorization for Deep Video-Based 3D Human Pose Estimation. arXiv."},{"key":"ref_29","unstructured":"Cheng, Y., Yang, B., Wang, B., Yan, W., and Tan, R.T. (November, January 27). Occlusion-Aware Networks for 3D Human Pose Estimation in Video. Proceedings of the IEEE International Conference on Computer Vision, Seoul, Korea."},{"key":"ref_30","doi-asserted-by":"crossref","first-page":"28936","DOI":"10.1109\/ACCESS.2018.2837654","article-title":"Integrating feature selection and feature extraction methods with deep learning to predict clinical outcome of breast cancer","volume":"6","author":"Zhang","year":"2018","journal-title":"IEEE Access"},{"key":"ref_31","doi-asserted-by":"crossref","unstructured":"Li, M., Chen, S., Chen, X., Zhang, Y., Wang, Y., and Tian, Q. (2019, January 16\u201320). Actional-structural graph convolutional networks for skeleton-based action recognition. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Long Beach, CA.","DOI":"10.1109\/CVPR.2019.00371"},{"key":"ref_32","unstructured":"Cai, Y., Ge, L., Liu, J., Cai, J., Cham, T.J., Yuan, J., and Thalmann, N.M. (November, January 27). Exploiting spatial-temporal relationships for 3d pose estimation via graph convolutional networks. Proceedings of the IEEE International Conference on Computer Vision, Seoul, Korea."},{"key":"ref_33","doi-asserted-by":"crossref","unstructured":"Zhang, Y., An, L., Yu, T., Li, X., Li, K., and Liu, Y. (2020). 4D Association Graph for Realtime Multi-person Motion Capture Using Multiple Video Cameras. arXiv.","DOI":"10.1109\/CVPR42600.2020.00140"},{"key":"ref_34","doi-asserted-by":"crossref","unstructured":"Liu, J., Guang, Y., and Rojas, J. (2020). GAST-Net: Graph Attention Spatio-temporal Convolutional Networks for 3D Human Pose Estimation in Video. arXiv.","DOI":"10.1109\/ICRA48506.2021.9561605"},{"key":"ref_35","doi-asserted-by":"crossref","unstructured":"Wang, L., Chen, Y., Guo, Z., Qian, K., Lin, M., Li, H., and Ren, J.S. (2019). Generalizing Monocular 3D Human Pose Estimation in the Wild. arXiv.","DOI":"10.1109\/ICCVW.2019.00497"},{"key":"ref_36","unstructured":"Misra, D. (2019). Mish: A Self Regularized Non-Monotonic Neural Activation Function. arXiv."},{"key":"ref_37","unstructured":"Nair, V., and Hinton, G.E. (2010, January 21\u201324). Rectified linear units improve restricted boltzmann machines. Proceedings of the 27th International Conference on Machine Learning (ICML-10), Haifa, Israel."},{"key":"ref_38","first-page":"405","article-title":"Noisy Softplus: A Biology Inspired Activation Function","volume":"Volume 9950","author":"Liu","year":"2016","journal-title":"International Conference on Neural Information Processing"},{"key":"ref_39","doi-asserted-by":"crossref","unstructured":"Wang, X., Girshick, R., Gupta, A., and He, K. (2018, January 18\u201322). Non-Local Neural Networks. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00813"},{"key":"ref_40","doi-asserted-by":"crossref","unstructured":"Sun, X., Shang, J., Liang, S., and Wei, Y. (2017, January 22\u201329). Compositional Human Pose Regression. Proceedings of the IEEE International Conference on Computer Vision (ICCV), Venice, Italy.","DOI":"10.1109\/ICCV.2017.284"},{"key":"ref_41","doi-asserted-by":"crossref","unstructured":"Fang, H., Xu, Y., Wang, W., Liu, X., and Zhu, S.C. (2018, January 2\u20137). Learning Pose Grammar to Encode Human Body Configuration for 3D Pose Estimation. Proceedings of the Thirty-Second AAAI Conference on Artificial Intelligence, New Orleans, LA, USA.","DOI":"10.1609\/aaai.v32i1.12270"}],"container-title":["Symmetry"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/2073-8994\/12\/7\/1116\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,11]],"date-time":"2025-10-11T09:47:21Z","timestamp":1760176041000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/2073-8994\/12\/7\/1116"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2020,7,4]]},"references-count":41,"journal-issue":{"issue":"7","published-online":{"date-parts":[[2020,7]]}},"alternative-id":["sym12071116"],"URL":"https:\/\/doi.org\/10.3390\/sym12071116","relation":{},"ISSN":["2073-8994"],"issn-type":[{"value":"2073-8994","type":"electronic"}],"subject":[],"published":{"date-parts":[[2020,7,4]]}}}