{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,11,15]],"date-time":"2025-11-15T10:23:41Z","timestamp":1763202221758,"version":"build-2065373602"},"reference-count":37,"publisher":"MDPI AG","issue":"3","license":[{"start":{"date-parts":[[2019,2,10]],"date-time":"2019-02-10T00:00:00Z","timestamp":1549756800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"name":"The National Natureal Science Foundation","award":["61762025"],"award-info":[{"award-number":["61762025"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Sensors"],"abstract":"<jats:p>In recent years, increasing human data comes from image sensors. In this paper, a novel approach combining convolutional pose machines (CPMs) with GoogLeNet is proposed for human pose estimation using image sensor data. The first stage of the CPMs directly generates a response map of each human skeleton\u2019s key points from images, in which we introduce some layers from the GoogLeNet. On the one hand, the improved model uses deeper network layers and more complex network structures to enhance the ability of low level feature extraction. On the other hand, the improved model applies a fine-tuning strategy, which benefits the estimation accuracy. Moreover, we introduce the inception structure to greatly reduce parameters of the model, which reduces the convergence time significantly. Extensive experiments on several datasets show that the improved model outperforms most mainstream models in accuracy and training time. The prediction efficiency of the improved model is improved by 1.023 times compared with the CPMs. At the same time, the training time of the improved model is reduced 3.414 times. This paper presents a new idea for future research.<\/jats:p>","DOI":"10.3390\/s19030718","type":"journal-article","created":{"date-parts":[[2019,2,12]],"date-time":"2019-02-12T03:18:20Z","timestamp":1549941500000},"page":"718","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":23,"title":["Improved Convolutional Pose Machines for Human Pose Estimation Using Image Sensor Data"],"prefix":"10.3390","volume":"19","author":[{"given":"Baohua","family":"Qiang","sequence":"first","affiliation":[{"name":"Guangxi Key Laboratory of Trusted Software, Guilin University of Electronic Technology, Guilin 541004, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Shihao","family":"Zhang","sequence":"additional","affiliation":[{"name":"Guangxi Key Laboratory of Trusted Software, Guilin University of Electronic Technology, Guilin 541004, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-1616-6868","authenticated-orcid":false,"given":"Yongsong","family":"Zhan","sequence":"additional","affiliation":[{"name":"Guangxi Key Laboratory of Trusted Software, Guilin University of Electronic Technology, Guilin 541004, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Wu","family":"Xie","sequence":"additional","affiliation":[{"name":"Guangxi Key Laboratory of Trusted Software, Guilin University of Electronic Technology, Guilin 541004, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Tian","family":"Zhao","sequence":"additional","affiliation":[{"name":"Guangxi Key Laboratory of Trusted Software, Guilin University of Electronic Technology, Guilin 541004, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2019,2,10]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","unstructured":"Wang, L., Zang, J.L., Zhang, Q.L., Niu, Z.X., Hua, G., and Zheng, N.N. (2018). Action Recognition by an Attention-Aware Temporal Weighted Convolutional Neural NetWork. Sensors, 18.","DOI":"10.3390\/s18071979"},{"key":"ref_2","doi-asserted-by":"crossref","unstructured":"Gong, W.J., Zhang, X.N., Gonezalez, J., Sobral, A., Bouwmans, T., Tu, C.H., and Zahzah, E.-H. (2016). Human Pose Estimation from Monocular Images: A Comprehensive Survey. Sensors, 16.","DOI":"10.3390\/s16121966"},{"key":"ref_3","first-page":"1","article-title":"Progress in two-dimensional human pose estimation","volume":"4","author":"Han","year":"2017","journal-title":"J. Xi\u2019an Univ. Posts Telecom."},{"key":"ref_4","doi-asserted-by":"crossref","unstructured":"Pishchulin, L., Insafutdinov, E., Tang, S., Andres, B., Andriluka, M., Gehler, P., and Schiele, B. (2016, January 27\u201330). Deepcut: Joint subset partition and labeling for multi person pose estimation. Proceedings of the 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.533"},{"key":"ref_5","doi-asserted-by":"crossref","unstructured":"Tompson, J., Goroshin, R., Jain, A., LeCun, Y., and Bregler, C. (2015, January 8\u201310). Efficient object localization using convolutional networks. Proceedings of the 2015 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Boston, MA, USA.","DOI":"10.1109\/CVPR.2015.7298664"},{"key":"ref_6","unstructured":"Tompson, J., Jain, A., LeCun, Y., and Bregler, C. (2014, January 8\u201313). Joints training of a convolutional network and a graphical model for human pose estimation. Proceedings of the 2014 International Conference on Neural Information Processing Systems (NIPS), Montreal, QC, Canada."},{"key":"ref_7","unstructured":"Wang, R. (2016, March 27). Human Posture Estimation based on Deep Convolution Neural Network. Available online: http:\/\/nvsm.cnki.net\/kns\/brief\/default_result.aspx."},{"key":"ref_8","doi-asserted-by":"crossref","unstructured":"Pfister, T., Charles, J., and Zisserman, A. (2015, January 11\u201316). Flowing ConvNets for Human Pose Estimation in Videos. Proceedings of the 2015 IEEE International Conference on Computer Vision (ICCV), Santiago, Chile.","DOI":"10.1109\/ICCV.2015.222"},{"key":"ref_9","doi-asserted-by":"crossref","unstructured":"Insafutdinov, E., Pishchulin, L., Andres, B., Andriluka, M., and Schiele, B. (2016, January 8\u201316). Deepercut: A deeper, stronger, and faster multi-person pose estimation model. Proceedings of the 2016 European Conference on Computer Vision (ECCV), Amsterdam, Netherlands.","DOI":"10.1007\/978-3-319-46466-4_3"},{"key":"ref_10","doi-asserted-by":"crossref","unstructured":"Wei, S.-E., Ramakrishna, V., Kanade, T., and Sheikh, Y. (2016, January 27\u201330). Convolutional Pose Machines. Proceedings of the 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.511"},{"key":"ref_11","unstructured":"(2018, November 11). MPII Human Pose Dataset. Available online: http:\/\/human-pose.mpi-inf.mpg.de."},{"key":"ref_12","unstructured":"(2018, November 11). Leeds Sports Pose. Available online: http:\/\/sam.johnson.io\/research\/lsp.html."},{"key":"ref_13","doi-asserted-by":"crossref","unstructured":"Newell, A., Yang, K., and Deng, J. (2016, January 8\u201316). Stacked hourglass networks for human pose estimation. Proceedings of the 2016 European Conference on Computer Vision (ECCV), Amsterdam, Netherlands.","DOI":"10.1007\/978-3-319-46484-8_29"},{"key":"ref_14","doi-asserted-by":"crossref","unstructured":"Chu, X., Yang, W., Ouyang, W., Ma, C., Yuille, A.L., and Wang, X. (2017, January 21\u201329). Multi-context attention for human pose estimation. Proceedings of the 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.601"},{"key":"ref_15","unstructured":"Chou, C., Chien, J., and Chen, H. (2017, January 21\u201329). Self adversarial training for human pose estimation. Proceedings of the 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Honolulu, HI, USA."},{"key":"ref_16","doi-asserted-by":"crossref","unstructured":"Yang, W., Li, S., Ouyang, W., Li, H., and Wang, X. (2017, January 22\u201329). Learning feature pyramids for human pose estimation. Proceedings of the 2017 IEEE International Conference on Computer Vision (ICCV), Venice, Italy.","DOI":"10.1109\/ICCV.2017.144"},{"key":"ref_17","doi-asserted-by":"crossref","unstructured":"Szegedy, C., Liu, W., Jia, Y., Sermanet, P., Reed, S., Anguelov, D., Erhan, D., Vanhoucke, V., and Rabinovich, A. (2015, January 8\u201310). Going deeper with convolutions. Proceedings of the 2015 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Boston, MA, USA.","DOI":"10.1109\/CVPR.2015.7298594"},{"key":"ref_18","doi-asserted-by":"crossref","unstructured":"Ramakrishna, V., Munoz, D., Hebert, M., Bagnell, J., and Sheikh, Y. (2014, January 6\u201312). Pose Machines: Articulated Pose Estimation via Inference Machines. Proceedings of the 2014 European Conference on Computer Vision (ECCV), Zurich, Switzerland.","DOI":"10.1007\/978-3-319-10605-2_3"},{"key":"ref_19","unstructured":"Krizhevsky, A., Sutskever, I., and Hinton, G.E. (2012, January 3\u20136). ImageNet classification with deep convolutional neural networks. Proceedings of the 2012 International Conference on Neural Information Processing Systems (NIPS), Lake Tahoe, NV, USA."},{"key":"ref_20","unstructured":"Jia, Y., Shelhamer, E., and Donahue, J. (2015, January 8\u201310). Caffe: Convolutional architecture for fast feature embedding. Proceedings of the 2015 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Boston, MA, USA."},{"key":"ref_21","first-page":"1229","article-title":"Review of Convolutional Neural Networks","volume":"40","author":"Zhou","year":"2017","journal-title":"J. Comput. Sci."},{"key":"ref_22","unstructured":"Lee, C.-Y., Xie, S., Gallagher, P., Zhang, Z., and Tu, Z. (2015, January 9\u201312). Deeply supervised nets. Proceedings of the 2015 International Conference on Artificial Intelligence and Statistics (AISTATS), San Diego, CA, USA."},{"key":"ref_23","doi-asserted-by":"crossref","first-page":"157","DOI":"10.1109\/72.279181","article-title":"Learning long-term dependencies with gradient descent is difficult","volume":"5","author":"Bengio","year":"1994","journal-title":"IEEE Trans. Neural Netw."},{"key":"ref_24","unstructured":"Bradley, D. (2010). Learning in Modular Systems. [Ph.D. Thesis, Robotics Institute, Carnegie Mellon University]."},{"key":"ref_25","unstructured":"Glorot, X., and Bengio, Y. (2010, January 13\u201315). Understanding the difficulty of training deep feedforward neural networks. Proceedings of the 2010 International Conference on Artificial Intelligence and Statistics (AISTATS), Sardinia, Italy."},{"key":"ref_26","unstructured":"Hochreiter, S., Bengio, Y., Frasconi, P., and Schmidhuber, J. (2018, October 10). Gradient Flow in Recurrent Nets: The Difficulty of Learning Long-Term Dependencies. Available online: http:\/\/citeseerx.ist.psu.edu\/viewdoc\/summary?doi=10.1.1.24.7321."},{"key":"ref_27","doi-asserted-by":"crossref","unstructured":"Deng, F., Pu, S.L., Chen, X.H., Shi, Y.S., Yuan, T., and Pu, S.Y. (2018). Hyperspectral Image Classification with Capsule Network Using Limited Training Samples. Sensors, 18.","DOI":"10.3390\/s18093153"},{"key":"ref_28","doi-asserted-by":"crossref","unstructured":"Mohamed, A., Hinton, G., and Penn, G. (2012, January 25\u201330). Understanding how deep belief networks perform acoustic modeling. Proceedings of the 2012 IEEE International Conference on Acoustics, Speech and Signal Processing, Kyoto, Japan.","DOI":"10.1109\/ICASSP.2012.6288863"},{"key":"ref_29","doi-asserted-by":"crossref","first-page":"30","DOI":"10.1109\/TASL.2011.2134090","article-title":"Context-Dependent Pre-Trained Deep Neural Networks for Large-Vocabulary Speech Recognition","volume":"20","author":"Dahl","year":"2012","journal-title":"IEEE Trans. Audio Speech Lang. Process."},{"key":"ref_30","unstructured":"(2018, November 11). The Extended Leeds Sports Pose. Available online: http:\/\/sam.johnson.io\/research\/lspet.html."},{"key":"ref_31","first-page":"21","article-title":"A Survey of Research Work on Neural Network Generalization and Structure Optimization Algorithms","volume":"19","author":"Wu","year":"2002","journal-title":"Appl. Res. Comput."},{"key":"ref_32","doi-asserted-by":"crossref","unstructured":"Lifshitz, I., Fetaya, E., and Ullman, S. (2016, January 8\u201316). Human pose estimation using deep consensus voting. Proceedings of the 2016 European Conference on Computer Vision (ECCV), Amsterdam, Netherlands.","DOI":"10.1007\/978-3-319-46475-6_16"},{"key":"ref_33","unstructured":"Tang, Z., Peng, X., Geng, S., Zhu, Y., and Metaxas, D. (2018, January 3\u20136). CU-Net: Coupled U-Nets. Proceedings of the 2018 British Machine Vision Conference (BMVC), Newcastle, UK."},{"key":"ref_34","doi-asserted-by":"crossref","unstructured":"Tang, Z., Peng, X., Geng, S., Wu, L., Zhang, S., and Metaxas, D. (2018, January 8\u201314). Quantized Densely Connected U-Nets for Efficient Landmark Localizetion. Proceedings of the 2018 European Conference on Computer Vision (ECCV), Munich, Germany.","DOI":"10.1007\/978-3-030-01219-9_21"},{"key":"ref_35","doi-asserted-by":"crossref","unstructured":"He, K., Zhang, X., Ren, S., and Sun, J. (2015, January 8\u201310). Deep Residual Learning for Image Recognition. Proceedings of the 2015 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Boston, MA, USA.","DOI":"10.1109\/CVPR.2016.90"},{"key":"ref_36","unstructured":"Loffe, S., and Szegedy, C. (2015, January 8\u201310). Batch Normalization: Accelerating Deep Network Traing by Reducing Internal Covariate Shift. Proceedings of the 2015 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Boston, MA, USA."},{"key":"ref_37","doi-asserted-by":"crossref","unstructured":"Szegedy, C., Vanhoucke, V., Loffe, S., Shlens, J., and Wojna, Z. (2015, January 8\u201310). Rethinking the Inception Architecture for Computer Vision. Proceedings of the 2015 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Boston, MA, USA.","DOI":"10.1109\/CVPR.2016.308"}],"container-title":["Sensors"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1424-8220\/19\/3\/718\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,11]],"date-time":"2025-10-11T12:30:59Z","timestamp":1760185859000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1424-8220\/19\/3\/718"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2019,2,10]]},"references-count":37,"journal-issue":{"issue":"3","published-online":{"date-parts":[[2019,2]]}},"alternative-id":["s19030718"],"URL":"https:\/\/doi.org\/10.3390\/s19030718","relation":{},"ISSN":["1424-8220"],"issn-type":[{"type":"electronic","value":"1424-8220"}],"subject":[],"published":{"date-parts":[[2019,2,10]]}}}