{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,12,20]],"date-time":"2025-12-20T22:26:35Z","timestamp":1766269595399,"version":"build-2065373602"},"reference-count":34,"publisher":"MDPI AG","issue":"21","license":[{"start":{"date-parts":[[2022,10,30]],"date-time":"2022-10-30T00:00:00Z","timestamp":1667088000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"name":"National Natural Science Foundation of China","award":["61877038","U2001205","20200201"],"award-info":[{"award-number":["61877038","U2001205","20200201"]}]},{"name":"Open Subjects of National Engineering Laboratory for Integrated Aero-Space-Ground-Ocean Big Data Application Technology, China","award":["61877038","U2001205","20200201"],"award-info":[{"award-number":["61877038","U2001205","20200201"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Sensors"],"abstract":"<jats:p>Solid developments have been seen in deep-learning-based pose estimation, but few works have explored performance in dense crowds, such as a classroom scene; furthermore, no specific knowledge is considered in the design of image augmentation for pose estimation. A masked autoencoder was shown to have a non-negligible capability in image reconstruction, where the masking mechanism that randomly drops patches forces the model to build unknown pixels from known pixels. Inspired by this self-supervised learning method, where the restoration of the feature loss induced by the mask is consistent with tackling the occlusion problem in classroom scenarios, we discovered that the transfer performance of the pre-trained weights could be used as a model-based augmentation to overcome the intractable occlusion in classroom pose estimation. In this study, we proposed a top-down pose estimation method that utilized the natural reconstruction capability of missing information of the MAE as an effective occluded image augmentation in a pose estimation task. The difference with the original MAE was that instead of using a 75% random mask ratio, we regarded the keypoint distribution probabilistic heatmap as a reference for masking, which we named Pose Mask. To test the performance of our method in heavily occluded classroom scenes, we collected a new dataset for pose estimation in classroom scenes named Class Pose and conducted many experiments, the results of which showed promising performance.<\/jats:p>","DOI":"10.3390\/s22218331","type":"journal-article","created":{"date-parts":[[2022,10,30]],"date-time":"2022-10-30T10:47:57Z","timestamp":1667126877000},"page":"8331","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":2,"title":["Pose Mask: A Model-Based Augmentation Method for 2D Pose Estimation in Classroom Scenes Using Surveillance Images"],"prefix":"10.3390","volume":"22","author":[{"given":"Shichang","family":"Liu","sequence":"first","affiliation":[{"name":"School of Computer Science, Shaanxi Normal University, Xi\u2019an 710119, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Miao","family":"Ma","sequence":"additional","affiliation":[{"name":"School of Computer Science, Shaanxi Normal University, Xi\u2019an 710119, China"},{"name":"Key Laboratory of Modern Teaching Technology, Ministry of Education, Xi\u2019an 710062, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Haiyang","family":"Li","sequence":"additional","affiliation":[{"name":"School of Computer Science, Shaanxi Normal University, Xi\u2019an 710119, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Hanyang","family":"Ning","sequence":"additional","affiliation":[{"name":"School of Computer Science, Shaanxi Normal University, Xi\u2019an 710119, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Min","family":"Wang","sequence":"additional","affiliation":[{"name":"National Engineering Laboratory for Integrated Aero-Space-Ground-Ocean Big Data Application Technology, Xi\u2019an 710072, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2022,10,30]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","unstructured":"Liu, M.Y., and Yuan, J.S. (2018, January 18\u201322). Recognizing human actions as the evolution of pose estimation maps. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00127"},{"key":"ref_2","doi-asserted-by":"crossref","first-page":"1675","DOI":"10.1007\/s11263-021-01446-y","article-title":"Mimetics: Towards understanding human actions out of context","volume":"129","author":"Philippe","year":"2021","journal-title":"IJCV"},{"key":"ref_3","first-page":"103055","article-title":"Human pose estimation and its application to action recognition: A survey","volume":"76","author":"Song","year":"2021","journal-title":"JVCIR"},{"key":"ref_4","doi-asserted-by":"crossref","unstructured":"Andriluka, M., Pishchulin, L., Gehler, P., and Schiele, B. (2014, January 23\u201328). 2D human pose estimation: New benchmark and state of the art analysis. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Columbus, OH, USA.","DOI":"10.1109\/CVPR.2014.471"},{"key":"ref_5","doi-asserted-by":"crossref","unstructured":"Bulat, A., Kossaifi, J., Tzimiropoulos, G., and Pantic, M. (2020, January 16\u201320). Toward fast and accurate human pose estimation via soft-gated skip connections. Proceedings of the IEEE International Conference on Automatic Face and Gesture Recognition, Buenos Aires, Argentina.","DOI":"10.1109\/FG47880.2020.00014"},{"key":"ref_6","doi-asserted-by":"crossref","unstructured":"Lin, T.Y., Maire, M., and Belongie, S. (2014, January 5\u201312). Microsoft coco: Common objects in context. Proceedings of the European Conference on Computer Vision, Zurich, Switzerland.","DOI":"10.1007\/978-3-319-10602-1_48"},{"key":"ref_7","doi-asserted-by":"crossref","unstructured":"Lin, F.C., Ngo, H.H., and Dow, C.R. (2021). Student behavior recognition system for the classroom environment based on skeleton pose estimation and person detection. Sensors, 21.","DOI":"10.3390\/s21165314"},{"key":"ref_8","unstructured":"(2022, September 22). GitHub. Available online: https:\/\/github.com\/CMU-Perceptual-Computing-Lab\/openpose."},{"key":"ref_9","doi-asserted-by":"crossref","unstructured":"Xu, X., and Xin, T. (2020, January 5\u20138). Classroom attention analysis based on multiple euler angles constraint and head pose estimation. Proceedings of the International Conference on Multimedia Modeling, Daejeon Metropolitan City, Korea.","DOI":"10.1007\/978-3-030-37731-1_27"},{"key":"ref_10","doi-asserted-by":"crossref","unstructured":"Yu, Z., Li, Y., and Liu, Y. (2022, January 22\u201327). Synpose: A Large-Scale and Densely Annotated Synthetic Dataset for Human Pose Estimation in Classroom. Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing, Marina Bay Sands Expo & Convention Center, Singapore.","DOI":"10.1109\/ICASSP43922.2022.9747453"},{"key":"ref_11","doi-asserted-by":"crossref","first-page":"139","DOI":"10.1145\/3422622","article-title":"Generative adversarial networks","volume":"63","author":"Goodfellow","year":"2020","journal-title":"Commun. ACM"},{"key":"ref_12","doi-asserted-by":"crossref","unstructured":"Toshev, A., and Szegedy, C. (2014, January 23\u201328). Deeppose: Human pose estimation via deep neural networks. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Columbus, OH, USA.","DOI":"10.1109\/CVPR.2014.214"},{"key":"ref_13","unstructured":"Nie, X., Feng, J., Zhang, J., and Yan, S. (November, January 27). Single-stage multi-person pose machines. Proceedings of the IEEE\/CVF International Conference on Computer Vision, Seoul, Korea."},{"key":"ref_14","doi-asserted-by":"crossref","unstructured":"Newell, A., Yang, K., and Deng, J. (2016, January 8\u201316). Stacked hourglass networks for human pose estimation. Proceedings of the European Conference on Computer Vision, Amsterdam, The Netherlands.","DOI":"10.1007\/978-3-319-46484-8_29"},{"key":"ref_15","doi-asserted-by":"crossref","unstructured":"Xiao, B., Wu, H., and Wei, Y. (2018, January 8\u201314). Simple baselines for human pose estimation and tracking. Proceedings of the European Conference on Computer Vision, Munich, Germany.","DOI":"10.1007\/978-3-030-01231-1_29"},{"key":"ref_16","doi-asserted-by":"crossref","unstructured":"Sun, K., Xiao, B., Liu, D., and Wang, J. (2019, January 16\u201320). Deep high-resolution representation learning for human pose estimation. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.00584"},{"key":"ref_17","unstructured":"Gidaris, S., Singh, P., and Komodakis, N. (May, January 30). Unsupervised representation learning by predicting image rotations. Proceedings of the International Conference on Learning Representations, Vancouver, BC, Canada."},{"key":"ref_18","doi-asserted-by":"crossref","unstructured":"He, K., Chen, X., Xie, S., Li, Y., Doll\u00e1r, P., and Girshick, R. (2022, January 16\u201324). Masked autoencoders are scalable vision learners. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, New Orleans, LA, USA.","DOI":"10.1109\/CVPR52688.2022.01553"},{"key":"ref_19","unstructured":"Devlin, J., Chang, M.W., Lee, K., and Toutanova, K. (2018). Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv."},{"key":"ref_20","unstructured":"Xu, H., Ding, S., Zhang, X., Xiong, H., and Tian, Q. (2022). Masked Autoencoders are Robust Data Augmentors. arXiv."},{"key":"ref_21","unstructured":"Xu, Y., Zhang, J., Zhang, Q., and Tao, D. (2022). ViTPose: Simple vision transformer baselines for human pose estimation. arXiv."},{"key":"ref_22","unstructured":"Dosovitskiy, A., Beyer, L., and Kolesnikov, A. (2020). An image is worth 16x16 words: Transformers for image recognition at scale. arXiv."},{"key":"ref_23","unstructured":"He, K., Zhang, X., Ren, S., and Sun, J. (July, January 26). Deep residual learning for image recognition. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA."},{"key":"ref_24","unstructured":"Yuan, Y., Fu, R., and Huang, L. (2021, January 6\u201314). Hrformer: High-resolution vision transformer for dense predict. Proceedings of the 2021 Conference on Neural Information Processing Systems, Vertual."},{"key":"ref_25","unstructured":"Dai, Z., Liu, H., Le, Q.V., and Tan, M. (2021). Coatnet: Marrying convolution and attention for all data sizes. arXiv."},{"key":"ref_26","unstructured":"Xiao, T., Singh, M., Mintun, M., Darrell, T., Doll\u00e1r, P., and Girshick, R. (2021). Early convolutions help transformers see better. arXiv."},{"key":"ref_27","doi-asserted-by":"crossref","unstructured":"Cai, Z., and Vasconcelos, N. (2018, January 18\u201322). Cascade r-cnn: Delving into high quality object detection. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00644"},{"key":"ref_28","doi-asserted-by":"crossref","unstructured":"Liu, Z., Mao, H., Wu, C.Y., Feichtenhofer, C., Darrell, T., and Xie, S. (2022, January 16\u201324). A convnet for the 2020s. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, New Orleans, LA, USA.","DOI":"10.1109\/CVPR52688.2022.01167"},{"key":"ref_29","unstructured":"Chen, K., Wang, J., and Pang, J. (2019). MMDetection: Open mmlab detection toolbox and benchmark. arXiv."},{"key":"ref_30","doi-asserted-by":"crossref","unstructured":"Zhang, H., Wu, C., and Zhang, Z. (2022, January 16\u201324). Resnest: Split-attention networks. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, New Orleans, LA, USA.","DOI":"10.1109\/CVPRW56347.2022.00309"},{"key":"ref_31","doi-asserted-by":"crossref","first-page":"652","DOI":"10.1109\/TPAMI.2019.2938758","article-title":"Res2net: A new multi-scale backbone architecture","volume":"43","author":"Gao","year":"2019","journal-title":"T-PAMI"},{"key":"ref_32","doi-asserted-by":"crossref","unstructured":"Zhang, F., Zhu, X., Dai, H., Ye, M., and Zhu, C. (2020, January 14\u201319). Distribution-aware coordinate representation for human pose estimation. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA.","DOI":"10.1109\/CVPR42600.2020.00712"},{"key":"ref_33","doi-asserted-by":"crossref","unstructured":"Liu, H., Liu, F., Fan, X., and Huang, D. (2021). Polarized self-attention: Towards high-quality pixel-wise regression. arXiv.","DOI":"10.1016\/j.neucom.2022.07.054"},{"key":"ref_34","doi-asserted-by":"crossref","unstructured":"Zhu, J.Y., Park, T., Isola, P., and Efros, A.A. (2017, January 12\u201314). Unpaired image-to-image translation using cycle-consistent adversarial networks. Proceedings of the IEEE International Conference on Computer Vision, Venice, Italy.","DOI":"10.1109\/ICCV.2017.244"}],"container-title":["Sensors"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1424-8220\/22\/21\/8331\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,11]],"date-time":"2025-10-11T01:06:04Z","timestamp":1760144764000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1424-8220\/22\/21\/8331"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,10,30]]},"references-count":34,"journal-issue":{"issue":"21","published-online":{"date-parts":[[2022,11]]}},"alternative-id":["s22218331"],"URL":"https:\/\/doi.org\/10.3390\/s22218331","relation":{},"ISSN":["1424-8220"],"issn-type":[{"type":"electronic","value":"1424-8220"}],"subject":[],"published":{"date-parts":[[2022,10,30]]}}}