{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,31]],"date-time":"2026-07-31T15:20:55Z","timestamp":1785511255050,"version":"3.56.0"},"reference-count":59,"publisher":"MDPI AG","issue":"2","license":[{"start":{"date-parts":[[2023,1,6]],"date-time":"2023-01-06T00:00:00Z","timestamp":1672963200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["62271409"],"award-info":[{"award-number":["62271409"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["62171381"],"award-info":[{"award-number":["62171381"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["62271409"],"award-info":[{"award-number":["62271409"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["62171381"],"award-info":[{"award-number":["62171381"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Remote Sensing"],"abstract":"<jats:p>Remote-vision-based image processing plays a vital role in the safety helmet and harness monitoring of construction sites, in which computer-vision-based automatic safety helmet and harness monitoring systems have attracted significant attention for practical applications. However, many problems have not been well solved in existing computer-vision-based systems, such as the shortage of safety helmet and harness monitoring datasets and the low accuracy of the detection algorithms. To address these issues, an attribute-knowledge-modeling-based safety helmet and harness monitoring system is constructed in this paper, which elegantly transforms safety state recognition into images\u2019 semantic attribute recognition. Specifically, a novel transformer-based end-to-end network with a self-attention mechanism is proposed to improve attribute recognition performance by making full use of the correlations between image features and semantic attributes, based on which a security recognition system is constructed by integrating detection, tracking, and attribute recognition. Experimental results for safety helmet and harness detection demonstrate that the accuracy and robustness of the proposed transformer-based attribute recognition algorithm obviously outperforms the state-of-the-art algorithms, and the presented system is robust to challenges such as pose variation, occlusion, and a cluttered background.<\/jats:p>","DOI":"10.3390\/rs15020347","type":"journal-article","created":{"date-parts":[[2023,1,9]],"date-time":"2023-01-09T04:47:08Z","timestamp":1673239628000},"page":"347","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":16,"title":["A Remote-Vision-Based Safety Helmet and Harness Monitoring System Based on Attribute Knowledge Modeling"],"prefix":"10.3390","volume":"15","author":[{"given":"Xiao","family":"Wu","sequence":"first","affiliation":[{"name":"School of Electronic and Information, Northwestern Polytechnical University, Xi\u2019an 710129, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-1429-5009","authenticated-orcid":false,"given":"Yupeng","family":"Li","sequence":"additional","affiliation":[{"name":"School of Electronic and Information, Northwestern Polytechnical University, Xi\u2019an 710129, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Jihui","family":"Long","sequence":"additional","affiliation":[{"name":"School of Electronic and Information, Northwestern Polytechnical University, Xi\u2019an 710129, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-3380-8957","authenticated-orcid":false,"given":"Shun","family":"Zhang","sequence":"additional","affiliation":[{"name":"School of Electronic and Information, Northwestern Polytechnical University, Xi\u2019an 710129, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Shuai","family":"Wan","sequence":"additional","affiliation":[{"name":"School of Electronic and Information, Northwestern Polytechnical University, Xi\u2019an 710129, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-8018-596X","authenticated-orcid":false,"given":"Shaohui","family":"Mei","sequence":"additional","affiliation":[{"name":"School of Electronic and Information, Northwestern Polytechnical University, Xi\u2019an 710129, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1968","published-online":{"date-parts":[[2023,1,6]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","first-page":"1040","DOI":"10.1080\/13467581.2021.1877141","article-title":"Analysis of safety risk factors of modular construction to identify accident trends","volume":"21","author":"Jeong","year":"2022","journal-title":"J. Asian Archit. Build. Eng."},{"key":"ref_2","unstructured":"(2019, July 06). OSHA, Available online: https:\/\/www.osha.gov\/Publications\/OSHA3252\/3252.html."},{"key":"ref_3","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1007\/s11432-020-3102-9","article-title":"Learning hyperspectral images from RGB images via a coarse-to-fine CNN","volume":"65","author":"Mei","year":"2022","journal-title":"Sci. China Inf. Sci."},{"key":"ref_4","doi-asserted-by":"crossref","first-page":"107458","DOI":"10.1016\/j.compeleceng.2021.107458","article-title":"Method based on the cross-layer attention mechanism and multiscale perception for safety helmet-wearing detection","volume":"95","author":"Han","year":"2021","journal-title":"Comput. Electr. Eng."},{"key":"ref_5","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1109\/TIM.2022.3218574","article-title":"Toward Efficient Safety Helmet Detection Based on YoloV5 With Hierarchical Positive Sample Selection and Box Density Filtering","volume":"71","author":"Li","year":"2022","journal-title":"IEEE Trans. Instrum. Meas."},{"key":"ref_6","doi-asserted-by":"crossref","first-page":"103085","DOI":"10.1016\/j.autcon.2020.103085","article-title":"Deep learning for site safety: Real-time detection of personal protective equipment","volume":"112","author":"Nath","year":"2020","journal-title":"Autom. Constr."},{"key":"ref_7","doi-asserted-by":"crossref","first-page":"102894","DOI":"10.1016\/j.autcon.2019.102894","article-title":"Automatic detection of hardhats worn by construction personnel: A deep learning approach and benchmark dataset","volume":"106","author":"Wu","year":"2019","journal-title":"Autom. Constr."},{"key":"ref_8","doi-asserted-by":"crossref","unstructured":"Fang, C., Xiang, H., Leng, C., Chen, J., and Yu, Q. (2022). Research on Real-Time Detection of Safety Harness Wearing of Workshop Personnel Based on YOLOv5 and OpenPose. Sustainability, 14.","DOI":"10.3390\/su14105872"},{"key":"ref_9","doi-asserted-by":"crossref","first-page":"53","DOI":"10.1016\/j.autcon.2018.02.018","article-title":"Falls from heights: A computer vision-based approach for safety harness detection","volume":"91","author":"Fang","year":"2018","journal-title":"Autom. Constr."},{"key":"ref_10","doi-asserted-by":"crossref","first-page":"108220","DOI":"10.1016\/j.patcog.2021.108220","article-title":"Pedestrian attribute recognition: A survey","volume":"121","author":"Wang","year":"2022","journal-title":"Pattern Recognit."},{"key":"ref_11","doi-asserted-by":"crossref","first-page":"439","DOI":"10.1016\/j.aei.2012.02.011","article-title":"Real-time construction worker posture analysis for ergonomics training","volume":"26","author":"Ray","year":"2012","journal-title":"Adv. Eng. Inform."},{"key":"ref_12","doi-asserted-by":"crossref","first-page":"239","DOI":"10.1016\/j.aei.2015.02.001","article-title":"Computer vision techniques for construction safety and health monitoring","volume":"29","author":"Seo","year":"2015","journal-title":"Adv. Eng. Inform."},{"key":"ref_13","doi-asserted-by":"crossref","first-page":"2","DOI":"10.1016\/j.autcon.2016.11.007","article-title":"Wearable IMU-based real-time motion warning system for construction workers\u2019 musculoskeletal disorders prevention","volume":"74","author":"Yan","year":"2017","journal-title":"Autom. Constr."},{"key":"ref_14","doi-asserted-by":"crossref","first-page":"424","DOI":"10.1016\/j.apergo.2017.03.016","article-title":"An evaluation of wearable sensor s and their placements for analyzing construction worker\u2019s trunk posture i n laboratory conditions","volume":"65","author":"Wonil","year":"2017","journal-title":"Appl. Erg."},{"key":"ref_15","doi-asserted-by":"crossref","first-page":"58","DOI":"10.1016\/j.autcon.2018.01.003","article-title":"Transfer learning and deep convolutional neural networks for safety guardrail detection in 2D images","volume":"89","author":"Kolar","year":"2018","journal-title":"Autom. Constr."},{"key":"ref_16","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1016\/j.autcon.2017.09.018","article-title":"Detecting non-hardhat-use by a deep learning method from far-field surveillance videos","volume":"85","author":"Fang","year":"2018","journal-title":"Autom. Constr."},{"key":"ref_17","doi-asserted-by":"crossref","first-page":"166603","DOI":"10.1109\/ACCESS.2021.3135662","article-title":"A novel implementation of an ai-based smart construction safety inspection protocol in the uae","volume":"9","author":"Shanti","year":"2021","journal-title":"IEEE Access"},{"key":"ref_18","doi-asserted-by":"crossref","unstructured":"Alrayes, F.S., Alotaibi, S.S., Alissa, K.A., Maashi, M., Alhogail, A., Alotaibi, N., Mohsen, H., and Motwakel, A. (2022). Artificial Intelligence-Based Secure Communication and Classification for Drone-Enabled Emergency Monitoring Systems. Drones, 6.","DOI":"10.3390\/drones6090222"},{"key":"ref_19","doi-asserted-by":"crossref","first-page":"364","DOI":"10.1016\/j.jsr.2022.09.011","article-title":"Real-time monitoring of work-at-height safety hazards in construction sites using drones and deep learning","volume":"83","author":"Shanti","year":"2022","journal-title":"J. Saf. Res."},{"key":"ref_20","doi-asserted-by":"crossref","unstructured":"Zhu, J., Liao, S., Lei, Z., Yi, D., and Li, S. (2013, January 2\u20138). Pedestrian attribute classification in surveillance: Database and evaluation. Proceedings of the IEEE International Conference on Computer Vision Workshops, Sydney, NSW, Australia.","DOI":"10.1109\/ICCVW.2013.51"},{"key":"ref_21","doi-asserted-by":"crossref","unstructured":"Deng, Y., Luo, P., Loy, C.C., and Tang, X. (2014, January 3\u20137). Pedestrian attribute recognition at far distance. Proceedings of the 22nd ACM International Conference on Multimedia, New York, NY, USA.","DOI":"10.1145\/2647868.2654966"},{"key":"ref_22","unstructured":"Zhao, X., Sang, L., Ding, G., Han, J., Di, N., and Yan, C. (February, January 27). Recurrent attention model for pedestrian attribute recognition. Proceedings of the AAAI Conference on Artificial Intelligence, Honolulu, HI, USA."},{"key":"ref_23","doi-asserted-by":"crossref","first-page":"6126","DOI":"10.1109\/TIP.2019.2919199","article-title":"Attention-based pedestrian attribute analysis","volume":"28","author":"Tan","year":"2019","journal-title":"IEEE Trans. Image Process."},{"key":"ref_24","unstructured":"Tang, C., Sheng, L., Zhang, Z., and Hu, X. (November, January 27). Improving pedestrian attribute recognition with weakly-supervised multi-scale attribute-specific localization. Proceedings of the IEEE\/CVF International Conference on Computer Vision, Seoul, Republic of Korea."},{"key":"ref_25","doi-asserted-by":"crossref","unstructured":"Liu, X., Zhao, H., Tian, M., Sheng, L., Shao, J., Yi, S., Yan, J., and Wang, X. (2017, January 22\u201329). Hydraplus-net: Attentive deep features for pedestrian analysis. Proceedings of the IEEE International Conference on Computer Vision, Venice, Italy.","DOI":"10.1109\/ICCV.2017.46"},{"key":"ref_26","doi-asserted-by":"crossref","unstructured":"Li, D., Chen, X., and Huang, K. (2015, January 3\u20136). Multi-attribute learning for pedestrian attribute recognition in surveillance scenarios. Proceedings of the 2015 3rd IAPR Asian Conference on Pattern Recognition (ACPR), Kuala Lumpur, Malaysia.","DOI":"10.1109\/ACPR.2015.7486476"},{"key":"ref_27","first-page":"5509612","article-title":"Hyperspectral image classification using attention-based bidirectional long short-term memory network","volume":"60","author":"Mei","year":"2021","journal-title":"IEEE Trans. Geosci. Remote Sens."},{"key":"ref_28","first-page":"5502012","article-title":"Accelerating convolutional neural network-based hyperspectral image classification by step activation quantization","volume":"60","author":"Mei","year":"2021","journal-title":"IEEE Trans. Geosci. Remote Sens."},{"key":"ref_29","unstructured":"Sarfraz, M.S., Schumann, A., Wang, Y., and Stiefelhagen, R. (2017). Deep view-sensitive pedestrian attribute inference in an end-to-end model. arXiv."},{"key":"ref_30","doi-asserted-by":"crossref","unstructured":"Wang, J., Zhu, X., Gong, S., and Li, W. (2017, January 22\u201329). Attribute recognition by joint recurrent learning of context and correlation. Proceedings of the IEEE International Conference on Computer Vision, Venice, Italy.","DOI":"10.1109\/ICCV.2017.65"},{"key":"ref_31","doi-asserted-by":"crossref","unstructured":"Zhao, X., Sang, L., Ding, G., Guo, Y., and Jin, X. (2018, January 13\u201319). Grouping attribute recognition for pedestrian with joint recurrent learning. Proceedings of the IJCAI, Stockholm, Sweden.","DOI":"10.24963\/ijcai.2018\/441"},{"key":"ref_32","unstructured":"Li, Q., Zhao, X., He, R., and Huang, K. (February, January 27). Visual-semantic graph reasoning for pedestrian attribute recognition. Proceedings of the AAAI conference on artificial intelligence, Honolulu, HI, USA."},{"key":"ref_33","doi-asserted-by":"crossref","unstructured":"Tan, Z., Yang, Y., Wan, J., Guo, G., and Li, S.Z. (2020, January 7\u201312). Relation-aware pedestrian attribute recognition with graph convolutional networks. Proceedings of the AAAI Conference on Artificial Intelligence, New York, NY, USA.","DOI":"10.1609\/aaai.v34i07.6883"},{"key":"ref_34","unstructured":"Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A.N., Kaiser, \u0141., and Polosukhin, I. (2017). Attention is all you need. Adv. Neural Inf. Process. Syst., 30."},{"key":"ref_35","unstructured":"Devlin, J., Chang, M.W., Lee, K., and Toutanova, K. (2018). Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv."},{"key":"ref_36","unstructured":"Lan, Z., Chen, M., Goodman, S., Gimpel, K., Sharma, P., and Soricut, R. (2019). Albert: A lite bert for self-supervised learning of language representations. arXiv."},{"key":"ref_37","doi-asserted-by":"crossref","unstructured":"Dai, Z., Yang, Z., Yang, Y., Carbonell, J., Le, Q.V., and Salakhutdinov, R. (2019). Transformer-xl: Attentive language models beyond a fixed-length context. arXiv.","DOI":"10.18653\/v1\/P19-1285"},{"key":"ref_38","first-page":"1877","article-title":"Language models are few-shot learners","volume":"33","author":"Brown","year":"2020","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_39","unstructured":"Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., and Gelly, S. (2020). An image is worth 16 \u00d7 16 words: Transformers for image recognition at scale. arXiv."},{"key":"ref_40","doi-asserted-by":"crossref","unstructured":"Carion, N., Massa, F., Synnaeve, G., Usunier, N., Kirillov, A., and Zagoruyko, S. (2020). End-to-end object detection with transformers. European Conference on Computer Vision, Springer.","DOI":"10.1007\/978-3-030-58452-8_13"},{"key":"ref_41","doi-asserted-by":"crossref","unstructured":"He, S., Luo, H., Wang, P., Wang, F., Li, H., and Jiang, W. Transreid: Transformer-based object re-identification. Proceedings of the Proceedings of the IEEE\/CVF International Conference on Computer Vision, Virtual Conference, 10\u201317 October 2021.","DOI":"10.1109\/ICCV48922.2021.01474"},{"key":"ref_42","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/3505244","article-title":"Transformers in vision: A survey","volume":"54","author":"Khan","year":"2022","journal-title":"ACM Comput. Surv. (CSUR)"},{"key":"ref_43","doi-asserted-by":"crossref","unstructured":"Gabeur, V., Sun, C., Alahari, K., and Schmid, C. (2020). Multi-modal transformer for video retrieval. European Conference on Computer Vision, Springer.","DOI":"10.1007\/978-3-030-58548-8_13"},{"key":"ref_44","doi-asserted-by":"crossref","unstructured":"Cornia, M., Stefanini, M., Baraldi, L., and Cucchiara, R. (2020, January 13\u201319). Meshed-memory transformer for image captioning. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA.","DOI":"10.1109\/CVPR42600.2020.01059"},{"key":"ref_45","unstructured":"Chen, S., Hong, Z., Liu, Y., Xie, G.S., Sun, B., Li, H., Peng, Q., Lu, K., and You, X. (March, January 22). Transzero: Attribute-guided transformer for zero-shot learning. Proceedings of the AAAI, Virtual Conference."},{"key":"ref_46","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1186\/s40537-019-0197-0","article-title":"A survey on image data augmentation for deep learning","volume":"6","author":"Shorten","year":"2019","journal-title":"J. Big Data"},{"key":"ref_47","doi-asserted-by":"crossref","unstructured":"He, K., Zhang, X., Ren, S., and Sun, J. (2016, January 27\u201330). Deep residual learning for image recognition. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.90"},{"key":"ref_48","doi-asserted-by":"crossref","unstructured":"Deng, J., Dong, W., Socher, R., Li, L.J., Li, K., and Fei-Fei, L. (2009, January 20\u201325). Imagenet: A large-scale hierarchical image database. Proceedings of the 2009 IEEE Conference on Computer Vision and Pattern Recognition, Miami, FL, USA.","DOI":"10.1109\/CVPR.2009.5206848"},{"key":"ref_49","unstructured":"Ba, J.L., Kiros, J.R., and Hinton, G.E. (2016). Layer normalization. arXiv."},{"key":"ref_50","unstructured":"Kingma, D.P., and Ba, J. (2014). Adam: A method for stochastic optimization. arXiv."},{"key":"ref_51","doi-asserted-by":"crossref","unstructured":"Lin, T.Y., Maire, M., Belongie, S., Hays, J., Perona, P., Ramanan, D., Doll\u00e1r, P., and Zitnick, C.L. (2014). Microsoft coco: Common objects in context. European Conference on Computer Vision, Springer.","DOI":"10.1007\/978-3-319-10602-1_48"},{"key":"ref_52","doi-asserted-by":"crossref","unstructured":"Padilla, R., Netto, S.L., and Da Silva, E.A. (2020, January 1\u20133). A survey on performance metrics for object-detection algorithms. Proceedings of the 2020 International Conference on Systems, Signals and Image Processing (IWSSIP), Niteroi, Brazil.","DOI":"10.1109\/IWSSIP48289.2020.9145130"},{"key":"ref_53","first-page":"580","article-title":"Multi-target tracking by learning local-to-global trajectory models","volume":"48","author":"Zhang","year":"2015","journal-title":"PR"},{"key":"ref_54","doi-asserted-by":"crossref","first-page":"96","DOI":"10.1007\/s11263-019-01212-1","article-title":"Tracking persons-of-interest via unsupervised representation adaptation","volume":"128","author":"Zhang","year":"2020","journal-title":"Int. J. Comput. Vis."},{"key":"ref_55","doi-asserted-by":"crossref","unstructured":"Ristani, E., Solera, F., Zou, R., Cucchiara, R., and Tomasi, C. (2016). Performance measures and a data set for multi-target, multi-camera tracking. European Conference on Computer Vision, Springer.","DOI":"10.1007\/978-3-319-48881-3_2"},{"key":"ref_56","doi-asserted-by":"crossref","first-page":"1575","DOI":"10.1109\/TIP.2018.2878349","article-title":"A richly annotated pedestrian dataset for person retrieval in real surveillance scenarios","volume":"28","author":"Li","year":"2018","journal-title":"IEEE Trans. Image Process."},{"key":"ref_57","doi-asserted-by":"crossref","unstructured":"Zagoruyko, S., and Komodakis, N. (2016). Wide residual networks. arXiv.","DOI":"10.5244\/C.30.87"},{"key":"ref_58","doi-asserted-by":"crossref","unstructured":"Huang, L., Wang, W., Chen, J., and Wei, X.Y. (2019, January 16\u201320). Attention on attention for image captioning. Proceedings of the IEEE\/CVF International Conference on Computer Vision, Long Beach, CA, USA.","DOI":"10.1109\/ICCV.2019.00473"},{"key":"ref_59","doi-asserted-by":"crossref","unstructured":"Pan, Y., Yao, T., Li, Y., and Mei, T. (2020, January 14\u201319). X-linear attention networks for image captioning. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Virtual Conference.","DOI":"10.1109\/CVPR42600.2020.01098"}],"container-title":["Remote Sensing"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/2072-4292\/15\/2\/347\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,10]],"date-time":"2025-10-10T18:01:35Z","timestamp":1760119295000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/2072-4292\/15\/2\/347"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,1,6]]},"references-count":59,"journal-issue":{"issue":"2","published-online":{"date-parts":[[2023,1]]}},"alternative-id":["rs15020347"],"URL":"https:\/\/doi.org\/10.3390\/rs15020347","relation":{},"ISSN":["2072-4292"],"issn-type":[{"value":"2072-4292","type":"electronic"}],"subject":[],"published":{"date-parts":[[2023,1,6]]}}}