{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2024,7,25]],"date-time":"2024-07-25T11:10:35Z","timestamp":1721905835318},"reference-count":41,"publisher":"National Library of Serbia","issue":"2","license":[{"start":{"date-parts":[[2023,1,1]],"date-time":"2023-01-01T00:00:00Z","timestamp":1672531200000},"content-version":"unspecified","delay-in-days":0,"URL":"http:\/\/creativecommons.org\/licenses\/by-nc-nd\/4.0\/"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["ComSIS","COMPUT SCI INF SYST","COMPUT SCI INFORM SY","COMPUTER SCI INFORM","COMSIS J"],"published-print":{"date-parts":[[2023]]},"abstract":"<jats:p>Recognizing pedestrian attributes has recently obtained increasing attention due to its great potential in person re-identification, recommendation system, and other applications. Existing methods have achieved good results, but these methods do not fully utilize region information and the correlation between attributes. This paper aims at proposing a robust pedestrian attribute recognition framework. Specifically, we first propose an end-to-end framework for attribute recognition. Secondly, spatial and semantic self-attention mechanism is used for key points localization and bounding boxes generation. Finally, a hierarchical recognition strategy is proposed, the whole region is used for the global attribute recognition, and the relevant regions are used for the local attribute recognition. Experimental results on two pedestrian attribute datasets PETA and RAP show that the mean recognition accuracy reaches 84.63% and 82.70%. The heatmap analysis shows that our method can effectively improve the spatial and the semantic correlation between attributes. Compared with existing methods, it can achieve better recognition effect.<\/jats:p>","DOI":"10.2298\/csis220815016f","type":"journal-article","created":{"date-parts":[[2023,3,1]],"date-time":"2023-03-01T08:28:09Z","timestamp":1677659289000},"page":"793-812","source":"Crossref","is-referenced-by-count":1,"title":["Pedestrian attribute recognition based on dual self-attention mechanism"],"prefix":"10.2298","volume":"20","author":[{"given":"Zhongkui","family":"Fan","sequence":"first","affiliation":[{"name":"School of Communication and Information Engineering, Shanghai University, Shanghai, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Ye-Peng","family":"Guan","sequence":"additional","affiliation":[{"name":"School of Communication and Information Engineering, Shanghai University, Shanghai, China + Key Laboratory of Advanced Displays and System Application, Ministry of Education, Shanghai, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1078","reference":[{"key":"ref1","doi-asserted-by":"crossref","unstructured":"Gordo A, Almazan J, Revaud J, et al. End-to-end learning of deep visual representations for image retrieval[J]. International Journal of Computer Vision, 2017, 124(2): 237-254.","DOI":"10.1007\/s11263-017-1016-8"},{"key":"ref2","doi-asserted-by":"crossref","unstructured":"Dubey S R. A decade survey of content based image retrieval using deep learning[J]. IEEE Transactions on Circuits and Systems for Video Technology, 2021.","DOI":"10.1109\/TCSVT.2021.3080920"},{"key":"ref3","doi-asserted-by":"crossref","unstructured":"Cffa B , Gg A , Jvdw A , et al. Saliency for fine-grained object recognition in domains with scarce training data[J]. Pattern Recognition, 2019, 94:62-73.","DOI":"10.1016\/j.patcog.2019.05.002"},{"key":"ref4","doi-asserted-by":"crossref","unstructured":"Liang, M., & Hu, X. (2015). Recurrent convolutional neural network for object recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition (pp. 3367-3375).","DOI":"10.1109\/CVPR.2015.7298958"},{"key":"ref5","doi-asserted-by":"crossref","unstructured":"Ding, S., Lin, L., Wang, G., & Chao, H. (2015). Deep feature learning with relative distance comparison for person re-identification. Pattern Recognition, 48(10), 2993-3003.","DOI":"10.1016\/j.patcog.2015.04.005"},{"key":"ref6","doi-asserted-by":"crossref","unstructured":"Ryan Layne, Timothy M Hospedales, Shaogang Gong, and Q Mary. Person re-identification by attributes. In Bmvc, volume 2, page 8, 2012.","DOI":"10.5244\/C.26.24"},{"key":"ref7","unstructured":"Jianqing Zhu, Shengcai Liao, Dong Yi, Zhen Lei,and Stan Z Li. Multi-label cnn based pedestrian attribute learning for soft biometrics. In ICB. IEEE, 2015."},{"key":"ref8","doi-asserted-by":"crossref","unstructured":"Yubin Deng, Ping Luo, Chen Change Loy, and Xiaoou Tang. Pedestrian attribute recognition at far distance. In ACM MM, 2014..","DOI":"10.1145\/2647868.2654966"},{"key":"ref9","doi-asserted-by":"crossref","unstructured":"Y. Deng, P. Luo, C. C. Loy, and X. Tang. Pedestrian attribute recognition at far distance. In Proc. ACM Multimedia, 2014.","DOI":"10.1145\/2647868.2654966"},{"key":"ref10","doi-asserted-by":"crossref","unstructured":"R. Layne, T. M. Hospedales, S. Gong, and Q. Mary. Person reidentification by attributes. In Proc. BMVC, 2012.","DOI":"10.5244\/C.26.24"},{"key":"ref11","doi-asserted-by":"crossref","unstructured":"J. Zhu, S. Liao, Z. Lei, D. Yi, and S. Z. Li. Pedestrian attribute classification in surveillance: Database and evaluation. In Proc. ICCV Workshops, 2013.","DOI":"10.1109\/ICCVW.2013.51"},{"key":"ref12","doi-asserted-by":"crossref","unstructured":"D. Li, X. Chen, and K. Huang, \u201cMulti-attribute learning for pedestrian attribute recognition in surveillance scenarios,\u201d in Pattern Recognition (ACPR), 2015 3rd IAPR Asian Conference on. IEEE, 2015, pp. 111-115.","DOI":"10.1109\/ACPR.2015.7486476"},{"key":"ref13","doi-asserted-by":"crossref","unstructured":"D. Li, X. Chen, Z. Zhang, and K. Huang, \u201cPose guided deep model for pedestrian attribute recognition in surveillance scenarios,\u201d in 2018 IEEE International Conference on Multimedia and Expo (ICME). IEEE, 2018, pp. 1-6.","DOI":"10.1109\/ICME.2018.8486604"},{"key":"ref14","doi-asserted-by":"crossref","unstructured":"N. Zhang, M. Paluri, M. Ranzato, T. Darrell, and L. Bourdev. Panda: Pose aligned networks for deep attribute modeling. In Proc. CVPR,2014.","DOI":"10.1109\/CVPR.2014.212"},{"key":"ref15","doi-asserted-by":"crossref","unstructured":"L. Bourdev and J. Malik, \u201cPoselets: Body part detectors trained using 3d human pose annotations,\u201d in Computer Vision, 2009 IEEE 12th International Conference on. IEEE, 2009, pp. 1365-1372.","DOI":"10.1109\/ICCV.2009.5459303"},{"key":"ref16","doi-asserted-by":"crossref","unstructured":"D. A. Vaquero, R. S. Feris, D. Tran, L. Brown, A. Hampapur, and M. Turk. Attribute-based people search in surveillance environments. In Workshop on Applications of Computer Vision (WACV), pages 1-8, 2009.","DOI":"10.1109\/WACV.2009.5403131"},{"key":"ref17","doi-asserted-by":"crossref","unstructured":"J.Q. Zhu, S.C. Liao, Z. Lei, D. Yi, S.Z. Li. Pedestrian Attribute Classification in Surveillance: Database and Evaluation. IEEE ICCV Workshop, 2013.","DOI":"10.1109\/ICCVW.2013.51"},{"key":"ref18","doi-asserted-by":"crossref","unstructured":"J.Q. Zhu, S.c. Liao, D.Yi, Z. Lei, S.Z. Li. Multi-Label CNN Based Pedestrian Attribute Learning for Soft Biometrics. IAPR ICB, 2015.","DOI":"10.1109\/ICB.2015.7139070"},{"key":"ref19","doi-asserted-by":"crossref","unstructured":"Y.Deng, P.Luo, c.c. Loy, X.O. Tang. Pedestrian Attribute Recognition at Far Distance. ACM MM, 2014.","DOI":"10.1145\/2647868.2654966"},{"key":"ref20","doi-asserted-by":"crossref","unstructured":"W. Chen, X. Chen, J. Zhang, and K. Huang. A multi-task deep network for person re-identification. In AAAI, 2017.","DOI":"10.1609\/aaai.v31i1.11201"},{"key":"ref21","doi-asserted-by":"crossref","unstructured":"P. Sudowe, H. Spitzer, B. Leibe. Person Attribute Recognition with a Jointly-trained Holistic CNN Model. IEEE ICCV Workshop, 2015.","DOI":"10.1109\/ICCVW.2015.51"},{"key":"ref22","doi-asserted-by":"crossref","unstructured":"A. H. Abdulnabi, G. Wang, J. Lu, and K. Jia, \u201cMulti-task cnn model for attribute prediction,\u201d IEEE Transactions on Multimedia, vol. 17, no. 11, pp. 1949-1959, 2015.","DOI":"10.1109\/TMM.2015.2477680"},{"key":"ref23","unstructured":"Yao, C. , et al. \"Hierarchical pedestrian attribute recognition based on adaptive region localization.\" IEEE International Conference on Multimedia & Expo Workshops IEEE, 2017."},{"key":"ref24","unstructured":"Feng, Z. , et al. \"Learning Spatial Regularization with Image-Level Supervisions for Multi-label Image Classification.\" IEEE (2017)."},{"key":"ref25","doi-asserted-by":"crossref","unstructured":"X. Liu, H. Zhao, M. Tian, L. Sheng, J. Shao, S. Yi, J. Yan, and X. Wang,\u201cHydraplus-net: Attentive deep features for pedestrian analysis,\u201d in Proceedings of the IEEE International Conference on Computer Vision, 2017, pp. 350-359.","DOI":"10.1109\/ICCV.2017.46"},{"key":"ref26","doi-asserted-by":"crossref","unstructured":"Li, Qiaozhe, et al. \"Pedestrian Attribute Recognition by Joint Visual-semantic Reasoning and Knowledge Distillation.\" IJCAI. 2019.","DOI":"10.24963\/ijcai.2019\/117"},{"key":"ref27","doi-asserted-by":"crossref","unstructured":"Huang, Lingli. \"Improved non-local means algorithm for image denoising.\" Journal of Computer and Communications 3.04 (2015): 23.","DOI":"10.4236\/jcc.2015.34003"},{"key":"ref28","doi-asserted-by":"crossref","unstructured":"Wang, X., Girshick, R., Gupta, A., & He, K. (2018). Non-local neural networks. In Proceedings of the IEEE conference on computer vision and pattern recognition (pp. 7794-7803).","DOI":"10.1109\/CVPR.2018.00813"},{"key":"ref29","unstructured":"Zhang, Han, et al. \"Self-attention generative adversarial networks.\" International conference on machine learning. PMLR, 2019."},{"key":"ref30","unstructured":"Vaswani, Ashish, et al. \"Attention is all you need.\" Advances in neural information processing systems. 2017."},{"key":"ref31","doi-asserted-by":"crossref","unstructured":"Hu, Jie, Li Shen, and Gang Sun. \"Squeeze-and-excitation networks.\" Proceedings of the IEEE conference on computer vision and pattern recognition. 2018.","DOI":"10.1109\/CVPR.2018.00745"},{"key":"ref32","doi-asserted-by":"crossref","unstructured":"Woo, Sanghyun, et al. \"Cbam: Convolutional block attention module.\" Proceedings of the European conference on computer vision (ECCV). 2018.","DOI":"10.1007\/978-3-030-01234-2_1"},{"key":"ref33","doi-asserted-by":"crossref","unstructured":"Chen, Tianlong, et al. \"Abd-net: Attentive but diverse person re-identification.\" Proceedings of the IEEE\/CVF International Conference on Computer Vision. 2019.","DOI":"10.1109\/ICCV.2019.00844"},{"key":"ref34","doi-asserted-by":"crossref","unstructured":"Tan, Zichang, et al. \"Attention-based pedestrian attribute analysis.\" IEEE transactions on image processing 28.12 (2019): 6126-6140.","DOI":"10.1109\/TIP.2019.2919199"},{"key":"ref35","doi-asserted-by":"crossref","unstructured":"Yang, Z. , et al. \"Gated Channel Transformation for Visual Recognition.\" 2020 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR) IEEE, 2020.","DOI":"10.1109\/CVPR42600.2020.01181"},{"key":"ref36","unstructured":"Q. L. X. Z. R. H. K. HUANG, \u201cVisual-semantic graph reasoning for pedestrian attribute recognition,\u201d in Association for the Advancement of Artificial Intelligence, AAAI, 2019."},{"key":"ref37","unstructured":"P. Liu, X. Liu, J. Yan, and J. Shao, \u201cLocalization guided learning for pedestrian attribute recognition\u201d 2018."},{"key":"ref38","unstructured":"Zheng S F, Tang J, Luo B, et. al. Multistage pedestrian attribute recognition method based on improved loss function. Pattern Recognition and Artificial Intelligence.2018.31(12):1085-1095."},{"key":"ref39","doi-asserted-by":"crossref","unstructured":"Zhao X, Sang L, Ding G, et al. Recurrent attention model for pedestrian attribute recognition. AAAI Press, 2019.","DOI":"10.1609\/aaai.v33i01.33019275"},{"key":"ref40","doi-asserted-by":"crossref","unstructured":"Ji Z, He E, Wang H et al. Image-attribute reciprocally guided attention network for pedestrian attribute recognition. Pattern Recognition Letters.2019, pp. 89-95.","DOI":"10.1016\/j.patrec.2019.01.010"},{"key":"ref41","unstructured":"J. Jia, H. Huang, W. Yang, X. Chen, and K. Huang,\u201cRethinking of Pedestrian Attribute Recognition: Realistic Datasets with Efficient Method,\u201c arXiv preprint arXiv:2005.11909, 2020."}],"container-title":["Computer Science and Information Systems"],"original-title":[],"language":"en","deposited":{"date-parts":[[2024,7,25]],"date-time":"2024-07-25T10:30:20Z","timestamp":1721903420000},"score":1,"resource":{"primary":{"URL":"https:\/\/doiserbia.nb.rs\/Article.aspx?ID=1820-02142300016F"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023]]},"references-count":41,"journal-issue":{"issue":"2","published-print":{"date-parts":[[2023]]}},"URL":"https:\/\/doi.org\/10.2298\/csis220815016f","relation":{},"ISSN":["1820-0214","2406-1018"],"issn-type":[{"value":"1820-0214","type":"print"},{"value":"2406-1018","type":"electronic"}],"subject":[],"published":{"date-parts":[[2023]]}}}