{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,4,21]],"date-time":"2026-04-21T22:48:30Z","timestamp":1776811710153,"version":"3.51.2"},"reference-count":30,"publisher":"European Society of Computational Methods in Sciences and Engineering","issue":"4","license":[{"start":{"date-parts":[[2023,7,1]],"date-time":"2023-07-01T00:00:00Z","timestamp":1688169600000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/journals.sagepub.com\/page\/policies\/text-and-data-mining-license"}],"content-domain":{"domain":["journals.sagepub.com"],"crossmark-restriction":true},"short-container-title":["Journal of Computational Methods in Sciences and Engineering"],"published-print":{"date-parts":[[2023,7]]},"abstract":"<jats:p>Holistic scene understanding is a challenging problem in computer vision. Most recent researches in this field were focusing on the object detection, the semantic segmentation and the relationship detection tasks. The attribute can provide meaningful information for the object instance, thus the object instance can be expressed more detail in the scene understanding. However, most researches in this field have been limited to several special conditions. Such as, several researches were just focusing on the attribute of special object class, because their solutions were aimed at a limited-scenarios, their methods are hardly to generalize in other scenarios. We also find that most of the research for multi-attribute detection task were only regarding each attribute as binary class and simply use the multi-binary-classifier method for the attribute detection. But these strategies above not consider the relation between each pair of the attributes, they will fall into trouble in the \u201cimperfect\u201d attribute dataset (which is labeled with the missing and incomplete annotations), and they will have low performance in the long-tail attribute class (which has lower rank of annotation and more missing labels). In this paper, we focus on the multi-attribute detection for a variant of object classes and take the relation between attributes into consideration. We propose a GRU-based model to detect a variable-length attribute sequence with a customized loss compute method to solve the \u201cimperfect\u201d attribute dataset problem. Furthermore, we perform ablative studies to prove the effectiveness of each part of our method. Finally, we compare our model with several existed multi-attribute detection methods on VG (Visual Genome) and CUB200 bird datasets to prove the superior performance of the proposed model.<\/jats:p>","DOI":"10.3233\/jcm-226762","type":"journal-article","created":{"date-parts":[[2023,5,10]],"date-time":"2023-05-10T09:07:04Z","timestamp":1683709624000},"page":"1913-1927","update-policy":"https:\/\/doi.org\/10.1177\/sage-journals-update-policy","source":"Crossref","is-referenced-by-count":0,"title":["Variable-length sequence model for attribute detection in the image"],"prefix":"10.66113","volume":"23","author":[{"given":"Xin","family":"Li","sequence":"first","affiliation":[{"name":"Southeast University, NanJing, Jiangsu, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Jiaming","family":"Gu","sequence":"additional","affiliation":[{"name":"National Engineering Research Center for Broadband Networks and Applications, Shanghai, Shanghai, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Xiaoyuan","family":"Lu","sequence":"additional","affiliation":[{"name":"Southeast University, NanJing, Jiangsu, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yan","family":"Ning","sequence":"additional","affiliation":[{"name":"National Engineering Research Center for Broadband Networks and Applications, Shanghai, Shanghai, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Liang","family":"Zhang","sequence":"additional","affiliation":[{"name":"Xidian University, Xi\u2019an, Shanxi, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Peiyi","family":"Shen","sequence":"additional","affiliation":[{"name":"Xidian University, Xi\u2019an, Shanxi, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Chaochen","family":"Gu","sequence":"additional","affiliation":[{"name":"Shanghai Jiao Tong University, Shanghai, Shanghai, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"55691","published-online":{"date-parts":[[2023,7]]},"reference":[{"key":"e_1_3_1_2_2","doi-asserted-by":"crossref","unstructured":"AntolS AgrawalA LuJ MitchellM BatraD Lawrence ZitnickC ParikhD. Vqa: Visual question answering in: Proceedings of the IEEE international conference on computer vision 2015; pp 2425-2433.","DOI":"10.1109\/ICCV.2015.279"},{"key":"e_1_3_1_3_2","doi-asserted-by":"publisher","DOI":"10.3115\/1219044.1219075"},{"key":"e_1_3_1_4_2","doi-asserted-by":"crossref","unstructured":"ChenZ WeiX WangP GuoY. sMulti-label image recognition with graph convolutional networks in: 2019 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR) 2019; pp. 5172-5181.","DOI":"10.1109\/CVPR.2019.00532"},{"key":"e_1_3_1_5_2","first-page":"3298","article-title":"Detecting visual relationships with deep relational networks","author":"Dai B","year":"2017","unstructured":"DaiB ZhangY LinD, Detecting visual relationships with deep relational networks, in: Computer Vision and Pattern Recognition (CVPR), 2017 IEEE Conference on, IEEE. 2017; pp. 3298-3308.","journal-title":"Computer Vision and Pattern Recognition (CVPR)"},{"key":"e_1_3_1_6_2","doi-asserted-by":"crossref","unstructured":"FarhadiA EndresI HoiemD ForsythD. Describing objects by their attributes in: Computer Vision and Pattern Recognition 2009. CVPR 2009 IEEE Conference on IEEE. 2009; pp. 1778-1785.","DOI":"10.1109\/CVPR.2009.5206772"},{"key":"e_1_3_1_7_2","article-title":"Deep imbalanced learning for face recognition and attribute prediction","author":"Huang C","year":"2018","unstructured":"HuangC LiY LoyCC TangX. Deep imbalanced learning for face recognition and attribute prediction. IEEE Trans Pattern Anal Mach Intell. 1806.00194; 2018.","journal-title":"IEEE Trans Pattern Anal Mach Intell. 1806.00194"},{"key":"e_1_3_1_8_2","doi-asserted-by":"crossref","unstructured":"JohnsonJ KarpathyA Fei-FeiL. Densecap: Fully convolutional localization networks for dense captioning in: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition 2016; pp. 4565-4574.","DOI":"10.1109\/CVPR.2016.494"},{"key":"e_1_3_1_9_2","doi-asserted-by":"publisher","DOI":"10.1007\/s11263-016-0981-7"},{"key":"e_1_3_1_10_2","doi-asserted-by":"crossref","unstructured":"LiD ChenX HuangK. Multi-attribute learning for pedestrian attribute recognition in surveillance scenarios in: ACPR 2015; pp. 111-115.","DOI":"10.1109\/ACPR.2015.7486476"},{"key":"e_1_3_1_11_2","doi-asserted-by":"crossref","unstructured":"LiY OuyangW ZhouB WangK WangX. Scene graph generation from objects phrases and region captions IEEE International Conference on Computer Vision 2017.","DOI":"10.1109\/ICCV.2017.142"},{"key":"e_1_3_1_12_2","doi-asserted-by":"crossref","unstructured":"LiangX HuZ ZhangH GanC XingEP. Recurrent topic-transition gan for visual paragraph generation. 2017 IEEE International Conference on Computer Vision (ICCV) 2017; pp. 3382-3391.","DOI":"10.1109\/ICCV.2017.364"},{"key":"e_1_3_1_13_2","doi-asserted-by":"crossref","unstructured":"LiangX LeeL XingEP. Deep variation-structured reinforcement learning for visual relationship and attribute detection in: Computer Vision and Pattern Recognition (CVPR) 2017 IEEE Conference on IEEE. 2017; pp. 4408-4417.","DOI":"10.1109\/CVPR.2017.469"},{"key":"e_1_3_1_14_2","doi-asserted-by":"crossref","unstructured":"LiangX ShenX FengJ LinL YanS. Semantic object parsing with graph lstm in: European Conference on Computer Vision Springer. 2016; pp. 125-143.","DOI":"10.1007\/978-3-319-46448-0_8"},{"key":"e_1_3_1_15_2","doi-asserted-by":"crossref","unstructured":"ChikontweP LeeHJ. Deep Multi-Task Network for Learning Person Identity and Attributes in IEEE Access 2018; 6: 60801-60811.","DOI":"10.1109\/ACCESS.2018.2875783"},{"key":"e_1_3_1_16_2","doi-asserted-by":"crossref","unstructured":"LongJ ShelhamerE DarrellT. Fully convolutional networks for semantic segmentation in: Proceedings of the IEEE conference on computer vision and pattern recognition 2015; pp 3431-3440.","DOI":"10.1109\/CVPR.2015.7298965"},{"key":"e_1_3_1_17_2","article-title":"Segment-based methods for facial attribute detection from partial faces","author":"Mahbub U","year":"2018","unstructured":"MahbubU SarkarS ChellappaR. Segment-based methods for facial attribute detection from partial faces. IEEE Transactions on Affective Computing, 2018.","journal-title":"IEEE Transactions on Affective Computing"},{"key":"e_1_3_1_18_2","doi-asserted-by":"crossref","unstructured":"HuangY ChenJ OuyangW WanW XueY. Image Captioning With End-to-End Attribute Detection and Subsequent Attributes Prediction in IEEE Transactions on Image Processing 2020; pp. 4013-4026.","DOI":"10.1109\/TIP.2020.2969330"},{"key":"e_1_3_1_19_2","doi-asserted-by":"crossref","unstructured":"PenningtonJ SocherR ManningCD. Glove: Global vectors for word representation ins: Empirical Methods in Natural Language Processing (EMNLP) 2014; pp. 1532-1543.","DOI":"10.3115\/v1\/D14-1162"},{"key":"e_1_3_1_20_2","doi-asserted-by":"crossref","unstructured":"RedmonJ DivvalaS GirshickR FarhadiA. You only look once: Unified real-time object detection in: The IEEE Conference on Computer Vision and Pattern Recognition (CVPR) 2016.","DOI":"10.1109\/CVPR.2016.91"},{"key":"e_1_3_1_21_2","unstructured":"RenS HeK GirshickR SunJ. Faster r-cnn: Towards real-time object detection with region proposal networks in: Advances in neural information processing systems 2015; pp. 91-99."},{"key":"e_1_3_1_22_2","doi-asserted-by":"crossref","unstructured":"GongP WangX ChengY WangZJ YuQ. Zero-Shot Classification Based on Multitask Mixed Attribute Relations and Attribute-Specific Features in IEEE Transactions on Cognitive and Developmental Systems 2020","DOI":"10.1109\/TCDS.2019.2902250"},{"key":"e_1_3_1_23_2","doi-asserted-by":"crossref","unstructured":"TaherkhaniF NasrabadiNM DawsonJ. A Deep Face Identification Network Enhanced by Facial Attributes Prediction IEEE\/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW) 2018; pp. 666-6667.","DOI":"10.1109\/CVPRW.2018.00097"},{"key":"e_1_3_1_24_2","doi-asserted-by":"crossref","unstructured":"WahC BransonS WelinderP PeronaP BelongieS. Multiclass recognition and part localization with humans in the loop 2011 International Conference on Computer Vision 2011; pp. 2524-2531.","DOI":"10.1109\/ICCV.2011.6126539"},{"key":"e_1_3_1_25_2","unstructured":"YangP SunX LiW MaS WuW WangH. SGM: sequence generation model for multi-label classification in: Proceedings of the 27th International Conference on Computational Linguistics COLING 2018 Santa Fe New Mexico USA August 20\u201326 2018; pp. 3915-3926."},{"key":"e_1_3_1_26_2","doi-asserted-by":"crossref","unstructured":"YuD FuJ MeiT RuiY. Multi-level attention networks for visual question answering in: Computer Vision and Pattern Recognition (CVPR) 2017 IEEE Conference on IEEE. 2017; pp. 4187-4195.","DOI":"10.1109\/CVPR.2017.446"},{"key":"e_1_3_1_27_2","unstructured":"ZakizadehR SasdelliM QianY VazquezE. Finetag: Multiattribute classification at fine-grained level in images 2018."},{"key":"e_1_3_1_28_2","doi-asserted-by":"crossref","unstructured":"ZhangH KyawZ ChangS-F ChuaT-S. Visual Translation Embedding Network for Visual Relation Detection IEEE Conference on Computer Vision and Pattern Recognition (CVPR) 2017; pp. 3107-3115.","DOI":"10.1109\/CVPR.2017.331"},{"key":"e_1_3_1_29_2","doi-asserted-by":"crossref","unstructured":"ZhuangN YanY ChenS WangH. Multi-task Learning of Cascaded CNN for Facial Attribute Classification 24th International Conference on Pattern Recognition (ICPR) 2018; pp. 2069-2074.","DOI":"10.1109\/ICPR.2018.8545271"},{"key":"e_1_3_1_30_2","doi-asserted-by":"crossref","unstructured":"XuY YinF XuW LinJ CuiS. Wireless Traffic Prediction With Scalable Gaussian Process: Framework Algorithms and Verification in: IEEE Journal on Selected Areas in Communications pp. 1291-1306.","DOI":"10.1109\/JSAC.2019.2904330"},{"key":"e_1_3_1_31_2","doi-asserted-by":"crossref","unstructured":"YinF LinZ KongQ XuY LiD. Sergios Theodoridis Shuguang Cui FedLoc: Federated Learning Framework for Data-Driven Cooperative Localization and Location Data Processing in: IEEE Open Journal of Signal Processing 2020; pp. 187-215.","DOI":"10.1109\/OJSP.2020.3036276"}],"container-title":["Journal of Computational Methods in Sciences and Engineering"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/journals.sagepub.com\/doi\/pdf\/10.3233\/JCM-226762","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/journals.sagepub.com\/doi\/full-xml\/10.3233\/JCM-226762","content-type":"application\/xml","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/journals.sagepub.com\/doi\/pdf\/10.3233\/JCM-226762","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,4,21]],"date-time":"2026-04-21T22:06:57Z","timestamp":1776809217000},"score":1,"resource":{"primary":{"URL":"https:\/\/journals.sagepub.com\/doi\/10.3233\/JCM-226762"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,7]]},"references-count":30,"journal-issue":{"issue":"4","published-print":{"date-parts":[[2023,7]]}},"alternative-id":["10.3233\/JCM-226762"],"URL":"https:\/\/doi.org\/10.3233\/jcm-226762","relation":{},"ISSN":["1472-7978","1875-8983"],"issn-type":[{"value":"1472-7978","type":"print"},{"value":"1875-8983","type":"electronic"}],"subject":[],"published":{"date-parts":[[2023,7]]}}}