{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,4,30]],"date-time":"2026-04-30T14:40:26Z","timestamp":1777560026225,"version":"3.51.4"},"reference-count":23,"publisher":"SAGE Publications","issue":"3","content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["AIC"],"published-print":{"date-parts":[[2023,8,21]]},"abstract":"<jats:p>Object localization has been the focus of research in Fine-Grained Visual Categorization (FGVC). With the aim of improving the accuracy and precision of object localization in multi-branch networks, as well as the robustness and universality of object localization methods, our study mainly focus on how to combine coordinate attention and feature activation map for target localization. The model in this paper is a three-branch model including raw branch, object branch and part branch. The images are fed directly into the raw branch. Coordinate Attention Object Localization Module (CAOLM) is used to localize and crop objects in the image to generate the input for the object branch. Attention Partial Proposal Module (APPM) is used to propose part regions at different scales. The three classes of input images undergo end-to-end weakly supervised learning through different branches of the network. The model expands the receptive field to capture multi-scale features by Selective Branch Atrous Spatial Pooling Pyramid (SB-ASPP). It can fuse the feature maps obtained from the raw branch and the object branch with Selective Branch Block (SBBlock), and the complete features of the raw branch are used to supplement the missing information of the object branch. Extensive experimental results on CUB-200-2011, FGVC-Aircraft and Stanford Cars datasets show that our method has the best classification performance on FGVC-Aircraft and also has competitive performance on other datasets. Few parameters and fast inference speed are also the advantages of our model.<\/jats:p>","DOI":"10.3233\/aic-220187","type":"journal-article","created":{"date-parts":[[2023,5,12]],"date-time":"2023-05-12T13:28:34Z","timestamp":1683898114000},"page":"205-217","source":"Crossref","is-referenced-by-count":2,"title":["Multi-branch selection fusion fine-grained classification algorithm based on coordinate attention localization"],"prefix":"10.1177","volume":"36","author":[{"given":"Feng","family":"Zhang","sequence":"first","affiliation":[{"name":"Department of Computer and Electronic Information, Guangxi University, Nanning, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Gaocai","family":"Wang","sequence":"additional","affiliation":[{"name":"Department of Computer and Electronic Information, Guangxi University, Nanning, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Man","family":"Wu","sequence":"additional","affiliation":[{"name":"School of Electrical Engineering, Guangxi University, Nanning, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Shuqiang","family":"Huang","sequence":"additional","affiliation":[{"name":"College of Cyber Security of Jinan University, Jinan University, Guangzhou, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"179","reference":[{"key":"10.3233\/AIC-220187_ref1","doi-asserted-by":"crossref","unstructured":"A.\u00a0Behera, Z.\u00a0Wharton, P.R.\u00a0Hewage and A.\u00a0Bera, Context-aware attentional pooling (cap) for fine-grained visual classification, in: Proceedings of the AAAI Conference on Artificial Intelligence, Vol.\u00a035, 2021, pp.\u00a0929\u2013937.","DOI":"10.1609\/aaai.v35i2.16176"},{"key":"10.3233\/AIC-220187_ref4","doi-asserted-by":"crossref","unstructured":"F.\u00a0Chollet, Xception: Deep learning with depthwise separable convolutions, in: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2017, pp.\u00a01251\u20131258.","DOI":"10.1109\/CVPR.2017.195"},{"key":"10.3233\/AIC-220187_ref5","doi-asserted-by":"crossref","unstructured":"J.\u00a0Fu, H.\u00a0Zheng and T.\u00a0Mei, Look closer to see better: Recurrent attention convolutional neural network for fine-grained image recognition, in: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2017, pp.\u00a04438\u20134446.","DOI":"10.1109\/CVPR.2017.476"},{"key":"10.3233\/AIC-220187_ref6","doi-asserted-by":"crossref","unstructured":"W.\u00a0Ge, X.\u00a0Lin and Y.\u00a0Yu, Weakly supervised complementary parts models for fine-grained image classification from the bottom up, in: Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, 2019, pp.\u00a03034\u20133043.","DOI":"10.1109\/CVPR.2019.00315"},{"key":"10.3233\/AIC-220187_ref8","doi-asserted-by":"crossref","unstructured":"K.\u00a0He, X.\u00a0Zhang, S.\u00a0Ren and J.\u00a0Sun, Deep residual learning for image recognition, in: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2016, pp.\u00a0770\u2013778.","DOI":"10.1109\/CVPR.2016.90"},{"key":"10.3233\/AIC-220187_ref9","doi-asserted-by":"crossref","unstructured":"Q.\u00a0Hou, D.\u00a0Zhou and J.\u00a0Feng, Coordinate attention for efficient mobile network design, in: Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, 2021, pp.\u00a013713\u201313722.","DOI":"10.1109\/CVPR46437.2021.01350"},{"key":"10.3233\/AIC-220187_ref10","doi-asserted-by":"crossref","unstructured":"J.\u00a0Hu, L.\u00a0Shen and G.\u00a0Sun, Squeeze-and-excitation networks, in: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2018, pp.\u00a07132\u20137141.","DOI":"10.1109\/CVPR.2018.00745"},{"key":"10.3233\/AIC-220187_ref11","unstructured":"S.\u00a0Ioffe and C.\u00a0Szegedy, Batch normalization: Accelerating deep network training by reducing internal covariate shift, in: International Conference on Machine Learning, PMLR, 2015, pp.\u00a0448\u2013456."},{"key":"10.3233\/AIC-220187_ref12","doi-asserted-by":"crossref","unstructured":"R.\u00a0Ji, L.\u00a0Wen, L.\u00a0Zhang, D.\u00a0Du, Y.\u00a0Wu, C.\u00a0Zhao, X.\u00a0Liu and F.\u00a0Huang, Attention convolutional binary neural tree for fine-grained visual categorization, in: Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, 2020, pp.\u00a010468\u201310477.","DOI":"10.1109\/CVPR42600.2020.01048"},{"key":"10.3233\/AIC-220187_ref13","doi-asserted-by":"crossref","unstructured":"J.\u00a0Krause, M.\u00a0Stark, J.\u00a0Deng and L.\u00a0Fei-Fei, 3d object representations for fine-grained categorization, in: Proceedings of the IEEE International Conference on Computer Vision Workshops, 2013, pp.\u00a0554\u2013561.","DOI":"10.1109\/ICCVW.2013.77"},{"key":"10.3233\/AIC-220187_ref14","doi-asserted-by":"crossref","unstructured":"X.\u00a0Li, W.\u00a0Wang, X.\u00a0Hu and J.\u00a0Yang, Selective kernel networks, in: Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, 2019, pp.\u00a0510\u2013519.","DOI":"10.1109\/CVPR.2019.00060"},{"key":"10.3233\/AIC-220187_ref16","unstructured":"V.\u00a0Nair and G.E.\u00a0Hinton, Rectified linear units improve restricted Boltzmann machines, in: Icml, 2010."},{"issue":"3","key":"10.3233\/AIC-220187_ref17","doi-asserted-by":"publisher","first-page":"1487","DOI":"10.1109\/TIP.2017.2774041","article-title":"Object-part attention model for fine-grained image classification","volume":"27","author":"Peng","year":"2017","journal-title":"IEEE Transactions on Image Processing"},{"key":"10.3233\/AIC-220187_ref19","doi-asserted-by":"crossref","unstructured":"Y.\u00a0Rao, G.\u00a0Chen, J.\u00a0Lu and J.\u00a0Zhou, Counterfactual attention learning for fine-grained visual categorization and re-identification, in: Proceedings of the IEEE\/CVF International Conference on Computer Vision, 2021, pp.\u00a01025\u20131034.","DOI":"10.1109\/ICCV48922.2021.00106"},{"key":"10.3233\/AIC-220187_ref20","doi-asserted-by":"crossref","unstructured":"C.\u00a0Szegedy, S.\u00a0Ioffe, V.\u00a0Vanhoucke and A.A.\u00a0Alemi, Inception-v4, inception-resnet and the impact of residual connections on learning, in: Thirty-First AAAI Conference on Artificial Intelligence, 2017.","DOI":"10.1609\/aaai.v31i1.11231"},{"key":"10.3233\/AIC-220187_ref22","doi-asserted-by":"publisher","DOI":"10.1109\/WACV.2018.00163"},{"issue":"4","key":"10.3233\/AIC-220187_ref23","doi-asserted-by":"publisher","first-page":"1470","DOI":"10.1007\/s10618-021-00765-5","article-title":"Smoothed dilated convolutions for improved dense prediction","volume":"35","author":"Wang","year":"2021","journal-title":"Data Mining and Knowledge Discovery"},{"issue":"6","key":"10.3233\/AIC-220187_ref24","doi-asserted-by":"publisher","first-page":"2868","DOI":"10.1109\/TIP.2017.2688133","article-title":"Selective convolutional descriptor aggregation for fine-grained image retrieval","volume":"26","author":"Wei","year":"2017","journal-title":"IEEE Transactions on Image Processing"},{"key":"10.3233\/AIC-220187_ref26","doi-asserted-by":"crossref","unstructured":"S.\u00a0Woo, J.\u00a0Park, J.-Y.\u00a0Lee and I.S.\u00a0Kweon, Cbam: Convolutional block attention module, in: Proceedings of the European Conference on Computer Vision (ECCV), 2018, pp.\u00a03\u201319.","DOI":"10.1007\/978-3-030-01234-2_1"},{"issue":"5","key":"10.3233\/AIC-220187_ref27","doi-asserted-by":"publisher","first-page":"1791","DOI":"10.1109\/TCYB.2018.2813971","article-title":"Deep attention-based spatially recursive networks for fine-grained visual recognition","volume":"49","author":"Wu","year":"2018","journal-title":"IEEE transactions on cybernetics"},{"key":"10.3233\/AIC-220187_ref28","doi-asserted-by":"crossref","unstructured":"S.\u00a0Xie, R.\u00a0Girshick, P.\u00a0Doll\u00e1r, Z.\u00a0Tu and K.\u00a0He, Aggregated residual transformations for deep neural networks, in: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2017, pp.\u00a01492\u20131500.","DOI":"10.1109\/CVPR.2017.634"},{"key":"10.3233\/AIC-220187_ref29","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-67832-6_12"},{"key":"10.3233\/AIC-220187_ref31","doi-asserted-by":"crossref","unstructured":"P.\u00a0Zhuang, Y.\u00a0Wang and Y.\u00a0Qiao, Learning attentive pairwise interaction for fine-grained classification, in: Proceedings of the AAAI Conference on Artificial Intelligence, Vol.\u00a034, 2020, pp.\u00a013130\u201313137.","DOI":"10.1609\/aaai.v34i07.7016"}],"container-title":["AI Communications"],"original-title":[],"link":[{"URL":"https:\/\/content.iospress.com\/download?id=10.3233\/AIC-220187","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,4,28]],"date-time":"2026-04-28T18:28:08Z","timestamp":1777400888000},"score":1,"resource":{"primary":{"URL":"https:\/\/journals.sagepub.com\/doi\/full\/10.3233\/AIC-220187"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,8,21]]},"references-count":23,"journal-issue":{"issue":"3"},"URL":"https:\/\/doi.org\/10.3233\/aic-220187","relation":{},"ISSN":["1875-8452","0921-7126"],"issn-type":[{"value":"1875-8452","type":"electronic"},{"value":"0921-7126","type":"print"}],"subject":[],"published":{"date-parts":[[2023,8,21]]}}}