{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,2,28]],"date-time":"2026-02-28T17:55:34Z","timestamp":1772301334704,"version":"3.50.1"},"reference-count":41,"publisher":"Springer Science and Business Media LLC","issue":"10","license":[{"start":{"date-parts":[[2022,10,7]],"date-time":"2022-10-07T00:00:00Z","timestamp":1665100800000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2022,10,7]],"date-time":"2022-10-07T00:00:00Z","timestamp":1665100800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"name":"The 2020 Key research and development Plan of Shaanxi Province","award":["2020ZDLSF04-02"],"award-info":[{"award-number":["2020ZDLSF04-02"]}]},{"DOI":"10.13039\/501100001775","name":"University of Technology Sydney","doi-asserted-by":"crossref","id":[{"id":"10.13039\/501100001775","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Multimed Tools Appl"],"published-print":{"date-parts":[[2023,4]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:p>Fine-Grained Visual Categorization (FGVC) aims to distinguish between extremely similar subordinate-level categories within the same basic-level category. Existing research has proven the great importance of the discriminative features in FGVC but ignored the contributions for correct classification from other features, and the extracted features always contain more information about the obvious regions but less about subtle regions. In this paper, firstly, a novel module named forcing module is proposed to force the network to extract more diverse features for FGVC, which generates a suppression mask based on the class activation maps to suppress the most distinguishable regions, so as to force the network to extract other secondary distinguishable features as the final features. The forcing module consists of the original branch and the forcing branch. The original branch focuses on the primary discriminative regions while the forcing branch focuses on secondary discriminative regions. Secondly, in order to solve the problem that information of small-scale distinguishable features is lost seriously after multi-layer down-sampling, according to the class activation maps of the first prediction, the object is cropped and scaled as the second input. To reduce the prediction error, the first and second prediction probabilities are fused as the final prediction result. Experimental results indicate that the proposed method not only outperforms the baseline model by a large margin (3.7%, 5.9%, 3.1% respectively) on CUB-200-2011, Stanford-Cars, and FGVC-Aircraft, but also achieves state-of-the-art performance on FGVC-Aircraft.<\/jats:p>","DOI":"10.1007\/s11042-022-13619-z","type":"journal-article","created":{"date-parts":[[2022,10,7]],"date-time":"2022-10-07T08:03:48Z","timestamp":1665129828000},"page":"14799-14813","update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":13,"title":["Learning enhanced features and inferring twice for fine-grained image classification"],"prefix":"10.1007","volume":"82","author":[{"given":"Xuan","family":"Nie","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Bosong","family":"Chai","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Luyao","family":"Wang","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Qiyu","family":"Liao","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-9581-8849","authenticated-orcid":false,"given":"Min","family":"Xu","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"297","published-online":{"date-parts":[[2022,10,7]]},"reference":[{"key":"13619_CR1","first-page":"1","volume":"20","author":"A Azulay","year":"2019","unstructured":"Azulay A, Weiss Y (2019) Why do deep convolutional networks generalize so poorly to small image transformations?. J Mach Learn Res 20:1\u201325","journal-title":"J Mach Learn Res"},{"key":"13619_CR2","doi-asserted-by":"crossref","unstructured":"Berg T, Belhumeur PN (2013) POOF: part-based one-vs.-one features for fine-grained categorization, face verification, and attribute estimation. In: 2013 IEEE conference on computer vision and pattern recognition, Portland, OR, USA, pp 955\u2013962","DOI":"10.1109\/CVPR.2013.128"},{"key":"13619_CR3","doi-asserted-by":"crossref","unstructured":"Chai Y, Lempitsky V, Zisserman A (2013) Symbiotic segmentation and part localization for fine-grained categorization. In: Proceedings of the 2013 IEEE international conference on computer vision, IEEE","DOI":"10.1109\/ICCV.2013.47"},{"key":"13619_CR4","doi-asserted-by":"publisher","first-page":"4683","DOI":"10.1109\/TIP.2020.2973812","volume":"29","author":"D Chang","year":"2020","unstructured":"Chang D, Ding Y, Xie J, Bhunia AK, Li X, Ma Z, Wu M, Guo J, Song Y-Z (2020) The devil is in the channels: mutual-channel loss for fine-grained image classification. IEEE Trans Image Process 29:4683\u20134695","journal-title":"IEEE Trans Image Process"},{"key":"13619_CR5","doi-asserted-by":"crossref","unstructured":"Chen Y, Bai Y, Zhang W, Mei T (2019) Destruction and construction learning for fine-grained image recognition. In: 2019 IEEE\/CVF conference on computer vision and pattern recognition (CVPR), Long Beach, CA, USA, pp 5152\u20135161","DOI":"10.1109\/CVPR.2019.00530"},{"key":"13619_CR6","doi-asserted-by":"crossref","unstructured":"Deng J, Dong W, Socher R, Li L, Li K, Li F-F (2009) Imagenet: a large-scale hierarchical image database. In: 2009 IEEE conference on computer vision and pattern recognition, Miami, FL, USA, pp 248\u2013255","DOI":"10.1109\/CVPR.2009.5206848"},{"key":"13619_CR7","doi-asserted-by":"crossref","unstructured":"Ding Y, Zhou Y, Zhu Y, Ye Q, Jiao J (2019) Selective sparse sampling for Fine-Grained image recognition. In: 2019 IEEE\/CVF international conference on computer vision (ICCV), Seoul, Korea (South), pp 6598\u20136607","DOI":"10.1109\/ICCV.2019.00670"},{"key":"13619_CR8","doi-asserted-by":"publisher","first-page":"101294","DOI":"10.1016\/j.jup.2021.101294","volume":"73","author":"G-F Fan","year":"2021","unstructured":"Fan G-F, Yu M, Dong S-Q, Yeh Y-H, Hong W-C (2021) Forecasting short-term electricity load using hybrid support vector regression with grey catastrophe and random forest modeling. Util Policy 73:101294","journal-title":"Util Policy"},{"key":"13619_CR9","doi-asserted-by":"crossref","unstructured":"Gao Y, Beijbom O, Zhang N, Darrell T (2016) Compact bilinear pooling. In: 2016 IEEE conference on computer vision and pattern recognition (CVPR), Las Vegas, NV, USA, pp 317\u2013326","DOI":"10.1109\/CVPR.2016.41"},{"key":"13619_CR10","doi-asserted-by":"crossref","unstructured":"Ge W, Lin X, Yu Y (2019) Weakly supervised complementary parts models for fine-grained image classification from the bottom up. In: 2019 IEEE\/CVF conference on computer vision and pattern recognition (CVPR), Long Beach, CA, USA, 2019, pp 3029\u20133038","DOI":"10.1109\/CVPR.2019.00315"},{"key":"13619_CR11","unstructured":"Glorot X, Bengio Y (2010) Understanding the difficulty of training deep feedforward neural networks. In: Proceedings of the thirteenth international conference on artificial intelligence and statistics, pp 249\u2013256"},{"key":"13619_CR12","doi-asserted-by":"crossref","unstructured":"Hanselmann H, Ney H (2020) ELOPE: fine-grained visual classification with efficient localization, pooling and embedding. In: 2020 IEEE Winter Conference on Applications of Computer Vision (WACV), Snowmass, CO, USA, pp 1236\u20131245","DOI":"10.1109\/WACV45572.2020.9093601"},{"key":"13619_CR13","doi-asserted-by":"crossref","unstructured":"He K, Zhang X, Ren S, Sun J (2016) Deep residual learning for image recognition. In: 2016 IEEE conference on computer vision and pattern recognition (CVPR), Las Vegas, NV, USA, pp 770\u2013778","DOI":"10.1109\/CVPR.2016.90"},{"key":"13619_CR14","doi-asserted-by":"crossref","unstructured":"Huang G, Liu Z, Van Der Maaten L, Weinberger KQ (2017) Densely connected convolutional networks. In: 2017 IEEE conference on computer vision and pattern recognition (CVPR), Honolulu, HI, USA, vol 2017. pp 2261\u20132269","DOI":"10.1109\/CVPR.2017.243"},{"key":"13619_CR15","unstructured":"Khosla A, Jayadevaprakash N, Yao B, Fei-Fei L (2012) Novel dataset for fine-grained image categorization: stanford dogs"},{"key":"13619_CR16","doi-asserted-by":"crossref","unstructured":"Kong S, Fowlkes C (2017) Low-rank bilinear pooling for fine-grained classification. In: 2017 IEEE conference on computer vision and pattern recognition (CVPR), Honolulu, HI, USA, pp 7025\u20137034","DOI":"10.1109\/CVPR.2017.743"},{"key":"13619_CR17","doi-asserted-by":"crossref","unstructured":"Krause J, Stark M, Deng J, Fei-Fei L (2013) 3D object representations for fine-grained categorization. In: 2013 IEEE international conference on computer vision workshops, Sydney, NSW, Australia, pp 554\u2013561","DOI":"10.1109\/ICCVW.2013.77"},{"key":"13619_CR18","doi-asserted-by":"crossref","unstructured":"Lin T-Y, RoyChowdhury A, Subhransu M (2015) Bilinear CNN Models for Fine-grained Visual Recognition. In: Proceedings of the IEEE international conference on computer vision, pp 1449\u20131457","DOI":"10.1109\/ICCV.2015.170"},{"key":"13619_CR19","doi-asserted-by":"crossref","unstructured":"Liu Z, Mao H, Wu C-Y, Feichtenhofer C, Darrell T, Xie S (2022) A ConvNet for the 2020s. arXiv:2201.03545","DOI":"10.1109\/CVPR52688.2022.01167"},{"key":"13619_CR20","unstructured":"Maji S, et al. (2013) Fine-grained visual classification of aircraft hal inria"},{"issue":"13","key":"13619_CR21","doi-asserted-by":"publisher","first-page":"3812","DOI":"10.1093\/nar\/gkg509","volume":"31","author":"PC Ng","year":"2003","unstructured":"Ng PC, Henikoff S (2003) SIFT: predicting amino acid changes that affect protein function. Nucleic Acids Res 31(13):3812\u20133814","journal-title":"Nucleic Acids Res"},{"issue":"1","key":"13619_CR22","first-page":"91","volume":"7","author":"EM Onyema","year":"2020","unstructured":"Onyema EM, Elhaj MAE, Bashir SG, Abdullahi I, Hauwa AA, Hayatu AA, Edeh MO, Abdullahi I (2020) Evaluation of the performance of K-nearest neighbor algorithm in determining student learning styles. Int J Innov Sci Eng Technol 7(1):91\u2013102","journal-title":"Int J Innov Sci Eng Technol"},{"key":"13619_CR23","unstructured":"Paszke A, Gross S, Chintala S, Chanan G, Yang E, DeVito Z, Lin Z, Desmaison A, Antiga L, Lerer A (2017) Automatic differentiation in pytorch"},{"key":"13619_CR24","doi-asserted-by":"crossref","unstructured":"Semih Kayhan O, van Gemert JC (2020) On Translation Invariance in CNNs: Convolutional Layers Can Exploit Absolute Spatial Location. In: 2020 IEEE\/CVF Conference on computer vision and pattern recognition (CVPR), Seattle, WA, USA, pp 14262\u201314273","DOI":"10.1109\/CVPR42600.2020.01428"},{"key":"13619_CR25","doi-asserted-by":"crossref","unstructured":"Sun G, Cholakkal H, Khan S, Khan F, Shao L (2020) Fine-grained recognition, accounting for subtle differences between similar classes. In: Proceedings of the AAAI conference on artificial intelligence, pp 12047\u201312054","DOI":"10.1609\/aaai.v34i07.6882"},{"key":"13619_CR26","doi-asserted-by":"crossref","unstructured":"Sun M, Yuan Y, Zhou F, Ding E (2018) Multi-attention multi-class constraint for fine-grained image recognition. In: Proceedings of the european conference on computer vision (ECCV), pp 805\u2013821","DOI":"10.1007\/978-3-030-01270-0_49"},{"key":"13619_CR27","doi-asserted-by":"crossref","unstructured":"Szegedy C, Liu W, Jia Y, Sermanet P, Reed S, Anguelov D, Erhan D, Vanhoucke V, Rabinovich A (2015) Going deeper with convolutions. In: 2015 IEEE conference on computer vision and pattern recognition (CVPR), Boston, MA, USA, pp 1\u20139","DOI":"10.1109\/CVPR.2015.7298594"},{"key":"13619_CR28","unstructured":"Tan M, Le Q (2019) Efficientnet: rethinking model scaling for convolutional neural networks. In: International conference on machine learning, PMLR, pp 6105\u20136114"},{"key":"13619_CR29","unstructured":"Wah C, Branson S, Welinder P, Perona P, Belongie S (2011) The caltech-ucsd birds-200-2011 dataset. California Institute of Technology"},{"key":"13619_CR30","doi-asserted-by":"publisher","first-page":"1414","DOI":"10.1109\/TNNLS.2020.3042276","volume":"33","author":"Q Wang","year":"2020","unstructured":"Wang Q, Huang W, Xiong Z, et al. (2020) Looking closer at the scene: multiscale representation learning for remote sensing image scene classification. IEEE Trans Neural Netw Learn Syst 33:1414\u20131428","journal-title":"IEEE Trans Neural Netw Learn Syst"},{"key":"13619_CR31","doi-asserted-by":"crossref","unstructured":"Wang Z, Wang S, Li H, Dou Z, Li J (2020) Graph-propagation based correlation learning for weakly supervised fine-grained image classification. In: Proceedings of the AAAI conference on artificial intelligence, pp 12289\u201312296","DOI":"10.1609\/aaai.v34i07.6912"},{"key":"13619_CR32","doi-asserted-by":"crossref","unstructured":"Xie L, Tian Q, Hong R, Yan S, Zhang B (2013) Hierarchical part matching for fine-grained visual categorization. In: 2013 IEEE international conference on computer vision, pp 1641\u20131648","DOI":"10.1109\/ICCV.2013.206"},{"key":"13619_CR33","doi-asserted-by":"publisher","first-page":"2722","DOI":"10.1109\/TIP.2021.3053459","volume":"30","author":"Z Xiong","year":"2021","unstructured":"Xiong Z, Yuan Y, Wang Q (2021) ASK: adaptively selecting key local features for RGB-d scene recognition. IEEE Trans Image Process 30:2722\u20132733","journal-title":"IEEE Trans Image Process"},{"key":"13619_CR34","doi-asserted-by":"crossref","unstructured":"Yang Z, Luo T, Wang D, Hu Z, Gao J, Wang L (2018) Learning to navigate for fine-grained classification. In: Proceedings of the european conference on computer vision (ECCV) pp 420\u2013435","DOI":"10.1007\/978-3-030-01264-9_26"},{"key":"13619_CR35","doi-asserted-by":"crossref","unstructured":"Yu C, Zhao X, Zheng Q, Zhang P, You X (2018) Hierarchical Bilinear Pooling for Fine-Grained Visual Recognition. In: Proceedings of the European conference on computer vision (ECCV), pp 574\u2013589","DOI":"10.1007\/978-3-030-01270-0_35"},{"key":"13619_CR36","unstructured":"Zhang R (2019) Making convolutional networks shift-invariant again. ICML"},{"key":"13619_CR37","doi-asserted-by":"crossref","unstructured":"Zhang L, Huang S, Liu W, Tao D (2019) Learning a mixture of granularity-specific experts for fine-grained categorization. In: 2019 IEEE\/CVF international conference on computer vision (ICCV), Seoul, Korea (South), pp 8330\u20138339","DOI":"10.1109\/ICCV.2019.00842"},{"key":"13619_CR38","doi-asserted-by":"crossref","unstructured":"Zhang N, et al. (2014) Part-based r-CNNs for Fine-Grained Category Detection. In: European conference on computer vision. Springer, Cham","DOI":"10.1007\/978-3-319-10590-1_54"},{"key":"13619_CR39","doi-asserted-by":"crossref","unstructured":"Zheng H, Fu J, Mei T, Luo J (2017) Learning multi-attention convolutional neural network for fine-grained image recognition. In: 2017 IEEE international conference on computer vision (ICCV), Venice, Italy, pp 5219\u20135227","DOI":"10.1109\/ICCV.2017.557"},{"key":"13619_CR40","doi-asserted-by":"crossref","unstructured":"Zheng H, Fu J, Zha Z-J, Luo J (2019) Looking for the devil in the details, learning trilinear attention sampling network for fine-grained image recognition. In: IEEE\/CVF conference on computer vision and pattern recognition (CVPR ), IEEE","DOI":"10.1109\/CVPR.2019.00515"},{"key":"13619_CR41","doi-asserted-by":"crossref","unstructured":"Zhuang P, Wang Y, Qiao Y (2020) Learning attentive pairwise interaction for fine-grained classification. In: Proceedings of the AAAI conference on artificial intelligence, pp 13130\u201313137","DOI":"10.1609\/aaai.v34i07.7016"}],"container-title":["Multimedia Tools and Applications"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s11042-022-13619-z.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s11042-022-13619-z\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s11042-022-13619-z.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2023,3,24]],"date-time":"2023-03-24T11:39:10Z","timestamp":1679657950000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s11042-022-13619-z"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,10,7]]},"references-count":41,"journal-issue":{"issue":"10","published-print":{"date-parts":[[2023,4]]}},"alternative-id":["13619"],"URL":"https:\/\/doi.org\/10.1007\/s11042-022-13619-z","relation":{},"ISSN":["1380-7501","1573-7721"],"issn-type":[{"value":"1380-7501","type":"print"},{"value":"1573-7721","type":"electronic"}],"subject":[],"published":{"date-parts":[[2022,10,7]]},"assertion":[{"value":"24 February 2021","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"20 May 2022","order":2,"name":"revised","label":"Revised","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"21 July 2022","order":3,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"7 October 2022","order":4,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"We have no conflict of interests to disclose.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"<!--Emphasis Type='Bold' removed-->Conflict of Interests"}}]}}