{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,14]],"date-time":"2026-07-14T22:28:06Z","timestamp":1784068086616,"version":"3.55.0"},"reference-count":40,"publisher":"Springer Science and Business Media LLC","issue":"2","license":[{"start":{"date-parts":[[2025,4,4]],"date-time":"2025-04-04T00:00:00Z","timestamp":1743724800000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2025,4,4]],"date-time":"2025-04-04T00:00:00Z","timestamp":1743724800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"name":"Young and Middle-Aged Scientific Research Foundation of Qinghai Normal University","award":["No. KJQN2022010"],"award-info":[{"award-number":["No. KJQN2022010"]}]}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Neural Process Lett"],"abstract":"<jats:title>Abstract<\/jats:title>\n          <jats:p>Attention mechanisms are critical tools for enhancing the performance of convolutional neural networks (CNNs), focusing on spatial and channel dimensions of feature maps, known as spatial attention and channel attention, respectively. While many advanced attention methods combine these dimensions to improve performance, particularly in downstream computer vision tasks, such methods often introduce significant computational overhead or fail to effectively capture long-range spatial dependencies alongside channel attention. To address these challenges, this paper proposes the sequential fusion attention (SFA) method, which introduces a complementary fusion strategy to integrate spatial and channel attention. Spatial attention leverages strip pooling to model long-range dependencies, while channel attention employs dynamic encoding to refine features. By utilizing a grouped processing approach, the SFA module achieves an optimal balance between computational efficiency and representation power. Extensive experiments on benchmark datasets demonstrate that SFA consistently outperforms state-of-the-art attention mechanisms, delivering competitive accuracy in image classification, object detection, and semantic segmentation tasks while maintaining reduced model complexity. This work underscores the potential of lightweight attention mechanisms in modern computer vision and paves the way for further innovations in resource-efficient neural network design. Our code is publicly available at the following URL: <jats:ext-link xmlns:xlink=\"http:\/\/www.w3.org\/1999\/xlink\" xlink:href=\"https:\/\/github.com\/Xuwei86\/SFA\" ext-link-type=\"uri\">https:\/\/github.com\/Xuwei86\/SFA<\/jats:ext-link>\n          <\/jats:p>","DOI":"10.1007\/s11063-025-11748-8","type":"journal-article","created":{"date-parts":[[2025,4,5]],"date-time":"2025-04-05T19:20:47Z","timestamp":1743880847000},"update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":10,"title":["SFA: Efficient Attention Mechanism for Superior CNN Performance"],"prefix":"10.1007","volume":"57","author":[{"given":"Wei","family":"Xu","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Yi","family":"Wan","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Dong","family":"Zhao","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"297","published-online":{"date-parts":[[2025,4,4]]},"reference":[{"key":"11748_CR1","doi-asserted-by":"crossref","unstructured":"Xie S, Girshick R, Doll\u00e1r P, Tu Z, He K (2017) Aggregated residual transformations for deep neural networks. In: Proceedings of the IEEE conference on computer vision and pattern recognition, pp 1492\u20131500","DOI":"10.1109\/CVPR.2017.634"},{"key":"11748_CR2","unstructured":"Howard AG, Zhu M, Chen B, Kalenichenko D, Wang W, Weyand T, Andreetto M, Adam H (2017) Mobilenets: efficient convolutional neural networks for mobile vision applications. arXiv preprint arXiv:1704.04861"},{"key":"11748_CR3","doi-asserted-by":"crossref","unstructured":"Ding X, Zhang X, Ma N, Han J, Ding G, Sun J (2021) Repvgg: making VGG-style convnets great again. In: Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition, pp 13733\u201313742","DOI":"10.1109\/CVPR46437.2021.01352"},{"key":"11748_CR4","doi-asserted-by":"crossref","unstructured":"Wang W, Dai J, Chen Z, Huang Z, Li Z, Zhu X, Hu X, Lu T, Lu L, Li H, et al (2023) Internimage: exploring large-scale vision foundation models with deformable convolutions. In: Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition, pp 14408\u201314419","DOI":"10.1109\/CVPR52729.2023.01385"},{"key":"11748_CR5","unstructured":"Vaswani A, Shazeer N, Parmar N, Uszkoreit J, Jones L, Gomez AN, Kaiser \u0141, Polosukhin I (2017) Attention Is All You Need. In Advances in Neural Information Processing Systems 30 pp 5998-6008. https:\/\/proceedings.neurips.cc\/paper\/2017\/file\/3f5ee243547dee91fbd053c1c4a845aa-Paper.pdf"},{"key":"11748_CR6","unstructured":"Devlin J, Chang M-W, Lee K, Toutanova K (2018) Bert: pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805"},{"key":"11748_CR7","unstructured":"Dosovitskiy A, Beyer L, Kolesnikov A, Weissenborn D, Zhai X, Unterthiner T (2020) Transformers for image recognition at scale. arXiv preprint arXiv:2010.11929"},{"key":"11748_CR8","doi-asserted-by":"crossref","unstructured":"Liu Z, Lin Y, Cao Y, Hu H, Wei Y, Zhang Z, Lin S, Guo B (2021) Swin transformer: hierarchical vision transformer using shifted windows. In: Proceedings of the IEEE\/CVF international conference on computer vision, pp 10012\u201310022","DOI":"10.1109\/ICCV48922.2021.00986"},{"key":"11748_CR9","unstructured":"Mehta S, Rastegari M (2021) Mobilevit: light-weight, general-purpose, and mobile-friendly vision transformer. arXiv preprint arXiv:2110.02178"},{"key":"11748_CR10","doi-asserted-by":"crossref","unstructured":"Liu Z, Mao H, Wu C-Y, Feichtenhofer C, Darrell T, Xie S (2022) A convnet for the 2020s. In: Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition, pp 11976\u201311986","DOI":"10.1109\/CVPR52688.2022.01167"},{"key":"11748_CR11","doi-asserted-by":"crossref","unstructured":"Woo S, Debnath S, Hu R, Chen X, Liu Z, Kweon IS, Xie S (2023) Convnext v2: co-designing and scaling convnets with masked autoencoders. In: Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition, pp 16133\u201316142","DOI":"10.1109\/CVPR52729.2023.01548"},{"key":"11748_CR12","unstructured":"Chen H, Wang Y, Guo J, Tao D (2023) Vanillanet: the power of minimalism in deep learning. arXiv preprint arXiv:2305.12972"},{"key":"11748_CR13","doi-asserted-by":"crossref","unstructured":"Wang W, Xie E, Li X, Fan D-P, Song K, Liang D, Lu T, Luo P, Shao L (2021) Pyramid vision transformer: a versatile backbone for dense prediction without convolutions. In: Proceedings of the IEEE\/CVF international conference on computer vision, pp 568\u2013578","DOI":"10.1109\/ICCV48922.2021.00061"},{"key":"11748_CR14","doi-asserted-by":"crossref","unstructured":"Bello I, Zoph B, Vaswani A, Shlens J, Le QV (2019) Attention augmented convolutional networks. In: Proceedings of the IEEE\/CVF international conference on computer vision, pp 3286\u20133295","DOI":"10.1109\/ICCV.2019.00338"},{"key":"11748_CR15","doi-asserted-by":"crossref","unstructured":"Hu J, Shen L, Sun G (2018) Squeeze-and-excitation networks, 7132\u20137141. In: The IEEE conference on computer vision and pattern recognition (CVPR). Salt Lake City, UT","DOI":"10.1109\/CVPR.2018.00745"},{"key":"11748_CR16","doi-asserted-by":"crossref","unstructured":"Wang Q, Wu B, Zhu P, Li P, Zuo W, Hu Q (2020) Eca-net: efficient channel attention for deep convolutional neural networks. In: Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition, pp 11534\u201311542","DOI":"10.1109\/CVPR42600.2020.01155"},{"key":"11748_CR17","unstructured":"Romero DW, Knigge DM, Gu A, Bekkers EJ, Gavves E, Tomczak JM, Hoogendoorn M (2022) Towards a general purpose CNN for long range dependencies in $$ n $$ d"},{"key":"11748_CR18","doi-asserted-by":"crossref","unstructured":"Hou Q, Zhang L, Cheng M-M, Feng J (2020) Strip pooling: rethinking spatial pooling for scene parsing. In: Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition, pp 4003\u20134012","DOI":"10.1109\/CVPR42600.2020.00406"},{"key":"11748_CR19","doi-asserted-by":"crossref","unstructured":"Sandler M, Howard A, Zhu M, Zhmoginov A, Chen L-C (2018) Mobilenetv2: inverted residuals and linear bottlenecks. In: Proceedings of the IEEE conference on computer vision and pattern recognition, pp 4510\u20134520","DOI":"10.1109\/CVPR.2018.00474"},{"key":"11748_CR20","doi-asserted-by":"crossref","unstructured":"He K, Zhang X, Ren S, Sun J (2016) Deep residual learning for image recognition. In: Proceedings of the IEEE conference on computer vision and pattern recognition, pp 770\u2013778","DOI":"10.1109\/CVPR.2016.90"},{"key":"11748_CR21","unstructured":"Chen L-C, Papandreou G, Schroff F, Adam H (2017) Rethinking atrous convolution for semantic image segmentation. arXiv preprint arXiv:1706.05587"},{"key":"11748_CR22","doi-asserted-by":"crossref","unstructured":"Deng J, Dong W, Socher R, Li L-J, Li K, Fei-Fei L (2009) Imagenet: a large-scale hierarchical image database. In: 2009 IEEE conference on computer vision and pattern recognition. IEEE, pp 248\u2013255","DOI":"10.1109\/CVPR.2009.5206848"},{"key":"11748_CR23","doi-asserted-by":"crossref","unstructured":"Lin T-Y, Maire M, Belongie S, Hays J, Perona P, Ramanan D, Doll\u00e1r P, Zitnick CL (2014) Microsoft coco: common objects in context. In: Computer vision\u2014ECCV 2014: 13th European conference, Zurich, Switzerland, September 6\u201312, 2014, proceedings, part V 13. Springer, pp 740\u2013755","DOI":"10.1007\/978-3-319-10602-1_48"},{"key":"11748_CR24","doi-asserted-by":"publisher","first-page":"98","DOI":"10.1007\/s11263-014-0733-5","volume":"111","author":"M Everingham","year":"2015","unstructured":"Everingham M, Eslami SA, Van Gool L, Williams CK, Winn J, Zisserman A (2015) The pascal visual object classes challenge: a retrospective. Int J Comput Vis 111:98\u2013136","journal-title":"Int J Comput Vis"},{"key":"11748_CR25","doi-asserted-by":"crossref","unstructured":"Szegedy C, Liu W, Jia Y, Sermanet P, Reed S, Anguelov D, Erhan D, Vanhoucke V, Rabinovich A (2015) Going deeper with convolutions. In: Proceedings of the IEEE conference on computer vision and pattern recognition, pp 1\u20139","DOI":"10.1109\/CVPR.2015.7298594"},{"key":"11748_CR26","doi-asserted-by":"crossref","unstructured":"Xue H, Liu C, Wan F, Jiao J, Ji X, Ye Q (2019) Danet: Divergent activation for weakly supervised object localization. In: Proceedings of the IEEE\/CVF international conference on computer vision, pp 6589\u20136598","DOI":"10.1109\/ICCV.2019.00669"},{"key":"11748_CR27","doi-asserted-by":"crossref","unstructured":"Li X, Wang W, Hu X, Yang J (2019) Selective kernel networks. In: Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition, pp 510\u2013519","DOI":"10.1109\/CVPR.2019.00060"},{"key":"11748_CR28","doi-asserted-by":"crossref","unstructured":"Qin Z, Zhang P, Wu F, Li X (2021) Fcanet: frequency channel attention networks. In: Proceedings of the IEEE\/CVF international conference on computer vision, pp 783\u2013792","DOI":"10.1109\/ICCV48922.2021.00082"},{"key":"11748_CR29","doi-asserted-by":"publisher","DOI":"10.1016\/j.eswa.2023.121338","volume":"237","author":"S Yan","year":"2024","unstructured":"Yan S, Shao H, Wang J, Zheng X, Liu B (2024) Liconvformer: a lightweight fault diagnosis framework using separable multiscale convolution and broadcast self-attention. Expert Syst Appl 237:121338","journal-title":"Expert Syst Appl"},{"key":"11748_CR30","doi-asserted-by":"crossref","unstructured":"Woo S, Park J, Lee J-Y, Kweon IS (2018) Cbam: convolutional block attention module. In: Proceedings of the European conference on computer vision (ECCV), pp 3\u201319","DOI":"10.1007\/978-3-030-01234-2_1"},{"key":"11748_CR31","doi-asserted-by":"crossref","unstructured":"Zhang Q-L, Yang Y-B (2021) Sa-net: shuffle attention for deep convolutional neural networks. In: ICASSP 2021-2021 IEEE international conference on acoustics, speech and signal processing (ICASSP). IEEE, pp 2235\u20132239","DOI":"10.1109\/ICASSP39728.2021.9414568"},{"key":"11748_CR32","doi-asserted-by":"crossref","unstructured":"Cao Y, Xu J, Lin S, Wei F, Hu H (2019) Gcnet: non-local networks meet squeeze-excitation networks and beyond. In: Proceedings of the IEEE\/CVF international conference on computer vision workshops","DOI":"10.1109\/ICCVW.2019.00246"},{"key":"11748_CR33","doi-asserted-by":"publisher","first-page":"429","DOI":"10.1016\/j.neunet.2023.12.003","volume":"171","author":"Y Cui","year":"2024","unstructured":"Cui Y, Knoll A (2024) Dual-domain strip attention for image restoration. Neural Netw 171:429\u2013439","journal-title":"Neural Netw"},{"key":"11748_CR34","doi-asserted-by":"crossref","unstructured":"Selvaraju RR, Cogswell M, Das A, Vedantam R, Parikh D, Batra D (2017) Grad-cam: visual explanations from deep networks via gradient-based localization. In: Proceedings of the IEEE international conference on computer vision, pp 618\u2013626","DOI":"10.1109\/ICCV.2017.74"},{"key":"11748_CR35","unstructured":"Xu W, Wan Y (2024) ELA: efficient local attention for deep convolutional neural networks"},{"key":"11748_CR36","unstructured":"Paszke A, Gross S, Massa F, Lerer A, Chintala S (2019) Pytorch: an imperative style, high-performance deep learning library"},{"key":"11748_CR37","unstructured":"Shaji AP, Hemalatha S Cosine annealing scheduler added with momentum factor: Classification of tomato plant diseases. Available at SSRN 4505950"},{"key":"11748_CR38","doi-asserted-by":"crossref","unstructured":"Hou Q, Zhou D, Feng J (2021) Coordinate attention for efficient mobile network design. In: Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition, pp 13713\u201313722","DOI":"10.1109\/CVPR46437.2021.01350"},{"key":"11748_CR39","unstructured":"Wan Q, Huang Z, Lu J, Yu G, Zhang L (2023) Seaformer: squeeze-enhanced axial transformer for mobile semantic segmentation. arXiv preprint arXiv:2301.13156"},{"key":"11748_CR40","unstructured":"Ge Z, Liu S, Wang F, Li Z, Sun J (2021) Yolox: exceeding yolo series in 2021. arXiv preprint arXiv:2107.08430"}],"container-title":["Neural Processing Letters"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s11063-025-11748-8.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s11063-025-11748-8\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s11063-025-11748-8.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,4,23]],"date-time":"2025-04-23T16:58:59Z","timestamp":1745427539000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s11063-025-11748-8"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,4,4]]},"references-count":40,"journal-issue":{"issue":"2","published-online":{"date-parts":[[2025,4]]}},"alternative-id":["11748"],"URL":"https:\/\/doi.org\/10.1007\/s11063-025-11748-8","relation":{},"ISSN":["1573-773X"],"issn-type":[{"value":"1573-773X","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,4,4]]},"assertion":[{"value":"24 February 2025","order":1,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"4 April 2025","order":2,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"All authors declare that they have no Conflict of interest.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Conflict of interest"}}],"article-number":"38"}}