{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,21]],"date-time":"2026-07-21T02:45:12Z","timestamp":1784601912337,"version":"3.55.0"},"reference-count":64,"publisher":"Springer Science and Business Media LLC","issue":"3","license":[{"start":{"date-parts":[[2025,5,27]],"date-time":"2025-05-27T00:00:00Z","timestamp":1748304000000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2025,5,27]],"date-time":"2025-05-27T00:00:00Z","timestamp":1748304000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Mach. Intell. Res."],"published-print":{"date-parts":[[2025,6]]},"abstract":"<jats:title>Abstract<\/jats:title>\n                  <jats:p>\n                    Efficiently capturing multi-scale local information and building long-range dependencies among pixels are essential for medical image segmentation because of the various sizes and shapes of the lesion regions or organs. In this paper, we propose the multi-scale cross-axis attention (MCA) mechanism to address these challenges through enhanced axial attention. To address the issues of insufficient learning of positional bias and limited long-distance interaction in axial attention caused by the small dataset, we propose using a dual cross-attention mechanism instead of axial attention to enhance global information capture. Meanwhile, to compensate for the lack of explicit attention to local information in axial attention, we use multiple convolutions of strip-shaped kernels with different kernel sizes in each axial attention path, which improves the efficiency of MCA in local information encoding. By integrating MCA into the multi-scale cross-axis attention network (MSCAN) backbone, we develop our network architecture, termed MCANet. With merely 4 M+ parameters, MCANet outperforms previous heavyweight approaches (e.g., swin transformer-based methods) across four challenging tasks: skin lesion segmentation, nuclei segmentation, abdominal multi-organ segmentation, and polyp segmentation. The code is available at\n                    <jats:ext-link xmlns:xlink=\"http:\/\/www.w3.org\/1999\/xlink\" xlink:href=\"https:\/\/github.com\/haoshao-nku\/medical_seg\" ext-link-type=\"uri\">https:\/\/github.com\/haoshao-nku\/medical_seg<\/jats:ext-link>\n                    .\n                  <\/jats:p>","DOI":"10.1007\/s11633-025-1552-6","type":"journal-article","created":{"date-parts":[[2025,5,27]],"date-time":"2025-05-27T02:14:54Z","timestamp":1748312094000},"page":"437-451","update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":40,"title":["MCANet: Medical Image Segmentation with Multi-scale Cross-axis Attention"],"prefix":"10.1007","volume":"22","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-0359-7629","authenticated-orcid":false,"given":"Hao","family":"Shao","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0006-7026-162X","authenticated-orcid":false,"given":"Quansheng","family":"Zeng","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-8388-8708","authenticated-orcid":false,"given":"Qibin","family":"Hou","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-0219-3443","authenticated-orcid":false,"given":"Jufeng","family":"Yang","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"297","published-online":{"date-parts":[[2025,5,27]]},"reference":[{"issue":"2","key":"1552_CR1","first-page":"145","volume":"7","author":"S Addimulam","year":"2022","unstructured":"S. Addimulam, M. A. Mohammed, R. K. Karanam, D. Ying, R. Pydipalli, B. Patel, M. A. Shajahan, N. Dhameliya, V. M. Natakam. Deep learning-enhanced image segmentation for medical diagnostics. Malaysian Journal of Medical and Biological Research, vol. 7, no. 2, pp. 145\u2013152, 2022","journal-title":"Malaysian Journal of Medical and Biological Research"},{"issue":"2","key":"1552_CR2","doi-asserted-by":"publisher","first-page":"1","DOI":"10.2352\/J.ImagingSci.Technol.2020.64.2.020508","volume":"64","author":"G T Du","year":"2020","unstructured":"G. T. Du, X. Cao, J. M. Liang, X. L. Chen, Y. H. Zhan. Medical image segmentation based on U-Net: A review. Journal of Imaging Science and Technology, vol. 64, no. 2, pp. 1\u201312, 2020. DOI: https:\/\/doi.org\/10.2352\/J.ImagingSci.Technol.2020.64.2.020508.","journal-title":"Journal of Imaging Science and Technology"},{"issue":"4","key":"1552_CR3","doi-asserted-by":"publisher","first-page":"1016","DOI":"10.1109\/TMI.2018.2876633","volume":"38","author":"Y K Huo","year":"2019","unstructured":"Y. K. Huo, Z. B. Xu, H. Moon, S. X. Bao, A. Assad, T. K. Moyo, M. R. Savona, R. G. Abramson, B. A. Landman. SynSeg-Net: Synthetic segmentation without target modality ground truth. IEEE Transactions on Medical Imaging, vol. 38, no. 4, pp. 1016\u20131025, 2019. DOI: https:\/\/doi.org\/10.1109\/tmi.2018.2876633.","journal-title":"IEEE Transactions on Medical Imaging"},{"issue":"7","key":"1552_CR4","doi-asserted-by":"publisher","first-page":"2494","DOI":"10.1109\/TMI.2020.2972701","volume":"39","author":"C Chen","year":"2020","unstructured":"C. Chen, Q. Dou, H. Chen, J. Qin, P. A. Heng. Unsupervised bidirectional cross-modality adaptation via deeply synergistic image and feature alignment for medical image segmentation. IEEE Transactions on Medical Imaging, vol. 39, no. 7, pp. 2494\u20132505, 2020. DOI: https:\/\/doi.org\/10.1109\/tmi.2020.2972701.","journal-title":"IEEE Transactions on Medical Imaging"},{"issue":"6","key":"1552_CR5","doi-asserted-by":"publisher","first-page":"531","DOI":"10.1007\/s11633-022-1371-y","volume":"19","author":"G P Ji","year":"2022","unstructured":"G. P. Ji, G. B. Xiao, Y. C. Chou, D. P. Fan, K. Zhao, G. Chen, L. Van Gool. Video polyp segmentation: A deep learning perspective. Machine Intelligence Research, vol. 19, no. 6, pp. 531\u2013549, 2022. DOI: https:\/\/doi.org\/10.1007\/s11633-022-1371-y.","journal-title":"Machine Intelligence Research"},{"key":"1552_CR6","doi-asserted-by":"publisher","first-page":"234","DOI":"10.1007\/978-3-319-24574-4_28","volume-title":"Proceedings of the 18th International Conference on Medical Image Computing and Computer-Assisted Intervention","author":"O Ronneberger","year":"2015","unstructured":"O. Ronneberger, P. Fischer, T. Brox. U-Net: Convolutional networks for biomedical image segmentation. In Proceedings of the 18th International Conference on Medical Image Computing and Computer-Assisted Intervention, Munich, Germany, pp. 234\u2013241, 2015. DOI: https:\/\/doi.org\/10.1007\/978-3-319-24574-4_28."},{"issue":"6","key":"1552_CR7","doi-asserted-by":"publisher","first-page":"1856","DOI":"10.1109\/TMI.2019.2959609","volume":"39","author":"Z W Zhou","year":"2020","unstructured":"Z. W. Zhou, M. M. R. Siddiquee, N. Tajbakhsh, J. M. Liang. UNet++: Redesigning skip connections to exploit multiscale features in image segmentation. IEEE Transactions on Medical Imaging, vol. 39, no. 6, pp. 1856\u20131867, 2020. DOI: https:\/\/doi.org\/10.1109\/TMI.2019.2959609.","journal-title":"IEEE Transactions on Medical Imaging"},{"key":"1552_CR8","volume-title":"Attention U-Net: Learning where to look for the pancreas","author":"O Oktay","year":"2018","unstructured":"O. Oktay, J. Schlemper, L. Le Folgoc, M. Lee, M. Heinrich, K. Misawa, K. Mori, S. McDonagh, N. Y. Hammerla, B. Kainz, B. Glocker, D. Rueckert. Attention U-Net: Learning where to look for the pancreas, [Online], Available: https:\/\/arxiv.org\/abs\/1804.03999, 2018."},{"key":"1552_CR9","doi-asserted-by":"publisher","first-page":"4002","DOI":"10.1109\/cvpr42600.2020.00406","volume-title":"Proceedings of IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Q B Hou","year":"2020","unstructured":"Q. B. Hou, L. Zhang, M. M. Cheng, J. S. Feng. Strip pooling: Rethinking spatial pooling for scene parsing. In Proceedings of IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Seattle, USA, pp. 4002\u20134011, 2020. DOI: https:\/\/doi.org\/10.1109\/cvpr42600.2020.00406."},{"key":"1552_CR10","doi-asserted-by":"publisher","first-page":"603","DOI":"10.1109\/iccv.2019.00069","volume-title":"Proceedings of IEEE\/CVF International Conference on Computer Vision","author":"Z L Huang","year":"2019","unstructured":"Z. L. Huang, X. G. Wang, L. C. Huang, C. Huang, Y. C. Wei, W. Y. Liu. CCNet: Criss-cross attention for semantic segmentation. In Proceedings of IEEE\/CVF International Conference on Computer Vision, Seoul, Republic of Korea, pp. 603\u2013612, 2019. DOI: https:\/\/doi.org\/10.1109\/iccv.2019.00069."},{"key":"1552_CR11","volume-title":"Proceedings of the 9th International Conference on Learning Representations","author":"A Dosovitskiy","year":"2021","unstructured":"A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. H. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, J. Uszkoreit, N. Houlsby. An image is worth 16 \u00d7 16 words: Transformers for image recognition at scale. In Proceedings of the 9th International Conference on Learning Representations, 2021."},{"key":"1552_CR12","doi-asserted-by":"publisher","first-page":"9992","DOI":"10.1109\/iccv48922.2021.00986","volume-title":"Proceedings of IEEE\/CVF International Conference on Computer Vision","author":"Z Liu","year":"2021","unstructured":"Z. Liu, Y. T. Lin, Y. Cao, H. Hu, Y. X. Wei, Z. Zhang, S. Lin, B. N. Guo. Swin transformer: Hierarchical vision transformer using shifted windows. In Proceedings of IEEE\/CVF International Conference on Computer Vision, Montreal, Canada, pp. 9992\u201310002, 2021. DOI: https:\/\/doi.org\/10.1109\/iccv48922.2021.00986."},{"key":"1552_CR13","volume-title":"TransUNet: Transformers make strong encoders for medical image segmentation","author":"J N Chen","year":"2021","unstructured":"J. N. Chen, Y. Y. Lu, Q. H. Yu, X. D. Luo, E. Adeli, Y. Wang, L. Lu, A. L. Yuille, Y. Y. Zhou. TransUNet: Transformers make strong encoders for medical image segmentation, [Online], Available: https:\/\/arxiv.org\/abs\/2102.04306, 2021."},{"key":"1552_CR14","volume-title":"Pyramid medical transformer for medical image segmentation","author":"Z Z Zhang","year":"2021","unstructured":"Z. Z. Zhang, W. X. Zhang. Pyramid medical transformer for medical image segmentation, [Online], Available: https:\/\/arxiv.org\/abs\/2104.14702, 2021."},{"key":"1552_CR15","doi-asserted-by":"publisher","first-page":"109","DOI":"10.1007\/978-3-030-87193-2_11","volume-title":"Proceedings of the 24th International Conference on Medical Image Computing and Computer Assisted Intervention","author":"W X Wang","year":"2021","unstructured":"W. X. Wang, C. Chen, M. Ding, H. Yu, S. Zha, J. Y. Li. TransBTS: Multimodal brain tumor segmentation using transformer. In Proceedings of the 24th International Conference on Medical Image Computing and Computer Assisted Intervention, Strasbourg, France, pp. 109\u2013119, 2021. DOI: https:\/\/doi.org\/10.1007\/978-3-030-87193-2_11."},{"key":"1552_CR16","doi-asserted-by":"publisher","first-page":"1748","DOI":"10.1109\/wacv51458.2022.00181","volume-title":"Proceedings of IEEE\/CVF Winter Conference on Applications of Computer Vision","author":"A Hatamizadeh","year":"2022","unstructured":"A. Hatamizadeh, Y. C. Tang, V. Nath, D. Yang, A. Myronenko, B. Landman, H. R. Roth, D. G. Xu. UNETR: Transformers for 3D medical image segmentation. In Proceedings of IEEE\/CVF Winter Conference on Applications of Computer Vision, Waikoloa, USA, pp. 1748\u20131758, 2022. DOI: https:\/\/doi.org\/10.1109\/wacv51458.2022.00181."},{"key":"1552_CR17","volume-title":"Axial attention in multidimensional transformers","author":"J Ho","year":"2019","unstructured":"J. Ho, N. Kalchbrenner, D. Weissenborn, T. Salimans. Axial attention in multidimensional transformers, [Online], Available: https:\/\/arxiv.org\/abs\/1912.12180, 2019."},{"key":"1552_CR18","doi-asserted-by":"publisher","first-page":"36","DOI":"10.1007\/978-3-030-87193-2_4","volume-title":"Proceedings of the 24th International Conference on Medical Image Computing and Computer Assisted Intervention","author":"J M J Valanarasu","year":"2021","unstructured":"J. M. J. Valanarasu, P. Oza, I. Hacihaliloglu, V. M. Patel. Medical transformer: Gated axial-attention for medical image segmentation. In Proceedings of the 24th International Conference on Medical Image Computing and Computer Assisted Intervention, Strasbourg, France, pp. 36\u201346, 2021. DOI: https:\/\/doi.org\/10.1007\/978-3-030-87193-2_4."},{"key":"1552_CR19","doi-asserted-by":"publisher","first-page":"108","DOI":"10.1007\/978-3-030-58548-8_7","volume-title":"Proceedings of the 16th European Conference on Computer Vision","author":"H Y Wang","year":"2020","unstructured":"H. Y. Wang, Y. K. Zhu, B. Green, H. Adam, A. Yuille, L. C. Chen. Axial-DeepLab: Stand-alone axial-attention for panoptic segmentation. In Proceedings of the 16th European Conference on Computer Vision, Glasgow, UK, pp. 108\u2013126, 2020. DOI: https:\/\/doi.org\/10.1007\/978-3-030-58548-8_7."},{"key":"1552_CR20","volume-title":"Proceedings of the 36th International Conference on Neural Information Processing Systems","author":"M H Guo","year":"2022","unstructured":"M. H. Guo, C. Z. Lu, Q. B. Hou, Z. N. Liu, M. M. Cheng, S. M. Hu. SegNeXt: Rethinking convolutional attention design for semantic segmentation. In Proceedings of the 36th International Conference on Neural Information Processing Systems, New Orleans, USA, Article number 84, 2022."},{"key":"1552_CR21","doi-asserted-by":"publisher","unstructured":"A. L. Lin, B. Z. Chen, J. Y. Xu, Z. Zhang, G. M. Lu, D. Zhang. DS-TransUNet: Dual swin transformer U-Net for medical image segmentation. IEEE Transactions on Instrumentation and Measurement, vol. 71, Article number 4005615, 2022. DOI: https:\/\/doi.org\/10.1109\/TIM.2022.3178991.","DOI":"10.1109\/TIM.2022.3178991"},{"key":"1552_CR22","doi-asserted-by":"publisher","unstructured":"Q. Xu, Z. C. Ma, N. He, W. T. Duan. DCSAU-Net: A deeper and more compact split-attention U-Net for medical image segmentation. Computers in Biology and Medicine, vol. 154, Article number 106626, 2023. DOI: https:\/\/doi.org\/10.1016\/j.compbiomed.2023.106626.","DOI":"10.1016\/j.compbiomed.2023.106626"},{"issue":"11","key":"1552_CR23","doi-asserted-by":"publisher","first-page":"9375","DOI":"10.1109\/TNNLS.2022.3159394","volume":"34","author":"N K Tomar","year":"2023","unstructured":"N. K. Tomar, D. Jha, M. A. Riegler, H. D. Johansen, D. Johansen, J. Rittscher, P. Halvorsen, S. Ali. FANet: A feedback attention network for improved biomedical image segmentation. IEEE Transactions on Neural Networks and Learning Systems, vol. 34, no. 11, pp. 9375\u20139388, 2023. DOI: https:\/\/doi.org\/10.1109\/TNNLS.2022.3159394.","journal-title":"IEEE Transactions on Neural Networks and Learning Systems"},{"key":"1552_CR24","volume-title":"Focal-UNet: UNet-like focal modulation for medical image segmentation","author":"M Naderi","year":"2022","unstructured":"M. Naderi, M. H. Givkashi, F. Piri, N. Karimi, S. Samavi. Focal-UNet: UNet-like focal modulation for medical image segmentation, [Online], Available: https:\/\/arxiv.org\/abs\/2212.09263, 2022."},{"key":"1552_CR25","doi-asserted-by":"publisher","first-page":"2441","DOI":"10.1609\/aaai.v36i3.20144","volume-title":"Proceedings of the 36th AAAI Conference on Artificial Intelligence","author":"H N Wang","year":"2022","unstructured":"H. N. Wang, P. Cao, J. Q. Wang, O. R. Zaiane. UCTrans-Net: Rethinking the skip connections in U-Net from a channel-wise perspective with transformer. In Proceedings of the 36th AAAI Conference on Artificial Intelligence, pp. 2441\u20132449, 2022. DOI: https:\/\/doi.org\/10.1609\/aaai.v36i3.20144."},{"key":"1552_CR26","doi-asserted-by":"publisher","first-page":"263","DOI":"10.1007\/978-3-030-59725-2_26","volume-title":"Proceedings of the 23rd International Conference on Medical Image Computing and Computer Assisted Intervention","author":"D P Fan","year":"2020","unstructured":"D. P. Fan, G. P. Ji, T. Zhou, G. Chen, H. Z. Fu, J. B. Shen, L. Shao. PraNet: Parallel reverse attention network for polyp segmentation. In Proceedings of the 23rd International Conference on Medical Image Computing and Computer Assisted Intervention, Lima, Peru, pp. 263\u2013273, 2020. DOI: https:\/\/doi.org\/10.1007\/978-3-030-59725-2_26."},{"key":"1552_CR27","doi-asserted-by":"publisher","first-page":"14","DOI":"10.1007\/978-3-030-87193-2_2","volume-title":"Proceedings of the 24th International Conference on Medical Image Computing and Computer Assisted Intervention","author":"Y D Zhang","year":"2021","unstructured":"Y. D. Zhang, H. Y. Liu, Q. Hu. TransFuse: Fusing transformers and CNNs for medical image segmentation. In Proceedings of the 24th International Conference on Medical Image Computing and Computer Assisted Intervention, Strasbourg, France, pp. 14\u201324, 2021. DOI: https:\/\/doi.org\/10.1007\/978-3-030-87193-2_2."},{"key":"1552_CR28","volume-title":"HarDNet-MSEG: A simple encoder-decoder polyp segmentation neural network that achieves over 0.9 mean dice and 86 FPS","author":"C H Huang","year":"2021","unstructured":"C. H. Huang, H. Y. Wu, Y. L. Lin. HarDNet-MSEG: A simple encoder-decoder polyp segmentation neural network that achieves over 0.9 mean dice and 86 FPS, [Online], Available: https:\/\/arxiv.org\/abs\/2101.07172, 2021."},{"key":"1552_CR29","doi-asserted-by":"publisher","first-page":"205","DOI":"10.1007\/978-3-031-25066-8_9","volume-title":"Proceedings of Computer Vision \u2013 ECCV Workshops","author":"H Cao","year":"2022","unstructured":"H. Cao, Y. Y. Wang, J. Chen, D. S. Jiang, X. P. Zhang, Q. Tian, M. N. Wang. Swin-Unet: UNet-like pure transformer for medical image segmentation. In Proceedings of Computer Vision \u2013 ECCV Workshops, Tel Aviv, Israel, pp. 205\u2013218, 2022. DOI: https:\/\/doi.org\/10.1007\/978-3-031-25066-8_9."},{"issue":"5","key":"1552_CR30","doi-asserted-by":"publisher","first-page":"1484","DOI":"10.1109\/TMI.2022.3230943","volume":"42","author":"X H Huang","year":"2023","unstructured":"X. H. Huang, Z. F. Deng, D. D. Li, X. G. Yuan, Y. Fu. MISSFormer: An effective transformer for 2D medical image segmentation. IEEE Transactions on Medical Imaging, vol. 42, no. 5, pp. 1484\u20131494, 2023. DOI: https:\/\/doi.org\/10.1109\/TMI.2022.3230943.","journal-title":"IEEE Transactions on Medical Imaging"},{"key":"1552_CR31","doi-asserted-by":"publisher","first-page":"327","DOI":"10.1109\/itme.2018.00080","volume-title":"Proceedings of the 9th International Conference on Information Technology in Medicine and Education","author":"X Xiao","year":"2018","unstructured":"X. Xiao, S. Lian, Z. M. Luo, S. Z. Li. Weighted Res-UNet for high-quality retina vessel segmentation. In Proceedings of the 9th International Conference on Information Technology in Medicine and Education, Hangzhou, China, pp. 327\u2013331, 2018. DOI: https:\/\/doi.org\/10.1109\/itme.2018.00080."},{"key":"1552_CR32","doi-asserted-by":"publisher","first-page":"4731","DOI":"10.1609\/aaai.v38i5.28274","volume-title":"Proceedings of the 38th AAAI Conference on Artificial Intelligence","author":"H Shao","year":"2024","unstructured":"H. Shao, Y. Zhang, Q. B. Hou. Polyper: Boundary sensitive polyp segmentation. In Proceedings of the 38th AAAI Conference on Artificial Intelligence, Vancouver, Canada, pp. 4731\u20134739, 2024. DOI: https:\/\/doi.org\/10.1609\/aaai.v38i5.28274."},{"key":"1552_CR33","doi-asserted-by":"publisher","first-page":"363","DOI":"10.1007\/978-3-030-59719-1_36","volume-title":"Proceedings of the 23rd International Conference on Medical Image Computing and Computer Assisted Intervention","author":"J M J Valanarasu","year":"2020","unstructured":"J. M. J. Valanarasu, V. A. Sindagi, I. Hacihaliloglu, V. M. Patel. KiU-Net: Towards accurate segmentation of biomedical images using over-complete representations. In Proceedings of the 23rd International Conference on Medical Image Computing and Computer Assisted Intervention, Lima, Peru, pp. 363\u2013373, 2020. DOI: https:\/\/doi.org\/10.1007\/978-3-030-59719-1_36."},{"issue":"6","key":"1552_CR34","doi-asserted-by":"publisher","first-page":"1596","DOI":"10.1109\/TMI.2022.3143833","volume":"41","author":"S Q Huang","year":"2022","unstructured":"S. Q. Huang, J. N. Li, Y. Z. Xiao, N. Shen, T. F. Xu. RT-Net: Relation transformer network for diabetic retinopathy multi-lesion segmentation. IEEE Transactions on Medical Imaging, vol. 41, no. 6, pp. 1596\u20131607, 2022. DOI: https:\/\/doi.org\/10.1109\/TMI.2022.3143833.","journal-title":"IEEE Transactions on Medical Imaging"},{"key":"1552_CR35","doi-asserted-by":"publisher","first-page":"78","DOI":"10.1007\/978-3-030-87193-2_8","volume-title":"Proceedings of the 24th International Conference on Medical Image Computing and Computer Assisted Intervention","author":"D Karimi","year":"2021","unstructured":"D. Karimi, S. D. Vasylechko, A. Gholipour. Convolution-free medical image segmentation using transformers. In Proceedings of the 24th International Conference on Medical Image Computing and Computer Assisted Intervention, Strasbourg, France, pp. 78\u201388, 2021. DOI: https:\/\/doi.org\/10.1007\/978-3-030-87193-2_8."},{"key":"1552_CR36","doi-asserted-by":"publisher","first-page":"3270","DOI":"10.1109\/wacv51458.2022.00333","volume-title":"Proceedings of IEEE\/CVF Winter Conference on Applications of Computer Vision","author":"X Y Yan","year":"2022","unstructured":"X. Y. Yan, H. Tang, S. L. Sun, H. Y. Ma, D. Y. Kong, X. H. Xie. AFTer-UNet: Axial fusion transformer UNet for medical image segmentation. In Proceedings of IEEE\/CVF Winter Conference on Applications of Computer Vision, Waikoloa, USA, pp. 3270\u20133280, 2022. DOI: https:\/\/doi.org\/10.1109\/wacv51458.2022.00333."},{"key":"1552_CR37","doi-asserted-by":"publisher","unstructured":"F. Shamshad, S. Khan, S. W. Zamir, M. H. Khan, M. Hayat, F. S. Khan, H. Z. Fu. Transformers in medical imaging: A survey. Medical Image Analysis, vol. 88, Article number 102802, 2023. DOI: https:\/\/doi.org\/10.1016\/j.media.2023.102802.","DOI":"10.1016\/j.media.2023.102802"},{"key":"1552_CR38","doi-asserted-by":"publisher","first-page":"519","DOI":"10.1007\/978-3-319-46487-9_32","volume-title":"Proceedings of the 14th European Conference on Computer Vision","author":"G Ghiasi","year":"2016","unstructured":"G. Ghiasi, C. C. Fowlkes. Laplacian pyramid reconstruction and refinement for semantic segmentation. In Proceedings of the 14th European Conference on Computer Vision, Amsterdam, The Netherlands, pp. 519\u2013534, 2016. DOI: https:\/\/doi.org\/10.1007\/978-3-319-46487-9_32."},{"key":"1552_CR39","doi-asserted-by":"publisher","first-page":"326","DOI":"10.1007\/978-3-030-87193-2_31","volume-title":"Proceedings of the 24th International Conference on Medical Image Computing and Computer Assisted Intervention","author":"Y F Ji","year":"2021","unstructured":"Y. F. Ji, R. M. Zhang, H. J. Wang, Z. Li, L. Y. Wu, S. T. Zhang, P. Luo. Multi-compound transformer for accurate biomedical image segmentation. In Proceedings of the 24th International Conference on Medical Image Computing and Computer Assisted Intervention, Strasbourg, France, pp. 326\u2013336, 2021. DOI: https:\/\/doi.org\/10.1007\/978-3-030-87193-2_31."},{"key":"1552_CR40","doi-asserted-by":"publisher","first-page":"171","DOI":"10.1007\/978-3-030-87199-4_16","volume-title":"Proceedings of the 24th International Conference on Medical Image Computing and Computer Assisted Intervention","author":"Y T Xie","year":"2021","unstructured":"Y. T. Xie, J. P. Zhang, C. H. Shen, Y. Xia. CoTr: Efficiently bridging CNN and transformer for 3D medical image segmentation. In Proceedings of the 24th International Conference on Medical Image Computing and Computer Assisted Intervention, Strasbourg, France, pp. 171\u2013180, 2021. DOI: https:\/\/doi.org\/10.1007\/978-3-030-87199-4_16."},{"issue":"10","key":"1552_CR41","doi-asserted-by":"publisher","first-page":"3349","DOI":"10.1109\/TPAMI.2020.2983686","volume":"43","author":"J D Wang","year":"2021","unstructured":"J. D. Wang, K. Sun, T. H. Cheng, B. R. Jiang, C. R. Deng, Y. Zhao, D. Liu, Y. D. Mu, M. K. Tan, X. G. Wang, W. Y. Liu, B. Xiao. Deep high-resolution representation learning for visual recognition. IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 43, no. 10, pp. 3349\u20133364, 2021. DOI: https:\/\/doi.org\/10.1109\/TPAMI.2020.2983686.","journal-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence"},{"key":"1552_CR42","doi-asserted-by":"publisher","first-page":"82","DOI":"10.1109\/cvpr.2019.00017","volume-title":"Proceedings of IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"C X Liu","year":"2019","unstructured":"C. X. Liu, L. C. Chen, F. Schroff, H. Adam, W. Hua, A. L. Yuille, Li F.-F. Auto-DeepLab: Hierarchical neural architecture search for semantic image segmentation. In Proceedings of IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Long Beach, USA, pp. 82\u201392, 2019. DOI: https:\/\/doi.org\/10.1109\/cvpr.2019.00017."},{"key":"1552_CR43","doi-asserted-by":"publisher","first-page":"7794","DOI":"10.1109\/cvpr.2018.00813","volume-title":"Proceedings of IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"X L Wang","year":"2018","unstructured":"X. L. Wang, R. Girshick, A. Gupta, K. M. He. Non-local neural networks. In Proceedings of IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Salt Lake City, USA, pp. 7794\u20137803, 2018. DOI: https:\/\/doi.org\/10.1109\/cvpr.2018.00813."},{"key":"1552_CR44","first-page":"6000","volume-title":"Proceedings of the 31st International Conference on Neural Information Processing Systems","author":"A Vaswani","year":"2017","unstructured":"A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, \u0141. Kaiser, I. Polosukhin. Attention is all you need. In Proceedings of the 31st International Conference on Neural Information Processing Systems, Long Beach, USA, pp. 6000\u20136010, 2017."},{"key":"1552_CR45","doi-asserted-by":"publisher","first-page":"7132","DOI":"10.1109\/cvpr.2018.00745","volume-title":"Proceedings of IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"J Hu","year":"2018","unstructured":"J. Hu, L. Shen, G. Sun. Squeeze-and-excitation networks. In Proceedings of IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Salt Lake City, USA, pp. 7132\u20137141, 2018. DOI: https:\/\/doi.org\/10.1109\/cvpr.2018.00745."},{"key":"1552_CR46","doi-asserted-by":"publisher","first-page":"3138","DOI":"10.1109\/wacv48630.2021.00318","volume-title":"Proceedings of IEEE Winter Conference on Applications of Computer Vision","author":"D Misra","year":"2021","unstructured":"D. Misra, T. Nalamada, A. U. Arasanipalai, Q. B. Hou. Rotate to attend: Convolutional triplet attention module. In Proceedings of IEEE Winter Conference on Applications of Computer Vision, Waikoloa, USA, pp. 3138\u20133147, 2021. DOI: https:\/\/doi.org\/10.1109\/wacv48630.2021.00318."},{"key":"1552_CR47","doi-asserted-by":"publisher","first-page":"13708","DOI":"10.1109\/cvpr46437.2021.01350","volume-title":"Proceedings of IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Q B Hou","year":"2021","unstructured":"Q. B. Hou, D. Q. Zhou, J. S. Feng. Coordinate attention for efficient mobile network design. In Proceedings of IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Nashville, USA, pp. 13708\u201313717, 2021. DOI: https:\/\/doi.org\/10.1109\/cvpr46437.2021.01350."},{"key":"1552_CR48","doi-asserted-by":"publisher","first-page":"763","DOI":"10.1109\/iccv48922.2021.00082","volume-title":"Proceedings of IEEE\/CVF International Conference on Computer Vision","author":"Z Q Qin","year":"2021","unstructured":"Z. Q. Qin, P. Y. Zhang, F. Wu, X. Li. FcaNet: Frequency channel attention networks. In Proceedings of IEEE\/CVF International Conference on Computer Vision, Montreal, Canada, pp. 763\u2013772, 2021. DOI: https:\/\/doi.org\/10.1109\/iccv48922.2021.00082."},{"key":"1552_CR49","volume-title":"Proceedings of the 35th International Conference on Neural Information Processing Systems","author":"E Z Xie","year":"2021","unstructured":"E. Z. Xie, W. H. Wang, Z. D. Yu, A. Anandkumar, J. M. Alvarez, P. Luo. SegFormer: Simple and efficient design for semantic segmentation with transformers. In Proceedings of the 35th International Conference on Neural Information Processing Systems, Article number 924, 2021."},{"key":"1552_CR50","doi-asserted-by":"publisher","first-page":"548","DOI":"10.1109\/iccv48922.2021.00061","volume-title":"Proceedings of IEEE\/CVF International Conference on Computer Vision","author":"W H Wang","year":"2021","unstructured":"W. H. Wang, E. Z. Xie, X. Li, D. P. Fan, K. T. Song, D. Liang, T. Lu, P. Luo, L. Shao. Pyramid vision transformer: A versatile backbone for dense prediction without convolutions. In Proceedings of IEEE\/CVF International Conference on Computer Vision, Montreal, Canada, pp. 548\u2013558, 2021. DOI: https:\/\/doi.org\/10.1109\/iccv48922.2021.00061."},{"key":"1552_CR51","doi-asserted-by":"publisher","unstructured":"Y. T. Shi, W. J. Wang, M. Z. Yuan, X. H. Wang. Self-paced dual-axis attention fusion network for retinal vessel segmentation. Electronics, vol. 12, no. 9, Article number 2107, 2023. DOI: https:\/\/doi.org\/10.3390\/electronics12092107.","DOI":"10.3390\/electronics12092107"},{"key":"1552_CR52","doi-asserted-by":"publisher","unstructured":"Z. Y. Ju, Z. C. Zhou, Z. X. Qi, C. Yi. H2MaT-Unet: Hierarchical hybrid multi-axis transformer based Unet for medical image segmentation. Computers in Biology and Medicine, vol. 174, Article number 108387, 2024. DOI: https:\/\/doi.org\/10.1016\/j.compbiomed.2024.108387.","DOI":"10.1016\/j.compbiomed.2024.108387"},{"issue":"4","key":"1552_CR53","doi-asserted-by":"publisher","first-page":"834","DOI":"10.1109\/TPAMI.2017.2699184","volume":"40","author":"L C Chen","year":"2018","unstructured":"L. C. Chen, G. Papandreou, I. Kokkinos, K. Murphy, A. L. Yuille. DeepLab: Semantic image segmentation with deep convolutional nets, atrous convolution, and fully connected CRFs. IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 40, no. 4, pp. 834\u2013848, 2018. DOI: https:\/\/doi.org\/10.1109\/TPAMI.2017.2699184.","journal-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence"},{"issue":"4","key":"1552_CR54","doi-asserted-by":"publisher","first-page":"815","DOI":"10.1109\/TPAMI.2018.2815688","volume":"41","author":"Q B Hou","year":"2019","unstructured":"Q. B. Hou, M. M. Cheng, X. W. Hu, A. Borji, Z. W. Tu, P. H. S. Torr. Deeply supervised salient object detection with short connections. IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 41, no. 4, pp. 815\u2013828, 2019. DOI: https:\/\/doi.org\/10.1109\/TPAMI.2018.2815688.","journal-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence"},{"key":"1552_CR55","volume-title":"Mmsegmentation: OpenMMLab semantic segmentation toolbox and benchmark","author":"M. Contributors","year":"2022","unstructured":"M. Contributors. Mmsegmentation: OpenMMLab semantic segmentation toolbox and benchmark, [Online], Available: https:\/\/github.com\/open-mmlab\/mmsegmentation, 2022."},{"key":"1552_CR56","doi-asserted-by":"publisher","first-page":"558","DOI":"10.1109\/cbms49503.2020.00111","volume-title":"Proceedings of the 33rd International Symposium on Computer-Based Medical Systems","author":"D Jha","year":"2020","unstructured":"D. Jha, M. A. Riegler, D. Johansen, P. Halvorsen, H. D. Johansen. DoubleU-Net: A deep convolutional neural network for medical image segmentation. In Proceedings of the 33rd International Symposium on Computer-Based Medical Systems, Rochester, USA, pp. 558\u2013564, 2020. DOI: https:\/\/doi.org\/10.1109\/cbms49503.2020.00111."},{"issue":"2","key":"1552_CR57","doi-asserted-by":"publisher","first-page":"203","DOI":"10.1038\/s41592-020-01008-z","volume":"18","author":"F Isensee","year":"2021","unstructured":"F. Isensee, P. F. Jaeger, S. A. A. Kohl, J. Petersen, K. H. Maier-Hein. nnU-Net: A self-configuring method for deep learning-based biomedical image segmentation. Nature Methods, vol. 18, no. 2, pp. 203\u2013211, 2021. DOI: https:\/\/doi.org\/10.1038\/s41592-020-01008-z.","journal-title":"Nature Methods"},{"key":"1552_CR58","doi-asserted-by":"publisher","first-page":"197","DOI":"10.1016\/j.media.2019.01.012","volume":"53","author":"J Schlemper","year":"2019","unstructured":"J. Schlemper, O. Oktay, M. Schaap, M. Heinrich, B. Kainz, B. Glocker, D. Rueckert. Attention gated networks: Learning to leverage salient regions in medical images. Medical Image Analysis, vol. 53, pp. 197\u2013207, 2019. DOI: https:\/\/doi.org\/10.1016\/j.media.2019.01.012.","journal-title":"Medical Image Analysis"},{"key":"1552_CR59","doi-asserted-by":"publisher","first-page":"6191","DOI":"10.1109\/wacv56688.2023.00614","volume-title":"Proceedings of IEEE\/CVF Winter Conference on Applications of Computer Vision","author":"M Heidari","year":"2023","unstructured":"M. Heidari, A. Kazerouni, M. Soltany, R. Azad, E. K. Aghdam, J. Cohen-Adad, D. Merhof. HiFormer: Hierarchical multi-scale representations using transformers for medical image segmentation. In Proceedings of IEEE\/CVF Winter Conference on Applications of Computer Vision, Waikoloa, USA, pp. 6191\u20136201, 2023. DOI: https:\/\/doi.org\/10.1109\/wacv56688.2023.00614."},{"key":"1552_CR60","doi-asserted-by":"publisher","first-page":"302","DOI":"10.1007\/978-3-030-32239-7_34","volume-title":"Proceedings of the 22nd International Conference on Medical Image Computing and Computer Assisted Intervention","author":"Y Q Fang","year":"2019","unstructured":"Y. Q. Fang, C. Chen, Y. X. Yuan, K. Y. Tong. Selective feature aggregation network with area-boundary constraints for polyp segmentation. In Proceedings of the 22nd International Conference on Medical Image Computing and Computer Assisted Intervention, Shenzhen, China, pp. 302\u2013310, 2019. DOI: https:\/\/doi.org\/10.1007\/978-3-030-32239-7_34."},{"key":"1552_CR61","volume-title":"Proceedings of the 5th International Conference on Learning Representations","author":"S Zagoruyko","year":"2017","unstructured":"S. Zagoruyko, N. Komodakis. Paying more attention to attention: Improving the performance of convolutional neural networks via attention transfer. In Proceedings of the 5th International Conference on Learning Representations, Toulon, France, 2017."},{"key":"1552_CR62","doi-asserted-by":"publisher","first-page":"9166","DOI":"10.1109\/iccv.2019.00926","volume-title":"Proceedings of IEEE\/CVF International Conference on Computer Vision","author":"X Li","year":"2019","unstructured":"X. Li, Z. S. Zhong, J. L. Wu, Y. B. Yang, Z. C. Lin, H. Liu. Expectation-maximization attention networks for semantic segmentation. In Proceedings of IEEE\/CVF International Conference on Computer Vision, Seoul, Republic of Korea, pp. 9166\u20139175, 2019. DOI: https:\/\/doi.org\/10.1109\/iccv.2019.00926."},{"key":"1552_CR63","volume-title":"Proceedings of the 9th International Conference on Learning Representations","author":"Z Y Geng","year":"2021","unstructured":"Z. Y. Geng, M. H. Guo, H. X. Chen, X. Li, K. Wei, Z. C. Lin. Is attention better than matrix decomposition? In Proceedings of the 9th International Conference on Learning Representations, 2021."},{"key":"1552_CR64","doi-asserted-by":"publisher","first-page":"770","DOI":"10.1109\/cvpr.2016.90","volume-title":"Proceedings of IEEE Conference on Computer Vision and Pattern Recognition","author":"K M He","year":"2016","unstructured":"K. M. He, X. Y. Zhang, S. Q. Ren, J. Sun. Deep residual learning for image recognition. In Proceedings of IEEE Conference on Computer Vision and Pattern Recognition, Las Vegas, USA, pp. 770\u2013778, 2016. DOI: https:\/\/doi.org\/10.1109\/cvpr.2016.90."}],"updated-by":[{"DOI":"10.1007\/s11633-026-1652-y","type":"correction","label":"Correction","source":"publisher","updated":{"date-parts":[[2026,5,15]],"date-time":"2026-05-15T00:00:00Z","timestamp":1778803200000}}],"container-title":["Machine Intelligence Research"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s11633-025-1552-6.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s11633-025-1552-6","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s11633-025-1552-6.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,5,21]],"date-time":"2026-05-21T11:02:54Z","timestamp":1779361374000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s11633-025-1552-6"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,5,27]]},"references-count":64,"journal-issue":{"issue":"3","published-print":{"date-parts":[[2025,6]]}},"alternative-id":["1552"],"URL":"https:\/\/doi.org\/10.1007\/s11633-025-1552-6","relation":{},"ISSN":["2731-538X","2731-5398"],"issn-type":[{"value":"2731-538X","type":"print"},{"value":"2731-5398","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,5,27]]},"assertion":[{"value":"13 May 2024","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"26 February 2025","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"27 May 2025","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"21 May 2026","order":5,"name":"change_date","label":"Change Date","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"Correction","order":6,"name":"change_type","label":"Change Type","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"A Correction to this paper has been published:","order":7,"name":"change_details","label":"Change Details","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"https:\/\/doi.org\/10.1007\/s11633-026-1652-y","URL":"https:\/\/doi.org\/10.1007\/s11633-026-1652-y","order":8,"name":"change_details","label":"Change Details","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"15 May 2026","order":9,"name":"change_date","label":"Change Date","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"Correction","order":10,"name":"change_type","label":"Change Type","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"A Correction to this paper has been published:","order":11,"name":"change_details","label":"Change Details","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"https:\/\/doi.org\/10.1007\/s11633-026-1652-y","URL":"https:\/\/doi.org\/10.1007\/s11633-026-1652-y","order":12,"name":"change_details","label":"Change Details","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"The authors declared that they have no conflicts of interest to this work.","order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations of conflict of interest"}}]}}