{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,15]],"date-time":"2026-07-15T18:24:31Z","timestamp":1784139871599,"version":"3.55.0"},"reference-count":41,"publisher":"Frontiers Media SA","license":[{"start":{"date-parts":[[2025,12,1]],"date-time":"2025-12-01T00:00:00Z","timestamp":1764547200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":["frontiersin.org"],"crossmark-restriction":true},"short-container-title":["Front. Neurorobot."],"abstract":"<jats:sec>\n                    <jats:title>Introduction<\/jats:title>\n                    <jats:p>The combination of CNN and Transformer has attracted much attention for medical image segmentation due to its superior performance at present. However, the segmentation performance is affected by limitations such as the local receptive field and static weights of CNN convolution operations, as well as insufficient information exchange between Transformer local regions.<\/jats:p>\n                  <\/jats:sec>\n                  <jats:sec>\n                    <jats:title>Methods<\/jats:title>\n                    <jats:p>To address these issues, an integrated attention mechanism and pyramid pooling network is proposed in this paper. Firstly, an efficient channel attention mechanism is embedded into CNN to extract more comprehensive image features. Then, CBAM_ASPP module is introduced into the bottleneck layer to obtain multi-scale context information. Finally, in order to address the limitations of traditional convolution, depthwise separable convolution is used to achieve a lightweight network.<\/jats:p>\n                  <\/jats:sec>\n                  <jats:sec>\n                    <jats:title>Results<\/jats:title>\n                    <jats:p>The experiments based on the Synapse multi organ segmentation dataset and ACDC dataset showed that the proposed IAP-TransUNet achieved Dice similarity coefficients (DSCs) of 78.85% and 90.46%, respectively. Compared with the state-of-the-art method, for the Synapse multi organ segmentation dataset, the Hausdorff distance was reduced by 2.92%. For the ACDC dataset, the segmentation accuracy of the left ventricle, myocardium, and right ventricle was improved by 0.14%, 1.89%, and 0.23%, respectively.<\/jats:p>\n                  <\/jats:sec>\n                  <jats:sec>\n                    <jats:title>Discussion<\/jats:title>\n                    <jats:p>The experimental results demonstrate that the proposed network has improved the effectiveness and shows strong performance on both CT and MRI data, which suggests its potential for generalization across different medical imaging modalities.<\/jats:p>\n                  <\/jats:sec>","DOI":"10.3389\/fnbot.2025.1706626","type":"journal-article","created":{"date-parts":[[2025,12,1]],"date-time":"2025-12-01T06:27:40Z","timestamp":1764570460000},"update-policy":"https:\/\/doi.org\/10.3389\/crossmark-policy","source":"Crossref","is-referenced-by-count":2,"title":["IAP-TransUNet: integration of the attention mechanism and pyramid pooling for medical image segmentation"],"prefix":"10.3389","volume":"19","author":[{"given":"Yuxuan","family":"Shi","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Fang","family":"Li","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Shuting","family":"Zhao","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Hongmeng","family":"Yu","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Xinrong","family":"Chen","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Quan","family":"Liu","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1965","published-online":{"date-parts":[[2025,12,1]]},"reference":[{"key":"B1","doi-asserted-by":"publisher","first-page":"123","DOI":"10.1007\/s13246-025-01635-w","article-title":"A review of image processing and analysis of computed tomography images using deep learning methods","volume":"48","author":"Anderson","year":"2025","journal-title":"Phys. Eng. Sci. Med."},{"key":"B2","first-page":"92","article-title":"\u201cOptimizing the dice score and jaccard index for medical image segmentation: theory and practice,\u201d","volume-title":"Medical Image Computing and Computer Assisted Intervention\u2013MICCAI 2019: 22nd International Conference, Shenzhen, China, October 13\u201317, 2019, Proceedings, Part II 22","author":"Bertels","year":"2019"},{"key":"B3","doi-asserted-by":"publisher","first-page":"15934","DOI":"10.1109\/TITS.2024.3417813","article-title":"Sdpt: Semantic-aware dimension-pooling transformer for image segmentation","volume":"25","author":"Cao","year":"2024","journal-title":"IEEE Transac. Intell. Transp. Syst."},{"key":"B4","article-title":"\u201cSwin-UNet: UNet-like pure transformer for medical image segmentation,\u201d","volume-title":"European Conference on Computer Vision","author":"Cao","year":"2022"},{"key":"B5","unstructured":"Chen\n              J.\n            \n            \n              Lu\n              Y.\n            \n            \n              Yu\n              Q.\n            \n            \n              Luo\n              X.\n            \n            \n              Adeli\n              E.\n            \n            \n              Wang\n              Y.\n            \n          \n          Transunet: Transformers make strong encoders for medical image segmentation. arXiv\n          \n          2021"},{"key":"B6","doi-asserted-by":"publisher","first-page":"834","DOI":"10.1109\/TPAMI.2017.2699184","article-title":"Deeplab: Semantic image segmentation with deep convolutional nets, atrous convolution, and fully connected crfs","volume":"40","author":"Chen","year":"2017","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"B7","first-page":"1251","article-title":"\u201cXception: deep learning with depthwise separable convolutions,\u201d","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern recognition","author":"Chollet","year":"2017"},{"key":"B8","unstructured":"An image is worth 16x16 words: Transformers for image recognition at scale\n          \n          \n            \n              Dosovitskiy\n              A.\n            \n            \n              Beyer\n              L.\n            \n            \n              Kolesnikov\n              A.\n            \n            \n              Weissenborn\n              D.\n            \n            \n              Zhai\n              X.\n            \n            \n              Unterthiner\n              T.\n            \n          \n          arXiv\n          \n          2020"},{"key":"B9","first-page":"656","article-title":"\u201cDomain adaptive relational reasoning for 3d multi-organ segmentation,\u201d","volume-title":"Medical Image Computing and Computer Assisted Intervention\u2013MICCAI 2020: 23rd International Conference, Lima, Peru, October 4\u20138, 2020, Proceedings, Part I 23","author":"Fu","year":"2020"},{"key":"B10","doi-asserted-by":"publisher","first-page":"568","DOI":"10.1109\/JBHI.2019.2912935","article-title":"Fully dense UNet for 2-D sparse photoacoustic tomography artifact removal","volume":"24","author":"Guan","year":"2019","journal-title":"IEEE J. Biomed. Health Inform."},{"key":"B11","doi-asserted-by":"publisher","first-page":"110542","DOI":"10.1016\/j.clinimag.2025.110542","article-title":"Deep learning based colorectal cancer detection in medical images: a comprehensive analysis of datasets, methods, and future directions","volume":"125","author":"G\u00fclmez","year":"2025","journal-title":"Clin. Imag."},{"key":"B12","doi-asserted-by":"publisher","first-page":"87","DOI":"10.1109\/TPAMI.2022.3152247","article-title":"A survey on vision transformer","volume":"45","author":"Han","year":"2022","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"B13","doi-asserted-by":"publisher","first-page":"1904","DOI":"10.1109\/TPAMI.2015.2389824","article-title":"Spatial pyramid pooling in deep convolutional networks for visual recognition","volume":"37","author":"He","year":"2015","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"B14","first-page":"770","article-title":"\u201cDeep residual learning for image recognition,\u201d","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition","author":"He","year":"2016"},{"key":"B15","doi-asserted-by":"publisher","first-page":"4806","DOI":"10.1109\/ACCESS.2019.2962617","article-title":"The real-world-weight cross-entropy loss function: modeling the costs of mislabeling","volume":"8","author":"Ho","year":"2019","journal-title":"IEEE Access"},{"key":"B16","first-page":"7132","article-title":"\u201cSqueeze-and-excitation networks,\u201d","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition","author":"Hu","year":"2018"},{"key":"B17","first-page":"1055","article-title":"\u201cUnet 3+: a full-scale connected unet for medical image segmentation,\u201d","volume-title":"ICASSP 2020-2020 IEEE international conference on acoustics, speech and signal processing (ICASSP)","author":"Huang","year":"2020"},{"key":"B18","doi-asserted-by":"publisher","first-page":"1281","DOI":"10.3390\/min14121281","article-title":"Res-UNet ensemble learning for semantic segmentation of mineral optical microscopy images","volume":"14","author":"Jiang","year":"2024","journal-title":"Minerals"},{"key":"B19","doi-asserted-by":"publisher","first-page":"862","DOI":"10.1007\/s00103-025-04093-7","article-title":"AI-based applications in medical image computing","volume":"68","author":"Kepp","year":"2025","journal-title":"Bundesgesundheitsblatt-Gesundheitsforschung-Gesundheitsschutz"},{"key":"B20","doi-asserted-by":"publisher","first-page":"101817","DOI":"10.1016\/j.comgeo.2021.101817","article-title":"Between shapes, using the Hausdorff distance","volume":"100","author":"Kreveld","year":"2022","journal-title":"Comput. Geometry"},{"key":"B21","doi-asserted-by":"publisher","first-page":"3503","DOI":"10.1021\/acs.jproteome.9b00411","article-title":"Fertility-GRU: identifying fertility-related proteins by incorporating deep-gated recurrent units and original position-specific scoring matrix profiles","volume":"18","author":"Le","year":"2019","journal-title":"J. Proteome Res."},{"key":"B22","first-page":"3431","article-title":"Fully convolutional networks for semantic segmentation","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition","author":"Long","year":"2015"},{"key":"B23","doi-asserted-by":"crossref","first-page":"565","DOI":"10.1109\/3DV.2016.79","article-title":"\u201cV-net: fully convolutional neural networks for volumetric medical image segmentation,\u201d","volume-title":"2016 Fourth International Conference on 3D Vision (3DV)","author":"Milletari","year":"2016"},{"key":"B24","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1109\/ICT-ROBOT.2018.8549917","article-title":"\u201cAn improvement for medical image analysis using data enhancement techniques in deep learning,\u201d","volume-title":"2018 International Conference on Information and Communication Technology Robotics (ICT-ROBOT)","author":"Namozov","year":"2018"},{"key":"B25","unstructured":"Oktay\n              O.\n            \n            \n              Schlemper\n              J.\n            \n            \n              Folgoc\n              L. L.\n            \n          \n          Attention u-net: Learning where to look for the pancreas. arXiv\n          \n          2018"},{"key":"B26","doi-asserted-by":"publisher","first-page":"14","DOI":"10.1186\/s12938-024-01212-4","article-title":"Advantages of transformer and its application for medical image segmentation: a survey","volume":"23","author":"Pu","year":"2024","journal-title":"Biomed. Eng. Online"},{"key":"B27","first-page":"234","article-title":"\u201cU-net: Convolutional networks for biomedical image segmentation,\u201d","volume-title":"Medical Image Computing and Computer-Assisted Intervention\u2013MICCAI 2015: 18th International Conference, Munich, Germany, October 5-9, 2015, Proceedings, Part III 18","author":"Ronneberger","year":"2015"},{"key":"B28","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1007\/s12065-020-00540-3","article-title":"Convolutional neural networks in medical image understanding: a survey","volume":"15","author":"Sarvamangala","year":"2022","journal-title":"Evol. Intell."},{"key":"B29","doi-asserted-by":"crossref","DOI":"10.1109\/BIBM62325.2024.10822087","article-title":"\u201cBiSeg-SAM: weakly-supervised post-processing framework for boosting binary segmentation in segment anything models,\u201d","volume-title":"2024 IEEE International Conference on Bioinformatics and Biomedicine (BIBM)","author":"Su","year":"2024"},{"key":"B30","doi-asserted-by":"publisher","first-page":"5998","DOI":"10.5555\/3295222.3295349","article-title":"Attention is all you need","volume":"30","author":"Vaswani","year":"2017","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"B31","doi-asserted-by":"publisher","first-page":"51","DOI":"10.1109\/MCE.2022.3181759","article-title":"Lightweight deep learning: an overview","volume":"11","author":"Wang","year":"2022","journal-title":"IEEE Consum. Electron. Magazine"},{"key":"B32","first-page":"11534","article-title":"\u201cECA-Net: efficient channel attention for deep convolutional neural networks,\u201d","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Wang","year":"2020"},{"key":"B33","first-page":"38","article-title":"\u201cTransformers: State-of-the-art natural language processing,\u201d","author":"Wolf","year":"2020","journal-title":"Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrations"},{"key":"B34","first-page":"3","article-title":"\u201cCbam: convolutional block attention module,\u201d","volume-title":"Proceedings of the European conference on computer vision (ECCV)","author":"Woo","year":"2018"},{"key":"B35","doi-asserted-by":"publisher","first-page":"92","DOI":"10.1016\/j.neucom.2020.04.157","article-title":"Convolutional neural networks for medical image analysis: state-of-the-art, comparisons, improvement and perspectives","volume":"444","author":"Yu","year":"2021","journal-title":"Neurocomputing"},{"key":"B36","unstructured":"Zhang\n              C.\n            \n            \n              Liao\n              Q.\n            \n            \n              Rakhlin\n              A.\n            \n          \n          Theory of deep learning IIb: optimization properties of SGD. arXiv\n          \n          2018"},{"key":"B37","doi-asserted-by":"publisher","first-page":"93","DOI":"10.1007\/s12204-021-2264-x","article-title":"Rethinking the dice loss for deep learning lesion segmentation in medical images","volume":"26","author":"Zhang","year":"2021","journal-title":"J. Shanghai Jiaotong Univ. Sci"},{"key":"B38","doi-asserted-by":"publisher","first-page":"40569","DOI":"10.1021\/acsomega.2c05881","article-title":"Improved prediction model of protein and peptide toxicity by integrating channel attention into a convolutional neural network and gated recurrent units","volume":"7","author":"Zhao","year":"2022","journal-title":"ACS Omega"},{"key":"B39","first-page":"6881","article-title":"\u201cRethinking semantic segmentation from a sequence-to-sequence perspective with transformers,\u201d","volume-title":"Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition","author":"Zheng","year":"2021"},{"key":"B40","doi-asserted-by":"publisher","first-page":"1856","DOI":"10.1109\/TMI.2019.2959609","article-title":"Unet++: Redesigning skip connections to exploit multiscale features in image segmentation","volume":"39","author":"Zhou","year":"2019","journal-title":"IEEE Trans. Med. Imag."},{"key":"B41","doi-asserted-by":"crossref","first-page":"791","DOI":"10.1109\/CYBER55403.2022.9907730","article-title":"\u201cSemantic Segmentation of FOD Using an Improved Deeplab V3+ Model,\u201d","volume-title":"2022 12th International Conference on CYBER Technology in Automation, Control, and Intelligent Systems (CYBER)","author":"Zhu","year":"2022"}],"container-title":["Frontiers in Neurorobotics"],"original-title":[],"link":[{"URL":"https:\/\/www.frontiersin.org\/articles\/10.3389\/fnbot.2025.1706626\/full","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,12,1]],"date-time":"2025-12-01T06:27:46Z","timestamp":1764570466000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.frontiersin.org\/articles\/10.3389\/fnbot.2025.1706626\/full"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,12,1]]},"references-count":41,"alternative-id":["10.3389\/fnbot.2025.1706626"],"URL":"https:\/\/doi.org\/10.3389\/fnbot.2025.1706626","relation":{},"ISSN":["1662-5218"],"issn-type":[{"value":"1662-5218","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,12,1]]},"article-number":"1706626"}}