{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,3,13]],"date-time":"2026-03-13T10:54:49Z","timestamp":1773399289831,"version":"3.50.1"},"reference-count":71,"publisher":"Institution of Engineering and Technology (IET)","issue":"1","license":[{"start":{"date-parts":[[2026,3,13]],"date-time":"2026-03-13T00:00:00Z","timestamp":1773360000000},"content-version":"vor","delay-in-days":71,"URL":"http:\/\/creativecommons.org\/licenses\/by-nc-nd\/4.0\/"},{"start":{"date-parts":[[2026,1,1]],"date-time":"2026-01-01T00:00:00Z","timestamp":1767225600000},"content-version":"tdm","delay-in-days":0,"URL":"http:\/\/doi.wiley.com\/10.1002\/tdm_license_1.1"}],"funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["62033009"],"award-info":[{"award-number":["62033009"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["52431012"],"award-info":[{"award-number":["52431012"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["ietresearch.onlinelibrary.wiley.com"],"crossmark-restriction":true},"short-container-title":["IET Image Processing"],"published-print":{"date-parts":[[2026,1]]},"abstract":"<jats:title>ABSTRACT<\/jats:title>\n                  <jats:p>\n                    Underwater images harbor abundant resources for marine science and applications. The fusion of Transformers and CNNs has emerged as a promising approach for salient object detection. These encoders extract hierarchical features at different scales, each exhibiting unique characteristics. Low\u2010level features primarily focus on detailed information, whereas deep features encapsulate higher level semantic information. Motivated by these observations, we design a hierarchical fusion network combining the Swin Transformer with a CNN for underwater salient object detection. Underwater images are preprocessed via an enhancement algorithm utilised to restore attenuated spectral characteristics. At the low\u2010level encoder stages, global features derived from the Swin Transformer and local features extracted from ResNet50 are integrated via an adaptive interactive fusion strategy, which prioritises detail preservation. At the high\u2010level encoder stages, long\u2010range dependency features are aggregated with local features via a cross\u2010attention fusion module, which emphasises semantic interaction. An upsampling and fusion decoder is designed to generate a coarse saliency map, densely integrating the hierarchical features while preserving fine details. Finally, a post\u2010refinement network is utilised to further enhance detection performance by leveraging boundary and contextual information, achieving an accurate salient object segmentation result. Experimental results demonstrate that our method achieves\n                    <jats:italic>F<\/jats:italic>\n                    \u2010measure scores of 0.9322, 0.9247, and 0.8332 on the USOD10K, USOD, and UFO\u2010120 datasets, highlighting the efficacy of hierarchical fusion of Swin Transformer and CNN for underwater salient object detection. The code of our method is available at\n                    <jats:ext-link xmlns:xlink=\"http:\/\/www.w3.org\/1999\/xlink\" xlink:href=\"https:\/\/github.com\/leckie711\/hfusod\">https:\/\/github.com\/leckie711\/hfusod<\/jats:ext-link>\n                    .\n                  <\/jats:p>","DOI":"10.1049\/ipr2.70331","type":"journal-article","created":{"date-parts":[[2026,3,13]],"date-time":"2026-03-13T09:08:18Z","timestamp":1773392898000},"update-policy":"https:\/\/doi.org\/10.1002\/crossmark_policy","source":"Crossref","is-referenced-by-count":0,"title":["HFUSOD: Hierarchical Fusion of Swin Transformer With CNN Network for Underwater Salient Object Detection"],"prefix":"10.1049","volume":"20","author":[{"given":"Weiliang","family":"Huang","sequence":"first","affiliation":[{"name":"Logistics Engineering College Shanghai Maritime University  Shanghai China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-7252-4952","authenticated-orcid":false,"given":"Daqi","family":"Zhu","sequence":"additional","affiliation":[{"name":"The School of Mechanical Engineering University of Shanghai for Science and Technology  Shanghai China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"265","published-online":{"date-parts":[[2026,3,13]]},"reference":[{"key":"e_1_2_11_2_1","doi-asserted-by":"publisher","DOI":"10.1007\/s00521\u2010023\u201009258\u20106"},{"key":"e_1_2_11_3_1","doi-asserted-by":"publisher","DOI":"10.1109\/TCSVT.2024.3412949"},{"key":"e_1_2_11_4_1","doi-asserted-by":"crossref","unstructured":"J.Luiten I. E.Zulfikar andB.Leibe \u201cUnOVOST: Unsupervised Offline Video Object Segmentation and Tracking \u201d inProceedings of the IEEE\/CVF Winter Conference on Applications of Computer Vision 2020 https:\/\/doi.org\/10.1109\/WACV45572.2020.9093285.","DOI":"10.1109\/WACV45572.2020.9093285"},{"key":"e_1_2_11_5_1","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1109\/TGRS.2024.3466389","article-title":"Visual Global\u2010Salient Guided Network for Remote Sensing Image\u2010Text Retrieval","volume":"62","author":"He Y.","year":"2024","journal-title":"IEEE Transactions on Geoscience and Remote Sensing"},{"key":"e_1_2_11_6_1","doi-asserted-by":"publisher","DOI":"10.1049\/ipr2.12972"},{"key":"e_1_2_11_7_1","doi-asserted-by":"publisher","DOI":"10.1109\/TETCI.2024.3359061"},{"key":"e_1_2_11_8_1","doi-asserted-by":"publisher","DOI":"10.1049\/ipr2.12922"},{"key":"e_1_2_11_9_1","doi-asserted-by":"crossref","first-page":"3735","DOI":"10.1007\/s00371-024-03630-w","article-title":"Underwater Image Restoration and Enhancement: A Comprehensive Review of Recent Trends, Challenges, and Applications","volume":"41","author":"Alsakar Y. M.","year":"2025","journal-title":"The Visual Computer"},{"key":"e_1_2_11_10_1","volume-title":"International Conference on Data Mining and Big Data","author":"Luo W.","year":"2024"},{"key":"e_1_2_11_11_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.inffus.2024.102536"},{"key":"e_1_2_11_12_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.inffus.2024.102634"},{"key":"e_1_2_11_13_1","doi-asserted-by":"crossref","unstructured":"M.Wang H.Ding J. H.Liew J.Liu Y.Zhao andY.Wei \u201cSegRefiner: Towards Model\u2010Agnostic Segmentation Refinement With Discrete Diffusion Process \u201d arXiv preprint arXiv:2312.12425 2023.","DOI":"10.52202\/075280-3492"},{"key":"e_1_2_11_14_1","doi-asserted-by":"crossref","unstructured":"L.Chang Y.Wang L.Deng B.Du andC.Xu \u201cWaterDiffusion: Learning a Prior\u2010Involved Unrolling Diffusion for Joint Underwater Saliency Detection and Visual Restoration \u201d inProceedings of the AAAI Conference on Artificial Intelligence 2025 https:\/\/doi.org\/10.1609\/aaai.v39i2.32196.","DOI":"10.1609\/aaai.v39i2.32196"},{"key":"e_1_2_11_15_1","doi-asserted-by":"crossref","unstructured":"A.Khan Z.Rauf A.Sohail et\u00a0al. \u201cA Survey of the Vision Transformers and Their CNN\u2010Transformer Based Variants \u201dArtificial Intelligence Review56 no. Suppl.3(2023):2917\u20132970 https:\/\/doi.org\/10.1007\/s10462\u2010023\u201010595\u20100.","DOI":"10.1007\/s10462-023-10595-0"},{"key":"e_1_2_11_16_1","doi-asserted-by":"crossref","unstructured":"A.Gulati J.Qin C.\u2010C.Chiu et\u00a0al. \u201cConformer: Convolution\u2010Augmented Transformer for Speech Recognition \u201d arXiv preprint arXiv:2005.08100 2020.","DOI":"10.21437\/Interspeech.2020-3015"},{"key":"e_1_2_11_17_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2024.3449959"},{"key":"e_1_2_11_18_1","first-page":"1","article-title":"Heterogeneous Feature Collaboration Network for Salient Object Detection in Optical Remote Sensing Images","volume":"62","author":"Liu Y.","year":"2024","journal-title":"IEEE Transactions on Geoscience and Remote Sensing"},{"key":"e_1_2_11_19_1","unstructured":"M. J.Islam R.Wang andJ.Sattar \u201cSVAM: Saliency\u2010Guided Visual Attention Modeling by Autonomous Underwater Robots \u201d arXiv preprint arXiv:2011.06252 2020."},{"key":"e_1_2_11_20_1","volume-title":"ICASSP 2022\u20132022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)","author":"Chen R.","year":"2022"},{"key":"e_1_2_11_21_1","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2023.3266163"},{"key":"e_1_2_11_22_1","doi-asserted-by":"crossref","unstructured":"Z.Zheng \u201cCoralSCOP: Segment any COral Image on This Planet \u201d inProceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition 2024 https:\/\/doi.org\/10.1109\/CVPR52733.2024.02661.","DOI":"10.1109\/CVPR52733.2024.02661"},{"key":"e_1_2_11_23_1","unstructured":"T.Yan Z.Wan X.Deng P.Zhang Y.Liu andH.Lu \u201cMAS\u2010SAM: Segment Any Marine Animal With Aggregated Features \u201d arXiv preprint arXiv:2404.15700 2024."},{"key":"e_1_2_11_24_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.inffus.2024.102806"},{"key":"e_1_2_11_25_1","doi-asserted-by":"publisher","DOI":"10.1109\/JOE.2023.3344154"},{"key":"e_1_2_11_26_1","doi-asserted-by":"crossref","unstructured":"Q.Wu \u201cEffiSeaNet: Pioneering Lightweight Network for Underwater Salient Object Detection \u201d inProceedings of the Asian Conference on Computer Vision 2024.","DOI":"10.1007\/978-981-96-0911-6_6"},{"key":"e_1_2_11_27_1","doi-asserted-by":"publisher","DOI":"10.1109\/TCSVT.2024.3491907"},{"key":"e_1_2_11_28_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.neucom.2022.10.081"},{"key":"e_1_2_11_29_1","doi-asserted-by":"publisher","DOI":"10.1109\/TCSVT.2023.3307693"},{"key":"e_1_2_11_30_1","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2023.3293759"},{"key":"e_1_2_11_31_1","unstructured":"Y.Yuan P.Gao andX.Tan \u201cM3Net: Multilevel Mixed and Multistage Attention Network for Salient Object Detection \u201d arXiv preprint arXiv:2309.08365 2023."},{"key":"e_1_2_11_32_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.imavis.2022.104595"},{"key":"e_1_2_11_33_1","doi-asserted-by":"publisher","DOI":"10.1109\/TCSVT.2023.3321190"},{"key":"e_1_2_11_34_1","doi-asserted-by":"publisher","DOI":"10.1109\/TMM.2023.3264883"},{"key":"e_1_2_11_35_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.neucom.2023.126329"},{"key":"e_1_2_11_36_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.neucom.2023.126560"},{"key":"e_1_2_11_37_1","doi-asserted-by":"publisher","DOI":"10.1007\/s00521\u2010022\u201007069\u20109"},{"key":"e_1_2_11_38_1","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1109\/TGRS.2024.3362836","article-title":"ASNet: Adaptive Semantic Network Based on Transformer\u2013CNN for Salient Object Detection in Optical Remote Sensing Images","volume":"62","author":"Yan R.","year":"2024","journal-title":"IEEE Transactions on Geoscience and Remote Sensing"},{"key":"e_1_2_11_39_1","doi-asserted-by":"publisher","DOI":"10.1109\/TCSVT.2022.3225865"},{"key":"e_1_2_11_40_1","doi-asserted-by":"crossref","unstructured":"X.Qin \u201cBASNet: Boundary\u2010Aware Salient Object Detection \u201d inProceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition 2019 https:\/\/doi.org\/10.1109\/CVPR.2019.00766.","DOI":"10.1109\/CVPR.2019.00766"},{"key":"e_1_2_11_41_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.eswa.2024.123903"},{"key":"e_1_2_11_42_1","unstructured":"Z.Tian H.Zhao M.Shu et\u00a0al. \u201cRegion Refinement Network for Salient Object Detection \u201d arXiv preprint arXiv:1906.11443 2019."},{"key":"e_1_2_11_43_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.eswa.2024.125278"},{"key":"e_1_2_11_44_1","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2003.819861"},{"key":"e_1_2_11_45_1","doi-asserted-by":"crossref","unstructured":"R.Zhang \u201cThe Unreasonable Effectiveness of Deep Features as a Perceptual Metric \u201d inProceedings of the IEEE Conference on Computer Vision and Pattern Recognition 2018 https:\/\/doi.org\/10.1109\/CVPR.2018.00068.","DOI":"10.1109\/CVPR.2018.00068"},{"key":"e_1_2_11_46_1","doi-asserted-by":"crossref","unstructured":"P.Zhang \u201cFantastic Animals and Where to Find Them: Segment Any Marine Animal With Dual SAM \u201d inProceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition 2024 https:\/\/doi.org\/10.1109\/CVPR52733.2024.00249.","DOI":"10.1109\/CVPR52733.2024.00249"},{"key":"e_1_2_11_47_1","doi-asserted-by":"crossref","unstructured":"S.Woo \u201cCBAM: Convolutional Block Attention Module \u201d inProceedings of the European Conference on Computer Vision (ECCV) 2018.","DOI":"10.1007\/978-3-030-01234-2_1"},{"key":"e_1_2_11_48_1","doi-asserted-by":"crossref","unstructured":"J.Hu L.Shen andG.Sun \u201cSqueeze\u2010and\u2010Excitation Networks \u201d inProceedings of the IEEE Conference on Computer Vision and Pattern Recognition 2018 https:\/\/doi.org\/10.1109\/CVPR.2018.00745.","DOI":"10.1109\/CVPR.2018.00745"},{"key":"e_1_2_11_49_1","doi-asserted-by":"crossref","unstructured":"W.Shi J.Caballero F.Husz\u00e1r et\u00a0al. \u201cReal\u2010Time Single Image and Video Super\u2010Resolution Using an Efficient Sub\u2010Pixel Convolutional Neural Network \u201d inProceedings of the IEEE Conference on Computer Vision and Pattern Recognition 2016 https:\/\/doi.org\/10.1109\/CVPR.2016.207.","DOI":"10.1109\/CVPR.2016.207"},{"key":"e_1_2_11_50_1","doi-asserted-by":"publisher","DOI":"10.1007\/s10479\u2010005\u20105724\u2010z"},{"key":"e_1_2_11_51_1","volume-title":"The Thirty\u2010Seventh Asilomar Conference on Signals, Systems & Computers, 2003","author":"Wang Z.","year":"2003"},{"key":"e_1_2_11_52_1","doi-asserted-by":"crossref","unstructured":"G.M\u00e1ttyus W.Luo andR.Urtasun \u201cDeepRoadMapper: Extracting Road Topology From Aerial Images \u201d inProceedings of the IEEE International Conference on Computer Vision 2017.","DOI":"10.1109\/ICCV.2017.372"},{"key":"e_1_2_11_53_1","unstructured":"M. J.Islam P.Luo andJ.Sattar \u201cSimultaneous Enhancement and Super\u2010Resolution of Underwater Imagery for Improved Visual Perception \u201d arXiv preprint arXiv:2002.01155 2020."},{"key":"e_1_2_11_54_1","doi-asserted-by":"publisher","DOI":"10.1109\/TCSVT.2023.3318672"},{"key":"e_1_2_11_55_1","doi-asserted-by":"publisher","DOI":"10.1109\/TCYB.2020.2969255"},{"key":"e_1_2_11_56_1","doi-asserted-by":"publisher","DOI":"10.1109\/TCSVT.2018.2870832"},{"key":"e_1_2_11_57_1","doi-asserted-by":"publisher","DOI":"10.1109\/TCSVT.2023.3264442"},{"key":"e_1_2_11_58_1","doi-asserted-by":"crossref","unstructured":"Z.Wu L.Su andQ.Huang \u201cCascaded Partial Decoder for Fast and Accurate Salient Object Detection \u201d inProceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition 2019 https:\/\/doi.org\/10.1109\/CVPR.2019.00403.","DOI":"10.1109\/CVPR.2019.00403"},{"key":"e_1_2_11_59_1","doi-asserted-by":"crossref","unstructured":"Z.Zhao \u201cComplementary Trilateral Decoder for Fast and Accurate Salient Object Detection \u201d inProceedings of the 29th ACM International Conference on Multimedia 2021 https:\/\/doi.org\/10.1145\/3474085.3475494.","DOI":"10.1145\/3474085.3475494"},{"key":"e_1_2_11_60_1","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2022.3164550"},{"key":"e_1_2_11_61_1","volume-title":"Computer Vision\u2013ECCV2020: 16th European Conference, Glasgow, UK, August 23\u201328, 2020, Proceedings, Part II 16","author":"Zhao X.","year":"2020"},{"key":"e_1_2_11_62_1","doi-asserted-by":"crossref","unstructured":"Z.Chen Q.Xu R.Cong andQ.Huang \u201cGlobal Context\u2010Aware Progressive Aggregation Network for Salient Object Detection \u201d inProceedings of the AAAI Conference on Artificial Intelligence 2020 https:\/\/doi.org\/10.1609\/aaai.v34i07.6633.","DOI":"10.1609\/aaai.v34i07.6633"},{"key":"e_1_2_11_63_1","doi-asserted-by":"crossref","unstructured":"J.Wei \u201cLabel Decoupling Framework for Salient Object Detection \u201d inProceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition 2020 https:\/\/doi.org\/10.1109\/CVPR42600.2020.01304.","DOI":"10.1109\/CVPR42600.2020.01304"},{"key":"e_1_2_11_64_1","doi-asserted-by":"crossref","unstructured":"Y.Pang \u201cMulti\u2010Scale Interactive Network for Salient Object Detection \u201d inProceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition 2020 https:\/\/doi.org\/10.1109\/CVPR42600.2020.00943.","DOI":"10.1109\/CVPR42600.2020.00943"},{"key":"e_1_2_11_65_1","doi-asserted-by":"crossref","unstructured":"M.Ma C.Xia andJ.Li \u201cPyramidal Feature Shrinking for Salient Object Detection \u201d inProceedings of the AAAI Conference on Artificial Intelligence 2021 https:\/\/doi.org\/10.1609\/aaai.v35i3.16331.","DOI":"10.1609\/aaai.v35i3.16331"},{"key":"e_1_2_11_66_1","doi-asserted-by":"crossref","unstructured":"J.\u2010J.Liu \u201cA Simple Pooling\u2010Based Design for Real\u2010Time Salient Object Detection \u201d inProceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition 2019 https:\/\/doi.org\/10.1109\/CVPR.2019.00404.","DOI":"10.1109\/CVPR.2019.00404"},{"key":"e_1_2_11_67_1","doi-asserted-by":"crossref","unstructured":"Z.Wu L.Su andQ.Huang \u201cStacked Cross Refinement Network for Edge\u2010Aware Salient Object Detection \u201d inProceedings of the IEEE\/CVF International Conference on Computer Vision 2019 https:\/\/doi.org\/10.1109\/ICCV.2019.00736.","DOI":"10.1109\/ICCV.2019.00736"},{"key":"e_1_2_11_68_1","doi-asserted-by":"crossref","unstructured":"Y.Wang \u201cPixels Regions and Objects: Multiple Enhancement for Salient Object Detection \u201d inProceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition 2023 https:\/\/doi.org\/10.1109\/CVPR52729.2023.00967.","DOI":"10.1109\/CVPR52729.2023.00967"},{"key":"e_1_2_11_69_1","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1109\/TGRS.2024.3425658","article-title":"Iterative Saliency Aggregation and Assignment Network for Efficient Salient Object Detection in Optical Remote Sensing Images","volume":"62","author":"Yao Z.","year":"2024","journal-title":"IEEE Transactions on Geoscience and Remote Sensing"},{"key":"e_1_2_11_70_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.patcog.2024.110328"},{"key":"e_1_2_11_71_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.eswa.2023.121778"},{"key":"e_1_2_11_72_1","doi-asserted-by":"publisher","DOI":"10.1109\/JOE.2023.3252760"}],"container-title":["IET Image Processing"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/ietresearch.onlinelibrary.wiley.com\/doi\/pdf\/10.1049\/ipr2.70331","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/ietresearch.onlinelibrary.wiley.com\/doi\/full-xml\/10.1049\/ipr2.70331","content-type":"application\/xml","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/ietresearch.onlinelibrary.wiley.com\/doi\/pdf\/10.1049\/ipr2.70331","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,3,13]],"date-time":"2026-03-13T09:08:27Z","timestamp":1773392907000},"score":1,"resource":{"primary":{"URL":"https:\/\/ietresearch.onlinelibrary.wiley.com\/doi\/10.1049\/ipr2.70331"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,1]]},"references-count":71,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2026,1]]}},"alternative-id":["10.1049\/ipr2.70331"],"URL":"https:\/\/doi.org\/10.1049\/ipr2.70331","archive":["Portico"],"relation":{},"ISSN":["1751-9659","1751-9667"],"issn-type":[{"value":"1751-9659","type":"print"},{"value":"1751-9667","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,1]]},"assertion":[{"value":"2025-09-08","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2026-03-04","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2026-03-13","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}],"article-number":"e70331"}}