{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,30]],"date-time":"2026-06-30T10:14:36Z","timestamp":1782814476252,"version":"3.54.5"},"reference-count":42,"publisher":"MDPI AG","issue":"5","license":[{"start":{"date-parts":[[2024,4,24]],"date-time":"2024-04-24T00:00:00Z","timestamp":1713916800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Symmetry"],"abstract":"<jats:p>Exploring the ocean\u2019s resources requires finding underwater objects, which is a challenging task due to blurry images and small, densely packed targets. To improve the accuracy of underwater target detection, we propose an enhanced version of the YOLOv7 network called YOLOv7-SN. Our goal is to optimize the effectiveness and accuracy of underwater target detection by introducing a series of innovations. We incorporate the channel attention module SE into the network\u2019s key part to improve the extraction of relevant features for underwater targets. We also introduce the RFE module with dilated convolution behind the backbone network to capture multi-scale information. Additionally, we use the Wasserstein distance as a new metric to replace the traditional loss function and address the challenge of small target detection. Finally, we employ probe heads carrying implicit knowledge to further enhance the model\u2019s accuracy. These methods aim to optimize the efficacy of underwater target detection and improve its ability to deal with the complexity and challenges of underwater environments. We conducted experiments on the URPC2020, and RUIE datasets. The results show that the mean accuracy (mAP) is improved by 5.9% and 3.9%, respectively, compared to the baseline model.<\/jats:p>","DOI":"10.3390\/sym16050514","type":"journal-article","created":{"date-parts":[[2024,4,24]],"date-time":"2024-04-24T07:38:51Z","timestamp":1713944331000},"page":"514","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":10,"title":["YOLOv7-SN: Underwater Target Detection Algorithm Based on Improved YOLOv7"],"prefix":"10.3390","volume":"16","author":[{"given":"Ming","family":"Zhao","sequence":"first","affiliation":[{"name":"School of Mathematical Sciences, Harbin Normal University, Harbin 150500, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Huibo","family":"Zhou","sequence":"additional","affiliation":[{"name":"School of Mathematical Sciences, Harbin Normal University, Harbin 150500, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Xue","family":"Li","sequence":"additional","affiliation":[{"name":"School of Mathematical Sciences, Harbin Normal University, Harbin 150500, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1968","published-online":{"date-parts":[[2024,4,24]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","first-page":"98841","DOI":"10.1109\/ACCESS.2019.2929932","article-title":"An overview of next-generation underwater target detection and tracking: An integrated underwater architecture","volume":"7","author":"Ghafoor","year":"2019","journal-title":"IEEE Access"},{"key":"ref_2","doi-asserted-by":"crossref","first-page":"108159","DOI":"10.1016\/j.compeleceng.2022.108159","article-title":"Underwater object detection using collaborative weakly supervision","volume":"102","author":"Cai","year":"2022","journal-title":"Comput. Electr. Eng."},{"key":"ref_3","doi-asserted-by":"crossref","first-page":"1153416","DOI":"10.3389\/fmars.2023.1153416","article-title":"Underwater target detection algorithm based on improved yolov4 with semidsconv and fiou loss function","volume":"10","author":"Zhang","year":"2023","journal-title":"Front. Mar. Sci."},{"key":"ref_4","unstructured":"Dalal, N., and Triggs, B. (2005, January 20\u201326). Histograms of oriented gradients for human detection. Proceedings of the 2005 IEEE Computer Society Conference on Computer Vision and Pattern Recognition (CVPR\u201905), San Diego, CA, USA."},{"key":"ref_5","doi-asserted-by":"crossref","unstructured":"Pinto, F., Torr, P.H., and Dokania, P.K. (2022, January 23\u201327). An impartial take to the cnn vs transformer robustness contest. Proceedings of the Computer Vision\u2013ECCV 2022: 17th European Conference, Tel Aviv, Israel.","DOI":"10.1007\/978-3-031-19778-9_27"},{"key":"ref_6","doi-asserted-by":"crossref","unstructured":"Yao, Y., Qiu, Z., and Zhong, M. (2019, January 20\u201322). Application of improved MobileNet-SSD on underwater sea cucumber detection robot. Proceedings of the 2019 IEEE 4th Advanced Information Technology, Electronic and Automation Control Conference (IAEAC), Chengdu, China.","DOI":"10.1109\/IAEAC47372.2019.8997970"},{"key":"ref_7","doi-asserted-by":"crossref","unstructured":"Redmon, J., Divvala, S., Girshick, R., and Farhadi, A. (2016, January 27\u201330). You only look once: Unified, real-time object detection. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.91"},{"key":"ref_8","doi-asserted-by":"crossref","unstructured":"Redmon, J., and Farhadi, A. (2017, January 21\u201326). YOLO9000: Better, faster, stronger. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.690"},{"key":"ref_9","unstructured":"Redmon, J., and Farhadi, A. (2018). Yolov3: An incremental improvement. arXiv."},{"key":"ref_10","unstructured":"Bochkovskiy, A., Wang, C.Y., and Liao, H.Y.M. (2020). Yolov4: Optimal speed and accuracy of object detection. arXiv."},{"key":"ref_11","doi-asserted-by":"crossref","unstructured":"Zhu, X., Lyu, S., Wang, X., and Zhao, Q. (2021, January 11\u201317). TPH-YOLOv5: Improved YOLOv5 based on transformer prediction head for object detection on drone-captured scenarios. Proceedings of the IEEE\/CVF International Conference on Computer Vision, Montreal, BC, Canada.","DOI":"10.1109\/ICCVW54120.2021.00312"},{"key":"ref_12","unstructured":"Li, C., Li, L., Jiang, H., Weng, K., Geng, Y., Li, L., Ke, Z., Li, Q., Cheng, M., and Nie, W. (2022). YOLOv6: A Single-Stage Object Detection Framework for Industrial Applications. arXiv."},{"key":"ref_13","doi-asserted-by":"crossref","unstructured":"Wang, C.-Y., Bochkovskiy, A., and Liao, H.-Y. (2022). YOLOv7: Trainable bag-of-freebies sets new state-of-the-art for real-time object detectors. arXiv.","DOI":"10.1109\/CVPR52729.2023.00721"},{"key":"ref_14","unstructured":"Li, X., Shang, M., Qin, H., and Chen, L. (2015, January 19\u201322). Fast accurate fish detection and recognition of underwater images with fast r-cnn. Proceedings of the OCEANS 2015-MTS\/IEEE Washington, DC, USA."},{"key":"ref_15","doi-asserted-by":"crossref","unstructured":"Girshick, R. (2015, January 7\u201313). Fast r-cnn. Proceedings of the IEEE International Conference on Computer Vision, Santiago, Chile.","DOI":"10.1109\/ICCV.2015.169"},{"key":"ref_16","doi-asserted-by":"crossref","unstructured":"Chen, L., Liu, Z., Tong, L., Jiang, Z., Wang, S., Dong, J., and Zhou, H. (2020, January 19\u201324). Underwater object detection using Invert Multi-Class Adaboost with deep learning. Proceedings of the 2020 International Joint Conference on Neural Networks (IJCNN), Glasgow, UK.","DOI":"10.1109\/IJCNN48605.2020.9207506"},{"key":"ref_17","doi-asserted-by":"crossref","first-page":"108415","DOI":"10.1016\/j.oceaneng.2020.108415","article-title":"Underwater targets classification using local wavelet acoustic pattern and Multi-Layer Perceptron neural network optimized by modified Whale Optimization Algorithm","volume":"219","author":"Qiao","year":"2021","journal-title":"Ocean Eng."},{"key":"ref_18","doi-asserted-by":"crossref","unstructured":"Lei, F., Tang, F., and Li, S. (2022). Underwater target detection algorithm based on improved YOLOv5. J. Mar. Sci. Eng., 10.","DOI":"10.3390\/jmse10030310"},{"key":"ref_19","unstructured":"Anasosalu Vasu, P.K., Gabriel, J., Zhu, J., Tuzel, O., and Ranjan, A. (2022). An Improved One Millisecond Mobile Backbone. arXiv."},{"key":"ref_20","unstructured":"PGao, P., Lu, J., Li, H., Mottaghi, R., and Kembhavi, A. (2021). Container: Context aggregation network. arXiv."},{"key":"ref_21","doi-asserted-by":"crossref","unstructured":"Doll\u00e1r, P., Singh, M., and Girshick, R. (2021). Fast and accurate model scaling. arXiv.","DOI":"10.1109\/CVPR46437.2021.00098"},{"key":"ref_22","unstructured":"Xu, K., Ba, J., Kiros, R., Cho, K., Courville, A., Salakhudinov, R., Zemel, R., and Bengio, Y. (2015, January 7\u20139). Show, attend and tell: Neural image caption generation with visual attention. Proceedings of the International Conference on Machine Learning, PMLR, Lille, France."},{"key":"ref_23","unstructured":"Tsotsos, J.K. (2021). A Computational Perspective on Visual Attention, MIT Press."},{"key":"ref_24","unstructured":"Mnih, V., Heess, N., and Graves, A. (2014, January 8\u201313). Recurrent models of visual attention. Proceedings of the 27th International Conference on Neural Information Processing Systems, Montreal, QC, Canada."},{"key":"ref_25","doi-asserted-by":"crossref","unstructured":"Bello, I., Zoph, B., Vaswani, A., Shlens, J., and Le, Q.V. (2019, January 27\u201328). Attention augmented convolutional networks. Proceedings of the IEEE\/CVF International Conference on Computer Vision, Seoul, Republic of Korea.","DOI":"10.1109\/ICCV.2019.00338"},{"key":"ref_26","doi-asserted-by":"crossref","unstructured":"Fu, J., Liu, J., Tian, H., Li, Y., Bao, Y., Fang, Z., and Lu, H. (2019, January 15\u201320). Dual attention network for scene segmentation. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.00326"},{"key":"ref_27","doi-asserted-by":"crossref","unstructured":"Shen, Z., and Nguyen, C. (December, January 29). Temporal 3D RetinaNet for fish detection. Proceedings of the 2020 Digital Image Computing: Techniques and Applications (DICTA), Melbourne, Australia.","DOI":"10.1109\/DICTA51227.2020.9363372"},{"key":"ref_28","doi-asserted-by":"crossref","unstructured":"Li, X., Wang, W., Hu, X., and Yang, J. (2019, January 15\u201320). Selective kernel networks. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.00060"},{"key":"ref_29","doi-asserted-by":"crossref","unstructured":"Woo, S., Park, J., Lee, J.Y., and Kweon, I.S. (2018, January 8\u201314). Cbam: Convolutional block attention module. Proceedings of the European Conference on Computer Vision (ECCV), Munich, Germany.","DOI":"10.1007\/978-3-030-01234-2_1"},{"key":"ref_30","doi-asserted-by":"crossref","unstructured":"Misra, D., Nalamada, T., Arasanipalai, A.U., and Hou, Q. (2021, January 3\u20138). Rotate to attend: Convolutional triplet attention module. Proceedings of the IEEE\/CVF Winter Conference on Applications of Computer Vision, Waikoloa, HI, USA.","DOI":"10.1109\/WACV48630.2021.00318"},{"key":"ref_31","doi-asserted-by":"crossref","unstructured":"Wang, L., Peng, J., and Sun, W. (2019). Spatial\u2013spectral squeeze-and-excitation residual network for hyperspectral image classification. Remote Sens., 11.","DOI":"10.3390\/rs11070884"},{"key":"ref_32","doi-asserted-by":"crossref","unstructured":"Rezatofighi, H., Tsoi, N., Gwak, J., Sadeghian, A., Reid, I., and Savarese, S. (2019, January 15\u201320). Generalized intersection over union: A metric and a loss for bounding box regression. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.00075"},{"key":"ref_33","doi-asserted-by":"crossref","first-page":"146","DOI":"10.1016\/j.neucom.2022.07.042","article-title":"Focal and efficient iou loss for accurate bounding box regression","volume":"506","author":"Zhang","year":"2022","journal-title":"Neurocomputing"},{"key":"ref_34","unstructured":"Wang, C.Y., Yeh, I.H., and Liao, H.Y.M. (2021). You only learn one representation: Unified network for multiple tasks. arXiv."},{"key":"ref_35","unstructured":"Yu, Z., Huang, H., Chen, W., Su, Y., Liu, Y., and Wang, X. (2021). Yolo-facev2: A scale and occlusion aware face detector. arXiv."},{"key":"ref_36","unstructured":"Yu, F., and Koltun, V. (2015). Multi-scale context aggregation by dilated convolutions. arXiv."},{"key":"ref_37","doi-asserted-by":"crossref","unstructured":"Zheng, Z., Wang, P., Liu, W., Li, J., Ye, R., and Ren, D. (2020, January 7\u201312). Distance-IoU loss: Faster and better learning for bounding box regression. Proceedings of the AAAI conference on artificial intelligence, New York, NY, USA.","DOI":"10.1609\/aaai.v34i07.6999"},{"key":"ref_38","unstructured":"Wang, J., Xu, C., Yang, W., and Yu, L. (2021). A normalized Gaussian Wasserstein distance for tiny object detection. arXiv."},{"key":"ref_39","doi-asserted-by":"crossref","unstructured":"Liu, W., Anguelov, D., Erhan, D., Szegedy, C., Reed, S., Fu, C.Y., and Berg, A.C. (2016, January 11\u201314). Ssd: Single shot multibox detector. Proceedings of the Computer Vision\u2013ECCV 2016: 14th European Conference, Amsterdam, The Netherlands. Proceedings, Part I 14.","DOI":"10.1007\/978-3-319-46448-0_2"},{"key":"ref_40","doi-asserted-by":"crossref","unstructured":"Lin, T.Y., Goyal, P., Girshick, R., He, K., and Doll\u00e1r, P. (2017, January 22\u201329). Focal loss for dense object detection. Proceedings of the IEEE International Conference on Computer Vision, Venice, Italy.","DOI":"10.1109\/ICCV.2017.324"},{"key":"ref_41","unstructured":"Ge, Z., Liu, S., Wang, F., Li, Z., and Sun, J. (2021). YOLOX: Exceeding YOLO series in 2021. arXiv."},{"key":"ref_42","doi-asserted-by":"crossref","first-page":"4861","DOI":"10.1109\/TCSVT.2019.2963772","article-title":"Real-World Underwater Enhancement: Challenges, Benchmarks, and Solutions Under Natural Light","volume":"30","author":"Liu","year":"2020","journal-title":"IEEE Trans. Circuits Syst. Video Technol."}],"container-title":["Symmetry"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/2073-8994\/16\/5\/514\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,10]],"date-time":"2025-10-10T14:33:16Z","timestamp":1760106796000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/2073-8994\/16\/5\/514"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,4,24]]},"references-count":42,"journal-issue":{"issue":"5","published-online":{"date-parts":[[2024,5]]}},"alternative-id":["sym16050514"],"URL":"https:\/\/doi.org\/10.3390\/sym16050514","relation":{},"ISSN":["2073-8994"],"issn-type":[{"value":"2073-8994","type":"electronic"}],"subject":[],"published":{"date-parts":[[2024,4,24]]}}}