{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,8]],"date-time":"2026-07-08T16:05:39Z","timestamp":1783526739416,"version":"3.55.0"},"reference-count":79,"publisher":"MDPI AG","issue":"9","license":[{"start":{"date-parts":[[2024,5,1]],"date-time":"2024-05-01T00:00:00Z","timestamp":1714521600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"name":"College Student Innovation and Entrepreneurship project of Hainan University","award":["Hdcxcyxm201704"],"award-info":[{"award-number":["Hdcxcyxm201704"]}]},{"name":"College Student Innovation and Entrepreneurship project of Hainan University","award":["623RC449"],"award-info":[{"award-number":["623RC449"]}]},{"name":"Hainan Provincial Natural Science Foundation of China","award":["Hdcxcyxm201704"],"award-info":[{"award-number":["Hdcxcyxm201704"]}]},{"name":"Hainan Provincial Natural Science Foundation of China","award":["623RC449"],"award-info":[{"award-number":["623RC449"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Sensors"],"abstract":"<jats:p>Underwater visual detection technology is crucial for marine exploration and monitoring. Given the growing demand for accurate underwater target recognition, this study introduces an innovative architecture, YOLOv8-MU, which significantly enhances the detection accuracy. This model incorporates the large kernel block (LarK block) from UniRepLKNet to optimize the backbone network, achieving a broader receptive field without increasing the model\u2019s depth. Additionally, the integration of C2fSTR, which combines the Swin transformer with the C2f module, and the SPPFCSPC_EMA module, which blends Cross-Stage Partial Fast Spatial Pyramid Pooling (SPPFCSPC) with attention mechanisms, notably improves the detection accuracy and robustness for various biological targets. A fusion block from DAMO-YOLO further enhances the multi-scale feature extraction capabilities in the model\u2019s neck. Moreover, the adoption of the MPDIoU loss function, designed around the vertex distance, effectively addresses the challenges of localization accuracy and boundary clarity in underwater organism detection. The experimental results on the URPC2019 dataset indicate that YOLOv8-MU achieves an mAP@0.5 of 78.4%, showing an improvement of 4.0% over the original YOLOv8 model. Additionally, on the URPC2020 dataset, it achieves 80.9%, and, on the Aquarium dataset, it reaches 75.5%, surpassing other models, including YOLOv5 and YOLOv8n, thus confirming the wide applicability and generalization capabilities of our proposed improved model architecture. Furthermore, an evaluation on the improved URPC2019 dataset demonstrates leading performance (SOTA), with an mAP@0.5 of 88.1%, further verifying its superiority on this dataset. These results highlight the model\u2019s broad applicability and generalization capabilities across various underwater datasets.<\/jats:p>","DOI":"10.3390\/s24092905","type":"journal-article","created":{"date-parts":[[2024,5,2]],"date-time":"2024-05-02T03:57:56Z","timestamp":1714622276000},"page":"2905","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":21,"title":["YOLOv8-MU: An Improved YOLOv8 Underwater Detector Based on a Large Kernel Block and a Multi-Branch Reparameterization Module"],"prefix":"10.3390","volume":"24","author":[{"ORCID":"https:\/\/orcid.org\/0009-0007-8035-734X","authenticated-orcid":false,"given":"Xing","family":"Jiang","sequence":"first","affiliation":[{"name":"School of Tropical Agriculture and Forestry (School of Agricultural and Rural, School of Rural Revitalization), Hainan University, Danzhou 571737, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0001-8975-3724","authenticated-orcid":false,"given":"Xiting","family":"Zhuang","sequence":"additional","affiliation":[{"name":"School of Tropical Agriculture and Forestry (School of Agricultural and Rural, School of Rural Revitalization), Hainan University, Danzhou 571737, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0007-4807-8017","authenticated-orcid":false,"given":"Jisheng","family":"Chen","sequence":"additional","affiliation":[{"name":"School of Tropical Agriculture and Forestry (School of Agricultural and Rural, School of Rural Revitalization), Hainan University, Danzhou 571737, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-0871-6874","authenticated-orcid":false,"given":"Jian","family":"Zhang","sequence":"additional","affiliation":[{"name":"School of Tropical Agriculture and Forestry (School of Agricultural and Rural, School of Rural Revitalization), Hainan University, Danzhou 571737, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0008-2607-2623","authenticated-orcid":false,"given":"Yiwen","family":"Zhang","sequence":"additional","affiliation":[{"name":"School of Tropical Agriculture and Forestry (School of Agricultural and Rural, School of Rural Revitalization), Hainan University, Danzhou 571737, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1968","published-online":{"date-parts":[[2024,5,1]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","first-page":"e08975","DOI":"10.1016\/j.heliyon.2022.e08975","article-title":"Projecting Future Changes in Distributions of Small-Scale Pelagic Fisheries of the Southern Colombian Pacific Ocean","volume":"8","author":"Selvaraj","year":"2022","journal-title":"Heliyon"},{"key":"ref_2","doi-asserted-by":"crossref","unstructured":"Shen, R., Zhao, Y., Cheng, H., Hu, S., Chen, S., and Ge, S. (2023). Surface-Related Multiples Elimination for Waterborne GPR Data. Remote Sens., 15.","DOI":"10.3390\/rs15133250"},{"key":"ref_3","doi-asserted-by":"crossref","first-page":"115051","DOI":"10.1016\/j.eswa.2021.115051","article-title":"Real-time nondestructive fish behavior detecting in mixed polyculture system using deep-learning and low-cost devices","volume":"178","author":"Hu","year":"2021","journal-title":"Expert Syst. Appl."},{"key":"ref_4","doi-asserted-by":"crossref","unstructured":"Wang, S., Liu, X., Yu, S., Zhu, X., Chen, B., and Sun, X. (2024). Design and Implementation of SSS-Based AUV Autonomous Online Object Detection System. Electronics, 13.","DOI":"10.3390\/electronics13061064"},{"key":"ref_5","doi-asserted-by":"crossref","unstructured":"Lee, M.-F.R., and Chen, Y.-C. (2023). Artificial Intelligence Based Object Detection and Tracking for a Small Underwater Robot. Processes, 11.","DOI":"10.3390\/pr11020312"},{"key":"ref_6","doi-asserted-by":"crossref","first-page":"108863","DOI":"10.1016\/j.automatica.2020.108863","article-title":"Distributed Implementation of Nonlinear Model Predictive Control for AUV Trajectory Tracking","volume":"115","author":"Shen","year":"2020","journal-title":"Automatica"},{"key":"ref_7","doi-asserted-by":"crossref","first-page":"1732","DOI":"10.1109\/TII.2020.2994586","article-title":"Intelligent Collaborative Navigation and Control for AUV Tracking","volume":"17","author":"Guo","year":"2020","journal-title":"IEEE Trans. Ind. Inform."},{"key":"ref_8","doi-asserted-by":"crossref","first-page":"478","DOI":"10.1007\/s12555-019-0673-5","article-title":"Current Estimation and Path Following for an Autonomous Underwater Vehicle (AUV) by Using a High-Gain Observer Based on an AUV Dynamic Model","volume":"19","author":"Kim","year":"2021","journal-title":"Int. J. Control Autom. Syst."},{"key":"ref_9","doi-asserted-by":"crossref","unstructured":"Wang, T., Ding, F., and Sun, Z. (2023). Visual-Aided Shared Control of Semi-Autonomous Underwater Vehicle for Efficient Underwater Grasping. J. Mar. Sci. Eng., 11.","DOI":"10.3390\/jmse11091837"},{"key":"ref_10","doi-asserted-by":"crossref","unstructured":"Zhang, Y., Jiang, Y., Qi, H., Zhao, M., Wang, Y., Wang, K., and Wei, F. (2023). An Underwater Human\u2013Robot Interaction Using a Visual\u2013Textual Model for Autonomous Underwater Vehicles. Sensors, 23.","DOI":"10.3390\/s23010197"},{"key":"ref_11","doi-asserted-by":"crossref","first-page":"86","DOI":"10.1109\/MNET.2019.1800425","article-title":"Localization and Data Collection in AUV-Aided Underwater Sensor Networks: Challenges and Opportunities","volume":"33","author":"Su","year":"2019","journal-title":"IEEE Netw."},{"key":"ref_12","doi-asserted-by":"crossref","first-page":"012079","DOI":"10.1088\/1757-899X\/1096\/1\/012079","article-title":"Implementation of Real-Time Edge Detection Using Canny and Sobel Algorithms","volume":"1096","author":"Lynn","year":"2021","journal-title":"IOP Conf. Ser. Mater. Sci. Eng."},{"key":"ref_13","unstructured":"Kurniati, F.T., Manongga, D.H., Sediyono, E., Prasetyo, S.Y.J., and Huizen, R.R. (2024). GLCM-Based Feature Combination for Extraction Model Optimization in Object Detection Using Machine Learning. arXiv."},{"key":"ref_14","doi-asserted-by":"crossref","first-page":"302","DOI":"10.1109\/LGRS.2019.2919755","article-title":"Fourier-Based Rotation-Invariant Feature Boosting: An Efficient Framework for Geospatial Object Detection","volume":"17","author":"Wu","year":"2020","journal-title":"IEEE Geosci. Remote Sensing Lett."},{"key":"ref_15","doi-asserted-by":"crossref","first-page":"104759","DOI":"10.1016\/j.engappai.2022.104759","article-title":"Enhancing Underwater Image via Adaptive Color and Contrast Enhancement, and Denoising","volume":"111","author":"Li","year":"2022","journal-title":"Eng. Appl. Artif. Intell."},{"key":"ref_16","doi-asserted-by":"crossref","first-page":"6216","DOI":"10.1364\/OE.449930","article-title":"Underwater Image Enhancement Using Adaptive Color Restoration and Dehazing","volume":"30","author":"Li","year":"2022","journal-title":"Opt. Express"},{"key":"ref_17","doi-asserted-by":"crossref","unstructured":"Jiang, W., Yang, X., Tong, F., Yang, Y., and Zhou, T. (2022). A Low-Complexity Underwater Acoustic Coherent Communication System for Small AUV. Remote Sens., 14.","DOI":"10.3390\/rs14143405"},{"key":"ref_18","doi-asserted-by":"crossref","first-page":"108926","DOI":"10.1016\/j.patcog.2022.108926","article-title":"SWIPENET: Object Detection in Noisy Underwater Scenes","volume":"132","author":"Chen","year":"2022","journal-title":"Pattern Recognit."},{"key":"ref_19","doi-asserted-by":"crossref","unstructured":"Dong, X., Qin, Y., Gao, Y., Fu, R., Liu, S., and Ye, Y. (2022). Attention-Based Multi-Level Feature Fusion for Object Detection in Remote Sensing Images. Remote Sens., 14.","DOI":"10.3390\/rs14153735"},{"key":"ref_20","doi-asserted-by":"crossref","first-page":"89687","DOI":"10.1109\/ACCESS.2022.3201086","article-title":"Thangka Image Segmentation Method Based on Enhanced Receptive Field","volume":"10","author":"Wang","year":"2022","journal-title":"IEEE Access"},{"key":"ref_21","doi-asserted-by":"crossref","first-page":"26357","DOI":"10.1109\/TGRS.2020.3009143","article-title":"Adaptive Effective Receptive Field Convolution for Semantic Segmentation of VHR Remote Sensing Images","volume":"59","author":"Chen","year":"2021","journal-title":"IEEE Trans. Geosci. Remote Sens."},{"key":"ref_22","doi-asserted-by":"crossref","first-page":"26357","DOI":"10.1109\/JSEN.2023.3318371","article-title":"RFRFlow: Recurrent Feature Refinement Network for Optical Flow Estimation","volume":"23","author":"Zhu","year":"2023","journal-title":"IEEE Sens. J."},{"key":"ref_23","doi-asserted-by":"crossref","first-page":"98854","DOI":"10.1109\/ACCESS.2019.2930293","article-title":"SKFlow: Optical Flow Estimation Using Selective Kernel Networks","volume":"7","author":"Zhai","year":"2019","journal-title":"IEEE Access"},{"key":"ref_24","doi-asserted-by":"crossref","first-page":"1442","DOI":"10.1109\/TIP.2023.3244647","article-title":"Domain Adaptation for Underwater Image Enhancement","volume":"32","author":"Wang","year":"2023","journal-title":"IEEE Trans. Image Process."},{"key":"ref_25","doi-asserted-by":"crossref","unstructured":"Zhao, S., Zheng, J., Sun, S., and Zhang, L. (2022). An improved YOLO algorithm for fast and accurate underwater object detection. Symmetry, 14.","DOI":"10.2139\/ssrn.4079287"},{"key":"ref_26","first-page":"12325","article-title":"Edge-guided Representation Learning for Underwater Object Detection","volume":"cit2","author":"Dai","year":"2024","journal-title":"CAAI Trans. Intel. Tech."},{"key":"ref_27","unstructured":"Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A.N., Kaiser, \u0141., and Polosukhin, I. (2017, January 4\u20139). Attention is all you need. Proceedings of the Thirty-First Annual Conference on Neural Information Processing Systems (NIPS), Long Beach, CA, USA."},{"key":"ref_28","doi-asserted-by":"crossref","unstructured":"Liu, Z., Lin, Y., Cao, Y., Hu, H., Wei, Y., Zhang, Z., Lin, S., and Guo, B. (2021, January 11\u201317). Swin Transformer: Hierarchical Vision Transformer Using Shifted Windows. Proceedings of the 2021 IEEE\/CVF International Conference on Computer Vision (ICCV), Montreal, QC, Canada.","DOI":"10.1109\/ICCV48922.2021.00986"},{"key":"ref_29","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1109\/LGRS.2020.3029584","article-title":"Underwater Acoustic Target Classification Based on Dense Convolutional Neural Network","volume":"19","author":"Doan","year":"2022","journal-title":"IEEE Geosci. Remote Sens. Lett."},{"key":"ref_30","unstructured":"Ding, X., Zhang, Y., Ge, Y., Zhao, S., Song, L., Yue, X., and Shan, Y. (2023). Unireplknet: A universal perception large-kernel convnet for audio, video, point cloud, time-series and image recognition. arXiv."},{"key":"ref_31","doi-asserted-by":"crossref","first-page":"1056300","DOI":"10.3389\/fmars.2022.1056300","article-title":"Underwater Object Detection Algorithm Based on Attention Mechanism and Cross-Stage Partial Fast Spatial Pyramidal Pooling","volume":"9","author":"Yan","year":"2022","journal-title":"Front. Mar. Sci."},{"key":"ref_32","unstructured":"Xu, X., Jiang, Y., Chen, W., Huang, Y., Zhang, Y., and Sun, X. (2022). DAMO-YOLO: A Report on Real-Time Object Detection Design. arXiv."},{"key":"ref_33","unstructured":"Siliang, M., and Yong, X. (2023). MPDIoU: A Loss for Efficient and Accurate Bounding Box Regression. arXiv."},{"key":"ref_34","doi-asserted-by":"crossref","unstructured":"Girshick, R., Donahue, J., Darrell, T., and Malik, J. (2014, January 23\u201328). Rich Feature Hierarchies for Accurate Object Detection and Semantic Segmentation. Proceedings of the 2014 IEEE Conference on Computer Vision and Pattern Recognition, Columbus, OH, USA.","DOI":"10.1109\/CVPR.2014.81"},{"key":"ref_35","doi-asserted-by":"crossref","first-page":"1137","DOI":"10.1109\/TPAMI.2016.2577031","article-title":"Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks","volume":"39","author":"Ren","year":"2017","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_36","doi-asserted-by":"crossref","first-page":"104190","DOI":"10.1016\/j.engappai.2021.104190","article-title":"Underwater Target Detection Based on Faster R-CNN and Adversarial Occlusion Network","volume":"100","author":"Zeng","year":"2021","journal-title":"Eng. Appl. Artif. Intell."},{"key":"ref_37","doi-asserted-by":"crossref","first-page":"150","DOI":"10.1016\/j.neucom.2023.01.088","article-title":"Boosting R-CNN: Reweighting R-CNN Samples by RPN\u2019s Error for Underwater Object Detection","volume":"530","author":"Song","year":"2023","journal-title":"Neurocomputing"},{"key":"ref_38","doi-asserted-by":"crossref","unstructured":"Hsia, C.-H., Chang, T.-H.W., Chiang, C.-Y., and Chan, H.-T. (2022). Mask R-CNN with New Data Augmentation Features for Smart Detection of Retail Products. Appl. Sci., 12.","DOI":"10.3390\/app12062902"},{"key":"ref_39","doi-asserted-by":"crossref","unstructured":"Redmon, J., Divvala, S., Girshick, R., and Farhadi, A. (2016, January 27\u201330). You Only Look Once: Unified, Real-Time Object Detection. Proceedings of the 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.91"},{"key":"ref_40","doi-asserted-by":"crossref","unstructured":"Redmon, J., and Farhadi, A. (2017, January 21\u201326). YOLO9000: Better, Faster, Stronger. Proceedings of the 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.690"},{"key":"ref_41","unstructured":"Redmon, J., and Farhadi, A. (2018). Yolov3: An incremental improvement. arXiv."},{"key":"ref_42","unstructured":"Bochkovskiy, A., Wang, C.Y., and Liao, H.Y.M. (2020). Yolov4: Optimal speed and accuracy of object detection. arXiv."},{"key":"ref_43","unstructured":"Jocher, G. (2022, December 22). YOLOv5 by Ultralytics. Available online: https:\/\/github.com\/ultralytics\/yolov5."},{"key":"ref_44","unstructured":"Li, C., Li, L., Jiang, H., Weng, K., Geng, Y., Li, L., Ke, Z., Li, Q., Cheng, M., and Nie, W. (2022). YOLOv6: A Single-Stage Object Detection Framework for Industrial Applications. arXiv."},{"key":"ref_45","doi-asserted-by":"crossref","unstructured":"Wang, C.-Y., Bochkovskiy, A., and Liao, H.-Y.M. (2023, January 17\u201324). YOLOv7: Trainable Bag-of-Freebies Sets New State-of-the-Art for Real-Time Object Detectors. Proceedings of the 2023 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Vancouver, BC, Canada.","DOI":"10.1109\/CVPR52729.2023.00721"},{"key":"ref_46","unstructured":"Jocher, G. (2023, February 15). YOLOv8 by Ultralytics. Available online: https:\/\/github.com\/ultralytics\/ultralytics."},{"key":"ref_47","doi-asserted-by":"crossref","unstructured":"Li, E., Wang, Q., Zhang, J., Zhang, W., Mo, H., and Wu, Y. (2023). Fish Detection under Occlusion Using Modified You Only Look Once v8 Integrating Real-Time Detection Transformer Features. Appl. Sci., 13.","DOI":"10.3390\/app132312645"},{"key":"ref_48","unstructured":"Liu, W., Anguelov, D., Erhan, D., Szegedy, C., Reed, S., Fu, C.-Y., and Berg, A.C. (2016). Lecture Notes in Computer Science, Springer International Publishing."},{"key":"ref_49","doi-asserted-by":"crossref","unstructured":"Lin, T.-Y., Goyal, P., Girshick, R., He, K., and Dollar, P. (2017, January 22\u201329). Focal Loss for Dense Object Detection. Proceedings of the 2017 IEEE International Conference on Computer Vision (ICCV), Venice, Italy.","DOI":"10.1109\/ICCV.2017.324"},{"key":"ref_50","doi-asserted-by":"crossref","unstructured":"Yang, S., Quan, Z., Nie, M., and Yang, W. (2021, January 10\u201317). TransPose: Keypoint Localization via Transformer. Proceedings of the 2021 IEEE\/CVF International Conference on Computer Vision (ICCV), Montreal, QC, Canada.","DOI":"10.1109\/ICCV48922.2021.01159"},{"key":"ref_51","doi-asserted-by":"crossref","unstructured":"Mao, W., Ge, Y., Shen, C., Tian, Z., Wang, X., Wang, Z., and den Hengel, A.v. (2022, January 23\u201327). Poseur: Direct Human Pose Regression with Transformers. Proceedings of the European Conference on Computer Vision, Tel Aviv, Israel.","DOI":"10.1007\/978-3-031-20068-7_5"},{"key":"ref_52","first-page":"38571","article-title":"ViTPose: Simple Vision Transformer Baselines for Human Pose Estimation","volume":"35","author":"Xu","year":"2022","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_53","doi-asserted-by":"crossref","unstructured":"Wang, Y., Guo, W., Zhao, S., Xue, B., Zhang, W., and Xing, Z. (2022). A Big Coal Block Alarm Detection Method for Scraper Conveyor Based on YOLO-BS. Sensors, 22.","DOI":"10.3390\/s22239052"},{"key":"ref_54","doi-asserted-by":"crossref","unstructured":"Srinivas, A., Lin, T.-Y., Parmar, N., Shlens, J., Abbeel, P., and Vaswani, A. (2021, January 20\u201325). Bottleneck Transformers for Visual Recognition. Proceedings of the 2021 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Nashville, TN, USA.","DOI":"10.1109\/CVPR46437.2021.01625"},{"key":"ref_55","doi-asserted-by":"crossref","first-page":"346","DOI":"10.1007\/978-3-319-10578-9_23","article-title":"Spatial Pyramid Pooling in Deep Convolutional Networks for Visual Recognition","volume":"Volume 8691","author":"He","year":"2014","journal-title":"Computer Vision\u2013ECCV 2014"},{"key":"ref_56","doi-asserted-by":"crossref","first-page":"834","DOI":"10.1109\/TPAMI.2017.2699184","article-title":"Deeplab: Semantic image segmentation with deep convolutional nets, atrous convolution, and fully connected crfs","volume":"40","author":"Chen","year":"2017","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_57","doi-asserted-by":"crossref","unstructured":"Wu, T., and Dong, Y. (2023). YOLO-SE: Improved YOLOv8 for Remote Sensing Object Detection and Recognition. Appl. Sci., 13.","DOI":"10.3390\/app132412977"},{"key":"ref_58","doi-asserted-by":"crossref","unstructured":"Yu, J., Jiang, Y., Wang, Z., Cao, Z., and Huang, T. (2016, January 15\u201319). UnitBox: An Advanced Object Detection Network. Proceedings of the 24th ACM International Conference on Multimedia, Amsterdam, The Netherlands.","DOI":"10.1145\/2964284.2967274"},{"key":"ref_59","doi-asserted-by":"crossref","unstructured":"Rezatofighi, H., Tsoi, N., Gwak, J., Sadeghian, A., Reid, I., and Savarese, S. (2019, January 15\u201320). Generalized Intersection Over Union: A Metric and a Loss for Bounding Box Regression. Proceedings of the 2019 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.00075"},{"key":"ref_60","first-page":"12993","article-title":"Distance-IoU Loss: Faster and Better Learning for Bounding Box Regression","volume":"34","author":"Zheng","year":"2020","journal-title":"Proc. Aaai Conf. Artif. Intell."},{"key":"ref_61","doi-asserted-by":"crossref","first-page":"146","DOI":"10.1016\/j.neucom.2022.07.042","article-title":"Focal and efficient IOU loss for accurate bounding box regression","volume":"506","author":"Zhang","year":"2022","journal-title":"Neurocomputing"},{"key":"ref_62","unstructured":"Tong, Z., Chen, Y., Xu, Z., and Yu, R. (2023). Wise-iou: Bounding box regression loss with dynamic focusing mechanism. arXiv."},{"key":"ref_63","doi-asserted-by":"crossref","unstructured":"Hu, J., Shen, L., and Sun, G. (2018, January 18\u201323). Squeeze-and-excitation networks. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00745"},{"key":"ref_64","unstructured":"Ioffe, S., and Szegedy, C. (2015, January 6\u201311). Batch normalization: Accelerating deep network training by reducing internal covariate shift. Proceedings of the 32nd International Conference on Machine Learning, Lille, France."},{"key":"ref_65","doi-asserted-by":"crossref","first-page":"111238","DOI":"10.1109\/ACCESS.2023.3321290","article-title":"Automatic Identifier of Socket for Electrical Vehicles Using SWIN-Transformer and SimAM Attention Mechanism-Based EVS YOLO","volume":"11","author":"Mahaadevan","year":"2023","journal-title":"IEEE Access"},{"key":"ref_66","doi-asserted-by":"crossref","first-page":"113936","DOI":"10.1016\/j.measurement.2023.113936","article-title":"STF-YOLO: A Small Target Detection Algorithm for UAV Remote Sensing Images Based on Improved SwinTransformer and Class Weighted Classification Decoupling Head","volume":"224","author":"Hui","year":"2024","journal-title":"Measurement"},{"key":"ref_67","doi-asserted-by":"crossref","unstructured":"Ouyang, D., He, S., Zhang, G., Luo, M., Guo, H., Zhan, J., and Huang, Z. (2023, January 4\u201310). Efficient Multi-Scale Attention Module with Cross-Spatial Learning. Proceedings of the ICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Rhodes Island, Greece.","DOI":"10.1109\/ICASSP49357.2023.10096516"},{"key":"ref_68","doi-asserted-by":"crossref","unstructured":"Xie, S., and Sun, H. (2023). Tea-YOLOv8s: A Tea Bud Detection Model Based on Deep Learning and Computer Vision. Sensors, 23.","DOI":"10.3390\/s23146576"},{"key":"ref_69","doi-asserted-by":"crossref","unstructured":"Yang, H., Min, Z., Zhang, Y., Wang, Z., and Jiang, D. (2021, January 10\u201314). An improved model-free finite control set predictive power control for PWM rectifiers. Proceedings of the 2021 IEEE Energy Conversion Congress and Exposition (ECCE), Vancouver, BC, Canada.","DOI":"10.1109\/ECCE47101.2021.9595084"},{"key":"ref_70","doi-asserted-by":"crossref","unstructured":"Hao, W., Ren, C., Han, M., Zhang, L., Li, F., and Liu, Z. (2023). Cattle Body Detection Based on YOLOv5-EMA for Precision Livestock Farming. Animals, 13.","DOI":"10.3390\/ani13223535"},{"key":"ref_71","unstructured":"Wang, C.Y., Liao, H.Y.M., and Yeh, I.H. (2022). Designing Network Design Strategies Through Gradient Path Analysis. arXiv."},{"key":"ref_72","doi-asserted-by":"crossref","unstructured":"Wang, C.-Y., Mark Liao, H.-Y., Wu, Y.-H., Chen, P.-Y., Hsieh, J.-W., and Yeh, I.-H. (2020, January 14\u201319). CSPNet: A New Backbone That Can Enhance Learning Capability of CNN. Proceedings of the 2020 IEEE\/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), Seattle, WA, USA.","DOI":"10.1109\/CVPRW50498.2020.00203"},{"key":"ref_73","doi-asserted-by":"crossref","unstructured":"Zhang, J., Chen, H., Yan, X., Zhou, K., Zhang, J., Zhang, Y., Jiang, H., and Shao, B. (2023). An Improved YOLOv5 Underwater Detector Based on an Attention Mechanism and Multi-Branch Reparameterization Module. Electronics, 12.","DOI":"10.3390\/electronics12122597"},{"key":"ref_74","doi-asserted-by":"crossref","unstructured":"Tian, Z., Shen, C., Chen, H., and He, T. (November, January 27). FCOS: Fully Convolutional One-Stage Object Detection. Proceedings of the 2019 IEEE\/CVF International Conference on Computer Vision (ICCV), Seoul, Republic of Korea.","DOI":"10.1109\/ICCV.2019.00972"},{"key":"ref_75","doi-asserted-by":"crossref","first-page":"3096","DOI":"10.1109\/TPAMI.2021.3050494","article-title":"Learning to Match Anchors for Visual Object Detection","volume":"44","author":"Zhang","year":"2022","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_76","unstructured":"Tan, M., and Le, Q.V. (2019, January 10\u201315). EfficientNet: Rethinking Model Scaling for Convolutional Neural Networks. Proceedings of the 36th International Conference on Machine Learning, Long Beach, CA, USA."},{"key":"ref_77","doi-asserted-by":"crossref","first-page":"107489","DOI":"10.1016\/j.asoc.2021.107489","article-title":"Edge Computing-Based Person Detection System for Top View Surveillance: Using CenterNet with Transfer Learning","volume":"107","author":"Ahmed","year":"2021","journal-title":"Appl. Soft Comput."},{"key":"ref_78","doi-asserted-by":"crossref","first-page":"68836","DOI":"10.1109\/ACCESS.2023.3287932","article-title":"Marine Organism Detection Based on Double Domains Augmentation and an Improved YOLOv7","volume":"11","author":"Zhang","year":"2023","journal-title":"IEEE Access"},{"key":"ref_79","doi-asserted-by":"crossref","first-page":"14881","DOI":"10.1007\/s00521-022-07264-8","article-title":"Refined Marine Object Detector with Attention-Based Spatial Pyramid Pooling Networks and Bidirectional Feature Fusion Strategy","volume":"34","author":"Xu","year":"2022","journal-title":"Neural Comput. Appl."}],"container-title":["Sensors"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1424-8220\/24\/9\/2905\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,10]],"date-time":"2025-10-10T14:38:25Z","timestamp":1760107105000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1424-8220\/24\/9\/2905"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,5,1]]},"references-count":79,"journal-issue":{"issue":"9","published-online":{"date-parts":[[2024,5]]}},"alternative-id":["s24092905"],"URL":"https:\/\/doi.org\/10.3390\/s24092905","relation":{},"ISSN":["1424-8220"],"issn-type":[{"value":"1424-8220","type":"electronic"}],"subject":[],"published":{"date-parts":[[2024,5,1]]}}}