{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,29]],"date-time":"2026-06-29T10:57:21Z","timestamp":1782730641945,"version":"3.54.5"},"reference-count":31,"publisher":"MDPI AG","issue":"10","license":[{"start":{"date-parts":[[2023,5,18]],"date-time":"2023-05-18T00:00:00Z","timestamp":1684368000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Sensors"],"abstract":"<jats:p>Underwater video object detection is a challenging task due to the poor quality of underwater videos, including blurriness and low contrast. In recent years, Yolo series models have been widely applied to underwater video object detection. However, these models perform poorly for blurry and low-contrast underwater videos. Additionally, they fail to account for the contextual relationships between the frame-level results. To address these challenges, we propose a video object detection model named UWV-Yolox. First, the Contrast Limited Adaptive Histogram Equalization method is used to augment the underwater videos. Then, a new CSP_CA module is proposed by adding Coordinate Attention to the backbone of the model to augment the representations of objects of interest. Next, a new loss function is proposed, including regression and jitter loss. Finally, a frame-level optimization module is proposed to optimize the detection results by utilizing the relationship between neighboring frames in videos, improving the video detection performance. To evaluate the performance of our model, We construct experiments on the UVODD dataset built in the paper, and select mAP@0.5 as the evaluation metric. The mAP@0.5 of the UWV-Yolox model reaches 89.0%, which is 3.2% better than the original Yolox model. Furthermore, compared with other object detection models, the UWV-Yolox model has more stable predictions for objects, and our improvements can be flexibly applied to other models.<\/jats:p>","DOI":"10.3390\/s23104859","type":"journal-article","created":{"date-parts":[[2023,5,18]],"date-time":"2023-05-18T07:35:50Z","timestamp":1684395350000},"page":"4859","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":17,"title":["UWV-Yolox: A Deep Learning Model for Underwater Video Object Detection"],"prefix":"10.3390","volume":"23","author":[{"given":"Haixia","family":"Pan","sequence":"first","affiliation":[{"name":"School of Software, Beihang University, Beijing 100191, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Jiahua","family":"Lan","sequence":"additional","affiliation":[{"name":"School of Software, Beihang University, Beijing 100191, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-5647-936X","authenticated-orcid":false,"given":"Hongqiang","family":"Wang","sequence":"additional","affiliation":[{"name":"School of Software, Beihang University, Beijing 100191, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Yanan","family":"Li","sequence":"additional","affiliation":[{"name":"School of Software, Beihang University, Beijing 100191, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Meng","family":"Zhang","sequence":"additional","affiliation":[{"name":"School of Software, Beihang University, Beijing 100191, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Mojie","family":"Ma","sequence":"additional","affiliation":[{"name":"School of Software, Beihang University, Beijing 100191, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Dongdong","family":"Zhang","sequence":"additional","affiliation":[{"name":"School of Software, Beihang University, Beijing 100191, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Xiaoran","family":"Zhao","sequence":"additional","affiliation":[{"name":"School of Software, Beihang University, Beijing 100191, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1968","published-online":{"date-parts":[[2023,5,18]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","first-page":"3212","DOI":"10.1109\/TNNLS.2018.2876865","article-title":"Object detection with deep learning: A review","volume":"30","author":"Zhao","year":"2019","journal-title":"IEEE Trans. Neural Netw. Learn. Syst."},{"key":"ref_2","unstructured":"Zuiderveld, K. (1994). Graphic Gems IV, Academic Press Professional."},{"key":"ref_3","first-page":"2","article-title":"Underwater Image Enhancement Using an Integrated Colour Model","volume":"34","author":"Iqbal","year":"2007","journal-title":"IAENG Int. J. Comput. Sci."},{"key":"ref_4","doi-asserted-by":"crossref","unstructured":"Huang, D., Wang, Y., Song, W., Sequeira, J., and Mavromatis, S. (2018, January 5\u20137). Shallow-water image enhancement using relative global histogram stretching based on adaptive parameter acquisition. Proceedings of the MultiMedia Modeling: 24th International Conference, MMM 2018, Bangkok, Thailand.","DOI":"10.1007\/978-3-319-73603-7_37"},{"key":"ref_5","doi-asserted-by":"crossref","unstructured":"Li, B., Peng, X., Wang, Z., Xu, J., and Feng, D. (2017, January 22\u201329). Aod-net: All-in-one dehazing network. Proceedings of the IEEE International Conference on Computer Vision, Venice, Italy.","DOI":"10.1109\/ICCV.2017.511"},{"key":"ref_6","doi-asserted-by":"crossref","unstructured":"Fu, M., Liu, H., Yu, Y., Chen, J., and Wang, K. (2021, January 20\u201325). Dw-gan: A discrete wavelet transform gan for nonhomogeneous dehazing. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Nashville, TN, USA.","DOI":"10.1109\/CVPRW53098.2021.00029"},{"key":"ref_7","doi-asserted-by":"crossref","first-page":"5277","DOI":"10.1007\/s00521-022-07964-1","article-title":"Toward visual quality enhancement of dehazing effect with improved Cycle-GAN","volume":"35","author":"Liu","year":"2022","journal-title":"Neural Comput. Appl."},{"key":"ref_8","doi-asserted-by":"crossref","unstructured":"Zhang, M., Xu, S., Song, W., He, Q., and Wei, Q. (2021). Lightweight underwater object detection based on yolo v4 and multi-scale attentional feature fusion. Remote Sens., 13.","DOI":"10.3390\/rs13224706"},{"key":"ref_9","doi-asserted-by":"crossref","unstructured":"Zhang, H., Wu, J., Yu, H., Wang, W., Zhang, Y., and Zhou, Y. (2021, January 20\u201321). An underwater fish individual recognition method based on improved YoloV4 and FaceNet. Proceedings of the 2021 20th International Conference on Ubiquitous Computing and Communications (IUCC\/CIT\/DSCI\/SmartCNS), London, UK.","DOI":"10.1109\/IUCC-CIT-DSCI-SmartCNS55181.2021.00042"},{"key":"ref_10","doi-asserted-by":"crossref","unstructured":"Li, S., Pan, B., Cheng, Y., Yan, X., Wang, C., and Yang, C. (2022, January 15\u201317). Underwater Fish Object Detection based on Attention Mechanism improved Ghost-YOLOv5. Proceedings of the 2022 7th International Conference on Intelligent Computing and Signal Processing (ICSP), Xi\u2019an, China.","DOI":"10.1109\/ICSP54964.2022.9778582"},{"key":"ref_11","doi-asserted-by":"crossref","first-page":"3195","DOI":"10.1109\/TNNLS.2021.3053249","article-title":"New generation deep learning for video object detection: A survey","volume":"33","author":"Jiao","year":"2021","journal-title":"IEEE Trans. Neural Netw. Learn. Syst."},{"key":"ref_12","unstructured":"Han, W., Khorrami, P., Paine, T.L., Ramachandran, P., Babaeizadeh, M., Shi, H., Li, J., Yan, S., and Huang, T.S. (2016). Seq-nms for Video Object Detection. arXiv."},{"key":"ref_13","unstructured":"Patraucean, V., Handa, A., and Cipolla, R. (2015). Spatio-Temporal Video Autoencoder with Differentiable Memory. arXiv."},{"key":"ref_14","doi-asserted-by":"crossref","unstructured":"Feichtenhofer, C., Pinz, A., and Zisserman, A. (2017, January 22\u201329). Detect to track and track to detect. Proceedings of the IEEE International Conference on Computer Vision, Venice, Italy.","DOI":"10.1109\/ICCV.2017.330"},{"key":"ref_15","unstructured":"Chai, Y. (November, January 27). Patchwork: A patch-wise attention network for efficient object detection and segmentation in video streams. Proceedings of the IEEE\/CVF International Conference on Computer Vision, Seoul, Republic of Korea."},{"key":"ref_16","unstructured":"Wang, T., Xiong, J., Xu, X., and Shi, Y. (February, January 27). SCNN: A general distribution based statistical convolutional neural network with application to video object detection. Proceedings of the AAAI Conference on Artificial Intelligence, Honolulu, HI, USA."},{"key":"ref_17","doi-asserted-by":"crossref","unstructured":"Hou, Q., Zhou, D., and Feng, J. (2021, January 20\u201325). Coordinate attention for efficient mobile network design. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Nashville, TN, USA.","DOI":"10.1109\/CVPR46437.2021.01350"},{"key":"ref_18","doi-asserted-by":"crossref","unstructured":"Kang, K., Ouyang, W., Li, H., and Wang, X. (2016, January 27\u201330). Object detection from video tubelets with convolutional neural networks. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.95"},{"key":"ref_19","doi-asserted-by":"crossref","unstructured":"He, L., Zhou, Q., Li, X., Niu, L., Cheng, G., Li, X., Liu, W., Tong, Y., Ma, L., and Zhang, L. (2021, January 20\u201324). End-to-end video object detection with spatial-temporal transformers. Proceedings of the 29th ACM International Conference on Multimedia, Chengdu, China.","DOI":"10.1145\/3474085.3475285"},{"key":"ref_20","doi-asserted-by":"crossref","unstructured":"Zhao, W., Zhang, J., Li, L., Barnes, N., Liu, N., and Han, J. (2021, January 20\u201325). Weakly supervised video salient object detection. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Nashville, TN, USA.","DOI":"10.1109\/CVPR46437.2021.01655"},{"key":"ref_21","doi-asserted-by":"crossref","unstructured":"Wen, G., Li, S., Liu, F., Luo, X., Er, M.J., Mahmud, M., and Wu, T. (2023). YOLOv5s-CA: A Modified YOLOv5s Network with Coordinate Attention for Underwater Target Detection. Sensors, 23.","DOI":"10.3390\/s23073367"},{"key":"ref_22","doi-asserted-by":"crossref","first-page":"400","DOI":"10.1214\/aoms\/1177729586","article-title":"A stochastic approximation method","volume":"22","author":"Robbins","year":"1951","journal-title":"Ann. Math. Stat."},{"key":"ref_23","unstructured":"Pedersen, M., Bruslund Haurum, J., Gade, R., and Moeslund, T.B. (2019, January 16\u201317). Detection of marine animals in a new underwater dataset with varying visibility. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition Workshops, Long Beach, CA, USA."},{"key":"ref_24","doi-asserted-by":"crossref","unstructured":"Jiang, L., Wang, Y., Jia, Q., Xu, S., Liu, Y., Fan, X., Li, H., Liu, R., Xue, X., and Wang, R. (2021, January 20\u201324). Underwater species detection using channel sharpening attention. Proceedings of the 29th ACM International Conference on Multimedia, Chengdu, China.","DOI":"10.1145\/3474085.3475563"},{"key":"ref_25","doi-asserted-by":"crossref","unstructured":"Liu, H., Song, P., and Ding, R. (2020, January 25\u201328). Towards domain generalization in underwater object detection. Proceedings of the 2020 IEEE International Conference on Image Processing (ICIP), Virtual Conference.","DOI":"10.1109\/ICIP40778.2020.9191364"},{"key":"ref_26","doi-asserted-by":"crossref","first-page":"211","DOI":"10.1007\/s11263-015-0816-y","article-title":"Imagenet large scale visual recognition challenge","volume":"115","author":"Russakovsky","year":"2015","journal-title":"Int. J. Comput. Vis."},{"key":"ref_27","doi-asserted-by":"crossref","unstructured":"Ancuti, C., Ancuti, C.O., Haber, T., and Bekaert, P. (2012, January 16\u201321). Enhancing underwater images and videos by fusion. Proceedings of the 2012 IEEE Conference on Computer Vision and Pattern Recognition, Providence, RI, USA.","DOI":"10.1109\/CVPR.2012.6247661"},{"key":"ref_28","doi-asserted-by":"crossref","first-page":"20368","DOI":"10.1109\/TITS.2022.3170328","article-title":"Cycle-snspgan: Towards real-world image dehazing via cycle spectral normalized soft likelihood estimation patch gan","volume":"23","author":"Wang","year":"2022","journal-title":"IEEE Trans. Intell. Transp. Syst."},{"key":"ref_29","doi-asserted-by":"crossref","unstructured":"Zhou, Q., Li, X., He, L., Yang, Y., Cheng, G., Tong, Y., Ma, L., and Tao, D. (2022). TransVOD: End-to-End Video Object Detection with Spatial-Temporal Transformers. arXiv.","DOI":"10.1109\/TPAMI.2022.3223955"},{"key":"ref_30","doi-asserted-by":"crossref","first-page":"150","DOI":"10.1016\/j.neucom.2023.01.088","article-title":"Boosting R-CNN: Reweighting R-CNN samples by RPN\u2019s error for underwater object detection","volume":"530","author":"Song","year":"2023","journal-title":"Neurocomputing"},{"key":"ref_31","unstructured":"Shi, Y., Wang, N., and Guo, X. (2022). YOLOV: Making Still Image Object Detectors Great at Video Object Detection. arXiv."}],"container-title":["Sensors"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1424-8220\/23\/10\/4859\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,10]],"date-time":"2025-10-10T19:37:31Z","timestamp":1760125051000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1424-8220\/23\/10\/4859"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,5,18]]},"references-count":31,"journal-issue":{"issue":"10","published-online":{"date-parts":[[2023,5]]}},"alternative-id":["s23104859"],"URL":"https:\/\/doi.org\/10.3390\/s23104859","relation":{},"ISSN":["1424-8220"],"issn-type":[{"value":"1424-8220","type":"electronic"}],"subject":[],"published":{"date-parts":[[2023,5,18]]}}}