{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,3,25]],"date-time":"2026-03-25T13:09:25Z","timestamp":1774444165323,"version":"3.50.1"},"reference-count":53,"publisher":"MDPI AG","issue":"4","license":[{"start":{"date-parts":[[2026,3,25]],"date-time":"2026-03-25T00:00:00Z","timestamp":1774396800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["J. Imaging"],"abstract":"<jats:p>Water-surface perception is critical for autonomous surface vehicle navigation, where reliable tracking of task-relevant objects is essential for safe and robust operation. Referring multi-object tracking (RMOT) provides a flexible tracking paradigm by allowing users to specify objects of interest through natural language. However, existing RMOT benchmarks are mainly designed for ground or satellite scenes and fail to capture the distinctive visual and semantic characteristics of water-surface environments, including strong reflections, severe illumination variations, weak motion constraints, and a high proportion of small objects. To address this gap, we introduce Refer-ASV, the first RMOT dataset tailored for ASV navigation in complex water-surface scenes. Refer-ASV is constructed from real-world ASV videos and features diverse navigation scenes and fine-grained vessel categories. To facilitate systematic evaluation on Refer-ASV, we further propose RAMOT, an end-to-end baseline framework that enhances visual\u2013language alignment throughout the tracking pipeline by improving visual\u2013language alignment and robustness in challenging maritime environments. Experimental results show that RAMOT achieves a HOTA score of 39.97 on Refer-ASV, outperforming existing methods. Additional experiments on Refer-KITTI demonstrate its generalization ability across different scenes.<\/jats:p>","DOI":"10.3390\/jimaging12040145","type":"journal-article","created":{"date-parts":[[2026,3,25]],"date-time":"2026-03-25T09:11:14Z","timestamp":1774429874000},"page":"145","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":0,"title":["Refer-ASV: Referring Multi-Object Tracking in Autonomous Surface Vehicle Navigation Scenes"],"prefix":"10.3390","volume":"12","author":[{"ORCID":"https:\/\/orcid.org\/0000-0003-3664-9410","authenticated-orcid":false,"given":"Bin","family":"Xue","sequence":"first","affiliation":[{"name":"State Key Laboratory of Multimodal Artificial Intelligence Systems, Institute of Automation, Chinese Academy of Sciences, Beijing 100190, China"},{"name":"School of Artificial Intelligence, University of Chinese Academy of Sciences, Beijing 101408, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-2119-090X","authenticated-orcid":false,"given":"Qiang","family":"Yu","sequence":"additional","affiliation":[{"name":"State Key Laboratory of Multimodal Artificial Intelligence Systems, Institute of Automation, Chinese Academy of Sciences, Beijing 100190, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-2256-8815","authenticated-orcid":false,"given":"Kun","family":"Ding","sequence":"additional","affiliation":[{"name":"State Key Laboratory of Multimodal Artificial Intelligence Systems, Institute of Automation, Chinese Academy of Sciences, Beijing 100190, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-1385-3224","authenticated-orcid":false,"given":"Ying","family":"Wang","sequence":"additional","affiliation":[{"name":"State Key Laboratory of Multimodal Artificial Intelligence Systems, Institute of Automation, Chinese Academy of Sciences, Beijing 100190, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-2089-9733","authenticated-orcid":false,"given":"Shiming","family":"Xiang","sequence":"additional","affiliation":[{"name":"State Key Laboratory of Multimodal Artificial Intelligence Systems, Institute of Automation, Chinese Academy of Sciences, Beijing 100190, China"},{"name":"School of Artificial Intelligence, University of Chinese Academy of Sciences, Beijing 101408, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-7433-4474","authenticated-orcid":false,"given":"Chunhong","family":"Pan","sequence":"additional","affiliation":[{"name":"State Key Laboratory of Multimodal Artificial Intelligence Systems, Institute of Automation, Chinese Academy of Sciences, Beijing 100190, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2026,3,25]]},"reference":[{"key":"ref_1","first-page":"37","article-title":"A survey on video detection and tracking of maritime vessels","volume":"20","author":"Moreira","year":"2014","journal-title":"Int. J. Recent Res. Appl. Stud."},{"key":"ref_2","doi-asserted-by":"crossref","unstructured":"Cheng, Y., Zhu, J., Jiang, M., Fu, J., Pang, C., Wang, P., Sankaran, K., Onabola, O., Liu, Y., and Liu, D. (2021, January 10\u201317). Flow: A dataset and benchmark for floating waste detection in inland waters. Proceedings of the IEEE International Conference on Computer Vision, Montreal, QC, Canada.","DOI":"10.1109\/ICCV48922.2021.01077"},{"key":"ref_3","doi-asserted-by":"crossref","first-page":"1433","DOI":"10.1109\/TIM.2019.2963515","article-title":"A low-cost unmanned surface vehicle for pervasive water quality monitoring","volume":"69","author":"Madeo","year":"2020","journal-title":"IEEE Trans. Instrum. Meas."},{"key":"ref_4","doi-asserted-by":"crossref","unstructured":"Wang, W., Shan, T., Leoni, P., Fern\u00e1ndez-Guti\u00e9rrez, D., Meyers, D., Ratti, C., and Rus, D. (2020\u201324, January 24). Roboat II: A novel autonomous surface vessel for urban environments. Proceedings of the IEEE International Conference on Intelligent Robots and Systems, Las Vegas, NV, USA.","DOI":"10.1109\/IROS45743.2020.9340712"},{"key":"ref_5","doi-asserted-by":"crossref","unstructured":"Verma, V., Gupta, D., Gupta, S., Uppal, M., Anand, D., Ortega-Mansilla, A., Alharithi, F.S., Almotiri, J., and Goyal, N. (2022). A deep learning-based intelligent garbage detection system using an unmanned aerial vehicle. Symmetry, 14.","DOI":"10.3390\/sym14050960"},{"key":"ref_6","doi-asserted-by":"crossref","first-page":"2781","DOI":"10.1007\/s11277-022-09841-5","article-title":"Compressive sensing node localization method using autonomous underwater vehicle network","volume":"126","author":"Kulandaivel","year":"2022","journal-title":"Wirel. Pers. Commun."},{"key":"ref_7","doi-asserted-by":"crossref","first-page":"574","DOI":"10.32629\/jai.v6i2.574","article-title":"Enhance traffic flow prediction with real-time vehicle data integration","volume":"6","author":"Jain","year":"2023","journal-title":"J. Auton. Intell."},{"key":"ref_8","doi-asserted-by":"crossref","first-page":"81","DOI":"10.5916\/jamet.2024.48.2.81","article-title":"Object detection for various types of vessels using the YOLO algorithm","volume":"48","author":"Park","year":"2024","journal-title":"J. Adv. Mar. Eng. Technol."},{"key":"ref_9","doi-asserted-by":"crossref","first-page":"117243","DOI":"10.1016\/j.oceaneng.2024.117243","article-title":"Sea-IoUTracker: A more stable and reliable maritime target tracking scheme for unmanned vessel platforms","volume":"299","author":"Guo","year":"2024","journal-title":"Ocean. Eng."},{"key":"ref_10","doi-asserted-by":"crossref","unstructured":"Wu, D., Han, W., Wang, T., Dong, X., Zhang, X., and Shen, J. (2023, January 17\u201324). Referring multi-object tracking. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Vancouver, BC, Canada.","DOI":"10.1109\/CVPR52729.2023.01406"},{"key":"ref_11","doi-asserted-by":"crossref","first-page":"16584","DOI":"10.1109\/TITS.2024.3415772","article-title":"WaterScenes: A multi-task 4D radar-camera fusion dataset and benchmarks for autonomous driving on water surfaces","volume":"25","author":"Yao","year":"2024","journal-title":"IEEE Trans. Intell. Transp. Syst."},{"key":"ref_12","doi-asserted-by":"crossref","unstructured":"Bengio, Y., Louradour, J., Collobert, R., and Weston, J. (2009, January 14\u201318). Curriculum learning. Proceedings of the International Conference on Machine Learning, Montreal, QC, Canada.","DOI":"10.1145\/1553374.1553380"},{"key":"ref_13","doi-asserted-by":"crossref","first-page":"103448","DOI":"10.1016\/j.artint.2020.103448","article-title":"Multiple object tracking: A literature review","volume":"293","author":"Luo","year":"2021","journal-title":"Artif. Intell."},{"key":"ref_14","doi-asserted-by":"crossref","unstructured":"Wojke, N., Bewley, A., and Paulus, D. (2017, January 17\u201320). Simple online and realtime tracking with a deep association metric. Proceedings of the IEEE International Conference on Image Processing, Beijing, China.","DOI":"10.1109\/ICIP.2017.8296962"},{"key":"ref_15","doi-asserted-by":"crossref","first-page":"4870","DOI":"10.1109\/TCSVT.2024.3524670","article-title":"Sparsetrack: Multi-object tracking by performing scene decomposition based on pseudo-depth","volume":"35","author":"Liu","year":"2025","journal-title":"IEEE Trans. Circuits Syst. Video Technol."},{"key":"ref_16","doi-asserted-by":"crossref","unstructured":"Zeng, L., Huang, Y., Lin, Y., Zhu, Z., Zhang, X., Liu, Y., Li, Y., and Zheng, Y. (2026). SORT-LFR: Revisiting sort for multi-object tracking in low-frame-rate videos. IEEE Trans. Multimed., early access.","DOI":"10.1109\/TMM.2026.3651130"},{"key":"ref_17","doi-asserted-by":"crossref","first-page":"3069","DOI":"10.1007\/s11263-021-01513-4","article-title":"Fairmot: On the fairness of detection and re-identification in multiple object tracking","volume":"129","author":"Zhang","year":"2021","journal-title":"Int. J. Comput. Vis."},{"key":"ref_18","doi-asserted-by":"crossref","first-page":"3182","DOI":"10.1109\/TIP.2022.3165376","article-title":"Rethinking the competition between detection and reid in multiobject tracking","volume":"31","author":"Liang","year":"2022","journal-title":"IEEE Trans. Image Process."},{"key":"ref_19","doi-asserted-by":"crossref","first-page":"15380","DOI":"10.1109\/TPAMI.2023.3301975","article-title":"Qdtrack: Quasi-dense similarity learning for appearance-only multiple object tracking","volume":"45","author":"Fischer","year":"2023","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_20","doi-asserted-by":"crossref","first-page":"5117","DOI":"10.1109\/TCSVT.2023.3249162","article-title":"Multi-object tracking: Decoupling features to solve the contradictory dilemma of feature requirements","volume":"33","author":"Jin","year":"2023","journal-title":"IEEE Trans. Circuits Syst. Video Technol."},{"key":"ref_21","doi-asserted-by":"crossref","first-page":"7820","DOI":"10.1109\/TPAMI.2022.3225078","article-title":"TransCenter: Transformers with dense representations for multiple-object tracking","volume":"45","author":"Xu","year":"2022","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_22","doi-asserted-by":"crossref","unstructured":"Zeng, F., Dong, B., Zhang, Y., Wang, T., Zhang, X., and Wei, Y. (2022, January 23\u201327). Motr: End-to-end multiple-object tracking with transformer. Proceedings of the European Conference on Computer Vision, Tel Aviv, Israel.","DOI":"10.1007\/978-3-031-19812-0_38"},{"key":"ref_23","doi-asserted-by":"crossref","first-page":"6571","DOI":"10.1109\/TCSVT.2023.3263884","article-title":"STDFormer: Spatial-temporal motion transformer for multiple object tracking","volume":"33","author":"Hu","year":"2023","journal-title":"IEEE Trans. Circuits Syst. Video Technol."},{"key":"ref_24","doi-asserted-by":"crossref","first-page":"125129","DOI":"10.1016\/j.eswa.2024.125129","article-title":"Lightweight multiobject ship tracking algorithm based on trajectory association and improved YOLOv7tiny","volume":"259","author":"Hao","year":"2025","journal-title":"Expert Syst. Appl."},{"key":"ref_25","doi-asserted-by":"crossref","first-page":"1993","DOI":"10.1109\/TITS.2016.2634580","article-title":"Video processing from electro-optical sensors for object detection and tracking in a maritime environment: A survey","volume":"18","author":"Prasad","year":"2017","journal-title":"IEEE Trans. Intell. Transp. Syst."},{"key":"ref_26","doi-asserted-by":"crossref","unstructured":"Choi, J., Cho, D., Lee, G., Kim, H., Yang, G., Kim, J., and Cho, Y. (2025, January 19\u201323). Polaris dataset: A maritime object detection and tracking dataset in Pohang canal. Proceedings of the IEEE International Conference on Robotics and Automation, Atlanta, GA, USA.","DOI":"10.1109\/ICRA55743.2025.11128583"},{"key":"ref_27","unstructured":"Yao, S., Guan, R., Ni, Y., Xu, S., Yue, Y., Zhu, X., and Liu, R.W. (2025, January 19\u201325). USVTrack: USV-based 4D radar-camera tracking dataset for autonomous driving in inland waterways. Proceedings of the IEEE International Conference on Intelligent Robots and Systems, Hangzhou, China."},{"key":"ref_28","doi-asserted-by":"crossref","first-page":"1029","DOI":"10.1109\/TCSVT.2025.3595760","article-title":"USVTrack: A benchmark for multi-object tracking in complex water surface scenes","volume":"36","author":"Xue","year":"2026","journal-title":"IEEE Trans. Circuits Syst. Video Technol."},{"key":"ref_29","unstructured":"Zhang, Y., Wu, D., Han, W., and Dong, X. (2024). Bootstrapping referring multi-object tracking. arXiv."},{"key":"ref_30","first-page":"5004613","article-title":"Multigranularity localization transformer with collaborative understanding for referring multiobject tracking","volume":"74","author":"Chen","year":"2025","journal-title":"IEEE Trans. Instrum. Meas."},{"key":"ref_31","doi-asserted-by":"crossref","first-page":"103349","DOI":"10.1016\/j.inffus.2025.103349","article-title":"Cognitive disentanglement for referring multi-object tracking","volume":"124","author":"Liang","year":"2025","journal-title":"Inf. Fusion"},{"key":"ref_32","doi-asserted-by":"crossref","unstructured":"Du, Y., Lei, C., Zhao, Z., and Su, F. (2024, January 16\u201322). iKUN: Speak to trackers without retraining. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA.","DOI":"10.1109\/CVPR52733.2024.01810"},{"key":"ref_33","doi-asserted-by":"crossref","unstructured":"Geiger, A., Lenz, P., and Urtasun, R. (2012, January 16\u201321). Are we ready for autonomous driving? the kitti vision benchmark suite. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Providence, RI, USA.","DOI":"10.1109\/CVPR.2012.6248074"},{"key":"ref_34","doi-asserted-by":"crossref","unstructured":"Sun, P., Cao, J., Jiang, Y., Yuan, Z., Bai, S., Kitani, K., and Luo, P. (2022, January 18\u201324). Dancetrack: Multi-object tracking in uniform appearance and diverse motion. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, New Orleans, LA, USA.","DOI":"10.1109\/CVPR52688.2022.02032"},{"key":"ref_35","unstructured":"Nguyen, P., Quach, K.G., Kitani, K., and Luu, K. (2023, January 10\u201316). Type-to-track: Retrieve any object via prompt-based tracking. Proceedings of the Advances in Neural Information Processing Systems, New Orleans, LA, USA."},{"key":"ref_36","doi-asserted-by":"crossref","unstructured":"Zhang, P., Zhang, Y., Shi, H., Wang, D., Liu, X., and Wang, L. (2025, January 27\u201331). Referring multi-object tracking in satellite videos: A new benchmark and baseline. Proceedings of the ACM International Conference on Multimedia, Dublin, Ireland.","DOI":"10.1145\/3746027.3758254"},{"key":"ref_37","doi-asserted-by":"crossref","first-page":"7275","DOI":"10.1109\/TITS.2025.3527011","article-title":"WaterVG: Waterway visual grounding based on text-guided vision and mmWave radar","volume":"26","author":"Guan","year":"2025","journal-title":"IEEE Trans. Intell. Transp. Syst."},{"key":"ref_38","doi-asserted-by":"crossref","unstructured":"Guan, R., Liu, J., Jia, L., Zhao, H., Yao, S., Zhu, X., Man, K.L., Lim, E.G., Smith, J., and Yue, Y. (2025, January 19\u201325). NanoMVG: USV-centric low-power multi-task visual grounding based on prompt-guided camera and 4D mmWave radar. Proceedings of the IEEE International Conference on Intelligent Robots and Systems, Hangzhou, China.","DOI":"10.1109\/IROS60139.2025.11246532"},{"key":"ref_39","doi-asserted-by":"crossref","unstructured":"Guan, R., Ouyang, N., Xu, T., Liang, S., Dai, W., Sun, Y., Gao, S., Lai, S., Yao, S., and Hu, X. (2026). Da yu: Towards USV-Based Image Captioning for Waterway Surveillance and Scene Understanding. IEEE Trans. Circuits Syst. Video Technol., early access.","DOI":"10.1109\/TCSVT.2026.3651269"},{"key":"ref_40","doi-asserted-by":"crossref","first-page":"845","DOI":"10.1007\/s11263-020-01393-0","article-title":"Motchallenge: A benchmark for single-camera multiple target tracking","volume":"129","author":"Dendorfer","year":"2021","journal-title":"Int. J. Comput. Vis."},{"key":"ref_41","unstructured":"Yang, A., Li, A., Yang, B., Zhang, B., Hui, B., Zheng, B., Yu, B., Gao, C., Huang, C., and Lv, C. (2025). Qwen3 technical report. arXiv."},{"key":"ref_42","doi-asserted-by":"crossref","unstructured":"He, K., Zhang, X., Ren, S., and Sun, J. (2016, January 27\u201330). Deep residual learning for image recognition. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.90"},{"key":"ref_43","unstructured":"Liu, Y., Ott, M., Goyal, N., Du, J., Joshi, M., Chen, D., Levy, O., Lewis, M., Zettlemoyer, L., and Stoyanov, V. (2019). Roberta: A robustly optimized bert pretraining approach. arXiv."},{"key":"ref_44","unstructured":"Zhu, X., Su, W., Lu, L., Li, B., Wang, X., and Dai, J. (2020). Deformable detr: Deformable transformers for end-to-end object detection. arXiv."},{"key":"ref_45","doi-asserted-by":"crossref","unstructured":"Lin, T.Y., Doll\u00e1r, P., Girshick, R., He, K., Hariharan, B., and Belongie, S. (2017, January 21\u201326). Feature pyramid networks for object detection. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.106"},{"key":"ref_46","doi-asserted-by":"crossref","unstructured":"Carion, N., Massa, F., Synnaeve, G., Usunier, N., Kirillov, A., and Zagoruyko, S. (2020). End-to-end object detection with transformers. Proceedings of the European Conference on Computer Vision, Springer.","DOI":"10.1007\/978-3-030-58452-8_13"},{"key":"ref_47","unstructured":"Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A.N., Kaiser, \u0141., and Polosukhin, I. (2017, January 4\u20139). Attention is all you need. Proceedings of the Advances in Neural Information Processing Systems, Long Beach, CA, USA."},{"key":"ref_48","doi-asserted-by":"crossref","first-page":"548","DOI":"10.1007\/s11263-020-01375-2","article-title":"Hota: A higher order metric for evaluating multi-object tracking","volume":"129","author":"Luiten","year":"2021","journal-title":"Int. J. Comput. Vis."},{"key":"ref_49","doi-asserted-by":"crossref","first-page":"246309","DOI":"10.1155\/2008\/246309","article-title":"Evaluating multiple object tracking performance: The clear mot metrics","volume":"2008","author":"Bernardin","year":"2008","journal-title":"EURASIP J. Image Video Process."},{"key":"ref_50","doi-asserted-by":"crossref","unstructured":"Ristani, E., Solera, F., Zou, R., Cucchiara, R., and Tomasi, C. (2016). Performance measures and a data set for multi-target, multi-camera tracking. Proceedings of the European Conference on Computer Vision, Springer.","DOI":"10.1007\/978-3-319-48881-3_2"},{"key":"ref_51","doi-asserted-by":"crossref","unstructured":"Lin, T.Y., Goyal, P., Girshick, R., He, K., and Doll\u00e1r, P. (2017, January 22\u201329). Focal loss for dense object detection. Proceedings of the IEEE International Conference on Computer Vision, Venice, Italy.","DOI":"10.1109\/ICCV.2017.324"},{"key":"ref_52","doi-asserted-by":"crossref","unstructured":"Rezatofighi, H., Tsoi, N., Gwak, J., Sadeghian, A., Reid, I., and Savarese, S. (2019, January 15\u201320). Generalized intersection over union: A metric and a loss for bounding box regression. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.00075"},{"key":"ref_53","doi-asserted-by":"crossref","unstructured":"He, W., Jian, Y., Lu, Y., and Wang, H. (2024, January 14\u201319). Visual-linguistic representation learning with deep cross-modality fusion for referring multi-object tracking. Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing, Seoul, Republic of Korea.","DOI":"10.1109\/ICASSP48485.2024.10447535"}],"container-title":["Journal of Imaging"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/2313-433X\/12\/4\/145\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,3,25]],"date-time":"2026-03-25T10:21:09Z","timestamp":1774434069000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/2313-433X\/12\/4\/145"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,3,25]]},"references-count":53,"journal-issue":{"issue":"4","published-online":{"date-parts":[[2026,4]]}},"alternative-id":["jimaging12040145"],"URL":"https:\/\/doi.org\/10.3390\/jimaging12040145","relation":{},"ISSN":["2313-433X"],"issn-type":[{"value":"2313-433X","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,3,25]]}}}