{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,12,12]],"date-time":"2025-12-12T13:48:37Z","timestamp":1765547317899,"version":"build-2065373602"},"reference-count":46,"publisher":"MDPI AG","issue":"16","license":[{"start":{"date-parts":[[2024,8,20]],"date-time":"2024-08-20T00:00:00Z","timestamp":1724112000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"name":"Double First-Class Innovation Research Project for the People\u2019s Public Security University of China","award":["2023SYL08"],"award-info":[{"award-number":["2023SYL08"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Sensors"],"abstract":"<jats:p>Referring video object segmentation (R-VOS) is a fundamental vision-language task which aims to segment the target referred by language expression in all video frames. Existing query-based R-VOS methods have conducted in-depth exploration of the interaction and alignment between visual and linguistic features but fail to transfer the information of the two modalities to the query vector with balanced intensities. Furthermore, most of the traditional approaches suffer from severe information loss in the process of multi-scale feature fusion, resulting in inaccurate segmentation. In this paper, we propose DCT, an end-to-end decoupled cross-modal transformer for referring video object segmentation, to better utilize multi-modal and multi-scale information. Specifically, we first design a Language-Guided Visual Enhancement Module (LGVE) to transmit discriminative linguistic information to visual features of all levels, performing an initial filtering of irrelevant background regions. Then, we propose a decoupled transformer decoder, using a set of object queries to gather entity-related information from both visual and linguistic features independently, mitigating the attention bias caused by feature size differences. Finally, the Cross-layer Feature Pyramid Network (CFPN) is introduced to preserve more visual details by establishing direct cross-layer communication. Extensive experiments have been carried out on A2D-Sentences, JHMDB-Sentences and Ref-Youtube-VOS. The results show that DCT achieves competitive segmentation accuracy compared with the state-of-the-art methods.<\/jats:p>","DOI":"10.3390\/s24165375","type":"journal-article","created":{"date-parts":[[2024,8,20]],"date-time":"2024-08-20T09:13:48Z","timestamp":1724145228000},"page":"5375","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":3,"title":["Decoupled Cross-Modal Transformer for Referring Video Object Segmentation"],"prefix":"10.3390","volume":"24","author":[{"given":"Ao","family":"Wu","sequence":"first","affiliation":[{"name":"School of Information and Cyber Security, People\u2019s Public Security University of China, Beijing 100038, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Rong","family":"Wang","sequence":"additional","affiliation":[{"name":"School of Information and Cyber Security, People\u2019s Public Security University of China, Beijing 100038, China"},{"name":"Key Laboratory of Security Prevention Technology and Risk Assessment of Ministry of Public Security, Beijing 100038, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Quange","family":"Tan","sequence":"additional","affiliation":[{"name":"School of Information and Cyber Security, People\u2019s Public Security University of China, Beijing 100038, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Zhenfeng","family":"Song","sequence":"additional","affiliation":[{"name":"School of Information and Cyber Security, People\u2019s Public Security University of China, Beijing 100038, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2024,8,20]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","unstructured":"Caelles, S., Maninis, K.-K., Pont-Tuset, J., Leal-Taix\u00e9, L., Cremers, D., and Van Gool, L. (2017, January 21\u201326). One-shot video object segmentation. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.565"},{"key":"ref_2","doi-asserted-by":"crossref","unstructured":"Maninis, K.-K., Caelles, S., Pont-Tuset, J., and Van Gool, L. (2018, January 18\u201323). Deep extreme cut: From extreme points to object segmentation. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00071"},{"key":"ref_3","doi-asserted-by":"crossref","unstructured":"Yang, Z., Wei, Y., and Yang, Y. (2020, January 23\u201328). Collaborative video object segmentation by foreground-background integration. Proceedings of the European Conference on Computer Vision, Online.","DOI":"10.1007\/978-3-030-58558-7_20"},{"key":"ref_4","doi-asserted-by":"crossref","unstructured":"Hu, R., Rohrbach, M., and Darrell, T. (2016). Segmentation from natural language expressions. Computer Vision\u2013ECCV 2016, Proceedings of the 14th European Conference, Amsterdam, The Netherlands, 11\u201314 October 2016, Part I 14, Springer International Publishing.","DOI":"10.1007\/978-3-319-46448-0_7"},{"key":"ref_5","unstructured":"Bellver, M., Ventura, C., Silberer, C., Kazakos, I., Torres, J., and Giro-i-Nieto, X. (2020). Refvos: A closer look at referring expressions for video object segmentation. arXiv."},{"key":"ref_6","doi-asserted-by":"crossref","unstructured":"Liu, C., Lin, Z., Shen, X., Yang, J., Lu, X., and Yuille, A. (2017, January 22\u201329). Recurrent multimodal interaction for referring image segmentation. Proceedings of the IEEE International Conference on Computer Vision, Venice, Italy.","DOI":"10.1109\/ICCV.2017.143"},{"key":"ref_7","doi-asserted-by":"crossref","unstructured":"Gavrilyuk, K., Ghodrati, A., Li, Z., and Snoek, C.G. (2018, January 18\u201323). Actor and action video segmentation from a sentence. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00624"},{"key":"ref_8","doi-asserted-by":"crossref","unstructured":"Wang, H., Deng, C., Ma, F., and Yang, Y. (2020, January 7\u201312). Context modulated dynamic networks for actor and action video segmentation with language queries. Proceedings of the AAAI Conference on Artificial Intelligence, New York, NY, USA.","DOI":"10.1609\/aaai.v34i07.6895"},{"key":"ref_9","doi-asserted-by":"crossref","unstructured":"Ye, L., Rochan, M., Liu, Z., and Wang, Y. (2019, January 15\u201320). Cross-modal self-attention network for referring image segmentation. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.01075"},{"key":"ref_10","doi-asserted-by":"crossref","unstructured":"Hu, Z., Feng, G., Sun, J., Zhang, L., and Lu, H. (2020, January 13\u201319). Bi-directional relationship inferring network for referring image segmentation. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA.","DOI":"10.1109\/CVPR42600.2020.00448"},{"key":"ref_11","doi-asserted-by":"crossref","unstructured":"Yang, Z., Wang, J., Tang, Y., Chen, K., Zhao, H., and Torr, P.H. (2022, January 18\u201324). Lavt: Language-aware vision transformer for referring image segmentation. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, New Orleans, LA, USA.","DOI":"10.1109\/CVPR52688.2022.01762"},{"key":"ref_12","doi-asserted-by":"crossref","unstructured":"Seo, S., Lee, J.-Y., and Han, B. (2020). Urvos: Unified referring video object segmentation network with a large-scale benchmark. Computer Vision\u2013ECCV 2020, Proceedings of the 16th European Conference, Glasgow, UK, 23\u201328 August 2020, Part XV 16, Springer International Publishing.","DOI":"10.1007\/978-3-030-58555-6_13"},{"key":"ref_13","doi-asserted-by":"crossref","unstructured":"Botach, A., Zheltonozhskii, E., and Baskin, C. (2022, January 18\u201324). End-to-end referring video object segmentation with multimodal transformers. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, New Orleans, LA, USA.","DOI":"10.1109\/CVPR52688.2022.00493"},{"key":"ref_14","doi-asserted-by":"crossref","unstructured":"Wu, J., Jiang, Y., Sun, P., Yuan, Z., and Luo, P. (2022, January 18\u201324). Language as queries for referring video object segmentation. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, New Orleans, LA, USA.","DOI":"10.1109\/CVPR52688.2022.00492"},{"key":"ref_15","doi-asserted-by":"crossref","unstructured":"Li, X., Wang, J., Xu, X., Li, X., Lu, Y., and Raj, B. (2022). R^2VOS: Robust Referring Video Object Segmentation via Relational Multimodal Cycle Consistency. arXiv.","DOI":"10.1109\/ICCV51070.2023.02032"},{"key":"ref_16","unstructured":"Yuan, L., Shi, M., and Yue, Z. (2023). LoSh: Long-Short Text Joint Prediction Network for Referring Video Object Segmentation. arXiv."},{"key":"ref_17","unstructured":"Feng, G., Zhang, L., Hu, Z., and Lu, H. (2022). Deeply Interleaved Two-Stream Encoder for Referring Video Segmentation. arXiv."},{"key":"ref_18","doi-asserted-by":"crossref","unstructured":"Lin, T.-Y., Doll\u00e1r, P., Girshick, R., He, K., Hariharan, B., and Belongie, S. (2017, January 21\u201326). Feature pyramid networks for object detection. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.106"},{"key":"ref_19","doi-asserted-by":"crossref","unstructured":"Carion, N., Massa, F., Synnaeve, G., Usunier, N., Kirillov, A., and Zagoruyko, S. (2020, January 23\u201328). End-to-end object detection with transformers. Proceedings of the European Conference on Computer Vision, Online.","DOI":"10.1007\/978-3-030-58452-8_13"},{"key":"ref_20","doi-asserted-by":"crossref","first-page":"4587","DOI":"10.1109\/TIP.2021.3072811","article-title":"Cross-layer feature pyramid network for salient object detection","volume":"30","author":"Li","year":"2021","journal-title":"IEEE Trans. Image Process."},{"key":"ref_21","doi-asserted-by":"crossref","unstructured":"Hui, T., Liu, S., Huang, S., Li, G., Yu, S., Zhang, F., and Han, J. (2020). Linguistic structure guided context modeling for referring image segmentation. Computer Vision\u2013ECCV 2020, Proceedings of the 16th European Conference, Glasgow, UK, 23\u201328 August 2020, Part X 16, Springer International Publishing.","DOI":"10.1007\/978-3-030-58607-2_4"},{"key":"ref_22","unstructured":"Liang, C., Wu, Y., Zhou, T., Wang, W., Yang, Z., Wei, Y., and Yang, Y. (2021). Rethinking cross-modal interaction from a top-down perspective for referring video object segmentation. arXiv."},{"key":"ref_23","doi-asserted-by":"crossref","first-page":"4474","DOI":"10.1109\/TIP.2022.3185487","article-title":"Actor and action modular network for text-based video segmentation","volume":"31","author":"Yang","year":"2022","journal-title":"IEEE Trans. Image Process."},{"key":"ref_24","doi-asserted-by":"crossref","unstructured":"Yu, L., Lin, Z., Shen, X., Yang, J., Lu, X., Bansal, M., and Berg, T.L. (2018, January 18\u201323). Mattnet: Modular attention network for referring expression comprehension. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00142"},{"key":"ref_25","unstructured":"Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A.N., Kaiser, \u0141., and Polosukhin, I. (2017). Attention is all you need. arXiv."},{"key":"ref_26","first-page":"1877","article-title":"Language models are few-shot learners","volume":"33","author":"Brown","year":"2020","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_27","unstructured":"Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K. (2018). Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv."},{"key":"ref_28","unstructured":"Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., and Gelly, S. (2020). An image is worth 16x16 words: Transformers for image recognition at scale. arXiv."},{"key":"ref_29","doi-asserted-by":"crossref","unstructured":"Liu, Z., Lin, Y., Cao, Y., Hu, H., Wei, Y., Zhang, Z., Lin, S., and Guo, B. (2021, January 11\u201317). Swin transformer: Hierarchical vision transformer using shifted windows. Proceedings of the IEEE\/CVF International Conference on Computer Vision, Montreal, BC, Canada.","DOI":"10.1109\/ICCV48922.2021.00986"},{"key":"ref_30","unstructured":"Radford, A., Kim, J.W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., and Clark, J. (2021, January 18\u201324). Learning transferable visual models from natural language supervision. Proceedings of the International Conference on Machine Learning, Online."},{"key":"ref_31","doi-asserted-by":"crossref","unstructured":"Ding, H., Liu, C., Wang, S., and Jiang, X. (2021, January 11\u201317). Vision-language transformer and query generation for referring segmentation. Proceedings of the IEEE\/CVF International Conference on Computer Vision, Montreal, BC, Canada.","DOI":"10.1109\/ICCV48922.2021.01601"},{"key":"ref_32","doi-asserted-by":"crossref","unstructured":"Kamath, A., Singh, M., LeCun, Y., Synnaeve, G., Misra, I., and Carion, N. (2021, January 11\u201317). Mdetr-modulated detection for end-to-end multi-modal understanding. Proceedings of the IEEE\/CVF International Conference on Computer Vision, Montreal, BC, Canada.","DOI":"10.1109\/ICCV48922.2021.00180"},{"key":"ref_33","doi-asserted-by":"crossref","unstructured":"Wang, Y., Xu, Z., Wang, X., Shen, C., Cheng, B., Shen, H., and Xia, H. (2021, January 20\u201325). End-to-end video instance segmentation with transformers. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Nashville, TN, USA.","DOI":"10.1109\/CVPR46437.2021.00863"},{"key":"ref_34","doi-asserted-by":"crossref","unstructured":"Liu, Z., Ning, J., Cao, Y., Wei, Y., Zhang, Z., Lin, S., and Hu, H. (2022, January 18\u201324). Video swin transformer. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, New Orleans, LA, USA.","DOI":"10.1109\/CVPR52688.2022.00320"},{"key":"ref_35","unstructured":"Liu, Y., Ott, M., Goyal, N., Du, J., Joshi, M., Chen, D., Levy, O., Lewis, M., Zettlemoyer, L., and Stoyanov, V. (2019). Roberta: A robustly optimized bert pretraining approach. arXiv."},{"key":"ref_36","doi-asserted-by":"crossref","first-page":"83","DOI":"10.1002\/nav.3800020109","article-title":"The Hungarian method for the assignment problem","volume":"2","author":"Kuhn","year":"1955","journal-title":"Nav. Res. Logist. Q."},{"key":"ref_37","doi-asserted-by":"crossref","unstructured":"Milletari, F., Navab, N., and Ahmadi, S.-A. (2016, January 25\u201328). V-net: Fully convolutional neural networks for volumetric medical image segmentation. Proceedings of the 2016 Fourth International Conference on 3D Vision (3DV), Stanford, CA, USA.","DOI":"10.1109\/3DV.2016.79"},{"key":"ref_38","doi-asserted-by":"crossref","unstructured":"Lin, T.-Y., Goyal, P., Girshick, R., He, K., and Doll\u00e1r, P. (2017, January 22\u201329). Focal loss for dense object detection. Proceedings of the IEEE International Conference on Computer Vision, Venice, Italy.","DOI":"10.1109\/ICCV.2017.324"},{"key":"ref_39","unstructured":"Xu, C., Xiong, C., and Corso, J.J. (2017). Action understanding with multiple classes of actors. arXiv."},{"key":"ref_40","doi-asserted-by":"crossref","unstructured":"Jhuang, H., Gall, J., Zuffi, S., Schmid, C., and Black, M.J. (2013, January 1\u20138). Towards understanding action recognition. Proceedings of the IEEE International Conference on Computer Vision, Sydney, Australia.","DOI":"10.1109\/ICCV.2013.396"},{"key":"ref_41","doi-asserted-by":"crossref","unstructured":"Xu, N., Yang, L., Fan, Y., Yang, J., Yue, D., Liang, Y., Price, B., Cohen, S., and Huang, T. (2018, January 8\u201314). Youtube-vos: Sequence-to-sequence video object segmentation. Proceedings of the European Conference on Computer Vision (ECCV), Munich, Germany.","DOI":"10.1007\/978-3-030-01228-1_36"},{"key":"ref_42","first-page":"3719","article-title":"Referring segmentation in images and videos with cross-modal self-attention network","volume":"44","author":"Ye","year":"2021","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_43","doi-asserted-by":"crossref","unstructured":"Wang, H., Deng, C., Yan, J., and Tao, D. (2019, January 27\u201328). Asymmetric cross-guided attention network for actor and action video segmentation from natural language query. Proceedings of the IEEE\/CVF International Conference on Computer Vision, Seoul, Republic of Korea.","DOI":"10.1109\/ICCV.2019.00404"},{"key":"ref_44","first-page":"4761","article-title":"Cross-modal progressive comprehension for referring segmentation","volume":"44","author":"Liu","year":"2021","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_45","unstructured":"Liang, C., Wu, Y., Luo, Y., and Yang, Y. (2021). Clawcranenet: Leveraging object-level relation for text-based video segmentation. arXiv."},{"key":"ref_46","doi-asserted-by":"crossref","unstructured":"Ding, Z., Hui, T., Huang, S., Liu, S., Luo, X., Huang, J., and Wei, X. (2021). Progressive multimodal interaction network for referring video object segmentation. 3rd Large-Scale Video Object Segm. Chall., 8.","DOI":"10.1109\/CVPR52688.2022.00491"}],"container-title":["Sensors"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1424-8220\/24\/16\/5375\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,10]],"date-time":"2025-10-10T15:39:42Z","timestamp":1760110782000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1424-8220\/24\/16\/5375"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,8,20]]},"references-count":46,"journal-issue":{"issue":"16","published-online":{"date-parts":[[2024,8]]}},"alternative-id":["s24165375"],"URL":"https:\/\/doi.org\/10.3390\/s24165375","relation":{},"ISSN":["1424-8220"],"issn-type":[{"type":"electronic","value":"1424-8220"}],"subject":[],"published":{"date-parts":[[2024,8,20]]}}}