{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,4,22]],"date-time":"2026-04-22T19:24:00Z","timestamp":1776885840244,"version":"3.51.2"},"reference-count":70,"publisher":"Association for Computing Machinery (ACM)","issue":"6","license":[{"start":{"date-parts":[[2023,11,14]],"date-time":"2023-11-14T00:00:00Z","timestamp":1699920000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"crossref","award":["62172417, 62272461, 62276266, 62106268"],"award-info":[{"award-number":["62172417, 62272461, 62276266, 62106268"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]},{"DOI":"10.13039\/501100004608","name":"Natural Science Foundation of Jiangsu Province","doi-asserted-by":"crossref","award":["BK20201346"],"award-info":[{"award-number":["BK20201346"]}],"id":[{"id":"10.13039\/501100004608","id-type":"DOI","asserted-by":"crossref"}]},{"name":"Xuzhou Key Research and Development Program","award":["KC22287"],"award-info":[{"award-number":["KC22287"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Intell. Syst. Technol."],"published-print":{"date-parts":[[2023,12,31]]},"abstract":"<jats:p>\n            Video Object Segmentation (VOS) methods have made many breakthroughs with the help of the continuous development and advancement of deep learning. However, the deep learning model is vulnerable to malicious adversarial attacks, which mislead the model to make wrong decisions by adding adversarial perturbation that humans cannot perceive to the input image. Threats to deep learning models remind us that video object segmentation methods are also vulnerable to attacks, thereby threatening their security. Therefore, we study adversarial attacks on the VOS task to better identify the vulnerabilities of the VOS method, which in turn provides an opportunity to improve its robustness. In this paper, we propose an attention-guided adversarial attack method, which uses spatial attention blocks to capture features with global dependencies to construct correlations between consecutive video frames, and performs multipath aggregation to effectively integrate spatial-temporal perturbation, thereby guiding the deconvolution network to generate adversarial examples with strong attack capability. Specifically, the class loss function is designed to enable the deconvolution network to better activate noise in other regions and suppress the activation related to the object class based on the enhanced feature map of the object class. At the same time, attentional feature loss is designed to enhance the transferability against attack. The experimental results on the DAVIS dataset show that the proposed attention-guided adversarial attack method can significantly reduce the segmentation accuracy of OSVOS, and the\n            <jats:italic>J<\/jats:italic>\n            &amp;\n            <jats:italic>F<\/jats:italic>\n            mean on DAVIS 2016 can reach 73.6% drop rate. The generated adversarial examples are also highly transferable to other video object segmentation models.\n          <\/jats:p>","DOI":"10.1145\/3617067","type":"journal-article","created":{"date-parts":[[2023,9,2]],"date-time":"2023-09-02T11:31:20Z","timestamp":1693654280000},"page":"1-22","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":2,"title":["Attention-guided Adversarial Attack for Video Object Segmentation"],"prefix":"10.1145","volume":"14","author":[{"given":"Rui","family":"Yao","sequence":"first","affiliation":[{"name":"School of Computer Science and Technology, China University of Mining and Technology, Engineering Research Center of Mine Digitization, Ministry of Education of the Peoples Republic of China, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Ying","family":"Chen","sequence":"additional","affiliation":[{"name":"School of Computer Science and Technology, China University of Mining and Technology, Engineering Research Center of Mine Digitization, Ministry of Education of the Peoples Republic of China, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yong","family":"Zhou","sequence":"additional","affiliation":[{"name":"School of Computer Science and Technology, China University of Mining and Technology, Engineering Research Center of Mine Digitization, Ministry of Education of the Peoples Republic of China, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Fuyuan","family":"Hu","sequence":"additional","affiliation":[{"name":"School of Electronic and Information Engineering, Suzhou University of Science and Technology, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Jiaqi","family":"Zhao","sequence":"additional","affiliation":[{"name":"School of Computer Science and Technology, China University of Mining and Technology, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Bing","family":"Liu","sequence":"additional","affiliation":[{"name":"School of Computer Science and Technology, China University of Mining and Technology, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Zhiwen","family":"Shao","sequence":"additional","affiliation":[{"name":"School of Computer Science and Technology, China University of Mining and Technology, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2023,11,14]]},"reference":[{"key":"e_1_3_1_2_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.565"},{"key":"e_1_3_1_3_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-58452-8_13"},{"key":"e_1_3_1_4_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.00940"},{"key":"e_1_3_1_5_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00130"},{"key":"e_1_3_1_6_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00774"},{"key":"e_1_3_1_7_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00957"},{"key":"e_1_3_1_8_2","doi-asserted-by":"publisher","DOI":"10.1145\/3474085.3475693"},{"key":"e_1_3_1_9_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v33i01.33018287"},{"key":"e_1_3_1_10_2","article-title":"Explaining and harnessing adversarial examples","author":"Goodfellow I. J.","year":"2014","unstructured":"I. J. Goodfellow, J. Shlens, and C. Szegedy. 2014. Explaining and harnessing adversarial examples. Computer Science (2014).","journal-title":"Computer Science"},{"key":"e_1_3_1_11_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV48922.2021.01066"},{"key":"e_1_3_1_12_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-58595-2_13"},{"key":"e_1_3_1_13_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00946"},{"key":"e_1_3_1_14_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00745"},{"key":"e_1_3_1_15_2","doi-asserted-by":"publisher","DOI":"10.1109\/WACV45572.2020.9093333"},{"key":"e_1_3_1_16_2","first-page":"325","article-title":"MaskRNN: Instance level video object segmentation","volume":"2017","author":"Hu Y. T.","year":"2018","unstructured":"Y. T. Hu, J. B. Huang, and A. G. Schwing. 2018. MaskRNN: Instance level video object segmentation. Advances in Neural Information Processing Systems 2017-December (2018), 325\u2013334.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_1_17_2","doi-asserted-by":"publisher","DOI":"10.1109\/JAS.2021.1004210"},{"key":"e_1_3_1_18_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.00890"},{"key":"e_1_3_1_19_2","article-title":"Adversarial attacks for optical flow-based action recognition classifiers","author":"Inkawhich Nathan","year":"2018","unstructured":"Nathan Inkawhich, Matthew Inkawhich, Yiran Chen, and Hai Li. 2018. Adversarial attacks for optical flow-based action recognition classifiers. arXiv preprint arXiv:1811.11875 (2018).","journal-title":"arXiv preprint arXiv:1811.11875"},{"key":"e_1_3_1_20_2","article-title":"Batch Normalization: Accelerating deep network training by reducing internal covariate shift","author":"Ioffe Sergey","year":"2014","unstructured":"Sergey Ioffe and Christian Szegedy. 2014. Batch Normalization: Accelerating deep network training by reducing internal covariate shift. arXiv preprint arXiv:1502.03167 (2014).","journal-title":"arXiv preprint arXiv:1502.03167"},{"key":"e_1_3_1_21_2","first-page":"19545","article-title":"Space-time correspondence as a contrastive random walk","volume":"33","author":"Jabri Allan","year":"2020","unstructured":"Allan Jabri, Andrew Owens, and Alexei Efros. 2020. Space-time correspondence as a contrastive random walk. Advances in Neural Information Processing Systems 33 (2020), 19545\u201319560.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_1_22_2","article-title":"Spatial transformer networks","volume":"28","author":"Jaderberg Max","year":"2015","unstructured":"Max Jaderberg, Karen Simonyan, Andrew Zisserman, and Koray Kavukcuoglu. 2015. Spatial transformer networks. Advances in Neural Information Processing Systems 28 (2015).","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_1_23_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.336"},{"key":"e_1_3_1_24_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR46437.2021.00664"},{"key":"e_1_3_1_25_2","doi-asserted-by":"publisher","DOI":"10.1145\/3343031.3351088"},{"key":"e_1_3_1_26_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.00916"},{"key":"e_1_3_1_27_2","doi-asserted-by":"publisher","DOI":"10.1007\/s11263-019-01164-6"},{"key":"e_1_3_1_28_2","volume-title":"2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR\u201917)","author":"Khoreva A.","year":"2017","unstructured":"A. Khoreva, F. Perazzi, R. Benenson, B. Schiele, and A. Sorkine-Hornung. 2017. Learning video object segmentation from static images. In 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR\u201917)."},{"key":"e_1_3_1_29_2","article-title":"ImageNet classification with deep convolutional neural networks","volume":"25","author":"Krizhevsky Alex","year":"2012","unstructured":"Alex Krizhevsky, Ilya Sutskever, and Geoffrey E. Hinton. 2012. ImageNet classification with deep convolutional neural networks. Advances in Neural Information Processing Systems 25 (2012).","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_1_30_2","doi-asserted-by":"publisher","DOI":"10.1145\/3394171.3413531"},{"key":"e_1_3_1_31_2","article-title":"Adversarial examples in the physical world","volume":"1607","author":"Kurakin A.","year":"2016","unstructured":"A. Kurakin, I. Goodfellow, and S. Bengio. 2016. Adversarial examples in the physical world. ArXiv abs\/1607.02533 (2016).","journal-title":"ArXiv"},{"key":"e_1_3_1_32_2","doi-asserted-by":"crossref","unstructured":"X. Li and C. C. Loy. 2018. Video object segmentation with joint re-identification and attention-aware mask propagation. (2018).","DOI":"10.1007\/978-3-030-01219-9_6"},{"key":"e_1_3_1_33_2","article-title":"Robust adversarial perturbation on deep proposal-based models","volume":"1809","author":"Li Y.","year":"2018","unstructured":"Y. Li, D. Tian, Mingching-Chang, X. Bian, and S. Lyu. 2018. Robust adversarial perturbation on deep proposal-based models. ArXiv abs\/1809.05962 (2018).","journal-title":"ArXiv"},{"key":"e_1_3_1_34_2","doi-asserted-by":"crossref","first-page":"673","DOI":"10.1109\/SP.2019.00023","volume-title":"2019 IEEE Symposium on Security and Privacy (SP\u201919)","author":"Ling Xiang","year":"2019","unstructured":"Xiang Ling, Shouling Ji, Jiaxu Zou, Jiannan Wang, Chunming Wu, Bo Li, and Ting Wang. 2019. DEEPSEC: A uniform platform for security analysis of deep learning model. In 2019 IEEE Symposium on Security and Privacy (SP\u201919). IEEE, 673\u2013690."},{"key":"e_1_3_1_35_2","article-title":"Attention-guided global-local adversarial learning for detail-preserving multi-exposure image fusion","author":"Liu Jinyuan","year":"2022","unstructured":"Jinyuan Liu, Jingjie Shang, Risheng Liu, and Xin Fan. 2022. Attention-guided global-local adversarial learning for detail-preserving multi-exposure image fusion. IEEE Transactions on Circuits and Systems for Video Technology (2022).","journal-title":"IEEE Transactions on Circuits and Systems for Video Technology"},{"key":"e_1_3_1_36_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2015.7298965"},{"key":"e_1_3_1_37_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.00374"},{"key":"e_1_3_1_38_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.00898"},{"key":"e_1_3_1_39_2","article-title":"Towards deep learning models resistant to adversarial attacks","volume":"1706","author":"Madry A.","year":"2017","unstructured":"A. Madry, A. Makelov, L. Schmidt, D. Tsipras, and A. Vladu. 2017. Towards deep learning models resistant to adversarial attacks. ArXiv abs\/1706.06083 (2017).","journal-title":"ArXiv"},{"key":"e_1_3_1_40_2","volume-title":"Advances in Neural Information Processing Systems","author":"Mnih Volodymyr","year":"2014","unstructured":"Volodymyr Mnih, Nicolas Heess, Alex Graves, and Koray Kavukcuoglu. 2014. Recurrent models of visual attention. In Advances in Neural Information Processing Systems, Z. Ghahramani, M. Welling, C. Cortes, N. Lawrence, and K. Q. Weinberger (Eds.), Vol. 27. Curran Associates, Inc.https:\/\/proceedings.neurips.cc\/paper\/2014\/file\/09c6c3783b4a70054da74f2538ed47c6-Paper.pdf"},{"key":"e_1_3_1_41_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.17"},{"key":"e_1_3_1_42_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-58558-7_36"},{"key":"e_1_3_1_43_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00770"},{"key":"e_1_3_1_44_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2019.00932"},{"key":"e_1_3_1_45_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR46437.2021.00830"},{"key":"e_1_3_1_46_2","volume-title":"Advances in Neural Information Processing Systems","author":"Paszke A.","year":"2019","unstructured":"A. Paszke, S. Gross, F. Massa, A. Lerer, and S. Chintala. 2019. PyTorch: An imperative style, high-performance deep learning library. In Advances in Neural Information Processing Systems, Vol. 32. https:\/\/proceedings.neurips.cc\/paper\/2019\/file\/bdbca288fee7f92f2bfa9f7012727740-Paper.pdf"},{"key":"e_1_3_1_47_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.85"},{"key":"e_1_3_1_48_2","article-title":"The 2017 DAVIS challenge on video object segmentation","volume":"1704","author":"Pont-Tuset Jordi","year":"2017","unstructured":"Jordi Pont-Tuset, Federico Perazzi, Sergi Caelles, Pablo Arbel\u00e1ez, Alexander Sorkine-Hornung, and Luc Van Gool. 2017. The 2017 DAVIS challenge on video object segmentation. ArXiv abs\/1704.00675 (2017).","journal-title":"ArXiv"},{"key":"e_1_3_1_49_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.00743"},{"key":"e_1_3_1_50_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-58542-6_38"},{"key":"e_1_3_1_51_2","article-title":"Intriguing properties of neural networks","author":"Szegedy C.","year":"2013","unstructured":"C. Szegedy, W. Zaremba, I. Sutskever, J. Bruna, D Erhan, I. Goodfellow, and R. Fergus. 2013. Intriguing properties of neural networks. Computer Science (2013).","journal-title":"Computer Science"},{"key":"e_1_3_1_52_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.00971"},{"key":"e_1_3_1_53_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-01261-8_24"},{"key":"e_1_3_1_54_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.00835"},{"key":"e_1_3_1_55_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.00142"},{"key":"e_1_3_1_56_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2019.00933"},{"key":"e_1_3_1_57_2","article-title":"Transferable adversarial attacks for image and video object detection","author":"Wei Xingxing","year":"2018","unstructured":"Xingxing Wei, Siyuan Liang, Ning Chen, and Xiaochun Cao. 2018. Transferable adversarial attacks for image and video object detection. arXiv preprint arXiv:1811.12641 (2018).","journal-title":"arXiv preprint arXiv:1811.12641"},{"key":"e_1_3_1_58_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v33i01.33018973"},{"key":"e_1_3_1_59_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v34i07.6918"},{"key":"e_1_3_1_60_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v36i3.20168"},{"key":"e_1_3_1_61_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2019.00492"},{"key":"e_1_3_1_62_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2017.153"},{"key":"e_1_3_1_63_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2015.164"},{"key":"e_1_3_1_64_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00680"},{"key":"e_1_3_1_65_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-58558-7_20"},{"key":"e_1_3_1_66_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2017.238"},{"key":"e_1_3_1_67_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2019.00403"},{"key":"e_1_3_1_68_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2018.2881114"},{"key":"e_1_3_1_69_2","doi-asserted-by":"publisher","DOI":"10.1109\/ACCESS.2021.3110798"},{"key":"e_1_3_1_70_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.00698"},{"key":"e_1_3_1_71_2","article-title":"Seeing isn\u2019t believing: Practical adversarial attack against object detectors","volume":"1812","author":"Zhao Y.","year":"2018","unstructured":"Y. Zhao, H. Zhu, R. Liang, Q. Shen, S. Zhang, and K. Chen. 2018. Seeing isn\u2019t believing: Practical adversarial attack against object detectors. ArXiv abs\/1812.10217 (2018).","journal-title":"ArXiv"}],"container-title":["ACM Transactions on Intelligent Systems and Technology"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3617067","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3617067","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T16:36:06Z","timestamp":1750178166000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3617067"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,11,14]]},"references-count":70,"journal-issue":{"issue":"6","published-print":{"date-parts":[[2023,12,31]]}},"alternative-id":["10.1145\/3617067"],"URL":"https:\/\/doi.org\/10.1145\/3617067","relation":{},"ISSN":["2157-6904","2157-6912"],"issn-type":[{"value":"2157-6904","type":"print"},{"value":"2157-6912","type":"electronic"}],"subject":[],"published":{"date-parts":[[2023,11,14]]},"assertion":[{"value":"2022-01-12","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2023-08-11","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2023-11-14","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}