{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,16]],"date-time":"2026-06-16T09:54:23Z","timestamp":1781603663327,"version":"3.54.5"},"reference-count":46,"publisher":"MDPI AG","issue":"6","license":[{"start":{"date-parts":[[2026,6,12]],"date-time":"2026-06-12T00:00:00Z","timestamp":1781222400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Entropy"],"abstract":"<jats:p>Vision Transformers (ViTs) have achieved strong performance in computer vision, but their adversarial robustness remains underexplored. Existing ViT-oriented attacks mainly rely on iterative optimization, leading to high generation cost and limited transferability. Moreover, most generative attacks target a single class, making them inefficient for multi-target scenarios. To address these issues, we propose Attention-Guided Multi-Target Generative Adversarial Attack (AMGAA). AMGAA leverages ViT self-attention to guide target feature fusion and adaptively selects important source patches for perturbation generation. It jointly optimizes adversarial, attention constraint, and total variation losses to improve targeted attack success while preserving visual naturalness. Experiments on CIFAR-10 and ImageNet show that AMGAA achieves average attack success rates (ASRs) of 43.2% and 39.0% in ImageNet single-target transfer attacks and CIFAR-10 multi-target attacks, respectively. Compared with the generative attack methods evaluated in our experiments, AMGAA improves ASR by 5.7 percentage points in the multi-target setting and by 8.1 percentage points in unknown-class generalization. AMGAA also obtains a low LPIPS of 0.018, indicating good visual imperceptibility. Ablation studies confirm the effectiveness of its key components.<\/jats:p>","DOI":"10.3390\/e28060680","type":"journal-article","created":{"date-parts":[[2026,6,16]],"date-time":"2026-06-16T09:06:52Z","timestamp":1781600812000},"page":"680","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":0,"title":["AMGAA: Attention-Guided Multi-Target Generative Adversarial Attack for Vision Transformers"],"prefix":"10.3390","volume":"28","author":[{"ORCID":"https:\/\/orcid.org\/0009-0008-5447-4734","authenticated-orcid":false,"given":"Dongbo","family":"Ou","sequence":"first","affiliation":[{"name":"School of Computer Science and Engineering, Jishou University, Jishou 416000, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Jintian","family":"Lu","sequence":"additional","affiliation":[{"name":"School of Computer Science and Engineering, Jishou University, Jishou 416000, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Shihui","family":"Zhou","sequence":"additional","affiliation":[{"name":"School of Computer Science and Engineering, Jishou University, Jishou 416000, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Ying","family":"Zeng","sequence":"additional","affiliation":[{"name":"School of Computer Science and Engineering, Jishou University, Jishou 416000, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Dongwan","family":"Liao","sequence":"additional","affiliation":[{"name":"School of Computer Science and Engineering, Jishou University, Jishou 416000, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Haoyin","family":"Liu","sequence":"additional","affiliation":[{"name":"School of Computer Science and Engineering, Jishou University, Jishou 416000, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Chao","family":"Yang","sequence":"additional","affiliation":[{"name":"School of Computer Science and Engineering, Jishou University, Jishou 416000, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Yingsheng","family":"He","sequence":"additional","affiliation":[{"name":"School of Computer Science and Engineering, Jishou University, Jishou 416000, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0009-8605-8767","authenticated-orcid":false,"given":"Jie","family":"Tian","sequence":"additional","affiliation":[{"name":"School of Computer Science and Engineering, Jishou University, Jishou 416000, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1968","published-online":{"date-parts":[[2026,6,12]]},"reference":[{"key":"ref_1","unstructured":"Pereira, F., Burges, C., Bottou, L., and Weinberger, K. (2012). ImageNet Classification with Deep Convolutional Neural Networks. Proceedings of the Advances in Neural Information Processing Systems, Curran Associates, Inc."},{"key":"ref_2","first-page":"40030","article-title":"Analyzing Vision Transformers for Image Classification in Class Embedding Space","volume":"Volume 36","author":"Oh","year":"2023","journal-title":"Proceedings of the Advances in Neural Information Processing Systems"},{"key":"ref_3","doi-asserted-by":"crossref","first-page":"13244","DOI":"10.1038\/s41598-024-63818-x","article-title":"A deep image classification model based on prior feature knowledge embedding and application in medical diagnosis","volume":"14","author":"Xu","year":"2024","journal-title":"Sci. Rep."},{"key":"ref_4","doi-asserted-by":"crossref","first-page":"3534","DOI":"10.1109\/TNNLS.2024.3351674","article-title":"Novel snap-layer MMPC scheme via neural dynamics equivalency and solver for redundant robot arms with five-layer physical limits","volume":"36","author":"Tang","year":"2024","journal-title":"IEEE Trans. Neural Netw. Learn. Syst."},{"key":"ref_5","doi-asserted-by":"crossref","first-page":"11189","DOI":"10.1109\/TNNLS.2024.3466296","article-title":"Neural dynamic fault-tolerant scheme for collaborative motion planning of dual-redundant robot manipulators","volume":"36","author":"Zhang","year":"2024","journal-title":"IEEE Trans. Neural Netw. Learn. Syst."},{"key":"ref_6","doi-asserted-by":"crossref","unstructured":"Xiao, L., Zhang, Y., Liao, B., Zhang, Z., Ding, L., and Jin, L. (2017). A velocity-level bi-criteria optimization scheme for coordinated path tracking of dual robot manipulators using recurrent neural network. Front. Neurorobot., 11.","DOI":"10.3389\/fnbot.2017.00047"},{"key":"ref_7","doi-asserted-by":"crossref","first-page":"106339","DOI":"10.1016\/j.neunet.2024.106339","article-title":"DCDLN: A densely connected convolutional dynamic learning network for malaria disease diagnosis","volume":"176","author":"Zhang","year":"2024","journal-title":"Neural Netw."},{"key":"ref_8","first-page":"2255","article-title":"A systematic review of YOLO-based object detection in medical imaging: Advances, challenges, and future directions","volume":"85","author":"Cai","year":"2025","journal-title":"Comput. Mater. Contin."},{"key":"ref_9","doi-asserted-by":"crossref","first-page":"2798","DOI":"10.1109\/JBHI.2020.3019505","article-title":"Adaptive feature selection guided deep forest for COVID-19 classification with chest CT","volume":"24","author":"Sun","year":"2020","journal-title":"IEEE J. Biomed. Health Inform."},{"key":"ref_10","unstructured":"Szegedy, C., Zaremba, W., Sutskever, I., Bruna, J., Erhan, D., Goodfellow, I., and Fergus, R. (2013). Intriguing properties of neural networks. arXiv."},{"key":"ref_11","unstructured":"Goodfellow, I.J., Shlens, J., and Szegedy, C. (2014). Explaining and harnessing adversarial examples. arXiv."},{"key":"ref_12","unstructured":"Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., and Gelly, S. (2020). An image is worth 16x16 words: Transformers for image recognition at scale. arXiv."},{"key":"ref_13","doi-asserted-by":"crossref","first-page":"2278","DOI":"10.1109\/5.726791","article-title":"Gradient-based learning applied to document recognition","volume":"86","author":"LeCun","year":"2002","journal-title":"Proc. IEEE"},{"key":"ref_14","unstructured":"Simonyan, K., and Zisserman, A. (2014). Very deep convolutional networks for large-scale image recognition. arXiv."},{"key":"ref_15","doi-asserted-by":"crossref","unstructured":"He, K., Zhang, X., Ren, S., and Sun, J. (2016). Deep residual learning for image recognition. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, IEEE.","DOI":"10.1109\/CVPR.2016.90"},{"key":"ref_16","doi-asserted-by":"crossref","unstructured":"Mahmood, K., Mahmood, R., and Van Dijk, M. (2021). On the robustness of vision transformers to adversarial examples. Proceedings of the IEEE\/CVF International Conference on Computer Vision, IEEE.","DOI":"10.1109\/ICCV48922.2021.00774"},{"key":"ref_17","unstructured":"Shao, R., Shi, Z., Yi, J., Chen, P.Y., and Hsieh, C.J. (2021). On the adversarial robustness of vision transformers. arXiv."},{"key":"ref_18","unstructured":"Fu, Y., Zhang, S., Wu, S., Wan, C., and Lin, Y. (2022). Patch-fool: Are vision transformers always robust against adversarial perturbations?. arXiv."},{"key":"ref_19","unstructured":"Joshi, A., Akula, S.C., Jagatap, G., and Hegde, C. (2026, June 09). A Few Adversarial Tokens Can Break Vision Transformers. Available online: https:\/\/robustart.github.io\/long_paper\/31.pdf."},{"key":"ref_20","unstructured":"Joshi, A., Jagatap, G., and Hegde, C. (2021). Adversarial token attacks on vision transformers. arXiv."},{"key":"ref_21","unstructured":"Shao, M. (2023). Random position adversarial patch for vision transformers. arXiv."},{"key":"ref_22","doi-asserted-by":"crossref","unstructured":"Carlini, N., and Wagner, D. (2017). Towards evaluating the robustness of neural networks. Proceedings of the 2017 IEEE Symposium on Security and Privacy (sp), IEEE.","DOI":"10.1109\/SP.2017.49"},{"key":"ref_23","unstructured":"Madry, A., Makelov, A., Schmidt, L., Tsipras, D., and Vladu, A. (2017). Towards deep learning models resistant to adversarial attacks. arXiv."},{"key":"ref_24","doi-asserted-by":"crossref","unstructured":"Dong, Y., Liao, F., Pang, T., Su, H., Zhu, J., Hu, X., and Li, J. (2018). Boosting adversarial attacks with momentum. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, IEEE.","DOI":"10.1109\/CVPR.2018.00957"},{"key":"ref_25","unstructured":"Croce, F., and Hein, M. (2020, January 12\u201318). Reliable evaluation of adversarial robustness with an ensemble of diverse parameter-free attacks. Proceedings of the International Conference on Machine Learning, Vienna, Austria."},{"key":"ref_26","first-page":"23296","article-title":"Intriguing properties of vision transformers","volume":"34","author":"Naseer","year":"2021","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_27","doi-asserted-by":"crossref","unstructured":"Gu, J., Tresp, V., and Qin, Y. (2022). Are vision transformers robust to patch perturbations?. Proceedings of the European Conference on Computer Vision, Springer.","DOI":"10.1007\/978-3-031-19775-8_24"},{"key":"ref_28","doi-asserted-by":"crossref","unstructured":"Wu, W., Su, Y., Chen, X., Zhao, S., King, I., Lyu, M.R., and Tai, Y.W. (2020). Boosting the transferability of adversarial samples via attention. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, IEEE.","DOI":"10.1109\/CVPR42600.2020.00124"},{"key":"ref_29","doi-asserted-by":"crossref","first-page":"7526","DOI":"10.1002\/int.22892","article-title":"Attention-guided black-box adversarial attacks with large-scale multiobjective evolutionary optimization","volume":"37","author":"Wang","year":"2022","journal-title":"Int. J. Intell. Syst."},{"key":"ref_30","first-page":"1","article-title":"Attend and attack: Attention guided adversarial attacks on visual question answering models","volume":"Volume 2","author":"Sharma","year":"2018","journal-title":"Proceedings of the Proc. Conf. Neural Inf. Process. Syst. Workshop Secur. Mach. Learn"},{"key":"ref_31","doi-asserted-by":"crossref","first-page":"5663","DOI":"10.1109\/TPAMI.2024.3367773","article-title":"Adaptive perturbation for adversarial attack","volume":"46","author":"Yuan","year":"2024","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_32","doi-asserted-by":"crossref","unstructured":"Duan, J., Qiu, L., He, G., Zhao, L., Zhang, Z., and Li, H. (2024). A region-adaptive local perturbation-based method for generating adversarial examples in synthetic aperture radar object detection. Remote Sens., 16.","DOI":"10.3390\/rs16060997"},{"key":"ref_33","doi-asserted-by":"crossref","unstructured":"Liu, J., Zhang, C., and Lyu, X. (2025). Boosting the transferability of adversarial examples via local mixup and adaptive step size. Proceedings of the ICASSP 2025\u20142025 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), IEEE.","DOI":"10.1109\/ICASSP49660.2025.10889335"},{"key":"ref_34","doi-asserted-by":"crossref","first-page":"116","DOI":"10.1007\/s40747-024-01748-x","article-title":"APDL: An adaptive step size method for white-box adversarial attacks","volume":"11","author":"Hu","year":"2025","journal-title":"Complex Intell. Syst."},{"key":"ref_35","doi-asserted-by":"crossref","unstructured":"Xiao, C., Li, B., Zhu, J.Y., He, W., Liu, M., and Song, D. (2018). Generating adversarial examples with adversarial networks. arXiv.","DOI":"10.24963\/ijcai.2018\/543"},{"key":"ref_36","doi-asserted-by":"crossref","unstructured":"Naseer, M., Khan, S., Hayat, M., Khan, F.S., and Porikli, F. (2021). On generating transferable targeted perturbations. Proceedings of the IEEE\/CVF International Conference on Computer Vision, IEEE.","DOI":"10.1109\/ICCV48922.2021.00761"},{"key":"ref_37","doi-asserted-by":"crossref","unstructured":"Han, J., Dong, X., Zhang, R., Chen, D., Zhang, W., Yu, N., Luo, P., and Wang, X. (2019). Once a man: Towards multi-target attack via learning multi-target adversarial network once. Proceedings of the IEEE\/CVF International Conference on Computer Vision, IEEE.","DOI":"10.1109\/ICCV.2019.00526"},{"key":"ref_38","doi-asserted-by":"crossref","unstructured":"Fang, H., Kong, J., Chen, B., Dai, T., Wu, H., and Xia, S.T. (2024). Clip-guided generative networks for transferable targeted adversarial attacks. Proceedings of the European Conference on Computer Vision, Springer.","DOI":"10.1007\/978-3-031-73390-1_1"},{"key":"ref_39","unstructured":"Krizhevsky, A. (2026, June 09). Learning Multiple Layers of Features from Tiny Images. Available online: https:\/\/www.cs.toronto.edu\/~kriz\/learning-features-2009-TR.pdf."},{"key":"ref_40","doi-asserted-by":"crossref","unstructured":"Deng, J., Dong, W., Socher, R., Li, L.J., Li, K., and Fei-Fei, L. (2009). Imagenet: A large-scale hierarchical image database. Proceedings of the 2009 IEEE Conference on Computer Vision and Pattern Recognition, IEEE.","DOI":"10.1109\/CVPR.2009.5206848"},{"key":"ref_41","doi-asserted-by":"crossref","unstructured":"Liu, Z., Lin, Y., Cao, Y., Hu, H., Wei, Y., Zhang, Z., Lin, S., and Guo, B. (2021). Swin transformer: Hierarchical vision transformer using shifted windows. Proceedings of the IEEE\/CVF International Conference on Computer Vision, IEEE.","DOI":"10.1109\/ICCV48922.2021.00986"},{"key":"ref_42","doi-asserted-by":"crossref","unstructured":"Mao, X., Qi, G., Chen, Y., Li, X., Duan, R., Ye, S., He, Y., and Xue, H. (2022). Towards robust vision transformer. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, IEEE.","DOI":"10.1109\/CVPR52688.2022.01173"},{"key":"ref_43","doi-asserted-by":"crossref","unstructured":"Huang, G., Liu, Z., Van Der Maaten, L., and Weinberger, K.Q. (2017). Densely connected convolutional networks. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, IEEE.","DOI":"10.1109\/CVPR.2017.243"},{"key":"ref_44","unstructured":"Loshchilov, I., and Hutter, F. (2017). Decoupled weight decay regularization. arXiv."},{"key":"ref_45","unstructured":"Kingma, D.P., and Ba, J. (2014). Adam: A method for stochastic optimization. arXiv."},{"key":"ref_46","doi-asserted-by":"crossref","unstructured":"Kim, K., Wu, B., Dai, X., Zhang, P., Yan, Z., Vajda, P., and Kim, S.J. (2021). Rethinking the Self-Attention in Vision Transformers. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR) Workshops, IEEE.","DOI":"10.1109\/CVPRW53098.2021.00342"}],"container-title":["Entropy"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1099-4300\/28\/6\/680\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,6,16]],"date-time":"2026-06-16T09:28:50Z","timestamp":1781602130000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1099-4300\/28\/6\/680"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,6,12]]},"references-count":46,"journal-issue":{"issue":"6","published-online":{"date-parts":[[2026,6]]}},"alternative-id":["e28060680"],"URL":"https:\/\/doi.org\/10.3390\/e28060680","relation":{},"ISSN":["1099-4300"],"issn-type":[{"value":"1099-4300","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,6,12]]}}}