{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,15]],"date-time":"2026-06-15T14:37:24Z","timestamp":1781534244779,"version":"3.54.5"},"reference-count":35,"publisher":"MDPI AG","issue":"5","license":[{"start":{"date-parts":[[2026,5,4]],"date-time":"2026-05-04T00:00:00Z","timestamp":1777852800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"name":"Zhejiang University of Water Resource and Electric Power","award":["88106323038"],"award-info":[{"award-number":["88106323038"]}]},{"award":["88106323038"],"award-info":[{"award-number":["88106323038"]}],"id":[{"id":"https:\/\/ror.org\/04dg5b632","id-type":"ROR","asserted-by":"publisher"}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Entropy"],"abstract":"<jats:p>Multimodal models such as CLIP and ALBEF essentially maximize cross-modal mutual information to align heterogeneous modalities, utilizing semantic consistency as an implicit prior. However, this alignment mechanism creates a structural vulnerability: the models rely heavily on invariant information coupling. In this work, we investigate this vulnerability and propose a symmetry-driven adversarial attack framework. Unlike standard methods that inject high-entropy unstructured noise, our approach designs collaborative perturbations by modeling semantic-consistent mappings between geometric image transformations and syntactic text variations. By explicitly exploiting the information redundancy inherent in cross-modal symmetries, our method effectively reduces the entropy of the adversarial search space. This reveals a fundamental trade-off between information invariance and robustness, achieving state-of-the-art attack success rates with imperceptible perturbations.<\/jats:p>","DOI":"10.3390\/e28050521","type":"journal-article","created":{"date-parts":[[2026,5,5]],"date-time":"2026-05-05T07:57:29Z","timestamp":1777967849000},"page":"521","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":0,"title":["Symmetry-Driven Multimodal Adversarial Attacks: An Information-Theoretic Perspective on Cross-Modal Invariance and Robustness"],"prefix":"10.3390","volume":"28","author":[{"given":"Jin","family":"Wei","sequence":"first","affiliation":[{"name":"School of Computer Science and Technology, Zhejiang University of Water Resources and Electric Power, Hangzhou 310018, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Xinyuan","family":"Wang","sequence":"additional","affiliation":[{"name":"Department of Accounting and Management Engineering, Hebei Institute of Mechanical and Electrical Technology, Xingtai 054000, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Liam","family":"Xu","sequence":"additional","affiliation":[{"name":"Institute of Cyberspace Technology, Hong Kong College of Technology, Hong Kong, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Yunfei","family":"Li","sequence":"additional","affiliation":[{"name":"School of Computer Science and Technology, Zhejiang University of Water Resources and Electric Power, Hangzhou 310018, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1968","published-online":{"date-parts":[[2026,5,4]]},"reference":[{"key":"ref_1","unstructured":"Radford, A., Kim, J.W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., and Clark, J. (2021, January 18\u201324). Learning Transferable Visual Models from Natural Language Supervision. Proceedings of the 38th International Conference on Machine Learning (ICML), PMLR, Virtual."},{"key":"ref_2","first-page":"9694","article-title":"Align before Fuse: Vision and Language Representation Learning with Momentum Distillation","volume":"Volume 34","author":"Li","year":"2021","journal-title":"Proceedings of the Advances in Neural Information Processing Systems"},{"key":"ref_3","doi-asserted-by":"crossref","first-page":"1341","DOI":"10.1109\/TITS.2020.2972974","article-title":"Deep Multi-Modal Object Detection and Semantic Segmentation for Autonomous Driving","volume":"22","author":"Feng","year":"2021","journal-title":"IEEE Trans. Intell. Transp. Syst."},{"key":"ref_4","doi-asserted-by":"crossref","unstructured":"Cao, H., Xu, Y., Yang, J., Yin, P., Yuan, S., and Lihua, X. (2023). Multi-Modal Continual Test-Time Adaptation for 3D Semantic Segmentation. arXiv.","DOI":"10.1109\/ICCV51070.2023.01724"},{"key":"ref_5","doi-asserted-by":"crossref","first-page":"1773","DOI":"10.1038\/s41591-022-01981-2","article-title":"Multimodal biomedical AI","volume":"28","author":"Acosta","year":"2022","journal-title":"Nat. Med."},{"key":"ref_6","unstructured":"Shah, D., Shah, A., Aneja, J., Liu, C., and Calandra, R. (2023). From Pixels to Language: A Survey on Multimodal Learning for Robotics. arXiv."},{"key":"ref_7","unstructured":"Ahn, M., Brohan, A., Brown, N., Carbajal, Y., Chebotar, Y., Cortes, O., David, B., Finn, C., Fu, C., and Gopalakrishnan, K. (2022, January 14\u201318). Do As I Can, Not As I Say: Grounding Language in Robotic Affordances. Proceedings of the Conference on Robot Learning (CoRL), Auckland, New Zealand."},{"key":"ref_8","unstructured":"Zhao, Y., Pang, T., Du, C., Yang, X., Li, C., Cheung, N., and Lin, M. (2024). On Evaluating Adversarial Robustness of Large Vision-Language Models. arXiv."},{"key":"ref_9","doi-asserted-by":"crossref","first-page":"151","DOI":"10.1007\/s11633-019-1211-x","article-title":"Adversarial Attacks and Defenses in Images, Graphs and Text: A Review","volume":"17","author":"Xu","year":"2020","journal-title":"Int. J. Autom. Comput."},{"key":"ref_10","doi-asserted-by":"crossref","unstructured":"Eykholt, K., Evtimov, I., Fernandes, E., Li, B., Rahmati, A., Xiao, C., Prakash, A., Kohno, T., and Song, D. (2018, January 18\u201323). Robust Physical-World Attacks on Deep Learning Visual Classification. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00175"},{"key":"ref_11","doi-asserted-by":"crossref","first-page":"1287","DOI":"10.1126\/science.aaw4399","article-title":"Adversarial attacks on medical machine learning","volume":"363","author":"Finlayson","year":"2019","journal-title":"Science"},{"key":"ref_12","doi-asserted-by":"crossref","unstructured":"Zhang, Z., Chen, S., Du, Y., Zhao, H., and Dai, J. (2022). Towards Adversarial Attack on Vision-Language Pre-training Models. Proceedings of the 30th ACM International Conference on Multimedia (MM \u201922), Association for Computing Machinery.","DOI":"10.1145\/3503161.3547801"},{"key":"ref_13","unstructured":"Bai, Y., Zhu, X., He, K., Wang, L., and Wu, Q. (2023). Adversarial Attack on Vision-Language Pre-trained Models. arXiv."},{"key":"ref_14","unstructured":"Yin, Z., Ye, M., Zhang, T., Du, T., Zhu, J., Liu, H., Chen, J., Wang, T., and Ma, F. (2023). VLATTACK: Multimodal Adversarial Attacks on Vision-Language Tasks via Pre-trained Models. Proceedings of the Advances in Neural Information Processing Systems (NeurIPS), Curran Associates, Inc."},{"key":"ref_15","doi-asserted-by":"crossref","unstructured":"Schlarmann, C., and Hein, M. (2023, January 2\u20136). On the Adversarial Robustness of Multi-Modal Foundation Models. Proceedings of the IEEE\/CVF International Conference on Computer Vision (ICCV) Workshops (ICCVW), Paris, France.","DOI":"10.1109\/ICCVW60793.2023.00395"},{"key":"ref_16","doi-asserted-by":"crossref","unstructured":"Weyl, H. (1952). Symmetry, Princeton University Press.","DOI":"10.1515\/9781400874347"},{"key":"ref_17","unstructured":"Bronstein, M.M., Bruna, J., Cohen, T., and Veli\u010dkovi\u0107, P. (2021). Geometric Deep Learning: Grids, Groups, Graphs, Geodesics, and Gauges. arXiv."},{"key":"ref_18","unstructured":"Chen, H., Zhang, Y., Dong, Y., Yang, X., Su, H., and Zhu, J. (2024, January 7\u201311). Rethinking Model Ensemble in Transfer-Based Adversarial Attacks. Proceedings of the International Conference on Learning Representations (ICLR), Vienna, Austria. Available online: https:\/\/proceedings.iclr.cc\/paper_files\/paper\/2024\/hash\/53fe824f289060ce705ed7c01dae59d2-Abstract-Conference.html."},{"key":"ref_19","doi-asserted-by":"crossref","first-page":"103853","DOI":"10.1016\/j.cose.2024.103853","article-title":"Black-box adversarial transferability: An empirical study in cybersecurity perspective","volume":"141","author":"Roshan","year":"2024","journal-title":"Comput. Secur."},{"key":"ref_20","unstructured":"Faghri, F., Fleet, D.J., Kiros, J.R., and Fidler, S. (2018). VSE++: Improving Visual-Semantic Embeddings with Hard Negatives. arXiv."},{"key":"ref_21","unstructured":"Goodfellow, I.J., Shlens, J., and Szegedy, C. (2014). Explaining and Harnessing Adversarial Examples. arXiv."},{"key":"ref_22","doi-asserted-by":"crossref","unstructured":"Singh, A., Hu, R., Goswami, V., Couairon, G., Galuba, W., Rohrbach, M., and Kiela, D. (2022, January 18\u201324). FLAVA: A Foundational Language and Vision Alignment Model. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), New Orleans, LA, USA.","DOI":"10.1109\/CVPR52688.2022.01519"},{"key":"ref_23","doi-asserted-by":"crossref","unstructured":"Lin, T.Y., Maire, M., Belongie, S., Hays, J., Perona, P., Ramanan, D., Doll\u00e1r, P., and Zitnick, C.L. (2014). Microsoft COCO: Common Objects in Context. Proceedings of the European Conference on Computer Vision (ECCV), Springer.","DOI":"10.1007\/978-3-319-10602-1_48"},{"key":"ref_24","doi-asserted-by":"crossref","first-page":"67","DOI":"10.1162\/tacl_a_00166","article-title":"From Image Descriptions to Visual Denotations","volume":"2","author":"Young","year":"2014","journal-title":"Trans. Assoc. Comput. Linguist."},{"key":"ref_25","doi-asserted-by":"crossref","unstructured":"Zhang, R., Isola, P., Efros, A.A., Shechtman, E., and Wang, O. (2018, January 18\u201323). The Unreasonable Effectiveness of Deep Features as a Perceptual Metric. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00068"},{"key":"ref_26","unstructured":"Wu, C.H., Shah, R., Koh, J.Y., Salakhutdinov, R., Fried, D., and Raghunathan, A. (2024). Dissecting Adversarial Robustness of Multimodal LM Agents. arXiv."},{"key":"ref_27","doi-asserted-by":"crossref","unstructured":"Sharif, M., Bhagavatula, S., Bauer, L., and Reiter, M.K. (2016, January 24\u201328). Accessorize to a Crime: Real and Stealthy Attacks on State-of-the-Art Face Recognition. Proceedings of the 2016 ACM SIGSAC Conference on Computer and Communications Security (CCS), Vienna, Austria.","DOI":"10.1145\/2976749.2978392"},{"key":"ref_28","unstructured":"Jin, D., Jin, Z., Wu, J.T., and Yang, Y. (2020, January 16\u201320). BERT-Attack: Adversarial Attack Against BERT Using BERT. Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), Online."},{"key":"ref_29","unstructured":"Cohen, T., and Welling, M. (2016, January 19\u201324). Group Equivariant Convolutional Networks. Proceedings of the 33rd International Conference on Machine Learning (ICML), New York, NY, USA."},{"key":"ref_30","unstructured":"Du, Y., Ding, M., Nie, L., Zhang, H., and Yang, Y. (2021, January 1\u20136). Data Augmentation using Symmetries of Natural Language. Proceedings of the Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021, Online Event."},{"key":"ref_31","doi-asserted-by":"crossref","first-page":"72","DOI":"10.1109\/TDSC.2018.2874243","article-title":"Detecting Adversarial Image Examples in Deep Neural Networks with Adaptive Noise Reduction","volume":"18","author":"Liang","year":"2021","journal-title":"IEEE Trans. Dependable Secur. Comput."},{"key":"ref_32","unstructured":"Madry, A., Makelov, A., Schmidt, L., Tsipras, D., and Vladu, A. (May, January 30). Towards Deep Learning Models Resistant to Adversarial Attacks. Proceedings of the International Conference on Learning Representations (ICLR), Vancouver, BC, Canada. Available online: https:\/\/openreview.net\/forum?id=rJzIBfZAb."},{"key":"ref_33","doi-asserted-by":"crossref","unstructured":"Sun, M., Li, Z., Xiao, C., Qiu, H., Kailkhura, B., Liu, M., and Li, B. (2021, January 10\u201317). Can Shape Structure Features Improve Model Robustness Under Diverse Adversarial Settings?. Proceedings of the IEEE\/CVF International Conference on Computer Vision (ICCV), Montreal, QC, Canada.","DOI":"10.1109\/ICCV48922.2021.00743"},{"key":"ref_34","doi-asserted-by":"crossref","unstructured":"Kumar, D., Kumar, C., Seah, C.W., Xia, S., and Shao, M. (2020). Finding Achilles\u2019 Heel: Adversarial Attack on Multi-modal Action Recognition. Proceedings of the 28th ACM International Conference on Multimedia (MM \u201920), Association for Computing Machinery.","DOI":"10.1145\/3394171.3413531"},{"key":"ref_35","doi-asserted-by":"crossref","unstructured":"Papineni, K., Roukos, S., Ward, T., and Zhu, W.J. (2002, January 7\u201312). BLEU: A Method for Automatic Evaluation of Machine Translation. Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics (ACL), Philadelphia, PA, USA.","DOI":"10.3115\/1073083.1073135"}],"container-title":["Entropy"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1099-4300\/28\/5\/521\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,5,16]],"date-time":"2026-05-16T04:20:50Z","timestamp":1778905250000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1099-4300\/28\/5\/521"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,5,4]]},"references-count":35,"journal-issue":{"issue":"5","published-online":{"date-parts":[[2026,5]]}},"alternative-id":["e28050521"],"URL":"https:\/\/doi.org\/10.3390\/e28050521","relation":{},"ISSN":["1099-4300"],"issn-type":[{"value":"1099-4300","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,5,4]]}}}