{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,12,10]],"date-time":"2025-12-10T04:10:26Z","timestamp":1765339826729,"version":"3.46.0"},"publisher-location":"New York, NY, USA","reference-count":44,"publisher":"ACM","funder":[{"name":"National Key R&D Program of China","award":["2021ZD0112804"],"award-info":[{"award-number":["2021ZD0112804"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2025,10,27]]},"DOI":"10.1145\/3746027.3755387","type":"proceedings-article","created":{"date-parts":[[2025,10,25]],"date-time":"2025-10-25T07:38:54Z","timestamp":1761377934000},"page":"4310-4318","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":0,"title":["FedAPT: Federated Adversarial Prompt Tuning for Vision-Language Models"],"prefix":"10.1145","author":[{"ORCID":"https:\/\/orcid.org\/0009-0003-6886-4630","authenticated-orcid":false,"given":"Kun","family":"Zhai","sequence":"first","affiliation":[{"name":"College of Computer Science and Artificial Intelligence, Fudan University, Shanghai, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-6199-529X","authenticated-orcid":false,"given":"Siheng","family":"Chen","sequence":"additional","affiliation":[{"name":"Shanghai Jiao Tong University, Shanghai, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-2099-4973","authenticated-orcid":false,"given":"Xingjun","family":"Ma","sequence":"additional","affiliation":[{"name":"College of Computer Science and Artificial Intelligence, Fudan University, Shanghai, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-1907-8567","authenticated-orcid":false,"given":"Yu-Gang","family":"Jiang","sequence":"additional","affiliation":[{"name":"College of Intelligent Robotics and Advanced Manufacturing, Fudan University, shanghai, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2025,10,27]]},"reference":[{"key":"e_1_3_2_1_1_1","doi-asserted-by":"crossref","unstructured":"Sikai Bai Jie Zhang Song Guo Shuaicheng Li Jingcai Guo Jun Hou Tao Han and Xiaocheng Lu. 2024. DiPrompT: Disentangled Prompt Tuning for Multiple Latent Domain Generalization in Federated Learning. In CVPR. 27284--27293.","DOI":"10.1109\/CVPR52733.2024.02576"},{"volume-title":"Food-101--mining discriminative components with random forests","author":"Bossard Lukas","key":"e_1_3_2_1_2_1","unstructured":"Lukas Bossard, Matthieu Guillaumin, and Luc Van Gool. 2014. Food-101--mining discriminative components with random forests. In ECCV. Springer, 446--461."},{"key":"e_1_3_2_1_3_1","doi-asserted-by":"crossref","unstructured":"Nicholas Carlini and David Wagner. 2017. Towards evaluating the robustness of neural networks. In 2017 ieee symp. on security and privacy (sp). Ieee 39--57.","DOI":"10.1109\/SP.2017.49"},{"key":"e_1_3_2_1_4_1","volume-title":"Federated learning of large language models with parameter-efficient prompt tuning and adaptive optimization. arXiv preprint arXiv:2310.15080","author":"Che Tianshi","year":"2023","unstructured":"Tianshi Che, Ji Liu, Yang Zhou, Jiaxiang Ren, Jiwen Zhou, Victor S Sheng, Huaiyu Dai, and Dejing Dou. 2023. Federated learning of large language models with parameter-efficient prompt tuning and adaptive optimization. arXiv preprint arXiv:2310.15080 (2023)."},{"key":"e_1_3_2_1_5_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP49357.2023.10097245"},{"key":"e_1_3_2_1_6_1","doi-asserted-by":"crossref","unstructured":"Mircea Cimpoi Subhransu Maji Iasonas Kokkinos Sammy Mohamed and Andrea Vedaldi. 2014. Describing textures in the wild. In CVPR. 3606--3613.","DOI":"10.1109\/CVPR.2014.461"},{"key":"e_1_3_2_1_7_1","volume-title":"Imagenet: A large-scale hierarchical image database. In CVPR. Ieee, 248--255.","author":"Deng Jia","year":"2009","unstructured":"Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. 2009. Imagenet: A large-scale hierarchical image database. In CVPR. Ieee, 248--255."},{"key":"e_1_3_2_1_8_1","doi-asserted-by":"crossref","unstructured":"Wenlong Deng Christos Thrampoulidis and Xiaoxiao Li. 2024. Unlocking the potential of prompt-tuning in bridging generalized and personalized federated learning. In CVPR. 6087--6097.","DOI":"10.1109\/CVPR52733.2024.00582"},{"key":"e_1_3_2_1_9_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2004.383"},{"key":"e_1_3_2_1_10_1","volume-title":"Promptfl: Let federated participants cooperatively learn prompts instead of models-federated learning in age of foundation model","author":"Guo Tao","year":"2023","unstructured":"Tao Guo, Song Guo, Junxiao Wang, Xueyang Tang, and Wenchao Xu. 2023. Promptfl: Let federated participants cooperatively learn prompts instead of models-federated learning in age of foundation model. IEEE Transactions on Mobile Computing (2023)."},{"key":"e_1_3_2_1_11_1","doi-asserted-by":"publisher","DOI":"10.1109\/JSTARS.2019.2918242"},{"key":"e_1_3_2_1_12_1","doi-asserted-by":"crossref","unstructured":"Dan Hendrycks Steven Basart Norman Mu Saurav Kadavath FrankWang Evan Dorundo Rahul Desai Tyler Zhu Samyak Parajuli Mike Guo et al. 2021. The many faces of robustness: A critical analysis of out-of-distribution generalization. In CVPR. 8340--8349.","DOI":"10.1109\/ICCV48922.2021.00823"},{"key":"e_1_3_2_1_13_1","doi-asserted-by":"crossref","unstructured":"Dan Hendrycks Kevin Zhao Steven Basart Jacob Steinhardt and Dawn Song. 2021. Natural adversarial examples. In CVPR. 15262--15271.","DOI":"10.1109\/CVPR46437.2021.01501"},{"volume-title":"Visual prompt tuning","author":"Jia Menglin","key":"e_1_3_2_1_14_1","unstructured":"Menglin Jia, Luming Tang, Bor-Chun Chen, Claire Cardie, Serge Belongie, Bharath Hariharan, and Ser-Nam Lim. 2022. Visual prompt tuning. In ECCV. Springer, 709--727."},{"key":"e_1_3_2_1_15_1","unstructured":"Xiaojun Jia Yong Zhang Baoyuan Wu Ke Ma Jue Wang and Xiaochun Cao. 2022. LAS-AT: adversarial training with learnable attack strategy. In CVPR. 13398--13408."},{"key":"e_1_3_2_1_16_1","volume-title":"Maple: Multi-modal prompt learning. In CVPR. 19113--19122.","author":"Khattak Muhammad Uzair","year":"2023","unstructured":"Muhammad Uzair Khattak, Hanoona Rasheed, Muhammad Maaz, Salman Khan, and Fahad Shahbaz Khan. 2023. Maple: Multi-modal prompt learning. In CVPR. 19113--19122."},{"key":"e_1_3_2_1_17_1","volume-title":"Torchattacks: A pytorch repository for adversarial attacks. arXiv preprint arXiv:2010.01950","author":"Kim Hoki","year":"2020","unstructured":"Hoki Kim. 2020. Torchattacks: A pytorch repository for adversarial attacks. arXiv preprint arXiv:2010.01950 (2020)."},{"key":"e_1_3_2_1_18_1","unstructured":"Klim Kireev Maksym Andriushchenko and Nicolas Flammarion. 2022. On the effectiveness of adversarial training against common corruptions. In Uncertainty in Artificial Intelligence. PMLR 1012--1021."},{"key":"e_1_3_2_1_19_1","doi-asserted-by":"crossref","unstructured":"Jonathan Krause Michael Stark Jia Deng and Li Fei-Fei. 2013. 3d object representations for fine-grained categorization. In CVPR. 554--561.","DOI":"10.1109\/ICCVW.2013.77"},{"key":"e_1_3_2_1_20_1","doi-asserted-by":"crossref","unstructured":"Lin Li Haoyan Guan Jianing Qiu and Michael Spratling. 2024. One prompt word is enough to boost adversarial robustness for pre-trained vision-language models. In CVPR. 24408--24419.","DOI":"10.1109\/CVPR52733.2024.02304"},{"key":"e_1_3_2_1_21_1","volume-title":"Towards deep learning models resistant to adversarial attacks. arXiv preprint arXiv:1706.06083","author":"Madry Aleksander","year":"2017","unstructured":"Aleksander Madry, Aleksandar Makelov, Ludwig Schmidt, Dimitris Tsipras, and Adrian Vladu. 2017. Towards deep learning models resistant to adversarial attacks. arXiv preprint arXiv:1706.06083 (2017)."},{"key":"e_1_3_2_1_22_1","volume-title":"Fine-grained visual classification of aircraft. arXiv preprint arXiv:1306.5151","author":"Maji Subhransu","year":"2013","unstructured":"Subhransu Maji, Esa Rahtu, Juho Kannala, Matthew Blaschko, and Andrea Vedaldi. 2013. Fine-grained visual classification of aircraft. arXiv preprint arXiv:1306.5151 (2013)."},{"key":"e_1_3_2_1_23_1","unstructured":"Brendan McMahan Eider Moore Daniel Ramage Seth Hampson and Blaise Aguera y Arcas. 2017. Communication-efficient learning of deep networks from decentralized data. In Artificial intelligence and statistics. PMLR 1273--1282."},{"key":"e_1_3_2_1_24_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICVGIP.2008.47"},{"volume-title":"Cats and dogs","author":"Parkhi Omkar M","key":"e_1_3_2_1_25_1","unstructured":"Omkar M Parkhi, Andrea Vedaldi, Andrew Zisserman, and CV Jawahar. 2012. Cats and dogs. In CVPR. IEEE, 3498--3505."},{"key":"e_1_3_2_1_26_1","volume-title":"True few-shot learning with language models. Advances in neural information processing systems 34","author":"Perez Ethan","year":"2021","unstructured":"Ethan Perez, Douwe Kiela, and Kyunghyun Cho. 2021. True few-shot learning with language models. Advances in neural information processing systems 34 (2021), 11054--11070."},{"key":"e_1_3_2_1_27_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.patcog.2022.108889"},{"key":"e_1_3_2_1_28_1","volume-title":"Madan Ravi Ganesh, Zhenzhen Li, Lu Peng, and Wan-Yi Lin.","author":"Qiu Chen","year":"2024","unstructured":"Chen Qiu, Xingyu Li, Chaithanya Kumar Mummadi, Madan Ravi Ganesh, Zhenzhen Li, Lu Peng, and Wan-Yi Lin. 2024. Federated text-driven prompt generation for vision-language models. In ICLR."},{"key":"e_1_3_2_1_29_1","volume-title":"Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al.","author":"Radford Alec","year":"2021","unstructured":"Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. 2021. Learning transferable visual models from natural language supervision. In ICML. PMLR, 8748--8763."},{"key":"e_1_3_2_1_30_1","volume-title":"International conf. on machine learning. PMLR, 5389--5400","author":"Recht Benjamin","year":"2019","unstructured":"Benjamin Recht, Rebecca Roelofs, Ludwig Schmidt, and Vaishaal Shankar. 2019. Do imagenet classifiers generalize to imagenet?. In International conf. on machine learning. PMLR, 5389--5400."},{"key":"e_1_3_2_1_31_1","volume-title":"Adversarial training in communication constrained federated learning. arXiv preprint arXiv:2103.01319","author":"Shah Devansh","year":"2021","unstructured":"Devansh Shah, Parijat Dube, Supriyo Chakraborty, and Ashish Verma. 2021. Adversarial training in communication constrained federated learning. arXiv preprint arXiv:2103.01319 (2021)."},{"key":"e_1_3_2_1_32_1","volume-title":"UCF101: A dataset of 101 human actions classes from videos in the wild. arXiv preprint arXiv:1212.0402","author":"Soomro K","year":"2012","unstructured":"K Soomro. 2012. UCF101: A dataset of 101 human actions classes from videos in the wild. arXiv preprint arXiv:1212.0402 (2012)."},{"key":"e_1_3_2_1_33_1","volume-title":"Intriguing properties of neural networks. arXiv preprint arXiv:1312.6199","author":"Szegedy C","year":"2013","unstructured":"C Szegedy. 2013. Intriguing properties of neural networks. arXiv preprint arXiv:1312.6199 (2013)."},{"key":"e_1_3_2_1_34_1","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v38i6.28345"},{"key":"e_1_3_2_1_35_1","volume-title":"Learning robust global representations by penalizing local predictive power. Advances in neural information processing systems 32","author":"Wang Haohan","year":"2019","unstructured":"Haohan Wang, Songwei Ge, Zachary Lipton, and Eric P Xing. 2019. Learning robust global representations by penalizing local predictive power. Advances in neural information processing systems 32 (2019)."},{"key":"e_1_3_2_1_36_1","volume-title":"Dual prompt tuning for domain-aware federated learning. arXiv preprint arXiv:2310.03103","author":"Shah Anshul","year":"2023","unstructured":"GuoyizheWei, FengWang, Anshul Shah, and Rama Chellappa. 2023. Dual prompt tuning for domain-aware federated learning. arXiv preprint arXiv:2310.03103 (2023)."},{"key":"e_1_3_2_1_37_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2022.acl-long.174"},{"volume-title":"Sun database: Large-scale scene recognition from abbey to zoo. In 2010 IEEE computer society conf. on computer vision and pattern recognition","author":"Xiao Jianxiong","key":"e_1_3_2_1_38_1","unstructured":"Jianxiong Xiao, James Hays, Krista A Ehinger, Aude Oliva, and Antonio Torralba. 2010. Sun database: Large-scale scene recognition from abbey to zoo. In 2010 IEEE computer society conf. on computer vision and pattern recognition. IEEE, 3485--3492."},{"key":"e_1_3_2_1_39_1","unstructured":"Cihang Xie Zhishuai Zhang Yuyin Zhou Song Bai Jianyu Wang Zhou Ren and Alan L Yuille. 2019. Improving transferability of adversarial examples with input diversity. In CVPR. 2730--2739."},{"key":"e_1_3_2_1_40_1","doi-asserted-by":"crossref","unstructured":"Fu-En Yang Chien-Yi Wang and Yu-Chiang Frank Wang. 2023. Efficient model personalization in federated learning via client-specific prompt generation. In CVPR. 19159--19168.","DOI":"10.1109\/ICCV51070.2023.01755"},{"key":"e_1_3_2_1_41_1","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v37i9.26331"},{"volume-title":"Adversarial prompt tuning for vision-language models","author":"Zhang Jiaming","key":"e_1_3_2_1_42_1","unstructured":"Jiaming Zhang, Xingjun Ma, Xin Wang, Lingyu Qiu, Jiaqi Wang, Yu-Gang Jiang, and Jitao Sang. 2025. Adversarial prompt tuning for vision-language models. In ECCV. Springer, 56--72."},{"key":"e_1_3_2_1_43_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP49357.2023.10095356"},{"key":"e_1_3_2_1_44_1","volume-title":"Few-Shot Adversarial Prompt Learning on Vision-Language Models. arXiv preprint arXiv:2403.14774","author":"Zhou Yiwei","year":"2024","unstructured":"Yiwei Zhou, Xiaobo Xia, Zhiwei Lin, Bo Han, and Tongliang Liu. 2024. Few-Shot Adversarial Prompt Learning on Vision-Language Models. arXiv preprint arXiv:2403.14774 (2024)."}],"event":{"name":"MM '25: The 33rd ACM International Conference on Multimedia","sponsor":["SIGMM ACM Special Interest Group on Multimedia"],"location":"Dublin Ireland","acronym":"MM '25"},"container-title":["Proceedings of the 33rd ACM International Conference on Multimedia"],"original-title":[],"link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3746027.3755387","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,12,10]],"date-time":"2025-12-10T04:08:50Z","timestamp":1765339730000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3746027.3755387"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,10,27]]},"references-count":44,"alternative-id":["10.1145\/3746027.3755387","10.1145\/3746027"],"URL":"https:\/\/doi.org\/10.1145\/3746027.3755387","relation":{},"subject":[],"published":{"date-parts":[[2025,10,27]]},"assertion":[{"value":"2025-10-27","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}