{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,22]],"date-time":"2026-07-22T03:41:48Z","timestamp":1784691708812,"version":"3.55.0"},"reference-count":52,"publisher":"Association for Computing Machinery (ACM)","issue":"3","license":[{"start":{"date-parts":[[2025,3,7]],"date-time":"2025-03-07T00:00:00Z","timestamp":1741305600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Multimedia Comput. Commun. Appl."],"published-print":{"date-parts":[[2025,3,31]]},"abstract":"<jats:p>Unsupervised Domain Adaptation (UDA) aims to transfer models trained on a labeled source domain to an unlabeled target domain. Due to the excellent generalization ability of Vision Language Models (VLMs) such as CLIP in downstream tasks, most recent methods apply CLIP to UDA tasks through learning domain-specific text prompts for source and target domains separately. However, these methods fail to dynamically adjust image features based on the characteristics of their respective domains, thereby limiting their alignment with domain-specific text prompts in CLIP\u2019s joint space, which is a key factor in improving classification performance in the target domain. To bridge this gap, we propose a Unified Text-Image Space Alignment with Cross-Modal Prompting (UTISA) framework for UDA. First, we introduce a Cross-Modal Prompt Learning (CMP) module to generate domain-specific image prompts and layer-specific image prompts for the visual branch to encode domain-specific knowledge globally and locally. Second, under the guidance of image prompts, we introduce a Domain-Aware Multi-Layer Feature Fusion (DMF) module to construct multi-layer domain features for each domain and enhance the image features with these multi-layer domain features, which enables the image features to better reflect the characteristics of their respective domains, thereby promoting their alignment with domain-specific text prompts. Moreover, we introduce a Perturbation-Driven Regularization (PDR) mechanism for the target domain to enhance the robustness and generalization of the model. The experiments demonstrate that UTISA achieves the best performance on three mainstream UDA benchmarks, including 87.9% on Office-Home, 90.9% on VisDA-2017, and 62.4% on DomainNet.<\/jats:p>","DOI":"10.1145\/3715699","type":"journal-article","created":{"date-parts":[[2025,2,5]],"date-time":"2025-02-05T14:50:41Z","timestamp":1738767041000},"page":"1-20","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":6,"title":["Unified Text-Image Space Alignment with Cross-Modal Prompting in CLIP for UDA"],"prefix":"10.1145","volume":"21","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-8923-6997","authenticated-orcid":false,"given":"Yifan","family":"Jiao","sequence":"first","affiliation":[{"name":"Nanjing University of Posts and Telecommunications, Nanjing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0009-6756-2079","authenticated-orcid":false,"given":"Chenglong","family":"Cai","sequence":"additional","affiliation":[{"name":"Nanjing University of Posts and Telecommunications, Nanjing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-5956-831X","authenticated-orcid":false,"given":"Bing-Kun","family":"Bao","sequence":"additional","affiliation":[{"name":"Nanjing University of Posts and Telecommunications, Nanjing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2025,3,7]]},"reference":[{"key":"e_1_3_1_2_2","doi-asserted-by":"publisher","DOI":"10.1007\/s10994-009-5152-4"},{"key":"e_1_3_1_3_2","first-page":"74127","article-title":"Multi-prompt alignment for multi-source unsupervised domain adaptation","author":"Chen Haoran","year":"2024","unstructured":"Haoran Chen, Xintong Han, Zuxuan Wu, and Yu-Gang Jiang. 2024. Multi-prompt alignment for multi-source unsupervised domain adaptation. Advances in Neural Information Processing Systems 36 (2024), 74127\u201374139.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_1_4_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2009.5206848"},{"key":"e_1_3_1_5_2","first-page":"4171","volume-title":"Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies","author":"Devlin Jacob","year":"2018","unstructured":"Jacob Devlin, Ming Wei Chang, Kenton Lee, and Kristina Toutanova. 2018. BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, 4171\u20134186."},{"key":"e_1_3_1_6_2","volume-title":"Proceedings of the International Conference on Learning Representations","author":"Dosovitskiy Alexey","year":"2021","unstructured":"Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al. 2021. An image is worth 16x16 words: Transformers for image recognition at scale. In Proceedings of the International Conference on Learning Representations."},{"key":"e_1_3_1_7_2","first-page":"5436","volume-title":"Proceedings of the 31st International Joint Conference on Artificial Intelligence","author":"Du Yifan","year":"2022","unstructured":"Yifan Du, Zikang Liu, Junyi Li, and Wayne Xin Zhao. 2022. A survey of vision-language pre-trained models. In Proceedings of the 31st International Joint Conference on Artificial Intelligence, 5436\u20135443."},{"key":"e_1_3_1_8_2","first-page":"18577","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision (ICCV)","author":"Fahes Mohammad","year":"2023","unstructured":"Mohammad Fahes, Tuan-Hung Vu, Andrei Bursuc, Patrick P\u00e9rez, and Raoul De Charette. 2023. P\u00d8DA: Prompt-driven zero-shot domain adaptation. In Proceedings of the IEEE\/CVF International Conference on Computer Vision (ICCV), 18577\u201318587."},{"key":"e_1_3_1_9_2","first-page":"1180","volume-title":"Proceedings of the International Conference on Machine Learning","author":"Ganin Yaroslav","year":"2014","unstructured":"Yaroslav Ganin and Victor Lempitsky. 2014. Unsupervised domain adaptation by backpropagation. In Proceedings of the International Conference on Machine Learning, 1180\u20131189."},{"key":"e_1_3_1_10_2","first-page":"587","volume-title":"Proceedings of the 16th European Conference on Computer Vision (ECCV \u201920)","author":"Gao Jian","year":"2020","unstructured":"Jian Gao, Yang Hua, Guosheng Hu, Chi Wang, and Neil M Robertson. 2020. Reducing distributional uncertainty by mutual information maximisation and transferable feature learning. In Proceedings of the 16th European Conference on Computer Vision (ECCV \u201920), Part XXIII, 587\u2013605."},{"issue":"1","key":"e_1_3_1_11_2","first-page":"1","article-title":"Domain adaptation via prompt learning","volume":"36","author":"Ge Chunjiang","year":"2023","unstructured":"Chunjiang Ge, Rui Huang, Mixue Xie, Zihang Lai, Shiji Song, Shuang Li, and Gao Huang. 2023. Domain adaptation via prompt learning. IEEE Transactions on Neural Networks and Learning Systems 36, 1 (2023), 1\u201311.","journal-title":"IEEE Transactions on Neural Networks and Learning Systems"},{"key":"e_1_3_1_12_2","first-page":"4893","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR)","author":"Kang Yi Yang Guoliang","year":"2019","unstructured":"Yi Yang Guoliang Kang, Lu Jiang, and Alexander G Hauptmann. 2019. Contrastive adaptation network for unsupervised domain adaptation. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 4893\u20134902."},{"key":"e_1_3_1_13_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.90"},{"key":"e_1_3_1_14_2","first-page":"4904","volume-title":"Proceedings of the International Conference on Machine Learning","author":"Jia Chao","year":"2021","unstructured":"Chao Jia, Yinfei Yang, Ye Xia, Yi Ting Chen, and Tom Duerig. 2021. Scaling up visual and vision-language representation learning with noisy text supervision. In Proceedings of the International Conference on Machine Learning, 4904\u20134916."},{"key":"e_1_3_1_15_2","doi-asserted-by":"publisher","DOI":"10.1162\/tacl_a_00324"},{"key":"e_1_3_1_16_2","first-page":"423","article-title":"The power of scale for parameter-efficient prompt tuning","volume":"8","author":"Jiang Zhengbao","year":"2020","unstructured":"Zhengbao Jiang, Frank F. Xu, Jun Araki, and Graham Neubig. 2020. The power of scale for parameter-efficient prompt tuning. Transactions of the Association for Computational Linguistics 8 (2020), 423\u2013438.","journal-title":"Transactions of the Association for Computational Linguistics"},{"key":"e_1_3_1_17_2","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2024.3384766"},{"key":"e_1_3_1_18_2","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2021.3049948"},{"issue":"6","key":"e_1_3_1_19_2","doi-asserted-by":"crossref","first-page":"7338","DOI":"10.1109\/TPAMI.2022.3218569","article-title":"Dual instance-consistent network for cross-domain object detection","volume":"45","author":"Jiao Yifan","year":"2022","unstructured":"Yifan Jiao, Hantao Yao, and Changsheng Xu. 2022. Dual instance-consistent network for cross-domain object detection. IEEE Transactions on Pattern Analysis and Machine Intelligence 45, 6 (2021), 7338\u20137352.","journal-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence"},{"key":"e_1_3_1_20_2","first-page":"19113","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Uzair Khattak Muhammad","year":"2023","unstructured":"Muhammad Uzair Khattak, Hanoona Rasheed, Muhammad Maaz, Salman Khan, and Fahad Shahbaz Khan. 2023. Maple: Multi-modal prompt learning. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, 19113\u201319122."},{"key":"e_1_3_1_21_2","doi-asserted-by":"crossref","first-page":"3045","DOI":"10.18653\/v1\/2021.emnlp-main.243","volume-title":"Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing","author":"Lester Brian","year":"2021","unstructured":"Brian Lester, Rami Al-Rfou, and Noah Constant. 2021. Prefix-tuning: Optimizing continuous prompts for generation. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, 3045\u20133059."},{"key":"e_1_3_1_22_2","unstructured":"Liunian Harold Li Mark Yatskar Da Yin Cho Jui Hsieh and Kai Wei Chang. 2019. VisualBERT: A simple and performant baseline for vision and language. arXiv:1908.03557. Retrieved from https:\/\/arxiv.org\/abs\/"},{"key":"e_1_3_1_23_2","first-page":"6028","volume-title":"Proceedings of the International Conference on Machine Learning","author":"Liang Jian","year":"2020","unstructured":"Jian Liang, Dapeng Hu, and Jiashi Feng. 2020. Do we really need to access the source data? Source hypothesis transfer for unsupervised domain adaptation. In Proceedings of the International Conference on Machine Learning, 6028\u20136039."},{"key":"e_1_3_1_24_2","doi-asserted-by":"publisher","DOI":"10.1145\/3560815"},{"key":"e_1_3_1_25_2","volume-title":"Proceedings of the 32nd International Conference on Neural Information Processing Systems, 1647\u20131657","author":"Long Mingsheng","year":"2018","unstructured":"Mingsheng Long, Zhangjie Cao, Jianmin Wang, and Michael I. Jordan. 2018. Conditional adversarial domain adaptation. In Proceedings of the 32nd International Conference on Neural Information Processing Systems, 1647\u20131657."},{"key":"e_1_3_1_26_2","first-page":"97","volume-title":"Proceedings of the International Conference on Machine Learning","author":"Long Mingsheng","year":"2015","unstructured":"Mingsheng Long and Jianmin Wang. 2015. Learning transferable features with deep adaptation networks. In Proceedings of the International Conference on Machine Learning, 97\u2013105."},{"key":"e_1_3_1_27_2","first-page":"2208","volume-title":"Proceedings of the International Conference on Machine Learning","author":"Long Mingsheng","year":"2017","unstructured":"Mingsheng Long, Han Zhu, Jianmin Wang, and Michael I. Jordan. 2017. Deep transfer learning with joint adaptation networks. In Proceedings of the International Conference on Machine Learning, 2208\u20132217."},{"key":"e_1_3_1_28_2","first-page":"1094","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR)","author":"Na Jaemin","year":"2021","unstructured":"Jaemin Na, Heechul Jung, Hyung Jin Chang, and Wonjun Hwang. 2021. FixBi: Bridging domain spaces for unsupervised domain adaptation. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 1094\u20131103."},{"key":"e_1_3_1_29_2","doi-asserted-by":"publisher","DOI":"10.1109\/TKDE.2009.191"},{"key":"e_1_3_1_30_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.00234"},{"key":"e_1_3_1_31_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2019.00149"},{"key":"e_1_3_1_32_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/D19-1250"},{"key":"e_1_3_1_33_2","first-page":"4748","volume-title":"Proceedings of the International Conference on Machine Learning","author":"Radford Alec","year":"2021","unstructured":"Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, and Jack Clark. 2021. Learning transferable visual models from natural language supervision. In Proceedings of the International Conference on Machine Learning, 4748\u20134763."},{"key":"e_1_3_1_34_2","first-page":"8748","volume-title":"Proceedings of the International Conference on Machine Learning","author":"Radford Alec","year":"2021","unstructured":"Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. 2021. Learning transferable visual models from natural language supervision. In Proceedings of the International Conference on Machine Learning, 8748\u20138763."},{"key":"e_1_3_1_35_2","article-title":"Improving language understanding by generative pre-training","author":"Radford Alec","year":"2018","unstructured":"Alec Radford, Karthik Narasimhan, Tim Salimans, and Ilya Sutskever. 2018. Improving language understanding by generative pre-training. Computer Science.","journal-title":"Computer Science"},{"key":"e_1_3_1_36_2","first-page":"4058","volume-title":"Proceedings of the AAAI Conference on Artificial Intelligence","author":"Shen Jian","year":"2017","unstructured":"Jian Shen, Yanru Qu, Weinan Zhang, and Yong Yu. 2017. Wasserstein distance guided representation learning for domain adaptation. In Proceedings of the AAAI Conference on Artificial Intelligence, 4058\u20134065."},{"key":"e_1_3_1_37_2","first-page":"4222","volume-title":"Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing","author":"Shin Taylor","year":"2020","unstructured":"Taylor Shin, Yasaman Razeghi, Robert L. Logan IV, Eric Wallace, and Sameer Singh. 2020. AutoPrompt: Eliciting knowledge from language models with automatically generated prompts. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing, 4222\u20134235."},{"key":"e_1_3_1_38_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCVW60793.2023.00470"},{"key":"e_1_3_1_39_2","first-page":"153","volume-title":"Domain Adaptation in Computer Vision Applications","author":"Sun Baochen","year":"2017","unstructured":"Baochen Sun. 2017. Correlation alignment for domain adaptation. In Domain Adaptation in Computer Vision Applications. Springer, 153\u2013171."},{"key":"e_1_3_1_40_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52688.2022.00705"},{"key":"e_1_3_1_41_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.00875"},{"key":"e_1_3_1_42_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2011.5995347"},{"key":"e_1_3_1_43_2","first-page":"6000","article-title":"Attention is all you need","author":"Vaswani Ashish","year":"2017","unstructured":"Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, \u0141ukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. Advances in Neural Information Processing Systems 30 (2017), 6000\u20136010.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_1_44_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.572"},{"issue":"5","key":"e_1_3_1_45_2","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/3400066","article-title":"A survey of unsupervised deep domain adaptation","volume":"11","author":"Wilson Garrett","year":"2018","unstructured":"Garrett Wilson and Diane J. Cook. 2018. A survey of unsupervised deep domain adaptation. ACM Transactions on Intelligent Systems and Technology 11, 5 (2018), 1\u201346.","journal-title":"ACM Transactions on Intelligent Systems and Technology"},{"key":"e_1_3_1_46_2","doi-asserted-by":"publisher","DOI":"10.1145\/3486251"},{"key":"e_1_3_1_47_2","volume-title":"Proceedings of the International Conference on Learning Representations","author":"Xu Tongkun","year":"2022","unstructured":"Tongkun Xu, Weihua Chen, Pichao Wang, Fan Wang, Hao Li, and Rong Jin. 2022. Cdtrans: Cross-domain transformer for unsupervised domain adaptation. In Proceedings of the International Conference on Learning Representations."},{"key":"e_1_3_1_48_2","first-page":"520","volume-title":"Proceedings of the IEEE\/CVF Winter Conference on Applications of Computer Vision","author":"Yang Jinyu","year":"2023","unstructured":"Jinyu Yang, Jingjing Liu, Ning Xu, and Junzhou Huang. 2023. TVT: Transferable vision transformer for unsupervised domain adaptation. In Proceedings of the IEEE\/CVF Winter Conference on Applications of Computer Vision, 520\u2013530."},{"key":"e_1_3_1_49_2","doi-asserted-by":"publisher","DOI":"10.1145\/2700286"},{"key":"e_1_3_1_50_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2024.3369699"},{"issue":"5","key":"e_1_3_1_51_2","doi-asserted-by":"crossref","first-page":"2775","DOI":"10.1109\/TPAMI.2020.3036956","article-title":"Unsupervised multi-class domain adaptation: Theory, algorithms, and practice","volume":"44","author":"Zhang Yabin","year":"2022","unstructured":"Yabin Zhang, Bin Deng, Hui Tang, Lei Zhang, and Kui Jia. 2022. Unsupervised multi-class domain adaptation: Theory, algorithms, and practice. IEEE Transactions on Pattern Analysis and Machine Intelligence 44, 5 (2022), 2775\u20132792.","journal-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence"},{"key":"e_1_3_1_52_2","doi-asserted-by":"publisher","DOI":"10.1007\/s11263-022-01653-1"},{"key":"e_1_3_1_53_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52729.2023.00347"}],"container-title":["ACM Transactions on Multimedia Computing, Communications, and Applications"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3715699","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3715699","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,19]],"date-time":"2025-06-19T01:19:14Z","timestamp":1750295954000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3715699"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,3,7]]},"references-count":52,"journal-issue":{"issue":"3","published-print":{"date-parts":[[2025,3,31]]}},"alternative-id":["10.1145\/3715699"],"URL":"https:\/\/doi.org\/10.1145\/3715699","relation":{},"ISSN":["1551-6857","1551-6865"],"issn-type":[{"value":"1551-6857","type":"print"},{"value":"1551-6865","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,3,7]]},"assertion":[{"value":"2024-09-11","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-01-20","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-03-07","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}