{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,19]],"date-time":"2026-06-19T15:01:47Z","timestamp":1781881307491,"version":"3.54.5"},"reference-count":56,"publisher":"Association for Computing Machinery (ACM)","issue":"11","funder":[{"name":"the National Key R&D Program of China","award":["2022YFE0138600"],"award-info":[{"award-number":["2022YFE0138600"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Multimedia Comput. Commun. Appl."],"published-print":{"date-parts":[[2025,11,30]]},"abstract":"<jats:p>Empowered by the continuous integration of social multimedia and artificial intelligence, the application scenarios of Information Retrieval (IR) progressively tend to be diversified and personalized. Currently, User-Generated Content (UGC) systems have great potential to handle the interactions between large-scale users and massive media contents. As an emerging multimedia IR, Fashion Compatibility Modeling (FCM) aims to predict the matching degree of each given outfit and provide complementary item recommendation for user queries. Although existing studies attempt to explore the FCM task from a multi-modal perspective with promising progress, they still fail to fully leverage the interactions between multi-modal information or ignore the item\u2013item contextual connectivities of intra-outfit. In this article, a novel FCM scheme is proposed based on Correlation-Aware Cross-Modal Attention Network. To better tackle these issues, our work mainly focuses on enhancing comprehensive multi-modal representations of fashion items by integrating the cross-modal collaborative contents and uncovering the contextual correlations. Since the multi-modal information of fashion items can deliver various semantic clues from multiple aspects, a modality-driven collaborative learning module is presented to explicitly model the interactions of modal consistency and complementarity via a co-attention mechanism. Considering the rich connections among numerous items in each outfit as contextual cues, a correlation-aware information aggregation module is further designed to adaptively capture significant intra-correlations of item\u2013item for characterizing the content-aware outfit representations. Experiments conducted on two real-world fashion datasets demonstrate the superiority of our approach over state-of-the-art methods.<\/jats:p>","DOI":"10.1145\/3698772","type":"journal-article","created":{"date-parts":[[2024,10,5]],"date-time":"2024-10-05T09:15:38Z","timestamp":1728119738000},"page":"1-24","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":10,"title":["Correlation-Aware Cross-Modal Attention Network for Fashion Compatibility Modeling in UGC Systems"],"prefix":"10.1145","volume":"21","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-7424-7204","authenticated-orcid":false,"given":"Kai","family":"Cui","sequence":"first","affiliation":[{"name":"Hubei Key Laboratory of Distributed System Security, Hubei Engineering Research Center on Big Data Security, School of Cyber Science and Engineering, Huazhong University of Science and Technology, Wuhan, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-8993-8940","authenticated-orcid":false,"given":"Shenghao","family":"Liu","sequence":"additional","affiliation":[{"name":"Hubei Key Laboratory of Distributed System Security, Hubei Engineering Research Center on Big Data Security, School of Cyber Science and Engineering, Huazhong University of Science and Technology, Wuhan, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0006-1960-9563","authenticated-orcid":false,"given":"Wei","family":"Feng","sequence":"additional","affiliation":[{"name":"China Nuclear Power Operation Technology Corporation, Ltd., Wuhan, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-2641-1708","authenticated-orcid":false,"given":"Xianjun","family":"Deng","sequence":"additional","affiliation":[{"name":"Hubei Key Laboratory of Distributed System Security, Hubei Engineering Research Center on Big Data Security, School of Cyber Science and Engineering, Huazhong University of Science and Technology, Wuhan, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0000-6297-2903","authenticated-orcid":false,"given":"Liangbin","family":"Gao","sequence":"additional","affiliation":[{"name":"Hubei Key Laboratory of Distributed System Security, Hubei Engineering Research Center on Big Data Security, School of Cyber Science and Engineering, Huazhong University of Science and Technology, Wuhan, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0005-0457-7397","authenticated-orcid":false,"given":"Minmin","family":"Cheng","sequence":"additional","affiliation":[{"name":"Hubei Key Laboratory of Distributed System Security, Hubei Engineering Research Center on Big Data Security, School of Cyber Science and Engineering, Huazhong University of Science and Technology, Wuhan, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-4590-3874","authenticated-orcid":false,"given":"Hongwei","family":"Lu","sequence":"additional","affiliation":[{"name":"Hubei Key Laboratory of Distributed System Security, Hubei Engineering Research Center on Big Data Security, School of Cyber Science and Engineering, Huazhong University of Science and Technology, Wuhan, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-7986-4244","authenticated-orcid":false,"given":"Laurence T.","family":"Yang","sequence":"additional","affiliation":[{"name":"Huazhong University of Science and Technology, School of Computer Science and Technology, Wuhan, Hubei, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2025,11,10]]},"reference":[{"key":"e_1_3_1_2_2","doi-asserted-by":"publisher","unstructured":"Gairui Bai Wei Xi Xiaopeng Hong Xinhui Liu Yang Yue and Songwen Zhao. 2023. Robust and rotation-equivariant contrastive learning. IEEE Transactions on Neural Networks and Learning Systems (2023). Retrieved from 10.1109\/TNNLS.2023.3243258","DOI":"10.1109\/TNNLS.2023.3243258"},{"key":"e_1_3_1_3_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2018.2798607"},{"key":"e_1_3_1_4_2","doi-asserted-by":"publisher","DOI":"10.1145\/3606368"},{"key":"e_1_3_1_5_2","volume-title":"Proceedings of the International Conference on Learning Representations","author":"Chen Yiyang","year":"2024","unstructured":"Yiyang Chen, Zhedong Zheng, Wei Ji, Leigang Qu, and Tat-Seng Chua. 2024. Composed image retrieval with text feedback via multi-grained uncertainty regularization. In Proceedings of the International Conference on Learning Representations. arXiv:2211.07394. Retrieved from https:\/\/arxiv.org\/abs\/2211.07394"},{"key":"e_1_3_1_6_2","first-page":"12617","volume-title":"Proceedings of IEEE Conference on Computer Vision and Pattern Recognition","author":"Cucurull Guillem","year":"2019","unstructured":"Guillem Cucurull, Perouz Taslakian, and David Vazquez. 2019. Context-aware visual compatibility prediction. In Proceedings of IEEE Conference on Computer Vision and Pattern Recognition, 12617\u201312626."},{"key":"e_1_3_1_7_2","doi-asserted-by":"publisher","DOI":"10.1145\/3308558.3313444"},{"key":"e_1_3_1_8_2","doi-asserted-by":"crossref","first-page":"1s","DOI":"10.1145\/3531017","article-title":"Disentangling features for fashion recommendation","volume":"19","author":"Divitiis Lavinia De","year":"2023","unstructured":"Lavinia De Divitiis, Federico Becattini, Claudio Baecchi, and Alberto Del Bimbo. 2023. Disentangling features for fashion recommendation. ACM Transactions on Multimedia Computing, Communications and Applications 19, 1s (2023), 1\u201321.","journal-title":"ACM Transactions on Multimedia Computing, Communications and Applications"},{"key":"e_1_3_1_9_2","doi-asserted-by":"publisher","DOI":"10.1109\/TNNLS.2022.3173295"},{"key":"e_1_3_1_10_2","doi-asserted-by":"publisher","DOI":"10.1109\/TMM.2024.3406196"},{"key":"e_1_3_1_11_2","doi-asserted-by":"publisher","DOI":"10.1109\/TMM.2023.3301224"},{"key":"e_1_3_1_12_2","first-page":"961","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition","author":"Fei Hao","year":"2024","unstructured":"Hao Fei, Shengqiong Wu, Wei Ji, Hanwang Zhang, and Tat-Seng Chua. 2024. Dysen-VDM: Empowering dynamics-aware text-to-video diffusion with LLMs. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 961\u2013970."},{"key":"e_1_3_1_13_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2020.3025062"},{"key":"e_1_3_1_14_2","first-page":"21","article-title":"Federated inverse reinforcement learning for smart ICUs with differential privacy","volume":"10","author":"Gong Wei","year":"2023","unstructured":"Wei Gong, Linxiao Cao, Yifei Zhu, Fang Zuo, Xin He, and Haoquan Zhou. 2023. Federated inverse reinforcement learning for smart ICUs with differential privacy. IEEE Internet of Things Journal 10, 21 (2023), 19117\u201319124.","journal-title":"IEEE Internet of Things Journal"},{"key":"e_1_3_1_15_2","first-page":"482","volume-title":"Proceedings of the International ACM SIGIR Conference on Research and Development in Information Retrieval","author":"Guan Weili","year":"2022","unstructured":"Weili Guan, Fangkai Jiao, Xuemeng Song, Haokun Wen, Chung-Hsing Yeh, and Xiaojun Chang. 2022. Personalized fashion compatibility modeling via metapath-guided heterogeneous graph learning. In Proceedings of the International ACM SIGIR Conference on Research and Development in Information Retrieval, 482\u2013491."},{"key":"e_1_3_1_16_2","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2022.3187290"},{"key":"e_1_3_1_17_2","first-page":"2299","volume-title":"Proceedings of the ACM International Conference on Multimedia","author":"Guan Weili","year":"2021","unstructured":"Weili Guan, Haokun Wen, Xuemeng Song, Chung-Hsing Yeh, Xiaojun Chang, and Liqiang Nie. 2021. Multimodal compatibility modeling via exploring the consistent and complementary correlations. In Proceedings of the ACM International Conference on Multimedia, 2299\u20132307."},{"key":"e_1_3_1_18_2","first-page":"15908","article-title":"Transformer in transformer","volume":"34","author":"Han Kai","year":"2021","unstructured":"Kai Han, An Xiao, Enhua Wu, Jianyuan Guo, Chunjing Xu, and Yunhe Wang. 2021. Transformer in transformer. In Advances in Neural Information Processing Systems, Vol. 34, 15908\u201315919.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_1_19_2","first-page":"1078","volume-title":"Proceedings of the ACM international Conference on Multimedia","author":"Han Xintong","year":"2017","unstructured":"Xintong Han, Zuxuan Wu, Yu-Gang Jiang, and Larry S. Davis. 2017. Learning fashion compatibility with bidirectional LSTMs. In Proceedings of the ACM international Conference on Multimedia, 1078\u20131086."},{"key":"e_1_3_1_20_2","first-page":"770","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition,","author":"He Kaiming","year":"2016","unstructured":"Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2016. Deep residual learning for image recognition. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 770\u2013778."},{"key":"e_1_3_1_21_2","doi-asserted-by":"publisher","DOI":"10.1145\/3581783.3612088"},{"key":"e_1_3_1_22_2","first-page":"12830","volume-title":"Proceedings of the AAAI Conference on Artificial Intelligence","volume":"38","author":"Ji Xinyuan","year":"2024","unstructured":"Xinyuan Ji, Zhaowei Zhu, Wei Xi, Olga Gadyatskaya, Zilong Song, Yong Cai, and Yang Liu. 2024. FedFixer: Mitigating heterogeneous label noise in federated learning. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 38, 12830\u201312838."},{"key":"e_1_3_1_23_2","doi-asserted-by":"crossref","first-page":"4229","DOI":"10.1145\/3580305.3599824","volume-title":"Proceedings of the ACM SIGKDD Conference on Knowledge Discovery and Data Mining","author":"Jiang Lin","year":"2023","unstructured":"Lin Jiang, Shuai Wang, Baoshen Guo, Hai Wang, Desheng Zhang, and Guang Wang. 2023. FairCod: A fairness-aware concurrent dispatch system for large-scale instant delivery services. In Proceedings of the ACM SIGKDD Conference on Knowledge Discovery and Data Mining, 4229\u20134238."},{"key":"e_1_3_1_24_2","doi-asserted-by":"publisher","DOI":"10.1109\/TMM.2023.3246796"},{"issue":"1","key":"e_1_3_1_25_2","doi-asserted-by":"crossref","first-page":"353","DOI":"10.1109\/JIOT.2023.3285601","article-title":"Multimodal high-order relationship inference network for fashion compatibility modeling in internet of multimedia things","volume":"11","author":"Jing Peiguang","year":"2023","unstructured":"Peiguang Jing, Kai Cui, Jing Zhang, Yun Li, and Yuting Su. 2023. Multimodal high-order relationship inference network for fashion compatibility modeling in internet of multimedia things. IEEE Internet of Things Journal 11, 1 (2023), 353\u2013365.","journal-title":"IEEE Internet of Things Journal"},{"key":"e_1_3_1_26_2","doi-asserted-by":"publisher","DOI":"10.1109\/TKDE.2023.3337077"},{"key":"e_1_3_1_27_2","first-page":"159","volume-title":"Proceedings of the International ACM SIGIR Conference on Research and Development in Information Retrieval","author":"Li Xingchen","year":"2020","unstructured":"Xingchen Li, Xiang Wang, Xiangnan He, Long Chen, Jun Xiao, and Tat-Seng Chua. 2020. Hierarchical fashion graph network for personalized outfit recommendation. In Proceedings of the International ACM SIGIR Conference on Research and Development in Information Retrieval, 159\u2013168."},{"key":"e_1_3_1_28_2","doi-asserted-by":"publisher","DOI":"10.1109\/TMM.2017.2690144"},{"key":"e_1_3_1_29_2","first-page":"1047","volume-title":"Proceedings of the International Joint Conference on Artificial Intelligence","author":"Li Zongyi","year":"2024","unstructured":"Zongyi Li, Jianbo Li, Yuxuan Shi, Hefei Ling, Jiazhong Chen, Runsheng Wang, and Shijuan Huang. 2024. Cross-modal generation and alignment via attribute-guided prompt for unsupervised text-based person retrieval. In Proceedings of the International Joint Conference on Artificial Intelligence. International Joint Conferences on Artificial Intelligence Organization, 1047\u20131055."},{"key":"e_1_3_1_30_2","doi-asserted-by":"publisher","DOI":"10.1145\/3632959"},{"key":"e_1_3_1_31_2","first-page":"1527","volume-title":"Proceedings of the AAAI Conference on Artificial Intelligence","volume":"36","author":"Li Zongyi","year":"2022","unstructured":"Zongyi Li, Yuxuan Shi, Hefei Ling, Jiazhong Chen, Qian Wang, and Fengfan Zhou. 2022. Reliability exploration with self-ensemble learning for domain adaptive person re-identification. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 36, 1527\u20131535."},{"key":"e_1_3_1_32_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCSVT.2024.3428589"},{"key":"e_1_3_1_33_2","first-page":"560","volume-title":"Proceedings of the ACM International Conference on Multimedia Retrieval","author":"Liao Shuiying","year":"2023","unstructured":"Shuiying Liao, Yujuan Ding, and P. Y. Mok. 2023. Recommendation of mix-and-match clothing by modeling indirect personal compatibility. In Proceedings of the ACM International Conference on Multimedia Retrieval, 560\u2013564."},{"key":"e_1_3_1_34_2","first-page":"3311","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition","author":"Lin Yen-Liang","year":"2020","unstructured":"Yen-Liang Lin, Son Tran, and Larry S. Davis. 2020. Fashion outfit complementary item retrieval. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 3311\u20133319."},{"key":"e_1_3_1_35_2","first-page":"10562","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition","author":"Lu Zhi","year":"2019","unstructured":"Zhi Lu, Yang Hu, Yunchao Jiang, Yan Chen, and Bing Zeng. 2019. Learning binary code for personalized fashion recommendation. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 10562\u201310570."},{"key":"e_1_3_1_36_2","doi-asserted-by":"publisher","DOI":"10.1145\/3637217"},{"key":"e_1_3_1_37_2","doi-asserted-by":"publisher","DOI":"10.1145\/2766462.2767830"},{"key":"e_1_3_1_38_2","doi-asserted-by":"publisher","DOI":"10.1109\/TMM.2021.3134164"},{"key":"e_1_3_1_39_2","doi-asserted-by":"publisher","DOI":"10.1109\/TKDE.2023.3266423"},{"key":"e_1_3_1_40_2","doi-asserted-by":"publisher","DOI":"10.1145\/3474085.3475537"},{"key":"e_1_3_1_41_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.ipm.2023.103540"},{"key":"e_1_3_1_42_2","first-page":"10373","volume-title":"Proceedings of the IEEE International Conference on Computer Vision","author":"Tan Reuben","year":"2019","unstructured":"Reuben Tan, Mariya I. Vasileva, Kate Saenko, and Bryan A. Plummer. 2019. Learning similarity conditions without explicit supervision. In Proceedings of the IEEE International Conference on Computer Vision, 10373\u201310382."},{"key":"e_1_3_1_43_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-01270-0_24"},{"key":"e_1_3_1_44_2","doi-asserted-by":"publisher","DOI":"10.1109\/JIOT.2018.2871445"},{"key":"e_1_3_1_45_2","first-page":"3522","volume-title":"Proceedings of the IEEE International Conference on Data Engineering.","author":"Wang Hai","year":"2023","unstructured":"Hai Wang, Shuai Wang, Yu Yang, and Desheng Zhang. 2023. GCRL: Efficient delivery area assignment for last-mile logistics with group-based cooperative reinforcement learning. In Proceedings of the IEEE International Conference on Data Engineering. IEEE, 3522\u20133534."},{"key":"e_1_3_1_46_2","first-page":"5696","volume-title":"Proceedings of the ACM International Conference on Multimedia","author":"Wang Longzheng","year":"2023","unstructured":"Longzheng Wang, Chuang Zhang, Hongbo Xu, Yongxiu Xu, Xiaohan Xu, and Siqi Wang. 2023. Cross-modal contrastive learning for multimodal fake news detection. In Proceedings of the ACM International Conference on Multimedia, 5696\u20135704."},{"key":"e_1_3_1_47_2","first-page":"329","volume-title":"Proceedings of the ACM International Conference on Multimedia","author":"Wang Xin","year":"2019","unstructured":"Xin Wang, Bo Wu, and Yueqi Zhong. 2019. Outfit compatibility prediction and diagnosis with multi-layered comparison network. In Proceedings of the ACM International Conference on Multimedia, 329\u2013337."},{"key":"e_1_3_1_48_2","doi-asserted-by":"publisher","DOI":"10.1145\/3343031.3350858"},{"key":"e_1_3_1_49_2","doi-asserted-by":"publisher","DOI":"10.1109\/TIFS.2024.3352415"},{"key":"e_1_3_1_50_2","first-page":"403","volume-title":"Proceedings of the AAAI Conference on Artificial Intelligence","volume":"33","author":"Yang Xun","year":"2019","unstructured":"Xun Yang, Yunshan Ma, Lizi Liao, Meng Wang, and Tat-Seng Chua. 2019. TransNFCM: Translation-based neural fashion compatibility modeling. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 33, 403\u2013410."},{"key":"e_1_3_1_51_2","doi-asserted-by":"publisher","DOI":"10.1145\/3425636"},{"issue":"1","key":"e_1_3_1_52_2","doi-asserted-by":"crossref","first-page":"2306970","DOI":"10.1080\/09540091.2024.2306970","article-title":"A long-short dual-mode knowledge distillation framework for empirical asset pricing models in digital financial networks","volume":"36","author":"Yi Yuanyuan","year":"2024","unstructured":"Yuanyuan Yi, Kai Cui, Minghua Xu, Lingzhi Yi, Kun Yi, Xinlei Zhou, Shenghao Liu, and Gefei Zhou. 2024. A long-short dual-mode knowledge distillation framework for empirical asset pricing models in digital financial networks. Connection Science 36, 1 (2024), 2306970.","journal-title":"Connection Science"},{"issue":"3","key":"e_1_3_1_53_2","doi-asserted-by":"crossref","first-page":"1132","DOI":"10.1109\/TNET.2022.3213913","article-title":"Multiprotocol backscatter with commodity radios for personal IoT sensors","volume":"31","author":"Yuan Longzhi","year":"2022","unstructured":"Longzhi Yuan, Qiwei Wang, Jia Zhao, and Wei Gong. 2022. Multiprotocol backscatter with commodity radios for personal IoT sensors. IEEE\/ACM Transactions on Networking 31, 3 (2022), 1132\u20131144.","journal-title":"IEEE\/ACM Transactions on Networking"},{"key":"e_1_3_1_54_2","doi-asserted-by":"publisher","DOI":"10.1109\/TMM.2023.3293335"},{"key":"e_1_3_1_55_2","doi-asserted-by":"publisher","DOI":"10.1109\/TMM.2021.3059514"},{"key":"e_1_3_1_56_2","doi-asserted-by":"publisher","DOI":"10.1145\/3383184"},{"key":"e_1_3_1_57_2","doi-asserted-by":"publisher","DOI":"10.1145\/3519030"}],"container-title":["ACM Transactions on Multimedia Computing, Communications, and Applications"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3698772","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,11,10]],"date-time":"2025-11-10T14:50:20Z","timestamp":1762786220000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3698772"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,11,10]]},"references-count":56,"journal-issue":{"issue":"11","published-print":{"date-parts":[[2025,11,30]]}},"alternative-id":["10.1145\/3698772"],"URL":"https:\/\/doi.org\/10.1145\/3698772","relation":{},"ISSN":["1551-6857","1551-6865"],"issn-type":[{"value":"1551-6857","type":"print"},{"value":"1551-6865","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,11,10]]},"assertion":[{"value":"2024-04-01","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2024-09-23","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-11-10","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}