{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,18]],"date-time":"2026-06-18T16:02:18Z","timestamp":1781798538894,"version":"3.54.5"},"reference-count":71,"publisher":"Association for Computing Machinery (ACM)","issue":"10","license":[{"start":{"date-parts":[[2024,10,29]],"date-time":"2024-10-29T00:00:00Z","timestamp":1730160000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"crossref","award":["62325206, 61936005, 62301276"],"award-info":[{"award-number":["62325206, 61936005, 62301276"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]},{"name":"Key Research and Development Program of Jiangsu Province","award":["BE2023016-4"],"award-info":[{"award-number":["BE2023016-4"]}]},{"name":"Opening Foundation of Key Laboratory of Computer Vision and System, Ministry of Education, Tianjin University of Technology, China."}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Multimedia Comput. Commun. Appl."],"published-print":{"date-parts":[[2024,10,31]]},"abstract":"<jats:p>\n            Visible-infrared person re-identification (VI-ReID) aims to match individuals across different modalities. Existing methods can learn class-separable features but still struggle with modality gaps within class due to the modality-specific information, which is discriminative in one modality but not present in another (e.g., a black striped shirt). The presence of the interfering information creates a spurious correlation with the class label, which hinders alignment across modalities. To this end, we propose an Unbiased feature learning method based on Causal inTervention for VI-ReID from three aspects. Firstly, through the proposed structural causal graph, we demonstrate that modality-specific information acts as a confounder that restricts the intra-class feature alignment. Secondly, we propose a causal intervention method to remove the confounder using an effective approximation of backdoor adjustment, which involves adjusting the spurious correlation between features and labels. Thirdly, we incorporate the proposed approximation method into the basic VI-ReID model. Specifically, the confounder can be removed by adjusting the extracted features with a set of weighted pre-trained class prototypes from different modalities, where the weight is adapted based on the features. Extensive experiments on the SYSU-MM01 and RegDB datasets demonstrate that our method outperforms state-of-the-art methods. Code is available at\n            <jats:ext-link xmlns:xlink=\"http:\/\/www.w3.org\/1999\/xlink\" ext-link-type=\"uri\" xlink:href=\"https:\/\/github.com\/NJUPT-MCC\/UCT\">https:\/\/github.com\/NJUPT-MCC\/UCT<\/jats:ext-link>\n            .\n          <\/jats:p>","DOI":"10.1145\/3674737","type":"journal-article","created":{"date-parts":[[2024,6,27]],"date-time":"2024-06-27T19:53:12Z","timestamp":1719517992000},"page":"1-20","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":5,"title":["Unbiased Feature Learning with Causal Intervention for Visible-Infrared Person Re-Identification"],"prefix":"10.1145","volume":"20","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-8051-3070","authenticated-orcid":false,"given":"Bowen","family":"Yuan","sequence":"first","affiliation":[{"name":"Nanjing University of Posts and Telecommunications, Nanjing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0006-5614-7003","authenticated-orcid":false,"given":"Jiahao","family":"Lu","sequence":"additional","affiliation":[{"name":"Nanjing University of Posts and Telecommunications, Nanjing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0004-2173-1694","authenticated-orcid":false,"given":"Sisi","family":"You","sequence":"additional","affiliation":[{"name":"Nanjing University of Posts and Telecommunications, Nanjing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-5956-831X","authenticated-orcid":false,"given":"Bing-Kun","family":"Bao","sequence":"additional","affiliation":[{"name":"Nanjing University of Posts and Telecommunications, Nanjing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2024,10,29]]},"reference":[{"key":"e_1_3_1_2_2","doi-asserted-by":"publisher","unstructured":"Yoshua Bengio Tristan Deleu Nasim Rahaman Rosemary Ke S\u00e9bastien Lachapelle Olexa Bilaniuk Anirudh Goyal and Christopher Pal. 2019. A meta-transfer objective for learning to disentangle causal mechanisms. arXiv:1901.10912. DOI: 10.48550\/arXiv.1901.10912.","DOI":"10.48550\/arXiv.1901.10912"},{"key":"e_1_3_1_3_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-01264-9_9"},{"key":"e_1_3_1_4_2","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2022.3141868"},{"key":"e_1_3_1_5_2","doi-asserted-by":"crossref","unstructured":"Long Chen Yuhang Zheng Yulei Niu Hanwang Zhang and Jun Xiao. 2021. Counterfactual samples synthesizing and training for robust visual question answering. arXiv:2110.01013.","DOI":"10.1109\/CVPR42600.2020.01081"},{"issue":"6","key":"e_1_3_1_6_2","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/3595183","article-title":"Identity featurCe disentanglement for visible-infrared person re-identification","volume":"19","author":"Chen Xiumei","year":"2023","unstructured":"Xiumei Chen, Xiangtao Zheng, and Xiaoqiang Lu. 2023. Identity featurCe disentanglement for visible-infrared person re-identification. ACM Transactions on Multimedia Computing, Communications and Applications 19, 6 (2023), 1\u201320.","journal-title":"ACM Transactions on Multimedia Computing, Communications and Applications"},{"key":"e_1_3_1_7_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR46437.2021.00065"},{"key":"e_1_3_1_8_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52729.2023.02179"},{"key":"e_1_3_1_9_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV48922.2021.01161"},{"key":"e_1_3_1_10_2","doi-asserted-by":"publisher","DOI":"10.1145\/3474085.3475643"},{"key":"e_1_3_1_11_2","volume-title":"Causal Inference in Statistics: A Primer","author":"Glymour Madelyn","year":"2016","unstructured":"Madelyn Glymour, Judea Pearl, and Nicholas P Jewell. 2016. Causal Inference in Statistics: A Primer. John Wiley & Sons."},{"key":"e_1_3_1_12_2","doi-asserted-by":"publisher","DOI":"10.1145\/3617375"},{"key":"e_1_3_1_13_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCSVT.2022.3147813"},{"key":"e_1_3_1_14_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v36i1.19987"},{"key":"e_1_3_1_15_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-031-19781-9_28"},{"key":"e_1_3_1_16_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.00321"},{"key":"e_1_3_1_17_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCSVT.2019.2963721"},{"key":"e_1_3_1_18_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52729.2023.01786"},{"issue":"1","key":"e_1_3_1_19_2","first-page":"1","article-title":"Dynamic weighted gradient reversal network for visible-infrared person re-identification","volume":"20","author":"Li Chenghua","year":"2023","unstructured":"Chenghua Li, Zongze Li, Jing Sun, Yun Zhang, Xiaoping Jiang, and Fan Zhang. 2023. Dynamic weighted gradient reversal network for visible-infrared person re-identification. ACM Transactions on Multimedia Computing, Communications and Applications 20, 1 (2023), 1\u201323.","journal-title":"ACM Transactions on Multimedia Computing, Communications and Applications"},{"key":"e_1_3_1_20_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v34i04.5891"},{"key":"e_1_3_1_21_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCSVT.2023.3340225"},{"key":"e_1_3_1_22_2","doi-asserted-by":"publisher","DOI":"10.1145\/3503161.3548035"},{"key":"e_1_3_1_23_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52688.2022.00294"},{"key":"e_1_3_1_24_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52688.2022.01751"},{"key":"e_1_3_1_25_2","doi-asserted-by":"publisher","DOI":"10.1109\/TMM.2020.3042080"},{"key":"e_1_3_1_26_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCSVT.2022.3168999"},{"key":"e_1_3_1_27_2","first-page":"1835","volume-title":"Proceedings of the AAAI Conference on Artificial Intelligence (AAAI \u201923)","author":"Lu Hu","year":"2022","unstructured":"Hu Lu, Xuezhang Zou, and Pingping Zhang. 2022. Learning progressive modality-shared transformers for effective visible-infrared person re-identification. In Proceedings of the AAAI Conference on Artificial Intelligence (AAAI \u201923). 1835\u20131843."},{"key":"e_1_3_1_28_2","first-page":"1","volume-title":"Proceedings of the Advances in Neural Information Processing Systems","volume":"31","author":"Magliacane Sara","year":"2018","unstructured":"Sara Magliacane, Thijs Van Ommen, Tom Claassen, Stephan Bongers, Philip Versteeg, and Joris M. Mooij. 2018. Domain adaptation by using causal inference to predict invariant conditional distributions. In Proceedings of the Advances in Neural Information Processing Systems, Vol. 31. 1\u201311."},{"key":"e_1_3_1_29_2","doi-asserted-by":"publisher","DOI":"10.1017\/S0266466603004109"},{"key":"e_1_3_1_30_2","doi-asserted-by":"publisher","DOI":"10.3390\/s17030605"},{"key":"e_1_3_1_31_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR46437.2021.01251"},{"key":"e_1_3_1_32_2","volume-title":"The Book OF Why: The New Science of Cause and Effect","author":"Pearl Judea","year":"2018","unstructured":"Judea Pearl and Dana Mackenzie. 2018. The Book OF Why: The New Science of Cause and Effect. Basic books."},{"key":"e_1_3_1_33_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.01087"},{"key":"e_1_3_1_34_2","doi-asserted-by":"publisher","DOI":"10.1145\/3534678.3539242"},{"key":"e_1_3_1_35_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2017.74"},{"key":"e_1_3_1_36_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.00081"},{"key":"e_1_3_1_37_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.00643"},{"key":"e_1_3_1_38_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-01225-0_30"},{"issue":"11","key":"e_1_3_1_39_2","first-page":"2579","article-title":"Visualizing data using t-SNE","volume":"9","author":"Maaten Laurens Van der","year":"2008","unstructured":"Laurens Van der Maaten and Geoffrey Hinton. 2008. Visualizing data using t-SNE. Journal of Machine Learning Research 9, 11 (2008). 2579\u20132605.","journal-title":"Journal of Machine Learning Research"},{"key":"e_1_3_1_40_2","first-page":"16451","volume-title":"Proceedings of the 35th International Conference on Neural Information Processing Systems","volume":"34","author":"K\u00fcgelgen Julius Von","year":"2021","unstructured":"Julius Von K\u00fcgelgen, Yash Sharma, Luigi Gresele, Wieland Brendel, Bernhard Sch\u00f6lkopf, Michel Besserve, and Francesco Locatello. 2021. Self-supervised learning with data augmentations provably isolates content from style. In Proceedings of the 35th International Conference on Neural Information Processing Systems, Vol.34. 16451\u201316467."},{"key":"e_1_3_1_41_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2019.00372"},{"key":"e_1_3_1_42_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.neucom.2021.08.053"},{"key":"e_1_3_1_43_2","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2022.3213193"},{"key":"e_1_3_1_44_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.01077"},{"issue":"3","key":"e_1_3_1_45_2","first-page":"3933","article-title":"Weakly-supervised video object grounding via causal intervention","volume":"45","author":"Wang Wei","year":"2023","unstructured":"Wei Wang, Junyu Gao, and Changsheng Xu. 2023. Weakly-supervised video object grounding via causal intervention. IEEE Transactions on Pattern Analysis and Machine Intelligence 45, 3 (2023), 3933\u20133948.","journal-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence"},{"key":"e_1_3_1_46_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR46437.2021.00304"},{"key":"e_1_3_1_47_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.00071"},{"key":"e_1_3_1_48_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2017.575"},{"key":"e_1_3_1_49_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV51070.2023.01021"},{"key":"e_1_3_1_50_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR46437.2021.00431"},{"key":"e_1_3_1_51_2","first-page":"2048","volume-title":"Proceedings of the International Conference on Machine Learning","author":"Xu Kelvin","year":"2015","unstructured":"Kelvin Xu, Jimmy Ba, Ryan Kiros, Kyunghyun Cho, Aaron Courville, Ruslan Salakhudinov, Rich Zemel, and Yoshua Bengio. 2015. Show, attend and tell: Neural image caption generation with visual attention. In Proceedings of the International Conference on Machine Learning. PMLR, 2048\u20132057."},{"issue":"11","key":"e_1_3_1_52_2","first-page":"12996","article-title":"Deconfounded image captioning: A causal retrospect","volume":"45","author":"Yang Xu","year":"2021","unstructured":"Xu Yang, Hanwang Zhang, and Jianfei Cai. 2021. Deconfounded image captioning: A causal retrospect. IEEE Transactions on Pattern Analysis and Machine Intelligence 45, 11 (2021), 12996\u201313010.","journal-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence"},{"key":"e_1_3_1_53_2","doi-asserted-by":"publisher","DOI":"10.1145\/3343031.3351043"},{"key":"e_1_3_1_54_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v32i1.12293"},{"key":"e_1_3_1_55_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-58520-4_14"},{"key":"e_1_3_1_56_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2021.3054775"},{"key":"e_1_3_1_57_2","doi-asserted-by":"publisher","DOI":"10.1109\/TIFS.2020.3001665"},{"key":"e_1_3_1_58_2","doi-asserted-by":"publisher","DOI":"10.24963\/ijcai.2018\/152"},{"key":"e_1_3_1_59_2","doi-asserted-by":"publisher","DOI":"10.1145\/3614434"},{"key":"e_1_3_1_60_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV51070.2023.01027"},{"key":"e_1_3_1_61_2","doi-asserted-by":"publisher","DOI":"10.1007\/s00530-021-00872-9"},{"key":"e_1_3_1_62_2","first-page":"2734","volume-title":"Proceedings of the Advances in Neural Information Processing Systems","volume":"33","author":"Yue Zhongqi","year":"2020","unstructured":"Zhongqi Yue, Hanwang Zhang, Qianru Sun, and Xian-Sheng Hua. 2020. Interventional few-shot learning. In Proceedings of the Advances in Neural Information Processing Systems, Vol. 33. 2734\u20132746."},{"key":"e_1_3_1_63_2","first-page":"12310","volume-title":"Proceedings of the International Conference on Machine Learning","author":"Zbontar Jure","year":"2021","unstructured":"Jure Zbontar, Li Jing, Ishan Misra, Yann LeCun, and St\u00e9phane Deny. 2021. Barlow twins: Self-supervised learning via redundancy reduction. In Proceedings of the International Conference on Machine Learning. PMLR, 12310\u201312320."},{"key":"e_1_3_1_64_2","first-page":"655","volume-title":"Proceedings of the Advances in Neural Information Processing Systems","volume":"33","author":"Zhang Dong","year":"2020","unstructured":"Dong Zhang, Hanwang Zhang, Jinhui Tang, Xian-Sheng Hua, and Qianru Sun. 2020. Causal intervention for weakly-supervised semantic segmentation. In Proceedings of the Advances in Neural Information Processing Systems, Vol. 33. 655\u2013666."},{"key":"e_1_3_1_65_2","doi-asserted-by":"publisher","DOI":"10.1145\/3473341"},{"key":"e_1_3_1_66_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52688.2022.00720"},{"key":"e_1_3_1_67_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52729.2023.00214"},{"key":"e_1_3_1_68_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-031-19781-9_27"},{"key":"e_1_3_1_69_2","doi-asserted-by":"publisher","DOI":"10.1109\/TMM.2022.3163847"},{"key":"e_1_3_1_70_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.imavis.2023.104791"},{"key":"e_1_3_1_71_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.00224"},{"key":"e_1_3_1_72_2","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2020.3004267"}],"container-title":["ACM Transactions on Multimedia Computing, Communications, and Applications"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3674737","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3674737","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,19]],"date-time":"2025-06-19T00:57:50Z","timestamp":1750294670000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3674737"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,10,29]]},"references-count":71,"journal-issue":{"issue":"10","published-print":{"date-parts":[[2024,10,31]]}},"alternative-id":["10.1145\/3674737"],"URL":"https:\/\/doi.org\/10.1145\/3674737","relation":{},"ISSN":["1551-6857","1551-6865"],"issn-type":[{"value":"1551-6857","type":"print"},{"value":"1551-6865","type":"electronic"}],"subject":[],"published":{"date-parts":[[2024,10,29]]},"assertion":[{"value":"2023-12-27","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2024-06-13","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2024-10-29","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}