{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,8,4]],"date-time":"2026-08-04T15:20:01Z","timestamp":1785856801422,"version":"3.56.0"},"reference-count":61,"publisher":"Association for Computing Machinery (ACM)","issue":"5","funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"crossref","award":["62302080"],"award-info":[{"award-number":["62302080"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]},{"DOI":"10.13039\/501100017691","name":"Guangxi Key Research and Development Program","doi-asserted-by":"crossref","award":["Guike AB24010112"],"award-info":[{"award-number":["Guike AB24010112"]}],"id":[{"id":"10.13039\/501100017691","id-type":"DOI","asserted-by":"crossref"}]},{"name":"National Foreign Expert Project of China","award":["S20240327"],"award-info":[{"award-number":["S20240327"]}]},{"name":"Sichuan Science and Technology Program","award":["2025HJRC0021"],"award-info":[{"award-number":["2025HJRC0021"]}]},{"name":"Sichuan Province Innovative Talent Funding Project for Postdoctoral Fellows","award":["BX202312"],"award-info":[{"award-number":["BX202312"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Multimedia Comput. Commun. Appl."],"published-print":{"date-parts":[[2026,5,31]]},"abstract":"<jats:p>\n                    Text-to-Image Person Retrieval (TIPR) is a cross-modal matching task designed to identify the person images that best correspond to a given textual description. The key difficulty in TIPR is to realize robust correspondence between the textual and visual modalities within a unified latent representation space. To address this challenge, prior approaches incorporate attention mechanisms for implicit cross-modal local alignment. However, they lack the ability to verify whether all local features are correctly aligned. Moreover, existing methods tend to emphasize the utilization of hard negative samples during model optimization to strengthen discrimination between positive and negative pairs, often neglecting incorrectly matched positive pairs. To mitigate these problems, we propose FMFA, a cross-modal Full-Mode Fine-Grained Alignment framework, which enhances global matching through Explicit Fine-Grained Alignment (EFA) and existing implicit relational reasoning\u2014hence the term \u201cfull-mode\u201d\u2014without introducing extra supervisory signals. In particular, we propose an Adaptive Similarity Distribution Matching (A-SDM) module to rectify unmatched positive sample pairs. A-SDM adaptively pulls the unmatched positive pairs closer in the joint embedding space, thereby achieving more precise global alignment. Additionally, we introduce an EFA module, which makes up for the lack of verification capability of implicit relational reasoning. EFA strengthens explicit cross-modal fine-grained interactions by sparsifying the similarity matrix and employs a hard coding method for local alignment. We evaluate our method on three public datasets, where it attains state-of-the-art results among all global matching methods. The code for our method is publicly accessible at\n                    <jats:ext-link xmlns:xlink=\"http:\/\/www.w3.org\/1999\/xlink\" ext-link-type=\"uri\" xlink:href=\"https:\/\/github.com\/yinhao1102\/FMFA\">https:\/\/github.com\/yinhao1102\/FMFA<\/jats:ext-link>\n                    .\n                  <\/jats:p>","DOI":"10.1145\/3786798","type":"journal-article","created":{"date-parts":[[2026,1,3]],"date-time":"2026-01-03T13:39:08Z","timestamp":1767447548000},"page":"1-21","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":2,"title":["Cross-Modal Full-Mode Fine-Grained Alignment for Text-to-Image Person Retrieval"],"prefix":"10.1145","volume":"22","author":[{"ORCID":"https:\/\/orcid.org\/0009-0004-9807-9251","authenticated-orcid":false,"given":"Hao","family":"Yin","sequence":"first","affiliation":[{"name":"Shenzhen Institute for Advanced Study, University of Electronic Science and Technology of China, Shenzhen, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-9684-4424","authenticated-orcid":false,"given":"Xin","family":"Man","sequence":"additional","affiliation":[{"name":"Shenzhen Institute for Advanced Study, University of Electronic Science and Technology of China, Shenzhen, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-0928-6899","authenticated-orcid":false,"given":"Feiyu","family":"Chen","sequence":"additional","affiliation":[{"name":"University of Electronic Science and Technology of China, Chengdu, China  and Sichuan Artificial Intelligence Research Institute, Yibin, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-2615-1555","authenticated-orcid":false,"given":"Jie","family":"Shao","sequence":"additional","affiliation":[{"name":"University of Electronic Science and Technology of China, Chengdu, China  and Sichuan Artificial Intelligence Research Institute, Yibin, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-2999-2088","authenticated-orcid":false,"given":"Heng Tao","family":"Shen","sequence":"additional","affiliation":[{"name":"University of Electronic Science and Technology of China, Chengdu, China  and Sichuan Artificial Intelligence Research Institute, Yibin, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2026,4,20]]},"reference":[{"key":"e_1_3_1_2_2","unstructured":"Jinze Bai Shuai Bai Shusheng Yang Shijie Wang Sinan Tan Peng Wang Junyang Lin Chang Zhou and Jingren Zhou. 2023. Qwen-VL: A versatile vision-language model for understanding localization text reading and beyond. arXiv:2308.12966. Retrieved from https:\/\/arxiv.org\/abs\/2308.12966"},{"key":"e_1_3_1_3_2","doi-asserted-by":"publisher","DOI":"10.24963\/ijcai.2023\/62"},{"key":"e_1_3_1_4_2","doi-asserted-by":"publisher","DOI":"10.5555\/3692070.3692230"},{"key":"e_1_3_1_5_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v38i1.27801"},{"key":"e_1_3_1_6_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52733.2024.02549"},{"key":"e_1_3_1_7_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52729.2023.01060"},{"key":"e_1_3_1_8_2","doi-asserted-by":"publisher","DOI":"10.1007\/s11263-025-02461-z"},{"key":"e_1_3_1_9_2","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2022.3169693"},{"key":"e_1_3_1_10_2","first-page":"4171","volume-title":"Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies","author":"Devlin Jacob","year":"2019","unstructured":"Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), 4171\u20134186."},{"key":"e_1_3_1_11_2","unstructured":"Zefeng Ding Changxing Ding Zhiyin Shao and Dacheng Tao. 2021. Semantically self-aligned network for text-to-image part-aware person re-identification. arXiv:2107.12666. Retrieved from https:\/\/arxiv.org\/abs\/2107.12666"},{"key":"e_1_3_1_12_2","doi-asserted-by":"publisher","DOI":"10.1145\/3721482"},{"key":"e_1_3_1_13_2","unstructured":"Chenyang Gao Guanyu Cai Xinyang Jiang Feng Zheng Jun Zhang Yifei Gong Pai Peng Xiaowei Guo and Xing Sun. 2021. Contextual non-local alignment over full-scale representation for text-based person search. arXiv:2101.03036. Retrieved from https:\/\/arxiv.org\/abs\/2101.03036"},{"key":"e_1_3_1_14_2","doi-asserted-by":"crossref","unstructured":"Xiao Han Sen He Li Zhang and Tao Xiang. 2021. Text-based person search with limited data. arXiv:2110.10807. Retrieved from https:\/\/arxiv.org\/abs\/2110.10807","DOI":"10.5244\/C.35.10"},{"key":"e_1_3_1_15_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52688.2022.01553"},{"key":"e_1_3_1_16_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.90"},{"key":"e_1_3_1_17_2","doi-asserted-by":"publisher","DOI":"10.1007\/s11263-024-02322-1"},{"key":"e_1_3_1_18_2","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2023.3337653"},{"key":"e_1_3_1_19_2","first-page":"18409","volume-title":"Proceedings of the 41st International Conference on Machine Learning","author":"Hoffmann David T.","year":"2024","unstructured":"David T. Hoffmann, Simon Schrodi, Jelena Bratuli\u0107, Nadine Behrmann, Volker Fischer, and Thomas Brox. 2024. Eureka-moments in transformers: Multi-step tasks reveal Softmax induced optimization problems. In Proceedings of the 41st International Conference on Machine Learning, 18409\u201318438."},{"key":"e_1_3_1_20_2","doi-asserted-by":"publisher","DOI":"10.1145\/3747185"},{"key":"e_1_3_1_21_2","doi-asserted-by":"publisher","DOI":"10.1109\/TMM.2022.3225754"},{"key":"e_1_3_1_22_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52729.2023.00273"},{"key":"e_1_3_1_23_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52734.2025.00861"},{"key":"e_1_3_1_24_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-01225-0_13"},{"key":"e_1_3_1_25_2","first-page":"19730","volume-title":"Proceedings of the 40th International Conference on Machine Learning","author":"Li Junnan","year":"2023","unstructured":"Junnan Li, Dongxu Li, Silvio Savarese, and Steven Hoi. 2023. BLIP-2: Bootstrapping language-image pre-training with frozen image encoders and large language models. In Proceedings of the 40th International Conference on Machine Learning, 19730\u201319742."},{"key":"e_1_3_1_26_2","first-page":"12888","volume-title":"Proceedings of the 39th International Conference on Machine Learning","author":"Li Junnan","year":"2022","unstructured":"Junnan Li, Dongxu Li, Caiming Xiong, and Steven Hoi. 2022. BLIP: Bootstrapping language-image pre-training for unified vision-language understanding and generation. In Proceedings of the 39th International Conference on Machine Learning, 12888\u201312900."},{"key":"e_1_3_1_27_2","first-page":"9694","article-title":"Align before fuse: Vision and language representation learning with momentum distillation","volume":"34","author":"Li Junnan","year":"2021","unstructured":"Junnan Li, Ramprasaath Selvaraju, Akhilesh Gotmare, Shafiq Joty, Caiming Xiong, and Steven Chu Hong Hoi. 2021. Align before fuse: Vision and language representation learning with momentum distillation. In Proceedings of the Advances in Neural Information Processing Systems, Vol. 34, 9694\u20139705.","journal-title":"Proceedings of the Advances in Neural Information Processing Systems"},{"key":"e_1_3_1_28_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.551"},{"key":"e_1_3_1_29_2","unstructured":"Haotian Liu Chunyuan Li Yuheng Li Bo Li Yuanhan Zhang Sheng Shen and Yong Jae Lee. 2024. LLaVA-NeXT: Improved reasoning OCR and world knowledge. Retrieved from https:\/\/llava-vl.github.io\/blog\/2024-01-30-llava-next\/"},{"key":"e_1_3_1_30_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v39i6.32608"},{"key":"e_1_3_1_31_2","first-page":"11525","article-title":"Object-centric learning with slot attention","volume":"33","author":"Locatello Francesco","year":"2020","unstructured":"Francesco Locatello, Dirk Weissenborn, Thomas Unterthiner, Aravindh Mahendran, Georg Heigold, Jakob Uszkoreit, Alexey Dosovitskiy, and Thomas Kipf. 2020. Object-centric learning with slot attention. In Proceedings of the Advances in Neural Information Processing Systems, Vol. 33, 11525\u201311538.","journal-title":"Proceedings of the Advances in Neural Information Processing Systems"},{"key":"e_1_3_1_32_2","first-page":"13","article-title":"ViLBERT: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks","volume":"32","author":"Lu Jiasen","year":"2019","unstructured":"Jiasen Lu, Dhruv Batra, Devi Parikh, and Stefan Lee. 2019. ViLBERT: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks. In Proceedings of the Advances in Neural Information Processing Systems, Vol. 32, 13\u201323.","journal-title":"Proceedings of the Advances in Neural Information Processing Systems"},{"key":"e_1_3_1_33_2","doi-asserted-by":"publisher","DOI":"10.1109\/TIFS.2025.3586488"},{"key":"e_1_3_1_34_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52729.2023.01860"},{"key":"e_1_3_1_35_2","unstructured":"Aaron van den Oord Yazhe Li and Oriol Vinyals. 2018. Representation learning with contrastive predictive coding. arXiv:1807.03748. Retrieved from https:\/\/arxiv.org\/abs\/1807.03748"},{"key":"e_1_3_1_36_2","first-page":"474","volume-title":"Proceedings of the European Conference on Computer Vision","author":"Park Jicheol","year":"2024","unstructured":"Jicheol Park, Dongwon Kim, Boseung Jeong, and Suha Kwak. 2024. PLOT: Text-based person search with part slot attention for corresponding part discovery. In Proceedings of the European Conference on Computer Vision, 474\u2013490."},{"key":"e_1_3_1_37_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52733.2024.02568"},{"key":"e_1_3_1_38_2","first-page":"8748","volume-title":"Proceedings of the 38th International Conference on Machine Learning","author":"Radford Alec","year":"2021","unstructured":"Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever. 2021. Learning transferable visual models from natural language supervision. In Proceedings of the 38th International Conference on Machine Learning, 8748\u20138763."},{"key":"e_1_3_1_39_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52688.2022.01042"},{"key":"e_1_3_1_40_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2015.7298682"},{"key":"e_1_3_1_41_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/P16-1162"},{"key":"e_1_3_1_42_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV51070.2023.01026"},{"key":"e_1_3_1_43_2","doi-asserted-by":"publisher","DOI":"10.1145\/3503161.3548028"},{"key":"e_1_3_1_44_2","first-page":"624","volume-title":"Proceedings of the European Conference on Computer Vision","author":"Shu Xiujun","year":"2022","unstructured":"Xiujun Shu, Wei Wen, Haoqian Wu, Keyu Chen, Yiran Song, Ruizhi Qiao, Bo Ren, and Xiao Wang. 2022. See finer, see more: Implicit modality alignment for text-based person retrieval. In Proceedings of the European Conference on Computer Vision, 624\u2013641."},{"key":"e_1_3_1_45_2","first-page":"447","volume-title":"Proceedings of the European Conference on Computer Vision","author":"Sogi Naoya","year":"2024","unstructured":"Naoya Sogi, Takashi Shibata, and Makoto Terao. 2024. Object-aware query perturbation for cross-modal image-text retrieval. In Proceedings of the European Conference on Computer Vision, 447\u2013464."},{"key":"e_1_3_1_46_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52733.2024.01621"},{"key":"e_1_3_1_47_2","doi-asserted-by":"publisher","DOI":"10.1177\/107769905303000401"},{"key":"e_1_3_1_48_2","doi-asserted-by":"publisher","DOI":"10.1145\/3123266.3123326"},{"key":"e_1_3_1_49_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/D19-1098"},{"key":"e_1_3_1_50_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-58610-2_24"},{"key":"e_1_3_1_51_2","doi-asserted-by":"publisher","DOI":"10.1109\/TNNLS.2024.3381347"},{"key":"e_1_3_1_52_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCSVT.2024.3392831"},{"key":"e_1_3_1_53_2","doi-asserted-by":"publisher","DOI":"10.1145\/3686160"},{"key":"e_1_3_1_54_2","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2023.3327924"},{"key":"e_1_3_1_55_2","doi-asserted-by":"publisher","DOI":"10.1109\/TNNLS.2023.3310118"},{"key":"e_1_3_1_56_2","doi-asserted-by":"publisher","DOI":"10.1145\/3581783.3611709"},{"key":"e_1_3_1_57_2","first-page":"135043","article-title":"Feature-level adversarial attacks and ranking disruption for visible-infrared person re-identification","volume":"37","author":"Yang Xi","year":"2024","unstructured":"Xi Yang, Huanling Liu, De Cheng, Nannan Wang, and Xinbo Gao. 2024. Feature-level adversarial attacks and ranking disruption for visible-infrared person re-identification. In Proceedings of the Advances in Neural Information Processing Systems, Vol. 37, 135043\u2013135061.","journal-title":"Proceedings of the Advances in Neural Information Processing Systems"},{"key":"e_1_3_1_58_2","volume-title":"Proceedings of the International Conference on Learning Representations","author":"Yao Lewei","year":"2022","unstructured":"Lewei Yao, Runhui Huang, Lu Hou, Guansong Lu, Minzhe Niu, Hang Xu, Xiaodan Liang, Zhenguo Li, Xin Jiang, and Chunjing Xu. 2022. FILIP: Fine-grained interactive language-image pre-training. In Proceedings of the International Conference on Learning Representations."},{"key":"e_1_3_1_59_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-01246-5_42"},{"key":"e_1_3_1_60_2","doi-asserted-by":"publisher","DOI":"10.1145\/3383184"},{"key":"e_1_3_1_61_2","doi-asserted-by":"publisher","DOI":"10.1145\/3474085.3475369"},{"key":"e_1_3_1_62_2","first-page":"45666","volume-title":"Proceedings of the Advances in Neural Information Processing Systems","volume":"37","author":"Zuo Jialong","year":"2024","unstructured":"Jialong Zuo, Jiahao Hong, Feng Zhang, Changqian Yu, Hanyu Zhou, Changxin Gao, Nong Sang, and Jingdong Wang. 2024. PLIP: Language-image pre-training for person representation learning. In Proceedings of the Advances in Neural Information Processing Systems, Vol. 37, 45666\u201345702."}],"container-title":["ACM Transactions on Multimedia Computing, Communications, and Applications"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3786798","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,5,21]],"date-time":"2026-05-21T13:49:24Z","timestamp":1779371364000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3786798"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,4,20]]},"references-count":61,"journal-issue":{"issue":"5","published-print":{"date-parts":[[2026,5,31]]}},"alternative-id":["10.1145\/3786798"],"URL":"https:\/\/doi.org\/10.1145\/3786798","relation":{},"ISSN":["1551-6857","1551-6865"],"issn-type":[{"value":"1551-6857","type":"print"},{"value":"1551-6865","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,4,20]]},"assertion":[{"value":"2025-08-15","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-12-13","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2026-04-20","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}