{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,17]],"date-time":"2026-07-17T23:17:54Z","timestamp":1784330274928,"version":"3.55.0"},"reference-count":57,"publisher":"Association for Computing Machinery (ACM)","issue":"5","license":[{"start":{"date-parts":[[2024,1,11]],"date-time":"2024-01-11T00:00:00Z","timestamp":1704931200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"crossref","award":["62172263 and 62002209"],"award-info":[{"award-number":["62172263 and 62002209"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]},{"DOI":"10.13039\/501100007129","name":"Natural Science Foundation of Shandong Province","doi-asserted-by":"crossref","award":["ZR2020QF042, ZR2020YQ47, and ZR2020QF111"],"award-info":[{"award-number":["ZR2020QF042, ZR2020YQ47, and ZR2020QF111"]}],"id":[{"id":"10.13039\/501100007129","id-type":"DOI","asserted-by":"crossref"}]},{"name":"Taishan Scholar Foundation of Shandong Province","award":["ts20190924"],"award-info":[{"award-number":["ts20190924"]}]},{"name":"CCF-Baidu Open Fund","award":["CCF-BAIDU OF2022008"],"award-info":[{"award-number":["CCF-BAIDU OF2022008"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Multimedia Comput. Commun. Appl."],"published-print":{"date-parts":[[2024,5,31]]},"abstract":"<jats:p>\n            Most cross-modal retrieval methods assume the multi-modal training data is complete and has a one-to-one correspondence. However, in the real world, multi-modal data generally suffers from missing modality information due to the uncertainty of data collection and storage processes, which limits the practical application of existing cross-modal retrieval methods. Although some solutions have been proposed to generate the missing modality data using a single pseudo sample, this may lead to incomplete semantic restoration and sub-optimal retrieval results due to the limited semantic information it provides. To address this challenge, this article proposes an Incomplete Cross-Modal Retrieval with Deep Correlation Transfer (ICMR-DCT) method that can robustly model incomplete multi-modal data and dynamically capture the adjacency semantic correlation for cross-modal retrieval. Specifically, we construct intra-modal graph attention-based auto-encoder to learn modality-invariant representations by performing semantic reconstruction through intra-modality adjacency correlation mining. Then, we design dual cross-modal alignment constraints to project multi-modal representations into a common semantic space, thus bridging the heterogeneous modality gap and enhancing the discriminability of the common representation. We further introduce semantic preservation to enhance adjacency semantic information and achieve cross-modal semantic correlation. Moreover, we propose a nearest-neighbor weighting integration strategy with cross-modal correlation transfer to generate the missing modality data according to inter-modality mapping relations and adjacency correlations between each sample and its neighbors, which improves the robustness of our method against incomplete multi-modal training data. Extensive experiments on three widely tested benchmark datasets demonstrate the superior performance of our method in cross-modal retrieval tasks under both complete and incomplete retrieval scenarios. Our used datasets and source codes are available at\n            <jats:ext-link xmlns:xlink=\"http:\/\/www.w3.org\/1999\/xlink\" ext-link-type=\"url\" xlink:href=\"https:\/\/github.com\/shidan0122\/DCT.git\">https:\/\/github.com\/shidan0122\/DCT.git<\/jats:ext-link>\n            .\n          <\/jats:p>","DOI":"10.1145\/3637442","type":"journal-article","created":{"date-parts":[[2023,12,13]],"date-time":"2023-12-13T11:42:50Z","timestamp":1702467770000},"page":"1-21","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":18,"title":["Incomplete Cross-Modal Retrieval with Deep Correlation Transfer"],"prefix":"10.1145","volume":"20","author":[{"ORCID":"https:\/\/orcid.org\/0000-0003-2773-9924","authenticated-orcid":false,"given":"Dan","family":"Shi","sequence":"first","affiliation":[{"name":"Shandong Normal University, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-2993-7142","authenticated-orcid":false,"given":"Lei","family":"Zhu","sequence":"additional","affiliation":[{"name":"Shandong Normal University, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-5504-2529","authenticated-orcid":false,"given":"Jingjing","family":"Li","sequence":"additional","affiliation":[{"name":"University of Electronic Science and Technology of China, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-1955-3804","authenticated-orcid":false,"given":"Guohua","family":"Dong","sequence":"additional","affiliation":[{"name":"Institute of Basic Medical Sciences, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-6259-7533","authenticated-orcid":false,"given":"Huaxiang","family":"Zhang","sequence":"additional","affiliation":[{"name":"Shandong Normal University, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2024,1,11]]},"reference":[{"key":"e_1_3_1_2_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.inffus.2021.10.013"},{"key":"e_1_3_1_3_2","first-page":"1247","volume-title":"Proceedings of ICML","volume":"28","author":"Andrew Galen","year":"2013","unstructured":"Galen Andrew, Raman Arora, Jeff A. Bilmes, and Karen Livescu. 2013. Deep canonical correlation analysis. In Proceedings of ICML, Vol. 28. 1247\u20131255."},{"key":"e_1_3_1_4_2","volume-title":"Proceedings of ICLR","author":"Bruna Joan","year":"2014","unstructured":"Joan Bruna, Wojciech Zaremba, Arthur Szlam, and Yann LeCun. 2014. Spectral networks and locally connected networks on graphs. In Proceedings of ICLR."},{"key":"e_1_3_1_5_2","first-page":"5410","volume-title":"Proceedings of IJCAI","author":"Cao Min","year":"2022","unstructured":"Min Cao, Shiping Li, Juntao Li, Liqiang Nie, and Min Zhang. 2022. Image-text retrieval: A survey on recent research and development. In Proceedings of IJCAI. 5410\u20135417."},{"key":"e_1_3_1_6_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2022.3165024"},{"key":"e_1_3_1_7_2","first-page":"1","volume-title":"Proceedings of IJCNN","author":"Chen Dong","year":"2020","unstructured":"Dong Chen, Miaomiao Cheng, Chen Min, and Liping Jing. 2020. Unsupervised deep imputed hashing for partial cross-modal retrieval. In Proceedings of IJCNN. 1\u20138."},{"issue":"4","key":"e_1_3_1_8_2","first-page":"Article 95, 23","article-title":"Cross-modal graph matching network for image-text retrieval","volume":"18","author":"Cheng Yuhao","year":"2022","unstructured":"Yuhao Cheng, Xiaoguang Zhu, Jiuchao Qian, Fei Wen, and Peilin Liu. 2022. Cross-modal graph matching network for image-text retrieval. ACM Trans. Multim. Comput. Commun. Appl. 18, 4 (2022), Article 95, 23 pages.","journal-title":"ACM Trans. Multim. Comput. Commun. Appl."},{"key":"e_1_3_1_9_2","first-page":"1","volume-title":"Proceedings of CIVR","author":"Chua Tat-Seng","year":"2009","unstructured":"Tat-Seng Chua, Jinhui Tang, Richang Hong, Haojie Li, Zhiping Luo, and Yantao Zheng. 2009. NUS-WIDE: A real-world web image database from National University of Singapore. In Proceedings of CIVR. 1\u20139."},{"key":"e_1_3_1_10_2","article-title":"Empirical evaluation of gated recurrent neural networks on sequence modeling","volume":"1412","author":"Chung Junyoung","year":"2014","unstructured":"Junyoung Chung, \u00c7aglar G\u00fcl\u00e7ehre, KyungHyun Cho, and Yoshua Bengio. 2014. Empirical evaluation of gated recurrent neural networks on sequence modeling. CoRR abs\/1412.3555 (2014).","journal-title":"CoRR"},{"key":"e_1_3_1_11_2","doi-asserted-by":"publisher","DOI":"10.1109\/TMM.2015.2508146"},{"issue":"2","key":"e_1_3_1_12_2","first-page":"Article 43, 22","article-title":"From selective deep convolutional features to compact binary representations for image retrieval","volume":"15","author":"Do Thanh-Toan","year":"2019","unstructured":"Thanh-Toan Do, Tuan Hoang, Dang-Khoa Le Tan, Huu Le, Tam V. Nguyen, and Ngai-Man Cheung. 2019. From selective deep convolutional features to compact binary representations for image retrieval. ACM Trans. Multim. Comput. Commun. Appl. 15, 2 (2019), Article 43, 22 pages.","journal-title":"ACM Trans. Multim. Comput. Commun. Appl."},{"key":"e_1_3_1_13_2","first-page":"7","volume-title":"Proceedings of ACM MM","author":"Feng Fangxiang","year":"2014","unstructured":"Fangxiang Feng, Xiaojie Wang, and Ruifan Li. 2014. Cross-modal retrieval with correspondence autoencoder. In Proceedings of ACM MM. 7\u201316."},{"key":"e_1_3_1_14_2","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2019.2941858"},{"key":"e_1_3_1_15_2","doi-asserted-by":"publisher","DOI":"10.1162\/0899766042321814"},{"key":"e_1_3_1_16_2","first-page":"594","volume-title":"Proceedings of SIGKDD","author":"Hou Zhenyu","year":"2022","unstructured":"Zhenyu Hou, Xiao Liu, Yukuo Cen, Yuxiao Dong, Hongxia Yang, Chunjie Wang, and Jie Tang. 2022. GraphMAE: Self-supervised masked graph autoencoders. In Proceedings of SIGKDD. 594\u2013604."},{"key":"e_1_3_1_17_2","first-page":"5403","volume-title":"Proceedings of CVPR","author":"Hu Peng","year":"2021","unstructured":"Peng Hu, Xi Peng, Hongyuan Zhu, Liangli Zhen, and Jie Lin. 2021. Learning cross-modal retrieval with noisy labels. In Proceedings of CVPR. 5403\u20135413."},{"key":"e_1_3_1_18_2","first-page":"141","volume-title":"Proceedings of ICMR","author":"Hu Zhikai","year":"2019","unstructured":"Zhikai Hu, Xin Liu, Xingzhi Wang, Yiu-Ming Cheung, Nannan Wang, and Yewang Chen. 2019. Triplet fusion network hashing for unpaired cross-modal retrieval. In Proceedings of ICMR. 141\u2013149."},{"key":"e_1_3_1_19_2","first-page":"3270","volume-title":"Proceedings of CVPR","author":"Jiang Qing-Yuan","year":"2017","unstructured":"Qing-Yuan Jiang and Wu-Jun Li. 2017. Deep cross-modal hashing. In Proceedings of CVPR. 3270\u20133278."},{"key":"e_1_3_1_20_2","first-page":"3283","volume-title":"Proceedings of ACM MM","author":"Jing Mengmeng","year":"2020","unstructured":"Mengmeng Jing, Jingjing Li, Lei Zhu, Ke Lu, Yang Yang, and Zi Huang. 2020. Incomplete cross-modal retrieval with dual-aligned variational autoencoders. In Proceedings of ACM MM. 3283\u20133291."},{"key":"e_1_3_1_21_2","volume-title":"Proceedings of ICLR","author":"Kingma Diederik P.","year":"2014","unstructured":"Diederik P. Kingma and Max Welling. 2014. Auto-encoding variational Bayes. In Proceedings of ICLR."},{"key":"e_1_3_1_22_2","first-page":"4242","volume-title":"Proceedings of CVPR","author":"Li Chao","year":"2018","unstructured":"Chao Li, Cheng Deng, Ning Li, Wei Liu, Xinbo Gao, and Dacheng Tao. 2018. Self-supervised adversarial hashing networks for cross-modal retrieval. In Proceedings of CVPR. 4242\u20134251."},{"key":"e_1_3_1_23_2","first-page":"1589","volume-title":"Proceedings of ACM MM","author":"Liu Hong","year":"2018","unstructured":"Hong Liu, Mingbao Lin, Shengchuan Zhang, Yongjian Wu, Feiyue Huang, and Rongrong Ji. 2018. Dense auto-encoder hashing for robust cross-modality retrieval. In Proceedings of ACM MM. 1589\u20131597."},{"key":"e_1_3_1_24_2","first-page":"1129","volume-title":"Proceedings of ACM MM","author":"Lu Xu","year":"2019","unstructured":"Xu Lu, Lei Zhu, Zhiyong Cheng, Jingjing Li, Xiushan Nie, and Huaxiang Zhang. 2019. Flexible online multi-modal hashing for large-scale multimedia retrieval. In Proceedings of ACM MM. 1129\u20131137."},{"key":"e_1_3_1_25_2","first-page":"3846","volume-title":"Proceedings of IJCAI","author":"Peng Yuxin","year":"2016","unstructured":"Yuxin Peng, Xin Huang, and Jinwei Qi. 2016. Cross-media shared representation by hierarchical learning with multiple deep networks. In Proceedings of IJCAI. 3846\u20133853."},{"issue":"1","key":"e_1_3_1_26_2","first-page":"Article 22, 24","article-title":"CM-GANs: Cross-modal generative adversarial networks for common representation learning","volume":"15","author":"Peng Yuxin","year":"2019","unstructured":"Yuxin Peng and Jinwei Qi. 2019. CM-GANs: Cross-modal generative adversarial networks for common representation learning. ACM Trans. Multim. Comput. Commun. Appl. 15, 1 (2019), Article 22, 24 pages.","journal-title":"ACM Trans. Multim. Comput. Commun. Appl."},{"key":"e_1_3_1_27_2","doi-asserted-by":"publisher","DOI":"10.1109\/TMM.2017.2742704"},{"key":"e_1_3_1_28_2","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2018.2852503"},{"key":"e_1_3_1_29_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2013.142"},{"key":"e_1_3_1_30_2","first-page":"8748","volume-title":"Proceedings of ICML","volume":"139","author":"Radford Alec","year":"2021","unstructured":"Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever. 2021. Learning transferable visual models from natural language supervision. In Proceedings of ICML, Vol. 139. 8748\u20138763."},{"key":"e_1_3_1_31_2","first-page":"139","volume-title":"Proceedings of NAACL","author":"Rashtchian Cyrus","year":"2010","unstructured":"Cyrus Rashtchian, Peter Young, Micah Hodosh, and Julia Hockenmaier. 2010. Collecting image annotations using Amazon\u2019s Mechanical Turk. In Proceedings of NAACL. 139\u2013147."},{"key":"e_1_3_1_32_2","first-page":"823","volume-title":"Proceedings of AISTATS","volume":"33","author":"Rasiwasia Nikhil","year":"2014","unstructured":"Nikhil Rasiwasia, Dhruv Mahajan, Vijay Mahadevan, and Gaurav Aggarwal. 2014. Cluster canonical correlation analysis. In Proceedings of AISTATS, Vol. 33. 823\u2013831."},{"key":"e_1_3_1_33_2","first-page":"251","volume-title":"Proceedings of ACM MM","author":"Rasiwasia Nikhil","year":"2010","unstructured":"Nikhil Rasiwasia, Jos\u00e9 Costa Pereira, Emanuele Coviello, Gabriel Doyle, Gert R. G. Lanckriet, Roger Levy, and Nuno Vasconcelos. 2010. A new approach to cross-modal multimedia retrieval. In Proceedings of ACM MM. 251\u2013260."},{"key":"e_1_3_1_34_2","first-page":"75","volume-title":"Proceedings of MM","author":"Semedo David","year":"2019","unstructured":"David Semedo and Jo\u00e3o Magalh\u00e3es. 2019. Cross-modal subspace learning with scheduled adaptive margin constraints. In Proceedings of MM. 75\u201383."},{"key":"e_1_3_1_35_2","first-page":"2160","volume-title":"Proceedings of CVPR","author":"Sharma Abhishek","year":"2012","unstructured":"Abhishek Sharma, Abhishek Kumar, Hal Daum\u00e9 III, and David W. Jacobs. 2012. Generalized multiview analysis: A discriminative latent space. In Proceedings of CVPR. 2160\u20132167."},{"key":"e_1_3_1_36_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCYB.2016.2606441"},{"key":"e_1_3_1_37_2","first-page":"1399","volume-title":"Proceedings of SIGIR","author":"Shraga Roee","year":"2020","unstructured":"Roee Shraga, Haggai Roitman, Guy Feigenblat, and Mustafa Canim. 2020. Web table retrieval using multimodal deep learning. In Proceedings of SIGIR. 1399\u20131408."},{"key":"e_1_3_1_38_2","first-page":"1","volume-title":"Proceedings of ICLR","author":"Velickovic Petar","year":"2018","unstructured":"Petar Velickovic, Guillem Cucurull, Arantxa Casanova, Adriana Romero, Pietro Li\u00f2, and Yoshua Bengio. 2018. Graph attention networks. In Proceedings of ICLR. 1\u201312."},{"key":"e_1_3_1_39_2","first-page":"154","volume-title":"Proceedings of ACM MM","author":"Wang Bokun","year":"2017","unstructured":"Bokun Wang, Yang Yang, Xing Xu, Alan Hanjalic, and Heng Tao Shen. 2017. Adversarial cross-modal retrieval. In Proceedings of ACM MM. 154\u2013162."},{"key":"e_1_3_1_40_2","first-page":"4300","volume-title":"Proceedings of ACM MM","author":"Wang Junsheng","year":"2022","unstructured":"Junsheng Wang, Tiantian Gong, Zhixiong Zeng, Changchang Sun, and Yan Yan. 2022. C \\({}^{\\mbox{3}}\\) CMR: Cross-modality cross-instance contrastive learning for cross-media retrieval. In Proceedings of ACM MM. 4300\u20134308."},{"key":"e_1_3_1_41_2","article-title":"A comprehensive survey on cross-modal retrieval","volume":"1607","author":"Wang Kaiye","year":"2016","unstructured":"Kaiye Wang, Qiyue Yin, Wei Wang, Shu Wu, and Liang Wang. 2016. A comprehensive survey on cross-modal retrieval. CoRR abs\/1607.06215 (2016).","journal-title":"CoRR"},{"key":"e_1_3_1_42_2","first-page":"3904","volume-title":"Proceedings of IJCAI","author":"Wang Qifan","year":"2015","unstructured":"Qifan Wang, Luo Si, and Bin Shen. 2015. Learning to hash on partial multi-modal data. In Proceedings of IJCAI. 3904\u20133910."},{"key":"e_1_3_1_43_2","first-page":"1","volume-title":"Proceedings of ICLR","author":"Wang Weiran","year":"2016","unstructured":"Weiran Wang and Karen Livescu. 2016. Large-scale approximate kernel canonical correlation analysis. In Proceedings of ICLR. 1\u201314."},{"issue":"3","key":"e_1_3_1_44_2","first-page":"Article 84, 19","article-title":"Eigenvector-based distance metric learning for image classification and retrieval","volume":"15","author":"Wang Zhangcheng","year":"2019","unstructured":"Zhangcheng Wang, Ya Li, Richang Hong, and Xinmei Tian. 2019. Eigenvector-based distance metric learning for image classification and retrieval. ACM Trans. Multim. Comput. Commun. Appl. 15, 3 (2019), Article 84, 19 pages.","journal-title":"ACM Trans. Multim. Comput. Commun. Appl."},{"key":"e_1_3_1_45_2","first-page":"3984","volume-title":"Proceedings of CVPR","author":"Wu Yiling","year":"2017","unstructured":"Yiling Wu, Shuhui Wang, and Qingming Huang. 2017. Online asymmetric similarity learning for cross-modal retrieval. In Proceedings of CVPR. 3984\u20133993."},{"key":"e_1_3_1_46_2","first-page":"825","volume-title":"Proceedings of ACM MM","author":"Wu Yiling","year":"2018","unstructured":"Yiling Wu, Shuhui Wang, and Qingming Huang. 2018. Learning semantic structure-preserved embeddings for cross-modal retrieval. In Proceedings of ACM MM. 825\u2013833."},{"key":"e_1_3_1_47_2","first-page":"1419","volume-title":"Proceedings of SIGIR","author":"Xu Xing","year":"2020","unstructured":"Xing Xu, Kaiyi Lin, Huimin Lu, Lianli Gao, and Heng Tao Shen. 2020. Correlated features synthesis and alignment for zero-shot cross-modal retrieval. In Proceedings of SIGIR. 1419\u20131428."},{"issue":"1","key":"e_1_3_1_48_2","first-page":"741","article-title":"Multi-modal discrete collaborative filtering for efficient cold-start recommendation","volume":"35","author":"Xu Yang","year":"2023","unstructured":"Yang Xu, Lei Zhu, Zhiyong Cheng, Jingjing Li, Zheng Zhang, and Huaxiang Zhang. 2023. Multi-modal discrete collaborative filtering for efficient cold-start recommendation. IEEE Trans. Knowl. Data Eng. 35, 1 (2023), 741\u2013755.","journal-title":"IEEE Trans. Knowl. Data Eng."},{"key":"e_1_3_1_49_2","first-page":"7444","volume-title":"Proceedings of AAAI","author":"Yan Sijie","year":"2018","unstructured":"Sijie Yan, Yuanjun Xiong, and Dahua Lin. 2018. Spatial temporal graph convolutional networks for skeleton-based action recognition. In Proceedings of AAAI. 7444\u20137452."},{"key":"e_1_3_1_50_2","first-page":"7541","volume-title":"Proceedings of CVPR","author":"Yang Erkun","year":"2022","unstructured":"Erkun Yang, Dongren Yao, Tongliang Liu, and Cheng Deng. 2022. Mutual quantization for cross-modal search with noisy labels. In Proceedings of CVPR. 7541\u20137550."},{"key":"e_1_3_1_51_2","article-title":"A comprehensive empirical study of vision-language pre-trained model for supervised cross-modal retrieval","volume":"2201","author":"Zeng Zhixiong","year":"2022","unstructured":"Zhixiong Zeng and Wenji Mao. 2022. A comprehensive empirical study of vision-language pre-trained model for supervised cross-modal retrieval. CoRR abs\/2201.02772 (2022).","journal-title":"CoRR"},{"key":"e_1_3_1_52_2","first-page":"5427","volume-title":"Proceedings of ACM MM","author":"Zeng Zhixiong","year":"2021","unstructured":"Zhixiong Zeng, Ying Sun, and Wenji Mao. 2021. MCCN: Multimodal coordinated clustering network for large-scale cross-modal retrieval. In Proceedings of ACM MM. 5427\u20135435."},{"key":"e_1_3_1_53_2","first-page":"1125","volume-title":"Proceedings of SIGIR","author":"Zeng Zhixiong","year":"2021","unstructured":"Zhixiong Zeng, Shuai Wang, Nan Xu, and Wenji Mao. 2021. PAN: Prototype-based adaptive network for robust cross-modal retrieval. In Proceedings of SIGIR. 1125\u20131134."},{"issue":"1","key":"e_1_3_1_54_2","first-page":"Article 2, 22 p","article-title":"HCMSL: Hybrid cross-modal similarity learning for cross-modal retrieval","volume":"17","author":"Zhang Chengyuan","year":"2021","unstructured":"Chengyuan Zhang, Jiayu Song, Xiaofeng Zhu, Lei Zhu, and Shichao Zhang. 2021. HCMSL: Hybrid cross-modal similarity learning for cross-modal retrieval. ACM Trans. Multim. Comput. Commun. Appl. 17, 1s (2021), Article 2, 22 pages.","journal-title":"ACM Trans. Multim. Comput. Commun. Appl."},{"issue":"1","key":"e_1_3_1_55_2","first-page":"Article 2, 26 p","article-title":"Deep learning-based multimedia analytics: A review","volume":"15","author":"Zhang Wei","year":"2019","unstructured":"Wei Zhang, Ting Yao, Shiai Zhu, and Abdulmotaleb El-Saddik. 2019. Deep learning-based multimedia analytics: A review. ACM Trans. Multim. Comput. Commun. Appl. 15, 1s (2019), Article 2, 26 pages.","journal-title":"ACM Trans. Multim. Comput. Commun. Appl."},{"key":"e_1_3_1_56_2","first-page":"10394","volume-title":"Proceedings of CVPR","author":"Zhen Liangli","year":"2019","unstructured":"Liangli Zhen, Peng Hu, Xu Wang, and Dezhong Peng. 2019. Deep supervised cross-modal retrieval. In Proceedings of CVPR. 10394\u201310403."},{"key":"e_1_3_1_57_2","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2020.2974065"},{"key":"e_1_3_1_58_2","doi-asserted-by":"publisher","DOI":"10.1109\/TKDE.2023.3282921"}],"container-title":["ACM Transactions on Multimedia Computing, Communications, and Applications"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3637442","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3637442","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,19]],"date-time":"2025-06-19T00:03:25Z","timestamp":1750291405000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3637442"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,1,11]]},"references-count":57,"journal-issue":{"issue":"5","published-print":{"date-parts":[[2024,5,31]]}},"alternative-id":["10.1145\/3637442"],"URL":"https:\/\/doi.org\/10.1145\/3637442","relation":{},"ISSN":["1551-6857","1551-6865"],"issn-type":[{"value":"1551-6857","type":"print"},{"value":"1551-6865","type":"electronic"}],"subject":[],"published":{"date-parts":[[2024,1,11]]},"assertion":[{"value":"2023-04-20","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2023-12-10","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2024-01-11","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}