{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,2,25]],"date-time":"2026-02-25T17:15:39Z","timestamp":1772039739115,"version":"3.50.1"},"publisher-location":"New York, NY, USA","reference-count":55,"publisher":"ACM","license":[{"start":{"date-parts":[[2022,10,10]],"date-time":"2022-10-10T00:00:00Z","timestamp":1665360000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"name":"National Natural Science Foundation of China","award":["61877006 62192784"],"award-info":[{"award-number":["61877006 62192784"]}]},{"name":"CAAI-Huawei MindSpore Open Fund","award":["CAAIXSJLJJ-2021-007B"],"award-info":[{"award-number":["CAAIXSJLJJ-2021-007B"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2022,10,10]]},"DOI":"10.1145\/3503161.3548391","type":"proceedings-article","created":{"date-parts":[[2022,10,10]],"date-time":"2022-10-10T15:43:12Z","timestamp":1665416592000},"page":"277-285","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":15,"title":["Semantic Structure Enhanced Contrastive Adversarial Hash Network for Cross-media Representation Learning"],"prefix":"10.1145","author":[{"given":"Meiyu","family":"Liang","sequence":"first","affiliation":[{"name":"Beijing Univeristy of Posts and Telecommunications, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Junping","family":"Du","sequence":"additional","affiliation":[{"name":"Beijing Univeristy of Posts and Telecommunications, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Xiaowen","family":"Cao","sequence":"additional","affiliation":[{"name":"Beijing Univeristy of Posts and Telecommunications, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yang","family":"Yu","sequence":"additional","affiliation":[{"name":"Beijing Univeristy of Posts and Telecommunications, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Kangkang","family":"Lu","sequence":"additional","affiliation":[{"name":"Beijing Univeristy of Posts and Telecommunications, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Zhe","family":"Xue","sequence":"additional","affiliation":[{"name":"Beijing Univeristy of Posts and Telecommunications, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Min","family":"Zhang","sequence":"additional","affiliation":[{"name":"Beijing Univeristy of Posts and Telecommunications, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2022,10,10]]},"reference":[{"key":"e_1_3_2_2_1_1","unstructured":"Vaswani A Shazeer N Parmar N and etal 2017. Attention is all you need. In NIPS. 6000--6010. Vaswani A Shazeer N Parmar N and et al. 2017. Attention is all you need. In NIPS. 6000--6010."},{"key":"e_1_3_2_2_2_1","doi-asserted-by":"crossref","unstructured":"M. M. Bronstein A. M. Bronstein F. Michel and N. Paragios. 2010. Data fusion through cross-modality metric learning using similarity-sensitive hashing. In CVPR. 3594--3601. M. M. Bronstein A. M. Bronstein F. Michel and N. Paragios. 2010. Data fusion through cross-modality metric learning using similarity-sensitive hashing. In CVPR. 3594--3601.","DOI":"10.1109\/CVPR.2010.5539928"},{"key":"e_1_3_2_2_3_1","doi-asserted-by":"crossref","unstructured":"Li C Deng C and Wang L. 2019. Coupled cycleGAN: unsupervised hashing network for cross-modal retrieval. In AAAI. 176--183. Li C Deng C and Wang L. 2019. Coupled cycleGAN: unsupervised hashing network for cross-modal retrieval. In AAAI. 176--183.","DOI":"10.1609\/aaai.v33i01.3301176"},{"key":"e_1_3_2_2_4_1","doi-asserted-by":"crossref","unstructured":"Yue Cao Mingsheng Long Jianmin Wang and Shichen Liu. 2017. Collective deep quantization for efficient cross-modal retrieval. In AAAI. 3974--3980. Yue Cao Mingsheng Long Jianmin Wang and Shichen Liu. 2017. Collective deep quantization for efficient cross-modal retrieval. In AAAI. 3974--3980.","DOI":"10.1609\/aaai.v31i1.11218"},{"key":"e_1_3_2_2_5_1","doi-asserted-by":"crossref","unstructured":"Li Chao Deng Cheng Li Ning Liu Wei Gao Xinbo and Tao Dacheng. 2018. Selfsupervised adversarial hashing networks for cross-modal retrieval. In CVPR. 4242--4251. Li Chao Deng Cheng Li Ning Liu Wei Gao Xinbo and Tao Dacheng. 2018. Selfsupervised adversarial hashing networks for cross-modal retrieval. In CVPR. 4242--4251.","DOI":"10.1109\/CVPR.2018.00446"},{"key":"e_1_3_2_2_6_1","doi-asserted-by":"crossref","unstructured":"Yudong Chen Sen Wang Jianglin Lu Zhi Chen Zheng Zhang and Zi Huang. 2021. Local graph convolutional networks for cross-modal hashing. In ACM MM. 1921--1928. Yudong Chen Sen Wang Jianglin Lu Zhi Chen Zheng Zhang and Zi Huang. 2021. Local graph convolutional networks for cross-modal hashing. In ACM MM. 1921--1928.","DOI":"10.1145\/3474085.3475346"},{"key":"e_1_3_2_2_7_1","doi-asserted-by":"publisher","DOI":"10.1109\/TCSVT.2019.2900171"},{"key":"e_1_3_2_2_8_1","doi-asserted-by":"crossref","unstructured":"Tat-Seng Chua Jinhui Tang Richang Hong Haojie Li Zhiping Luo and Yantao Zheng. 2009. NUS-WIDE: a real-world web image database from national university of singapore. In CIVR. 1--9. Tat-Seng Chua Jinhui Tang Richang Hong Haojie Li Zhiping Luo and Yantao Zheng. 2009. NUS-WIDE: a real-world web image database from national university of singapore. In CIVR. 1--9.","DOI":"10.1145\/1646396.1646452"},{"key":"e_1_3_2_2_9_1","unstructured":"Jiequan Cui Zhisheng Zhong Shu Liu Bei Yu and Jiaya Jia. 2021. Parametric contrastive learning. In ICCV. 695--704. Jiequan Cui Zhisheng Zhong Shu Liu Bei Yu and Jiaya Jia. 2021. Parametric contrastive learning. In ICCV. 695--704."},{"key":"e_1_3_2_2_10_1","doi-asserted-by":"crossref","unstructured":"Debidatta Dwibedi Yusuf Aytar Jonathan Tompson Pierre Sermanet and Andrew Zisserman. 2021. With a little help from my friends: nearest-neighbor contrastive learning of visual representations. In ICCV. 9568--9577. Debidatta Dwibedi Yusuf Aytar Jonathan Tompson Pierre Sermanet and Andrew Zisserman. 2021. With a little help from my friends: nearest-neighbor contrastive learning of visual representations. In ICCV. 9568--9577.","DOI":"10.1109\/ICCV48922.2021.00945"},{"key":"e_1_3_2_2_11_1","doi-asserted-by":"crossref","unstructured":"Fangxiang Feng Xiaojie Wang and Ruifan Li. 2014. Cross-modal retrieval with correspondence autoencoder. In ACM MM. 7--16. Fangxiang Feng Xiaojie Wang and Ruifan Li. 2014. Cross-modal retrieval with correspondence autoencoder. In ACM MM. 7--16.","DOI":"10.1145\/2647868.2654902"},{"key":"e_1_3_2_2_12_1","unstructured":"I. J. Goodfellow J. Pouget-Abadie M. Mirza X. Bing and Y. Bengio. 2014. Generative adversarial nets. MIT Press (2014). I. J. Goodfellow J. Pouget-Abadie M. Mirza X. Bing and Y. Bengio. 2014. Generative adversarial nets. MIT Press (2014)."},{"key":"e_1_3_2_2_13_1","unstructured":"Kaiming He Haoqi Fan YuxinWu Saining Xie and Ross Girshick. 2020. Momentum contrast for unsupervised visual representation learning. In CVPR. 9729--9738. Kaiming He Haoqi Fan YuxinWu Saining Xie and Ross Girshick. 2020. Momentum contrast for unsupervised visual representation learning. In CVPR. 9729--9738."},{"key":"e_1_3_2_2_14_1","doi-asserted-by":"publisher","DOI":"10.1109\/TMM.2017.2760101"},{"key":"e_1_3_2_2_15_1","volume-title":"Lew","author":"Huiskes Mark J.","year":"2008","unstructured":"Mark J. Huiskes and Michael S . Lew . 2008 . The MIR Flickr retrieval evaluation. In MIR. 39--43. Mark J. Huiskes and Michael S. Lew. 2008. The MIR Flickr retrieval evaluation. In MIR. 39--43."},{"key":"e_1_3_2_2_16_1","unstructured":"Yuqi Huo Manli Zhang Guangzhen Liu and Haoyu Lu. 2021. WenLan: Bridging Vision and Language by Large-Scale Multi-Modal Pre-Training. In CoRR abs\/2103.06561. Yuqi Huo Manli Zhang Guangzhen Liu and Haoyu Lu. 2021. WenLan: Bridging Vision and Language by Large-Scale Multi-Modal Pre-Training. In CoRR abs\/2103.06561."},{"key":"e_1_3_2_2_17_1","volume-title":"Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 4171--4186","author":"Devlin","unstructured":"Devlin J, Chang Mingwei , Lee K , and et al. 2019. Bert: pre-training of deep bidirectional transformers for language understanding . In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 4171--4186 . Devlin J, Chang Mingwei, Lee K, and et al. 2019. Bert: pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 4171--4186."},{"key":"e_1_3_2_2_18_1","doi-asserted-by":"crossref","unstructured":"Qing-Yuan Jiang and Wu-Jun Li. 2017. Deep cross-modal hashing. In CVPR. 3270--3278. Qing-Yuan Jiang and Wu-Jun Li. 2017. Deep cross-modal hashing. In CVPR. 3270--3278.","DOI":"10.1109\/CVPR.2017.348"},{"key":"e_1_3_2_2_19_1","doi-asserted-by":"crossref","unstructured":"Chen Jiayi and Zhang Aidong. 2020. HGMF: heterogeneous graph-based fusion for multimodal data with incompleteness. In KDD. 1295--1305. Chen Jiayi and Zhang Aidong. 2020. HGMF: heterogeneous graph-based fusion for multimodal data with incompleteness. In KDD. 1295--1305.","DOI":"10.1145\/3394486.3403182"},{"key":"e_1_3_2_2_20_1","unstructured":"Thomas N Kipf and Max Welling. 2017. Semi-supervised classification with graph convolutional networks. In ICLR. 1--14. Thomas N Kipf and Max Welling. 2017. Semi-supervised classification with graph convolutional networks. In ICLR. 1--14."},{"key":"e_1_3_2_2_21_1","doi-asserted-by":"crossref","unstructured":"Z Lin G Ding and J Wang. 2015. Semantics preserving hashing for cross-view retrieval. In CVPR. 3864--3872. Z Lin G Ding and J Wang. 2015. Semantics preserving hashing for cross-view retrieval. In CVPR. 3864--3872.","DOI":"10.1109\/CVPR.2015.7299011"},{"key":"e_1_3_2_2_22_1","doi-asserted-by":"publisher","DOI":"10.1007\/s10489-019-01511-7"},{"key":"e_1_3_2_2_23_1","unstructured":"Jiawei Liu Zheng-Jun Zha and Richang Hong. 2019. Deep adversarial graph attention convolution network for text-based person search. In ACM MM. 665--673. Jiawei Liu Zheng-Jun Zha and Richang Hong. 2019. Deep adversarial graph attention convolution network for text-based person search. In ACM MM. 665--673."},{"key":"e_1_3_2_2_24_1","unstructured":"M. Long Y. Cao J. Wang and P. S. Yu. 2016. Composite correlation quantization for efficient multimodal retrieval. ACM SIGIR (2016) 579--588. M. Long Y. Cao J. Wang and P. S. Yu. 2016. Composite correlation quantization for efficient multimodal retrieval. ACM SIGIR (2016) 579--588."},{"key":"e_1_3_2_2_25_1","doi-asserted-by":"publisher","DOI":"10.1109\/TMM.2020.2969792"},{"key":"e_1_3_2_2_26_1","first-page":"3634","article-title":"Cross-media semantic correlation learning based on deep hash network and semantic expansion for social network cross-media search","volume":"31","author":"Meiyu Liang","year":"2020","unstructured":"Liang Meiyu , Du Junping , and Yang Congxian . 2020 . Cross-media semantic correlation learning based on deep hash network and semantic expansion for social network cross-media search . IEEE TNNLS 31 , 9 (2020), 3634 -- 3648 . Liang Meiyu, Du Junping, and Yang Congxian. 2020. Cross-media semantic correlation learning based on deep hash network and semantic expansion for social network cross-media search. IEEE TNNLS 31, 9 (2020), 3634--3648.","journal-title":"IEEE TNNLS"},{"key":"e_1_3_2_2_27_1","first-page":"2372","article-title":"An overview of cross-media retrieval: concepts, methodologies, benchmarks, and challenges","volume":"28","author":"Peng Y.","year":"2018","unstructured":"Y. Peng , X. Huang , and Y. Zhao . 2018 . An overview of cross-media retrieval: concepts, methodologies, benchmarks, and challenges . IEEE TCSVT 28 , 9 (2018), 2372 -- 2385 . Y. Peng, X. Huang, and Y. Zhao. 2018. An overview of cross-media retrieval: concepts, methodologies, benchmarks, and challenges. IEEE TCSVT 28, 9 (2018), 2372--2385.","journal-title":"IEEE TCSVT"},{"key":"e_1_3_2_2_28_1","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2018.2852503"},{"key":"e_1_3_2_2_29_1","doi-asserted-by":"publisher","DOI":"10.1145\/3284750"},{"key":"e_1_3_2_2_30_1","doi-asserted-by":"publisher","DOI":"10.3115\/v1\/D14-1162"},{"key":"e_1_3_2_2_31_1","doi-asserted-by":"crossref","unstructured":"Shengsheng Qian Dizhan Xue Huaiwen Zhang Quan Fang and Changsheng Xu. 2021. Dual adversarial graph neural networks for multi-label cross-modal retrieval. In AAAI. 2440--2448. Shengsheng Qian Dizhan Xue Huaiwen Zhang Quan Fang and Changsheng Xu. 2021. Dual adversarial graph neural networks for multi-label cross-modal retrieval. In AAAI. 2440--2448.","DOI":"10.1609\/aaai.v35i3.16345"},{"key":"e_1_3_2_2_32_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2016.2577031"},{"key":"e_1_3_2_2_33_1","doi-asserted-by":"crossref","unstructured":"Zhanjian Shen Deming Zhai Xianming Liu and Junjun Jiang. 2020. Semi-supervised graph convolutional hashing network for large-scale cross-modal retrieval. In ICIP. 2366--2370. Zhanjian Shen Deming Zhai Xianming Liu and Junjun Jiang. 2020. Semi-supervised graph convolutional hashing network for large-scale cross-modal retrieval. In ICIP. 2366--2370.","DOI":"10.1109\/ICIP40778.2020.9190641"},{"key":"e_1_3_2_2_34_1","first-page":"79","article-title":"A survey of cross-media analysis and reasoning technology research","volume":"48","author":"Shuhui Wang","year":"2021","unstructured":"Wang Shuhui , Yan Xu , and Huangqingming. 2021 . A survey of cross-media analysis and reasoning technology research . Computer Science 48 , 3 (2021), 79 -- 86 . Wang Shuhui, Yan Xu, and Huangqingming. 2021. A survey of cross-media analysis and reasoning technology research. Computer Science 48, 3 (2021), 79--86.","journal-title":"Computer Science"},{"key":"e_1_3_2_2_35_1","unstructured":"Wolf T Debut L Sanh V and etal 2019. Huggingface's transformers: State-of-the-art natural language processing. arXiv preprint arXiv:1910.03771 (2019). Wolf T Debut L Sanh V and et al. 2019. Huggingface's transformers: State-of-the-art natural language processing. arXiv preprint arXiv:1910.03771 (2019)."},{"key":"e_1_3_2_2_36_1","unstructured":"Petar Veli\u010dkovi\u0107 Guillem Cucurull Arantxa Casanova Adriana Romero Pietro Li\u00f2 and Yoshua Bengio. 2018. Graph attention networks. arXiv:1710.10903 [stat.ML] Petar Veli\u010dkovi\u0107 Guillem Cucurull Arantxa Casanova Adriana Romero Pietro Li\u00f2 and Yoshua Bengio. 2018. Graph attention networks. arXiv:1710.10903 [stat.ML]"},{"key":"e_1_3_2_2_37_1","doi-asserted-by":"crossref","unstructured":"Bokun Wang Yang Yang Xing Xu Alan Hanjalic and Heng Tao Shen. 2017. Adversarial cross-Modal retrieval. In ACM MM. 154--162. Bokun Wang Yang Yang Xing Xu Alan Hanjalic and Heng Tao Shen. 2017. Adversarial cross-Modal retrieval. In ACM MM. 154--162.","DOI":"10.1145\/3123266.3123326"},{"key":"e_1_3_2_2_38_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2018.2861000"},{"key":"e_1_3_2_2_39_1","first-page":"449","article-title":"Cross-modal retrieval with CNN visual features: a new baseline","volume":"47","author":"Wei Y.","year":"2017","unstructured":"Y. Wei , Y. Zhao , C. Lu , S. Wei , L. Liu , Z. Zhu , and S. Yan . 2017 . Cross-modal retrieval with CNN visual features: a new baseline . IEEE T CYBERNETICS 47 , 2 (2017), 449 -- 460 . Y. Wei, Y. Zhao, C. Lu, S. Wei, L. Liu, Z. Zhu, and S. Yan. 2017. Cross-modal retrieval with CNN visual features: a new baseline. IEEE T CYBERNETICS 47, 2 (2017), 449--460.","journal-title":"IEEE T CYBERNETICS"},{"key":"e_1_3_2_2_40_1","doi-asserted-by":"crossref","unstructured":"Gu Wendel Gu Xiaoyan Jingzi Gu Li Bo Xiong Zhi and Wang Weiping. 2019. Adversary guided asymmetric hashing for cross-modal retrieval. In ICMR. 159--167. Gu Wendel Gu Xiaoyan Jingzi Gu Li Bo Xiong Zhi and Wang Weiping. 2019. Adversary guided asymmetric hashing for cross-modal retrieval. In ICMR. 159--167.","DOI":"10.1145\/3323873.3325045"},{"key":"e_1_3_2_2_41_1","unstructured":"He Xiaodong Buehler C and etal 2018. Bottom-up and top-down attention for image captioning and visual question answering. In CVPR. 6077--6086. He Xiaodong Buehler C and et al. 2018. Bottom-up and top-down attention for image captioning and visual question answering. In CVPR. 6077--6086."},{"key":"e_1_3_2_2_42_1","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2020.2963957"},{"key":"e_1_3_2_2_43_1","doi-asserted-by":"crossref","unstructured":"Ruiqing Xu Chao Li Junchi Yan Cheng Deng and Xianglong Liu. 2019. Graph convolutional network hashing for cross-modal retrieval. In IJCAI. 982--988. Ruiqing Xu Chao Li Junchi Yan Cheng Deng and Xianglong Liu. 2019. Graph convolutional network hashing for cross-modal retrieval. In IJCAI. 982--988.","DOI":"10.24963\/ijcai.2019\/138"},{"key":"e_1_3_2_2_44_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.ins.2019.01.018"},{"key":"e_1_3_2_2_45_1","doi-asserted-by":"crossref","unstructured":"Shi Y You X and Zheng F. 2019. Equally-guided discriminative hashing for crossmodal retrieval. In AAAI. 4767--4773. Shi Y You X and Zheng F. 2019. Equally-guided discriminative hashing for crossmodal retrieval. In AAAI. 4767--4773.","DOI":"10.24963\/ijcai.2019\/662"},{"key":"e_1_3_2_2_46_1","doi-asserted-by":"crossref","unstructured":"Erkun Yang Cheng Deng Wei Liu Xianglong Liu Dacheng Tao and Xinbo Gao. 2017. Pairwise relationship guided deep hashing for cross-modal retrieval. In AAAI. 1618--1625. Erkun Yang Cheng Deng Wei Liu Xianglong Liu Dacheng Tao and Xinbo Gao. 2017. Pairwise relationship guided deep hashing for cross-modal retrieval. In AAAI. 1618--1625.","DOI":"10.1609\/aaai.v31i1.10719"},{"key":"e_1_3_2_2_47_1","doi-asserted-by":"crossref","unstructured":"Jun Yu Hao Zhou Yibing Zhan and Dacheng Tao. 2021. Deep graph-neighbor coherence preserving network for unsupervised cross-modal hashing. In AAAI. 4626--4634. Jun Yu Hao Zhou Yibing Zhan and Dacheng Tao. 2021. Deep graph-neighbor coherence preserving network for unsupervised cross-modal hashing. In AAAI. 4626--4634.","DOI":"10.1609\/aaai.v35i5.16592"},{"key":"e_1_3_2_2_48_1","doi-asserted-by":"crossref","unstructured":"Xin Yuan Zhe Lin Jason Kuen and el al. 2021. Multimodal contrastive training for visual representation learning. In CVPR. 6995--7004. Xin Yuan Zhe Lin Jason Kuen and el al. 2021. Multimodal contrastive training for visual representation learning. In CVPR. 6995--7004.","DOI":"10.1109\/CVPR46437.2021.00692"},{"key":"e_1_3_2_2_49_1","doi-asserted-by":"crossref","unstructured":"D. Zhang andW.-J. Li. 2014. Large-scale supervised multimodal hashing with semantic correlation maximization. In AAAI. 2177--2183. D. Zhang andW.-J. Li. 2014. Large-scale supervised multimodal hashing with semantic correlation maximization. In AAAI. 2177--2183.","DOI":"10.1609\/aaai.v28i1.8995"},{"key":"e_1_3_2_2_50_1","doi-asserted-by":"crossref","unstructured":"Jian Zhang Yuxin Peng and Mingkuan Yuan. 2018. Unsupervised generative adversarial cross-modal hashing. In AAAI. 539--546. Jian Zhang Yuxin Peng and Mingkuan Yuan. 2018. Unsupervised generative adversarial cross-modal hashing. In AAAI. 539--546.","DOI":"10.1609\/aaai.v32i1.11263"},{"key":"e_1_3_2_2_51_1","doi-asserted-by":"crossref","unstructured":"Pengfei Zhang Jiasheng Duan Zi Huang and Hongzhi Yin. 2021. Joint-teaching: learning to refine knowledge for resource-constrained unsupervised cross-modal retrieval. In ACM MM. 1517--1525. Pengfei Zhang Jiasheng Duan Zi Huang and Hongzhi Yin. 2021. Joint-teaching: learning to refine knowledge for resource-constrained unsupervised cross-modal retrieval. In ACM MM. 1517--1525.","DOI":"10.1145\/3474085.3475286"},{"key":"e_1_3_2_2_52_1","doi-asserted-by":"crossref","unstructured":"Xi Zhang Siyu Zhou Jiashi Feng Hanjiang Lai Bo Li Yan Pan Jian Yin and Shuicheng Yan. 2018. HashGAN: attention-aware deep adversarial hashing for cross modal retrieval. In ECCV. Xi Zhang Siyu Zhou Jiashi Feng Hanjiang Lai Bo Li Yan Pan Jian Yin and Shuicheng Yan. 2018. HashGAN: attention-aware deep adversarial hashing for cross modal retrieval. In ECCV.","DOI":"10.1007\/978-3-030-01267-0_36"},{"key":"e_1_3_2_2_53_1","doi-asserted-by":"crossref","unstructured":"Y Zhuang Z Yu WWang F Wu S Tang and J Shao. 2014. Cross-media hashing with neural networks. In ACM MM. 901--904. Y Zhuang Z Yu WWang F Wu S Tang and J Shao. 2014. Cross-media hashing with neural networks. In ACM MM. 901--904.","DOI":"10.1145\/2647868.2655059"},{"key":"e_1_3_2_2_54_1","first-page":"862","article-title":"Cross-modal video clip retrieval based on visual-text relationship alignment","volume":"50","author":"Zhuo Chen","year":"2020","unstructured":"Chen Zhuo , Du Hao , Wu Yufei , Xu Tong , and Chen Enhong . 2020 . Cross-modal video clip retrieval based on visual-text relationship alignment . SCI CHINA INFORM SCI 50 , 6 (2020), 862 -- 876 . Chen Zhuo, Du Hao,Wu Yufei, Xu Tong, and Chen Enhong. 2020. Cross-modal video clip retrieval based on visual-text relationship alignment. SCI CHINA INFORM SCI 50, 6 (2020), 862--876.","journal-title":"SCI CHINA INFORM SCI"},{"key":"e_1_3_2_2_55_1","doi-asserted-by":"crossref","unstructured":"Mohammadreza Zolfaghari Yi Zhu Peter Gehler and Thomas Brox. 2021. CrossCLR: cross-modal contrastive learning for multi-modal video representations. In ICCV. 1430--1439. Mohammadreza Zolfaghari Yi Zhu Peter Gehler and Thomas Brox. 2021. CrossCLR: cross-modal contrastive learning for multi-modal video representations. In ICCV. 1430--1439.","DOI":"10.1109\/ICCV48922.2021.00148"}],"event":{"name":"MM '22: The 30th ACM International Conference on Multimedia","location":"Lisboa Portugal","acronym":"MM '22","sponsor":["SIGMM ACM Special Interest Group on Multimedia"]},"container-title":["Proceedings of the 30th ACM International Conference on Multimedia"],"original-title":[],"link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3503161.3548391","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3503161.3548391","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T19:00:44Z","timestamp":1750186844000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3503161.3548391"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,10,10]]},"references-count":55,"alternative-id":["10.1145\/3503161.3548391","10.1145\/3503161"],"URL":"https:\/\/doi.org\/10.1145\/3503161.3548391","relation":{},"subject":[],"published":{"date-parts":[[2022,10,10]]},"assertion":[{"value":"2022-10-10","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}