{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,6]],"date-time":"2026-06-06T16:02:44Z","timestamp":1780761764617,"version":"3.54.1"},"publisher-location":"New York, NY, USA","reference-count":45,"publisher":"ACM","license":[{"start":{"date-parts":[[2022,6,27]],"date-time":"2022-06-27T00:00:00Z","timestamp":1656288000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["U1711262,61472428"],"award-info":[{"award-number":["U1711262,61472428"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]},{"name":"Public Computing Cloud, Renmin University of China"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2022,6,27]]},"DOI":"10.1145\/3512527.3531364","type":"proceedings-article","created":{"date-parts":[[2022,6,23]],"date-time":"2022-06-23T22:23:32Z","timestamp":1656023012000},"page":"527-535","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":14,"title":["Self-Lifting: A Novel Framework for Unsupervised Voice-Face Association Learning"],"prefix":"10.1145","author":[{"given":"Guangyu","family":"Chen","sequence":"first","affiliation":[{"name":"Renmin University of China, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Deyuan","family":"Zhang","sequence":"additional","affiliation":[{"name":"Shenyang Aerospace University, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Tao","family":"Liu","sequence":"additional","affiliation":[{"name":"Renmin University of China, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Xiaoyong","family":"Du","sequence":"additional","affiliation":[{"name":"Renmin University of China, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2022,6,27]]},"reference":[{"key":"e_1_3_2_2_1_1","doi-asserted-by":"publisher","DOI":"10.1145\/3197517.3201357"},{"key":"e_1_3_2_2_2_1","volume-title":"MM '20: The 28th ACM International Conference on Multimedia .","unstructured":"2020. Hearing like Seeing: Improving Voice-Face Interactions and Associations via Adversarial Deep Semantic Matching Network . In MM '20: The 28th ACM International Conference on Multimedia . 2020. Hearing like Seeing: Improving Voice-Face Interactions and Associations via Adversarial Deep Semantic Matching Network. In MM '20: The 28th ACM International Conference on Multimedia ."},{"key":"e_1_3_2_2_3_1","volume-title":"Listen and Learn. In 2017 IEEE International Conference on Computer Vision (ICCV) .","author":"Arandjelovic R","unstructured":"R Arandjelovic and A. Zisserman . 2017. Look , Listen and Learn. In 2017 IEEE International Conference on Computer Vision (ICCV) . R Arandjelovic and A. Zisserman. 2017. Look, Listen and Learn. In 2017 IEEE International Conference on Computer Vision (ICCV) ."},{"key":"e_1_3_2_2_4_1","volume-title":"IEEE International Conference on Automatic Face & Gesture Recognition","author":"Cao Q.","year":"2017","unstructured":"Q. Cao , L. Shen , W. Xie , O. M. Parkhi , and A. Zisserman . 2017. VGGFace2: A dataset for recognising faces across pose and age . IEEE International Conference on Automatic Face & Gesture Recognition ( 2017 ). Q. Cao, L. Shen, W. Xie, O. M. Parkhi, and A. Zisserman. 2017. VGGFace2: A dataset for recognising faces across pose and age. IEEE International Conference on Automatic Face & Gesture Recognition (2017)."},{"key":"e_1_3_2_2_5_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-01264-9_9"},{"key":"e_1_3_2_2_6_1","volume-title":"Unsupervised learning of visual features by contrasting cluster assignments. arXiv preprint arXiv:2006.09882","author":"Caron Mathilde","year":"2020","unstructured":"Mathilde Caron , Ishan Misra , Julien Mairal , Priya Goyal , Piotr Bojanowski , and Armand Joulin . 2020. Unsupervised learning of visual features by contrasting cluster assignments. arXiv preprint arXiv:2006.09882 ( 2020 ). Mathilde Caron, Ishan Misra, Julien Mairal, Priya Goyal, Piotr Bojanowski, and Armand Joulin. 2020. Unsupervised learning of visual features by contrasting cluster assignments. arXiv preprint arXiv:2006.09882 (2020)."},{"key":"e_1_3_2_2_7_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR46437.2021.01549"},{"key":"e_1_3_2_2_8_1","first-page":"1086","article-title":"VoxCeleb2","volume":"2018","author":"Chung Joon Son","year":"2018","unstructured":"Joon Son Chung , Arsha Nagrani , and Andrew Zisserman . 2018 . VoxCeleb2 : Deep Speaker Recognition. In Proc. Interspeech 2018. 1086 -- 1090 . https:\/\/doi.org\/10.21437\/Interspeech.2018--1929 10.21437\/Interspeech.2018--1929 Joon Son Chung, Arsha Nagrani, and Andrew Zisserman. 2018. VoxCeleb2: Deep Speaker Recognition. In Proc. Interspeech 2018. 1086--1090. https:\/\/doi.org\/10.21437\/Interspeech.2018--1929","journal-title":"Deep Speaker Recognition. In Proc. Interspeech"},{"key":"e_1_3_2_2_9_1","volume-title":"Asian conference on computer vision. Springer, 251--263","author":"Chung Joon Son","year":"2016","unstructured":"Joon Son Chung and Andrew Zisserman . 2016 . Out of time: automated lip sync in the wild . In Asian conference on computer vision. Springer, 251--263 . Joon Son Chung and Andrew Zisserman. 2016. Out of time: automated lip sync in the wild. In Asian conference on computer vision. Springer, 251--263."},{"key":"e_1_3_2_2_10_1","volume-title":"Out of Time: Automated Lip Sync in the Wild. In Asian Conference on Computer Vision .","author":"Chung J. S.","unstructured":"J. S. Chung and A. Zisserman . 2017 . Out of Time: Automated Lip Sync in the Wild. In Asian Conference on Computer Vision . J. S. Chung and A. Zisserman. 2017. Out of Time: Automated Lip Sync in the Wild. In Asian Conference on Computer Vision ."},{"key":"e_1_3_2_2_11_1","doi-asserted-by":"crossref","unstructured":"B. Desplanques J. Thienpondt and K. Demuynck. 2020. ECAPA-TDNN: Emphasized Channel Attention Propagation and Aggregation in TDNN Based Speaker Verification. In Interspeech 2020 .  B. Desplanques J. Thienpondt and K. Demuynck. 2020. ECAPA-TDNN: Emphasized Channel Attention Propagation and Aggregation in TDNN Based Speaker Verification. In Interspeech 2020 .","DOI":"10.21437\/Interspeech.2020-2650"},{"key":"e_1_3_2_2_12_1","doi-asserted-by":"crossref","unstructured":"Fangxiang Feng Xiaojie Wang Ruifan Li and Ibrar Ahmad. 2015. Correspondence autoencoders for cross-modal retrieval.  Fangxiang Feng Xiaojie Wang Ruifan Li and Ibrar Ahmad. 2015. Correspondence autoencoders for cross-modal retrieval.","DOI":"10.1145\/2808205"},{"key":"e_1_3_2_2_13_1","volume-title":"Zhaohan Daniel Guo, Mohammad Gheshlaghi Azar, et al.","author":"Grill Jean-Bastien","year":"2020","unstructured":"Jean-Bastien Grill , Florian Strub , Florent Altch\u00e9 , Corentin Tallec , Pierre H Richemond , Elena Buchatskaya , Carl Doersch , Bernardo Avila Pires , Zhaohan Daniel Guo, Mohammad Gheshlaghi Azar, et al. 2020 . Bootstrap your own latent: A new approach to self-supervised learning. arXiv preprint arXiv:2006.07733 (2020). Jean-Bastien Grill, Florian Strub, Florent Altch\u00e9, Corentin Tallec, Pierre H Richemond, Elena Buchatskaya, Carl Doersch, Bernardo Avila Pires, Zhaohan Daniel Guo, Mohammad Gheshlaghi Azar, et al. 2020. Bootstrap your own latent: A new approach to self-supervised learning. arXiv preprint arXiv:2006.07733 (2020)."},{"key":"e_1_3_2_2_14_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2006.100"},{"key":"e_1_3_2_2_15_1","doi-asserted-by":"publisher","DOI":"10.2307\/2346830"},{"key":"e_1_3_2_2_16_1","volume-title":"Putting a face to the voice: Fusing audio and visual signals across a video to determine speakers. arXiv preprint arXiv:1706.00079","author":"Hoover Ken","year":"2017","unstructured":"Ken Hoover , Sourish Chaudhuri , Caroline Pantofaru , Malcolm Slaney , and Ian Sturdy . 2017. Putting a face to the voice: Fusing audio and visual signals across a video to determine speakers. arXiv preprint arXiv:1706.00079 ( 2017 ). Ken Hoover, Sourish Chaudhuri, Caroline Pantofaru, Malcolm Slaney, and Ian Sturdy. 2017. Putting a face to the voice: Fusing audio and visual signals across a video to determine speakers. arXiv preprint arXiv:1706.00079 (2017)."},{"key":"e_1_3_2_2_17_1","doi-asserted-by":"publisher","DOI":"10.1145\/3240508.3240601"},{"key":"e_1_3_2_2_18_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.knosys.2019.05.017"},{"key":"e_1_3_2_2_19_1","volume-title":"Billion-scale similarity search with GPUs. arXiv preprint arXiv:1702.08734","author":"Johnson Jeff","year":"2017","unstructured":"Jeff Johnson , Matthijs Douze , and Herv\u00e9 J\u00e9gou . 2017. Billion-scale similarity search with GPUs. arXiv preprint arXiv:1702.08734 ( 2017 ). Jeff Johnson, Matthijs Douze, and Herv\u00e9 J\u00e9gou. 2017. Billion-scale similarity search with GPUs. arXiv preprint arXiv:1702.08734 (2017)."},{"key":"e_1_3_2_2_20_1","doi-asserted-by":"crossref","unstructured":"C. Kim H. V. Shin T. H. Oh A. Kaspar and W. Matusik. 2019. On Learning Associations of Faces and Voices.  C. Kim H. V. Shin T. H. Oh A. Kaspar and W. Matusik. 2019. On Learning Associations of Faces and Voices.","DOI":"10.1007\/978-3-030-20873-8_18"},{"key":"e_1_3_2_2_21_1","volume-title":"Adam: A Method for Stochastic Optimization. Computer Science","author":"Kingma D.","year":"2014","unstructured":"D. Kingma and J. Ba . 2014 . Adam: A Method for Stochastic Optimization. Computer Science (2014). D. Kingma and J. Ba. 2014. Adam: A Method for Stochastic Optimization. Computer Science (2014)."},{"key":"e_1_3_2_2_22_1","volume-title":"So Does Audio: Revealing the Hidden Face via Cross-Modality Transfer","author":"Kong Chenqi","year":"2021","unstructured":"Chenqi Kong , Baoliang Chen , Wenhan Yang , Haoliang Li , Peilin Chen , and Shiqi Wang . 2021. Appearance Matters , So Does Audio: Revealing the Hidden Face via Cross-Modality Transfer . IEEE Transactions on Circuits and Systems for Video Technology ( 2021 ). Chenqi Kong, Baoliang Chen, Wenhan Yang, Haoliang Li, Peilin Chen, and Shiqi Wang. 2021. Appearance Matters, So Does Audio: Revealing the Hidden Face via Cross-Modality Transfer. IEEE Transactions on Circuits and Systems for Video Technology (2021)."},{"key":"e_1_3_2_2_23_1","doi-asserted-by":"crossref","unstructured":"Mavica and W. Lauren. 2013. Matching Voice and Face Identity From Static Images. Journal of Experimental Psychology (2013).  Mavica and W. Lauren. 2013. Matching Voice and Face Identity From Static Images. Journal of Experimental Psychology (2013).","DOI":"10.1037\/a0030945"},{"key":"e_1_3_2_2_24_1","unstructured":"K. G. Munhall and E. Vatikiotis-Bateson. 1998. The moving face during speech communication. (1998).  K. G. Munhall and E. Vatikiotis-Bateson. 1998. The moving face during speech communication. (1998)."},{"key":"e_1_3_2_2_25_1","volume-title":"PyTorch Metric Learning. arxiv","author":"Musgrave Kevin","year":"2008","unstructured":"Kevin Musgrave , Serge Belongie , and Ser-Nam Lim . 2020. PyTorch Metric Learning. arxiv : 2008 .09164 [cs.CV] Kevin Musgrave, Serge Belongie, and Ser-Nam Lim. 2020. PyTorch Metric Learning. arxiv: 2008.09164 [cs.CV]"},{"key":"e_1_3_2_2_26_1","doi-asserted-by":"crossref","unstructured":"A. Nagrani S. Albanie and A. Zisserman. 2018a. Learnable PINs: Cross-Modal Embeddings for Person Identity. Springer Cham (2018).  A. Nagrani S. Albanie and A. Zisserman. 2018a. Learnable PINs: Cross-Modal Embeddings for Person Identity. Springer Cham (2018).","DOI":"10.1007\/978-3-030-01261-8_5"},{"key":"e_1_3_2_2_27_1","volume-title":"2018 IEEE\/CVF Conference on Computer Vision and Pattern Recognition .","author":"Nagrani A.","unstructured":"A. Nagrani , S. Albanie , and A. Zisserman . 2018b. Seeing Voices and Hearing Faces: Cross-modal biometric matching . In 2018 IEEE\/CVF Conference on Computer Vision and Pattern Recognition . A. Nagrani, S. Albanie, and A. Zisserman. 2018b. Seeing Voices and Hearing Faces: Cross-modal biometric matching. In 2018 IEEE\/CVF Conference on Computer Vision and Pattern Recognition ."},{"key":"e_1_3_2_2_28_1","doi-asserted-by":"crossref","unstructured":"A. Nagrani J. S. Chung and A. Zisserman. 2017. VoxCeleb: a large-scale speaker identification dataset. In Interspeech .  A. Nagrani J. S. Chung and A. Zisserman. 2017. VoxCeleb: a large-scale speaker identification dataset. In Interspeech .","DOI":"10.21437\/Interspeech.2017-950"},{"key":"e_1_3_2_2_29_1","doi-asserted-by":"crossref","unstructured":"S. Nawaz M. K. Janjua I. Gallo A. Mahmood and A. Calefati. 2019. Deep Latent Space Learning for Cross-modal Mapping of Audio and Visual Signals. (2019).  S. Nawaz M. K. Janjua I. Gallo A. Mahmood and A. Calefati. 2019. Deep Latent Space Learning for Cross-modal Mapping of Audio and Visual Signals. (2019).","DOI":"10.1109\/DICTA47822.2019.8945863"},{"key":"e_1_3_2_2_30_1","volume-title":"PyTorch: An Imperative Style","author":"Paszke Adam","unstructured":"Adam Paszke , Sam Gross , Francisco Massa , Adam Lerer , James Bradbury , Gregory Chanan , Trevor Killeen , Zeming Lin , Natalia Gimelshein , Luca Antiga , Alban Desmaison , Andreas Kopf , Edward Yang , Zachary DeVito , Martin Raison , Alykhan Tejani , Sasank Chilamkurthy , Benoit Steiner , Lu Fang , Junjie Bai , and Soumith Chintala . 2019. PyTorch: An Imperative Style , High-Performance Deep Learning Library . In Advances in Neural Information Processing Systems 32, H. Wallach, H. Larochelle, A. Beygelzimer, F. dtextquotesingle Alch\u00e9-Buc, E. Fox, and R. Garnett (Eds.). Curran Associates, Inc., 8024--8035. http:\/\/papers.neurips.cc\/paper\/9015-pytorch-an-imperative-style-high-performance-deep-learning-library.pdf Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, Alban Desmaison, Andreas Kopf, Edward Yang, Zachary DeVito, Martin Raison, Alykhan Tejani, Sasank Chilamkurthy, Benoit Steiner, Lu Fang, Junjie Bai, and Soumith Chintala. 2019. PyTorch: An Imperative Style, High-Performance Deep Learning Library. In Advances in Neural Information Processing Systems 32, H. Wallach, H. Larochelle, A. Beygelzimer, F. dtextquotesingle Alch\u00e9-Buc, E. Fox, and R. Garnett (Eds.). Curran Associates, Inc., 8024--8035. http:\/\/papers.neurips.cc\/paper\/9015-pytorch-an-imperative-style-high-performance-deep-learning-library.pdf"},{"key":"e_1_3_2_2_31_1","volume-title":"Renato De Mori, and Yoshua Bengio","author":"Ravanelli Mirco","year":"2021","unstructured":"Mirco Ravanelli , Titouan Parcollet , Peter Plantinga , Aku Rouhe , Samuele Cornell , Loren Lugosch , Cem Subakan , Nauman Dawalatabad , Abdelwahab Heba , Jianyuan Zhong , Ju-Chieh Chou , Sung-Lin Yeh , Szu-Wei Fu , Chien-Feng Liao , Elena Rastorgueva , Fran\u00e7ois Grondin , William Aris , Hwidong Na , Yan Gao , Renato De Mori, and Yoshua Bengio . 2021 . SpeechBrain: A General-Purpose Speech Toolkit . arxiv: 2106.04624 [eess.AS] arXiv:2106.04624. Mirco Ravanelli, Titouan Parcollet, Peter Plantinga, Aku Rouhe, Samuele Cornell, Loren Lugosch, Cem Subakan, Nauman Dawalatabad, Abdelwahab Heba, Jianyuan Zhong, Ju-Chieh Chou, Sung-Lin Yeh, Szu-Wei Fu, Chien-Feng Liao, Elena Rastorgueva, Fran\u00e7ois Grondin, William Aris, Hwidong Na, Yan Gao, Renato De Mori, and Yoshua Bengio. 2021. SpeechBrain: A General-Purpose Speech Toolkit. arxiv: 2106.04624 [eess.AS] arXiv:2106.04624."},{"key":"e_1_3_2_2_32_1","first-page":"868","article-title":"Matching novel face and voice identity using static and dynamic facial images. Attention, Perception","volume":"78","author":"Smith Hmj","year":"2016","unstructured":"Hmj Smith , A. K. Dunn , T. Baguley , and P. C. Stacey . 2016 . Matching novel face and voice identity using static and dynamic facial images. Attention, Perception , Psychophysics , Vol. 78 , 3 (2016), 868 -- 879 . Hmj Smith, A. K. Dunn, T. Baguley, and P. C. Stacey. 2016. Matching novel face and voice identity using static and dynamic facial images. Attention, Perception, Psychophysics, Vol. 78, 3 (2016), 868--879.","journal-title":"Psychophysics"},{"key":"e_1_3_2_2_33_1","volume-title":"Circle Loss: A Unified Perspective of Pair Similarity Optimization","author":"Sun Y.","year":"2020","unstructured":"Y. Sun , C. Cheng , Y. Zhang , C. Zhang , L. Zheng , Z. Wang , and Y. Wei . 2020 . Circle Loss: A Unified Perspective of Pair Similarity Optimization . In IEEE . Y. Sun, C. Cheng, Y. Zhang, C. Zhang, L. Zheng, Z. Wang, and Y. Wei. 2020. Circle Loss: A Unified Perspective of Pair Similarity Optimization. In IEEE ."},{"key":"e_1_3_2_2_34_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2015.7298594"},{"key":"e_1_3_2_2_35_1","volume-title":"Rohan Kumar Das, and Haizhou Li","author":"Tao Ruijie","year":"2020","unstructured":"Ruijie Tao , Rohan Kumar Das, and Haizhou Li . 2020 . Audio-visual speaker recognition with a cross-modal discriminative network. arXiv preprint arXiv:2008.03894 (2020). Ruijie Tao, Rohan Kumar Das, and Haizhou Li. 2020. Audio-visual speaker recognition with a cross-modal discriminative network. arXiv preprint arXiv:2008.03894 (2020)."},{"key":"e_1_3_2_2_36_1","volume-title":"Learning Discriminative Joint Embeddings for Efficient Face and Voice Association","author":"Wang Rui","year":"1881","unstructured":"Rui Wang , Xin Liu , Yiu-ming Cheung, Kai Cheng , Nannan Wang , and Wentao Fan . 2020. Learning Discriminative Joint Embeddings for Efficient Face and Voice Association . Association for Computing Machinery , New York, NY, USA , 1881 --1884. https:\/\/doi.org\/10.1145\/3397271.3401302 10.1145\/3397271.3401302 Rui Wang, Xin Liu, Yiu-ming Cheung, Kai Cheng, Nannan Wang, and Wentao Fan. 2020. Learning Discriminative Joint Embeddings for Efficient Face and Voice Association. Association for Computing Machinery, New York, NY, USA, 1881--1884. https:\/\/doi.org\/10.1145\/3397271.3401302"},{"key":"e_1_3_2_2_37_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.00516"},{"key":"e_1_3_2_2_38_1","doi-asserted-by":"crossref","unstructured":"P. Wen Q. Xu Y Jiang Z. Yang and Q. Huang. 2021. Seeking the Shape of Sound: An Adaptive Framework for Learning Voice-Face Association. (2021).  P. Wen Q. Xu Y Jiang Z. Yang and Q. Huang. 2021. Seeking the Shape of Sound: An Adaptive Framework for Learning Voice-Face Association. (2021).","DOI":"10.1109\/CVPR46437.2021.01608"},{"key":"e_1_3_2_2_39_1","unstructured":"Y. Wen M. A. Ismail W. Liu B. Raj and R. Singh. 2018. Disjoint Mapping Network for Cross-modal Matching of Voices and Faces. (2018).  Y. Wen M. A. Ismail W. Liu B. Raj and R. Singh. 2018. Disjoint Mapping Network for Cross-modal Matching of Voices and Faces. (2018)."},{"key":"e_1_3_2_2_40_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2017.309"},{"key":"e_1_3_2_2_41_1","volume-title":"Voice-Face Cross-modal Matching and Retrieval: A Benchmark. arxiv","author":"Xiong Chuyuan","year":"1911","unstructured":"Chuyuan Xiong , Deyuan Zhang , Tao Liu , and Xiaoyong Du. 2019. Voice-Face Cross-modal Matching and Retrieval: A Benchmark. arxiv : 1911 .09338 [cs.CV] Chuyuan Xiong, Deyuan Zhang, Tao Liu, and Xiaoyong Du. 2019. Voice-Face Cross-modal Matching and Retrieval: A Benchmark. arxiv: 1911.09338 [cs.CV]"},{"key":"e_1_3_2_2_42_1","doi-asserted-by":"crossref","unstructured":"H. Yehia P. Rubin and E. Vatikiotis-Bateson. 1998. Quantitative association of vocal-tract and facial behavior. Elsevier Science Publishers B. V. (1998).  H. Yehia P. Rubin and E. Vatikiotis-Bateson. 1998. Quantitative association of vocal-tract and facial behavior. Elsevier Science Publishers B. V. (1998).","DOI":"10.1016\/S0167-6393(98)00048-X"},{"key":"e_1_3_2_2_43_1","volume-title":"Barlow twins: Self-supervised learning via redundancy reduction. arXiv preprint arXiv:2103.03230","author":"Zbontar Jure","year":"2021","unstructured":"Jure Zbontar , Li Jing , Ishan Misra , Yann LeCun , and St\u00e9phane Deny . 2021. Barlow twins: Self-supervised learning via redundancy reduction. arXiv preprint arXiv:2103.03230 ( 2021 ). Jure Zbontar, Li Jing, Ishan Misra, Yann LeCun, and St\u00e9phane Deny. 2021. Barlow twins: Self-supervised learning via redundancy reduction. arXiv preprint arXiv:2103.03230 (2021)."},{"key":"e_1_3_2_2_44_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.01064"},{"key":"e_1_3_2_2_45_1","unstructured":"A. Zheng M. Hu B. Jiang Y. Huang and B. Luo. 2021. Adversarial-Metric Learning for Audio-Visual Cross-Modal Matching. IEEE Transactions on Multimedia Vol. PP 99 (2021) 1--1.  A. Zheng M. Hu B. Jiang Y. Huang and B. Luo. 2021. Adversarial-Metric Learning for Audio-Visual Cross-Modal Matching. IEEE Transactions on Multimedia Vol. PP 99 (2021) 1--1."}],"event":{"name":"ICMR '22: International Conference on Multimedia Retrieval","location":"Newark NJ USA","acronym":"ICMR '22","sponsor":["SIGMM ACM Special Interest Group on Multimedia"]},"container-title":["Proceedings of the 2022 International Conference on Multimedia Retrieval"],"original-title":[],"link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3512527.3531364","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3512527.3531364","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T19:30:12Z","timestamp":1750188612000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3512527.3531364"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,6,27]]},"references-count":45,"alternative-id":["10.1145\/3512527.3531364","10.1145\/3512527"],"URL":"https:\/\/doi.org\/10.1145\/3512527.3531364","relation":{},"subject":[],"published":{"date-parts":[[2022,6,27]]},"assertion":[{"value":"2022-06-27","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}