{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,4,22]],"date-time":"2026-04-22T20:26:02Z","timestamp":1776889562379,"version":"3.51.2"},"publisher-location":"New York, NY, USA","reference-count":49,"publisher":"ACM","license":[{"start":{"date-parts":[[2018,10,15]],"date-time":"2018-10-15T00:00:00Z","timestamp":1539561600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2018,10,15]]},"DOI":"10.1145\/3240508.3240601","type":"proceedings-article","created":{"date-parts":[[2018,10,18]],"date-time":"2018-10-18T13:52:08Z","timestamp":1539870728000},"page":"1011-1019","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":39,"title":["Face-Voice Matching using Cross-modal Embeddings"],"prefix":"10.1145","author":[{"given":"Shota","family":"Horiguchi","sequence":"first","affiliation":[{"name":"Hitachi, Ltd., Tokyo, Japan"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Naoyuki","family":"Kanda","sequence":"additional","affiliation":[{"name":"Hitachi, Ltd., Tokyo, Japan"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Kenji","family":"Nagamatsu","sequence":"additional","affiliation":[{"name":"Hitachi, Ltd., Tokyo, Japan"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2018,10,15]]},"reference":[{"key":"e_1_3_2_1_1_1","doi-asserted-by":"publisher","DOI":"10.1007\/11677482_3"},{"key":"e_1_3_2_1_2_1","doi-asserted-by":"crossref","unstructured":"Llu\u00eds Castrej\u00f3n Yusuf Aytar Carl Vondrick Hamed Pirsiavash and Antonio Torralba. 2016. Learning Aligned Cross-Modal Representations from Weakly Aligned Data. In CVPR . 2940--2949. Llu\u00eds Castrej\u00f3n Yusuf Aytar Carl Vondrick Hamed Pirsiavash and Antonio Torralba. 2016. Learning Aligned Cross-Modal Representations from Weakly Aligned Data. In CVPR . 2940--2949.","DOI":"10.1109\/CVPR.2016.321"},{"key":"e_1_3_2_1_3_1","doi-asserted-by":"publisher","DOI":"10.1049\/cp:19970924"},{"key":"e_1_3_2_1_4_1","doi-asserted-by":"publisher","DOI":"10.1121\/1.2229005"},{"key":"e_1_3_2_1_5_1","doi-asserted-by":"publisher","DOI":"10.1109\/TASL.2010.2064307"},{"key":"e_1_3_2_1_6_1","doi-asserted-by":"crossref","unstructured":"Xiao Fang Najim Dehak and James Glass. 2013. Bayesian Distance Metric Learning on i-vector for Speaker Verification. In Interspeech . 2514--2518. Xiao Fang Najim Dehak and James Glass. 2013. Bayesian Distance Metric Learning on i-vector for Speaker Verification. In Interspeech . 2514--2518.","DOI":"10.21437\/Interspeech.2013-421"},{"key":"e_1_3_2_1_7_1","doi-asserted-by":"publisher","DOI":"10.1109\/TMM.2004.827503"},{"key":"e_1_3_2_1_8_1","volume-title":"IEEE TPAMI","volume":"39","author":"Gebru Israel D.","year":"2017"},{"key":"e_1_3_2_1_9_1","unstructured":"Yandong Guo and Lei Zhang. 2017. One-shot Face Recognition by Promoting Underrepresented Classes. arXiv:1707.05574. (2017). Yandong Guo and Lei Zhang. 2017. One-shot Face Recognition by Promoting Underrepresented Classes. arXiv:1707.05574. (2017)."},{"key":"e_1_3_2_1_10_1","doi-asserted-by":"crossref","unstructured":"Yandong Guo Lei Zhang Yuxiao Hu Xiaodong He and Jianfeng Gao. 2016. MS-Celeb-1M: A Dataset and Benchmark for Large Scale Face Recognition. In ECCV . 87--102. Yandong Guo Lei Zhang Yuxiao Hu Xiaodong He and Jianfeng Gao. 2016. MS-Celeb-1M: A Dataset and Benchmark for Large Scale Face Recognition. In ECCV . 87--102.","DOI":"10.1007\/978-3-319-46487-9_6"},{"key":"e_1_3_2_1_11_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2006.100"},{"key":"e_1_3_2_1_12_1","doi-asserted-by":"crossref","unstructured":"Kaiming He Xiangyu Zhang Shaoqing Ren and Jian Sun. 2016. Deep Residual Learning for Image Recognition. In CVPR. 770--778. Kaiming He Xiangyu Zhang Shaoqing Ren and Jian Sun. 2016. Deep Residual Learning for Image Recognition. In CVPR. 770--778.","DOI":"10.1109\/CVPR.2016.90"},{"key":"e_1_3_2_1_13_1","unstructured":"Weipeng He Petr Motlicek and Jean-Marc Odobez. 2018. Deep neural networks for multiple speaker detection and localization. In ICRA . Weipeng He Petr Motlicek and Jean-Marc Odobez. 2018. Deep neural networks for multiple speaker detection and localization. In ICRA ."},{"key":"e_1_3_2_1_14_1","unstructured":"Ken Hoover Sourish Chaudhuri Caroline Pantofaru Malcolm Slaney and Ian Sturdy. 2017. Putting a Face to the Voice: Fusing Audio and Visual Signals Across a Video to Determine Speakers. arXiv:1706.00079. (2017). Ken Hoover Sourish Chaudhuri Caroline Pantofaru Malcolm Slaney and Ian Sturdy. 2017. Putting a Face to the Voice: Fusing Audio and Visual Signals Across a Video to Determine Speakers. arXiv:1706.00079. (2017)."},{"key":"e_1_3_2_1_15_1","unstructured":"Itseez. 2015. Open Source Computer Vision Library. https:\/\/github.com\/itseez\/opencv . (2015). Itseez. 2015. Open Source Computer Vision Library. https:\/\/github.com\/itseez\/opencv . (2015)."},{"key":"e_1_3_2_1_16_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.cub.2003.09.005"},{"key":"e_1_3_2_1_17_1","doi-asserted-by":"publisher","DOI":"10.1109\/TASL.2006.881693"},{"key":"e_1_3_2_1_18_1","doi-asserted-by":"publisher","DOI":"10.1109\/TASL.2007.894527"},{"key":"e_1_3_2_1_19_1","doi-asserted-by":"publisher","DOI":"10.1109\/TASL.2008.925147"},{"key":"e_1_3_2_1_20_1","doi-asserted-by":"crossref","unstructured":"Volkan Kilicc and Wenwu Wang. 2017. Audio-Visual Speaker Tracking. Motion Tracking and Gesture Recognition. InTech Chapter 03. Volkan Kilicc and Wenwu Wang. 2017. Audio-Visual Speaker Tracking. Motion Tracking and Gesture Recognition. InTech Chapter 03.","DOI":"10.5772\/intechopen.68146"},{"key":"e_1_3_2_1_21_1","doi-asserted-by":"publisher","DOI":"10.1016\/S0022-1031(02)00510-3"},{"key":"e_1_3_2_1_22_1","first-page":"159","article-title":"Crossmodal Source Identification in Speech Perception","volume":"16","author":"Lachs Lorin","year":"2004","journal-title":"Ecologial Psycology"},{"key":"e_1_3_2_1_23_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICIP.2002.1038171"},{"key":"e_1_3_2_1_24_1","first-page":"2579","article-title":"Visualizing data using t-SNE","volume":"9","author":"van der Maaten Laurens","year":"2008","journal-title":"JMLR"},{"key":"e_1_3_2_1_25_1","first-page":"307","article-title":"Matching Voice and Face Identity From Static Images","volume":"39","author":"Mavica Lauren W.","year":"2013","journal-title":"Journal of Experimental Psychology"},{"key":"e_1_3_2_1_26_1","first-page":"214","article-title":"Robust Multimodal Person Identification With Limited Training Data","volume":"43","author":"McLaughlin Niall","year":"2013","journal-title":"IEEE HMS"},{"key":"e_1_3_2_1_27_1","unstructured":"Gianluca Monaci. 2011. Towards Real-Time Audiovisual Speaker Localization. In EUSIPCO . 1055--1059. Gianluca Monaci. 2011. Towards Real-Time Audiovisual Speaker Localization. In EUSIPCO . 1055--1059."},{"key":"e_1_3_2_1_28_1","unstructured":"Mozilla. 2017. Common Voice. https:\/\/voice.mozilla.org\/en. (2017). Mozilla. 2017. Common Voice. https:\/\/voice.mozilla.org\/en. (2017)."},{"key":"e_1_3_2_1_29_1","doi-asserted-by":"crossref","unstructured":"Arsha Nagrani Samuel Albanie and Andrew Zisserman. 2018. Seeing Voices and Hearing Faces: Cross-modal biometric matching. In CVPR . 8427--8436. Arsha Nagrani Samuel Albanie and Andrew Zisserman. 2018. Seeing Voices and Hearing Faces: Cross-modal biometric matching. In CVPR . 8427--8436.","DOI":"10.1109\/CVPR.2018.00879"},{"key":"e_1_3_2_1_30_1","doi-asserted-by":"crossref","unstructured":"Arsha Nagrani Joon Son Chung and Andrew Zisserman. 2017. VoxCeleb: A large-scale speaker identification dataset. In Interspeech . 2616--2620. Arsha Nagrani Joon Son Chung and Andrew Zisserman. 2017. VoxCeleb: A large-scale speaker identification dataset. In Interspeech . 2616--2620.","DOI":"10.21437\/Interspeech.2017-950"},{"key":"e_1_3_2_1_31_1","volume-title":"Librispeech: An ASR corpus based on public domain audio books. In ICASSP . 5206--5210.","author":"Panayotov Vassil","year":"2015"},{"key":"e_1_3_2_1_32_1","doi-asserted-by":"crossref","unstructured":"Amaia Salvador Nicholas Hynes Yusuf Aytar Javier Marin Ferda Ofli Ingmar Weber and Antonio Torralba. 2017. Learning Cross-modal Embeddings for Cooking Recipes and Food Images. In CVPR . 3020--3028. Amaia Salvador Nicholas Hynes Yusuf Aytar Javier Marin Ferda Ofli Ingmar Weber and Antonio Torralba. 2017. Learning Cross-modal Embeddings for Cooking Recipes and Food Images. In CVPR . 3020--3028.","DOI":"10.1109\/CVPR.2017.327"},{"key":"e_1_3_2_1_33_1","doi-asserted-by":"crossref","unstructured":"Jordi Sanchez-Riera Xavier Alameda-Pineda Johannes Wienke Antoine Deleforge Soraya Arias Jan vC ech Sebastian Wrede and Radu Horaud. 2012. Online multimodal speaker detection for humanoid robots. In Humanoids. IEEE 126--133. Jordi Sanchez-Riera Xavier Alameda-Pineda Johannes Wienke Antoine Deleforge Soraya Arias Jan vC ech Sebastian Wrede and Radu Horaud. 2012. Online multimodal speaker detection for humanoid robots. In Humanoids. IEEE 126--133.","DOI":"10.1109\/HUMANOIDS.2012.6651509"},{"key":"e_1_3_2_1_34_1","doi-asserted-by":"crossref","unstructured":"Florian Schroff Dmitry Kalenichenko and James Philbin. 2015. FaceNet: A unified embedding for face recognition and clustering. In CVPR . 815--823. Florian Schroff Dmitry Kalenichenko and James Philbin. 2015. FaceNet: A unified embedding for face recognition and clustering. In CVPR . 815--823.","DOI":"10.1109\/CVPR.2015.7298682"},{"key":"e_1_3_2_1_35_1","doi-asserted-by":"publisher","DOI":"10.1109\/TSP.2004.838930"},{"key":"e_1_3_2_1_36_1","doi-asserted-by":"publisher","DOI":"10.1109\/TCSVT.2008.2009262"},{"key":"e_1_3_2_1_37_1","unstructured":"Harriet M. J. Smith. 2016. Matching novel face and voice identity using static and dynamic facial images . Ph.D. Dissertation. Nottingham Trent University. Harriet M. J. Smith. 2016. Matching novel face and voice identity using static and dynamic facial images . Ph.D. Dissertation. Nottingham Trent University."},{"key":"e_1_3_2_1_38_1","doi-asserted-by":"publisher","DOI":"10.1177\/1474704916630317"},{"key":"e_1_3_2_1_39_1","doi-asserted-by":"publisher","DOI":"10.3758\/s13414-015-1045-8"},{"key":"e_1_3_2_1_40_1","unstructured":"Kihyuk Sohn. 2016. Improved Deep Metric Learning with Multi-class N-pair Loss Objective. In NIPS . 1857--1865. Kihyuk Sohn. 2016. Improved Deep Metric Learning with Multi-class N-pair Loss Objective. In NIPS . 1857--1865."},{"key":"e_1_3_2_1_41_1","doi-asserted-by":"crossref","unstructured":"Hyun Oh Song Stefanie Jegelka Vivek Rathod and Kevin Murphy. 2017. Deep Metric Learning via Facility Location. In CVPR. 5382--5390. Hyun Oh Song Stefanie Jegelka Vivek Rathod and Kevin Murphy. 2017. Deep Metric Learning via Facility Location. In CVPR. 5382--5390.","DOI":"10.1109\/CVPR.2017.237"},{"key":"e_1_3_2_1_42_1","doi-asserted-by":"crossref","unstructured":"Hyun Oh Song Yu Xiang Stefanie Jegelka and Silvio Savarese. 2016. Deep Metric Learning via Lifted Structured Feature Embedding. In CVPR. 4004--4012. Hyun Oh Song Yu Xiang Stefanie Jegelka and Silvio Savarese. 2016. Deep Metric Learning via Lifted Structured Feature Embedding. In CVPR. 4004--4012.","DOI":"10.1109\/CVPR.2016.434"},{"key":"e_1_3_2_1_43_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2014.220"},{"key":"e_1_3_2_1_44_1","doi-asserted-by":"crossref","unstructured":"Ryu Takeda and Kazunori Komatani. 2016. Sound source localization based on deep neural networks with directional activate function exploiting phase information. In ICASSP. 405--409. Ryu Takeda and Kazunori Komatani. 2016. Sound source localization based on deep neural networks with directional activate function exploiting phase information. In ICASSP. 405--409.","DOI":"10.1109\/ICASSP.2016.7471706"},{"key":"e_1_3_2_1_45_1","volume-title":"Research Blog: Launching the Speech Commands Dataset. https:\/\/research.googleblog.com\/2017\/08\/launching-speech-commands-dataset.html.","author":"Warden Pete","year":"2017"},{"key":"e_1_3_2_1_46_1","unstructured":"Kilian Q Weinberger John Blitzer and Lawrence K Saul. 2006. Distance metric learning for large margin nearest neighbor classification. In NIPS . 1473--1480. Kilian Q Weinberger John Blitzer and Lawrence K Saul. 2006. Distance metric learning for large margin nearest neighbor classification. In NIPS . 1473--1480."},{"key":"e_1_3_2_1_47_1","doi-asserted-by":"crossref","unstructured":"Yandong Wen Kaipeng Zhang Zhifeng Li and Yu Qiao. 2016. A discriminative feature learning approach for deep face recognition. In ECCV . 499--515. Yandong Wen Kaipeng Zhang Zhifeng Li and Yu Qiao. 2016. A discriminative feature learning approach for deep face recognition. In ECCV . 499--515.","DOI":"10.1007\/978-3-319-46478-7_31"},{"key":"e_1_3_2_1_48_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2011.5995566"},{"key":"e_1_3_2_1_49_1","doi-asserted-by":"crossref","unstructured":"Xiao Zhang Zhiyuan Fang Yandong Wen Zhifeng Li and Yu Qiao. 2017. Range loss for deep face recognition with long-tailed training data. In ICCV . 5409--5418. Xiao Zhang Zhiyuan Fang Yandong Wen Zhifeng Li and Yu Qiao. 2017. Range loss for deep face recognition with long-tailed training data. In ICCV . 5409--5418.","DOI":"10.1109\/ICCV.2017.578"}],"event":{"name":"MM '18: ACM Multimedia Conference","location":"Seoul Republic of Korea","acronym":"MM '18","sponsor":["SIGMM ACM Special Interest Group on Multimedia"]},"container-title":["Proceedings of the 26th ACM international conference on Multimedia"],"original-title":[],"link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3240508.3240601","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3240508.3240601","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,4,3]],"date-time":"2026-04-03T20:40:37Z","timestamp":1775248837000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3240508.3240601"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2018,10,15]]},"references-count":49,"alternative-id":["10.1145\/3240508.3240601","10.1145\/3240508"],"URL":"https:\/\/doi.org\/10.1145\/3240508.3240601","relation":{},"subject":[],"published":{"date-parts":[[2018,10,15]]},"assertion":[{"value":"2018-10-15","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}