{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,10,27]],"date-time":"2025-10-27T21:06:39Z","timestamp":1761599199008,"version":"3.41.0"},"reference-count":50,"publisher":"Association for Computing Machinery (ACM)","issue":"2","license":[{"start":{"date-parts":[[2020,5,22]],"date-time":"2020-05-22T00:00:00Z","timestamp":1590105600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"DOI":"10.13039\/501100012226","name":"Fundamental Research Funds for the Central Universities","doi-asserted-by":"crossref","award":["2019XD-A12"],"award-info":[{"award-number":["2019XD-A12"]}],"id":[{"id":"10.13039\/501100012226","id-type":"DOI","asserted-by":"crossref"}]},{"DOI":"10.13039\/501100001809","name":"National Nature Science Foundation of China","doi-asserted-by":"crossref","award":["61876135, 61801335 and U1611461"],"award-info":[{"award-number":["61876135, 61801335 and U1611461"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]},{"name":"National Key R8D Program of China","award":["2017YFC0803700"],"award-info":[{"award-number":["2017YFC0803700"]}]},{"DOI":"10.13039\/501100012239","name":"Hubei Province Technological Innovation Major Project","doi-asserted-by":"crossref","award":["2018AAA062"],"award-info":[{"award-number":["2018AAA062"]}],"id":[{"id":"10.13039\/501100012239","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Multimedia Comput. Commun. Appl."],"published-print":{"date-parts":[[2020,5,31]]},"abstract":"<jats:p>Person search with one portrait, which attempts to search the targets in arbitrary scenes using one portrait image at a time, is an essential yet unexplored problem in the multimedia field. Existing approaches, which predominantly depend on the visual information of persons, cannot solve problems when there are variations in the person\u2019s appearance caused by complex environments and changes in pose, makeup, and clothing. In contrast to existing methods, in this article, we propose an associative multimodality index for person search with face, body, and voice information. In the offline stage, an associative network is proposed to learn the relationships among face, body, and voice information. It can adaptively estimate the weights of each embedding to construct an appropriate representation. The multimodality index can be built by using these representations, which exploit the face and voice as long-term keys and the body appearance as a short-term connection. In the online stage, through the multimodality association in the index, we can retrieve all targets depending only on the facial features of the query portrait. Furthermore, to evaluate our multimodality search framework and facilitate related research, we construct the Cast Search in Movies with Voice (CSM-V) dataset, a large-scale benchmark that contains 127K annotated voices corresponding to tracklets from 192 movies. According to extensive experiments on the CSM-V dataset, the proposed multimodality person search framework outperforms the state-of-the-art methods.<\/jats:p>","DOI":"10.1145\/3380549","type":"journal-article","created":{"date-parts":[[2020,5,25]],"date-time":"2020-05-25T22:07:21Z","timestamp":1590444441000},"page":"1-20","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":15,"title":["Listen, Look, and Find the One"],"prefix":"10.1145","volume":"16","author":[{"ORCID":"https:\/\/orcid.org\/0000-0003-0770-9891","authenticated-orcid":false,"given":"Xiao","family":"Wang","sequence":"first","affiliation":[{"name":"NERCMS, School of Computer Sicence, Wuhan University"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Wu","family":"Liu","sequence":"additional","affiliation":[{"name":"AI Research of JD.com"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Jun","family":"Chen","sequence":"additional","affiliation":[{"name":"NERCMS, School of Computer Science, Wuhan University"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Xiaobo","family":"Wang","sequence":"additional","affiliation":[{"name":"AI Research of JD.com"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Chenggang","family":"Yan","sequence":"additional","affiliation":[{"name":"Hangzhou Dianzi University"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Tao","family":"Mei","sequence":"additional","affiliation":[{"name":"AI Research of JD.com"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2020,5,22]]},"reference":[{"key":"e_1_2_1_1_1","doi-asserted-by":"publisher","DOI":"10.1109\/TMM.2017.2751969"},{"volume-title":"Proceedings of the CCIS. 480--484","author":"Li S.","key":"e_1_2_1_2_1"},{"key":"e_1_2_1_3_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2017.2709749"},{"key":"e_1_2_1_4_1","doi-asserted-by":"publisher","DOI":"10.1109\/TMM.2018.2866370"},{"key":"e_1_2_1_5_1","doi-asserted-by":"publisher","DOI":"10.1109\/TMM.2018.2872897"},{"key":"e_1_2_1_6_1","doi-asserted-by":"publisher","DOI":"10.1007\/s12021-018-9362-4"},{"volume-title":"Proceedings of the ECCV. 425--441","author":"Huang Q.","key":"e_1_2_1_7_1"},{"key":"e_1_2_1_8_1","unstructured":"C. Loy D. Lin and W. Ouyang. 2018. WIDER face and pedestrian challenge: http:\/\/wider-challenge.org\/. arXiv:1902.06854 (2018).  C. Loy D. Lin and W. Ouyang. 2018. WIDER face and pedestrian challenge: http:\/\/wider-challenge.org\/. arXiv:1902.06854 (2018)."},{"key":"e_1_2_1_9_1","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2017.2675341"},{"key":"e_1_2_1_10_1","doi-asserted-by":"publisher","DOI":"10.1109\/TIFS.2018.2866295"},{"key":"e_1_2_1_11_1","first-page":"1534","article-title":"KinectFaceDB: A kinect database for face recognition","volume":"44","author":"Rui M.","year":"2017","journal-title":"IEEE Trans. Syst. Man, Cybern."},{"key":"e_1_2_1_12_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2017.2652466"},{"key":"e_1_2_1_13_1","first-page":"1","article-title":"Person reidentification via discrepancy matrix and matrix metric","volume":"1","author":"Wang Z.","year":"2017","journal-title":"IEEE Trans. Cybern."},{"key":"e_1_2_1_14_1","doi-asserted-by":"crossref","unstructured":"A. Torfi N. Nasrabadi and J. Dawson. 2017. Text-independent speaker verification using 3D convolutional neural networks. arXiv:1705.09422 (2017).  A. Torfi N. Nasrabadi and J. Dawson. 2017. Text-independent speaker verification using 3D convolutional neural networks. arXiv:1705.09422 (2017).","DOI":"10.1109\/ICME.2018.8486441"},{"key":"e_1_2_1_15_1","unstructured":"WVU multimodal dataset. Retrieved from http:\/\/biic.wvu.edu.  WVU multimodal dataset. Retrieved from http:\/\/biic.wvu.edu."},{"volume-title":"Proceedings of the CVPR. 8427--8436","author":"Nagrani A.","key":"e_1_2_1_16_1"},{"volume-title":"Proceedings of the CVPR. 3415--3424","author":"Xiao T.","key":"e_1_2_1_17_1"},{"volume-title":"Proceedings of the CVPR. 1367--1376","author":"Zheng L.","key":"e_1_2_1_18_1"},{"volume-title":"Proceedings of the ICCV. 493--501","author":"Liu H.","key":"e_1_2_1_19_1"},{"key":"e_1_2_1_20_1","doi-asserted-by":"crossref","unstructured":"B. Munjal S. Amin F. Tombari and F. Galasso. 2019. Query-guided end-to-end person search. (2019) 811--820.  B. Munjal S. Amin F. Tombari and F. Galasso. 2019. Query-guided end-to-end person search. (2019) 811--820.","DOI":"10.1109\/CVPR.2019.00090"},{"volume-title":"Proceedings of the ACM Multimedia. 1--10","author":"Horiguchi S.","key":"e_1_2_1_21_1"},{"volume-title":"Proceedings of the CVPR. 87--97","author":"Gan C.","key":"e_1_2_1_22_1"},{"key":"e_1_2_1_23_1","doi-asserted-by":"crossref","unstructured":"R. Arandjelovic and A. Zisserman. 2017. Look listen and learn. arXiv:1705.08168 (2017).  R. Arandjelovic and A. Zisserman. 2017. Look listen and learn. arXiv:1705.08168 (2017).","DOI":"10.1109\/ICCV.2017.73"},{"volume-title":"Proceedings of the ICCV. 7053--7062","author":"Gan C.","key":"e_1_2_1_24_1"},{"volume-title":"Proceedings of the ECCV. 570--586","author":"Zhao H.","key":"e_1_2_1_25_1"},{"volume-title":"Proceedings of the IJCB. 1--9.","author":"Zhang S.","key":"e_1_2_1_26_1"},{"volume-title":"FDDB: A benchmark for face detection in unconstrained settings. In UMass Amherst Technical Report. 1--6.","year":"2010","author":"Jain V.","key":"e_1_2_1_27_1"},{"volume-title":"Proceedings of the CVPR. 2235--2245","author":"Feng Z.","key":"e_1_2_1_28_1"},{"volume-title":"Proceedings of the ICCV Workshops. 1--7.","author":"Sagonas C.","key":"e_1_2_1_29_1"},{"volume-title":"Proceedings of the ECCV. 765--780","author":"Wang F.","key":"e_1_2_1_30_1"},{"key":"e_1_2_1_31_1","unstructured":"Retrieved from http:\/\/trillionpairs.deepglint.com\/overview. ([n. d.]).  Retrieved from http:\/\/trillionpairs.deepglint.com\/overview. ([n. d.])."},{"volume-title":"Proceedings of the CVPR. 770--778","author":"He K.","key":"e_1_2_1_32_1"},{"volume-title":"Proceedings of the AAAI. 7138--7145","author":"Liu K.","key":"e_1_2_1_33_1"},{"volume-title":"Accelerating deep network training by reducing internal covariate shift. CoRR.abs\/1502.03167","year":"2015","author":"Normalization B.","key":"e_1_2_1_34_1"},{"volume-title":"Proceedings of the CVPR. 7834--7843","author":"Long X.","key":"e_1_2_1_35_1"},{"key":"e_1_2_1_36_1","doi-asserted-by":"crossref","unstructured":"A. Nagrani J. Chung and A. Zisserman. 2017. VoxCeleb: A large-scale speaker identification dataset. (2017) 2616--2620.  A. Nagrani J. Chung and A. Zisserman. 2017. VoxCeleb: A large-scale speaker identification dataset. (2017) 2616--2620.","DOI":"10.21437\/Interspeech.2017-950"},{"key":"e_1_2_1_37_1","unstructured":"J. Chung A. Nagrani and A. Zisserman. VoxCeleb2: Deep speaker recognition. ([n. d.]) 1086--1090.  J. Chung A. Nagrani and A. Zisserman. VoxCeleb2: Deep speaker recognition. ([n. d.]) 1086--1090."},{"key":"e_1_2_1_38_1","doi-asserted-by":"publisher","DOI":"10.1038\/nrn755"},{"key":"e_1_2_1_39_1","unstructured":"H. Ke D. Chen T. Shah X. Liu X. Zhang L. Zhang and X. Li. 2018. Cloud aided online EEG classification system for brain healthcare: A case study of depression evaluation with a lightweight CNN. Softw: Pract Exper. (2018) 1--15.  H. Ke D. Chen T. Shah X. Liu X. Zhang L. Zhang and X. Li. 2018. Cloud aided online EEG classification system for brain healthcare: A case study of depression evaluation with a lightweight CNN. Softw: Pract Exper. (2018) 1--15."},{"key":"e_1_2_1_40_1","first-page":"44","article-title":"Hearing faces and seeing voices: Amodal coding of person identity in the human brain. Sci","volume":"108","author":"Hasan B.","year":"2016","journal-title":"Rep."},{"key":"e_1_2_1_41_1","article-title":"Incremental factorization of big time series data with blind factor approximation","volume":"10","author":"Chen D.","year":"2019","journal-title":"IEEE Trans. Knowl. Data Eng. DOI"},{"key":"e_1_2_1_42_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPDS.2016.2613054"},{"volume-title":"Proceedings of the ACM Multimedia. 1618--1626","author":"Liu W.","key":"e_1_2_1_43_1"},{"volume-title":"Proceedings of the ACM Multimedia. 887--896","author":"Liu W.","key":"e_1_2_1_44_1"},{"volume-title":"Proceedings of the BigComp. 1--8.","author":"Liu J.","key":"e_1_2_1_45_1"},{"volume-title":"Proceedings of the ECCV. 868--884","author":"Zheng L.","key":"e_1_2_1_46_1"},{"key":"e_1_2_1_47_1","doi-asserted-by":"publisher","DOI":"10.1007\/s11263-015-0816-y"},{"key":"e_1_2_1_48_1","unstructured":"Y. Liu P. Shi B. Peng H. Yan Y. Zhou B. Han Y. Zheng C. Lin J. Jiang and Y. Fan. 2018. iQIYI-VID: A large dataset for multi-modal person identification. arXiv:1811.07548 (2018).  Y. Liu P. Shi B. Peng H. Yan Y. Zhou B. Han Y. Zheng C. Lin J. Jiang and Y. Fan. 2018. iQIYI-VID: A large dataset for multi-modal person identification. arXiv:1811.07548 (2018)."},{"key":"e_1_2_1_49_1","doi-asserted-by":"publisher","DOI":"10.1109\/LSP.2016.2603342"},{"key":"e_1_2_1_50_1","doi-asserted-by":"publisher","DOI":"10.2352\/ISSN.2470-1173.2016.11.IMAWM-463"}],"container-title":["ACM Transactions on Multimedia Computing, Communications, and Applications"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3380549","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3380549","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T21:31:32Z","timestamp":1750195892000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3380549"}},"subtitle":["Robust Person Search with Multimodality Index"],"short-title":[],"issued":{"date-parts":[[2020,5,22]]},"references-count":50,"journal-issue":{"issue":"2","published-print":{"date-parts":[[2020,5,31]]}},"alternative-id":["10.1145\/3380549"],"URL":"https:\/\/doi.org\/10.1145\/3380549","relation":{},"ISSN":["1551-6857","1551-6865"],"issn-type":[{"type":"print","value":"1551-6857"},{"type":"electronic","value":"1551-6865"}],"subject":[],"published":{"date-parts":[[2020,5,22]]},"assertion":[{"value":"2019-10-01","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2020-01-01","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2020-05-22","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}