{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,4,27]],"date-time":"2026-04-27T10:39:50Z","timestamp":1777286390454,"version":"3.51.4"},"reference-count":81,"publisher":"Springer Science and Business Media LLC","issue":"10","license":[{"start":{"date-parts":[[2022,1,5]],"date-time":"2022-01-05T00:00:00Z","timestamp":1641340800000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2022,1,5]],"date-time":"2022-01-05T00:00:00Z","timestamp":1641340800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"DOI":"10.13039\/501100003407","name":"ministero dell\u2019istruzione, dell\u2019universit\u00e0 e della ricerca","doi-asserted-by":"publisher","award":["D48G18000150006"],"award-info":[{"award-number":["D48G18000150006"]}],"id":[{"id":"10.13039\/501100003407","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Neural Comput &amp; Applic"],"published-print":{"date-parts":[[2022,5]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:p>Face-based video retrieval (FBVR) is the task of retrieving videos that containing the same face shown in the query image. In this article, we present the first end-to-end FBVR pipeline that is able to operate on large datasets of unconstrained, multi-shot, multi-person videos. We adapt an existing audiovisual recognition dataset to the task of FBVR and use it to evaluate our proposed pipeline. We compare a number of deep learning models for shot detection, face detection, and face feature extraction as part of our pipeline on a validation dataset made of more than 4000 videos. We obtain 97.25% mean average precision on an independent test set, composed of more than 1000 videos. The pipeline is able to extract features from videos at<jats:inline-formula><jats:alternatives><jats:tex-math>$$\\sim $$<\/jats:tex-math><mml:math xmlns:mml=\"http:\/\/www.w3.org\/1998\/Math\/MathML\"><mml:mo>\u223c<\/mml:mo><\/mml:math><\/jats:alternatives><\/jats:inline-formula>7 times the real-time speed, and it is able to perform a query on thousands of videos in less than 0.5 s.<\/jats:p>","DOI":"10.1007\/s00521-021-06875-x","type":"journal-article","created":{"date-parts":[[2022,1,5]],"date-time":"2022-01-05T22:03:01Z","timestamp":1641420181000},"page":"7489-7506","update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":8,"title":["A comparison of deep learning models for end-to-end face-based video retrieval in unconstrained videos"],"prefix":"10.1007","volume":"34","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-5221-636X","authenticated-orcid":false,"given":"Gioele","family":"Ciaparrone","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Leonardo","family":"Chiariglione","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Roberto","family":"Tagliaferri","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"297","published-online":{"date-parts":[[2022,1,5]]},"reference":[{"key":"6875_CR1","doi-asserted-by":"publisher","unstructured":"Herrmann C, Beyerer J (2015) Face retrieval on large-scale video data. In: 2015 12th conference on computer and robot vision, pp 192\u2013199. IEEE . https:\/\/doi.org\/10.1109\/CRV.2015.32","DOI":"10.1109\/CRV.2015.32"},{"key":"6875_CR2","doi-asserted-by":"publisher","unstructured":"Li Y, Wang R, Shan S, Chen X (2015) Hierarchical hybrid statistic based video binary code and its application to face retrieval in TV-series. In: 2015 11th IEEE international conference and workshops on automatic face and gesture recognition (FG), vol.\u00a01, pp 1\u20138. IEEE . https:\/\/doi.org\/10.1109\/FG.2015.7163089","DOI":"10.1109\/FG.2015.7163089"},{"key":"6875_CR3","doi-asserted-by":"publisher","unstructured":"Li Y, Wang R, Huang Z, Shan S, Chen X (2015) Face video retrieval with image query via hashing across Euclidean space and Riemannian manifold. In: Proceedings of the IEEE conference on computer vision and pattern recognition, pp 4758\u20134767. https:\/\/doi.org\/10.1109\/CVPR.2015.7299108","DOI":"10.1109\/CVPR.2015.7299108"},{"issue":"12","key":"6875_CR4","doi-asserted-by":"publisher","first-page":"5905","DOI":"10.1109\/TIP.2016.2616297","volume":"25","author":"Y Li","year":"2016","unstructured":"Li Y, Wang R, Cui Z, Shan S, Chen X (2016) Spatial pyramid covariance-based compact video code for robust face retrieval in TV-series. IEEE Trans Image Process 25(12):5905\u20135919. https:\/\/doi.org\/10.1109\/TIP.2016.2616297","journal-title":"IEEE Trans Image Process"},{"key":"6875_CR5","doi-asserted-by":"publisher","unstructured":"Jing C, Dong Z, Pei M, Jia Y (2017) Fusing appearance features and correlation features for face video retrieval. In: Pacific rim conference on multimedia, pp 150\u2013160. Springer . https:\/\/doi.org\/10.1007\/978-3-319-77383-4_15","DOI":"10.1007\/978-3-319-77383-4_15"},{"key":"6875_CR6","doi-asserted-by":"publisher","first-page":"357","DOI":"10.1016\/j.patcog.2018.04.014","volume":"81","author":"Z Dong","year":"2018","unstructured":"Dong Z, Jing C, Pei M, Jia Y (2018) Deep CNN based binary hash video representations for face retrieval. Pattern Recogn 81:357\u2013369. https:\/\/doi.org\/10.1016\/j.patcog.2018.04.014","journal-title":"Pattern Recogn"},{"key":"6875_CR7","doi-asserted-by":"publisher","unstructured":"Chung JS, Nagrani A, (2018) isserman A VoxCeleb2: deep speaker recognition. In: Proceedings of the 19th annual conference of the international speech communication association, vol\u00a01, pp 1086\u20131090 . https:\/\/doi.org\/10.21437\/Interspeech.2018-1929","DOI":"10.21437\/Interspeech.2018-1929"},{"key":"6875_CR8","doi-asserted-by":"publisher","unstructured":"Arandjelovi\u0107, O., Zisserman, A.: Automatic face recognition for film character retrieval in feature-length films. In: 2005 IEEE computer society conference on computer vision and pattern recognition (CVPR\u201905), vol\u00a01, pp 860\u2013867. IEEE (2005). https:\/\/doi.org\/10.1109\/CVPR.2005.81","DOI":"10.1109\/CVPR.2005.81"},{"key":"6875_CR9","doi-asserted-by":"publisher","unstructured":"Arandjelovi\u0107 O, Zisserman A On film character retrieval in feature-length films. In: Interactive video, pp 89\u2013105. Springer (2006). https:\/\/doi.org\/10.1007\/978-3-540-33215-2_5","DOI":"10.1007\/978-3-540-33215-2_5"},{"key":"6875_CR10","doi-asserted-by":"publisher","unstructured":"Sivic J, Everingham M, Zisserman A (2005) erson spotting: video shot retrieval for face sets. In: International conference on image and video retrieval, pp 226\u2013236. Springer . https:\/\/doi.org\/10.1007\/11526346_26","DOI":"10.1007\/11526346_26"},{"key":"6875_CR11","doi-asserted-by":"publisher","unstructured":"Sivic J, Zisserman A (2003) ideo Google: A text retrieval approach to object matching in videos. In: Proceedings ninth IEEE international conference on computer vision, p 1470. IEEE . https:\/\/doi.org\/10.1109\/ICCV.2003.1238663","DOI":"10.1109\/ICCV.2003.1238663"},{"key":"6875_CR12","doi-asserted-by":"publisher","unstructured":"Perronnin F, S\u00e1nchez J, Mensink T (2010) mproving the fisher kernel for large-scale image classification. In: European conference on computer vision, pp 143\u2013156. Springer . https:\/\/doi.org\/10.1007\/978-3-642-15561-1_11","DOI":"10.1007\/978-3-642-15561-1_11"},{"key":"6875_CR13","doi-asserted-by":"publisher","unstructured":"Li Y, Wang R, Cui Z, Shan S, Chen X (2014) ompact video code and its application to robust face retrieval in TV-series. In: Proceedings of the British machine vision conference, pp 1\u201312. BMVA Press . https:\/\/doi.org\/10.5244\/C.28.93","DOI":"10.5244\/C.28.93"},{"key":"6875_CR14","doi-asserted-by":"publisher","unstructured":"Wang R, Guo H, Davis LS, Dai Q (2012) ovariance discriminative learning: A natural and efficient approach to image set classification. In: 2012 IEEE conference on computer vision and pattern recognition, pp 2496\u20132503. IEEE . https:\/\/doi.org\/10.1109\/CVPR.2012.6247965","DOI":"10.1109\/CVPR.2012.6247965"},{"key":"6875_CR15","doi-asserted-by":"crossref","unstructured":"Dong Z, Jia S, Wu T, Pei M (2016)Face video retrieval via deep learning of binary hash representations. In: Thirtieth AAAI conference on artificial intelligence, pp 3471\u20133477","DOI":"10.1609\/aaai.v30i1.10445"},{"key":"6875_CR16","first-page":"1097","volume":"25","author":"A Krizhevsky","year":"2012","unstructured":"Krizhevsky A, Sutskever I, Hinton GE (2012) ImageNet classification with deep convolutional neural networks. Adv Neural Inf Process Syst 25:1097\u20131105","journal-title":"Adv Neural Inf Process Syst"},{"key":"6875_CR17","doi-asserted-by":"publisher","unstructured":"Qiao S, Wang R, Shan S, Chen X (2016) ep video code for efficient face video retrieval. In: Asian conference on computer vision, pp 296\u2013312. Springer . https:\/\/doi.org\/10.1007\/978-3-319-54187-7_20","DOI":"10.1007\/978-3-319-54187-7_20"},{"key":"6875_CR18","doi-asserted-by":"publisher","DOI":"10.1016\/j.patcog.2020.107754","author":"S Qiao","year":"2020","unstructured":"Qiao S, Wang R, Shan S, Chen X (2020) eep video code for efficient face video retrieval. Pattern Recognit. https:\/\/doi.org\/10.1016\/j.patcog.2020.107754","journal-title":"Pattern Recognit"},{"key":"6875_CR19","doi-asserted-by":"publisher","first-page":"1299","DOI":"10.1109\/TIP.2019.2940683","volume":"29","author":"S Qiao","year":"2019","unstructured":"Qiao S, Wang R, Shan S, Chen X (2019) Deep heterogeneous hashing for face video retrieval. IEEE Trans Image Process 29:1299\u20131312. https:\/\/doi.org\/10.1109\/TIP.2019.2940683","journal-title":"IEEE Trans Image Process"},{"key":"6875_CR20","doi-asserted-by":"publisher","unstructured":"Wang R, Qiao S, Shan S, Chen X (2020) Hybrid video and image hashing for robust face retrieval. In: 2020 15th IEEE international conference on automatic face and gesture recognition, pp 186\u2013193 . https:\/\/doi.org\/10.1109\/FG47880.2020.00028","DOI":"10.1109\/FG47880.2020.00028"},{"issue":"21","key":"6875_CR21","doi-asserted-by":"publisher","first-page":"22169","DOI":"10.1007\/s11042-017-4962-9","volume":"76","author":"M M\u00fchling","year":"2017","unstructured":"M\u00fchling M, Korfhage N, M\u00fcller E, Otto C, Springstein M, Langelage T, Veith U, Ewerth R, Freisleben B (2017) Deep learning for content-based video retrieval in film and television production. Multimed Tools Appl 76(21):22169\u201322194. https:\/\/doi.org\/10.1007\/s11042-017-4962-9","journal-title":"Multimed Tools Appl"},{"key":"6875_CR22","first-page":"91","volume":"28","author":"S Ren","year":"2015","unstructured":"Ren S, He K, Girshick R, Sun J (2015) Faster R-CNN: Towards real-time object detection with region proposal networks. Adv Neural Inf Process Syst 28:91\u201399","journal-title":"Adv Neural Inf Process Syst"},{"key":"6875_CR23","unstructured":"Yi D, Lei Z, Liao S, Li SZ (2014) Earning face representation from scratch. arXiv preprint arXiv:1411.7923"},{"key":"6875_CR24","doi-asserted-by":"publisher","unstructured":"Fang X, Zou Y (2019) Ake the best of face clues in iQIYI celebrity video identification challenge 2019. In: Proceedings of the 27th ACM international conference on multimedia, pp. 2526\u20132530 . https:\/\/doi.org\/10.1145\/3343031.3356056","DOI":"10.1145\/3343031.3356056"},{"key":"6875_CR25","unstructured":"2019 iQIYI celebrity video identification challenge. http:\/\/challenge.ai.iqiyi.com\/detail?raceId=5c767dc41a6fa0ccf53922e6. Accessed: 20 Oct 2020"},{"key":"6875_CR26","doi-asserted-by":"publisher","DOI":"10.1016\/j.dsp.2020.102809","author":"M Taskiran","year":"2020","unstructured":"Taskiran M, Kahraman N, Erdem CE (2020) ace recognition: Past, present and future (a review). Digit Signal Process. https:\/\/doi.org\/10.1016\/j.dsp.2020.102809","journal-title":"Digit Signal Process"},{"key":"6875_CR27","doi-asserted-by":"publisher","first-page":"102805","DOI":"10.1016\/j.cviu.2019.102805","volume":"189","author":"G Guo","year":"2019","unstructured":"Guo G, Zhang N (2019) A survey on deep learning based face recognition. Comput Vis Image Underst 189:102805. https:\/\/doi.org\/10.1016\/j.cviu.2019.102805","journal-title":"Comput Vis Image Underst"},{"key":"6875_CR28","doi-asserted-by":"crossref","unstructured":"Taigman Y, Yang M, Ranzato M, Wolf, L (2014) Deepface: closing the gap to human-level performance in face verification. In: Proceedings of the IEEE conference on computer vision and pattern recognition, pp 1701\u20131708","DOI":"10.1109\/CVPR.2014.220"},{"key":"6875_CR29","unstructured":"Huang GB, Mattar M, Berg T, Learned-Miller E (2008) Labeled faces in the wild: a database for studying face recognition in unconstrained environments. Workshop on faces in \u201creal-life\u201d images: detection. alignment, and recognition. Erik Learned-Miller and Andras Ferencz and Fr\u00e9d\u00e9ric Jurie, Marseille, France, pp 7\u201349"},{"key":"6875_CR30","doi-asserted-by":"publisher","unstructured":"Kumar N, Berg AC, Belhumeur PN, Nayar SK (2009) Attribute and simile classifiers for face verification. In: 2009 IEEE 12th international conference on computer vision, pp 365\u2013372. IEEE . https:\/\/doi.org\/10.1109\/ICCV.2009.5459250","DOI":"10.1109\/ICCV.2009.5459250"},{"key":"6875_CR31","doi-asserted-by":"publisher","unstructured":"Sun Y, Wang X, Tang X (2014) Deep learning face representation from predicting 10,000 classes. In: Proceedings of the IEEE conference on computer vision and pattern recognition, pp 1891\u20131898 . https:\/\/doi.org\/10.1109\/CVPR.2014.244","DOI":"10.1109\/CVPR.2014.244"},{"key":"6875_CR32","doi-asserted-by":"publisher","unstructured":"Chen D, Cao X, Wang L, Wen F, Sun J (2012) Bayesian face revisited: a joint formulation. In: European conference on computer vision, pp 566\u2013579. Springer . https:\/\/doi.org\/10.1007\/978-3-642-33712-3_41","DOI":"10.1007\/978-3-642-33712-3_41"},{"key":"6875_CR33","unstructured":"Sun Y, Chen Y, Wang X, Tang X (2014) Deep learning face representation by joint identification-verification. In: Advances in neural information processing systems, pp 1988\u20131996"},{"key":"6875_CR34","unstructured":"Sun Y, Liang D, Wang X, Tang X (2015) eepID3: face recognition with very deep neural networks. arXiv preprint arXiv:1502.00873"},{"key":"6875_CR35","doi-asserted-by":"crossref","unstructured":"Szegedy C, Liu W, Jia Y, Sermanet P, Reed S, Anguelov D, Erhan D, Vanhoucke V, Rabinovich A (2015) Going deeper with convolutions. In: Proceedings of the IEEE conference on computer vision and pattern recognition, pp 1\u20139","DOI":"10.1109\/CVPR.2015.7298594"},{"key":"6875_CR36","doi-asserted-by":"crossref","unstructured":"Schroff F, Kalenichenko D, Philbin J (2015) FaceNet: a unified embedding for face recognition and clustering. In: Proceedings of the IEEE conference on computer vision and pattern recognition, pp 815\u2013823 (2015)","DOI":"10.1109\/CVPR.2015.7298682"},{"key":"6875_CR37","doi-asserted-by":"publisher","unstructured":"Wolf L, Hassner T, Maoz I (2011) Face recognition in unconstrained videos with matched background similarity. In: CVPR 2011, pp 529\u2013534. IEEE. https:\/\/doi.org\/10.1109\/CVPR.2011.5995566","DOI":"10.1109\/CVPR.2011.5995566"},{"key":"6875_CR38","doi-asserted-by":"crossref","unstructured":"Parkhi OM, Vedaldi A, Zisserman A (2015) Deep face recognition. In: Proceedings of the British machine vision conference, p 41.1-41.12","DOI":"10.5244\/C.29.41"},{"key":"6875_CR39","doi-asserted-by":"publisher","unstructured":"Cao Q, Shen L, Xie W, Parkhi OM, Zisserman A (2018) VggFace2: a dataset for recognising faces across pose and age. In: 2018 13th IEEE international conference on automatic face & gesture recognition (FG 2018), pp 67\u201374. IEEE. https:\/\/doi.org\/10.1109\/FG.2018.00020","DOI":"10.1109\/FG.2018.00020"},{"key":"6875_CR40","doi-asserted-by":"publisher","unstructured":"Klare BF, Klein B, Taborsky E, Blanton A, Cheney J, Allen K, Grother P, Mah A, Jain AK (2015) Pushing the frontiers of unconstrained face detection and recognition: Iarpa Janus Benchmark A. In: Proceedings of the IEEE conference on computer vision and pattern recognition, pp 1931\u20131939. https:\/\/doi.org\/10.1109\/CVPR.2015.7298803","DOI":"10.1109\/CVPR.2015.7298803"},{"key":"6875_CR41","doi-asserted-by":"crossref","unstructured":"Wen Y, Zhang K, Li Z, Qiao Y (2016) A discriminative feature learning approach for deep face recognition. In: European conference on computer vision, pp 499\u2013515. Springer","DOI":"10.1007\/978-3-319-46478-7_31"},{"key":"6875_CR42","doi-asserted-by":"crossref","unstructured":"Qi C, Su F (2017) Contrastive-center loss for deep neural networks. In: 2017 IEEE international conference on image processing (ICIP), pp 2851\u20132855. IEEE","DOI":"10.1109\/ICIP.2017.8296803"},{"key":"6875_CR43","unstructured":"Liu W, Wen Y, Yu Z, Yang M (2016) Large-margin Softmax loss for convolutional neural networks. In: Proceedings of The 33rd international conference on machine learning, proceedings of machine learning research"},{"key":"6875_CR44","unstructured":"Liu Y, Li H, Wang X (2017) Rethinking feature discrimination and polymerization for large-scale recognition. arXiv preprint arXiv:1710.00870"},{"key":"6875_CR45","doi-asserted-by":"crossref","unstructured":"Liu W, Wen Y, Yu Z, Li M, Raj B, Song, L (2017) SphereFace: Deep hypersphere embedding for face recognition. In: Proceedings of the IEEE conference on computer vision and pattern recognition, pp 212\u2013220","DOI":"10.1109\/CVPR.2017.713"},{"key":"6875_CR46","doi-asserted-by":"crossref","unstructured":"Wang H, Wang Y, Zhou Z, Ji X, Gong D, Zhou J, Li Z, Liu W (2018) CosFace: large margin cosine loss for deep face recognition. In: Proceedings of the IEEE conference on computer vision and pattern recognition, pp 5265\u20135274 (2018)","DOI":"10.1109\/CVPR.2018.00552"},{"key":"6875_CR47","doi-asserted-by":"publisher","unstructured":"Deng J, Guo J, Xue N, Zafeiriou S (2019) ArcFace: additive angular margin loss for deep face recognition. In: Proceedings of the IEEE conference on computer vision and pattern recognition, pp 4690\u20134699. https:\/\/doi.org\/10.1109\/CVPR.2019.00482","DOI":"10.1109\/CVPR.2019.00482"},{"key":"6875_CR48","doi-asserted-by":"publisher","unstructured":"Kemelmacher-Shlizerman I, Seitz SM, Miller D, Brossard E (2016) The MegaFace benchmark: 1 million faces for recognition at scale. In: Proceedings of the IEEE conference on computer vision and pattern recognition, pp 4873\u20134882. https:\/\/doi.org\/10.1109\/CVPR.2016.527","DOI":"10.1109\/CVPR.2016.527"},{"key":"6875_CR49","doi-asserted-by":"publisher","unstructured":"Whitelam C, Taborsky E, Blanton A, Maze B, Adams J, Miller T, Kalka N, Jain AK, Duncan JA, Allen K, et\u00a0al. (2017) Iarpa Janus Benchmark-B face dataset. In: Proceedings of the IEEE conference on computer vision and pattern recognition workshops, pp 90\u201398. https:\/\/doi.org\/10.1109\/CVPRW.2017.87","DOI":"10.1109\/CVPRW.2017.87"},{"key":"6875_CR50","doi-asserted-by":"publisher","unstructured":"Maze B, Adams J, Duncan JA, Kalka N, Miller T, Otto C, Jain AK, Niggel WT, Anderson J, Cheney J et\u00a0al. (2018) Iarpa Janus Benchmark-C: Face dataset and protocol. In: 2018 international conference on biometrics (ICB), pp 158\u2013165. IEEE. https:\/\/doi.org\/10.1109\/ICB2018.2018.00033","DOI":"10.1109\/ICB2018.2018.00033"},{"key":"6875_CR51","unstructured":"Liu Y, Peng B, Shi P, Yan H, Zhou Y, Han B, Zheng Y, Lin C, Jiang J, Fan Y et\u00a0al. (2018) iQIYI-VID: A large dataset for multi-modal person identification. arXiv preprint arXiv:1811.07548"},{"key":"6875_CR52","doi-asserted-by":"publisher","unstructured":"Rao Y, Lin J, Lu J, Zhou J (2017) Learning discriminative aggregation network for video-based face recognition. In: Proceedings of the IEEE international conference on computer vision, pp 3781\u20133790 (2017). https:\/\/doi.org\/10.1109\/ICCV.2017.408","DOI":"10.1109\/ICCV.2017.408"},{"key":"6875_CR53","doi-asserted-by":"publisher","unstructured":"Rao Y, Lu J, Zhou J (2017) Attention-aware deep reinforcement learning for video face recognition. In: Proceedings of the IEEE international conference on computer vision, pp 3931\u20133940. https:\/\/doi.org\/10.1109\/ICCV.2017.424","DOI":"10.1109\/ICCV.2017.424"},{"issue":"4","key":"6875_CR54","doi-asserted-by":"publisher","first-page":"1002","DOI":"10.1109\/TPAMI.2017.2700390","volume":"40","author":"C Ding","year":"2017","unstructured":"Ding C, Tao D (2017) Trunk-branch ensemble convolutional neural networks for video-based face recognition. IEEE Trans Pattern Anal Mach Intell 40(4):1002\u20131014. https:\/\/doi.org\/10.1109\/TPAMI.2017.2700390","journal-title":"IEEE Trans Pattern Anal Mach Intell"},{"issue":"3","key":"6875_CR55","doi-asserted-by":"publisher","first-page":"194","DOI":"10.1109\/TBIOM.2020.2973504","volume":"2","author":"J Zheng","year":"2020","unstructured":"Zheng J, Ranjan R, Chen CH, Chen JC, Castillo CD, Chellappa R (2020) An automatic system for unconstrained video-based face recognition. IEEE Trans Biom Behav Identity Sci 2(3):194\u2013209. https:\/\/doi.org\/10.1109\/TBIOM.2020.2973504","journal-title":"IEEE Trans Biom Behav Identity Sci"},{"key":"6875_CR56","doi-asserted-by":"publisher","unstructured":"Chen JC, Lin WA, Zheng J, Chellappa R (2018) A real-time multi-task single shot face detector. In: 2018 25th IEEE international conference on image processing (ICIP), pp 176\u2013180. IEEE. https:\/\/doi.org\/10.1109\/ICIP.2018.8451649","DOI":"10.1109\/ICIP.2018.8451649"},{"key":"6875_CR57","doi-asserted-by":"publisher","unstructured":"Ranjan R, Sankaranarayanan S, Castillo CD, Chellappa R (2017) An all-in-one convolutional neural network for face analysis. In: 2017 12th IEEE international conference on automatic face & gesture recognition (FG 2017), pp 17\u201324. IEEE. https:\/\/doi.org\/10.1109\/FG.2017.137","DOI":"10.1109\/FG.2017.137"},{"issue":"1","key":"6875_CR58","doi-asserted-by":"publisher","first-page":"66","DOI":"10.1109\/MSP.2017.2764116","volume":"35","author":"R Ranjan","year":"2018","unstructured":"Ranjan R, Sankaranarayanan S, Bansal A, Bodla N, Chen JC, Patel VM, Castillo CD, Chellappa R (2018) Deep learning for understanding faces: machines may be just as good, or better, than humans. IEEE Signal Process Mag 35(1):66\u201383. https:\/\/doi.org\/10.1109\/MSP.2017.2764116","journal-title":"IEEE Signal Process Mag"},{"key":"6875_CR59","doi-asserted-by":"publisher","unstructured":"Kalka ND, Maze B, Duncan JA, O\u2019Connor K, Elliott S, Hebert K, Bryan J, Jain AK (2018) IJB\u2013S: IARPA Janus surveillance video benchmark. In: 2018 IEEE 9th international conference on biometrics theory, applications and systems (BTAS), pp 1\u20139. IEEE (2018). https:\/\/doi.org\/10.1109\/BTAS.2018.8698584","DOI":"10.1109\/BTAS.2018.8698584"},{"issue":"34\u201347","key":"6875_CR60","first-page":"4","volume":"4","author":"P Viola","year":"2001","unstructured":"Viola P, Jones M et al (2001) Robust real-time object detection. Int J Comput Vis 4(34\u201347):4","journal-title":"Int J Comput Vis"},{"key":"6875_CR61","doi-asserted-by":"publisher","unstructured":"Bansal A, Nanduri A, Castillo CD, Ranjan R, Chellappa R (2017) Umdfaces: an annotated face dataset for training deep networks. In: 2017 IEEE international joint conference on biometrics (IJCB), pp 464\u2013473. IEEE (2017). https:\/\/doi.org\/10.1109\/BTAS.2017.8272731","DOI":"10.1109\/BTAS.2017.8272731"},{"key":"6875_CR62","doi-asserted-by":"publisher","unstructured":"Bansal A, Castillo C, Ranjan R, Chellappa R (2017) The do\u2019s and don\u2019ts for CNN-based face verification. In: Proceedings of the IEEE international conference on computer vision workshops, pp 2545\u20132554. https:\/\/doi.org\/10.1109\/ICCVW.2017.299","DOI":"10.1109\/ICCVW.2017.299"},{"key":"6875_CR63","unstructured":"UMDFaces. http:\/\/umdfaces.io\/. Accessed 19 Nov 2020"},{"key":"6875_CR64","doi-asserted-by":"publisher","unstructured":"Liu Y, Shi P, Peng B, Yan H, Zhou Y, Han B, Zheng Y, Lin C, Jiang J, Fan Y et\u00a0al (2019) iQIYI celebrity video identification challenge. In: Proceedings of the 27th ACM international conference on multimedia, pp 2516\u20132520. https:\/\/doi.org\/10.1145\/3343031.3356081","DOI":"10.1145\/3343031.3356081"},{"key":"6875_CR65","doi-asserted-by":"publisher","unstructured":"Nagrani A, Chung JS, Zisserman A (2017) VoxCeleb: a large-scale speaker identification dataset. In: Proceedings of the 18th annual conference of the international speech communication association, pp 2616\u20132620 (2017). https:\/\/doi.org\/10.21437\/Interspeech.2017-950","DOI":"10.21437\/Interspeech.2017-950"},{"key":"6875_CR66","doi-asserted-by":"publisher","unstructured":"Sivic J, Everingham M, Zisserman A (2009) \u2018Who are you?\u201d\u2014learning person specific classifiers from video. In: 2009 IEEE conference on computer vision and pattern recognition, pp 1145\u20131152. IEEE. https:\/\/doi.org\/10.1109\/CVPR.2009.5206513","DOI":"10.1109\/CVPR.2009.5206513"},{"key":"6875_CR67","doi-asserted-by":"publisher","unstructured":"Bauml M, Tapaswi M, Stiefelhagen R (2013) Semi-supervised learning with constraints for person identification in multimedia data. In: Proceedings of the IEEE conference on computer vision and pattern recognition, pp 3602\u20133609. https:\/\/doi.org\/10.1109\/CVPR.2013.462","DOI":"10.1109\/CVPR.2013.462"},{"key":"6875_CR68","doi-asserted-by":"crossref","unstructured":"Nagrani A, Zisserman A (2018) From benedict cumberbatch to sherlock holmes: character identification in TV series without a script. arXiv preprint arXiv:1801.10442","DOI":"10.5244\/C.31.107"},{"key":"6875_CR69","doi-asserted-by":"publisher","unstructured":"Huang Q, Liu W, Lin D (2018) Person search in videos with one portrait through visual and temporal links. In: Proceedings of the European conference on computer vision (ECCV), pp 425\u2013441 (2018). https:\/\/doi.org\/10.1007\/978-3-030-01261-8_26","DOI":"10.1007\/978-3-030-01261-8_26"},{"key":"6875_CR70","doi-asserted-by":"publisher","unstructured":"Teng S, Tan W, Zhang W (2007) Cooperative shot boundary detection for video. In: International conference on computer supported cooperative work in design, pp 99\u2013110. Springer. https:\/\/doi.org\/10.1007\/978-3-540-92719-8_10","DOI":"10.1007\/978-3-540-92719-8_10"},{"key":"6875_CR71","doi-asserted-by":"publisher","first-page":"61","DOI":"10.1016\/j.neucom.2019.11.023","volume":"381","author":"G Ciaparrone","year":"2020","unstructured":"Ciaparrone G, S\u00e1nchez FL, Tabik S, Troiano L, Tagliaferri R, Herrera F (2020) Deep learning in video multi-object tracking: a survey. Neurocomputing 381:61\u201388. https:\/\/doi.org\/10.1016\/j.neucom.2019.11.023","journal-title":"Neurocomputing"},{"key":"6875_CR72","doi-asserted-by":"publisher","unstructured":"Guo Y, Zhang L, Hu Y, He X, Gao J (2016) MS-Celeb-1M: a dataset and benchmark for large-scale face recognition. In: European conference on computer vision, pp 87\u2013102. Springer. https:\/\/doi.org\/10.1007\/978-3-319-46487-9_6","DOI":"10.1007\/978-3-319-46487-9_6"},{"key":"6875_CR73","doi-asserted-by":"publisher","unstructured":"Baraldi L, Grana C, Cucchiara R (2015) Shot and scene detection via hierarchical clustering for re-using broadcast video. In: International conference on computer analysis of images and patterns, pp 801\u2013811. Springer. https:\/\/doi.org\/10.1007\/978-3-319-23192-1_67","DOI":"10.1007\/978-3-319-23192-1_67"},{"key":"6875_CR74","unstructured":"Sou\u010dek T, Loko\u010d J (2020) TransNet V2: an effective deep network architecture for fast shot transition detection. arXiv preprint arXiv:2008.04838"},{"key":"6875_CR75","unstructured":"Sou\u010dek T, Moravec J, Loko\u010d J (2019) TransNet: a deep network for fast detection of common shot transitions. arXiv preprint arXiv:1906.03363"},{"issue":"10","key":"6875_CR76","doi-asserted-by":"publisher","first-page":"1499","DOI":"10.1109\/LSP.2016.2603342","volume":"23","author":"K Zhang","year":"2016","unstructured":"Zhang K, Zhang Z, Li Z, Qiao Y (2016) Joint face detection and alignment using multitask cascaded convolutional networks. IEEE Signal Process Lett 23(10):1499\u20131503. https:\/\/doi.org\/10.1109\/LSP.2016.2603342","journal-title":"IEEE Signal Process Lett"},{"key":"6875_CR77","doi-asserted-by":"crossref","unstructured":"Najibi M, Samangouei P, Chellappa R, Davis LS (2017) SSH: single stage headless face detector. In: Proceedings of the IEEE international conference on computer vision, pp 4875\u20134884","DOI":"10.1109\/ICCV.2017.522"},{"key":"6875_CR78","doi-asserted-by":"crossref","unstructured":"Deng J, Guo J, Ververas E, Kotsia I, Zafeiriou S (2020) RetinaFace: single-shot multi-level face localisation in the wild. In: Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition, pp 5203\u20135212","DOI":"10.1109\/CVPR42600.2020.00525"},{"key":"6875_CR79","unstructured":"sklearn.metrics.average_precision_score \u2013 scikit-learn 0.23.2 documentation. https:\/\/scikit-learn.org\/stable\/modules\/generated\/sklearn.metrics.average_precision_score.html. Accessed 09 Dec 2020"},{"key":"6875_CR80","doi-asserted-by":"publisher","unstructured":"Bochinski E, Eiselein V, Sikora T (2017) High-speed tracking-by-detection without using image information. In: 2017 14th IEEE international conference on advanced video and signal based surveillance (AVSS), pp 1\u20136. IEEE. https:\/\/doi.org\/10.1109\/AVSS.2017.8078516","DOI":"10.1109\/AVSS.2017.8078516"},{"key":"6875_CR81","unstructured":"deepinsight\/insightface: face analysis project on MXNet. https:\/\/github.com\/deepinsight\/insightface. Accessed: 15 Dec 2020"}],"container-title":["Neural Computing and Applications"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s00521-021-06875-x.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s00521-021-06875-x\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s00521-021-06875-x.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2023,1,22]],"date-time":"2023-01-22T00:22:26Z","timestamp":1674346946000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s00521-021-06875-x"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,1,5]]},"references-count":81,"journal-issue":{"issue":"10","published-print":{"date-parts":[[2022,5]]}},"alternative-id":["6875"],"URL":"https:\/\/doi.org\/10.1007\/s00521-021-06875-x","relation":{},"ISSN":["0941-0643","1433-3058"],"issn-type":[{"value":"0941-0643","type":"print"},{"value":"1433-3058","type":"electronic"}],"subject":[],"published":{"date-parts":[[2022,1,5]]},"assertion":[{"value":"5 July 2021","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"15 December 2021","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"5 January 2022","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declaration"}},{"value":"The authors declare that they have no conflicts of interest.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Conflict of interest"}}]}}