{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,2,8]],"date-time":"2026-02-08T11:12:46Z","timestamp":1770549166057,"version":"3.49.0"},"reference-count":34,"publisher":"Springer Science and Business Media LLC","issue":"4","license":[{"start":{"date-parts":[[2021,9,12]],"date-time":"2021-09-12T00:00:00Z","timestamp":1631404800000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2021,9,12]],"date-time":"2021-09-12T00:00:00Z","timestamp":1631404800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Pattern Anal Applic"],"published-print":{"date-parts":[[2021,11]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:p>Nowadays, face detection and head pose estimation have a lot of application such as face recognition, aiding in gaze estimation and modeling attention. For these two tasks, it is usually to design two different models. However, the head pose estimation model often depends on the region of interest (ROI) detected in advance, which means that a serial face detector is needed. Even the lightest face detector will slow down the whole forward inference time and cannot achieve real-time performance when detecting the head pose of multiple people. We can see that both face detection and head pose estimation need face features, so a shared face feature map can be used between them. In this paper, a multi-task learning model is proposed that can solve both problems simultaneously. We directly detect the location of the center point of the bounding box of face; at this location, we calculate the size of the bounding box of face and the head attitude. We evaluate our model\u2019s performance on the AFLW. The proposed model has great competitiveness with the multi-stage face attribute analysis model, and our model can achieve real-time performance.<\/jats:p>","DOI":"10.1007\/s10044-021-01026-3","type":"journal-article","created":{"date-parts":[[2021,9,12]],"date-time":"2021-09-12T10:04:01Z","timestamp":1631441041000},"page":"1745-1755","update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":13,"title":["TRFH: towards real-time face detection and head pose estimation"],"prefix":"10.1007","volume":"24","author":[{"given":"Shicun","family":"Chen","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-6650-6790","authenticated-orcid":false,"given":"Yong","family":"Zhang","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Baocai","family":"Yin","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Boyue","family":"Wang","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"297","published-online":{"date-parts":[[2021,9,12]]},"reference":[{"issue":"1\u20132","key":"1026_CR1","doi-asserted-by":"publisher","first-page":"123","DOI":"10.1007\/BF01450852","volume":"15","author":"DF DeMenthon","year":"1995","unstructured":"DeMenthon DF, Davis LS (1995) Model-based object pose in 25 lines of code. Int J Comput Vision 15(1\u20132):123\u2013141","journal-title":"Int J Comput Vision"},{"key":"1026_CR2","unstructured":"Deng J, Guo J, Xue N, Zafeiriou S, Arcface: additive angular margin loss for deep face recognition. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 4690\u20134699"},{"key":"1026_CR3","doi-asserted-by":"crossref","unstructured":"Deng J, Guo J, Zhou Y, et al (2019) Retinaface: single-stage dense face localisation in the wild. arXiv preprint arXiv:1905.00641","DOI":"10.1109\/CVPR42600.2020.00525"},{"key":"1026_CR4","doi-asserted-by":"crossref","unstructured":"Fanelli G, Weise T, Gall J, Van\u00a0Gool L, Real time head pose estimation from consumer depth cameras. In: Joint pattern recognition symposium, pp. 101\u2013110. Springer","DOI":"10.1007\/978-3-642-23123-0_11"},{"key":"1026_CR5","unstructured":"He K, Gkioxari G, Doll\u00e1r P, Girshick R, Mask r-cnn. In: Proceedings of the IEEE international conference on computer vision, pp. 2961\u20132969"},{"key":"1026_CR6","unstructured":"He K, Zhang X, Ren S, Sun J, Deep residual learning for image recognition. In: Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 770\u2013778"},{"key":"1026_CR7","unstructured":"Jain V, Learned-Miller E (2010) Fddb: a benchmark for face detection in unconstrained settings. Tech. rep, UMass Amherst technical report"},{"key":"1026_CR8","unstructured":"Krizhevsky A, Sutskever I, Hinton GE, Imagenet classification with deep convolutional neural networks. Adva Neural Inform Process Syst 1097\u20131105"},{"key":"1026_CR9","doi-asserted-by":"crossref","unstructured":"Li H, Hua G, Lin Z, Brandt J, Yang J (2013) Probabilistic elastic part model for unsupervised face detector adaptation. In: Proceedings of the IEEE international conference on computer vision, pp. 793\u2013800","DOI":"10.1109\/ICCV.2013.103"},{"key":"1026_CR10","doi-asserted-by":"crossref","unstructured":"Liu W, Anguelov D, Erhan D, Szegedy C, Reed S, Fu CY, Berg AC, Ssd: Single shot multibox detector. In: European conference on computer vision, pp. 21\u201337. Springer","DOI":"10.1007\/978-3-319-46448-0_2"},{"key":"1026_CR11","doi-asserted-by":"crossref","unstructured":"Koestinger M, Wohlhart P, Roth PM, Bischof H (2011) Annotated Facial Landmarks in the Wild: a large-scale, real-world database for facial landmark localization. In: Proc. First IEEE International Workshop on Benchmarking Facial Image Analysis Technologies","DOI":"10.1109\/ICCVW.2011.6130513"},{"key":"1026_CR12","doi-asserted-by":"crossref","unstructured":"Mathias M, Benenson R, Pedersoli M, Van\u00a0Gool L, Face detection without bells and whistles. In: European conference on computer vision, pp. 720\u2013735. Springer","DOI":"10.1007\/978-3-319-10593-2_47"},{"key":"1026_CR13","unstructured":"Najibi M, Samangouei P, Chellappa R, Davis LS, Ssh: Single stage headless face detector. In: Proceedings of the IEEE international conference on computer vision, pp. 4875\u20134884"},{"issue":"5\u20136","key":"1026_CR14","doi-asserted-by":"publisher","first-page":"359","DOI":"10.1016\/S0262-8856(02)00008-2","volume":"20","author":"J Ng","year":"2002","unstructured":"Ng J, Gong S (2002) Composite support vector machines for detection of faces across views and pose estimation. Image Vision Comput 20(5\u20136):359\u2013368","journal-title":"Image Vision Comput"},{"key":"1026_CR15","unstructured":"Pan H, Han H, Shan S, Chen X, Mean-variance loss for deep age estimation from a face. In: Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 5285\u20135294"},{"issue":"1","key":"1026_CR16","doi-asserted-by":"publisher","first-page":"121","DOI":"10.1109\/TPAMI.2017.2781233","volume":"41","author":"R Ranjan","year":"2017","unstructured":"Ranjan R, Patel VM, Chellappa R (2017) Hyperface: a deep multi-task learning framework for face detection, landmark localization, pose estimation, and gender recognition. IEEE Trans Pattern Anal Mach Intell 41(1):121\u2013135","journal-title":"IEEE Trans Pattern Anal Mach Intell"},{"key":"1026_CR17","unstructured":"Ren S, He K, Girshick R, Sun J. Faster r-cnn: towards real-time object detection with region proposal networks. Adv Neural Inform Process Syst 91\u201399"},{"key":"1026_CR18","doi-asserted-by":"crossref","unstructured":"Rowley HA, Baluja S, Kanade T (1998) Rotation invariant neural network-based face detection. In: Proceedings IEEE computer society conference on computer vision and pattern recognition (Cat. No. 98CB36231), pp. 38\u201344. IEEE","DOI":"10.21236\/ADA341629"},{"issue":"1","key":"1026_CR19","doi-asserted-by":"publisher","first-page":"23","DOI":"10.1109\/34.655647","volume":"20","author":"HA Rowley","year":"1998","unstructured":"Rowley HA, Baluja S, Kanade T (1998) Neural network-based face detection. IEEE Trans Pattern Anal Mach Intell 20(1):23\u201338","journal-title":"IEEE Trans Pattern Anal Mach Intell"},{"key":"1026_CR20","unstructured":"Ruiz N, Chong E, Rehg JM, Fine-grained head pose estimation without keypoints. In: Proceedings of the IEEE conference on computer vision and pattern recognition workshops, pp. 2074\u20132083"},{"key":"1026_CR21","unstructured":"Schroff F, Kalenichenko D, Philbin J, Facenet: a unified embedding for face recognition and clustering. In: Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 815\u2013823"},{"issue":"12","key":"1026_CR22","doi-asserted-by":"publisher","first-page":"807","DOI":"10.1016\/S0262-8856(00)00096-2","volume":"19","author":"J Sherrah","year":"2001","unstructured":"Sherrah J, Gong S, Ong EJ (2001) Face distributions in similarity space under varying head pose. Image Vision Comput 19(12):807\u2013819","journal-title":"Image Vision Comput"},{"key":"1026_CR23","unstructured":"Simonyan K, Zisserman A (2014) Very deep convolutional networks for large-scale image recognition. arXiv preprint arXiv:1409.1556"},{"key":"1026_CR24","unstructured":"Sun K, Zhao Y, Jiang B, et al (2019) High-resolution representations for labeling pixels and regions. arXiv preprint arXiv:1904.04514"},{"issue":"2","key":"1026_CR25","doi-asserted-by":"publisher","first-page":"137","DOI":"10.1023\/B:VISI.0000013087.49260.fb","volume":"57","author":"P Viola","year":"2004","unstructured":"Viola P, Jones MJ (2004) Robust real-time face detection. Int J Comput Vision 57(2):137\u2013154","journal-title":"Int J Comput Vision"},{"key":"1026_CR26","unstructured":"Wang H, Li Z, Ji X, et al (2017) Face r-cnn[J]. arXiv preprint arXiv:1706.01061"},{"key":"1026_CR27","unstructured":"Yan J, Lei Z, Wen L, Li SZ, The fastest deformable part model for object detection. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 2497\u20132504"},{"issue":"10","key":"1026_CR28","doi-asserted-by":"publisher","first-page":"790","DOI":"10.1016\/j.imavis.2013.12.004","volume":"32","author":"J Yan","year":"2014","unstructured":"Yan J, Zhang X, Lei Z, Li SZ (2014) Face detection by structural models. Image Vision Comput 32(10):790\u2013799","journal-title":"Image Vision Comput"},{"key":"1026_CR29","unstructured":"Yang B, Yan J, Lei Z, Li SZ, Aggregate channel features for multi-view face detection. In: IEEE international joint conference on biometrics, pp. 1\u20138. IEEE"},{"key":"1026_CR30","unstructured":"Yang TY, Chen YT, Lin YY, Chuang YY, Fsa-net: Learning fine-grained structure aggregation for head pose estimation from a single image. In: Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 1087\u20131096"},{"key":"1026_CR31","unstructured":"Yashunin D, Baydasov T, Vlasov R (2020) MaskFace: multi-task face and landmark detector. arXiv preprint arXiv:2005.09412"},{"issue":"10","key":"1026_CR32","doi-asserted-by":"publisher","first-page":"1499","DOI":"10.1109\/LSP.2016.2603342","volume":"23","author":"K Zhang","year":"2016","unstructured":"Zhang K, Zhang Z, Li Z, et al (2016) Joint face detection and alignment using multitask cascaded convolutional networks. IEEE Signal Process Lett 23(10):1499\u20131503","journal-title":"IEEE Signal Process Lett"},{"key":"1026_CR33","unstructured":"Zhou X, Wang D, Kr\u00e4henb\u00fchl P (2019) Objects as points. arXiv preprint arXiv:1904.07850"},{"key":"1026_CR34","unstructured":"Zhu X, Ramanan D (2012) Face detection, pose estimation, and landmark localization in the wild. In: 2012 IEEE conference on computer vision and pattern recognition, pp. 2879\u20132886. IEEE"}],"container-title":["Pattern Analysis and Applications"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10044-021-01026-3.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s10044-021-01026-3\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10044-021-01026-3.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2021,10,23]],"date-time":"2021-10-23T08:21:55Z","timestamp":1634977315000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s10044-021-01026-3"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2021,9,12]]},"references-count":34,"journal-issue":{"issue":"4","published-print":{"date-parts":[[2021,11]]}},"alternative-id":["1026"],"URL":"https:\/\/doi.org\/10.1007\/s10044-021-01026-3","relation":{},"ISSN":["1433-7541","1433-755X"],"issn-type":[{"value":"1433-7541","type":"print"},{"value":"1433-755X","type":"electronic"}],"subject":[],"published":{"date-parts":[[2021,9,12]]},"assertion":[{"value":"31 December 2020","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"11 August 2021","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"12 September 2021","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}}]}}