{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T04:30:39Z","timestamp":1750221039155,"version":"3.41.0"},"reference-count":50,"publisher":"Association for Computing Machinery (ACM)","issue":"4","license":[{"start":{"date-parts":[[2018,10,10]],"date-time":"2018-10-10T00:00:00Z","timestamp":1539129600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"DOI":"10.13039\/501100012226","name":"Fundamental Research Funds for the Central Universities of China","doi-asserted-by":"crossref","id":[{"id":"10.13039\/501100012226","id-type":"DOI","asserted-by":"crossref"}]},{"DOI":"10.13039\/501100003453","name":"Natural Science Foundation of Guangdong Province","doi-asserted-by":"crossref","award":["2017A030311029 and 2016B010109002"],"award-info":[{"award-number":["2017A030311029 and 2016B010109002"]}],"id":[{"id":"10.13039\/501100003453","id-type":"DOI","asserted-by":"crossref"}]},{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"crossref","award":["61673402, 61273270 and 60802069"],"award-info":[{"award-number":["61673402, 61273270 and 60802069"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]},{"name":"Science and Technology Program of Guangzhou, China","award":["201704020180"],"award-info":[{"award-number":["201704020180"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Multimedia Comput. Commun. Appl."],"published-print":{"date-parts":[[2018,11,30]]},"abstract":"<jats:p>In this article, a two-stage refinement network is proposed for facial landmarks detection on unconstrained conditions. Our model can be divided into two modules, namely the Head Attribude Classifier (HAC) module and the Domain-Specific Refinement (DSR) module. Given an input facial image, HAC adopts multi-task learning mechanism to detect the head pose and obtain an initial shape. Based on the obtained head pose, DSR designs three different CNN-based refinement networks trained by specific domain, respectively, and automatically selects the most approximate network for the landmarks refinement. Different from existing two-stage models, HAC combines head pose prediction with facial landmarks estimation to improve the accuracy of head pose prediction, as well as obtaining a robust initial shape. Moreover, an adaptive sub-network training strategy applied in the DSR module can effectively solve the issue of traditional multi-view methods that an improperly selected sub-network may result in alignment failure. The extensive experimental results on two public datasets, AFLW and 300W, confirm the validity of our model.<\/jats:p>","DOI":"10.1145\/3241059","type":"journal-article","created":{"date-parts":[[2018,10,10]],"date-time":"2018-10-10T13:30:46Z","timestamp":1539178246000},"page":"1-19","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":4,"title":["Joint Head Attribute Classifier and Domain-Specific Refinement Networks for Face Alignment"],"prefix":"10.1145","volume":"14","author":[{"given":"Junfeng","family":"Zhang","sequence":"first","affiliation":[{"name":"School of Electronics and Information Technology, Sun Yat-Sen University, Guangzhou, Guangdong, Peoples Republic of China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-4884-323X","authenticated-orcid":false,"given":"Haifeng","family":"Hu","sequence":"additional","affiliation":[{"name":"School of Electronics and Information Technology, Sun Yat-Sen University, Guangzhou, Guangdong, Peoples Republic of China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2018,10,10]]},"reference":[{"key":"e_1_2_1_1_1","volume-title":"Martinez","author":"Benitez-Quiroz C. Fabian","year":"2016","unstructured":"C. Fabian Benitez-Quiroz , Ramprakash Srinivasan , and Aleix M . Martinez . 2016 . EmotioNet: An accurate, real-time algorithm for the automatic annotation of a million facial expressions in the wild. In Computer Vision and Pattern Recognition . 5562--557. C. Fabian Benitez-Quiroz, Ramprakash Srinivasan, and Aleix M. Martinez. 2016. EmotioNet: An accurate, real-time algorithm for the automatic annotation of a million facial expressions in the wild. In Computer Vision and Pattern Recognition. 5562--557."},{"key":"e_1_2_1_2_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2013.191"},{"key":"e_1_2_1_3_1","doi-asserted-by":"publisher","DOI":"10.1006\/cviu.1995.1004"},{"key":"e_1_2_1_4_1","doi-asserted-by":"publisher","DOI":"10.1109\/34.927467"},{"volume-title":"Proceedings of the IEEE International Conference on Automatic Face and Gesture Recognition. 227","author":"Cootes T. F.","key":"e_1_2_1_5_1","unstructured":"T. F. Cootes , K. Walker , and C. J. Taylor . 2002. View-based active appearance models . In Proceedings of the IEEE International Conference on Automatic Face and Gesture Recognition. 227 . T. F. Cootes, K. Walker, and C. J. Taylor. 2002. View-based active appearance models. In Proceedings of the IEEE International Conference on Automatic Face and Gesture Recognition. 227."},{"volume-title":"Proceedings of the British Machine Vision Conference. 929--938","author":"Cristinacce David","key":"e_1_2_1_6_1","unstructured":"David Cristinacce and Timothy F. Cootes . 2006. Feature detection and tracking with constrained local models . In Proceedings of the British Machine Vision Conference. 929--938 . David Cristinacce and Timothy F. Cootes. 2006. Feature detection and tracking with constrained local models. In Proceedings of the British Machine Vision Conference. 929--938."},{"key":"e_1_2_1_7_1","volume-title":"Joint multi-view face alignment in the wild. arXiv preprint arXiv:1708.06023","author":"Deng Jiankang","year":"2017","unstructured":"Jiankang Deng , George Trigeorgis , Yuxiang Zhou , and Stefanos Zafeiriou . 2017. Joint multi-view face alignment in the wild. arXiv preprint arXiv:1708.06023 ( 2017 ). Jiankang Deng, George Trigeorgis, Yuxiang Zhou, and Stefanos Zafeiriou. 2017. Joint multi-view face alignment in the wild. arXiv preprint arXiv:1708.06023 (2017)."},{"key":"e_1_2_1_8_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00045"},{"volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 21--26","author":"Dou Pengfei","key":"e_1_2_1_9_1","unstructured":"Pengfei Dou , Shishir K. Shah , and Ioannis A. Kakadiaris . 2017. End-to-end 3D face reconstruction with deep neural networks . In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 21--26 . Pengfei Dou, Shishir K. Shah, and Ioannis A. Kakadiaris. 2017. End-to-end 3D face reconstruction with deep neural networks. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 21--26."},{"key":"e_1_2_1_10_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.392"},{"key":"e_1_2_1_11_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-10605-2_36"},{"key":"e_1_2_1_12_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2015.421"},{"key":"e_1_2_1_13_1","doi-asserted-by":"crossref","unstructured":"Amin Jourabloo and Xiaoming Liu. 2016. Large-pose face alignment via CNN-based dense 3D model fitting. In Computer Vision and Pattern Recognition.  Amin Jourabloo and Xiaoming Liu. 2016. Large-pose face alignment via CNN-based dense 3D model fitting. In Computer Vision and Pattern Recognition.","DOI":"10.1109\/CVPR.2016.454"},{"key":"e_1_2_1_14_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2014.241"},{"key":"e_1_2_1_15_1","volume-title":"Guosheng Hu, and William Christmas.","author":"Kittler Josef","year":"2016","unstructured":"Josef Kittler , Patrik Huber , Zhen Hua Feng , Guosheng Hu, and William Christmas. 2016 . 3D Morphable Face Models and Their Applications. Springer International Publishing . Josef Kittler, Patrik Huber, Zhen Hua Feng, Guosheng Hu, and William Christmas. 2016. 3D Morphable Face Models and Their Applications. Springer International Publishing."},{"volume-title":"Proceedings of the International Conference on Neural Information Processing Systems. 1097--1105","author":"Krizhevsky Alex","key":"e_1_2_1_16_1","unstructured":"Alex Krizhevsky , Ilya Sutskever , and Geoffrey E. Hinton . 2012. ImageNet classification with deep convolutional neural networks . In Proceedings of the International Conference on Neural Information Processing Systems. 1097--1105 . Alex Krizhevsky, Ilya Sutskever, and Geoffrey E. Hinton. 2012. ImageNet classification with deep convolutional neural networks. In Proceedings of the International Conference on Neural Information Processing Systems. 1097--1105."},{"volume-title":"Proceedings of the International Conference on Neural Information Processing Systems. 950--957","author":"Krogh Anders","key":"e_1_2_1_17_1","unstructured":"Anders Krogh and John A. Hertz . 1991. A simple weight decay can improve generalization . In Proceedings of the International Conference on Neural Information Processing Systems. 950--957 . Anders Krogh and John A. Hertz. 1991. A simple weight decay can improve generalization. In Proceedings of the International Conference on Neural Information Processing Systems. 950--957."},{"key":"e_1_2_1_18_1","volume-title":"Proceedings of the IEEE International Conference on Computer Vision Workshops. 2144--2151","author":"Martin","year":"2012","unstructured":"Martin K?stinger, Paul Wohlhart , Peter M. Roth , and Horst Bischof . 2012 . Annotated facial landmarks in the wild: A large-scale, real-world database for facial landmark localization . In Proceedings of the IEEE International Conference on Computer Vision Workshops. 2144--2151 . Martin K?stinger, Paul Wohlhart, Peter M. Roth, and Horst Bischof. 2012. Annotated facial landmarks in the wild: A large-scale, real-world database for facial landmark localization. In Proceedings of the IEEE International Conference on Computer Vision Workshops. 2144--2151."},{"key":"e_1_2_1_19_1","doi-asserted-by":"publisher","DOI":"10.1109\/5.726791"},{"key":"e_1_2_1_20_1","volume-title":"Unconstrained facial landmark localization with backbone-branches fully-convolutional networks. arXiv preprint arXiv:1507.03409","author":"Liang Zhujin","year":"2015","unstructured":"Zhujin Liang , Shengyong Ding , and Liang Lin . 2015. Unconstrained facial landmark localization with backbone-branches fully-convolutional networks. arXiv preprint arXiv:1507.03409 ( 2015 ). Zhujin Liang, Shengyong Ding, and Liang Lin. 2015. Unconstrained facial landmark localization with backbone-branches fully-convolutional networks. arXiv preprint arXiv:1507.03409 (2015)."},{"key":"e_1_2_1_21_1","volume-title":"Improving person re-identification by attribute and identity learning. arXiv preprint arXiv:1703.07220","author":"Lin Yutian","year":"2017","unstructured":"Yutian Lin , Liang Zheng , Zhedong Zheng , Yu Wu , and Yi Yang . 2017. Improving person re-identification by attribute and identity learning. arXiv preprint arXiv:1703.07220 ( 2017 ). Yutian Lin, Liang Zheng, Zhedong Zheng, Yu Wu, and Yi Yang. 2017. Improving person re-identification by attribute and identity learning. arXiv preprint arXiv:1703.07220 (2017)."},{"key":"e_1_2_1_22_1","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2002.999679"},{"key":"e_1_2_1_23_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCVW.2017.190"},{"key":"e_1_2_1_24_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.393"},{"key":"e_1_2_1_25_1","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2016.2518867"},{"key":"e_1_2_1_26_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCVW.2013.59"},{"key":"e_1_2_1_27_1","volume-title":"Very deep convolutional networks for large-scale image recognition. arXiv preprint arXiv:1409.1556","author":"Simonyan Karen","year":"2014","unstructured":"Karen Simonyan and Andrew Zisserman . 2014. Very deep convolutional networks for large-scale image recognition. arXiv preprint arXiv:1409.1556 ( 2014 ). Karen Simonyan and Andrew Zisserman. 2014. Very deep convolutional networks for large-scale image recognition. arXiv preprint arXiv:1409.1556 (2014)."},{"key":"e_1_2_1_28_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2013.446"},{"key":"e_1_2_1_29_1","unstructured":"Yi Sun Yuheng Chen Xiaogang Wang and Xiaoou Tang. 2014. Deep learning face representation by joint identification-verification. In Advances in Neural Information Processing Systems. 1988--1996.   Yi Sun Yuheng Chen Xiaogang Wang and Xiaoou Tang. 2014. Deep learning face representation by joint identification-verification. In Advances in Neural Information Processing Systems. 1988--1996."},{"key":"e_1_2_1_30_1","doi-asserted-by":"crossref","unstructured":"Georgios Tzimiropoulos. 2015. Project-out cascaded regression with an application to face alignment. In Computer Vision and Pattern Recognition. 3659--3667.  Georgios Tzimiropoulos. 2015. Project-out cascaded regression with an application to face alignment. In Computer Vision and Pattern Recognition. 3659--3667.","DOI":"10.1109\/CVPR.2015.7298989"},{"key":"e_1_2_1_31_1","doi-asserted-by":"crossref","unstructured":"Robert Walecki Ognjen Rudovic Vladimir Pavlovic and Maja Pantic. 2016. Copula ordinal regression for joint estimation of facial action unit intensity. In Computer Vision and Pattern Recognition. 4902--4910.  Robert Walecki Ognjen Rudovic Vladimir Pavlovic and Maja Pantic. 2016. Copula ordinal regression for joint estimation of facial action unit intensity. In Computer Vision and Pattern Recognition. 4902--4910.","DOI":"10.1109\/CVPR.2016.530"},{"key":"e_1_2_1_32_1","doi-asserted-by":"publisher","DOI":"10.1007\/s11263-013-0667-3"},{"key":"e_1_2_1_33_1","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2016.2515987"},{"key":"e_1_2_1_34_1","doi-asserted-by":"crossref","unstructured":"Yue Wu and Qiang Ji. 2016. Constrained joint cascade regression framework for simultaneous facial action unit recognition and facial landmark detection. In Computer Vision and Pattern Recognition. 3400--3408.  Yue Wu and Qiang Ji. 2016. Constrained joint cascade regression framework for simultaneous facial action unit recognition and facial landmark detection. In Computer Vision and Pattern Recognition. 3400--3408.","DOI":"10.1109\/CVPR.2016.370"},{"key":"e_1_2_1_35_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-46448-0_4"},{"key":"e_1_2_1_36_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2013.75"},{"key":"e_1_2_1_37_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2015.7298882"},{"volume-title":"Proceedings of the IEEE International Conference on Automatic Face and Gesture Recognition. 642--649","author":"Xu Xiang","key":"e_1_2_1_38_1","unstructured":"Xiang Xu and Ioannis A. Kakadiaris . 2017. Joint head pose estimation and face alignment framework using global and local CNN features . In Proceedings of the IEEE International Conference on Automatic Face and Gesture Recognition. 642--649 . Xiang Xu and Ioannis A. Kakadiaris. 2017. Joint head pose estimation and face alignment framework using global and local CNN features. In Proceedings of the IEEE International Conference on Automatic Face and Gesture Recognition. 642--649."},{"key":"e_1_2_1_39_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPRW.2017.253"},{"key":"e_1_2_1_40_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.554"},{"key":"e_1_2_1_41_1","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2017.2765830"},{"key":"e_1_2_1_42_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-46454-1_4"},{"key":"e_1_2_1_43_1","doi-asserted-by":"crossref","unstructured":"Junfeng Zhang and Haifeng Hu. 2018. Exemplar-based cascaded stacked auto-encoder networks for robust face alignment. Computer Vision and Image Understanding.  Junfeng Zhang and Haifeng Hu. 2018. Exemplar-based cascaded stacked auto-encoder networks for robust face alignment. Computer Vision and Image Understanding.","DOI":"10.1016\/j.cviu.2018.05.002"},{"key":"e_1_2_1_44_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-10605-2_1"},{"key":"e_1_2_1_45_1","doi-asserted-by":"publisher","DOI":"10.1109\/LSP.2016.2603342"},{"key":"e_1_2_1_46_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-10599-4_7"},{"key":"e_1_2_1_47_1","doi-asserted-by":"publisher","DOI":"10.1145\/3159171"},{"key":"e_1_2_1_48_1","volume-title":"Change Loy Chen, and Xiaoou Tang","author":"Zhu Shizhan","year":"2015","unstructured":"Shizhan Zhu , Cheng Li , Change Loy Chen, and Xiaoou Tang . 2015 . Face alignment by coarse-to-fine shape searching. In Computer Vision and Pattern Recognition . 4998--5006. Shizhan Zhu, Cheng Li, Change Loy Chen, and Xiaoou Tang. 2015. Face alignment by coarse-to-fine shape searching. In Computer Vision and Pattern Recognition. 4998--5006."},{"key":"e_1_2_1_49_1","volume-title":"Change Loy Chen, and Xiaoou Tang","author":"Zhu Shizhan","year":"2016","unstructured":"Shizhan Zhu , Cheng Li , Change Loy Chen, and Xiaoou Tang . 2016 . Unconstrained face alignment via cascaded compositional learning. In Computer Vision and Pattern Recognition . 3409--3417. Shizhan Zhu, Cheng Li, Change Loy Chen, and Xiaoou Tang. 2016. Unconstrained face alignment via cascaded compositional learning. In Computer Vision and Pattern Recognition. 3409--3417."},{"volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 146--155","author":"Zhu Xiangyu","key":"e_1_2_1_50_1","unstructured":"Xiangyu Zhu , Zhen Lei , Xiaoming Liu , Hailin Shi , and Stan Z. Li . 2016. Face alignment across large poses: A 3D solution . In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 146--155 . Xiangyu Zhu, Zhen Lei, Xiaoming Liu, Hailin Shi, and Stan Z. Li. 2016. Face alignment across large poses: A 3D solution. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 146--155."}],"container-title":["ACM Transactions on Multimedia Computing, Communications, and Applications"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3241059","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3241059","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T00:43:46Z","timestamp":1750207426000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3241059"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2018,10,10]]},"references-count":50,"journal-issue":{"issue":"4","published-print":{"date-parts":[[2018,11,30]]}},"alternative-id":["10.1145\/3241059"],"URL":"https:\/\/doi.org\/10.1145\/3241059","relation":{},"ISSN":["1551-6857","1551-6865"],"issn-type":[{"type":"print","value":"1551-6857"},{"type":"electronic","value":"1551-6865"}],"subject":[],"published":{"date-parts":[[2018,10,10]]},"assertion":[{"value":"2018-03-01","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2018-07-01","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2018-10-10","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}