{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,6]],"date-time":"2026-05-06T20:01:05Z","timestamp":1778097665395,"version":"3.51.4"},"reference-count":42,"publisher":"Association for Computing Machinery (ACM)","issue":"1","license":[{"start":{"date-parts":[[2020,2,17]],"date-time":"2020-02-17T00:00:00Z","timestamp":1581897600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"DOI":"10.13039\/501100003453","name":"Natural Science Foundation of Guangdong","doi-asserted-by":"crossref","award":["2017A030311029"],"award-info":[{"award-number":["2017A030311029"]}],"id":[{"id":"10.13039\/501100003453","id-type":"DOI","asserted-by":"crossref"}]},{"name":"National Key R8D Program of China","award":["2018YFB1601101 and 2018YFB1601100"],"award-info":[{"award-number":["2018YFB1601101 and 2018YFB1601100"]}]},{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"crossref","award":["61673402, 61273270, and 60802069"],"award-info":[{"award-number":["61673402, 61273270, and 60802069"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]},{"name":"Science and Technology Program of Guangzhou","award":["201704020180"],"award-info":[{"award-number":["201704020180"]}]},{"DOI":"10.13039\/501100012226","name":"Fundamental Research Funds for the Central Universities of China","doi-asserted-by":"crossref","award":["19lgjc03"],"award-info":[{"award-number":["19lgjc03"]}],"id":[{"id":"10.13039\/501100012226","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Multimedia Comput. Commun. Appl."],"published-print":{"date-parts":[[2020,2,29]]},"abstract":"<jats:p>Facial landmark detection aims to locate keypoints for facial images, which typically suffer from variations caused by arbitrary pose, diverse facial expressions, and partial occlusion. In this article, we propose a coarse-to-fine framework that joins a stacked hourglass network and salient region attention refinement for robust face alignment. To achieve this goal, we first present a multi-scale region learning module to analyze the structure information at a different facial region and extract a strong discriminative deep feature. Then we employ a stacked hourglass network for heatmap regression and initial facial landmarks prediction. Specifically, the stacked hourglass network introduces an improved Inception-ResNet unit as a basic building block, which can effectively improve the receptive field and learn contextual feature representations. Meanwhile, a novel loss function takes into account global weights and local weights to make the heatmap regression more accurate. Different from existing heatmap regression models, we present a salient region attention refinement module to extract a precise feature based on the heatmap regression, and utilize the filtered feature for landmarks refinement to achieve accurate prediction. Extensive experimental results of several challenging datasets (including 300 Faces in the Wild, Caltech Occluded Faces in the Wild, and Annotated Facial Landmarks Faces in the Wild) confirm that our approach can achieve more competitive performance than the most advanced algorithms.<\/jats:p>","DOI":"10.1145\/3374760","type":"journal-article","created":{"date-parts":[[2020,3,4]],"date-time":"2020-03-04T10:23:32Z","timestamp":1583317412000},"page":"1-18","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":8,"title":["Joint Stacked Hourglass Network and Salient Region Attention Refinement for Robust Face Alignment"],"prefix":"10.1145","volume":"16","author":[{"given":"Junfeng","family":"Zhang","sequence":"first","affiliation":[{"name":"Sun Yat-Sen University, Guangzhou, Guangdong, People's Republic of China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-4884-323X","authenticated-orcid":false,"given":"Haifeng","family":"Hu","sequence":"additional","affiliation":[{"name":"Sun Yat-Sen University, Guangzhou, Guangdong, People's Republic of China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Guobin","family":"Shen","sequence":"additional","affiliation":[{"name":"Sun Yat-Sen University, Guangzhou, Guangdong, People's Republic of China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2020,2,17]]},"reference":[{"key":"e_1_2_1_1_1","doi-asserted-by":"crossref","unstructured":"Ankan Bansal Carlos Castillo Rajeev Ranjan and Rama Chellappa. 2017. The do\u2019s and don\u2019ts for CNN-based face verification. arXiv:1705.07426.  Ankan Bansal Carlos Castillo Rajeev Ranjan and Rama Chellappa. 2017. The do\u2019s and don\u2019ts for CNN-based face verification. arXiv:1705.07426.","DOI":"10.1109\/ICCVW.2017.299"},{"key":"e_1_2_1_2_1","volume-title":"Proceedings of the 2016 Conference on Computer Vision and Pattern Recognition. 5562--5570","author":"Benitez-Quiroz C. Fabian"},{"key":"e_1_2_1_3_1","doi-asserted-by":"publisher","DOI":"10.5244\/C.30.86"},{"key":"e_1_2_1_4_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-46478-7_44"},{"key":"e_1_2_1_5_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2017.400"},{"key":"e_1_2_1_6_1","volume-title":"Proceedings of the 2014 IEEE International Conference on Computer Vision. 1513--1520","author":"Burgosartizzu Xavier P.","year":"2014"},{"key":"e_1_2_1_7_1","doi-asserted-by":"crossref","unstructured":"Xiao Chu Wei Yang Wanli Ouyang Cheng Ma Alan L. Yuille and Xiaogang Wang. 2017. Multi-context attention for human pose estimation. arXiv:1702.07432.  Xiao Chu Wei Yang Wanli Ouyang Cheng Ma Alan L. Yuille and Xiaogang Wang. 2017. Multi-context attention for human pose estimation. arXiv:1702.07432.","DOI":"10.1109\/CVPR.2017.601"},{"key":"e_1_2_1_8_1","doi-asserted-by":"publisher","DOI":"10.1109\/34.927467"},{"key":"e_1_2_1_9_1","doi-asserted-by":"publisher","DOI":"10.1006\/cviu.1995.1004"},{"key":"e_1_2_1_10_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2010.5540094"},{"key":"e_1_2_1_11_1","volume-title":"Proceedings of the 2017 IEEE Conference on Computer Vision and Pattern Recognition. 21--26","author":"Dou Pengfei"},{"key":"e_1_2_1_12_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.392"},{"key":"e_1_2_1_13_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.90"},{"key":"e_1_2_1_14_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2015.421"},{"key":"e_1_2_1_15_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2014.241"},{"key":"e_1_2_1_16_1","doi-asserted-by":"publisher","DOI":"10.1109\/5.726791"},{"key":"e_1_2_1_17_1","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2016.2627807"},{"key":"e_1_2_1_18_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCVW.2017.190"},{"key":"e_1_2_1_19_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.393"},{"key":"e_1_2_1_20_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCVW.2011.6130513"},{"key":"e_1_2_1_21_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-46484-8_29"},{"key":"e_1_2_1_22_1","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2016.2518867"},{"key":"e_1_2_1_23_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCVW.2013.59"},{"key":"e_1_2_1_24_1","unstructured":"Karen Simonyan and Andrew Zisserman. 2014. Very deep convolutional networks for large-scale image recognition. arXiv:1409.1556.  Karen Simonyan and Andrew Zisserman. 2014. Very deep convolutional networks for large-scale image recognition. arXiv:1409.1556."},{"key":"e_1_2_1_25_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2013.446"},{"key":"e_1_2_1_26_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-01219-9_21"},{"key":"e_1_2_1_27_1","unstructured":"Jonathan J. Tompson Arjun Jain Yann LeCun and Christoph Bregler. 2014. Joint training of a convolutional network and a graphical model for human pose estimation. In Advances in Neural Information Processing Systems. 1799--1807.  Jonathan J. Tompson Arjun Jain Yann LeCun and Christoph Bregler. 2014. Joint training of a convolutional network and a graphical model for human pose estimation. In Advances in Neural Information Processing Systems. 1799--1807."},{"key":"e_1_2_1_28_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.453"},{"key":"e_1_2_1_29_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2015.7298989"},{"key":"e_1_2_1_30_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2013.79"},{"key":"e_1_2_1_31_1","volume-title":"Proceedings of the 2018 European Conference on Computer Vision.585--601","author":"Valle Roberto"},{"key":"e_1_2_1_32_1","doi-asserted-by":"publisher","DOI":"10.1007\/s11263-013-0667-3"},{"key":"e_1_2_1_33_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00227"},{"key":"e_1_2_1_34_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-46448-0_4"},{"key":"e_1_2_1_35_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2013.75"},{"key":"e_1_2_1_36_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPRW.2017.253"},{"key":"e_1_2_1_37_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-46454-1_4"},{"key":"e_1_2_1_38_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-10605-2_1"},{"key":"e_1_2_1_39_1","doi-asserted-by":"publisher","DOI":"10.1109\/LSP.2016.2603342"},{"key":"e_1_2_1_40_1","volume-title":"Proceedings of the 2015 Conference on Computer Vision and Pattern Recognition. 4998--5006","author":"Zhu Shizhan","year":"2015"},{"key":"e_1_2_1_41_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.371"},{"key":"e_1_2_1_42_1","volume-title":"Proceedings of the 2016 IEEE Conference on Computer Vision and Pattern Recognition. 146--155","author":"Zhu Xiangyu"}],"container-title":["ACM Transactions on Multimedia Computing, Communications, and Applications"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3374760","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3374760","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T22:33:08Z","timestamp":1750199588000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3374760"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2020,2,17]]},"references-count":42,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2020,2,29]]}},"alternative-id":["10.1145\/3374760"],"URL":"https:\/\/doi.org\/10.1145\/3374760","relation":{},"ISSN":["1551-6857","1551-6865"],"issn-type":[{"value":"1551-6857","type":"print"},{"value":"1551-6865","type":"electronic"}],"subject":[],"published":{"date-parts":[[2020,2,17]]},"assertion":[{"value":"2019-06-01","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2019-12-01","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2020-02-17","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}