{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,5]],"date-time":"2026-05-05T09:08:59Z","timestamp":1777972139925,"version":"3.51.4"},"reference-count":74,"publisher":"Springer Science and Business Media LLC","issue":"11","license":[{"start":{"date-parts":[[2025,8,23]],"date-time":"2025-08-23T00:00:00Z","timestamp":1755907200000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2025,8,23]],"date-time":"2025-08-23T00:00:00Z","timestamp":1755907200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"DOI":"10.13039\/501100001775","name":"University of Technology Sydney","doi-asserted-by":"crossref","id":[{"id":"10.13039\/501100001775","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Int J Comput Vis"],"published-print":{"date-parts":[[2025,11]]},"abstract":"<jats:title>Abstract<\/jats:title>\n                  <jats:p>\n                    Despite the similar structures of human faces, existing face alignment methods cannot learn unified knowledge from multiple datasets with different landmark annotations. The limited training samples in a single dataset commonly result in fragile robustness in this field. To mitigate knowledge discrepancies among different datasets and train a task-agnostic unified face alignment (TUFA) framework, this paper presents a strategy to unify knowledge from multiple datasets. Specifically, we calculate a mean face shape for each dataset. To explicitly align these mean shapes on an interpretable plane based on their semantics, each shape is then incorporated with a group of semantic alignment embeddings. The 2D coordinates of these aligned shapes can be viewed as the anchors of the plane. By encoding them into structure prompts and further regressing the corresponding facial landmarks using image features, a mapping from the plane to the target faces is finally established, which unifies the learning target of different datasets. Consequently, multiple datasets can be utilized to boost the generalization ability of the model. The successful mitigation of discrepancies also enhances the efficiency of knowledge transferring to a novel dataset, significantly boosts the performance of few-shot face alignment. Additionally, the interpretable plane endows TUFA with a task-agnostic characteristic, enabling it to locate landmarks unseen during training in a zero-shot manner. Extensive experiments are carried on seven benchmarks and the results demonstrate an impressive improvement in face alignment brought by knowledge discrepancies mitigation. The code is available at\n                    <jats:ext-link xmlns:xlink=\"http:\/\/www.w3.org\/1999\/xlink\" xlink:href=\"https:\/\/github.com\/Jiahao-UTS\/TUFA\" ext-link-type=\"uri\">https:\/\/github.com\/Jiahao-UTS\/TUFA<\/jats:ext-link>\n                  <\/jats:p>","DOI":"10.1007\/s11263-025-02520-5","type":"journal-article","created":{"date-parts":[[2025,8,23]],"date-time":"2025-08-23T17:59:56Z","timestamp":1755971996000},"page":"7985-8005","update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":1,"title":["Mitigating Knowledge Discrepancies among Multiple Datasets for Task-agnostic Unified Face Alignment"],"prefix":"10.1007","volume":"133","author":[{"ORCID":"https:\/\/orcid.org\/0000-0001-9628-9563","authenticated-orcid":false,"given":"Jiahao","family":"Xia","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-9581-8849","authenticated-orcid":false,"given":"Min","family":"Xu","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-2408-8302","authenticated-orcid":false,"given":"Wenjian","family":"Huang","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-9317-0268","authenticated-orcid":false,"given":"Jianguo","family":"Zhang","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-0021-3634","authenticated-orcid":false,"given":"Haimin","family":"Zhang","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-4526-6297","authenticated-orcid":false,"given":"Chunxia","family":"Xiao","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"297","published-online":{"date-parts":[[2025,8,23]]},"reference":[{"key":"2520_CR1","doi-asserted-by":"crossref","unstructured":"Asthana, A., Marks, T.K., Jones, M.J., Tieu, K.H., & Rohith, M. (2011). Fully automatic pose-invariant face recognition via 3d pose normalization. Proc. ieee int. conf. comput. vis. (p.937-944).","DOI":"10.1109\/ICCV.2011.6126336"},{"key":"2520_CR2","unstructured":"Ba, J.L., Kiros, J.R., Hinton, G.E. (2016). Layer normalization. arXiv:1607.06450"},{"key":"2520_CR3","doi-asserted-by":"crossref","unstructured":"Belhumeur, P.N., Jacobs, D.W., Kriegman, D.J., & Kumar, N. (2011). Localizing parts of faces using a consensus of exemplars. Proc. ieee conf. comput. vis. pattern recognit. (p.545-552).","DOI":"10.1109\/CVPR.2011.5995602"},{"key":"2520_CR4","doi-asserted-by":"crossref","unstructured":"Browatzki, B., & Wallraven, C. (2020). 3fabrec: Fast few-shot face alignment by reconstruction. Proc. ieee conf. comput. vis. pattern recognit. (p.6109-6119).","DOI":"10.1109\/CVPR42600.2020.00615"},{"key":"2520_CR5","doi-asserted-by":"crossref","unstructured":"Bulat, A., & Tzimiropoulos, G. (2017). How far are we from solving the 2d & 3d face alignment problem? (and a dataset of 230,000 3d facial landmarks). Proc. ieee int. conf. comput. vis. (p.1021-1030).","DOI":"10.1109\/ICCV.2017.116"},{"key":"2520_CR6","doi-asserted-by":"crossref","unstructured":"Burgos-Artizzu, X.P., Perona, P., & Doll\u00e1r, P. (2013). Robust face landmark estimation under occlusion. Proc. ieee int. conf. comput. vis. (p.1513-1520).","DOI":"10.1109\/ICCV.2013.191"},{"key":"2520_CR7","doi-asserted-by":"crossref","unstructured":"Caron, M., Touvron, H., Misra, I., Jegou, H., Mairal, J., Bojanowski, P., & Joulin, A. (2021). Emerging properties in self-supervised vision transformers. Proc. ieee int. conf. comput. vis. (p.9630-9640).","DOI":"10.1109\/ICCV48922.2021.00951"},{"key":"2520_CR8","doi-asserted-by":"crossref","unstructured":"Chen, L., Su, H., & Ji, Q. (2019). Face alignment with kernel density deep neural network. Proc. ieee int. conf. comput. vis. (p.6991-7001).","DOI":"10.1109\/ICCV.2019.00709"},{"issue":"6","key":"2520_CR9","doi-asserted-by":"publisher","first-page":"681","DOI":"10.1109\/34.927467","volume":"23","author":"T Cootes","year":"2001","unstructured":"Cootes, T., Edwards, G., & Taylor, C. (2001). Active appearance models. IEEE Trans. Pattern Anal. Mach. Intell., 23(6), 681\u2013685. https:\/\/doi.org\/10.1109\/34.927467","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"issue":"1","key":"2520_CR10","doi-asserted-by":"publisher","first-page":"38","DOI":"10.1006\/cviu.1995.1004","volume":"61","author":"T Cootes","year":"1995","unstructured":"Cootes, T., Taylor, C., Cooper, D., & Graham, J. (1995). Active shape models-their training and application. Comput. Vis. Image Understand., 61(1), 38\u201359. https:\/\/doi.org\/10.1006\/cviu.1995.1004","journal-title":"Comput. Vis. Image Understand."},{"key":"2520_CR11","doi-asserted-by":"publisher","unstructured":"Cristinacce, D., & Cootes, T.F. (2006). Feature detection and tracking with constrained local models. Proc. british mach. vis. conf. (p.95.1-95.10). (https:\/\/doi.org\/10.5244\/C.20.95)","DOI":"10.5244\/C.20.95"},{"key":"2520_CR12","doi-asserted-by":"crossref","unstructured":"Deng, J., Dong, W., Socher, R., Li, L.-J., Li, K., & Fei-Fei, L. (2009). Imagenet: A large-scale hierarchical image database. Proc. ieee conf. comput. vis. pattern recognit. (p.248-255).","DOI":"10.1109\/CVPR.2009.5206848"},{"key":"2520_CR13","unstructured":"Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., & Houlsby, N. (2021). An image is worth 16x16 words: Transformers for image recognition at scale. Proc. int. conf. learn. representations."},{"key":"2520_CR14","doi-asserted-by":"crossref","unstructured":"Feng, Z., Kittler, J., Christmas, W., Huber, P., Wu, X. (2017). Dynamic attention-controlled cascaded shape regression exploiting training data augmentation and fuzzy-set sample weighting. Proc. ieee conf. comput. vis. pattern recognit. (p.3681-3690).","DOI":"10.1109\/CVPR.2017.392"},{"key":"2520_CR15","doi-asserted-by":"crossref","unstructured":"Feng, Z.-H., Kittler, J., Christmas, W., Huber, P., Wu, X.-J. (2017). Dynamic attention-controlled cascaded shape regression exploiting training data augmentation and fuzzy-set sample weighting. Proc. ieee conf. comput. vis. pattern recognit. (p.3681-3690).","DOI":"10.1109\/CVPR.2017.392"},{"key":"2520_CR16","doi-asserted-by":"crossref","unstructured":"Ghiasi, G., & Fowlkes, C.C. (2014). Occlusion coherence: Localizing occluded faces with a hierarchical deformable part model. Proc. ieee conf. comput. vis. pattern recognit. (p.1899-1906).","DOI":"10.1109\/CVPR.2014.306"},{"key":"2520_CR17","doi-asserted-by":"crossref","unstructured":"Guler, R.A., Neverova, N., & Kokkinos, I. (2018). DensePose: Dense Human Pose Estimation in the Wild . Proc. ieee conf. comput. vis. pattern recognit. (p.7297-7306).","DOI":"10.1109\/CVPR.2018.00762"},{"key":"2520_CR18","doi-asserted-by":"crossref","unstructured":"He, X., Bharaj, G., Ferman, D., Rhodin, H., & Garrido, P. (2023). Few-shot geometry-aware keypoint localization. Proc. ieee conf. comput. vis. pattern recognit. (p.21337-21348).","DOI":"10.1109\/CVPR52729.2023.02044"},{"key":"2520_CR19","unstructured":"He, X., Wandt, B., & Rhodin, H. (2022). Autolink: Self-supervised learning of human skeletons and object outlines by linking keypoints. S.\u00a0Koyejo, S.\u00a0Mohamed, A.\u00a0Agarwal, D.\u00a0Belgrave, K.\u00a0Cho, and A.\u00a0Oh (Eds.), Proc. advances neural inf. process. syst. (Vol.\u00a035, pp. 36123\u201336141). Curran Associates, Inc."},{"issue":"4","key":"2520_CR20","doi-asserted-by":"publisher","first-page":"4071","DOI":"10.1109\/TPAMI.2022.3199617","volume":"45","author":"G Huang","year":"2023","unstructured":"Huang, G., Laradji, I., V\u00e1zquez, D., Lacoste-Julien, S., & Rodr\u00edguez, P. (2023). A survey of self-supervised and few-shot object detection. IEEE Trans. Pattern Anal. Mach. Intell., 45(4), 4071\u20134089. https:\/\/doi.org\/10.1109\/TPAMI.2022.3199617","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"2520_CR21","doi-asserted-by":"crossref","unstructured":"Huang, Y., Yang, H., Li, C., Kim, J., & Wei, F. (2021). Adnet: Leveraging error-bias towards normal direction in face alignment. Proc. ieee int. conf. comput. vis. (p.3060-3070).","DOI":"10.1109\/ICCV48922.2021.00307"},{"key":"2520_CR22","unstructured":"Jakab, T., Gupta, A., Bilen, H., & Vedaldi, A. (2018). Unsupervised learning of object landmarks through conditional image generation. S.\u00a0Bengio, H.\u00a0Wallach, H.\u00a0Larochelle, K.\u00a0Grauman, N.\u00a0Cesa-Bianchi, and R.\u00a0Garnett (Eds.), Proc. advances neural inf. process. syst. (Vol.\u00a031). Curran Associates, Inc."},{"issue":"1","key":"2520_CR23","doi-asserted-by":"publisher","first-page":"229","DOI":"10.1109\/TPAMI.2022.3153611","volume":"45","author":"S Jiang","year":"2023","unstructured":"Jiang, S., Zhu, Y., Liu, C., Song, X., Li, X., & Min, W. (2023). Dataset bias in few-shot image recognition. IEEE Trans. Pattern Anal. Mach. Intell., 45(1), 229\u2013246. https:\/\/doi.org\/10.1109\/TPAMI.2022.3153611","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"2520_CR24","doi-asserted-by":"publisher","DOI":"10.1007\/s11263-021-01521-4","author":"H Jin","year":"2021","unstructured":"Jin, H., Liao, S., & Shao, L. (2021). Pixel-in-pixel net: Towards efficient facial landmark detection in the wild. Int. J. Comput. Vis. https:\/\/doi.org\/10.1007\/s11263-021-01521-4","journal-title":"Int. J. Comput. Vis."},{"key":"2520_CR25","doi-asserted-by":"crossref","unstructured":"Kowalski, M., Naruniec, J., & Trzcinski, T. (2017). Deep alignment network: A convolutional neural network for robust face alignment. Proc. ieee conf. comput. vis. pattern recognit. workshops (p.2034-2043).","DOI":"10.1109\/CVPRW.2017.254"},{"key":"2520_CR26","doi-asserted-by":"crossref","unstructured":"Kumar, A., Marks, T.K., Mou, W., Wang, Y., Jones, M., Cherian, A., & Feng, C. (2020). Luvli face alignment: Estimating landmarks\u2019 location, uncertainty, and visibility likelihood. Proc. ieee conf. comput. vis. pattern recognit. (p.8233-8243).","DOI":"10.1109\/CVPR42600.2020.00826"},{"key":"2520_CR27","doi-asserted-by":"crossref","unstructured":"K\u00f6stinger, M., Wohlhart, P., Roth, P.M., & Bischof, H. (2011). Annotated facial landmarks in the wild: A large-scale, real-world database for facial landmark localization. Proc. ieee int. conf. comput. vis. workshops (p.2144-2151).","DOI":"10.1109\/ICCVW.2011.6130513"},{"key":"2520_CR28","doi-asserted-by":"crossref","unstructured":"Lan, X., Hu, Q., & Cheng, J. (2021). Revisting quantization error in face alignment. Proc. ieee int. conf. comput. vis. workshops (p.1521-1530).","DOI":"10.1109\/ICCVW54120.2021.00177"},{"key":"2520_CR29","doi-asserted-by":"publisher","unstructured":"Lan, X., Hu, Q., & Cheng, J. (2022). Atf: An alternating training framework for weakly supervised face alignment. IEEE Trans. Multim., Early Access, 1-1, https:\/\/doi.org\/10.1109\/TMM.2022.3164798","DOI":"10.1109\/TMM.2022.3164798"},{"key":"2520_CR30","doi-asserted-by":"crossref","unstructured":"Le, V., Brandt, J., Lin, Z., Bourdev, L., & Huang, T.S. (2012). Interactive facial feature localization. Proc. eur. conf. comput. vis. (p.679-692).","DOI":"10.1007\/978-3-642-33712-3_49"},{"key":"2520_CR31","doi-asserted-by":"crossref","unstructured":"Li, W., Lu, Y., Zheng, K., Liao, H., Lin, C., Luo, J., & Miao, S. (2020). Structured landmark detection via topology-adapting deep graph learning. Proc. eur. conf. comput. vis. (pp. 266\u2013283). Cham: Springer International Publishing.","DOI":"10.1007\/978-3-030-58545-7_16"},{"key":"2520_CR32","doi-asserted-by":"publisher","first-page":"5313","DOI":"10.1109\/TIP.2021.3082319","volume":"30","author":"C Lin","year":"2021","unstructured":"Lin, C., Zhu, B., Wang, Q., Liao, R., Qian, C., Lu, J., & Zhou, J. (2021). Structure-coherent deep feature learning for robust face alignment. IEEE Trans. Image Process., 30, 5313\u20135326. https:\/\/doi.org\/10.1109\/TIP.2021.3082319","journal-title":"IEEE Trans. Image Process."},{"issue":"3","key":"2520_CR33","doi-asserted-by":"publisher","first-page":"679","DOI":"10.1109\/TPAMI.2018.2885298","volume":"42","author":"H Liu","year":"2020","unstructured":"Liu, H., Lu, J., Guo, M., Wu, S., & Zhou, J. (2020). Learning reasoning-decision networks for robust face alignment. IEEE Trans. Pattern Anal. Mach. Intell., 42(3), 679\u2013693. https:\/\/doi.org\/10.1109\/TPAMI.2018.2885298","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"issue":"11","key":"2520_CR34","doi-asserted-by":"publisher","first-page":"1941","DOI":"10.1109\/TPAMI.2008.238","volume":"31","author":"X Liu","year":"2009","unstructured":"Liu, X. (2009). Discriminative face alignment. IEEE Trans. Pattern Anal. Mach. Intell., 31(11), 1941\u20131954. https:\/\/doi.org\/10.1109\/TPAMI.2008.238","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"2520_CR35","doi-asserted-by":"crossref","unstructured":"Liu, Z., Luo, P., Wang, X., & Tang, X. (2015). Deep learning face attributes in the wild. Proc. ieee int. conf. comput. vis. (p.3730-3738).","DOI":"10.1109\/ICCV.2015.425"},{"key":"2520_CR36","doi-asserted-by":"crossref","unstructured":"Lorenz, D., Bereska, L., Milbich, T., & Ommer, B. (2019). Unsupervised part-based disentangling of object shape and appearance. Proc. ieee conf. comput. vis. pattern recognit. (p.10947-10956).","DOI":"10.1109\/CVPR.2019.01121"},{"key":"2520_CR37","unstructured":"Loshchilov, I., & Hutter, F. (2019). Decoupled weight decay regularization. Proc. int. conf. learn. representations."},{"key":"2520_CR38","doi-asserted-by":"crossref","unstructured":"Lv, J., Shao, X., Xing, J., Cheng, C., & Zhou, X. (2017). A deep regression architecture with two-stage re-initialization for high performance facial landmark detection. Proc. ieee conf. comput. vis. pattern recognit. (p.3691-3700).","DOI":"10.1109\/CVPR.2017.393"},{"issue":"7","key":"2520_CR39","doi-asserted-by":"publisher","first-page":"8390","DOI":"10.1109\/TPAMI.2023.3234212","volume":"45","author":"D Mallis","year":"2023","unstructured":"Mallis, D., Sanchez, E., Bell, M., & Tzimiropoulos, G. (2023). From keypoints to object landmarks via self-training correspondence: A novel approach to unsupervised landmark discovery. IEEE Trans. Pattern Anal. Mach. Intell., 45(7), 8390\u20138404. https:\/\/doi.org\/10.1109\/TPAMI.2023.3234212","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"2520_CR40","doi-asserted-by":"crossref","unstructured":"Newell, A., Yang, K., & Deng, J. (2016). Stacked hourglass networks for human pose estimation. Proc. eur. conf. comput. vis. (pp. 483\u2013499).","DOI":"10.1007\/978-3-319-46484-8_29"},{"issue":"2","key":"2520_CR41","doi-asserted-by":"publisher","first-page":"848","DOI":"10.1109\/TPAMI.2020.3002500","volume":"44","author":"N Otberdout","year":"2022","unstructured":"Otberdout, N., Daoudi, M., Kacem, A., Ballihi, L., & Berretti, S. (2022). Dynamic facial expression generation on hilbert hypersphere with conditional wasserstein generative adversarial nets. IEEE Trans. Pattern Anal. Mach. Intell., 44(2), 848\u2013863. https:\/\/doi.org\/10.1109\/TPAMI.2020.3002500","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"issue":"4","key":"2520_CR42","doi-asserted-by":"publisher","first-page":"4051","DOI":"10.1109\/TPAMI.2022.3191696","volume":"45","author":"F Pourpanah","year":"2023","unstructured":"Pourpanah, F., Abdar, M., Luo, Y., Zhou, X., Wang, R., Lim, C. P., & Wu, Q. M. J. (2023). A review of generalized zero-shot learning methods. IEEE Trans. Pattern Anal. Mach. Intell., 45(4), 4051\u20134070. https:\/\/doi.org\/10.1109\/TPAMI.2022.3191696","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"2520_CR43","unstructured":"Prados-Torreblanca, A., Buenaposada, J.M., & Baumela, L. (2022). Shape preserving facial landmarks with graph attention networks. Proc. british mach. vis. conf."},{"key":"2520_CR44","doi-asserted-by":"crossref","unstructured":"Qian, S., Sun, K., Wu, W., Qian, C., & Jia, J. (2019). Aggregation via separation: Boosting facial landmark detector with semi-supervised style translation. Proc. ieee int. conf. comput. vis. (p.10152-10162).","DOI":"10.1109\/ICCV.2019.01025"},{"key":"2520_CR45","doi-asserted-by":"crossref","unstructured":"Sagonas, C., Tzimiropoulos, G., Zafeiriou, S., & Pantic, M. (2013a). 300 faces in-the-wild challenge: The first facial landmark localization challenge. Proc. ieee int. conf. comput. vis. workshops (p.397-403).","DOI":"10.1109\/ICCVW.2013.59"},{"key":"2520_CR46","doi-asserted-by":"crossref","unstructured":"Sagonas, C., Tzimiropoulos, G., Zafeiriou, S., & Pantic, M. (2013b). A semi-automatic methodology for facial landmark annotation. Proc. ieee conf. comput. vis. pattern recognit. workshops (p.896-903).","DOI":"10.1109\/CVPRW.2013.132"},{"key":"2520_CR47","doi-asserted-by":"crossref","unstructured":"Tai, Y., Liang, Y., Liu, X., Duan, L., Li, J., Wang, C., & Chen, Y. (2019). Towards highly accurate and stable face alignment for high-resolution videos. Proc. aaai conf. artif. intell. (Vol.\u00a033, pp. 8893\u20138900).","DOI":"10.1609\/aaai.v33i01.33018893"},{"issue":"10","key":"2520_CR48","doi-asserted-by":"publisher","first-page":"2594","DOI":"10.1109\/TPAMI.2019.2932979","volume":"42","author":"AB Tanfous","year":"2020","unstructured":"Tanfous, A. B., Drira, H., & Amor, B. B. (2020). Sparse coding of shape trajectories for facial expression and action recognition. IEEE Trans. Pattern Anal. Mach. Intell., 42(10), 2594\u20132607. https:\/\/doi.org\/10.1109\/TPAMI.2019.2932979","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"issue":"3","key":"2520_CR49","doi-asserted-by":"publisher","first-page":"644","DOI":"10.1007\/s11263-022-01722-5","volume":"131","author":"H Tang","year":"2023","unstructured":"Tang, H., Shao, L., Torr, P. H., & Sebe, N. (2023). Bipartite graph reasoning gans for person pose and facial image synthesis. Int. J. Comput. Vis., 131(3), 644\u2013658. https:\/\/doi.org\/10.1007\/s11263-022-01722-5","journal-title":"Int. J. Comput. Vis."},{"issue":"8","key":"2520_CR50","doi-asserted-by":"publisher","first-page":"2038","DOI":"10.1109\/TPAMI.2019.2907634","volume":"42","author":"Z Tang","year":"2020","unstructured":"Tang, Z., Peng, X., Li, K., & Metaxas, D. N. (2020). Towards efficient u-nets: A coupled and quantized approach. IEEE Trans. Pattern Anal. Mach. Intell., 42(8), 2038\u20132050. https:\/\/doi.org\/10.1109\/TPAMI.2019.2907634","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"2520_CR51","doi-asserted-by":"crossref","unstructured":"Thewlis, J., Bilen, H., & Vedaldi, A. (2017). Unsupervised learning of object landmarks by factorized spatial embeddings. Proc. ieee int. conf. comput. vis. (p.3229-3238).","DOI":"10.1109\/ICCV.2017.348"},{"key":"2520_CR52","doi-asserted-by":"crossref","unstructured":"Trigeorgis, G., Snape, P., Nicolaou, M.A., Antonakos, E., & Zafeiriou, S. (2016). Mnemonic descent method: A recurrent process applied for end-to-end face alignment. Proc. ieee conf. comput. vis. pattern recognit. (p.4177-4187).","DOI":"10.1109\/CVPR.2016.453"},{"key":"2520_CR53","doi-asserted-by":"crossref","unstructured":"Valle, R., Buenaposada, J.M., Vald\u00e9s, A., & Baumela, L. (2018). A deeply-initialized coarse-to-fine ensemble of regression trees for face alignment. Proc. eur. conf. comput. vis. (pp. 609\u2013624). Cham.","DOI":"10.1007\/978-3-030-01264-9_36"},{"issue":"10","key":"2520_CR54","doi-asserted-by":"publisher","first-page":"3349","DOI":"10.1109\/TPAMI.2020.2983686","volume":"43","author":"J Wang","year":"2021","unstructured":"Wang, J., Sun, K., Cheng, T., Jiang, B., Deng, C., Zhao, Y., & Xiao, B. (2021). Deep high-resolution representation learning for visual recognition. IEEE Trans. Pattern Anal. Mach. Intell., 43(10), 3349\u20133364. https:\/\/doi.org\/10.1109\/TPAMI.2020.2983686","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"2520_CR55","doi-asserted-by":"crossref","unstructured":"Wang, X., Bo, L., & Fuxin, L. (2019). Adaptive wing loss for robust face alignment via heatmap regression. Proc. ieee int. conf. comput. vis. (p.6970-6980).","DOI":"10.1109\/ICCV.2019.00707"},{"key":"2520_CR56","doi-asserted-by":"crossref","unstructured":"Wu, W., Qian, C., Yang, S., Wang, Q., Cai, Y., & Zhou, Q. (2018). Look at boundary: A boundary-aware face alignment algorithm. Proc. ieee conf. comput. vis. pattern recognit. (p.2129-2138).","DOI":"10.1109\/CVPR.2018.00227"},{"key":"2520_CR57","doi-asserted-by":"crossref","unstructured":"Wu, W., & Yang, S. (2017). Leveraging intra and inter-dataset variations for robust face alignment. Proc. ieee conf. comput. vis. pattern recognit. workshops (p.2096-2105).","DOI":"10.1109\/CVPRW.2017.261"},{"key":"2520_CR58","doi-asserted-by":"crossref","unstructured":"Xia, J., Qu, W., Huang, W., Zhang, J., Wang, X., & Xu, M. (2022). Sparse local patch transformer for robust face alignment and landmarks inherent relation learning. Proc. ieee conf. comput. vis. pattern recognit. (p.4052-4061).","DOI":"10.1109\/CVPR52688.2022.00402"},{"issue":"8","key":"2520_CR59","doi-asserted-by":"publisher","first-page":"10358","DOI":"10.1109\/TPAMI.2023.3260926","volume":"45","author":"J Xia","year":"2023","unstructured":"Xia, J., Xu, M., Zhang, H., Zhang, J., Huang, W., Cao, H., & Wen, S. (2023). Robust face alignment via inherent relation learning and uncertainty estimation. IEEE Trans. Pattern Anal. Mach. Intell., 45(8), 10358\u201310375. https:\/\/doi.org\/10.1109\/TPAMI.2023.3260926","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"2520_CR60","doi-asserted-by":"crossref","unstructured":"Xiao, B., Wu, H., & Wei, Y. (2018). Simple baselines for human pose estimation and tracking. Proc. eur. conf. comput. vis. (pp. 472\u2013487).","DOI":"10.1007\/978-3-030-01231-1_29"},{"key":"2520_CR61","doi-asserted-by":"crossref","unstructured":"Xiao, S., Feng, J., Xing, J., Lai, H., Yan, S., & Kassim, A. (2016). Robust facial landmark detection via recurrent attentive-refinement networks. Proc. eur. conf. comput. vis. (pp. 57\u201372). Cham.","DOI":"10.1007\/978-3-319-46448-0_4"},{"key":"2520_CR62","doi-asserted-by":"crossref","unstructured":"Xiong, X., & De\u00a0la Torre, F. (2013). Supervised descent method and its applications to face alignment. Proc. ieee conf. comput. vis. pattern recognit. (p.532-539).","DOI":"10.1109\/CVPR.2013.75"},{"key":"2520_CR63","doi-asserted-by":"crossref","unstructured":"Xu, X., Meng, Q., Qin, Y., Guo, J., Zhao, C., Zhou, F., & Lei, Z. (2021). Searching for alignment in face recognition. Proc. aaai conf. artif. intell. (Vol.\u00a035, pp. 3065\u20133073).","DOI":"10.1609\/aaai.v35i4.16415"},{"key":"2520_CR64","doi-asserted-by":"crossref","unstructured":"Yang, J., Liu, Q., & Zhang, K. (2017). Stacked hourglass network for robust facial landmark localisation. Proc. ieee conf. comput. vis. pattern recognit. workshops (p.2025-2033).","DOI":"10.1109\/CVPRW.2017.253"},{"key":"2520_CR65","doi-asserted-by":"crossref","unstructured":"Yang, S., Luo, P., Loy, C.-C., & Tang, X. (2016). Wider face: A face detection benchmark. Proc. ieee conf. comput. vis. pattern recognit.","DOI":"10.1109\/CVPR.2016.596"},{"key":"2520_CR66","doi-asserted-by":"crossref","unstructured":"Zhang, H., Xu, L., Lai, S., Shao, W., Zheng, N., Luo, P., & Zhang, K. (2024). Open-vocabulary animal keypoint detection with semantic-feature matching. Int. J. Comput. Vis.,1\u201318,. DOI: https:\/\/doi.org\/10.1007\/s11263-024-02126-3","DOI":"10.1007\/s11263-024-02126-3"},{"key":"2520_CR67","doi-asserted-by":"publisher","unstructured":"Zhang, J., Hu, H., & Feng, S. (2020). Robust facial landmark detection via heatmap-offset regression. IEEE Trans. Image Process.,29, 5050\u20135064. https:\/\/doi.org\/10.1109\/TIP.2020.2976765","DOI":"10.1109\/TIP.2020.2976765"},{"key":"2520_CR68","doi-asserted-by":"crossref","unstructured":"Zhang, Y., Guo, Y., Jin, Y., Luo, Y., He, Z., & Lee, H. (2018). Unsupervised discovery of object landmarks as structural representations. Proc. ieee conf. comput. vis. pattern recognit. (p.2694-2703).","DOI":"10.1109\/CVPR.2018.00285"},{"key":"2520_CR69","doi-asserted-by":"crossref","unstructured":"Zheng, Y., Yang, H., Zhang, T., Bao, J., Chen, D., Huang, Y., & Wen, F. (2022). General facial representation learning in a visual-linguistic manner. Proc. ieee conf. comput. vis. pattern recognit. (p.18697-18709).","DOI":"10.1109\/CVPR52688.2022.01814"},{"key":"2520_CR70","doi-asserted-by":"crossref","unstructured":"Zhou, Z., Li, H., Liu, H., Wang, N., Yu, G., & Ji, R. (2023). Star loss: Reducing semantic ambiguity in facial landmark detection. Proc. ieee conf. comput. vis. pattern recognit. (p.15475-15484).","DOI":"10.1109\/CVPR52729.2023.01485"},{"key":"2520_CR71","doi-asserted-by":"crossref","unstructured":"Zhu, C., Li, X., Li, J., Dai, S. (2021). Improving robustness of facial landmark detection by defending against adversarial attacks. Proc. ieee int. conf. comput. vis. (p.11731-11740).","DOI":"10.1109\/ICCV48922.2021.01154"},{"key":"2520_CR72","doi-asserted-by":"crossref","unstructured":"Zhu, C., Wan, X., Xie, S., Li, X., Gu, Y. (2022). Occlusion-robust face alignment using a viewpoint-invariant hierarchical network architecture. Proc. ieee conf. comput. vis. pattern recognit. (p.11102-11111).","DOI":"10.1109\/CVPR52688.2022.01083"},{"key":"2520_CR73","doi-asserted-by":"crossref","unstructured":"Zhu, S., Li, C., Loy, C.C., Tang, X. (2015). Face alignment by coarse-to-fine shape searching. Proc. ieee conf. comput. vis. pattern recognit. (p.4998-5006).","DOI":"10.1109\/CVPR.2015.7299134"},{"key":"2520_CR74","doi-asserted-by":"crossref","unstructured":"Zhu, X., & Ramanan, D. (2012). Face detection, pose estimation, and landmark localization in the wild. Proc. ieee conf. comput. vis. pattern recognit. (p.2879-2886).","DOI":"10.1109\/CVPR.2012.6248014"}],"container-title":["International Journal of Computer Vision"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s11263-025-02520-5.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s11263-025-02520-5\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s11263-025-02520-5.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,11,12]],"date-time":"2025-11-12T06:28:17Z","timestamp":1762928897000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s11263-025-02520-5"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,8,23]]},"references-count":74,"journal-issue":{"issue":"11","published-print":{"date-parts":[[2025,11]]}},"alternative-id":["2520"],"URL":"https:\/\/doi.org\/10.1007\/s11263-025-02520-5","relation":{},"ISSN":["0920-5691","1573-1405"],"issn-type":[{"value":"0920-5691","type":"print"},{"value":"1573-1405","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,8,23]]},"assertion":[{"value":"1 October 2024","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"25 June 2025","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"23 August 2025","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"15 September 2025","order":5,"name":"change_date","label":"Change Date","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"Update","order":6,"name":"change_type","label":"Change Type","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"The article has been corrected","order":7,"name":"change_details","label":"Change Details","group":{"name":"ArticleHistory","label":"Article History"}}]}}