{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,5]],"date-time":"2026-05-05T00:21:21Z","timestamp":1777940481569,"version":"3.51.4"},"reference-count":52,"publisher":"Springer Science and Business Media LLC","issue":"1","license":[{"start":{"date-parts":[[2026,5,2]],"date-time":"2026-05-02T00:00:00Z","timestamp":1777680000000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2026,5,2]],"date-time":"2026-05-02T00:00:00Z","timestamp":1777680000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"DOI":"10.13039\/501100003451","name":"Universidad del Pa\u00eds Vasco","doi-asserted-by":"crossref","id":[{"id":"10.13039\/501100003451","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Cogn Comput"],"published-print":{"date-parts":[[2026,12]]},"abstract":"<jats:title>Abstract<\/jats:title>\n                  <jats:p>\n                    Facial Beauty Prediction (FBP) is a long-standing and inherently challenging task in computer vision, primarily due to the subjective and multifaceted nature of human aesthetic judgment. Perceived facial beauty is influenced by a combination of global facial harmony, local feature attractiveness, and cultural or individual biases\u2014factors that are often difficult to quantify or model explicitly. While deep learning has significantly advanced the field, existing approaches frequently rely on convolutional backbones that emphasize local texture details but fail to capture the subtle, holistic relationships among facial components that define aesthetic perception. To address these limitations, we introduce TransFBP, a novel Vision Transformer (ViT)-based framework specifically designed for interpretable and human-aligned facial beauty assessment. The proposed model incorporates two major innovations. First, we develop a Cross-Attention Head that acts as a dynamic filter, enabling the network to automatically focus on the most visually meaningful facial areas\u2014such as the eyes, lips, and overall symmetry\u2014while suppressing irrelevant background information. This design enables the model to focus adaptively on semantically meaningful areas\u2014such as the eyes, lips, and overall facial symmetry\u2014thereby offering interpretable insights into what drives the aesthetic predictions. Second, to mitigate overfitting and enhance generalization, we propose Attention-Guided TransMix, a two-stage semantic data augmentation strategy. In the first stage, the method generates challenging hybrid samples by conditionally mixing images from opposite ends of the beauty score distribution, encouraging the model to learn discriminative features across a wide aesthetic spectrum. In the second stage, the model\u2019s own attention maps are leveraged to generate a semantically grounded supervisory score for each mixed image, ensuring that the augmented samples remain perceptually meaningful and score-consistent. We comprehensively evaluate TransFBP on the FBP5500 dataset, where our method achieves a state-of-the-art Pearson Correlation Coefficient (PCC) of 0.9291, surpassing existing approaches. The strong empirical results validate the effectiveness of our cross-attention mechanism and attention-guided augmentation strategy. Moreover, the interpretability of our attention maps provides valuable transparency into the model\u2019s decision process, paving the way for more explainable, reliable, and ethically aligned AI systems in aesthetic perception tasks. The code will be available at\n                    <jats:ext-link xmlns:xlink=\"http:\/\/www.w3.org\/1999\/xlink\" xlink:href=\"https:\/\/github.com\/DjameleddineBoukhari\/transFBP\" ext-link-type=\"uri\">https:\/\/github.com\/DjameleddineBoukhari\/transFBP<\/jats:ext-link>\n                    We introduced TransFBP, a novel Transformer-based framework that advances both the performance and interpretability of facial beauty prediction. The proposed approach is characterized by two key innovations: a Cross-Attention Head, which enables the model to dynamically integrate the most salient and contextually relevant facial features, and an Attention-Guided TransMix augmentation strategy, which enhances regularization by generating semantically consistent and challenging training samples.\n                  <\/jats:p>","DOI":"10.1007\/s12559-026-10582-x","type":"journal-article","created":{"date-parts":[[2026,5,2]],"date-time":"2026-05-02T05:51:07Z","timestamp":1777701067000},"update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":0,"title":["Enhancing Facial Beauty Prediction with a Cross-Attention Vision Transformer and Attention-Guided Augmentation"],"prefix":"10.1007","volume":"18","author":[{"given":"D. E.","family":"Boukhari","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"F.","family":"Dornaika","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"297","published-online":{"date-parts":[[2026,5,2]]},"reference":[{"key":"10582_CR1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-32598-9","volume-title":"Computer Models for Facial Beauty Analysis","author":"D Zhang","year":"2016","unstructured":"Zhang D, Chen F, Xu Y. Computer Models for Facial Beauty Analysis. Switzerland: Springer International Publishing; 2016."},{"issue":"4","key":"10582_CR2","doi-asserted-by":"publisher","first-page":"340","DOI":"10.1093\/ejo\/cji042","volume":"27","author":"H Knight","year":"2005","unstructured":"Knight H, Keith O. Ranking facial attractiveness. European J Orthodont. 2005;27(4):340\u20138.","journal-title":"European J Orthodont"},{"key":"10582_CR3","doi-asserted-by":"crossref","unstructured":"Boukhari DE, Chemsa A, Ajgou R, et al. An Ensemble of Deep Convolutional Neural Networks Models for Facial Beauty Prediction. J Adv Computat Intell Intell Inf 2023;27(5).","DOI":"10.20965\/jaciii.2023.p1209"},{"key":"10582_CR4","first-page":"20245","volume":"8","author":"J Gan","year":"2020","unstructured":"Gan J, Xiang L, Zhai Y, et al. 2M BeautyNet: Facial beauty prediction based on multi-task transfer learning. EEE Access. 2020;8:20245\u201356.","journal-title":"EEE Access"},{"issue":"8","key":"10582_CR5","doi-asserted-by":"publisher","first-page":"2710","DOI":"10.1016\/j.patcog.2007.11.022","volume":"41","author":"K Schmid","year":"2008","unstructured":"Schmid K, Marx D, Samal A. Computation of a face attractiveness index based on neoclassical canons, symmetry, and golden ratios. Pattern Recogn. 2008;41(8):2710\u20137.","journal-title":"Pattern Recogn."},{"key":"10582_CR6","doi-asserted-by":"crossref","unstructured":"Liang L, Lin L, Jin L, et al. SCUT-FBP5500: A diverse benchmark dataset for multi-paradigm facial beauty prediction. In: 24th International Conference on Pattern Recognition (ICPR). Beijing, China; 2018. p. 1598-1603.","DOI":"10.1109\/ICPR.2018.8546038"},{"issue":"3","key":"10582_CR7","doi-asserted-by":"publisher","first-page":"757","DOI":"10.1097\/01.prs.0000207382.60636.1c","volume":"118","author":"M Bashour","year":"2006","unstructured":"Bashour M. An objective system for measuring facial attractiveness. Plastic Reconstruct Surg. 2006;118(3):757\u201374.","journal-title":"Plastic Reconstruct Surg."},{"key":"10582_CR8","doi-asserted-by":"crossref","unstructured":"He K, Zhang X, Ren S, et al. Deep residual learning for image recognition. In: IEEE Conference on Computer Vision and Pattern Recognition (CVPR). Las Vegas, NV, USA; 2016. p. 770-778.","DOI":"10.1109\/CVPR.2016.90"},{"key":"10582_CR9","unstructured":"Simonyan K. Very deep convolutional networks for large-scale image recognition. 2014. arXiv preprint arXiv:1409.1556."},{"key":"10582_CR10","doi-asserted-by":"crossref","unstructured":"Boukhari DE, Chemsa A, Baarir Z-E. MobileViT architecture for Facial Beauty Prediction. In: 2024 International Conference on Telecommunications and Intelligent Systems (ICTIS). IEEE. 2024.","DOI":"10.1109\/ICTIS62692.2024.10894508"},{"key":"10582_CR11","unstructured":"Dosovitskiy A, Beyer L, Kolesnikov A, et al. An image is worth 16x16 words: Transformers for image recognition at scale. 2020. arXiv preprint arXiv: 2010.11929."},{"issue":"200","key":"10582_CR12","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1145\/3505244","volume":"54","author":"S Khan","year":"2022","unstructured":"Khan S, Naseer M, Hayat M, et al. Transformers in vision: A survey. ACM Comput Surv. 2022;54(200):1\u201341.","journal-title":"ACM Comput Surv"},{"key":"10582_CR13","doi-asserted-by":"crossref","unstructured":"Yun S, et al. Cutmix: Regularization strategy to train strong classifiers with localizable features. Proceed IEEE\/CVF Intern Conf Comput Vision. 2019.","DOI":"10.1109\/ICCV.2019.00612"},{"key":"10582_CR14","doi-asserted-by":"crossref","unstructured":"Ali R, Cui H. Artificial intelligence in facial measurement: a new era of symmetry and proportions analysis. Aesthet Plast Surg 2025;1\u201313.","DOI":"10.1007\/s00266-025-04746-7"},{"key":"10582_CR15","unstructured":"Iandola F, et al. Densenet: Implementing efficient convnet descriptor pyramids. 2014. arXiv preprint arXiv:1404.1869."},{"issue":"1","key":"10582_CR16","first-page":"125","volume":"56","author":"DE Boukhari","year":"2023","unstructured":"Boukhari DE, et al. Facial Beauty Prediction Using an Ensemble of Deep Convolutional Neural Networks. Eng Proceed. 2023;56(1):125.","journal-title":"Eng Proceed"},{"key":"10582_CR17","first-page":"2021","volume":"1","author":"I Lebedeva","year":"1922","unstructured":"Lebedeva I, Guo Y, Ying F. Transfer learning adaptive facial attractiveness assessment. J Phys: Conf Ser. 1922;1:2021.","journal-title":"J Phys: Conf Ser."},{"issue":"1","key":"10582_CR18","doi-asserted-by":"publisher","first-page":"10","DOI":"10.14500\/aro.11850","volume":"13","author":"AH Ibrahem","year":"2025","unstructured":"Ibrahem AH, Abdulazeez AM. A comprehensive review of facial beauty prediction using multi-task learning and facial attributes. Aro-the Scientif J Koya Univ. 2025;13(1):10\u201321.","journal-title":"Aro-the Scientif J Koya Univ."},{"issue":"4","key":"10582_CR19","doi-asserted-by":"publisher","first-page":"35","DOI":"10.1007\/s13735-025-00387-3","volume":"14","author":"DE Boukhari","year":"2025","unstructured":"Boukhari DE, Chemsa A. Enhancing Facial Beauty Prediction via a Dual-Pathway Hybrid Architecture Integrating Vmamba and ViT. Intern J Multimed Inf Retriev. 2025;14(4):35.","journal-title":"Intern J Multimed Inf Retriev"},{"issue":"325","key":"10582_CR20","first-page":"1","volume":"23","author":"L Carratino","year":"2022","unstructured":"Carratino L, et al. On mixup regularization. J Mach Learn Res. 2022;23(325):1\u201331.","journal-title":"J Mach Learn Res"},{"key":"10582_CR21","unstructured":"Zhang H, et al. mixup: Beyond empirical risk minimization. 2017. arXiv preprint arXiv:1710.09412."},{"key":"10582_CR22","unstructured":"DeVries T, Taylor GW. Improved regularization of convolutional neural networks with cutout. 2017. arXiv preprint arXiv:1708.04552."},{"key":"10582_CR23","unstructured":"Uddin AFM, et al. Saliencymix: A saliency guided data augmentation strategy for better regularization. 2020. arXiv preprint arXiv:2006.01791."},{"key":"10582_CR24","unstructured":"Vaswani A, et al. Attention is all you need. Adv Neural Inf Process Syst. 2017;30."},{"key":"10582_CR25","unstructured":"Hatamizadeh A, et al. Global context vision transformers. Intern Conf Mach Learn. 2023. PMLR."},{"key":"10582_CR26","unstructured":"Loshchilov I, Hutter F. Fixing weight decay regularization in adam. 2017;5:5. arXiv preprint arXiv:1711.05101."},{"key":"10582_CR27","doi-asserted-by":"publisher","first-page":"5099","DOI":"10.1007\/s00500-023-07963-x","volume":"27","author":"F Dornaika","year":"2023","unstructured":"Dornaika F. Multi-similarity semi-supervised manifold embedding for facial attractiveness scoring. Soft Comput. 2023;27:5099\u2013108.","journal-title":"Soft Comput"},{"key":"10582_CR28","doi-asserted-by":"crossref","unstructured":"Eddine Boukhari D, Chemsa A, Baarir Z-E. Facial Beauty Prediction Using Global Context Vision Transformer. In: International Symposium on iNnovative Informatics of Biskra (ISNIB). Algeria: Biskra; 2025. p. 2025.","DOI":"10.1109\/ISNIB64820.2025.10983768"},{"key":"10582_CR29","doi-asserted-by":"crossref","unstructured":"Shi S, Gao F, Meng X, et al. Improving Facial Attractiveness Prediction via Co-attention Learning. In: IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). Brighton: UK; 2019. p. 4045-4049.","DOI":"10.1109\/ICASSP.2019.8683112"},{"key":"10582_CR30","doi-asserted-by":"crossref","unstructured":"Gan J, Xie X, Zhai Y, He G, Mai C, Luo H. Facial Beauty Prediction Fusing Transfer Learning and Broad Learning System. Soft Comput. 2022.","DOI":"10.21203\/rs.3.rs-1349480\/v1"},{"key":"10582_CR31","doi-asserted-by":"crossref","unstructured":"Peng T, Li M, Chen F, et al. Geometric prior guided hybrid deep neural network for facial beauty analysis. In: CAAI Transactions on Intelligence Technology. 2023. p. 1\u201314.","DOI":"10.1049\/cit2.12197"},{"key":"10582_CR32","doi-asserted-by":"crossref","unstructured":"Xie D, Liang L, Jin L, et al. Scut-fbp: A benchmark dataset for facial beauty perception. In: IEEE International Conference on Systems, Man, and Cybernetics. Hong Kong, China; 2015. p. 1821\u20131826.","DOI":"10.1109\/SMC.2015.319"},{"key":"10582_CR33","doi-asserted-by":"publisher","first-page":"112009","DOI":"10.1016\/j.engappai.2025.112009","volume":"161","author":"DE Boukhari","year":"2025","unstructured":"Boukhari DE, et al. A comprehensive review of facial beauty prediction using deep learning techniques. Eng Appl Artif Intell. 2025;161:112009.","journal-title":"Eng Appl Artif Intell"},{"issue":"1","key":"10582_CR34","first-page":"1","volume":"2","author":"J Saeed","year":"2021","unstructured":"Saeed J, Abdulazeez AM. Facial beauty prediction and analysis based on deep convolutional neural network: a review. J Soft Comput Data Min. 2021;2(1):1\u201312.","journal-title":"J Soft Comput Data Min"},{"key":"10582_CR35","doi-asserted-by":"crossref","unstructured":"Cao K, Choi K, Jung H, et al. Deep learning for facial beauty prediction. Information. 2020;11(8).","DOI":"10.3390\/info11080391"},{"key":"10582_CR36","doi-asserted-by":"crossref","unstructured":"Bougourzi F, Dornaika F, Taleb-Ahmed A. Deep learning based face beauty prediction via dynamic robust losses and ensemble regression. Knowl-Based Syst. 2022;242(108246).","DOI":"10.1016\/j.knosys.2022.108246"},{"key":"10582_CR37","unstructured":"Krizhevsky A, Sutskever I, Hinton GE. Imagenet classification with deep convolutional neural networks. Adv Neural Inf Process Syst. 2012;25."},{"key":"10582_CR38","unstructured":"Lin L, Liang L, Jin L. Regression Guided by Relative Ranking Using Convolutional Neural Network (R3CNN) for Facial Beauty Prediction. IEEE Trans Affect Comput. 2019;1."},{"key":"10582_CR39","doi-asserted-by":"crossref","unstructured":"Xie S, Girshick R, Doll\u00e1r P, Tu Z, He K. Aggregated Residual Transformations for Deep Neural Networks. In: Proceedings of the 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Honolulu, HI, USA, 21\u201326. 2017. p. 5987\u20135995.","DOI":"10.1109\/CVPR.2017.634"},{"key":"10582_CR40","doi-asserted-by":"crossref","unstructured":"Xu J, Jin L, Liang L, et al. Facial attractiveness prediction using psychologically inspired convolutional neural network (PI-CNN). In: IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), New Orleans, LA, USA; 2017. p. 1657\u20131661.","DOI":"10.1109\/ICASSP.2017.7952438"},{"issue":"8","key":"10582_CR41","doi-asserted-by":"publisher","first-page":"2196","DOI":"10.1109\/TMM.2017.2780762","volume":"20","author":"Y-Y Fan","year":"2017","unstructured":"Fan Y-Y, et al. Label distribution-based facial attractiveness computation by deep residual learning. IEEE Trans Multimed. 2017;20(8):2196\u2013208.","journal-title":"IEEE Trans Multimed"},{"key":"10582_CR42","doi-asserted-by":"crossref","unstructured":"Lin L, Liang L, Jin L, et al. Attribute-Aware Convolutional Neural Networks for Facial Beauty Prediction. In: Proceedings of the twenty-eighth international joint conference on artificial intelligence. 2019. p.847\u2013853.","DOI":"10.24963\/ijcai.2019\/119"},{"issue":"1","key":"10582_CR43","doi-asserted-by":"publisher","first-page":"55","DOI":"10.18280\/ts.400105","volume":"40","author":"JN Saeed","year":"2023","unstructured":"Saeed JN, Abdulazeez AM, Ibrahim DA. An ensemble dcnns-based regression model for automatic facial beauty prediction and analyzation. Traitement du Signal. 2023;40(1):55.","journal-title":"Traitement du Signal"},{"issue":"2","key":"10582_CR44","doi-asserted-by":"publisher","first-page":"239","DOI":"10.1587\/transinf.2023EDL8058","volume":"107","author":"Z Sun","year":"2024","unstructured":"Sun Z, et al. Dynamic attentive convolution for facial beauty prediction. IEICE Trans Inf Syst. 2024;107(2):239\u201343.","journal-title":"IEICE Trans Inf Syst."},{"key":"10582_CR45","volume-title":"Face processing: Psychological, neuropsychological, and applied perspectives","author":"G Hole","year":"2010","unstructured":"Hole G, Bourne V. Face processing: Psychological, neuropsychological, and applied perspectives. USA: Oxford University Press; 2010."},{"key":"10582_CR46","doi-asserted-by":"crossref","unstructured":"Xiao L, et al. Exploring cognitive and aesthetic causality for multimodal aspect-based sentiment analysis. IEEE Trans Affect Comput. 2025.","DOI":"10.1109\/TAFFC.2025.3565506"},{"key":"10582_CR47","doi-asserted-by":"crossref","unstructured":"Kruk J, Ziems C, Yang D. Impressions: Visual semiotics and aesthetic impact understanding. In: Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023.","DOI":"10.18653\/v1\/2023.emnlp-main.755"},{"key":"10582_CR48","doi-asserted-by":"crossref","unstructured":"Xiao L, et al. Vanessa: Visual connotation and aesthetic attributes understanding network for multimodal aspect-based sentiment analysis. In: Findings of the Association for Computational Linguistics: EMNLP 2024. 2024.","DOI":"10.18653\/v1\/2024.findings-emnlp.671"},{"issue":"1","key":"10582_CR49","first-page":"19932","volume":"14","author":"E Stamkou","year":"2024","unstructured":"Stamkou E, et al. Emotional palette: a computational mapping of aesthetic experiences evoked by visual art. Scientif Rep. 2024;14(1):19932.","journal-title":"Scientif Rep"},{"key":"10582_CR50","doi-asserted-by":"publisher","first-page":"102304","DOI":"10.1016\/j.inffus.2024.102304","volume":"106","author":"L Xiao","year":"2024","unstructured":"Xiao L, et al. Atlantis: Aesthetic-oriented multiple granularities fusion network for joint multimodal aspect-based sentiment analysis. Inf Fusion. 2024;106:102304.","journal-title":"Inf Fusion"},{"key":"10582_CR51","unstructured":"Liu J, et al. Tokenmix: Rethinking image mixing for data augmentation in vision transformers. European Conf Comput Vision."},{"key":"10582_CR52","doi-asserted-by":"crossref","unstructured":"Chen J-N, et al. Transmix: Attend to mix for vision transformers. In: Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition. Cham: Springer Nature Switzerland; 2022.","DOI":"10.1109\/CVPR52688.2022.01182"}],"container-title":["Cognitive Computation"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s12559-026-10582-x.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s12559-026-10582-x","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s12559-026-10582-x.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,5,2]],"date-time":"2026-05-02T05:51:15Z","timestamp":1777701075000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s12559-026-10582-x"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,5,2]]},"references-count":52,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2026,12]]}},"alternative-id":["10582"],"URL":"https:\/\/doi.org\/10.1007\/s12559-026-10582-x","relation":{},"ISSN":["1866-9956","1866-9964"],"issn-type":[{"value":"1866-9956","type":"print"},{"value":"1866-9964","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,5,2]]},"assertion":[{"value":"14 October 2025","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"18 April 2026","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"2 May 2026","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"This article does not contain any studies with human participants or animals performed by any of the authors.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Ethical Approval"}},{"value":"The authors declare no competing interests.","order":3,"name":"Ethics","group":{"name":"EthicsHeading","label":"Competing interests"}}],"article-number":"41"}}