{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,11,2]],"date-time":"2025-11-02T06:09:14Z","timestamp":1762063754911,"version":"build-2065373602"},"reference-count":59,"publisher":"MDPI AG","issue":"9","license":[{"start":{"date-parts":[[2022,8,30]],"date-time":"2022-08-30T00:00:00Z","timestamp":1661817600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Future Internet"],"abstract":"<jats:p>Facial Expression Recognition (FER) can achieve an understanding of the emotional changes of a specific target group. The relatively small dataset related to facial expression recognition and the lack of a high accuracy of expression recognition are both a challenge for researchers. In recent years, with the rapid development of computer technology, especially the great progress of deep learning, more and more convolutional neural networks have been developed for FER research. Most of the convolutional neural performances are not good enough when dealing with the problems of overfitting from too-small datasets and noise, due to expression-independent intra-class differences. In this paper, we propose a Dual Path Stacked Attention Network (DPSAN) to better cope with the above challenges. Firstly, the features of key regions in faces are extracted using segmentation, and irrelevant regions are ignored, which effectively suppresses intra-class differences. Secondly, by providing the global image and segmented local image regions as training data for the integrated dual path model, the overfitting problem of the deep network due to a lack of data can be effectively mitigated. Finally, this paper also designs a stacked attention module to weight the fused feature maps according to the importance of each part for expression recognition. For the cropping scheme, this paper chooses to adopt a cropping method based on the fixed four regions of the face image, to segment out the key image regions and to ignore the irrelevant regions, so as to improve the efficiency of the algorithm computation. The experimental results on the public datasets, CK+ and FERPLUS, demonstrate the effectiveness of DPSAN, and its accuracy reaches the level of current state-of-the-art methods on both CK+ and FERPLUS, with 93.2% and 87.63% accuracy on the CK+ dataset and FERPLUS dataset, respectively.<\/jats:p>","DOI":"10.3390\/fi14090258","type":"journal-article","created":{"date-parts":[[2022,8,30]],"date-time":"2022-08-30T21:25:18Z","timestamp":1661894718000},"page":"258","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":4,"title":["Facial Expression Recognition Using Dual Path Feature Fusion and Stacked Attention"],"prefix":"10.3390","volume":"14","author":[{"ORCID":"https:\/\/orcid.org\/0000-0003-3026-4237","authenticated-orcid":false,"given":"Hongtao","family":"Zhu","sequence":"first","affiliation":[{"name":"School of Computer Engineering and Science, Shanghai University, Shanghai 200444, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Huahu","family":"Xu","sequence":"additional","affiliation":[{"name":"School of Computer Engineering and Science, Shanghai University, Shanghai 200444, China"},{"name":"Shanghai Shangda Hairun Information System Co., Ltd., Shanghai 200072, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Xiaojin","family":"Ma","sequence":"additional","affiliation":[{"name":"School of Computer Engineering and Science, Shanghai University, Shanghai 200444, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Minjie","family":"Bian","sequence":"additional","affiliation":[{"name":"School of Computer Engineering and Science, Shanghai University, Shanghai 200444, China"},{"name":"Shanghai Shangda Hairun Information System Co., Ltd., Shanghai 200072, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2022,8,30]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","first-page":"97","DOI":"10.1109\/34.908962","article-title":"Recognizing action units for facial expression analysis","volume":"23","author":"Tian","year":"2001","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_2","doi-asserted-by":"crossref","unstructured":"Darwin, C., and Prodger, P. (1998). The Expression of the Emotions in Man and Animals, Oxford University Press.","DOI":"10.1093\/oso\/9780195112719.002.0002"},{"key":"ref_3","doi-asserted-by":"crossref","unstructured":"Dhall, A., Kaur, A., Goecke, R., and Gedeon, T. (2018, January 16\u201320). Emotiw 2018: Audio-Video, student engagement and group-level affect prediction. Proceedings of the 20th ACM International Conference on Multimodal Interaction, Boulder, CO, USA.","DOI":"10.1145\/3242969.3264993"},{"key":"ref_4","doi-asserted-by":"crossref","unstructured":"Fabian Benitez-Quiroz, C., Srinivasan, R., and Martinez, A.M. (2016, January 27\u201330). Emotionet: An accurate, real-time algorithm for the automatic annotation of a million facial expressions in the wild. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.600"},{"key":"ref_5","unstructured":"Dominguez-Catena, I., Paternain, D., and Galar, M. (2022). Assessing Demographic Bias Transfer from Dataset to Model: A Case Study in Facial Expression Recognition. arXiv."},{"key":"ref_6","doi-asserted-by":"crossref","unstructured":"Li, S., Deng, W., and Du, J. (2017, January 21\u201326). Reliable Crowdsourcing and Deep Locality-Preserving Learning for Expression Recognition in the Wild. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.277"},{"key":"ref_7","doi-asserted-by":"crossref","first-page":"18","DOI":"10.1109\/TAFFC.2017.2740923","article-title":"AffectNet: A Database for Facial Expression, Valence, and Arousal Computing in the Wild","volume":"10","author":"Mollahosseini","year":"2017","journal-title":"IEEE Trans. Affect. Comput."},{"key":"ref_8","doi-asserted-by":"crossref","unstructured":"Cai, J., Meng, Z., Khan, A.S., Li, Z., O\u2019Reilly, J., and Tong, Y. (2018, January 15\u201319). Island loss for learning discriminative features in facial expression recognition. Proceedings of the 2018 13th IEEE International Conference on Automatic Face & Gesture Recognition (FG 2018), Xi\u2019an, China.","DOI":"10.1109\/FG.2018.00051"},{"key":"ref_9","doi-asserted-by":"crossref","unstructured":"Hou, C., Ai, J., Lin, Y., Guan, C., Li, J., and Zhu, W. (2022). Evaluation of Online Teaching Quality Based on Facial Expression Recognition. Future Internet, 14.","DOI":"10.3390\/fi14060177"},{"key":"ref_10","doi-asserted-by":"crossref","unstructured":"Sangerm\u00e1n Jim\u00e9nez, M.A., Ponce, P., and V\u00e1zquez-Cano, E. (2021). YouTube Videos in the Virtual Flipped Classroom Model using Brain Signals and Facial Expressions. Future Internet, 13.","DOI":"10.3390\/fi13090224"},{"key":"ref_11","doi-asserted-by":"crossref","first-page":"2439","DOI":"10.1109\/TIP.2018.2886767","article-title":"Occlusion Aware Facial Expression Recognition Using CNN with Attention Mechanism","volume":"28","author":"Li","year":"2018","journal-title":"IEEE Trans. Image Process."},{"key":"ref_12","doi-asserted-by":"crossref","unstructured":"Liu, K., Zhang, M., and Pan, Z. (2016, January 28\u201330). Facial expression recognition with CNN ensemble. Proceedings of the 2016 International Conference on Cyberworlds (CW), Chongqing, China.","DOI":"10.1109\/CW.2016.34"},{"key":"ref_13","doi-asserted-by":"crossref","unstructured":"Roy, S., and Etemad, A. (2022). Analysis of Semi-Supervised Methods for Facial Expression Recognition. arXiv.","DOI":"10.1109\/ACII55700.2022.9953876"},{"key":"ref_14","doi-asserted-by":"crossref","unstructured":"Siqueira, H., Magg, S., and Wermter, S. (2020, January 3). Efficient Facial Feature Learning with Wide Ensemble-Based Convolutional Neural Networks. Proceedings of the AAAI Conference on Artificial Intelligence, Palo Alto, CA, USA.","DOI":"10.1609\/aaai.v34i04.6037"},{"key":"ref_15","doi-asserted-by":"crossref","unstructured":"Wang, K., Peng, X., Yang, J., Lu, S., and Qiao, Y. (2020, January 13\u201319). Suppressing uncertainties for large-scale facial expression recognition. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA.","DOI":"10.1109\/CVPR42600.2020.00693"},{"key":"ref_16","doi-asserted-by":"crossref","first-page":"4057","DOI":"10.1109\/TIP.2019.2956143","article-title":"Region Attention Networks for Pose and Occlusion Robust Facial Expression Recognition","volume":"29","author":"Wang","year":"2020","journal-title":"IEEE Trans. Image Process."},{"key":"ref_17","doi-asserted-by":"crossref","unstructured":"Yu, Z., and Zhang, C. (2015, January 9\u201313). Image Based Static Facial Expression Recognition with Multiple Deep Network Learning. Proceedings of the 2015 ACM on International Conference on Multimodal Interaction, Seattle, WA, USA.","DOI":"10.1145\/2818346.2830595"},{"key":"ref_18","doi-asserted-by":"crossref","unstructured":"Gloor, P.A., Fronzetti Colladon, A., Altuntas, E., Cetinkaya, C., Kaiser, M.F., Ripperger, L., and Schaefer, T. (2022). Your Face Mirrors Your Deepest Beliefs\u2014Predicting Personality and Morals through Facial Emotion Recognition. Future Internet, 14.","DOI":"10.3390\/fi14010005"},{"key":"ref_19","doi-asserted-by":"crossref","unstructured":"Lucey, P., Cohn, J.F., Kanade, T., Saragih, J., Ambadar, Z., and Matthews, I. (2010, January 13\u201318). The Extended Cohn-Kanade Dataset (CK+): A Complete Dataset for Action Unit and Emotion-Specified Expression. Proceedings of the 2010 IEEE Computer Society Conference on Computer Vision and Pattern Recognition-Workshops, San Francisco, CA, USA.","DOI":"10.1109\/CVPRW.2010.5543262"},{"key":"ref_20","doi-asserted-by":"crossref","unstructured":"Goodfellow, I.J., Erhan, D., Carrier, P.L., Courville, A., Mirza, M., Hamner, B., Cukierski, W., Tang, Y., Thaler, D., and Bengio, Y. (2013, January 18\u201322). Challenges in representation learning: A report on three machine learning contests. Proceedings of the International Conference on Neural Information Processing, Bangkok, Thailand.","DOI":"10.1007\/978-3-642-42051-1_16"},{"key":"ref_21","doi-asserted-by":"crossref","unstructured":"Barsoum, E., Zhang, C., Ferrer, C.C., and Zhang, Z. (2016, January 12\u201316). Training deep networks for facial expression recognition with crowd-sourced label distribution. Proceedings of the 18th ACM International Conference on Multimodal Interaction, Tokyo, Japan.","DOI":"10.1145\/2993148.2993165"},{"key":"ref_22","unstructured":"Li, S., and Deng, W. (2020). Deep Facial Expression Recognition: A Survey. IEEE Trans. Affect. Comput."},{"key":"ref_23","first-page":"20","article-title":"Openface: A general-purpose face recognition library with mobile applications","volume":"6","author":"Amos","year":"2016","journal-title":"CMU Sch. Comput. Sci."},{"key":"ref_24","doi-asserted-by":"crossref","unstructured":"Fang, M., Boutros, F., and Damer, N. (2022). Unsupervised Face Morphing Attack Detection via Self-paced Anomaly Detection. arXiv.","DOI":"10.1109\/IJCB54206.2022.10008003"},{"key":"ref_25","doi-asserted-by":"crossref","unstructured":"Neto, P.C., Boutros, F., Pinto, J.R., Damer, N., Sequeira, A.F., Cardoso, J.S., Bengherabi, M., Bousnat, A., Boucheta, S., and Menotti, D. (2022). OCFR 2022: Competition on Occluded Face Recognition from Synthetically Generated Structure-Aware Occlusions. arXiv.","DOI":"10.1109\/IJCB54206.2022.10007963"},{"key":"ref_26","doi-asserted-by":"crossref","unstructured":"He, K., Zhang, X., Ren, S., and Sun, J. (2016, January 27\u201330). Deep residual learning for image recognition. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.90"},{"key":"ref_27","doi-asserted-by":"crossref","unstructured":"Thakur, N., and Han, C.Y. (2021). Indoor Localization for Personalized Ambient Assisted Living of Multiple Users in Multi-Floor Smart Environments. Big Data Cogn. Comput., 5.","DOI":"10.3390\/bdcc5030042"},{"key":"ref_28","doi-asserted-by":"crossref","unstructured":"Guerra, B.M.V., Schmid, M., Beltrami, G., and Ramat, S. (2022). Neural Networks for Automatic Posture Recognition in Ambient-Assisted Living. Sensors, 22.","DOI":"10.3390\/s22072609"},{"key":"ref_29","doi-asserted-by":"crossref","first-page":"915","DOI":"10.1109\/TPAMI.2007.1110","article-title":"Dynamic Texture Recognition using Local Binary Patterns with an Application to Facial Expressions","volume":"29","author":"Zhao","year":"2007","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_30","unstructured":"Zhong, L., Liu, Q., Yang, P., Liu, B., Huang, J., and Metaxas, D. (2012, January 16\u201321). Learning active facial patches for expression analysis. Proceedings of the 2012 IEEE Conference on Computer Vision and Pattern Recognition, Providence, RI, USA."},{"key":"ref_31","unstructured":"Dalal, N., and Triggs, B. (2005, January 20\u201326). Histograms of oriented gradients for human detection. Proceedings of the 2005 IEEE computer society Conference on Computer Vision and Pattern Recognition (CVPR\u201905), San Diego, CA, USA."},{"key":"ref_32","doi-asserted-by":"crossref","first-page":"803","DOI":"10.1016\/j.imavis.2008.08.005","article-title":"Facial expression recognition based on Local Binary Patterns: A comprehensive study","volume":"27","author":"Shan","year":"2009","journal-title":"Image Vis. Comput."},{"key":"ref_33","doi-asserted-by":"crossref","first-page":"91","DOI":"10.1023\/B:VISI.0000029664.99615.94","article-title":"Distinctive image features from scale-invariant keypoints","volume":"60","author":"Lowe","year":"2004","journal-title":"Int. J. Comput. Vis."},{"key":"ref_34","doi-asserted-by":"crossref","first-page":"1021","DOI":"10.1007\/s00371-011-0611-x","article-title":"3D facial expression recognition using SIFT descriptors of automatically detected keypoints","volume":"27","author":"Berretti","year":"2011","journal-title":"Vis. Comput."},{"key":"ref_35","doi-asserted-by":"crossref","first-page":"645","DOI":"10.1186\/s40064-015-1427-3","article-title":"Facial expression recognition and histograms of oriented gradients: A comprehensive study","volume":"4","author":"Leo","year":"2015","journal-title":"SpringerPlus"},{"key":"ref_36","unstructured":"Shan, C., Gong, S., and McOwan, P. (2005, January 11\u201314). Robust facial expression recognition using local binary patterns. Proceedings of the IEEE International Conference on Image Processing 2005, Genoa, Italy."},{"key":"ref_37","doi-asserted-by":"crossref","unstructured":"Thakur, N., and Han, C.Y. (2021). Country-Specific Interests towards Fall Detection from 2004\u20132021: An Open Access Dataset and Research Questions. Data, 6.","DOI":"10.3390\/data6080092"},{"key":"ref_38","unstructured":"Wang, Z., Wang, G., Huang, B., Xiong, Z., Hong, Q., Wu, H., Yi, P., Jiang, K., Wang, N., and Pei, Y. (2020). Masked face recognition dataset and application. arXiv."},{"key":"ref_39","doi-asserted-by":"crossref","first-page":"580","DOI":"10.1162\/jocn.2006.18.4.580","article-title":"Specialized Face Perception Mechanisms Extract Both Part and Spacing Information: Evidence from Developmental Prosopagnosia","volume":"18","author":"Yovel","year":"2006","journal-title":"J. Cogn. Neurosci."},{"key":"ref_40","doi-asserted-by":"crossref","first-page":"384","DOI":"10.1037\/0003-066X.48.4.384","article-title":"Facial expression and emotion","volume":"48","author":"Ekman","year":"1993","journal-title":"Am. Psychol."},{"key":"ref_41","doi-asserted-by":"crossref","first-page":"1052","DOI":"10.1016\/j.imavis.2007.11.004","article-title":"An analysis of facial expression recognition under partial facial image occlusion","volume":"26","author":"Kotsia","year":"2008","journal-title":"Image Vis. Comput."},{"key":"ref_42","doi-asserted-by":"crossref","unstructured":"Ciregan, D., Meier, U., and Schmidhuber, J. (2012, January 16\u201321). Multi-Column deep neural networks for image classification. Proceedings of the 2012 IEEE Conference on Computer Vision and Pattern Recognition, Providence, RI, USA.","DOI":"10.1109\/CVPR.2012.6248110"},{"key":"ref_43","doi-asserted-by":"crossref","unstructured":"He, K., Zhang, X., Ren, S., and Sun, J. (2016). Identity Mappings in Deep Residual Networks. European Conference on Computer Vision, Springer.","DOI":"10.1007\/978-3-319-46493-0_38"},{"key":"ref_44","doi-asserted-by":"crossref","first-page":"3174","DOI":"10.1007\/s11263-021-01521-4","article-title":"Pixel-in-Pixel Net: Towards Efficient Facial Landmark Detection in the Wild","volume":"129","author":"Jin","year":"2021","journal-title":"Int. J. Comput. Vis."},{"key":"ref_45","doi-asserted-by":"crossref","first-page":"2016","DOI":"10.1109\/TIP.2021.3049955","article-title":"Adaptively Learning Facial Expression Representation via C-F Labels and Distillation","volume":"30","author":"Li","year":"2021","journal-title":"IEEE Trans. Image Process."},{"key":"ref_46","doi-asserted-by":"crossref","unstructured":"Bargal, S.A., Barsoum, E., Ferrer, C.C., and Zhang, C. (2016, January 12\u201316). Emotion recognition in the wild from videos using images. Proceedings of the 18th ACM International Conference on Multimodal Interaction, Tokyo, Japan.","DOI":"10.1145\/2993148.2997627"},{"key":"ref_47","doi-asserted-by":"crossref","unstructured":"Zhang, J., Kan, M., Shan, S., and Chen, X. (2016, January 27\u201330). Occlusion-Free Face Alignment: Deep Regression Networks Coupled with De-Corrupt AutoEncoders. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.373"},{"key":"ref_48","doi-asserted-by":"crossref","unstructured":"Woo, S., Park, J., Lee, J.Y., and Kweon, I.S. (2018, January 17\u201324). Cbam: Convolutional block attention module. Proceedings of the European Conference on Computer Vision (ECCV), Munich, Germany.","DOI":"10.1007\/978-3-030-01234-2_1"},{"key":"ref_49","unstructured":"Paszke, A., Gross, S., Massa, F., Lerer, A., Bradbury, J., Chanan, G., Killeen, T., Lin, Z., Gimelshein, N., and Chintala, S. (2019). Pytorch: An imperative style, high-performance deep learning library. Adv. Neural Inf. Processing Syst., 32."},{"key":"ref_50","unstructured":"Abadi, M., Agarwal, A., Barham, P., Brevdo, E., Chen, Z., Citro, C., Corrado, G.S., Davis, A., Dean, J., and Zheng, X. (2016). Tensorflow: Large-Scale machine learning on heterogeneous distributed systems. arXiv."},{"key":"ref_51","unstructured":"Croci, M.L., Sengupta, U., and Juniper, M.P. (2021). Online parameter inference for the simulation of a Bunsen flame using heteroscedastic Bayesian neural network ensembles. arXiv."},{"key":"ref_52","doi-asserted-by":"crossref","unstructured":"Qureshi, A.S., and Roos, T. (2021). Transfer Learning with Ensembles of Deep Neural Networks for Skin Cancer Detection in Imbalanced Data Sets. arXiv.","DOI":"10.1007\/s11063-022-11049-4"},{"key":"ref_53","first-page":"29","article-title":"Evaluating Deep Neural Network Ensembles by Majority Voting Cum Meta-Learning Scheme","volume":"Volume 410","author":"Jain","year":"2021","journal-title":"Soft Computing and Signal Processing"},{"key":"ref_54","unstructured":"Liu, M., Li, S., Shan, S., and Chen, X. (2013, January 22\u201326). Au-Aware deep networks for facial expression recognition. Proceedings of the 2013 10th IEEE International Conference and Workshops on Automatic Face and Gesture Recognition (FG), Shanghai, China."},{"key":"ref_55","doi-asserted-by":"crossref","first-page":"643","DOI":"10.1016\/j.neucom.2017.08.043","article-title":"Facial expression recognition via learning deep sparse autoencoders","volume":"273","author":"Zeng","year":"2018","journal-title":"Neurocomputing"},{"key":"ref_56","doi-asserted-by":"crossref","unstructured":"Ding, H., Zhou, S.K., and Chellappa, R. (June, January 30). Facenet2expnet: Regularizing a deep face recognition net for expression recognition. Proceedings of the 2017 12th IEEE International Conference on Automatic Face & Gesture Recognition (FG 2017), Washington, DC, USA.","DOI":"10.1109\/FG.2017.23"},{"key":"ref_57","doi-asserted-by":"crossref","first-page":"78000","DOI":"10.1109\/ACCESS.2019.2921220","article-title":"Recognizing Facial Expressions Using a Shallow Convolutional Neural Network","volume":"7","author":"Miao","year":"2019","journal-title":"IEEE Access"},{"key":"ref_58","doi-asserted-by":"crossref","first-page":"321","DOI":"10.1007\/s42979-020-00325-6","article-title":"The FaceChannel: A Fast and Furious Deep Neural Network for Facial Expression Recognition","volume":"1","author":"Barros","year":"2020","journal-title":"SN Comput. Sci."},{"key":"ref_59","doi-asserted-by":"crossref","first-page":"544","DOI":"10.1109\/TAFFC.2018.2880201","article-title":"Facial Expression Recognition with Identity and Emotion Joint Learning","volume":"12","author":"Li","year":"2018","journal-title":"IEEE Trans. Affect. Comput."}],"container-title":["Future Internet"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1999-5903\/14\/9\/258\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,11]],"date-time":"2025-10-11T00:20:04Z","timestamp":1760142004000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1999-5903\/14\/9\/258"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,8,30]]},"references-count":59,"journal-issue":{"issue":"9","published-online":{"date-parts":[[2022,9]]}},"alternative-id":["fi14090258"],"URL":"https:\/\/doi.org\/10.3390\/fi14090258","relation":{},"ISSN":["1999-5903"],"issn-type":[{"type":"electronic","value":"1999-5903"}],"subject":[],"published":{"date-parts":[[2022,8,30]]}}}