{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,7]],"date-time":"2026-07-07T04:31:14Z","timestamp":1783398674559,"version":"3.54.6"},"reference-count":52,"publisher":"MDPI AG","issue":"7","license":[{"start":{"date-parts":[[2023,3,24]],"date-time":"2023-03-24T00:00:00Z","timestamp":1679616000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["62176018"],"award-info":[{"award-number":["62176018"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]},{"name":"Beijing Key Laboratory of Robot Bionics and Function Research","award":["62176018"],"award-info":[{"award-number":["62176018"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Sensors"],"abstract":"<jats:p>There are problems associated with facial expression recognition (FER), such as facial occlusion and head pose variations. These two problems lead to incomplete facial information in images, making feature extraction extremely difficult. Most current methods use prior knowledge or fixed-size patches to perform local cropping, thereby enhancing the ability to acquire fine-grained features. However, the former requires extra data processing work and is prone to errors; the latter destroys the integrity of local features. In this paper, we propose a local Sliding Window Attention Network (SWA-Net) for FER. Specifically, we propose a sliding window strategy for feature-level cropping, which preserves the integrity of local features and does not require complex preprocessing. Moreover, the local feature enhancement module mines fine-grained features with intraclass semantics through a multiscale depth network. The adaptive local feature selection module is introduced to prompt the model to find more essential local features. Extensive experiments demonstrate that our SWA-Net model achieves a comparable performance to that of state-of-the-art methods with scores of 90.03% on RAF-DB, 89.22% on FERPlus, 63.97% on AffectNet.<\/jats:p>","DOI":"10.3390\/s23073424","type":"journal-article","created":{"date-parts":[[2023,3,24]],"date-time":"2023-03-24T06:34:07Z","timestamp":1679639647000},"page":"3424","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":11,"title":["Facial Expression Recognition Using Local Sliding Window Attention"],"prefix":"10.3390","volume":"23","author":[{"given":"Shuang","family":"Qiu","sequence":"first","affiliation":[{"name":"School of Electrical and Information Engineering, Beijing University of Civil Engineering and Architecture, Beijing 100044, China"},{"name":"Beijing Key Laboratory of Robot Bionics and Function Research, Beijing 100044, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Guangzhe","family":"Zhao","sequence":"additional","affiliation":[{"name":"School of Electrical and Information Engineering, Beijing University of Civil Engineering and Architecture, Beijing 100044, China"},{"name":"Beijing Key Laboratory of Robot Bionics and Function Research, Beijing 100044, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Xiao","family":"Li","sequence":"additional","affiliation":[{"name":"School of Electronics and Information Engineering, Zhongyuan University of Technology, Zhengzhou 450007, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-8987-7578","authenticated-orcid":false,"given":"Xueping","family":"Wang","sequence":"additional","affiliation":[{"name":"School of Electrical and Information Engineering, Beijing University of Civil Engineering and Architecture, Beijing 100044, China"},{"name":"Beijing Key Laboratory of Robot Bionics and Function Research, Beijing 100044, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1968","published-online":{"date-parts":[[2023,3,24]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","unstructured":"Chowdary, M.K., Nguyen, T.N., and Hemanth, D.J. (2021). Deep learning-based facial emotion recognition for human\u2013computer interaction applications. Neural Comput. Appl., 1\u201318.","DOI":"10.1007\/s00521-021-06012-8"},{"key":"ref_2","doi-asserted-by":"crossref","first-page":"4501","DOI":"10.1016\/j.ijleo.2015.08.185","article-title":"Driver fatigue recognition based on facial expression analysis using local binary patterns","volume":"126","author":"Zhang","year":"2015","journal-title":"Optik"},{"key":"ref_3","doi-asserted-by":"crossref","first-page":"5619","DOI":"10.1109\/TII.2022.3141400","article-title":"Impact of deep learning approaches on facial expression recognition in healthcare industries","volume":"18","author":"Bisogni","year":"2022","journal-title":"IEEE Trans. Ind. Inform."},{"key":"ref_4","unstructured":"Ruiz-Garcia, A., Webb, N., Palade, V., Eastwood, M., and Elshaw, M. Deep learning for real time facial expression recognition in social robots. Proceedings of the International Conference on Neural Information Processing (ICONIP)."},{"key":"ref_5","doi-asserted-by":"crossref","first-page":"2058","DOI":"10.1109\/TAFFC.2022.3208309","article-title":"FLEPNet: Feature Level Ensemble Parallel Network for Facial Expression Recognition","volume":"13","author":"Karnati","year":"2022","journal-title":"IEEE Trans. Affect. Comput."},{"key":"ref_6","doi-asserted-by":"crossref","unstructured":"Ruan, D., Yan, Y., Lai, S., Chai, Z., Shen, C., and Wang, H. (2021, January 19\u201325). Feature Decomposition and Reconstruction Learning for Effective Facial Expression Recognition. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Virtual.","DOI":"10.1109\/CVPR46437.2021.00757"},{"key":"ref_7","doi-asserted-by":"crossref","unstructured":"Lucey, P., Cohn, J.F., Kanade, T., Saragih, J., Ambadar, Z., and Matthews, I. (2010, January 13\u201318). The Extended Cohn-Kanade Dataset (CK+): A complete dataset for action unit and emotion-specified expression. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), San Francisco, CA, USA.","DOI":"10.1109\/CVPRW.2010.5543262"},{"key":"ref_8","unstructured":"Pantic, M., Valstar, M., Rademaker, R., and Maat, L. (2005, January 6\u20139). Web-based database for facial expression analysis. Proceedings of the IEEE International Conference on Multimedia and Expo (ICME), Amsterdam, The Netherlands."},{"key":"ref_9","unstructured":"Lyons, M., Akamatsu, S., Kamachi, M., and Gyoba, J. (1998, January 14\u201316). Coding facial expressions with Gabor wavelets. Proceedings of the IEEE International Conference on Automatic Face and Gesture Recognition (FG), Nara, Japan."},{"key":"ref_10","doi-asserted-by":"crossref","unstructured":"Goodfellow, I.J., Erhan, D., Carrier, P.L., Courville, A., Mirza, M., Hamner, B., Cukierski, W., Tang, Y., Thaler, D., and Lee, D.H. (2013, January 3\u20137). Challenges in representation learning: A report on three machine learning contests. Proceedings of the International Conference on Neural Information Processing (ICONIP), Daegu, Republic of Korea.","DOI":"10.1007\/978-3-642-42051-1_16"},{"key":"ref_11","doi-asserted-by":"crossref","unstructured":"Li, S., Deng, W., and Du, J. (2017, January 21\u201326). Reliable Crowdsourcing and Deep Locality-Preserving Learning for Expression Recognition in the Wild. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.277"},{"key":"ref_12","doi-asserted-by":"crossref","first-page":"18","DOI":"10.1109\/TAFFC.2017.2740923","article-title":"AffectNet: A Database for Facial Expression, Valence, and Arousal Computing in the Wild","volume":"10","author":"Mollahosseini","year":"2019","journal-title":"IEEE Trans. Affect. Comput."},{"key":"ref_13","doi-asserted-by":"crossref","unstructured":"Zhang, Y., Wang, C., Ling, X., and Deng, W. (2022, January 23\u201327). Learn from all: Erasing attention consistency for noisy label facial expression recognition. Proceedings of the European Conference on Computer Vision (ECCV), Tel Aviv, Israel.","DOI":"10.1007\/978-3-031-19809-0_24"},{"key":"ref_14","doi-asserted-by":"crossref","unstructured":"Xue, F., Wang, Q., and Guo, G. (2021, January 11\u201317). TransFER: Learning Relation-aware Facial Expression Representations with Transformers. Proceedings of the IEEE International Conference on Computer Vision (ICCV), Virtual Event.","DOI":"10.1109\/ICCV48922.2021.00358"},{"key":"ref_15","doi-asserted-by":"crossref","first-page":"3190","DOI":"10.1109\/TCSVT.2021.3103782","article-title":"Self-Supervised Exclusive-Inclusive Interactive Learning for Multi-Label Facial Expression Recognition in the Wild","volume":"32","author":"Li","year":"2022","journal-title":"IEEE Trans. Circuits Syst. Video Technol."},{"key":"ref_16","doi-asserted-by":"crossref","first-page":"6253","DOI":"10.1109\/TCSVT.2022.3165321","article-title":"Adaptive Multilayer Perceptual Attention Network for Facial Expression Recognition","volume":"32","author":"Liu","year":"2022","journal-title":"IEEE Trans. Circuits Syst. Video Technol."},{"key":"ref_17","doi-asserted-by":"crossref","first-page":"4057","DOI":"10.1109\/TIP.2019.2956143","article-title":"Region Attention Networks for Pose and Occlusion Robust Facial Expression Recognition","volume":"29","author":"Wang","year":"2020","journal-title":"IEEE Trans. Image Process."},{"key":"ref_18","doi-asserted-by":"crossref","first-page":"2439","DOI":"10.1109\/TIP.2018.2886767","article-title":"Occlusion Aware Facial Expression Recognition Using CNN With Attention Mechanism","volume":"28","author":"Li","year":"2019","journal-title":"IEEE Trans. Image Process."},{"key":"ref_19","doi-asserted-by":"crossref","first-page":"6544","DOI":"10.1109\/TIP.2021.3093397","article-title":"Learning Deep Global Multi-Scale and Local Attention Features for Facial Expression Recognition in the Wild","volume":"30","author":"Zhao","year":"2021","journal-title":"IEEE Trans. Image Process."},{"key":"ref_20","doi-asserted-by":"crossref","first-page":"3812","DOI":"10.1093\/nar\/gkg509","article-title":"SIFT: Predicting amino acid changes that affect protein function","volume":"31","author":"Ng","year":"2003","journal-title":"Nucleic Acids Res."},{"key":"ref_21","unstructured":"Dalal, N., and Triggs, B. (2005, January 20\u201326). Histograms of oriented gradients for human detection. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), San Diego, CA, USA."},{"key":"ref_22","doi-asserted-by":"crossref","first-page":"467","DOI":"10.1109\/TIP.2002.999679","article-title":"Gabor feature based classification using the enhanced fisher linear discriminant model for face recognition","volume":"11","author":"Liu","year":"2002","journal-title":"IEEE Trans. Image Process."},{"key":"ref_23","unstructured":"Fasel, B. (2002, January 11\u201315). Robust face analysis using convolutional neural networks. Proceedings of the 2002 International Conference on Pattern Recognition, Quebec City, QC, Canada."},{"key":"ref_24","unstructured":"Tang, Y. (2013). Deep learning using linear support vector machines. arXiv."},{"key":"ref_25","doi-asserted-by":"crossref","unstructured":"Kahou, S.E., Pal, C., Bouthillier, X., Froumenty, P., G\u00fcl\u00e7ehre, \u00c7., Memisevic, R., Vincent, P., Courville, A., Bengio, Y., and Ferrari, R.C. (2013, January 9\u201313). Combining modality specific deep neural networks for emotion recognition in video. Proceedings of the 15th ACM International Conference on Multimodal Interaction (ICMI), Sydney, Australia.","DOI":"10.1145\/2522848.2531745"},{"key":"ref_26","doi-asserted-by":"crossref","first-page":"1195","DOI":"10.1109\/TAFFC.2020.2981446","article-title":"Deep facial expression recognition: A survey","volume":"13","author":"Li","year":"2020","journal-title":"IEEE Trans. Affect. Comput."},{"key":"ref_27","first-page":"17616","article-title":"Relative Uncertainty Learning for Facial Expression Recognition","volume":"34","author":"Zhang","year":"2021","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_28","doi-asserted-by":"crossref","unstructured":"Zou, X., Yan, Y., Xue, J.H., Chen, S., and Wang, H. (2022, January 23\u201327). Learn-to-Decompose: Cascaded Decomposition Network for Cross-Domain Few-Shot Facial Expression Recognition. Proceedings of the European Conference on Computer Vision (ECCV), Tel Aviv, Israel.","DOI":"10.1007\/978-3-031-19800-7_40"},{"key":"ref_29","doi-asserted-by":"crossref","first-page":"1057","DOI":"10.1109\/TAFFC.2020.2988264","article-title":"Facial expression recognition with deeply-supervised attention network","volume":"13","author":"Fan","year":"2020","journal-title":"IEEE Trans. Affect. Comput."},{"key":"ref_30","doi-asserted-by":"crossref","first-page":"1868","DOI":"10.1109\/TAFFC.2022.3197761","article-title":"Disentangling Identity and Pose for Facial Expression Recognition","volume":"13","author":"Jiang","year":"2022","journal-title":"IEEE Trans. Affect. Comput."},{"key":"ref_31","first-page":"1320","article-title":"MAFONN-EP: A minimal angular feature oriented neural network based emotion prediction system in image processing","volume":"34","author":"Krithika","year":"2022","journal-title":"J. King Saud-Univ.-Comput. Inf. Sci."},{"key":"ref_32","doi-asserted-by":"crossref","unstructured":"He, K., Zhang, X., Ren, S., and Sun, J. (2016, January 27\u201330). Deep residual learning for image recognition. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.90"},{"key":"ref_33","doi-asserted-by":"crossref","unstructured":"Woo, S., Park, J., Lee, J.Y., and Kweon, I.S. (2018, January 8\u201314). Cbam: Convolutional block attention module. Proceedings of the European Conference on Computer Vision (ECCV), Munich, Germany.","DOI":"10.1007\/978-3-030-01234-2_1"},{"key":"ref_34","doi-asserted-by":"crossref","unstructured":"Wen, Y., Zhang, K., Li, Z., and Qiao, Y. (2016, January 11\u201314). A discriminative feature learning approach for deep face recognition. Proceedings of the European Conference on Computer Vision (ECCV), Amsterdam, The Netherlands.","DOI":"10.1007\/978-3-319-46478-7_31"},{"key":"ref_35","doi-asserted-by":"crossref","unstructured":"Emery, A.E., Muntoni, F., and Quinlivan, R. (2015). Duchenne Muscular Dystrophy, Oxford Monographs on Medical Genetics.","DOI":"10.1093\/med\/9780199681488.001.0001"},{"key":"ref_36","doi-asserted-by":"crossref","first-page":"325","DOI":"10.1109\/TAFFC.2017.2731763","article-title":"Automatic analysis of facial actions: A survey","volume":"10","author":"Martinez","year":"2017","journal-title":"IEEE Trans. Affect. Comput."},{"key":"ref_37","doi-asserted-by":"crossref","unstructured":"Barsoum, E., Zhang, C., Ferrer, C.C., and Zhang, Z. (2016, January 12\u201316). Training deep networks for facial expression recognition with crowd-sourced label distribution. Proceedings of the 18th ACM International Conference on Multimodal Interaction (ICMI), Tokyo, Japan.","DOI":"10.1145\/2993148.2993165"},{"key":"ref_38","doi-asserted-by":"crossref","unstructured":"Guo, Y., Zhang, L., Hu, Y., He, X., and Gao, J. (2016, January 27\u201330). Ms-celeb-1m: A dataset and benchmark for large-scale face recognition. Proceedings of the European Conference on Computer Vision (ECCV), Las Vegas, NV, USA.","DOI":"10.1007\/978-3-319-46487-9_6"},{"key":"ref_39","doi-asserted-by":"crossref","unstructured":"Zeng, J., Shan, S., and Chen, X. (2018, January 8\u201314). Facial expression recognition with inconsistently annotated datasets. Proceedings of the European Conference on Computer Vision (ECCV), Munich, Germany.","DOI":"10.1007\/978-3-030-01261-8_14"},{"key":"ref_40","doi-asserted-by":"crossref","unstructured":"Wang, K., Peng, X., Yang, J., Lu, S., and Qiao, Y. (2020, January 14\u201319). Suppressing uncertainties for large-scale facial expression recognition. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA.","DOI":"10.1109\/CVPR42600.2020.00693"},{"key":"ref_41","doi-asserted-by":"crossref","unstructured":"She, J., Hu, Y., Shi, H., Wang, J., Shen, Q., and Mei, T. (2021, January 19\u201325). Dive into ambiguity: Latent distribution mining and pairwise uncertainty estimation for facial expression recognition. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Virtual.","DOI":"10.1109\/CVPR46437.2021.00618"},{"key":"ref_42","doi-asserted-by":"crossref","unstructured":"Ma, F., Sun, B., and Li, S. (2021). Facial Expression Recognition with Visual Transformers and Attentional Selective Fusion. IEEE Trans. Affect. Comput., 1\u201313.","DOI":"10.1109\/TAFFC.2021.3122146"},{"key":"ref_43","doi-asserted-by":"crossref","unstructured":"Zeng, D., Lin, Z., Yan, X., Liu, Y., Wang, F., and Tang, B. (2022, January 19\u201324). Face2Exp: Combating Data Biases for Facial Expression Recognition. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), New Orleans, LA, USA.","DOI":"10.1109\/CVPR52688.2022.01965"},{"key":"ref_44","doi-asserted-by":"crossref","first-page":"108737","DOI":"10.1016\/j.patcog.2022.108737","article-title":"Improving the Facial Expression Recognition and Its Interpretability via Generating Expression Pattern-map","volume":"129","author":"Zhang","year":"2022","journal-title":"Pattern Recognit."},{"key":"ref_45","doi-asserted-by":"crossref","unstructured":"Parkhi, O.M., Vedaldi, A., and Zisserman, A. (2015, January 7\u201310). Deep face recognition. Proceedings of the British Machine Vision Conference 2015, Swansea, UK.","DOI":"10.5244\/C.29.41"},{"key":"ref_46","doi-asserted-by":"crossref","unstructured":"Arnaud, E., Dapogny, A., and Bailly, K. (2022). Thin: Throwable information networks and application for facial expression recognition in the wild. IEEE Trans. Affect. Comput.","DOI":"10.1109\/TAFFC.2022.3144439"},{"key":"ref_47","doi-asserted-by":"crossref","unstructured":"Fan, X., Deng, Z., Wang, K., Peng, X., and Qiao, Y. (2020, January 25\u201328). Learning discriminative representation for facial expression recognition from uncertainties. Proceedings of the IEEE International Conference on Image Processing (ICIP), Virtual.","DOI":"10.1109\/ICIP40778.2020.9190643"},{"key":"ref_48","doi-asserted-by":"crossref","unstructured":"Farzaneh, A.H., and Qi, X. (2020, January 13\u201319). Discriminant distribution-agnostic loss for facial expression recognition in the wild. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), Seattle, WA, USA.","DOI":"10.1109\/CVPRW50498.2020.00211"},{"key":"ref_49","doi-asserted-by":"crossref","first-page":"3178","DOI":"10.1109\/TCSVT.2021.3103760","article-title":"Learning informative and discriminative features for facial expression recognition in the wild","volume":"32","author":"Li","year":"2021","journal-title":"IEEE Trans. Circuits Syst. Video Technol."},{"key":"ref_50","doi-asserted-by":"crossref","first-page":"690","DOI":"10.1109\/TCSVT.2021.3063052","article-title":"Triplet loss with multistage outlier suppression and class-pair margins for facial expression recognition","volume":"32","author":"Xie","year":"2021","journal-title":"IEEE Trans. Circuits Syst. Video Technol."},{"key":"ref_51","doi-asserted-by":"crossref","unstructured":"Hayale, W., Negi, P.S., and Mahoor, M. (2021). Deep Siamese neural networks for facial expression recognition in the wild. IEEE Trans. Affect. Comput.","DOI":"10.1109\/TAFFC.2021.3077248"},{"key":"ref_52","doi-asserted-by":"crossref","unstructured":"Selvaraju, R.R., Cogswell, M., Das, A., Vedantam, R., Parikh, D., and Batra, D. (2017, January 22\u201329). Grad-cam: Visual explanations from deep networks via gradient-based localization. Proceedings of the IEEE International Conference on Computer Vision (ICCV), Venice, Italy.","DOI":"10.1109\/ICCV.2017.74"}],"container-title":["Sensors"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1424-8220\/23\/7\/3424\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,10]],"date-time":"2025-10-10T19:02:18Z","timestamp":1760122938000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1424-8220\/23\/7\/3424"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,3,24]]},"references-count":52,"journal-issue":{"issue":"7","published-online":{"date-parts":[[2023,4]]}},"alternative-id":["s23073424"],"URL":"https:\/\/doi.org\/10.3390\/s23073424","relation":{},"ISSN":["1424-8220"],"issn-type":[{"value":"1424-8220","type":"electronic"}],"subject":[],"published":{"date-parts":[[2023,3,24]]}}}