{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,6]],"date-time":"2026-07-06T16:36:24Z","timestamp":1783355784490,"version":"3.54.6"},"reference-count":52,"publisher":"MDPI AG","issue":"13","license":[{"start":{"date-parts":[[2024,6,26]],"date-time":"2024-06-26T00:00:00Z","timestamp":1719360000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"name":"Jilin Provincial Science and Technology Development Program","award":["20230401104YY"],"award-info":[{"award-number":["20230401104YY"]}]},{"name":"Jilin Provincial Science and Technology Development Program","award":["20210201083GX"],"award-info":[{"award-number":["20210201083GX"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Sensors"],"abstract":"<jats:p>Convolutional neural networks (CNNs) have made significant progress in the field of facial expression recognition (FER). However, due to challenges such as occlusion, lighting variations, and changes in head pose, facial expression recognition in real-world environments remains highly challenging. At the same time, methods solely based on CNN heavily rely on local spatial features, lack global information, and struggle to balance the relationship between computational complexity and recognition accuracy. Consequently, the CNN-based models still fall short in their ability to address FER adequately. To address these issues, we propose a lightweight facial expression recognition method based on a hybrid vision transformer. This method captures multi-scale facial features through an improved attention module, achieving richer feature integration, enhancing the network\u2019s perception of key facial expression regions, and improving feature extraction capabilities. Additionally, to further enhance the model\u2019s performance, we have designed the patch dropping (PD) module. This module aims to emulate the attention allocation mechanism of the human visual system for local features, guiding the network to focus on the most discriminative features, reducing the influence of irrelevant features, and intuitively lowering computational costs. Extensive experiments demonstrate that our approach significantly outperforms other methods, achieving an accuracy of 86.51% on RAF-DB and nearly 70% on FER2013, with a model size of only 3.64 MB. These results demonstrate that our method provides a new perspective for the field of facial expression recognition.<\/jats:p>","DOI":"10.3390\/s24134153","type":"journal-article","created":{"date-parts":[[2024,6,26]],"date-time":"2024-06-26T09:29:33Z","timestamp":1719394173000},"page":"4153","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":13,"title":["Enhanced Hybrid Vision Transformer with Multi-Scale Feature Integration and Patch Dropping for Facial Expression Recognition"],"prefix":"10.3390","volume":"24","author":[{"ORCID":"https:\/\/orcid.org\/0000-0003-2450-5217","authenticated-orcid":false,"given":"Nianfeng","family":"Li","sequence":"first","affiliation":[{"name":"College of Computer Science and Technology, Changchun University, No. 6543, Satellite Road, Changchun 130022, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Yongyuan","family":"Huang","sequence":"additional","affiliation":[{"name":"College of Computer Science and Technology, Changchun University, No. 6543, Satellite Road, Changchun 130022, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Zhenyan","family":"Wang","sequence":"additional","affiliation":[{"name":"College of Computer Science and Technology, Changchun University, No. 6543, Satellite Road, Changchun 130022, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Ziyao","family":"Fan","sequence":"additional","affiliation":[{"name":"College of Computer Science and Technology, Changchun University, No. 6543, Satellite Road, Changchun 130022, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Xinyuan","family":"Li","sequence":"additional","affiliation":[{"name":"College of Computer Science and Technology, Changchun University, No. 6543, Satellite Road, Changchun 130022, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Zhiguo","family":"Xiao","sequence":"additional","affiliation":[{"name":"College of Computer Science and Technology, Changchun University, No. 6543, Satellite Road, Changchun 130022, China"},{"name":"School of Computer Science Technology, Beijing Institute of Technology, Beijing 100811, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1968","published-online":{"date-parts":[[2024,6,26]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","unstructured":"Alharbi, M., and Huang, S. (2020, January 28\u201330). A Survey of Incorporating Affective Computing for Human-System co-Adaptation. Proceedings of the 2nd World Symposium on Software Engineering, Xiamen, China.","DOI":"10.1145\/3425329.3425343"},{"key":"ref_2","doi-asserted-by":"crossref","first-page":"1195","DOI":"10.1109\/TAFFC.2020.2981446","article-title":"Deep facial expression recognition: A survey","volume":"13","author":"Li","year":"2020","journal-title":"IEEE Trans. Affect. Comput."},{"key":"ref_3","doi-asserted-by":"crossref","first-page":"177","DOI":"10.1016\/j.patcog.2019.03.019","article-title":"Deep multi-path convolutional neural network joint with salient region attention for facial expression recognition","volume":"92","author":"Xie","year":"2019","journal-title":"Pattern Recognit."},{"key":"ref_4","doi-asserted-by":"crossref","first-page":"176","DOI":"10.1049\/iet-ipr.2019.0293","article-title":"Fusing HOG and convolutional neural network spatial\u2013temporal features for video-based facial expression recognition","volume":"14","author":"Pan","year":"2020","journal-title":"IET Image Process."},{"key":"ref_5","doi-asserted-by":"crossref","first-page":"4057","DOI":"10.1109\/TIP.2019.2956143","article-title":"Region attention networks for pose and occlusion robust facial expression recognition","volume":"29","author":"Wang","year":"2020","journal-title":"IEEE Trans. Image Process."},{"key":"ref_6","doi-asserted-by":"crossref","first-page":"2439","DOI":"10.1109\/TIP.2018.2886767","article-title":"Occlusion aware facial expression recognition using CNN with attention mechanism","volume":"28","author":"Li","year":"2018","journal-title":"IEEE Trans. Image Process."},{"key":"ref_7","doi-asserted-by":"crossref","first-page":"451","DOI":"10.1109\/TAFFC.2020.3031602","article-title":"Facial expression recognition in the wild using multi-level features and attention mechanisms","volume":"14","author":"Li","year":"2020","journal-title":"IEEE Trans. Affect. Comput."},{"key":"ref_8","doi-asserted-by":"crossref","unstructured":"Li, Y., Zeng, J., Shan, S., and Chen, X. (2018, January 20\u201324). Patch-Gated CNN for Occlusion-Aware Facial Expression Recognition. Proceedings of the 2018 24th International Conference on Pattern Recognition (ICPR), Beijing, China.","DOI":"10.1109\/ICPR.2018.8545853"},{"key":"ref_9","doi-asserted-by":"crossref","first-page":"1057","DOI":"10.1109\/TAFFC.2020.2988264","article-title":"Facial expression recognition with deeply-supervised attention network","volume":"13","author":"Fan","year":"2020","journal-title":"IEEE Trans. Affect. Comput."},{"key":"ref_10","doi-asserted-by":"crossref","first-page":"340","DOI":"10.1016\/j.neucom.2020.06.014","article-title":"Attention mechanism-based CNN for facial expression recognition","volume":"411","author":"Li","year":"2020","journal-title":"Neurocomputing"},{"key":"ref_11","unstructured":"Mehta, S., and Rastegari, M. (2021). Mobilevit: Light-weight, general-purpose, and mobile-friendly vision transformer. arXiv."},{"key":"ref_12","doi-asserted-by":"crossref","first-page":"467","DOI":"10.1109\/TIP.2002.999679","article-title":"Gabor feature based classification using the enhanced fisher linear discriminant model for face recognition","volume":"11","author":"Liu","year":"2002","journal-title":"IEEE Trans. Image Process."},{"key":"ref_13","doi-asserted-by":"crossref","first-page":"803","DOI":"10.1016\/j.imavis.2008.08.005","article-title":"Facial expression recognition based on local binary patterns: A comprehensive study","volume":"27","author":"Shan","year":"2009","journal-title":"Image Vis. Comput."},{"key":"ref_14","unstructured":"Dalal, N., and Triggs, B. (2005, January 20\u201325). Histograms of Oriented Gradients for Human Detection. Proceedings of the 2005 IEEE Computer Society Conference on Computer Vision and Pattern Recognition (CVPR\u201905), San Diego, CA, USA."},{"key":"ref_15","doi-asserted-by":"crossref","unstructured":"Cai, J., Meng, Z., Khan, A.S., Li, Z., O\u2019Reilly, J., and Tong, Y. (2018, January 15\u201319). Island Loss for Learning Discriminative Features in Facial Expression Recognition. Proceedings of the 2018 13th IEEE International Conference on Automatic Face & Gesture Recognition (FG 2018), Xi\u2019an, China.","DOI":"10.1109\/FG.2018.00051"},{"key":"ref_16","doi-asserted-by":"crossref","unstructured":"Pan, B., Wang, S., and Xia, B. (2019, January 21\u201325). Occluded Facial Expression Recognition Enhanced through Privileged Information. Proceedings of the 27th ACM International Conference on Multimedia, Nice, France.","DOI":"10.1145\/3343031.3351049"},{"key":"ref_17","doi-asserted-by":"crossref","first-page":"881","DOI":"10.1109\/TAFFC.2020.2973158","article-title":"A deeper look at facial expression dataset bias","volume":"13","author":"Li","year":"2020","journal-title":"IEEE Trans. Affect. Comput."},{"key":"ref_18","doi-asserted-by":"crossref","first-page":"1483","DOI":"10.1007\/s11277-022-09616-y","article-title":"Facial expression recognition based on spatial and channel attention mechanisms","volume":"125","author":"Yao","year":"2022","journal-title":"Wirel. Pers. Commun."},{"key":"ref_19","unstructured":"Iandola, F.N., Han, S., Moskewicz, M.W., Ashraf, K., Dally, W.J., and Keutzer, K. (2016). SqueezeNet: AlexNet-level accuracy with 50x fewer parameters and <0.5 MB model size. arXiv."},{"key":"ref_20","unstructured":"Howard, A.G., Zhu, M., Chen, B., Kalenichenko, D., Wang, W., Weyand, T., Andreetto, M., and Adam, H. (2017). Mobilenets: Efficient convolutional neural networks for mobile vision applications. arXiv."},{"key":"ref_21","doi-asserted-by":"crossref","unstructured":"Sandler, M., Howard, A., Zhu, M., Zhmoginov, A., and Chen, L.C. (2018, January 19\u201323). Mobilenetv2: Inverted Residuals and Linear Bottlenecks. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00474"},{"key":"ref_22","unstructured":"Howard, A., Sandler, M., Chu, G., Chen, L.C., Chen, B., Tan, M., Wang, W., Zhu, Y., Pang, R., and Vasudevan, V. (November, January 27). Searching for Mobilenetv3. Proceedings of the IEEE\/CVF International Conference on Computer Vision, Seoul, Republic of Korea."},{"key":"ref_23","doi-asserted-by":"crossref","unstructured":"Zhang, X., Zhou, X., Lin, M., and Sun, J. (2018, January 18\u201323). Shufflenet: An Extremely Efficient Convolutional Neural Network for Mobile Devices. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00716"},{"key":"ref_24","doi-asserted-by":"crossref","unstructured":"Ma, N., Zhang, X., Zheng, H.T., and Sun, J. (2018, January 8\u201314). Shufflenet v2: Practical Guidelines for Efficient CNN Architecture Design. Proceedings of the European Conference on computer Vision (ECCV) 2018, Munich, Germany.","DOI":"10.1007\/978-3-030-01264-9_8"},{"key":"ref_25","doi-asserted-by":"crossref","first-page":"4435","DOI":"10.1016\/j.aej.2021.09.066","article-title":"A-MobileNet: An approach of facial expression recognition","volume":"61","author":"Nan","year":"2022","journal-title":"Alex. Eng. J."},{"key":"ref_26","doi-asserted-by":"crossref","unstructured":"Han, B., Hu, M., Wang, X., and Ren, F. (2022). A triple-structure network model based upon MobileNet V1 and multi-loss function for facial expression recognition. Symmetry, 14.","DOI":"10.3390\/sym14102055"},{"key":"ref_27","doi-asserted-by":"crossref","unstructured":"Zhang, L.Q., Liu, Z.T., and Jiang, C.S. (2022, January 25\u201327). An Improved SimAM Based CNN for Facial Expression Recognition. Proceedings of the 2022 41st Chinese Control Conference (CCC), Hefei, China.","DOI":"10.23919\/CCC55666.2022.9902045"},{"key":"ref_28","doi-asserted-by":"crossref","first-page":"103018","DOI":"10.1016\/j.jvcir.2020.103018","article-title":"Facial expression recognition using frequency multiplication network with uniform rectangular features","volume":"75","author":"Zhou","year":"2021","journal-title":"J. Vis. Commun. Image Represent."},{"key":"ref_29","doi-asserted-by":"crossref","unstructured":"Cotter, S.F. (2020, January 4\u20136). MobiExpressNet: A Deep Learning Network for Face Expression Recognition on Smart Phones. Proceedings of the 2020 IEEE International Conference on Consumer Electronics (ICCE), Las Vegas, NV, USA.","DOI":"10.1109\/ICCE46568.2020.9042973"},{"key":"ref_30","unstructured":"Goodfellow, I.J., Erhan, D., Carrier, P.L., Courville, A., Mirza, M., Hamner, B., Cukierski, W., Tang, Y., Thaler, D., and Lee, D.H. (2013, January 3\u20137). Challenges in Representation Learning: A Report on Three Machine Learning Contests. Proceedings of the Neural Information Processing: 20th International Conference, ICONIP 2013, Daegu, Republic of Korea. Proceedings, Part III 20."},{"key":"ref_31","doi-asserted-by":"crossref","unstructured":"Li, S., Deng, W., and Du, J. (2017, January 21\u201326). Reliable Crowdsourcing and Deep Locality-Preserving Learning for Expression Recognition in the Wild. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.277"},{"key":"ref_32","doi-asserted-by":"crossref","unstructured":"He, K., Zhang, X., Ren, S., and Sun, J. (2016, January 27\u201330). Deep Residual Learning for Image Recognition. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.90"},{"key":"ref_33","unstructured":"Simonyan, K., and Zisserman, A. (2014). Very deep convolutional networks for large-scale image recognition. arXiv."},{"key":"ref_34","doi-asserted-by":"crossref","first-page":"84","DOI":"10.1145\/3065386","article-title":"ImageNet classification with deep convolutional neural networks","volume":"60","author":"Krizhevsky","year":"2017","journal-title":"Commun. ACM"},{"key":"ref_35","doi-asserted-by":"crossref","unstructured":"Ghosh, S., Dhall, A., and Sebe, N. (2018, January 7\u201310). Automatic Group Affect Analysis in Images via Visual Attribute and Feature Networks. Proceedings of the 2018 25th IEEE International Conference on Image Processing (ICIP), Athens, Greece.","DOI":"10.1109\/ICIP.2018.8451242"},{"key":"ref_36","doi-asserted-by":"crossref","unstructured":"Hua, C.H., Huynh-The, T., Seo, H., and Lee, S. (2020, January 3\u20135). Convolutional Network with Densely Backward Attention for Facial Expression recognition. Proceedings of the 2020 14th International Conference on Ubiquitous Information Management and Communication (IMCOM), Taichung, Taiwan.","DOI":"10.1109\/IMCOM48794.2020.9001686"},{"key":"ref_37","first-page":"356","article-title":"Reliable crowdsourcing and deep locality-preserving learning for unconstrained facial expression recognition","volume":"28","author":"Shan","year":"2018","journal-title":"IEEE Trans. Image Process."},{"key":"ref_38","doi-asserted-by":"crossref","first-page":"725","DOI":"10.1109\/LSP.2020.2989670","article-title":"Accurate and reliable facial expression recognition using advanced softmax loss with fixed weights","volume":"27","author":"Jiang","year":"2020","journal-title":"IEEE Signal Process. Lett."},{"key":"ref_39","unstructured":"Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., and Gelly, S. (2020). An image is worth 16x16 words: Transformers for image recognition at scale. arXiv."},{"key":"ref_40","doi-asserted-by":"crossref","unstructured":"Bousaid, R., El Hajji, M., and Es-Saady, Y. (2022, January 12\u201314). Facial Emotions Recognition Using Vit and Transfer Learning. Proceedings of the 2022 5th International Conference on Advanced Communication Technologies and Networking (CommNet), Marrakech, Morocco.","DOI":"10.1109\/CommNet56067.2022.9993933"},{"key":"ref_41","doi-asserted-by":"crossref","first-page":"1236","DOI":"10.1109\/TAFFC.2021.3122146","article-title":"Facial expression recognition with visual transformers and attentional selective fusion","volume":"14","author":"Ma","year":"2021","journal-title":"IEEE Trans. Affect. Comput."},{"key":"ref_42","doi-asserted-by":"crossref","first-page":"35","DOI":"10.1016\/j.ins.2021.08.043","article-title":"Facial expression recognition with grid-wise attention and visual transformer","volume":"580","author":"Huang","year":"2021","journal-title":"Inf. Sci."},{"key":"ref_43","doi-asserted-by":"crossref","first-page":"109554","DOI":"10.1016\/j.foodcont.2022.109554","article-title":"Grading and fraud detection of saffron via learning-to-augment incorporated Inception-v4 CNN","volume":"147","author":"Momeny","year":"2023","journal-title":"Food Control"},{"key":"ref_44","doi-asserted-by":"crossref","first-page":"549","DOI":"10.1007\/s10489-020-01855-5","article-title":"E-FCNN for tiny facial expression recognition","volume":"51","author":"Shao","year":"2021","journal-title":"Appl. Intell."},{"key":"ref_45","doi-asserted-by":"crossref","unstructured":"Mollahosseini, A., Chan, D., and Mahoor, M.H. (2016, January 7\u201310). Going Deeper in Facial Expression Recognition Using Deep Neural Networks. Proceedings of the 2016 IEEE Winter Conference on Applications of Computer Vision (WACV), Lake Placid, NY, USA.","DOI":"10.1109\/WACV.2016.7477450"},{"key":"ref_46","unstructured":"Chen, C.F., Panda, R., and Fan, Q. (2021). Regionvit: Regional-to-local attention for vision transformers. arXiv."},{"key":"ref_47","doi-asserted-by":"crossref","unstructured":"Yuan, L., Chen, Y., Wang, T., Yu, W., Shi, Y., Jiang, Z.H., Tay, F.E., Feng, J., and Yan, S. (2021, January 10\u201317). Tokens-to-Token vit: Training Vision Transformers from Scratch on Imagenet. Proceedings of the IEEE\/CVF International Conference on Computer Vision, Montreal, QC, Canada.","DOI":"10.1109\/ICCV48922.2021.00060"},{"key":"ref_48","unstructured":"Zhou, D., Kang, B., Jin, X., Yang, L., Lian, X., Jiang, Z., Hou, Q., and Feng, J. (2021). Deepvit: Towards deeper vision transformer. arXiv."},{"key":"ref_49","doi-asserted-by":"crossref","unstructured":"Chen, C.F.R., Fan, Q., and Panda, R. (2021, January 10\u201317). Crossvit: Cross-Attention Multi-Scale Vision Transformer for Image Classification. Proceedings of the IEEE\/CVF International Conference on Computer Vision, Montreal, QC, Canada.","DOI":"10.1109\/ICCV48922.2021.00041"},{"key":"ref_50","unstructured":"Han, Q., Fan, Z., Dai, Q., Sun, L., Cheng, M.M., Liu, J., and Wang, J. (2021). On the connection between local attention and dynamic depth-wise convolution. arXiv."},{"key":"ref_51","doi-asserted-by":"crossref","unstructured":"Liu, Z., Lin, Y., Cao, Y., Hu, H., Wei, Y., Zhang, Z., Lin, S., and Guo, B. (2021, January 10\u201317). Swin Transformer: Hierarchical Vision Transformer Using Shifted Windows. Proceedings of the IEEE\/CVF International Conference on Computer Vision, Montreal, QC, Canada.","DOI":"10.1109\/ICCV48922.2021.00986"},{"key":"ref_52","unstructured":"Zhou, J., Wang, P., Wang, F., Liu, Q., Li, H., and Jin, R. (2021). Elsa: Enhanced local self-attention for vision transformer. arXiv."}],"container-title":["Sensors"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1424-8220\/24\/13\/4153\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,10]],"date-time":"2025-10-10T15:04:58Z","timestamp":1760108698000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1424-8220\/24\/13\/4153"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,6,26]]},"references-count":52,"journal-issue":{"issue":"13","published-online":{"date-parts":[[2024,7]]}},"alternative-id":["s24134153"],"URL":"https:\/\/doi.org\/10.3390\/s24134153","relation":{},"ISSN":["1424-8220"],"issn-type":[{"value":"1424-8220","type":"electronic"}],"subject":[],"published":{"date-parts":[[2024,6,26]]}}}