{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,17]],"date-time":"2026-07-17T21:58:04Z","timestamp":1784325484594,"version":"3.55.0"},"reference-count":41,"publisher":"MDPI AG","issue":"10","license":[{"start":{"date-parts":[[2022,5,13]],"date-time":"2022-05-13T00:00:00Z","timestamp":1652400000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"name":"Bisa Research Grant of Keimyung University","award":["20210735"],"award-info":[{"award-number":["20210735"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Sensors"],"abstract":"<jats:p>In recent image classification approaches, a vision transformer (ViT) has shown an excellent performance beyond that of a convolutional neural network. A ViT achieves a high classification for natural images because it properly preserves the global image features. Conversely, a ViT still has many limitations in facial expression recognition (FER), which requires the detection of subtle changes in expression, because it can lose the local features of the image. Therefore, in this paper, we propose Squeeze ViT, a method for reducing the computational complexity by reducing the number of feature dimensions while increasing the FER performance by concurrently combining global and local features. To measure the FER performance of Squeeze ViT, experiments were conducted on lab-controlled FER datasets and a wild FER dataset. Through comparative experiments with previous state-of-the-art approaches, we proved that the proposed method achieves an excellent performance on both types of datasets.<\/jats:p>","DOI":"10.3390\/s22103729","type":"journal-article","created":{"date-parts":[[2022,5,15]],"date-time":"2022-05-15T09:48:22Z","timestamp":1652608102000},"page":"3729","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":55,"title":["Facial Expression Recognition Based on Squeeze Vision Transformer"],"prefix":"10.3390","volume":"22","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-7452-3897","authenticated-orcid":false,"given":"Sangwon","family":"Kim","sequence":"first","affiliation":[{"name":"Department of Computer Engineering, Keimyung University, Daegu 42601, Korea"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Jaeyeal","family":"Nam","sequence":"additional","affiliation":[{"name":"Department of Computer Engineering, Keimyung University, Daegu 42601, Korea"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-7284-0768","authenticated-orcid":false,"given":"Byoung Chul","family":"Ko","sequence":"additional","affiliation":[{"name":"Department of Computer Engineering, Keimyung University, Daegu 42601, Korea"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1968","published-online":{"date-parts":[[2022,5,13]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","first-page":"7071","DOI":"10.1038\/s41598-021-86345-5","article-title":"Emotion detection using electroencephalography signals and a zero-time windowing-based epoch estimation and relevant electrode identification","volume":"11","author":"Gannouni","year":"2021","journal-title":"Sci. Rep."},{"key":"ref_2","doi-asserted-by":"crossref","unstructured":"Hasnul, M.A., Aziz, N.A.A., Alelyani, S., Mohana, M., and Aziz, A.A. (2021). Electrocardiogram-Based Emotion Recognition Systems and Their Applications in Healthcare\u2014A Review. Sensors, 21.","DOI":"10.3390\/s21155015"},{"key":"ref_3","doi-asserted-by":"crossref","first-page":"329","DOI":"10.3389\/fpsyg.2020.00329","article-title":"A Comparison of the Affectiva iMotions Facial Expression Analysis Software with EMG for Identifying Facial Expressions of Emotion","volume":"11","author":"Kulke","year":"2020","journal-title":"Front. Psychol."},{"key":"ref_4","doi-asserted-by":"crossref","first-page":"1454","DOI":"10.1073\/pnas.1322355111","article-title":"Compound facial expressions of emotion","volume":"111","author":"Du","year":"2014","journal-title":"Proc. Natl. Acad. Sci. USA"},{"key":"ref_5","doi-asserted-by":"crossref","first-page":"482","DOI":"10.1109\/TIFS.2020.3007327","article-title":"Fine-Grained Facial Expression Recognition in the Wild","volume":"16","author":"Liang","year":"2020","journal-title":"IEEE Trans. Inf. Forensics Secur."},{"key":"ref_6","doi-asserted-by":"crossref","first-page":"7803","DOI":"10.1007\/s11042-016-3418-y","article-title":"Facial expression recognition based on local region specific features and support vector machines","volume":"76","author":"Ghimire","year":"2017","journal-title":"Multimed. Tools Appl."},{"key":"ref_7","doi-asserted-by":"crossref","first-page":"7714","DOI":"10.3390\/s130607714","article-title":"Geometric feature-based facial expression recognition in image sequences using multi-class AdaBoost and support vector machines","volume":"13","author":"Ghimire","year":"2013","journal-title":"Sensors"},{"key":"ref_8","doi-asserted-by":"crossref","unstructured":"Jeong, M., and Ko, B.C. (2018). Driver\u2019s Facial Expression Recognition in Real-Time for Safe Driving. Sensors, 18.","DOI":"10.3390\/s18124270"},{"key":"ref_9","unstructured":"Zhao, S., Cai, H., Liu, H., Zhang, J., and Chen, S. (2018, January 3\u20136). Feature Selection Mechanism in CNNs for Facial Expression Recognition. Proceedings of the British Machine Vision Conference (BMVC), Newcastle, UK."},{"key":"ref_10","unstructured":"Fan, Y., Li, V., and Lam, J.C. (2020). Facial Expression Recognition with Deeply-Supervised Attention Network. IEEE Trans. Affect. Comput."},{"key":"ref_11","doi-asserted-by":"crossref","unstructured":"Xu, T., White, J., Kalkan, S., and Gunes, H. (2020, January 23\u201328). Investigating bias and fairness in facial expression recognition. Proceedings of the European Conference on Computer Vision (ECCV), Virtual.","DOI":"10.1007\/978-3-030-65414-6_35"},{"key":"ref_12","doi-asserted-by":"crossref","unstructured":"Wang, K., Peng, X., Yang, J., Lu, S., and Qiao, Y. (2020, January 14\u201319). Suppressing uncertainties for large-scale facial expression recognition. Proceedings of the IEEE International Conference on Computer Vision and Pattern Recognition (CVPR), Virtual.","DOI":"10.1109\/CVPR42600.2020.00693"},{"key":"ref_13","doi-asserted-by":"crossref","unstructured":"Minaee, S., Minaei, M., and Abdolrashidi, A. (2021). Deep-Emotion: Facial Expression Recognition Using Attentional Convolutional Network. Sensors, 21.","DOI":"10.3390\/s21093046"},{"key":"ref_14","doi-asserted-by":"crossref","unstructured":"Ko, B.C. (2018). A Brief Review of Facial Emotion Recognition Based on Visual Information. Sensors, 18.","DOI":"10.3390\/s18020401"},{"key":"ref_15","doi-asserted-by":"crossref","unstructured":"Greche, L., and Es-Sbai, N. (April, January 30). Automatic System for Facial Expression Recognition Based Histogram of Oriented Gradient and Normalized Cross Correlation. Proceedings of the 2016 International Conference on Information Technology for Organizations Development (IT4OD), Fez, Morocco.","DOI":"10.1109\/IT4OD.2016.7479316"},{"key":"ref_16","doi-asserted-by":"crossref","first-page":"340","DOI":"10.1109\/TAFFC.2014.2346515","article-title":"Intra-class variation reduction using training expression images for sparse representation based facial expression recognition","volume":"5","author":"Lee","year":"2014","journal-title":"IEEE Trans. Affect. Comput."},{"key":"ref_17","unstructured":"Krizhevsky, A., Sutskever, I., and Hinton, G.E. (2012, January 3\u20138). ImageNet Classification with Deep Convolutional Neural Networks. Proceedings of the Annual Conference on Neural Information Processing Systems (NeurIPS), Lake Tahoe, NV, USA."},{"key":"ref_18","unstructured":"Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., and Gelly, S. (2021, January 3\u20137). An Image is Worth 16 \u00d7 16 Words: Transformers for Image Recognition at Scale. Proceedings of the Ninth International Conference on Learning Representations (ICLR), Virtual."},{"key":"ref_19","doi-asserted-by":"crossref","unstructured":"Dai, Z., Cai, B., Lin, Y., and Chen, J. (2021, January 19\u201328). UP-DETR: Unsupervised Pre-Training for Object Detection With Transformers. Proceedings of the IEEE International Conference on Computer Vision and Pattern Recognition (CVPR), Virtual.","DOI":"10.1109\/CVPR46437.2021.00165"},{"key":"ref_20","doi-asserted-by":"crossref","unstructured":"Chen, X., Yan, B., Zhu, J., Wang, D., Yang, X., and Lu, H. (2021, January 19\u201328). Transformer Tracking. Proceedings of the IEEE International Conference on Computer Vision and Pattern Recognition (CVPR), Virtual.","DOI":"10.1109\/CVPR46437.2021.00803"},{"key":"ref_21","doi-asserted-by":"crossref","unstructured":"Esser, P., Rombach, R., and Ommer, B. (2021, January 19\u201328). Taming Transformers for High-Resolution Image Synthesis. Proceedings of the IEEE International Conference on Computer Vision and Pattern Recognition (CVPR), Virtual.","DOI":"10.1109\/CVPR46437.2021.01268"},{"key":"ref_22","doi-asserted-by":"crossref","unstructured":"Xue, F., Wang, Q., and Guo, G. (2021, January 11\u201317). TransFER: Learning Relation-aware Facial Expression Representations with Transformers. Proceedings of the IEEE International Conference on Computer Vision (ICCV), Virtual.","DOI":"10.1109\/ICCV48922.2021.00358"},{"key":"ref_23","unstructured":"Aouayeb, M., Hamidouche, W., Soladie, C., Kpalma, K., and Seguier, R. (2021). Learning Vision Transformer with Squeeze and Excitation for Facial Expression Recognition. arXiv."},{"key":"ref_24","doi-asserted-by":"crossref","first-page":"824592","DOI":"10.3389\/fnbot.2021.824592","article-title":"Progressive Multi-Scale Vision Transformer for Facial Action Unit Detection","volume":"12","author":"Wang","year":"2022","journal-title":"Front. Neurorobot."},{"key":"ref_25","doi-asserted-by":"crossref","unstructured":"Ma, F., Sun, B., and Li, S. (2022). Facial Expression Recognition with Visual Transformers and Attentional Selective Fusion. IEEE Trans. Affect. Comput.","DOI":"10.1109\/TAFFC.2021.3122146"},{"key":"ref_26","doi-asserted-by":"crossref","unstructured":"Bulat, A., and Tzimiropoulos, G. (2017, January 22\u201329). How far are we from solving the 2d & 3d face alignment problem?(and a dataset of 230,000 3d facial landmarks). Proceedings of the IEEE International Conference on Computer Vision (ICCV), Venice, Italy.","DOI":"10.1109\/ICCV.2017.116"},{"key":"ref_27","doi-asserted-by":"crossref","unstructured":"Woo, S., Park, J., Lee, J.Y., and Kweon, I.S. (2018, January 8\u201314). Cbam: Convolutional block attention module. Proceedings of the European Conference on Computer Vision (ECCV), Munich, Germany.","DOI":"10.1007\/978-3-030-01234-2_1"},{"key":"ref_28","unstructured":"Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A.N., Kaiser, \u0141., and Polosukhin, I. (2017, January 4\u20139). Attention is all you need. Proceedings of the Annual Conference on Neural Information Processing Systems (NeurIPS), Long Beach, CA, USA."},{"key":"ref_29","unstructured":"Touvron, H., Cord, M., Douze, M., Massa, F., Sablayrolles, A., and J\u00e9gou, H. (2021, January 18\u201324). Training data-efficient image transformers & distillation through attention. Proceedings of the International Conference on Machine Learning (ICML), Virtual."},{"key":"ref_30","doi-asserted-by":"crossref","unstructured":"Liu, Z., Lin, Y., Cao, Y., Hu, H., Wei, Y., Zhang, Z., Lin, S., and Guo, B. (2021, January 19\u201328). Swin transformer: Hierarchical vision transformer using shifted windows. Proceedings of the IEEE International Conference on Computer Vision and Pattern Recognition (CVPR), Virtual.","DOI":"10.1109\/ICCV48922.2021.00986"},{"key":"ref_31","doi-asserted-by":"crossref","unstructured":"Lucey, P., Cohn, J.F., Kanade, T., Saragih, J., Ambadar, Z., and Matthews, I. (2010, January 13\u201318). The extended cohn-kanade dataset (ck+): A complete dataset for action unit and emotion-specified expression. Proceedings of the IEEE International Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), San Francisco, CA, USA.","DOI":"10.1109\/CVPRW.2010.5543262"},{"key":"ref_32","unstructured":"Valstar, M., and Pantic, M. (2010, January 17\u201323). Induced disgust, happiness and surprise: An addition to the mmi facial expression database. Proceedings of the 3rd International Workshop on EMOTION (Satellite of LREC): Corpora for Research on Emotion and Affect, Valletta, Malta."},{"key":"ref_33","doi-asserted-by":"crossref","unstructured":"Li, S., Deng, W., and Du, J. (2017, January 21\u201326). Reliable Crowdsourcing and Deep Locality-Preserving Learning for Expression Recognition in the Wild. Proceedings of the IEEE International Conference on Computer Vision and Pattern Recognition (CVPR), Hawaii, HO, USA.","DOI":"10.1109\/CVPR.2017.277"},{"key":"ref_34","doi-asserted-by":"crossref","unstructured":"Zeng, J., Shan, S., and Chen, X. (2018, January 8\u201314). Facial expression recognition with inconsistently annotated datasets. Proceedings of the European Conference on Computer Vision (ECCV), Munich, Germany.","DOI":"10.1007\/978-3-030-01261-8_14"},{"key":"ref_35","doi-asserted-by":"crossref","unstructured":"Chen, Y., Wang, J., Chen, S., Shi, Z., and Cai, J. (2019, January 1\u20134). Facial Motion Prior Networks for Facial Expression Recognition. Proceedings of the IEEE Visual Communications and Image Processing (VCIP), Sydney, Australia.","DOI":"10.1109\/VCIP47243.2019.8965826"},{"key":"ref_36","doi-asserted-by":"crossref","unstructured":"Chen, S., Wang, J., Chen, Y., Shi, Z., Geng, X., and Rui, Y. (2020, January 14\u201319). Label distribution learning on auxiliary label space graphs for facial expression recognition. Proceedings of the IEEE International Conference on Computer Vision and Pattern Recognition (CVPR), Virtual.","DOI":"10.1109\/CVPR42600.2020.01400"},{"key":"ref_37","doi-asserted-by":"crossref","unstructured":"Ruan, D., Yan, Y., Lai, S., Chai, Z., Shen, C., and Wang, H. (2021, January 19\u201328). Feature Decomposition and Reconstruction Learning for Effective Facial Expression Recognition. Proceedings of the IEEE International Conference on Computer Vision and Pattern Recognition (CVPR), Virtual.","DOI":"10.1109\/CVPR46437.2021.00757"},{"key":"ref_38","unstructured":"Mao, S., Shi, G., Gou, S., Yan, D., Jiao, L., and Xiong, L. (2022). Adaptively Lighting up Facial Expression Crucial Regions via Local Non-Local Joint Network. arXiv."},{"key":"ref_39","doi-asserted-by":"crossref","unstructured":"She, J., Hu, Y., Shi, H., Wang, J., Shen, Q., and Mei, T. (2021, January 19\u201328). Dive into Ambiguity: Latent Distribution Mining and Pairwise Uncertainty Estimation for Facial Expression Recognition. Proceedings of the IEEE International Conference on Computer Vision and Pattern Recognition (CVPR), Virtual.","DOI":"10.1109\/CVPR46437.2021.00618"},{"key":"ref_40","unstructured":"Zhang, Y., Wang, C., and Deng, W. (2021, January 6\u201314). Relative Uncertainty Learning for Facial Expression Recognition. Proceedings of the Annual Conference on Neural Information Processing Systems (NeurIPS), Virtual."},{"key":"ref_41","doi-asserted-by":"crossref","unstructured":"Guo, Y., Zhang, L., Hu, Y., He, X., and Gao, J. (2016, January 8\u201316). Ms-celeb-1m: A dataset and benchmark for large-scale face recognition. Proceedings of the European Conference on Computer Vision (ECCV), Amsterdam, The Netherlands.","DOI":"10.1007\/978-3-319-46487-9_6"}],"container-title":["Sensors"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1424-8220\/22\/10\/3729\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,10]],"date-time":"2025-10-10T23:10:25Z","timestamp":1760137825000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1424-8220\/22\/10\/3729"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,5,13]]},"references-count":41,"journal-issue":{"issue":"10","published-online":{"date-parts":[[2022,5]]}},"alternative-id":["s22103729"],"URL":"https:\/\/doi.org\/10.3390\/s22103729","relation":{},"ISSN":["1424-8220"],"issn-type":[{"value":"1424-8220","type":"electronic"}],"subject":[],"published":{"date-parts":[[2022,5,13]]}}}