{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,13]],"date-time":"2026-06-13T21:31:10Z","timestamp":1781386270097,"version":"3.54.1"},"reference-count":36,"publisher":"MDPI AG","issue":"12","license":[{"start":{"date-parts":[[2023,6,16]],"date-time":"2023-06-16T00:00:00Z","timestamp":1686873600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Sensors"],"abstract":"<jats:p>In the billions of faces that are shaped by thousands of different cultures and ethnicities, one thing remains universal: the way emotions are expressed. To take the next step in human\u2013machine interactions, a machine (e.g., a humanoid robot) must be able to clarify facial emotions. Allowing systems to recognize micro-expressions affords the machine a deeper dive into a person\u2019s true feelings, which will take human emotion into account while making optimal decisions. For instance, these machines will be able to detect dangerous situations, alert caregivers to challenges, and provide appropriate responses. Micro-expressions are involuntary and transient facial expressions capable of revealing genuine emotions. We propose a new hybrid neural network (NN) model capable of micro-expression recognition in real-time applications. Several NN models are first compared in this study. Then, a hybrid NN model is created by combining a convolutional neural network (CNN), a recurrent neural network (RNN, e.g., long short-term memory (LSTM)), and a vision transformer. The CNN can extract spatial features (within a neighborhood of an image), whereas the LSTM can summarize temporal features. In addition, a transformer with an attention mechanism can capture sparse spatial relations residing in an image or between frames in a video clip. The inputs of the model are short facial videos, while the outputs are the micro-expressions recognized from the videos. The NN models are trained and tested with publicly available facial micro-expression datasets to recognize different micro-expressions (e.g., happiness, fear, anger, surprise, disgust, sadness). Score fusion and improvement metrics are also presented in our experiments. The results of our proposed models are compared with that of literature-reported methods tested on the same datasets. The proposed hybrid model performs the best, where score fusion can dramatically increase recognition performance.<\/jats:p>","DOI":"10.3390\/s23125650","type":"journal-article","created":{"date-parts":[[2023,6,16]],"date-time":"2023-06-16T08:56:01Z","timestamp":1686905761000},"page":"5650","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":23,"title":["Facial Micro-Expression Recognition Enhanced by Score Fusion and a Hybrid Model from Convolutional LSTM and Vision Transformer"],"prefix":"10.3390","volume":"23","author":[{"given":"Yufeng","family":"Zheng","sequence":"first","affiliation":[{"name":"Department of Data Science, University of Mississippi Medical Center, Jackson, MS 39216, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-6894-6108","authenticated-orcid":false,"given":"Erik","family":"Blasch","sequence":"additional","affiliation":[{"name":"MOVEJ Analytics, Fairborn, OH 45324, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1968","published-online":{"date-parts":[[2023,6,16]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","first-page":"205","DOI":"10.1196\/annals.1280.010","article-title":"Darwin, deception, and facial expression","volume":"1000","author":"Ekman","year":"2003","journal-title":"Ann. N. Y. Acad. Sci."},{"key":"ref_2","doi-asserted-by":"crossref","unstructured":"Zhang, L., and Arandjelovi\u0107, O. (2021). Review of Automatic Microexpression Recognition in the Past Decade. Mach. Learn. Knowl. Extr., 3.","DOI":"10.3390\/make3020021"},{"key":"ref_3","doi-asserted-by":"crossref","first-page":"124","DOI":"10.1037\/h0030377","article-title":"Constants across cultures in the face and emotion","volume":"17","author":"Ekman","year":"1971","journal-title":"J. Pers. Soc. Psychol."},{"key":"ref_4","unstructured":"Ekman, P. (2009). The Philosophy of Deception, Oxford University Press."},{"key":"ref_5","doi-asserted-by":"crossref","first-page":"88","DOI":"10.1080\/00332747.1969.11023575","article-title":"Nonverbal leakage and clues to deception","volume":"32","author":"Ekman","year":"1969","journal-title":"Psychiatry"},{"key":"ref_6","doi-asserted-by":"crossref","unstructured":"Bhushan, B. (2015). Study of facial micro-expressions in psychology. Underst. Facial Expr. Commun., 265\u2013286.","DOI":"10.1007\/978-81-322-1934-7_13"},{"key":"ref_7","unstructured":"Ekman, P. (2009). Telling Lies: Clues to Deceit in the Marketplace, Politics, and Marriage (Revised Edition), WW Norton & Company."},{"key":"ref_8","doi-asserted-by":"crossref","first-page":"508","DOI":"10.1111\/j.1467-9280.2008.02116.x","article-title":"Reading between the lies: Identifying concealed and falsified emotions in universal facial expressions","volume":"19","author":"Porter","year":"2008","journal-title":"Psychol. Sci."},{"key":"ref_9","unstructured":"Frank, M., Herbasz, M., Sinuk, K., Keller, A., and Nolan, C. (2009). The Annual Meeting of the International Communication Association, Sheraton."},{"key":"ref_10","doi-asserted-by":"crossref","first-page":"52","DOI":"10.1037\/0033-2909.95.1.52","article-title":"The neuropsychology of facial expression: A review of the neurological and psychological mechanisms for producing facial expressions","volume":"95","author":"Rinn","year":"1984","journal-title":"Psychol. Bull."},{"key":"ref_11","doi-asserted-by":"crossref","first-page":"181","DOI":"10.1007\/s11031-011-9212-2","article-title":"Evidence for training the ability to read microexpressions of emotion","volume":"35","author":"Matsumoto","year":"2011","journal-title":"Motiv. Emot."},{"key":"ref_12","doi-asserted-by":"crossref","unstructured":"Polikovsky, S., Kameda, Y., and Ohta, Y. (2009, January 3). Facial micro-expressions recognition using high speed camera and 3D-gradient descriptor. Proceedings of the 3rd International Conference on Imaging for Crime Detection and Prevention (ICDP 2009), London, UK.","DOI":"10.1049\/ic.2009.0244"},{"key":"ref_13","doi-asserted-by":"crossref","first-page":"673","DOI":"10.1002\/bsl.729","article-title":"From flawed self-assessment to blatant whoppers: The utility of voluntary and involuntary behavior in detecting deception","volume":"24","author":"Ekman","year":"2006","journal-title":"Behav. Sci. Law"},{"key":"ref_14","doi-asserted-by":"crossref","first-page":"119","DOI":"10.1007\/s10919-010-0102-1","article-title":"Executing facial control during deception situations","volume":"35","author":"Hurley","year":"2011","journal-title":"J. Nonverbal Behav."},{"key":"ref_15","first-page":"5826","article-title":"Video-based Facial Micro-Expression Analysis: A Survey of Datasets, Features and Algorithms","volume":"44","author":"Ben","year":"2021","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_16","doi-asserted-by":"crossref","first-page":"299","DOI":"10.1109\/TAFFC.2015.2485205","article-title":"A Main Directional Mean Optical Flow Feature for Spontaneous Micro-Expression Recognition","volume":"7","author":"Liu","year":"2015","journal-title":"IEEE Trans. Affect. Comput."},{"key":"ref_17","doi-asserted-by":"crossref","first-page":"251","DOI":"10.1016\/j.neucom.2018.05.107","article-title":"Micro-expression recognition with small sample size by transferring long-term convolutional neural network","volume":"312","author":"Wang","year":"2018","journal-title":"Neurocomputing"},{"key":"ref_18","doi-asserted-by":"crossref","first-page":"1973","DOI":"10.1109\/TAFFC.2022.3213509","article-title":"Short and Long Range Relation Based Spatio-Temporal Transformer for Micro-Expression Recognition","volume":"13","author":"Zhang","year":"2022","journal-title":"IEEE Trans. Affect. Comput."},{"key":"ref_19","doi-asserted-by":"crossref","first-page":"116","DOI":"10.1109\/TAFFC.2016.2573832","article-title":"SAMM: A Spontaneous Micro-Facial Movement Dataset","volume":"9","author":"Davison","year":"2016","journal-title":"IEEE Trans. Affect. Comput."},{"key":"ref_20","doi-asserted-by":"crossref","unstructured":"Li, X., Pfister, T., Huang, X., and Zhao, G. (2013, January 22\u201326). A spontaneous microexpression database: Inducement, collection and baseline. Proceedings of the 2013 10th IEEE International Conference and Workshops on Automatic Face and Gesture Recognition (FG), Shanghai, China.","DOI":"10.1109\/FG.2013.6553717"},{"key":"ref_21","unstructured":"Yan, W.-J., Wu, Q., Liu, Y.-J., Wang, S.-J., and Fu, X. (2013, January 22\u201326). CASME database: A dataset of spontaneous micro-expressions collected from neutralized faces. Proceedings of the 2013 10th IEEE International Conference and Workshops on Automatic Face and Gesture Recognition (FG), Shanghai, China."},{"key":"ref_22","doi-asserted-by":"crossref","unstructured":"Yan, W.-J., Li, X., Wang, S.-J., Zhao, G., Liu, Y.-J., Chen, Y.-H., and Fu, X. (2014). CASME II: An Improved Spontaneous Micro-Expression Database and the Baseline Evaluation. PLoS ONE, 9.","DOI":"10.1371\/journal.pone.0086041"},{"key":"ref_23","doi-asserted-by":"crossref","first-page":"424","DOI":"10.1109\/TAFFC.2017.2654440","article-title":"CAS(ME): A Database for Spontaneous Macro-Expression and Micro-Expression Spotting and Recognition","volume":"9","author":"Qu","year":"2017","journal-title":"IEEE Trans. Affect. Comput."},{"key":"ref_24","first-page":"126","article-title":"Facial action coding system (FACS): A technique for the measurement of facial actions","volume":"47","author":"Ekman","year":"1978","journal-title":"Riv. Psichiatr."},{"key":"ref_25","unstructured":"Krizhevsky, A., Sutskever, I., and Hinton, G.E. (2012, January 3\u20136). Imagenet classification with deep convolutional neural networks. Proceedings of the 25th International Conference on Neural Information Processing Systems, Lake Tahoe, NV, USA."},{"key":"ref_26","doi-asserted-by":"crossref","unstructured":"He, K., Zhang, X., Ren, S., and Sun, J. (26\u20131, January 26). Deep Residual Learning for Image Recognition. Proceedings of the 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.90"},{"key":"ref_27","doi-asserted-by":"crossref","unstructured":"Chollet, F. (2017, January 21\u201326). Xception: Deep Learning with Depthwise Separable Convolutions. Proceedings of the 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.195"},{"key":"ref_28","unstructured":"Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K. (2018). BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. arXiv."},{"key":"ref_29","unstructured":"Wensel, J., Ullah, H., and Munir, A. (2022, August 16). ViT-ReT: Vision and Recurrent Transformer Neural Networks for Human Activity Recognition in Videos. Available online: https:\/\/arxiv.org\/pdf\/2208.07929.pdf."},{"key":"ref_30","doi-asserted-by":"crossref","first-page":"281","DOI":"10.1109\/34.982906","article-title":"A theoretical study on six classifier fusion strategies","volume":"24","author":"Kuncheva","year":"2002","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_31","doi-asserted-by":"crossref","first-page":"861","DOI":"10.1016\/S0031-3203(01)00103-0","article-title":"Decision-level fusion in fingerprint verification","volume":"35","author":"Prabhakar","year":"2002","journal-title":"Pattern Recognit."},{"key":"ref_32","doi-asserted-by":"crossref","unstructured":"Ulery, B., Hicklin, A., Watson, C., Fellner, W., and Hallinan, P. (2006, September 01). Studies of Biometric Fusion. NIST Interagency Report, Available online: https:\/\/tsapps.nist.gov\/publication\/get_pdf.cfm?pub_id=50872.","DOI":"10.6028\/NIST.IR.7346"},{"key":"ref_33","first-page":"106","article-title":"An Exploration of the Impacts of Three Factors in Multimodal Biometric Score Fusion: Score Modality, Recognition Method, and Fusion Process","volume":"9","author":"Zheng","year":"2015","journal-title":"J. Adv. Inf. Fusion"},{"key":"ref_34","doi-asserted-by":"crossref","unstructured":"Zheng, Y., Blasch, E., and Liu, Z. (2018). Multispectral Image Fusion and Colorization, SPIE Press.","DOI":"10.1117\/3.2316455"},{"key":"ref_35","doi-asserted-by":"crossref","first-page":"121","DOI":"10.1023\/A:1009715923555","article-title":"A tutorial on support vector machines for pattern recognition","volume":"2","author":"Burges","year":"1998","journal-title":"Data Min. Knowl. Discov."},{"key":"ref_36","doi-asserted-by":"crossref","unstructured":"Hastie, T., Tibshirani, R., and Friedman, J.H. (2009). The Elements of Statistical Learning: Data Mining, Inference, and Prediction, Springer.","DOI":"10.1007\/978-0-387-84858-7"}],"container-title":["Sensors"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1424-8220\/23\/12\/5650\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,10]],"date-time":"2025-10-10T19:56:39Z","timestamp":1760126199000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1424-8220\/23\/12\/5650"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,6,16]]},"references-count":36,"journal-issue":{"issue":"12","published-online":{"date-parts":[[2023,6]]}},"alternative-id":["s23125650"],"URL":"https:\/\/doi.org\/10.3390\/s23125650","relation":{},"ISSN":["1424-8220"],"issn-type":[{"value":"1424-8220","type":"electronic"}],"subject":[],"published":{"date-parts":[[2023,6,16]]}}}