{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,8]],"date-time":"2026-05-08T15:59:59Z","timestamp":1778255999336,"version":"3.51.4"},"reference-count":54,"publisher":"MDPI AG","issue":"7","license":[{"start":{"date-parts":[[2019,6,29]],"date-time":"2019-06-29T00:00:00Z","timestamp":1561766400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Entropy"],"abstract":"<jats:p>Automatic emotion recognition has become an important trend in many artificial intelligence (AI) based applications and has been widely explored in recent years. Most research in the area of automated emotion recognition is based on facial expressions or speech signals. Although the influence of the emotional state on body movements is undeniable, this source of expression is still underestimated in automatic analysis. In this paper, we propose a novel method to recognise seven basic emotional states\u2014namely, happy, sad, surprise, fear, anger, disgust and neutral\u2014utilising body movement. We analyse motion capture data under seven basic emotional states recorded by professional actor\/actresses using Microsoft Kinect v2 sensor. We propose a new representation of affective movements, based on sequences of body joints. The proposed algorithm creates a sequential model of affective movement based on low level features inferred from the spacial location and the orientation of joints within the tracked skeleton. In the experimental results, different deep neural networks were employed and compared to recognise the emotional state of the acquired motion sequences. The experimental results conducted in this work show the feasibility of automatic emotion recognition from sequences of body gestures, which can serve as an additional source of information in multimodal emotion recognition.<\/jats:p>","DOI":"10.3390\/e21070646","type":"journal-article","created":{"date-parts":[[2019,7,1]],"date-time":"2019-07-01T03:23:59Z","timestamp":1561951439000},"page":"646","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":86,"title":["Emotion Recognition from Skeletal Movements"],"prefix":"10.3390","volume":"21","author":[{"given":"Tomasz","family":"Sapi\u0144ski","sequence":"first","affiliation":[{"name":"Institute of Mechatronics and Information Systems Lodz University of Technology, 90-924 Lodz, Poland"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-3416-5554","authenticated-orcid":false,"given":"Dorota","family":"Kami\u0144ska","sequence":"additional","affiliation":[{"name":"Institute of Mechatronics and Information Systems Lodz University of Technology, 90-924 Lodz, Poland"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-4282-7052","authenticated-orcid":false,"given":"Adam","family":"Pelikant","sequence":"additional","affiliation":[{"name":"Institute of Mechatronics and Information Systems Lodz University of Technology, 90-924 Lodz, Poland"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-8460-5717","authenticated-orcid":false,"given":"Gholamreza","family":"Anbarjafari","sequence":"additional","affiliation":[{"name":"iCV Lab, Institute of Technology, University of Tartu, 51014 Tartu, Estonia"},{"name":"Faculty of Engineering, Hasan Kalyoncu University, 27000 Sahinbey, Gaziantep, Turkey"},{"name":"Institute of Digital Technologies, Loughborough University London, London E15 2GZ, UK"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2019,6,29]]},"reference":[{"key":"ref_1","unstructured":"Ekman, P. (2002). Facial action coding system (FACS). A Human Face, Available online: https:\/\/www.cs.cmu.edu\/~face\/facs.htm."},{"key":"ref_2","unstructured":"Pease, A., McIntosh, J., and Cullen, P. (1981). Body Language, Malor Books. Camel."},{"key":"ref_3","unstructured":"Izdebski, K. (2008). Emotions in the Human Voice, Volume 3: Culture and Perception, Plural Publishing."},{"key":"ref_4","doi-asserted-by":"crossref","first-page":"2067","DOI":"10.1109\/TPAMI.2008.26","article-title":"Emotion recognition based on physiological changes in music listening","volume":"30","author":"Kim","year":"2008","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_5","unstructured":"Ekman, P. (2012). Emotions Revealed: Understanding Faces and Feelings, Hachette."},{"key":"ref_6","doi-asserted-by":"crossref","first-page":"45","DOI":"10.1111\/spc3.12083","article-title":"Emotional mimicry: Why and when we mimic emotions","volume":"8","author":"Hess","year":"2014","journal-title":"Soc. Personal. Psychol. Compass"},{"key":"ref_7","unstructured":"Kulkarni, K., Corneanu, C., Ofodile, I., Escalera, S., Baro, X., Hyniewska, S., Allik, J., and Anbarjafari, G. (2018). Automatic recognition of facial displays of unfelt emotions. IEEE Trans. Affect. Comput."},{"key":"ref_8","doi-asserted-by":"crossref","unstructured":"Mehrabian, A. (2017). Nonverbal Communication, Routledge.","DOI":"10.4324\/9781351308724"},{"key":"ref_9","unstructured":"Mehrabian, A. (1971). Silent Messages, Wadsworth."},{"key":"ref_10","doi-asserted-by":"crossref","first-page":"98","DOI":"10.1016\/j.inffus.2017.02.003","article-title":"A review of affective computing: From unimodal analysis to multimodal fusion","volume":"37","author":"Poria","year":"2017","journal-title":"Inf. Fusion"},{"key":"ref_11","unstructured":"Corneanu, C., Noroozi, F., Kaminska, D., Sapinski, T., Escalera, S., and Anbarjafari, G. (2018). Survey on emotional body gesture recognition. IEEE Trans. Affect. Comput."},{"key":"ref_12","doi-asserted-by":"crossref","first-page":"24","DOI":"10.1016\/j.jvcir.2013.04.007","article-title":"Sequence of the most informative joints (smij): A new representation for human skeletal action recognition","volume":"25","author":"Ofli","year":"2014","journal-title":"J. Vis. Commun. Image Represent."},{"key":"ref_13","unstructured":"Gunes, H., and Piccardi, M. (2005, January 12). Affect recognition from face and body: Early fusion vs. late fusion. Proceedings of the 2005 IEEE International Conference on Systems, Man and Cybernetics, Waikoloa, HI, USA."},{"key":"ref_14","doi-asserted-by":"crossref","unstructured":"Ofodile, I., Helmi, A., Clap\u00e9s, A., Avots, E., Peensoo, K.M., Valdma, S.M., Valdmann, A., Valtna-Lukner, H., Omelkov, S., and Escalera, S. (2019). Action Recognition Using Single-Pixel Time-of-Flight Detection. Entropy, 21.","DOI":"10.3390\/e21040414"},{"key":"ref_15","doi-asserted-by":"crossref","unstructured":"Kipp, M., and Martin, J.C. (2009, January 10\u201312). Gesture and emotion: Can basic gestural form features discriminate emotions?. Proceedings of the 3rd International Conference on Affective Computing and Intelligent Interaction and Workshops (ACII 2009), Amsterdam, The Netherlands.","DOI":"10.1109\/ACII.2009.5349544"},{"key":"ref_16","doi-asserted-by":"crossref","unstructured":"Bernhardt, D., and Robinson, P. (2009). Detecting emotions from connected action sequences. Visual Informatics: Bridging Research and Practice, Proceedings of the International Visual Informatics Conference (IVIC 2009), Kuala Lumpur, Malaysia, 11\u201313 November 2009, Springer.","DOI":"10.1007\/978-3-642-05036-7_1"},{"key":"ref_17","doi-asserted-by":"crossref","unstructured":"Rasti, P., Uiboupin, T., Escalera, S., and Anbarjafari, G. (2016). Convolutional neural network super resolution for face recognition in surveillance monitoring. Articulated Motion and Deformable Objects (AMDO 2016), Springer.","DOI":"10.1007\/978-3-319-41778-3_18"},{"key":"ref_18","doi-asserted-by":"crossref","first-page":"25","DOI":"10.1155\/2009\/482585","article-title":"Data fusion boosted face recognition based on probability distribution functions in different colour channels","volume":"2009","author":"Demirel","year":"2009","journal-title":"Eurasip J. Adv. Signal Process."},{"key":"ref_19","doi-asserted-by":"crossref","unstructured":"Litvin, A., Nasrollahi, K., Ozcinar, C., Guerrero, S.E., Moeslund, T.B., and Anbarjafari, G. (2019). A Novel Deep Network Architecture for Reconstructing RGB Facial Images from Thermal for Face Recognition. Multimed. Tools Appl.","DOI":"10.1007\/s11042-019-7667-4"},{"key":"ref_20","doi-asserted-by":"crossref","unstructured":"Nasrollahi, K., Escalera, S., Rasti, P., Anbarjafari, G., Baro, X., Escalante, H.J., and Moeslund, T.B. (2015, January 10\u201313). Deep learning based super-resolution for improved action recognition. Proceedings of the IEEE 2015 International Conference on Image Processing Theory, Tools and Applications (IPTA), Orleans, France.","DOI":"10.1109\/IPTA.2015.7367098"},{"key":"ref_21","doi-asserted-by":"crossref","unstructured":"Glowinski, D., Mortillaro, M., Scherer, K., Dael, N., Volpe, G., and Camurri, A. (2015, January 21\u201324). Towards a minimal representation of affective gestures. Proceedings of the 2015 International Conference on Affective Computing and Intelligent Interaction (ACII), Xi\u2019an, China.","DOI":"10.1109\/ACII.2015.7344616"},{"key":"ref_22","unstructured":"Castellano, G. (2008). Movement Expressivity Analysis in Affective Computers: From Recognition to Expression of Emotion. [Ph.D. Thesis, Department of Communication, Computer and System Sciences, University of Genoa]. (Unpublished)."},{"key":"ref_23","doi-asserted-by":"crossref","unstructured":"Kaza, K., Psaltis, A., Stefanidis, K., Apostolakis, K.C., Thermos, S., Dimitropoulos, K., and Daras, P. (2016). Body motion analysis for emotion recognition in serious games. Universal Access in Human-Computer Interaction, Proceedings of the International Conference on Universal Access in Human-Computer Interaction, Toronto, ON, Canada, 17\u201322 July 2016, Springer.","DOI":"10.1007\/978-3-319-40244-4_4"},{"key":"ref_24","doi-asserted-by":"crossref","first-page":"1027","DOI":"10.1109\/TSMCB.2010.2103557","article-title":"Automatic recognition of non-acted affective postures","volume":"41","author":"Kleinsmith","year":"2011","journal-title":"IEEE Trans. Syst. Man, Cybern. Part B (Cybern.)"},{"key":"ref_25","doi-asserted-by":"crossref","first-page":"199","DOI":"10.1109\/TCIAIG.2012.2202663","article-title":"Continuous recognition of player\u2019s affective body expression as dynamic quality of aesthetic experience","volume":"4","author":"Savva","year":"2012","journal-title":"IEEE Trans. Comput. Intell. Games"},{"key":"ref_26","doi-asserted-by":"crossref","first-page":"621","DOI":"10.1007\/s12369-014-0243-1","article-title":"Recognizing emotions conveyed by human gait","volume":"6","author":"Venture","year":"2014","journal-title":"Int. J. Soc. Robot."},{"key":"ref_27","doi-asserted-by":"crossref","first-page":"454","DOI":"10.1109\/THMS.2014.2310953","article-title":"Affective movement recognition based on generative and discriminative stochastic dynamic models","volume":"44","author":"Samadani","year":"2014","journal-title":"IEEE Trans. Hum. Mach. Syst."},{"key":"ref_28","doi-asserted-by":"crossref","first-page":"140","DOI":"10.1016\/j.neunet.2015.09.009","article-title":"Multimodal emotional state recognition using sequence-dependent deep hierarchical features","volume":"72","author":"Barros","year":"2015","journal-title":"Neural Netw."},{"key":"ref_29","doi-asserted-by":"crossref","unstructured":"Gunes, H., and Piccardi, M. (2006, January 20\u201324). A bimodal face and body gesture database for automatic analysis of human nonverbal affective behavior. Proceedings of the IEEE 18th International Conference on Pattern Recognition (ICPR 2006), Hong Kong, China.","DOI":"10.1109\/ICPR.2006.39"},{"key":"ref_30","unstructured":"Li, B., Bai, B., and Han, C. (2018). Upper body motion recognition based on key frame and random forest regression. Multimed. Tools Appl., 1\u201316."},{"key":"ref_31","doi-asserted-by":"crossref","unstructured":"Sapi\u0144ski, T., Kami\u0144ska, D., Pelikant, A., Ozcinar, C., Avots, E., and Anbarjafari, G. (2018). Multimodal Database of Emotional Speech, Video and Gestures. Pattern Recognition and Information Forensics, Proceedings of the International Conference on Pattern Recognitionm, Beijing, China, 20\u201324 August 2018, Springer.","DOI":"10.1007\/978-3-030-05792-3_15"},{"key":"ref_32","doi-asserted-by":"crossref","first-page":"124","DOI":"10.1037\/h0030377","article-title":"Constants across cultures in the face and emotion","volume":"17","author":"Ekman","year":"1971","journal-title":"J. Personal. Soc. Psychol."},{"key":"ref_33","unstructured":"(2018, January 11). Microsoft Kinect. Available online: https:\/\/msdn.microsoft.com\/."},{"key":"ref_34","unstructured":"Bulut, E., and Capin, T. (2007). Key frame extraction from motion capture data by curve saliency. Comput. Animat. Soc. Agents, 119. Available online: https:\/\/s3.amazonaws.com\/academia.edu.documents\/42103016\/casa.pdf?response-content-disposition=inline%3B%20filename%3DKey_frame_extraction_from_motion_capture.pdf&X-Amz-Algorithm=AWS4-HMAC-SHA256&X-Amz-Credential=AKIAIWOWYYGZ2Y53UL3A%2F20190629%2Fus-east-1%2Fs3%2Faws4_request&X-Amz-Date=20190629T015324Z&X-Amz-Expires=3600&X-Amz-SignedHeaders=host&X-Amz-Signature=7c38895c4f79ebe3faf97dc8839ec237a2851828bd91bc26c8518cabfce692d6."},{"key":"ref_35","doi-asserted-by":"crossref","first-page":"355","DOI":"10.1016\/0004-3702(87)90070-1","article-title":"Three-dimensional object recognition from single two-dimensional images","volume":"31","author":"Lowe","year":"1987","journal-title":"Artif. Intell."},{"key":"ref_36","doi-asserted-by":"crossref","first-page":"1047","DOI":"10.3390\/ijerph7031047","article-title":"Leg length, body proportion, and health: a review with a note on beauty","volume":"7","author":"Bogin","year":"2010","journal-title":"Int. J. Environ. Res. Public Health"},{"key":"ref_37","unstructured":"Ioffe, S., and Szegedy, C. (2015). Batch normalization: Accelerating deep network training by reducing internal covariate shift. arXiv."},{"key":"ref_38","doi-asserted-by":"crossref","unstructured":"Laurent, C., Pereyra, G., Brakel, P., Zhang, Y., and Bengio, Y. (2016, January 20\u201325). Batch normalized recurrent neural networks. Proceedings of the IEEE 2016 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Shanghai, China.","DOI":"10.1109\/ICASSP.2016.7472159"},{"key":"ref_39","doi-asserted-by":"crossref","first-page":"1464","DOI":"10.1109\/23.589532","article-title":"Importance of input data normalization for the application of neural networks to complex industrial problems","volume":"44","author":"Sola","year":"1997","journal-title":"IEEE Trans. Nucl. Sci."},{"key":"ref_40","unstructured":"Noroozi, F., Marjanovic, M., Njegus, A., Escalera, S., and Anbarjafari, G. (2018). A Study of Language and Classifier-independent Feature Analysis for Vocal Emotion Recognition. arXiv."},{"key":"ref_41","doi-asserted-by":"crossref","unstructured":"Avots, E., Sapi\u0144ski, T., Bachmann, M., and Kami\u0144ska, D. (2018). Audiovisual emotion recognition in wild. Mach. Vis. Appl., 1\u201311.","DOI":"10.1007\/s00138-018-0960-9"},{"key":"ref_42","unstructured":"Krizhevsky, A., Sutskever, I., and Hinton, G.E. (2012). Imagenet classification with deep convolutional neural networks. Adv. Neural Inf. Process. Syst., 1097\u20131105."},{"key":"ref_43","doi-asserted-by":"crossref","first-page":"107","DOI":"10.1142\/S0218488598000094","article-title":"The vanishing gradient problem during learning recurrent neural nets and problem solutions","volume":"6","author":"Hochreiter","year":"1998","journal-title":"Int. J. Uncertain. Fuzziness Knowl. Based Syst."},{"key":"ref_44","doi-asserted-by":"crossref","first-page":"234","DOI":"10.1109\/TMM.2018.2856094","article-title":"Exploiting recurrent neural networks and leap motion controller for the recognition of sign language and semaphoric hand gestures","volume":"21","author":"Avola","year":"2018","journal-title":"IEEE Trans. Multimed."},{"key":"ref_45","first-page":"190","article-title":"Training and analysing deep recurrent neural networks","volume":"1","author":"Hermans","year":"2013","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_46","unstructured":"Kingma, D.P., and Ba, J. (2014). Adam: A Method for Stochastic Optimization. CoRR, Available online: https:\/\/arxiv.org\/abs\/1412.6980."},{"key":"ref_47","doi-asserted-by":"crossref","first-page":"19","DOI":"10.1007\/s10479-005-5724-z","article-title":"A tutorial on the cross-entropy method","volume":"134","author":"Kroese","year":"2005","journal-title":"Ann. Oper. Res."},{"key":"ref_48","unstructured":"Pham, H.H., Khoudour, L., Crouzil, A., Zegers, P., and Velastin, S.A. (2019, June 28). Learning and recognizing human action from skeleton movement with deep residual neural networks. Available online: https:\/\/arxiv.org\/abs\/1803.07780."},{"key":"ref_49","unstructured":"He, K., Zhang, X., Ren, S., and Sun, J. (July, January 26). Deep residual learning for image recognition. Proceedings of the IEEE Conference On Computer Vision And Pattern Recognition, Las Vegas, NV, USA."},{"key":"ref_50","unstructured":"Holmes, G., Donkin, A., and Witten, I.H. (December, January 29). Weka: A machine learning workbench. Proceedings of the ANZIIS \u201994\u2014Australian New Zealnd Intelligent Information Systems Conference, Brisbane, Australia."},{"key":"ref_51","doi-asserted-by":"crossref","unstructured":"G\u00fcler, R.A., Neverova, N., and Kokkinos, I. (2018). Densepose: Dense human pose estimation in the wild. arXiv.","DOI":"10.1109\/CVPR.2018.00762"},{"key":"ref_52","doi-asserted-by":"crossref","unstructured":"Zhang, S., Liu, X., and Xiao, J. (2017, January 24\u201331). On geometric features for skeleton-based action recognition using multilayer lstm networks. Proceedings of the 2017 IEEE Winter Conference on Applications of Computer Vision (WACV), Santa Rosa, CA, USA.","DOI":"10.1109\/WACV.2017.24"},{"key":"ref_53","doi-asserted-by":"crossref","unstructured":"Song, S., Lan, C., Xing, J., Zeng, W., and Liu, J. (2017, January 4\u20139). An end-to-end spatio-temporal attention model for human action recognition from skeleton data. Proceedings of the Thirty-First AAAI Conference on Artificial Intelligence, San Francisco, CA, USA.","DOI":"10.1609\/aaai.v31i1.11212"},{"key":"ref_54","unstructured":"Minh, T.L., Inoue, N., and Shinoda, K. (2018). A fine-to-coarse convolutional neural network for 3d human action recognition. arXiv."}],"container-title":["Entropy"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1099-4300\/21\/7\/646\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,11]],"date-time":"2025-10-11T13:02:29Z","timestamp":1760187749000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1099-4300\/21\/7\/646"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2019,6,29]]},"references-count":54,"journal-issue":{"issue":"7","published-online":{"date-parts":[[2019,7]]}},"alternative-id":["e21070646"],"URL":"https:\/\/doi.org\/10.3390\/e21070646","relation":{},"ISSN":["1099-4300"],"issn-type":[{"value":"1099-4300","type":"electronic"}],"subject":[],"published":{"date-parts":[[2019,6,29]]}}}