{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,1,15]],"date-time":"2026-01-15T14:53:44Z","timestamp":1768488824282,"version":"3.49.0"},"reference-count":28,"publisher":"Association for Computing Machinery (ACM)","issue":"2","license":[{"start":{"date-parts":[[2023,9,26]],"date-time":"2023-09-26T00:00:00Z","timestamp":1695686400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Multimedia Comput. Commun. Appl."],"published-print":{"date-parts":[[2024,2,29]]},"abstract":"<jats:p>A fascinating challenge in robotics-human interaction is imitating the emotion recognition capability of humans to robots with the aim to make human-robotics interaction natural, genuine and intuitive. To achieve the natural interaction in affective robots, human-machine interfaces, and autonomous vehicles, understanding our attitudes and opinions is very important, and it provides a practical and feasible path to realize the connection between machine and human. Multimodal interface that includes voice along with facial expression can manifest a large range of nuanced emotions compared to purely textual interfaces and provide a great value to improve the intelligence level of effective communication. Interfaces that fail to manifest or ignore user emotions may significantly impact the performance and risk being perceived as cold, socially inept, untrustworthy, and incompetent. To equip a child well for life, we need to help our children identify their feelings, manage them well, and express their needs in healthy, respectful, and direct ways. Early identification of emotional deficits can help to prevent low social functioning in children. In this work, we analyzed the child\u2019s spontaneous behavior using multimodal facial expression and voice signal presenting multimodal transformer-based last feature fusion for facial behavior analysis in children to extract contextualized representations from RGB video sequence and Hematoxylin and eosin video sequence and then using these representations followed by pairwise concatenations of contextualized representations using cross-feature fusion technique to predict users emotions. To validate the performance of the proposed framework, we have performed experiments with the different pairwise concatenations of contextualized representations that showed significantly better performance than state-of-the-art method. Besides, we perform t-distributed stochastic neighbor embedding visualization to visualize the discriminative feature in lower dimension space and probability density estimation to visualize the prediction capability of our proposed model.<\/jats:p>","DOI":"10.1145\/3539577","type":"journal-article","created":{"date-parts":[[2022,5,26]],"date-time":"2022-05-26T15:10:29Z","timestamp":1653577829000},"page":"1-17","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":5,"title":["Spontaneous Facial Behavior Analysis Using Deep Transformer-based Framework for Child\u2013computer Interaction"],"prefix":"10.1145","volume":"20","author":[{"ORCID":"https:\/\/orcid.org\/0000-0003-3102-1595","authenticated-orcid":false,"given":"Abdul","family":"Qayyum","sequence":"first","affiliation":[{"name":"Department of Computer Science and Engineering, Universit\u00e9 Bourgogne, France"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-3930-6600","authenticated-orcid":false,"given":"Imran","family":"Razzak","sequence":"additional","affiliation":[{"name":"School of Computer Science and Engineering, University of New South Wales, Sydney, Australia"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-5727-3697","authenticated-orcid":false,"given":"M.","family":"Tanveer","sequence":"additional","affiliation":[{"name":"Department of Mathematics, Indian Institute of Technology Indore, Simrol, Indore, India"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-4444-5776","authenticated-orcid":false,"given":"Moona","family":"Mazher","sequence":"additional","affiliation":[{"name":"Department of Computer Engineering and Mathematics, University Rovira i Virgili, Spain"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2023,9,26]]},"reference":[{"key":"e_1_3_2_2_2","doi-asserted-by":"publisher","DOI":"10.1109\/ACCESS.2018.2885279"},{"key":"e_1_3_2_3_2","first-page":"836","volume-title":"Proceedings of the IEEE International Conference on Pervasive Computing and Communications Workshops (PerCom Workshops)","author":"Albraikan Amani","year":"2018","unstructured":"Amani Albraikan, Diana P. Tob\u00f3n, and Abdulmotaleb El Saddik. 2018. Hyper-parameter optimization for emotion detection using physiological signals. In Proceedings of the IEEE International Conference on Pervasive Computing and Communications Workshops (PerCom Workshops). IEEE, 836\u2013841."},{"issue":"19","key":"e_1_3_2_4_2","doi-asserted-by":"crossref","first-page":"8402","DOI":"10.1109\/JSEN.2018.2867221","article-title":"Toward user-independent emotion recognition using physiological signals","volume":"19","author":"Albraikan Amani","year":"2018","unstructured":"Amani Albraikan, Diana P. Tob\u00f3n, and Abdulmotaleb El Saddik. 2018. Toward user-independent emotion recognition using physiological signals. IEEE Sensors J. 19, 19 (2018), 8402\u20138412.","journal-title":"IEEE Sensors J."},{"key":"e_1_3_2_5_2","first-page":"1","volume-title":"Affect and Emotion in Human-Computer Interaction","author":"Beale Russell","year":"2008","unstructured":"Russell Beale and Christian Peter. 2008. The role of affect and emotion in HCI. In Affect and Emotion in Human-Computer Interaction. Springer, 1\u201311."},{"key":"e_1_3_2_6_2","doi-asserted-by":"publisher","DOI":"10.1145\/3176649"},{"key":"e_1_3_2_7_2","doi-asserted-by":"crossref","first-page":"198","DOI":"10.1016\/j.neulet.2019.01.032","article-title":"Facial emotion recognition in deaf children: Evidence from event-related potentials and event-related spectral perturbation analysis","volume":"703","author":"Gu Huang","year":"2019","unstructured":"Huang Gu, Qiong Chen, Xiaoli Xing, Junfeng Zhao, and Xiaoming Li. 2019. Facial emotion recognition in deaf children: Evidence from event-related potentials and event-related spectral perturbation analysis. Neuroscience Lett. 703 (2019), 198\u2013204.","journal-title":"Neuroscience Lett."},{"key":"e_1_3_2_8_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.imavis.2019.02.004"},{"key":"e_1_3_2_9_2","doi-asserted-by":"crossref","unstructured":"Yelin Kim and Emily Mower Provost. 2015. Emotion recognition during speech using dynamics of multiple regions of the face. (2015).","DOI":"10.1145\/2808204"},{"key":"e_1_3_2_10_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.ijcci.2020.100203"},{"issue":"1","key":"e_1_3_2_11_2","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/3464382","article-title":"Facial-expression-aware emotional color transfer based on convolutional neural network","volume":"18","author":"Liu Shiguang","year":"2022","unstructured":"Shiguang Liu, Huixin Wang, and Min Pei. 2022. Facial-expression-aware emotional color transfer based on convolutional neural network. ACM Trans. Multim. Comput., Commun. Applic. 18, 1 (2022), 1\u201319.","journal-title":"ACM Trans. Multim. Comput., Commun. Applic."},{"key":"e_1_3_2_12_2","article-title":"Evaluation of interpretability for deep learning algorithms in EEG emotion recognition: A case study in autism","author":"Mayor-Torres Juan Manuel","year":"2021","unstructured":"Juan Manuel Mayor-Torres, Sara Medina-DeVilliers, Tessa Clarkson, Matthew D. Lerner, and Giuseppe Riccardi. 2021. Evaluation of interpretability for deep learning algorithms in EEG emotion recognition: A case study in autism. arXiv preprint arXiv:2111.13208 (2021).","journal-title":"arXiv preprint arXiv:2111.13208"},{"key":"e_1_3_2_13_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.bspc.2019.101721"},{"key":"e_1_3_2_14_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.compeleceng.2016.04.009"},{"key":"e_1_3_2_15_2","doi-asserted-by":"publisher","DOI":"10.1145\/3311747"},{"key":"e_1_3_2_16_2","doi-asserted-by":"crossref","first-page":"104375","DOI":"10.1016\/j.imavis.2022.104375","article-title":"Progressive ShallowNet for large scale dynamic and spontaneous facial behaviour analysis in children","author":"Qayyum Abdul","year":"2022","unstructured":"Abdul Qayyum, Imran Razzak, Nour Moustafa, and Mona Mazhar. 2022. Progressive ShallowNet for large scale dynamic and spontaneous facial behaviour analysis in children. Image Vis. Comput. (2022), 104375.","journal-title":"Image Vis. Comput."},{"key":"e_1_3_2_17_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.imavis.2008.08.005"},{"key":"e_1_3_2_18_2","first-page":"1","article-title":"Emotion-sensitive human-computer interaction (HCI): State of the art\u2014Seminar paper","author":"Voeffray S.","year":"2011","unstructured":"S. Voeffray. 2011. Emotion-sensitive human-computer interaction (HCI): State of the art\u2014Seminar paper. Emot. Recog. (2011), 1\u20134.","journal-title":"Emot. Recog."},{"key":"e_1_3_2_19_2","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2019.2956143"},{"key":"e_1_3_2_20_2","doi-asserted-by":"publisher","DOI":"10.1145\/3355397"},{"key":"e_1_3_2_21_2","doi-asserted-by":"publisher","DOI":"10.1145\/3152118"},{"key":"e_1_3_2_22_2","doi-asserted-by":"publisher","DOI":"10.1145\/2993148.2997639"},{"key":"e_1_3_2_23_2","doi-asserted-by":"publisher","DOI":"10.1145\/3490686"},{"key":"e_1_3_2_24_2","first-page":"222","volume-title":"Proceedings of the European Conference on Computer Vision (ECCV)","author":"Zeng Jiabei","year":"2018","unstructured":"Jiabei Zeng, Shiguang Shan, and Xilin Chen. 2018. Facial expression recognition with inconsistently annotated datasets. In Proceedings of the European Conference on Computer Vision (ECCV). 222\u2013237."},{"issue":"1","key":"e_1_3_2_25_2","first-page":"1","article-title":"Deep learning\u2013based multimedia analytics: A review","volume":"15","author":"Zhang Wei","year":"2019","unstructured":"Wei Zhang, Ting Yao, Shiai Zhu, and Abdulmotaleb El Saddik. 2019. Deep learning\u2013based multimedia analytics: A review. ACM Trans. Multim. Comput., Commun. Applic. 15, 1s (2019), 1\u201326.","journal-title":"ACM Trans. Multim. Comput., Commun. Applic."},{"issue":"1","key":"e_1_3_2_26_2","first-page":"1","article-title":"CovLets: A second-order descriptor for modeling multiple features","volume":"16","author":"Zhang Zhaoxin","year":"2020","unstructured":"Zhaoxin Zhang, Changyong Guo, Fanzhi Meng, Taizhong Xu, and Junkai Huang. 2020. CovLets: A second-order descriptor for modeling multiple features. ACM Trans. Multim. Comput., Commun. Applic. 16, 1s (2020), 1\u201314.","journal-title":"ACM Trans. Multim. Comput., Commun. Applic."},{"key":"e_1_3_2_27_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2007.1110"},{"issue":"1","key":"e_1_3_2_28_2","doi-asserted-by":"crossref","first-page":"38","DOI":"10.1109\/TSMCB.2010.2044788","article-title":"Graph-preserving sparse nonnegative matrix factorization with application to facial expression recognition","volume":"41","author":"Zhi Ruicong","year":"2010","unstructured":"Ruicong Zhi, Markus Flierl, Qiuqi Ruan, and W. Bastiaan Kleijn. 2010. Graph-preserving sparse nonnegative matrix factorization with application to facial expression recognition. IEEE Trans. Syst., Man, Cyber., Part B (Cyber.) 41, 1 (2010), 38\u201352.","journal-title":"IEEE Trans. Syst., Man, Cyber., Part B (Cyber.)"},{"key":"e_1_3_2_29_2","first-page":"2562","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition","author":"Zhong Lin","year":"2012","unstructured":"Lin Zhong, Qingshan Liu, Peng Yang, Bo Liu, Junzhou Huang, and Dimitris N. Metaxas. 2012. Learning active facial patches for expression analysis. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. IEEE, 2562\u20132569."}],"container-title":["ACM Transactions on Multimedia Computing, Communications, and Applications"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3539577","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3539577","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T17:51:29Z","timestamp":1750182689000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3539577"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,9,26]]},"references-count":28,"journal-issue":{"issue":"2","published-print":{"date-parts":[[2024,2,29]]}},"alternative-id":["10.1145\/3539577"],"URL":"https:\/\/doi.org\/10.1145\/3539577","relation":{},"ISSN":["1551-6857","1551-6865"],"issn-type":[{"value":"1551-6857","type":"print"},{"value":"1551-6865","type":"electronic"}],"subject":[],"published":{"date-parts":[[2023,9,26]]},"assertion":[{"value":"2021-10-16","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2022-05-16","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2023-09-26","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}