{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,9,4]],"date-time":"2025-09-04T14:18:34Z","timestamp":1756995514808,"version":"3.41.0"},"reference-count":46,"publisher":"Association for Computing Machinery (ACM)","issue":"5","license":[{"start":{"date-parts":[[2024,1,22]],"date-time":"2024-01-22T00:00:00Z","timestamp":1705881600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"name":"European Union under the Italian National Recovery and Resilience Plan (NRRP) of NextGenerationEU"},{"name":"Sustainable Mobility Center","award":["CN_00000023"],"award-info":[{"award-number":["CN_00000023"]}]},{"name":"Dottorati e contratti di ricerca su tematiche dell\u2019innovazione","award":["1062 on 10.08.2021"],"award-info":[{"award-number":["1062 on 10.08.2021"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Multimedia Comput. Commun. Appl."],"published-print":{"date-parts":[[2024,5,31]]},"abstract":"<jats:p>In this article, we propose a Hand Gesture Recognition (HGR) system based on a novel deep transformer (DT) neural network for media player control. The extracted hand skeleton features are processed by separate transformers for each finger in isolation to better identify the finger characteristics to drive the following classification. The achieved HGR accuracy (0.853) outperforms state-of-the-art HGR approaches when tested on the popular NVIDIA dataset. Moreover, we conducted a subjective assessment involving 30 people to evaluate the Quality of Experience (QoE) provided by the proposed DT-HGR for controlling a media player application compared with two traditional input devices, i.e., mouse and keyboard. The assessment participants were asked to evaluate objective (accuracy) and subjective (physical fatigue, usability, pragmatic quality, and hedonic quality) measurements. We found that (i) the accuracy of DT-HGR is very high (91.67%), only slightly lower than that of traditional alternative interaction modalities; and that (ii) the perceived quality for DT-HGR in terms of satisfaction, comfort, and interactivity is very high, with an average Mean Opinion Score (MOS) value as high as 4.4, whereas the alternative approaches did not reach 3.8, which encourages a more pervasive adoption of the natural gesture interaction.<\/jats:p>","DOI":"10.1145\/3638560","type":"journal-article","created":{"date-parts":[[2023,12,25]],"date-time":"2023-12-25T11:39:17Z","timestamp":1703504357000},"page":"1-22","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":5,"title":["Controlling Media Player with Hands: A Transformer Approach and a Quality of Experience Assessment"],"prefix":"10.1145","volume":"20","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-8745-1327","authenticated-orcid":false,"given":"Alessandro","family":"Floris","sequence":"first","affiliation":[{"name":"DIEE, University of Cagliari, Italy and CNIT, University of Cagliari, Italy"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-0792-1200","authenticated-orcid":false,"given":"Simone","family":"Porcu","sequence":"additional","affiliation":[{"name":"DIEE, University of Cagliari, Italy and CNIT, University of Cagliari, Italy"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-1350-3574","authenticated-orcid":false,"given":"Luigi","family":"Atzori","sequence":"additional","affiliation":[{"name":"DIEE, University of Cagliari, Italy and CNIT, University of Cagliari, Italy"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2024,1,22]]},"reference":[{"key":"e_1_3_1_2_2","doi-asserted-by":"crossref","first-page":"1165","DOI":"10.1109\/CVPR.2019.00126","volume-title":"2019 IEEE Conference on Computer Vision and Pattern Recognition (CVPR)","author":"Abavisani Mahdi","year":"2019","unstructured":"Mahdi Abavisani, Hamid Reza Vaezi Joze, and Vishal M. Patel. 2019. Improving the performance of unimodal dynamic hand-gesture recognition with multimodal training. In 2019 IEEE Conference on Computer Vision and Pattern Recognition (CVPR). 1165\u20131174."},{"key":"e_1_3_1_3_2","doi-asserted-by":"publisher","DOI":"10.1109\/NCVPRIPG.2015.7490026"},{"key":"e_1_3_1_4_2","doi-asserted-by":"publisher","DOI":"10.1109\/CICT.2013.6558300"},{"key":"e_1_3_1_5_2","doi-asserted-by":"publisher","DOI":"10.1109\/TMM.2018.2856094"},{"key":"e_1_3_1_6_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.patcog.2022.108762"},{"issue":"3","key":"e_1_3_1_7_2","first-page":"1","article-title":"Quality of experience for unified communications: A survey","volume":"30","author":"Husi\u0107 Jasmina Barakovi\u0107","year":"2019","unstructured":"Jasmina Barakovi\u0107 Husi\u0107, Sabina Barakovi\u0107, Enida Cero, Nina Slamnik, Merima O\u0107uz, Azer Dedovi\u0107, and Osman Zup\u010di\u0107. 2019. Quality of experience for unified communications: A survey. International Journal of Network Management 30, 3 (2019), 1\u201325.","journal-title":"International Journal of Network Management"},{"key":"e_1_3_1_8_2","doi-asserted-by":"publisher","DOI":"10.1145\/2071536.2071540"},{"key":"e_1_3_1_9_2","volume-title":"CVPR Workshop on Computer Vision for Augmented and Virtual Reality","author":"Bazarevsky Valentin","year":"2019","unstructured":"Valentin Bazarevsky, Yury Kartynnik, Andrey Vakunov, Karthik Raveendran, and Matthias Grundmann. 2019. BlazeFace: Sub-millisecond neural face detection on mobile GPUs. In CVPR Workshop on Computer Vision for Augmented and Virtual Reality."},{"key":"e_1_3_1_10_2","doi-asserted-by":"publisher","DOI":"10.1249\/00005768-198205000-00012"},{"key":"e_1_3_1_11_2","first-page":"1","volume-title":"Usability Evaluation In Industry (1st. ed.)","author":"Brooke J.","year":"1996","unstructured":"J. Brooke. 1996. Usability Evaluation In Industry (1st. ed.). CRC Press, Chapter SUS: A \u2018Quick and Dirty\u2019 Usability Scale, 1\u20136."},{"key":"e_1_3_1_12_2","doi-asserted-by":"publisher","DOI":"10.1145\/1873951.1874055"},{"key":"e_1_3_1_13_2","doi-asserted-by":"crossref","first-page":"623","DOI":"10.1109\/3DV50981.2020.00072","volume-title":"2020 International Conference on 3D Vision (3DV)","author":"D\u2019Eusanio Andrea","year":"2020","unstructured":"Andrea D\u2019Eusanio, Alessandro Simoni, Stefano Pini, Guido Borghi, Roberto Vezzani, and Rita Cucchiara. 2020. A transformer-based network for dynamic hand gesture recognition. In 2020 International Conference on 3D Vision (3DV). 623\u2013632."},{"key":"e_1_3_1_14_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.comnet.2022.108781"},{"issue":"18","key":"e_1_3_1_15_2","doi-asserted-by":"crossref","first-page":"17421","DOI":"10.1109\/JSEN.2021.3059685","article-title":"Dynamic hand gesture recognition based on 3D hand pose estimation for human\u2013robot interaction","volume":"22","author":"Gao Qing","year":"2022","unstructured":"Qing Gao, Yongquan Chen, Zhaojie Ju, and Yi Liang. 2022. Dynamic hand gesture recognition based on 3D hand pose estimation for human\u2013robot interaction. IEEE Sensors Journal 22, 18 (2022), 17421\u201317430.","journal-title":"IEEE Sensors Journal"},{"key":"e_1_3_1_16_2","doi-asserted-by":"crossref","first-page":"289","DOI":"10.1109\/3DV.2019.00040","volume-title":"2019 International Conference on 3D Vision (3DV)","author":"Gupta Vikram","year":"2019","unstructured":"Vikram Gupta, Sai Kumar Dwivedi, Rishabh Dabral, and Arjun Jain. 2019. Progression modelling for online and early gesture detection. In 2019 International Conference on 3D Vision (3DV). 289\u2013297."},{"key":"e_1_3_1_17_2","first-page":"1","volume-title":"2019 11th Int. Conf. on Electronics, Computers and Artificial Intelligence (ECAI)","author":"Iorga Cristian","year":"2019","unstructured":"Cristian Iorga and Victor-Emil Neagoe. 2019. A deep CNN approach with transfer learning for image recognition. In 2019 11th Int. Conf. on Electronics, Computers and Artificial Intelligence (ECAI). 1\u20136."},{"key":"e_1_3_1_18_2","unstructured":"ITU. 2008. Subjective video quality assessment methods for multimedia applications. Recommendation ITU-T P.910."},{"key":"e_1_3_1_19_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCONS.2018.8662901"},{"key":"e_1_3_1_20_2","doi-asserted-by":"publisher","DOI":"10.1109\/FG.2019.8756576"},{"key":"e_1_3_1_21_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-540-89350-9_6"},{"key":"e_1_3_1_22_2","doi-asserted-by":"publisher","DOI":"10.1002\/9781118706237.ch9"},{"key":"e_1_3_1_23_2","volume-title":"Qualinet White Paper on Definitions of Quality of Experience (2012)","author":"Callet Patrick Le","year":"2012","unstructured":"Patrick Le Callet, Sebastian M\u00f6ller, and Andrew Perkis. 2012. Qualinet White Paper on Definitions of Quality of Experience (2012). European Network on Quality of Experience in Multimedia Systems and Services (COST Action IC 1003), Lausanne, Switzerland, Version 1.2, March 2013."},{"key":"e_1_3_1_24_2","doi-asserted-by":"publisher","DOI":"10.3724\/SP.J.2096-5796.2018.0006"},{"key":"e_1_3_1_25_2","first-page":"1","volume-title":"2022 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR)","author":"Liu Shuying","year":"2022","unstructured":"Shuying Liu, Wenbin Wu, Jiaxian Wu, and Yue Lin. 2022. Spatial-temporal parallel transformer for arm-hand dynamic estimation. In 2022 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 1\u20136."},{"key":"e_1_3_1_26_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.456"},{"key":"e_1_3_1_27_2","doi-asserted-by":"crossref","first-page":"79","DOI":"10.1109\/MysuruCon52639.2021.9641567","volume-title":"2021 IEEE Mysore Sub Section International Conference (MysuruCon)","author":"Nagalapuram Gayathri Devi","year":"2021","unstructured":"Gayathri Devi Nagalapuram, S. Roopashree, D. Varshashree, D. Dheeraj, and Donal Jovian Nazareth. 2021. Controlling media player with hand gestures using convolutional neural network. In 2021 IEEE Mysore Sub Section International Conference (MysuruCon). 79\u201386."},{"key":"e_1_3_1_28_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.patcog.2017.10.033"},{"key":"e_1_3_1_29_2","doi-asserted-by":"publisher","DOI":"10.1109\/TITS.2014.2337331"},{"key":"e_1_3_1_30_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICAdTE.2013.6524715"},{"key":"e_1_3_1_31_2","doi-asserted-by":"publisher","DOI":"10.1109\/ACCESS.2021.3140175"},{"key":"e_1_3_1_32_2","doi-asserted-by":"publisher","DOI":"10.1109\/GEM.2018.8516520"},{"key":"e_1_3_1_33_2","doi-asserted-by":"crossref","first-page":"620","DOI":"10.1007\/978-3-319-58071-5_47","volume-title":"Human-Computer Interaction. User Interface Design, Development and Multimodality","author":"Pirker Johanna","year":"2017","unstructured":"Johanna Pirker, Mathias Pojer, Andreas Holzinger, and Christian G\u00fctl. 2017. Gesture-based interactions in video games with the leap motion controller. In Human-Computer Interaction. User Interface Design, Development and Multimodality, Masaaki Kurosu (Ed.). Springer International Publishing, Cham, 620\u2013633."},{"key":"e_1_3_1_34_2","doi-asserted-by":"publisher","DOI":"10.5120\/1495-2012"},{"key":"e_1_3_1_35_2","doi-asserted-by":"publisher","DOI":"10.1007\/s10462-012-9356-9"},{"key":"e_1_3_1_36_2","first-page":"4645","volume-title":"Proc. of the 30th IEEE Conf. on Computer Vision and Pattern Recognition (CVPR)","author":"Simon Tomas","year":"2017","unstructured":"Tomas Simon, Hanbyul Joo, Iain Matthews, and Yaser Sheikh. 2017. Hand keypoint detection in single images using multiview bootstrapping. In Proc. of the 30th IEEE Conf. on Computer Vision and Pattern Recognition (CVPR). 4645\u20134653."},{"key":"e_1_3_1_37_2","doi-asserted-by":"publisher","DOI":"10.5555\/2968826.2968890"},{"key":"e_1_3_1_38_2","first-page":"4489","volume-title":"Proc. of the 2015 IEEE Int. Conf. on Computer Vision (ICCV\u201915)","author":"Tran Du","year":"2015","unstructured":"Du Tran, Lubomir Bourdev, Rob Fergus, Lorenzo Torresani, and Manohar Paluri. 2015. Learning spatiotemporal features with 3D convolutional networks. In Proc. of the 2015 IEEE Int. Conf. on Computer Vision (ICCV\u201915). IEEE Computer Society, USA, 4489\u20134497."},{"key":"e_1_3_1_39_2","first-page":"36","volume-title":"HEALTHINFO 2018, The Third Int. Conf. on Informatics and Assistive Technologies for Health-Care, Medical Support and Wellbeing","author":"Trojaniello Diana","year":"2018","unstructured":"Diana Trojaniello, Alessia Cristiano, Stela Musteata, and Alberto Sanna. 2018. Evaluating real-time hand gesture recognition for automotive applications in elderly population: Cognitive load, user experience and usability degree. In HEALTHINFO 2018, The Third Int. Conf. on Informatics and Assistive Technologies for Health-Care, Medical Support and Wellbeing. 36\u201341."},{"key":"e_1_3_1_40_2","first-page":"2567","volume-title":"INTERSPEECH 2009, 10th Annual Conf. of the Int. Speech Communication Association","author":"Turunen Markku","year":"2009","unstructured":"Markku Turunen, Jaakko Hakulinen, Aleksi Melto, Tomi Heimonen, Tuuli Laivo, and Juho Hella. 2009. SUXES \u2014User experience evaluation method for spoken and multimodal interaction. In INTERSPEECH 2009, 10th Annual Conf. of the Int. Speech Communication Association. 2567\u20132570."},{"key":"e_1_3_1_41_2","doi-asserted-by":"crossref","first-page":"36","DOI":"10.1007\/978-3-642-34182-3_4","volume-title":"Gesture and Sign Language in Human-Computer Interaction and Embodied Communication","author":"Beurden Maurice H. P. H. van","year":"2012","unstructured":"Maurice H. P. H. van Beurden, Wijnand A. Ijsselsteijn, and Yvonne A. W. de Kort. 2012. User experience of gesture based interfaces: A comparison with traditional interaction methods on pragmatic and hedonic qualities. In Gesture and Sign Language in Human-Computer Interaction and Embodied Communication, Eleni Efthimiou, Georgios Kouroupetroglou, and Stavroula-Evita Fotinea (Eds.). Springer, Berlin, 36\u201347."},{"key":"e_1_3_1_42_2","article-title":"Attention is all you need","volume":"30","author":"Vaswani Ashish","year":"2017","unstructured":"Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, \u0141ukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. Advances in Neural Information Processing Systems 30 (2017).","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_1_43_2","first-page":"219","volume-title":"International Journal of Computer Vision 119","author":"Wang Heng","year":"2016","unstructured":"Heng Wang, Dan Oneata, Jakob Verbeek, and Cordelia Schmid. 2016. A robust and efficient video representation for action recognition. In International Journal of Computer Vision 119. MIT Press, 219\u2013238."},{"key":"e_1_3_1_44_2","doi-asserted-by":"publisher","DOI":"10.1007\/s12193-011-0088-y"},{"key":"e_1_3_1_45_2","doi-asserted-by":"crossref","first-page":"6469","DOI":"10.1109\/CVPR.2018.00677","volume-title":"2018 IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Yang Xiaodong","year":"2018","unstructured":"Xiaodong Yang, Pavlo Molchanov, and Jan Kautz. 2018. Making convolutional networks recurrent for visual sequence learning. In 2018 IEEE\/CVF Conference on Computer Vision and Pattern Recognition. 6469\u20136478."},{"key":"e_1_3_1_46_2","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2021.3087348"},{"key":"e_1_3_1_47_2","volume-title":"CVPR Workshop on Computer Vision for Augmented and Virtual Reality","author":"Zhang Fan","year":"2020","unstructured":"Fan Zhang, Valentin Bazarevsky, Andrey Vakunov, Andrei Tkachenka, George Sung, Chuo-Ling Chang, and Matthias Grundmann. 2020. MediaPipe hands: On-device real-time hand tracking. In CVPR Workshop on Computer Vision for Augmented and Virtual Reality."}],"container-title":["ACM Transactions on Multimedia Computing, Communications, and Applications"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3638560","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3638560","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T22:50:49Z","timestamp":1750287049000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3638560"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,1,22]]},"references-count":46,"journal-issue":{"issue":"5","published-print":{"date-parts":[[2024,5,31]]}},"alternative-id":["10.1145\/3638560"],"URL":"https:\/\/doi.org\/10.1145\/3638560","relation":{},"ISSN":["1551-6857","1551-6865"],"issn-type":[{"type":"print","value":"1551-6857"},{"type":"electronic","value":"1551-6865"}],"subject":[],"published":{"date-parts":[[2024,1,22]]},"assertion":[{"value":"2023-06-05","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2023-12-20","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2024-01-22","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}