{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,1]],"date-time":"2026-05-01T10:15:42Z","timestamp":1777630542037,"version":"3.51.4"},"reference-count":47,"publisher":"MDPI AG","issue":"3","license":[{"start":{"date-parts":[[2023,3,21]],"date-time":"2023-03-21T00:00:00Z","timestamp":1679356800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["62201438"],"award-info":[{"award-number":["62201438"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Symmetry"],"abstract":"<jats:p>Imagining recognition of behaviors from video sequences for a machine is full of challenges but meaningful. This work aims to predict students\u2019 behavior in an experimental class, which relies on the symmetry idea from reality to annotated reality centered on the feature space. A heteromorphic ensemble algorithm is proposed to make the obtained features more aggregated and reduce the computational burden. Namely, the deep learning models are improved to obtain feature vectors representing gestures from video frames and the classification algorithm is optimized for behavior recognition. So, the symmetric idea is realized by decomposing the task into three schemas including hand detection and cropping, hand joints feature extraction, and gesture classification. Firstly, a new detector method named YOLOv4-specific tiny detection (STD) is proposed by reconstituting the YOLOv4-tiny model, which could produce two outputs with some attention mechanism leveraging context information. Secondly, the efficient pyramid squeeze attention (EPSA) net is integrated into EvoNorm-S0 and the spatial pyramid pool (SPP) layer to obtain the hand joint position information. Lastly, the D\u2013S theory is used to fuse two classifiers, support vector machine (SVM) and random forest (RF), to produce a mixed classifier named S\u2013R. Eventually, the synergetic effects of our algorithm are shown by experiments on self-created datasets with a high average recognition accuracy of 89.6%.<\/jats:p>","DOI":"10.3390\/sym15030769","type":"journal-article","created":{"date-parts":[[2023,3,21]],"date-time":"2023-03-21T04:48:01Z","timestamp":1679374081000},"page":"769","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":1,"title":["A Novel Heteromorphic Ensemble Algorithm for Hand Pose Recognition"],"prefix":"10.3390","volume":"15","author":[{"given":"Shiruo","family":"Liu","sequence":"first","affiliation":[{"name":"School of Electronic Engineering, Xidian University, Xi\u2019an 710071, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-7895-1899","authenticated-orcid":false,"given":"Xiaoguang","family":"Yuan","sequence":"additional","affiliation":[{"name":"School of Electronic Engineering, Xidian University, Xi\u2019an 710071, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Wei","family":"Feng","sequence":"additional","affiliation":[{"name":"School of Electronic Engineering, Xidian University, Xi\u2019an 710071, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-1129-5601","authenticated-orcid":false,"given":"Aifeng","family":"Ren","sequence":"additional","affiliation":[{"name":"School of Electronic Engineering, Xidian University, Xi\u2019an 710071, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Zhenyong","family":"Hu","sequence":"additional","affiliation":[{"name":"School of Electronic Engineering, Xidian University, Xi\u2019an 710071, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Zuheng","family":"Ming","sequence":"additional","affiliation":[{"name":"Laboraory of Information Processing and Transmission, L2TI, Institut Galil\u00e9e, University Paris XIII, 93079 Villetaneuse, France"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Adnan","family":"Zahid","sequence":"additional","affiliation":[{"name":"School of Engineering, University of Glasgow, Glasgow G12 8QQ, UK"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-7097-9969","authenticated-orcid":false,"given":"Qammer","family":"Abbasi","sequence":"additional","affiliation":[{"name":"School of Engineering, University of Glasgow, Glasgow G12 8QQ, UK"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Shuo","family":"Wang","sequence":"additional","affiliation":[{"name":"School of Electronic Engineering, Xidian University, Xi\u2019an 710071, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2023,3,21]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","unstructured":"Swindells, C., Quinn, K.I., Dill, J., and Tory, M.K. (2002, January 27\u201330). That one there! Pointing to establish device identity. Proceedings of the ACM Symposium on User Interface Software and Technology, Paris, France.","DOI":"10.1145\/571985.572007"},{"key":"ref_2","doi-asserted-by":"crossref","unstructured":"Nickel, K., and Stiefelhagen, R. (2003, January 5\u20137). Pointing gesture recognition based on 3D-tracking of face, hands and head orientation. Proceedings of the International Conference on Multimodal Interaction, Vancouver, BC, Canada.","DOI":"10.1145\/958432.958460"},{"key":"ref_3","doi-asserted-by":"crossref","unstructured":"Goza, S.M., Ambrose, R.O., Diftler, M.A., and Spain, I.M. (2004, January 24\u201329). Telepresence Control of the NASA\/DARPA Robonaut on a Mobility Platform. Proceedings of the CHI 2004 Conference on Human Factors in Computing Systems, Vienna, Austria.","DOI":"10.1145\/985692.985771"},{"key":"ref_4","doi-asserted-by":"crossref","first-page":"825","DOI":"10.1109\/TRA.2003.817093","article-title":"FAce MOUSe: A novel human-machine interface for controlling the position of a laparoscope","volume":"19","author":"Nishikawa","year":"2003","journal-title":"IEEE Trans. Robot. Autom."},{"key":"ref_5","doi-asserted-by":"crossref","first-page":"302","DOI":"10.1086\/502200","article-title":"Bacterial Contamination of Computer Keyboards in a Teaching Hospital","volume":"24","author":"Schultz","year":"2003","journal-title":"Infect. Control Hosp. Epidemiol."},{"key":"ref_6","doi-asserted-by":"crossref","first-page":"461","DOI":"10.1109\/TSMCC.2008.923862","article-title":"A Survey of Glove-Based Systems and Their Applications","volume":"38","author":"Dipietro","year":"2008","journal-title":"IEEE Trans. Syst. Man Cybern. Part C"},{"key":"ref_7","doi-asserted-by":"crossref","first-page":"173","DOI":"10.1016\/j.mejo.2018.01.014","article-title":"Wearable technologies for hand joints monitoring for rehabilitation: A survey","volume":"88","author":"Rashid","year":"2019","journal-title":"Microelectron. J."},{"key":"ref_8","doi-asserted-by":"crossref","unstructured":"Chen, W., Yu, C., Tu, C., Lyu, Z., Tang, J., Ou, S., Fu, Y., and Xue, Z. (2020). A Survey on Hand Pose Estimation with Wearable Sensors and Computer-Vision-Based Methods. Sensors, 20.","DOI":"10.3390\/s20041074"},{"key":"ref_9","doi-asserted-by":"crossref","first-page":"28121","DOI":"10.1007\/s11042-018-5971-z","article-title":"A systematic literature review on vision based gesture recognition techniques","volume":"77","author":"Ahmad","year":"2018","journal-title":"Multimed. Tools. Appl."},{"key":"ref_10","doi-asserted-by":"crossref","first-page":"2368","DOI":"10.1109\/TITS.2014.2337331","article-title":"Hand Gesture Recognition in Real Time for Automotive Interfaces: A Multimodal Vision-Based Approach and Evaluations","volume":"15","author":"Trivedi","year":"2014","journal-title":"IEEE Trans. Intell. Transp. Syst."},{"key":"ref_11","doi-asserted-by":"crossref","unstructured":"Devineau, G., Moutarde, F., Xi, W., and Yang, J. (2018, January 15\u201319). Deep Learning for Hand Gesture Recognition on Skeletal Data. Proceedings of the 2018 13th IEEE International Conference on Automatic Face & Gesture Recognition (FG 2018), Xi\u2019an, China.","DOI":"10.1109\/FG.2018.00025"},{"key":"ref_12","doi-asserted-by":"crossref","unstructured":"Liu, J., Liu, Y., Wang, Y., Prinet, V., Xiang, S., and Pan, C. (2020, January 13\u201319). Decoupled Representation Learning for Skeleton-Based Gesture Recognition. Proceedings of the 2020 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA.","DOI":"10.1109\/CVPR42600.2020.00579"},{"key":"ref_13","doi-asserted-by":"crossref","first-page":"137","DOI":"10.1023\/B:VISI.0000013087.49260.fb","article-title":"Robust real-time face detection","volume":"57","author":"Viola","year":"2004","journal-title":"Int. J. Comput. Vis."},{"key":"ref_14","doi-asserted-by":"crossref","first-page":"1904","DOI":"10.1109\/TPAMI.2015.2389824","article-title":"Spatial Pyramid Pooling in Deep Convolutional Networks for Visual Recognition","volume":"37","author":"He","year":"2015","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_15","doi-asserted-by":"crossref","first-page":"114499","DOI":"10.1016\/j.eswa.2020.114499","article-title":"Selective spatiotemporal features learning for dynamic gesture recognition","volume":"169","author":"Tang","year":"2021","journal-title":"Expert Syst. Appl."},{"key":"ref_16","unstructured":"Rajput, D.S., Reddy, T.S.K., and Raju, D.N. (2018). Deep Learning and Neural Networks, IGI Global."},{"key":"ref_17","unstructured":"Bochkovskiy, A., Wang, C.Y., and Liao, H.M. (2020). YOLOv4: Optimal Speed and Accuracy of Object Detection. arXiv."},{"key":"ref_18","unstructured":"Simonyan, K., and Zisserman, A. (2014). Very Deep Convolutional Networks for Large-Scale Image Recognition. arXiv."},{"key":"ref_19","doi-asserted-by":"crossref","unstructured":"He, K., Zhang, X., Ren, S., and Sun, J. (2016, January 27\u201330). Deep Residual Learning for Image Recognition. Proceedings of the 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.90"},{"key":"ref_20","unstructured":"Vaswani, A., Shazeer, N.M., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A.N., Kaiser, L., and Polosukhin, I. (2017, January 4\u20139). Attention is All you Need. Proceedings of the NIPS, Long Beach, CA, USA."},{"key":"ref_21","doi-asserted-by":"crossref","first-page":"2011","DOI":"10.1109\/TPAMI.2019.2913372","article-title":"Squeeze-and-Excitation Networks","volume":"42","author":"Hu","year":"2020","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_22","unstructured":"Zhang, H., Zu, K., Lu, J., Zou, Y., and Meng, D. (2021). EPSANet: An Efficient Pyramid Split Attention Block on Convolutional Neural Network. arXiv."},{"key":"ref_23","doi-asserted-by":"crossref","unstructured":"Haroon, M., Altaf, S., Ahmad, S., Zaindin, M., Huda, S., and Iqbal, S. (2022). Hand Gesture Recognition with Symmetric Pattern under Diverse Illuminated Conditions Using Artificial Neural Network. Symmetry, 14.","DOI":"10.3390\/sym14102045"},{"key":"ref_24","doi-asserted-by":"crossref","first-page":"15803","DOI":"10.1007\/s11042-020-10446-y","article-title":"Techno-regulation and intelligent safeguards","volume":"80","author":"Zaccagnino","year":"2021","journal-title":"Multimed. Tools Appl."},{"key":"ref_25","doi-asserted-by":"crossref","first-page":"119614","DOI":"10.1016\/j.eswa.2023.119614","article-title":"Touchscreen gestures as images. A transfer learning approach for soft biometric traits recognition","volume":"219","author":"Guarino","year":"2023","journal-title":"Expert Syst. Appl."},{"key":"ref_26","doi-asserted-by":"crossref","unstructured":"Hussain, S., Saxena, R., Han, X., Khan, J.A., and Shin, H. (2017, January 5\u20138). Hand gesture recognition using deep learning. Proceedings of the 2017 International SoC Design Conference (ISOCC), Seoul, Republic of Korea.","DOI":"10.1109\/ISOCC.2017.8368821"},{"key":"ref_27","doi-asserted-by":"crossref","first-page":"1670","DOI":"10.3390\/sym7041670","article-title":"Application of Assistive Computer Vision Methods to Oyama Karate Techniques Recognition","volume":"7","author":"Hachaj","year":"2015","journal-title":"Symmetry"},{"key":"ref_28","doi-asserted-by":"crossref","unstructured":"Khan, M.S., and Zualkernan, I.A. (2020, January 19\u201321). Using Convolutional Neural Networks for Smart Classroom Observation. Proceedings of the 2020 International Conference on Artificial Intelligence in Information and Communication (ICAIIC), Fukuoka, Japan.","DOI":"10.1109\/ICAIIC48513.2020.9065260"},{"key":"ref_29","doi-asserted-by":"crossref","unstructured":"Ren, X., and Yang, D. (2021, January 20\u201322). Student Behavior Detection Based on YOLOv4-Bi. Proceedings of the 2021 IEEE International Conference on Computer Science, Artificial Intelligence and Electronic Engineering (CSAIEE), Online.","DOI":"10.1109\/CSAIEE54046.2021.9543310"},{"key":"ref_30","doi-asserted-by":"crossref","first-page":"101","DOI":"10.1016\/j.patrec.2013.10.010","article-title":"Combining multiple depth-based descriptors for hand gesture recognition","volume":"50","author":"Dominio","year":"2014","journal-title":"Pattern Recognit. Lett."},{"key":"ref_31","doi-asserted-by":"crossref","first-page":"283","DOI":"10.1016\/j.ijleo.2017.11.158","article-title":"Light Invariant Real-Time Robust Hand Gesture Recognition","volume":"159","author":"Chaudhary","year":"2018","journal-title":"Optik"},{"key":"ref_32","doi-asserted-by":"crossref","first-page":"889","DOI":"10.1007\/s00138-018-0969-0","article-title":"Abnormal gesture recognition based on multi-model fusion strategy","volume":"30","author":"Lin","year":"2018","journal-title":"Mach. Vis. Appl."},{"key":"ref_33","doi-asserted-by":"crossref","unstructured":"Zhang, Y.C. (2018, January 26\u201329). Gesture Recognition System Based on Improved Stacked Hourglass Structure. Proceedings of the 2018 International Conference on Computer, Communications and Mechatronics Engineering (CCME 2018), Cuernavaca, Mexico.","DOI":"10.12783\/dtcse\/ccme2018\/28570"},{"key":"ref_34","doi-asserted-by":"crossref","unstructured":"Zhang, Z., Wu, B., and Jiang, Y. (2022, January 15\u201317). Gesture Recognition System Based on Improved YOLO v3. Proceedings of the 2022 7th International Conference on Intelligent Computing and Signal Processing (ICSP), Xi\u2019an, China.","DOI":"10.1109\/ICSP54964.2022.9778394"},{"key":"ref_35","unstructured":"Redmon, J., and Farhadi, A. (2018). YOLOv3: An Incremental Improvement. arXiv."},{"key":"ref_36","doi-asserted-by":"crossref","first-page":"2278","DOI":"10.1109\/5.726791","article-title":"Gradient-based learning applied to document recognition","volume":"86","author":"Lecun","year":"1998","journal-title":"Proc. IEEE"},{"key":"ref_37","doi-asserted-by":"crossref","unstructured":"Szegedy, C., Liu, W., Jia, Y., Sermanet, P., Reed, S.E., Anguelov, D., Erhan, D., Vanhoucke, V., and Rabinovich, A. (2015, January 7\u201312). Going Deeper with Convolutions. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Boston, MA, USA.","DOI":"10.1109\/CVPR.2015.7298594"},{"key":"ref_38","unstructured":"Howard, A.G., Zhu, M., Chen, B., Kalenichenko, D., Wang, W., Weyand, T., Andreetto, M., and Adam, H. (2017). MobileNets: Efficient Convolutional Neural Networks for Mobile Vision Applications. arXiv."},{"key":"ref_39","doi-asserted-by":"crossref","unstructured":"Wang, C.Y., Bochkovskiy, A., and Liao, H.y. (2021, January 20\u201325). Scaled-YOLOv4: Scaling Cross Stage Partial Network. Proceedings of the 2021 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Nashville, TN, USA.","DOI":"10.1109\/CVPR46437.2021.01283"},{"key":"ref_40","unstructured":"Liu, H., Brock, A., Simonyan, K., and Le, Q.V. (2020). Evolving Normalization-Activation Layers. arXiv."},{"key":"ref_41","doi-asserted-by":"crossref","unstructured":"Seeland, M., and M\u00e4der, P. (2021). Multi-view classification with convolutional neural networks. PLoS ONE, 16.","DOI":"10.1371\/journal.pone.0245230"},{"key":"ref_42","doi-asserted-by":"crossref","unstructured":"Simon, T., Joo, H., Matthews, I.A., and Sheikh, Y. (2017, January 21\u201326). Hand Keypoint Detection in Single Images using Multiview Bootstrapping. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.494"},{"key":"ref_43","doi-asserted-by":"crossref","first-page":"292","DOI":"10.1109\/JBHI.2019.2909688","article-title":"TSE-CNN: A Two-Stage End-to-End CNN for Human Activity Recognition","volume":"24","author":"Huang","year":"2020","journal-title":"IEEE J. Biomed. Health. Inf."},{"key":"ref_44","first-page":"3133","article-title":"Do We Need Hundreds of Classifiers to Solve Real World Classification Problems?","volume":"15","author":"Cernadas","year":"2014","journal-title":"J. Mach. Learn. Res."},{"key":"ref_45","doi-asserted-by":"crossref","first-page":"1345","DOI":"10.3390\/s110201345","article-title":"A Trust Evaluation Algorithm for Wireless Sensor Networks Based on Node Behaviors and D-S Evidence Theory","volume":"11","author":"Feng","year":"2011","journal-title":"Sensors"},{"key":"ref_46","doi-asserted-by":"crossref","unstructured":"Narasimhaswamy, S., Wei, Z., Wang, Y., Zhang, J., and Nguyen, M.H. (November, January 27). Contextual Attention for Hand Detection in the Wild. Proceedings of the 2019 IEEE\/CVF International Conference on Computer Vision (ICCV), Seoul, Republic of Korea.","DOI":"10.1109\/ICCV.2019.00966"},{"key":"ref_47","doi-asserted-by":"crossref","first-page":"25","DOI":"10.1016\/j.imavis.2018.12.001","article-title":"Large-scale multiview 3D hand pose dataset","volume":"81","author":"Cazorla","year":"2019","journal-title":"Image Vis. Comput."}],"container-title":["Symmetry"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/2073-8994\/15\/3\/769\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,10]],"date-time":"2025-10-10T18:59:32Z","timestamp":1760122772000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/2073-8994\/15\/3\/769"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,3,21]]},"references-count":47,"journal-issue":{"issue":"3","published-online":{"date-parts":[[2023,3]]}},"alternative-id":["sym15030769"],"URL":"https:\/\/doi.org\/10.3390\/sym15030769","relation":{},"ISSN":["2073-8994"],"issn-type":[{"value":"2073-8994","type":"electronic"}],"subject":[],"published":{"date-parts":[[2023,3,21]]}}}