{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,12]],"date-time":"2026-06-12T01:05:15Z","timestamp":1781226315197,"version":"3.54.1"},"reference-count":49,"publisher":"MDPI AG","issue":"24","license":[{"start":{"date-parts":[[2019,12,9]],"date-time":"2019-12-09T00:00:00Z","timestamp":1575849600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Sensors"],"abstract":"<jats:p>With the recent growth of Smart TV technology, the demand for unique and beneficial applications motivates the study of a unique gesture-based system for a smart TV-like environment. Combining movie recommendation, social media platform, call a friend application, weather updates, chatting app, and tourism platform into a single system regulated by natural-like gesture controller is proposed to allow the ease of use and natural interaction. Gesture recognition problem solving was designed through 24 gestures of 13 static and 11 dynamic gestures that suit to the environment. Dataset of a sequence of RGB and depth images were collected, preprocessed, and trained in the proposed deep learning architecture. Combination of three-dimensional Convolutional Neural Network (3DCNN) followed by Long Short-Term Memory (LSTM) model was used to extract the spatio-temporal features. At the end of the classification, Finite State Machine (FSM) communicates the model to control the class decision results based on application context. The result suggested the combination data of depth and RGB to hold 97.8% of accuracy rate on eight selected gestures, while the FSM has improved the recognition rate from 89% to 91% in a real-time performance.<\/jats:p>","DOI":"10.3390\/s19245429","type":"journal-article","created":{"date-parts":[[2019,12,9]],"date-time":"2019-12-09T11:22:51Z","timestamp":1575890571000},"page":"5429","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":65,"title":["Dynamic Hand Gesture Recognition Using 3DCNN and LSTM with FSM Context-Aware Model"],"prefix":"10.3390","volume":"19","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-2090-3608","authenticated-orcid":false,"given":"Noorkholis Luthfil","family":"Hakim","sequence":"first","affiliation":[{"name":"Department of Computer Science and Information Engineering, National Central University, Taoyuan 32001, Taiwan"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Timothy K.","family":"Shih","sequence":"additional","affiliation":[{"name":"Department of Computer Science and Information Engineering, National Central University, Taoyuan 32001, Taiwan"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Sandeli Priyanwada","family":"Kasthuri Arachchi","sequence":"additional","affiliation":[{"name":"Department of Computer Science and Information Engineering, National Central University, Taoyuan 32001, Taiwan"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Wisnu","family":"Aditya","sequence":"additional","affiliation":[{"name":"Department of Computer Science and Information Engineering, National Central University, Taoyuan 32001, Taiwan"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Yi-Cheng","family":"Chen","sequence":"additional","affiliation":[{"name":"Department of Information Management, National Central University, Taoyuan 32001, Taiwan"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-0401-8473","authenticated-orcid":false,"given":"Chih-Yang","family":"Lin","sequence":"additional","affiliation":[{"name":"Department of Electrical Engineering, Yuan Ze University, Taoyuan 32003, Taiwan"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1968","published-online":{"date-parts":[[2019,12,9]]},"reference":[{"key":"ref_1","unstructured":"Grant, H., and Lai, C.-K. (1998, January 13\u201316). Simulation modeling with artificial reality technology (SMART): An integration of virtual reality and simulation modeling. Proceedings of the 1998 Winter Simulation Conference, Washington, DC, USA."},{"key":"ref_2","unstructured":"Guo, Z. (2011, January 19\u201320). Research of hand positioning and gesture recognition based on binocular vision. Proceedings of the 2011 IEEE International Symposium on VR Innovation, Singapore."},{"key":"ref_3","unstructured":"Lee, S.-H., Sohn, M.-K., Kim, D.-J., Kim, B., and Kim, H. (2013, January 11\u201314). Smart TV interaction system using face and hand gesture recognition. Proceedings of the 2013 IEEE International Conference on Consumer Electronics (ICCE), Las Vegas, NV, USA."},{"key":"ref_4","doi-asserted-by":"crossref","first-page":"543","DOI":"10.1109\/ACCESS.2015.2432679","article-title":"An elicitation study on gesture preferences and memorability toward a practical hand-gesture vocabulary for smart televisions","volume":"3","author":"Dong","year":"2015","journal-title":"IEEE Access"},{"key":"ref_5","unstructured":"Huang, J., Zhou, W., Li, H., and Li, W. (July, January 29). Sign language recognition using 3d convolutional neural networks. Proceedings of the 2015 IEEE International Conference on Multimedia and Expo (ICME), Turin, Italy."},{"key":"ref_6","doi-asserted-by":"crossref","unstructured":"Cui, R., Liu, H., and Zhang, C. (2017, January 21\u201326). Recurrent convolutional neural networks for continuous sign language recognition by staged optimization. Proceedings of the 2017 IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.175"},{"key":"ref_7","doi-asserted-by":"crossref","first-page":"355","DOI":"10.1016\/j.ergon.2017.02.004","article-title":"Gesture recognition for human-robot collaboration: A review","volume":"68","author":"Liu","year":"2018","journal-title":"Int. J. Ind. Ergon."},{"key":"ref_8","doi-asserted-by":"crossref","first-page":"583","DOI":"10.1007\/s10846-014-0039-4","article-title":"Online dynamic gesture recognition for human robot interaction","volume":"77","author":"Xu","year":"2015","journal-title":"J. Intell. Robot. Syst."},{"key":"ref_9","first-page":"438","article-title":"Virtual guitar: Using real-time finger tracking for musical instruments","volume":"18","author":"Hakim","year":"2019","journal-title":"Int. J. Comput. Sci. Eng."},{"key":"ref_10","doi-asserted-by":"crossref","first-page":"7019","DOI":"10.1109\/ACCESS.2017.2788558","article-title":"Real-time continuous detection and recognition of subject-specific smart TV gestures via fusion of depth and inertial sensing","volume":"6","author":"Dawar","year":"2018","journal-title":"IEEE Access"},{"key":"ref_11","doi-asserted-by":"crossref","first-page":"30","DOI":"10.1109\/38.250916","article-title":"A survey of glove-based input","volume":"14","author":"Sturman","year":"1994","journal-title":"IEEE Comput. Graph. Appl."},{"key":"ref_12","doi-asserted-by":"crossref","first-page":"63","DOI":"10.1145\/1531326.1531369","article-title":"Real-time hand-tracking with a color glove","volume":"28","author":"Wang","year":"2009","journal-title":"ACM Trans. Graph."},{"key":"ref_13","doi-asserted-by":"crossref","unstructured":"Mummadi, C., Leo, F., Verma, K., Kasireddy, S., Scholl, P., Kempfle, J., and Laerhoven, K. (2018). Real-Time and Embedded Detection of Hand Gestures with an IMU-Based Glove. Informatics, 5.","DOI":"10.3390\/informatics5020028"},{"key":"ref_14","doi-asserted-by":"crossref","unstructured":"Georgi, M., Amma, C., and Schultz, T. (2015, January 12\u201315). Recognizing Hand and Finger Gestures with IMU based Motion and EMG based Muscle Activity Sensing. Proceedings of the 2015 International Joint Conference on Biomedical Engineering Systems and Technologies, Lisbon, Portugal.","DOI":"10.5220\/0005276900990108"},{"key":"ref_15","doi-asserted-by":"crossref","unstructured":"Arachchi, S.K., Hakim, N.L., Hsu, H.-H., Klimenko, S.V., and Shih, T.K. (2018, January 16\u201318). Real-time static and dynamic gesture recognition using mixed space features for 3D virtual world\u2019s interactions. Proceedings of the 2018 32nd International Conference on Advanced Information Networking and Applications Workshops (WAINA), Cracow, Poland.","DOI":"10.1109\/WAINA.2018.00157"},{"key":"ref_16","doi-asserted-by":"crossref","first-page":"1659","DOI":"10.1109\/TCSVT.2015.2469551","article-title":"Survey on 3D hand gesture recognition","volume":"26","author":"Cheng","year":"2015","journal-title":"IEEE Trans. Circuits Syst. Video Technol."},{"key":"ref_17","doi-asserted-by":"crossref","first-page":"3941","DOI":"10.1007\/s00521-016-2294-8","article-title":"Deep learning in vision-based static hand gesture recognition","volume":"28","author":"Oyedotun","year":"2017","journal-title":"Neural Comput. Appl."},{"key":"ref_18","doi-asserted-by":"crossref","unstructured":"Molchanov, P., Gupta, S., Kim, K., and Kautz, J. (2015, January 7\u201312). Hand gesture recognition with 3D convolutional neural networks. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition Workshops, Boston, MA, USA.","DOI":"10.1109\/CVPRW.2015.7301342"},{"key":"ref_19","unstructured":"Ren, Z., Meng, J., Yuan, J., and Zhang, Z. (December, January 28). Robust hand gesture recognition with kinect sensor. Proceedings of the 19th ACM International Conference on Multimedia, Scottsdale, AZ, USA."},{"key":"ref_20","doi-asserted-by":"crossref","first-page":"1110","DOI":"10.1109\/TMM.2013.2246148","article-title":"Robust part-based hand gesture recognition using kinect sensor","volume":"15","author":"Ren","year":"2013","journal-title":"IEEE Trans. Multimed."},{"key":"ref_21","unstructured":"Li, Y. (2012, January 22\u201324). Hand gesture recognition using Kinect. Proceedings of the 2012 IEEE International Conference on Computer Science and Automation Engineering, Beijing, China."},{"key":"ref_22","doi-asserted-by":"crossref","unstructured":"Melax, S., Keselman, L., and Orsten, S. (2013, January 29\u201331). Dynamics based 3D skeletal hand tracking. Proceedings of the Graphics Interface 2013, Regina, SK, Canada.","DOI":"10.1145\/2448196.2448232"},{"key":"ref_23","doi-asserted-by":"crossref","first-page":"97","DOI":"10.1016\/j.jvcir.2015.01.015","article-title":"A hand gesture recognition technique for human-computer interaction","volume":"28","year":"2015","journal-title":"J. Vis. Commun. Image Represent."},{"key":"ref_24","doi-asserted-by":"crossref","first-page":"305","DOI":"10.1109\/TIM.2015.2498560","article-title":"Static and dynamic hand gesture recognition in depth data using dynamic time warping","volume":"65","author":"Plouffe","year":"2015","journal-title":"IEEE Trans. Instrum. Meas."},{"key":"ref_25","doi-asserted-by":"crossref","unstructured":"Wu, Y.-K., Wang, H.-C., Chang, L.-C., and Li, K.-C. (2013). Using HMMs and depth information for signer-independent sign language recognition. Multi-Disciplinary Trends in Artificial Intelligence, Springer.","DOI":"10.1007\/978-3-642-44949-9_8"},{"key":"ref_26","doi-asserted-by":"crossref","unstructured":"Chen, Q., Georganas, N.D., and Petriu, E.M. (2007, January 1\u20133). Real-time vision-based hand gesture recognition using haar-like features. Proceedings of the 2007 IEEE Instrumentation & Measurement Technology Conference IMTC 2007, Warsaw, Poland.","DOI":"10.1109\/IMTC.2007.379068"},{"key":"ref_27","doi-asserted-by":"crossref","first-page":"60","DOI":"10.1145\/1897816.1897838","article-title":"Vision-based hand-gesture applications","volume":"54","author":"Wachs","year":"2011","journal-title":"Commun. ACM"},{"key":"ref_28","unstructured":"Ren, Z., Meng, J., and Yuan, J. (2011, January 13\u201316). Depth camera based hand gesture recognition and its applications in human-computer-interaction. Proceedings of the 2011 8th International Conference on Information, Communications & Signal Processing, Singapore."},{"key":"ref_29","doi-asserted-by":"crossref","unstructured":"Yang, C., Jang, Y., Beh, J., Han, D., and Ko, H. (2012, January 13\u201316). Gesture recognition using depth-based hand tracking for contactless controller application. Proceedings of the 2012 IEEE International Conference on Consumer Electronics (ICCE), Las Vegas, NV, USA.","DOI":"10.1109\/ICCE.2012.6161876"},{"key":"ref_30","doi-asserted-by":"crossref","first-page":"29","DOI":"10.1016\/j.cviu.2015.05.010","article-title":"Histogram of 3D facets: A depth descriptor for human action and hand gesture recognition","volume":"139","author":"Zhang","year":"2015","journal-title":"Comput. Vis. Image Underst."},{"key":"ref_31","doi-asserted-by":"crossref","unstructured":"Escobedo, E., and Camara, G. (2016, January 4\u20137). A new approach for dynamic gesture recognition using skeleton trajectory representation and histograms of cumulative magnitudes. Proceedings of the 2016 29th SIBGRAPI Conference on Graphics, Patterns and Images (SIBGRAPI), Sao Paulo, Brazil.","DOI":"10.1109\/SIBGRAPI.2016.037"},{"key":"ref_32","unstructured":"Simonyan, K., and Zisserman, A. (2014, January 8\u201313). Two-stream convolutional networks for action recognition in videos. Proceedings of the 27th International Conference on Neural Information Processing Systems, Montreal, QC, Canada."},{"key":"ref_33","doi-asserted-by":"crossref","unstructured":"Wang, L., Qiao, Y., and Tang, X. (2015, January 7\u201312). Action recognition with trajectory-pooled deep-convolutional descriptors. Proceedings of the 2015 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Boston, MA, USA.","DOI":"10.1109\/CVPR.2015.7299059"},{"key":"ref_34","unstructured":"Feichtenhofer, C., Pinz, A., and Wildes, R. (2016, January 5\u201310). Spatiotemporal residual networks for video action recognition. Proceedings of the 30th International Conference on Neural Information Processing Systems, Barcelona, Spain."},{"key":"ref_35","doi-asserted-by":"crossref","unstructured":"Shahroudy, A., Liu, J., Ng, T.-T., and Wang, G. (2016, January 27\u201330). Ntu rgb+ d: A large scale dataset for 3d human activity analysis. Proceedings of the 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.115"},{"key":"ref_36","doi-asserted-by":"crossref","unstructured":"Singh, B., Marks, T.K., Jones, M., Tuzel, O., and Shao, M. (2016, January 27\u201330). A multi-stream bi-directional recurrent neural network for fine-grained action detection. Proceedings of the 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.216"},{"key":"ref_37","doi-asserted-by":"crossref","first-page":"430","DOI":"10.1007\/s11263-016-0957-7","article-title":"Beyond temporal pooling: Recurrence and temporal convolutions for gesture recognition in video","volume":"126","author":"Pigou","year":"2018","journal-title":"Int. J. Comput. Vis."},{"key":"ref_38","unstructured":"Du, Y., Wang, W., and Wang, L. (2015, January 7\u201312). Hierarchical recurrent neural network for skeleton based action recognition. Proceedings of the 2015 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Boston, MA, USA."},{"key":"ref_39","doi-asserted-by":"crossref","unstructured":"Veeriah, V., Zhuang, N., and Qi, G.-J. (2015, January 7\u201313). Differential recurrent neural networks for action recognition. Proceedings of the 2015 IEEE International Conference on Computer Vision (ICCV\u201915), Santiago, Chile.","DOI":"10.1109\/ICCV.2015.460"},{"key":"ref_40","doi-asserted-by":"crossref","first-page":"84","DOI":"10.1145\/3065386","article-title":"Imagenet classification with deep convolutional neural networks","volume":"60","author":"Krizhevsky","year":"2017","journal-title":"Commun. ACM"},{"key":"ref_41","unstructured":"Pigou, L., Dieleman, S., Kindermans, P.-J., and Schrauwen, B. (2014). Sign Language Recognition Using Convolutional Neural Networks, Springer."},{"key":"ref_42","doi-asserted-by":"crossref","unstructured":"Tran, D., Bourdev, L., Fergus, R., Torresani, L., and Paluri, M. (2015, January 7\u201313). Learning spatiotemporal features with 3d convolutional networks. Proceedings of the 2015 IEEE International Conference on Computer Vision (ICCV), Santiago, Chile.","DOI":"10.1109\/ICCV.2015.510"},{"key":"ref_43","doi-asserted-by":"crossref","unstructured":"Molchanov, P., Yang, X., Gupta, S., Kim, K., Tyree, S., and Kautz, J. (2016, January 27\u201330). Online detection and classification of dynamic hand gestures with recurrent 3d convolutional neural network. Proceedings of the 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.456"},{"key":"ref_44","doi-asserted-by":"crossref","first-page":"101","DOI":"10.1049\/ip-vis:19941058","article-title":"Visual gesture recognition","volume":"141","author":"Davis","year":"1994","journal-title":"IEE Proc. Vis. Image Signal Process"},{"key":"ref_45","doi-asserted-by":"crossref","first-page":"1805","DOI":"10.1016\/S0031-3203(99)00175-2","article-title":"Visual understanding of dynamic hand gestures","volume":"33","author":"Yeasin","year":"2000","journal-title":"Pattern Recognit."},{"key":"ref_46","unstructured":"Hong, P., Turk, M., and Huang, T.S. (2000, January 28\u201330). Gesture modeling and recognition using finite state machines. Proceedings of the Fourth IEEE International Conference on Automatic Face and Gesture Recognition, Grenoble, France."},{"key":"ref_47","unstructured":"Wan, J., Zhao, Y., Zhou, S., Guyon, I., Escalera, S., and Li, S.Z. (July, January 26). Chalearn looking at people RGB-D isolated and continuous datasets for gesture recognition. Proceedings of the 2006 IEEE Conference on Computer Vision and Pattern Recognition Workshops, Las Vegas, NV, USA."},{"key":"ref_48","unstructured":"Kovac, J., Peer, P., and Solina, F. (2003, January 22\u201324). Human skin color clustering for face detection. Proceedings of the IEEE Region 8 EUROCON 2003, Computer as a Tool, Ljubljana, Slovenia."},{"key":"ref_49","unstructured":"(2019, November 08). Data Science Bootcamp. Available online: https:\/\/medium.com\/data-science-bootcamp\/understand-the-softmax-function-in-minutes-f3a59641e86d."}],"container-title":["Sensors"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1424-8220\/19\/24\/5429\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,11]],"date-time":"2025-10-11T13:40:47Z","timestamp":1760190047000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1424-8220\/19\/24\/5429"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2019,12,9]]},"references-count":49,"journal-issue":{"issue":"24","published-online":{"date-parts":[[2019,12]]}},"alternative-id":["s19245429"],"URL":"https:\/\/doi.org\/10.3390\/s19245429","relation":{},"ISSN":["1424-8220"],"issn-type":[{"value":"1424-8220","type":"electronic"}],"subject":[],"published":{"date-parts":[[2019,12,9]]}}}