{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,10,18]],"date-time":"2025-10-18T20:58:57Z","timestamp":1760821137399,"version":"build-2065373602"},"reference-count":45,"publisher":"MDPI AG","issue":"2","license":[{"start":{"date-parts":[[2020,1,18]],"date-time":"2020-01-18T00:00:00Z","timestamp":1579305600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Sensors"],"abstract":"<jats:p>Gesture spotting is an essential task for recognizing finger gestures used to control in-car touchless interfaces. Automated methods to achieve this task require to detect video segments where gestures are observed, to discard natural behaviors of users\u2019 hands that may look as target gestures, and be able to work online. In this paper, we address these challenges with a recurrent neural architecture for online finger gesture spotting. We propose a multi-stream network merging hand and hand-location features, which help to discriminate target gestures from natural movements of the hand, since these may not happen in the same 3D spatial location. Our multi-stream recurrent neural network (RNN) recurrently learns semantic information, allowing to spot gestures online in long untrimmed video sequences. In order to validate our method, we collect a finger gesture dataset in an in-vehicle scenario of an autonomous car. 226 videos with more than 2100 continuous instances were captured with a depth sensor. On this dataset, our gesture spotting approach outperforms state-of-the-art methods with an improvement of about 10% and 15% of recall and precision, respectively. Furthermore, we demonstrated that by combining with an existing gesture classifier (a 3D Convolutional Neural Network), our proposal achieves better performance than previous hand gesture recognition methods.<\/jats:p>","DOI":"10.3390\/s20020528","type":"journal-article","created":{"date-parts":[[2020,1,20]],"date-time":"2020-01-20T04:27:09Z","timestamp":1579494429000},"page":"528","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":11,"title":["Finger Gesture Spotting from Long Sequences Based on Multi-Stream Recurrent Neural Networks"],"prefix":"10.3390","volume":"20","author":[{"ORCID":"https:\/\/orcid.org\/0000-0003-4945-8314","authenticated-orcid":false,"given":"Gibran","family":"Benitez-Garcia","sequence":"first","affiliation":[{"name":"Toyota Technological Institute, Nagoya 468-8511, Japan"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-8868-8094","authenticated-orcid":false,"given":"Muhammad","family":"Haris","sequence":"additional","affiliation":[{"name":"Toyota Technological Institute, Nagoya 468-8511, Japan"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yoshiyuki","family":"Tsuda","sequence":"additional","affiliation":[{"name":"DENSO CORPORATION, Kariya 448-8661, Japan"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Norimichi","family":"Ukita","sequence":"additional","affiliation":[{"name":"Toyota Technological Institute, Nagoya 468-8511, Japan"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2020,1,18]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","unstructured":"Kendon, A. (1980). Gesticulation and Speech: Two Aspects of the Process of Utterance. The Relationship of Verbal and Nonverbal Communication, Mouton.","DOI":"10.1515\/9783110813098.207"},{"key":"ref_2","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1007\/s10462-012-9356-9","article-title":"Vision based hand gesture recognition for human computer interaction: A survey","volume":"43","author":"Rautaray","year":"2015","journal-title":"Artif. Intell. Rev."},{"key":"ref_3","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1016\/j.cviu.2016.09.001","article-title":"Computer vision for assistive technologies","volume":"154","author":"Leo","year":"2017","journal-title":"Comput. Vis. Image Underst."},{"key":"ref_4","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1007\/s10055-016-0293-9","article-title":"Industry use of virtual reality in product design and manufacturing: A survey","volume":"21","author":"Berg","year":"2017","journal-title":"Virtual Real."},{"key":"ref_5","unstructured":"Pickering, C.A., Burnham, K.J., and Richardson, M.J. (2007, January 28\u201329). A research study of hand gesture recognition technologies and applications for human vehicle interaction. Proceedings of the 3rd Institution of Engineering and Technology Conference on Automotive Electronics, Warwick, UK."},{"key":"ref_6","doi-asserted-by":"crossref","first-page":"136","DOI":"10.1177\/0018720809336542","article-title":"Skill acquisition while operating in-vehicle information systems: Interface design determines the level of safety-relevant distractions","volume":"51","author":"Jahn","year":"2009","journal-title":"Hum. Factors"},{"key":"ref_7","doi-asserted-by":"crossref","unstructured":"Parada-Loira, F., Gonz\u00e1lez-Agulla, E., and Alba-Castro, J.L. (2014, January 8\u201311). Hand gestures to control infotainment equipment in cars. Proceedings of the 2014 IEEE Intelligent Vehicles Symposium, Dearborn, MI, USA.","DOI":"10.1109\/IVS.2014.6856614"},{"key":"ref_8","doi-asserted-by":"crossref","unstructured":"Zengeler, N., Kopinski, T., and Handmann, U. (2019). Hand gesture recognition in automotive human\u2013machine interaction using depth cameras. Sensors, 19.","DOI":"10.3390\/s19010059"},{"key":"ref_9","doi-asserted-by":"crossref","first-page":"634","DOI":"10.25046\/aj020381","article-title":"Augmented Reality Prototype HUD for Passenger Infotainment in a Vehicular Environment","volume":"2","author":"Wang","year":"2017","journal-title":"Adv. Sci. Technol. Eng. Syst. J."},{"key":"ref_10","doi-asserted-by":"crossref","unstructured":"Wang, S., Charissis, V., Lagoo, R., Campbell, J., and Harrison, D.K. (2019, January 11\u201313). Reducing Driver Distraction by Utilizing Augmented Reality Head-Up Display System for Rear Passengers. Proceedings of the 2019 IEEE International Conference on Consumer Electronics (ICCE), Las Vegas, NV, USA.","DOI":"10.1109\/ICCE.2019.8661927"},{"key":"ref_11","doi-asserted-by":"crossref","unstructured":"Charissis, V., and Naef, M. (2007, January 13\u201315). Evaluation of prototype automotive head-up display interface: Testing driver\u2019s focusing ability through a VR simulation. Proceedings of the 2007 IEEE Intelligent Vehicles Symposium, Istanbul, Turkey.","DOI":"10.1109\/IVS.2007.4290174"},{"key":"ref_12","doi-asserted-by":"crossref","unstructured":"Wang, P., Li, W., Liu, S., Gao, Z., Tang, C., and Ogunbona, P. (2016, January 4\u20138). Large-scale isolated gesture recognition using convolutional neural networks. Proceedings of the 23rd International Conference on Pattern Recognition (ICPR), Canc\u00fan, Mexico.","DOI":"10.1109\/ICPR.2016.7899599"},{"key":"ref_13","doi-asserted-by":"crossref","unstructured":"Miao, Q., Li, Y., Ouyang, W., Ma, Z., Xu, X., Shi, W., and Cao, X. (2017, January 22\u201329). Multimodal gesture recognition based on the resc3d network. Proceedings of the IEEE International Conference on Computer Vision (ICCV), Venice, Italy.","DOI":"10.1109\/ICCVW.2017.360"},{"key":"ref_14","doi-asserted-by":"crossref","first-page":"4517","DOI":"10.1109\/ACCESS.2017.2684186","article-title":"Multimodal gesture recognition using 3-D convolution and convolutional LSTM","volume":"5","author":"Zhu","year":"2017","journal-title":"IEEE Access"},{"key":"ref_15","doi-asserted-by":"crossref","unstructured":"Narayana, P., Beveridge, R., and Draper, B.A. (2018, January 18\u201322). Gesture recognition: Focus on the hands. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00549"},{"key":"ref_16","doi-asserted-by":"crossref","unstructured":"Roitberg, A., Pollert, T., Haurilet, M., Martin, M., and Stiefelhagen, R. (2019, January 16\u201320). Analysis of Deep Fusion Strategies for Multi-modal Gesture Recognition. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), Long Beach, CA, USA.","DOI":"10.1109\/CVPRW.2019.00029"},{"key":"ref_17","doi-asserted-by":"crossref","first-page":"1011","DOI":"10.1109\/TMM.2018.2869278","article-title":"Continuous Gesture Segmentation and Recognition using 3DCNN and Convolutional LSTM","volume":"21","author":"Zhu","year":"2018","journal-title":"IEEE Trans. Multimed."},{"key":"ref_18","doi-asserted-by":"crossref","unstructured":"Narayana, P., Beveridge, J.R., and Draper, B. (2019, January 14\u201319). Continuous Gesture Recognition through Selective Temporal Fusion. Proceedings of the International Joint Conference on Neural Networks (IJCNN), Budapest, Hungary.","DOI":"10.1109\/IJCNN.2019.8852385"},{"key":"ref_19","doi-asserted-by":"crossref","unstructured":"Asadi-Aghbolaghi, M., Clap\u00e9s, A., Bellantonio, M., Escalante, H.J., Ponce-L\u00f3pez, V., Bar\u00f3, X., Guyon, I., Kasaei, S., and Escalera, S. (2017). Deep learning for action and gesture recognition in image sequences: A survey. Gesture Recognition, Springer.","DOI":"10.1007\/978-3-319-57021-1_19"},{"key":"ref_20","unstructured":"Becattini, F., Uricchio, T., Seidenari, L., Del Bimbo, A., and Ballan, L. (2018, January 8\u201314). Am I Done? Predicting Action Progress in Videos. Proceedings of the European Conference on Computer Vision Workshops (ECCVW), Munich, Germany."},{"key":"ref_21","doi-asserted-by":"crossref","unstructured":"Zolfaghari, M., Singh, K., and Brox, T. (2018, January 8\u201314). ECO: Efficient Convolutional Network for Online Video Understanding. Proceedings of the European Conference on Computer Vision (ECCV), Munich, Germany.","DOI":"10.1007\/978-3-030-01216-8_43"},{"key":"ref_22","doi-asserted-by":"crossref","unstructured":"Lin, T., Zhao, X., Su, H., Wang, C., and Yang, M. (2018, January 8\u201314). Bsn: Boundary sensitive network for temporal action proposal generation. Proceedings of the European Conference on Computer Vision (ECCV), Munich, Germany.","DOI":"10.1007\/978-3-030-01225-0_1"},{"key":"ref_23","unstructured":"Lin, T., Liu, X., Li, X., Ding, E., and Wen, S. (November, January 27). BMN: Boundary-Matching Network for Temporal Action Proposal Generation. Proceedings of the IEEE International Conference on Computer Vision (ICCV), Seoul, South Korea."},{"key":"ref_24","doi-asserted-by":"crossref","unstructured":"Long, F., Yao, T., Qiu, Z., Tian, X., Luo, J., and Mei, T. (2019, January 16\u201320). Gaussian Temporal Awareness Networks for Action Localization. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.00043"},{"key":"ref_25","doi-asserted-by":"crossref","unstructured":"Escorcia, V., Heilbron, F.C., Niebles, J.C., and Ghanem, B. (2016, January 8\u201316). Daps: Deep action proposals for action understanding. Proceedings of the European Conference on Computer Vision (ECCV), Amsterdam, The Netherlands.","DOI":"10.1007\/978-3-319-46487-9_47"},{"key":"ref_26","doi-asserted-by":"crossref","unstructured":"Buch, S., Escorcia, V., Shen, C., Ghanem, B., and Niebles, J.C. (2017, January 21\u201326). Sst: Single-stream temporal action proposals. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.675"},{"key":"ref_27","doi-asserted-by":"crossref","unstructured":"Narayana, P., Beveridge, J.R., and Draper, B. (2019, January 14\u201319). Analyzing Multi-Channel Networks for Gesture Recognition. Proceedings of the International Joint Conference on Neural Networks (IJCNN), Budapest, Hungary.","DOI":"10.1109\/IJCNN.2019.8851991"},{"key":"ref_28","doi-asserted-by":"crossref","unstructured":"Liu, Z., Chai, X., Liu, Z., and Chen, X. (2017, January 22\u201329). Continuous gesture recognition with hand-oriented spatiotemporal feature. Proceedings of the IEEE International Conference on Computer Vision Workshop (ICCVW), Venice, Italy.","DOI":"10.1109\/ICCVW.2017.361"},{"key":"ref_29","unstructured":"Liu, R., Lehman, J., Molino, P., Such, F.P., Frank, E., Sergeev, A., and Yosinski, J. (2018, January 3\u20138). An intriguing failing of convolutional neural networks and the coordconv solution. Proceedings of the Advances in Neural Information Processing Systems (NeurIPS), Montr\u00e9al, QC, Canada."},{"key":"ref_30","unstructured":"Ren, S., He, K., Girshick, R., and Sun, J. (2015, January 7\u201312). Faster R-CNN: Towards real-time object detection with region proposal networks. Proceedings of the Advances in Neural Information Processing Systems (NeurIPS), Montr\u00e9al, QC, Canada."},{"key":"ref_31","unstructured":"Karpathy, A., Johnson, J., and Li, F.-F. (2015). Visualizing and understanding recurrent networks. arXiv."},{"key":"ref_32","doi-asserted-by":"crossref","unstructured":"Karpathy, A., Toderici, G., Shetty, S., Leung, T., Sukthankar, R., and Li, F.-F. (2014, January 23\u201328). Large-scale Video Classification with Convolutional Neural Networks. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Columbus, OH, USA.","DOI":"10.1109\/CVPR.2014.223"},{"key":"ref_33","unstructured":"Simonyan, K., and Zisserman, A. (2014, January 8\u201313). Two-stream convolutional networks for action recognition in videos. Proceedings of the Advances in Neural Information Processing Systems (NeurIPS), Montreal, QC, Canada."},{"key":"ref_34","doi-asserted-by":"crossref","unstructured":"Hong, J., Cho, B., Hong, Y.W., and Byun, H. (2019). Contextual Action Cues from Camera Sensor for Multi-Stream Action Recognition. Sensors, 19.","DOI":"10.3390\/s19061382"},{"key":"ref_35","doi-asserted-by":"crossref","unstructured":"Chen, X., Wang, G., Guo, H., Zhang, C., Wang, H., and Zhang, L. (2019). MFA-Net: Motion Feature Augmented Network for Dynamic Hand Gesture Recognition from Skeletal Data. Sensors, 19.","DOI":"10.3390\/s19020239"},{"key":"ref_36","doi-asserted-by":"crossref","unstructured":"Wan, J., Escalera, S., Anbarjafari, G., Escalante, H.J., Bar\u00f3, X., Guyon, I., Madadi, M., Allik, J., Gorbova, J., and Lin, C. (2017, January 22\u201329). Results and Analysis of ChaLearn LAP Multi-modal Isolated and Continuous Gesture Recognition, and Real Versus Fake Expressed Emotions Challenges. Proceedings of the IEEE International Conference on Computer Vision Workshop (ICCVW), Venice, Italy.","DOI":"10.1109\/ICCVW.2017.377"},{"key":"ref_37","doi-asserted-by":"crossref","unstructured":"Wang, H., Wang, P., Song, Z., and Li, W. (2017, January 22\u201329). Large-Scale Multimodal Gesture Segmentation and Recognition Based on Convolutional Neural Networks. Proceedings of the IEEE International Conference on Computer Vision Workshop (ICCVW), Venice, Italy.","DOI":"10.1109\/ICCVW.2017.371"},{"key":"ref_38","doi-asserted-by":"crossref","unstructured":"Benitez-Garcia, G., Haris, M., Tsuda, Y., and Ukita, N. (2019, January 27\u201331). Similar Finger Gesture Recognition using Triplet-loss Networks. Proceedings of the Sixteenth IAPR International Conference on Machine Vision Applications (MVA), Tokyo, Japan.","DOI":"10.23919\/MVA.2019.8757973"},{"key":"ref_39","doi-asserted-by":"crossref","unstructured":"K\u00f6p\u00fckl\u00fc, O., Gunduz, A., K\u00f6se, N., and Rigoll, G. (2019, January 14\u201318). Real-time Hand Gesture Detection and Classification Using Convolutional Neural Networks. Proceedings of the 14th IEEE International Conference on Automatic Face & Gesture Recognition (FG), Lille, France.","DOI":"10.1109\/FG.2019.8756576"},{"key":"ref_40","doi-asserted-by":"crossref","unstructured":"He, K., Zhang, X., Ren, S., and Sun, J. (2016, January 27\u201330). Deep residual learning for image recognition. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.90"},{"key":"ref_41","unstructured":"Chung, J., Gulcehre, C., Cho, K., and Bengio, Y. (2014). Empirical evaluation of gated recurrent neural networks on sequence modeling. arXiv."},{"key":"ref_42","doi-asserted-by":"crossref","first-page":"303","DOI":"10.1007\/s11263-009-0275-4","article-title":"The pascal visual object classes (voc) challenge","volume":"88","author":"Everingham","year":"2010","journal-title":"Int. J. Comput. Vis."},{"key":"ref_43","doi-asserted-by":"crossref","unstructured":"Alwassel, H., Caba Heilbron, F., Escorcia, V., and Ghanem, B. (2018, January 8\u201314). Diagnosing error in temporal action detectors. Proceedings of the European Conference on Computer Vision (ECCV), Munich, Germany.","DOI":"10.1007\/978-3-030-01219-9_16"},{"key":"ref_44","doi-asserted-by":"crossref","unstructured":"Hara, K., Kataoka, H., and Satoh, Y. (2018, January 18\u201322). Can spatiotemporal 3d cnns retrace the history of 2d cnns and imagenet?. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00685"},{"key":"ref_45","unstructured":"Chao, P., Kao, C.Y., Ruan, Y.S., Huang, C.H., and Lin, Y.L. (November, January 27). Hardnet: A low memory traffic network. Proceedings of the IEEE International Conference on Computer Vision (ICCV), Seoul, Korea."}],"container-title":["Sensors"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1424-8220\/20\/2\/528\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,13]],"date-time":"2025-10-13T13:29:45Z","timestamp":1760362185000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1424-8220\/20\/2\/528"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2020,1,18]]},"references-count":45,"journal-issue":{"issue":"2","published-online":{"date-parts":[[2020,1]]}},"alternative-id":["s20020528"],"URL":"https:\/\/doi.org\/10.3390\/s20020528","relation":{},"ISSN":["1424-8220"],"issn-type":[{"type":"electronic","value":"1424-8220"}],"subject":[],"published":{"date-parts":[[2020,1,18]]}}}