{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,22]],"date-time":"2026-07-22T16:06:58Z","timestamp":1784736418605,"version":"3.55.0"},"reference-count":41,"publisher":"MDPI AG","issue":"19","license":[{"start":{"date-parts":[[2020,10,1]],"date-time":"2020-10-01T00:00:00Z","timestamp":1601510400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Sensors"],"abstract":"<jats:p>The quality of recognition systems for continuous utterances in signed languages could be largely advanced within the last years. However, research efforts often do not address specific linguistic features of signed languages, as e.g., non-manual expressions. In this work, we evaluate the potential of a single video camera-based recognition system with respect to the latter. For this, we introduce a two-stage pipeline based on two-dimensional body joint positions extracted from RGB camera data. The system first separates the data flow of a signed expression into meaningful word segments on the base of a frame-wise binary Random Forest. Next, every segment is transformed into image-like shape and classified with a Convolutional Neural Network. The proposed system is then evaluated on a data set of continuous sentence expressions in Japanese Sign Language with a variation of non-manual expressions. Exploring multiple variations of data representations and network parameters, we are able to distinguish word segments of specific non-manual intonations with 86% accuracy from the underlying body joint movement data. Full sentence predictions achieve a total Word Error Rate of 15.75%. This marks an improvement of 13.22% as compared to ground truth predictions obtained from labeling insensitive towards non-manual content. Consequently, our analysis constitutes an important contribution for a better understanding of mixed manual and non-manual content in signed communication.<\/jats:p>","DOI":"10.3390\/s20195621","type":"journal-article","created":{"date-parts":[[2020,10,1]],"date-time":"2020-10-01T09:04:12Z","timestamp":1601543052000},"page":"5621","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":24,"title":["Recognition of Non-Manual Content in Continuous Japanese Sign Language"],"prefix":"10.3390","volume":"20","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-9530-4714","authenticated-orcid":false,"given":"Heike","family":"Brock","sequence":"first","affiliation":[{"name":"Honda Research Institute Japan Co., Ltd., Wako-shi, Saitama 351-0188, Japan"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Iva","family":"Farag","sequence":"additional","affiliation":[{"name":"Faculty of Sciences and Engineering, Saarland University, 66123 Saarbr\u00fccken, Germany"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-6134-4558","authenticated-orcid":false,"given":"Kazuhiro","family":"Nakadai","sequence":"additional","affiliation":[{"name":"Honda Research Institute Japan Co., Ltd., Wako-shi, Saitama 351-0188, Japan"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1968","published-online":{"date-parts":[[2020,10,1]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","unstructured":"Kawas, S., Karalis, G., Wen, T., and Ladner, R.E. (2016, January 23\u201326). Improving Real-Time Captioning Experiences for Deaf and Hard of Hearing Students. Proceedings of the 18th International ACM SIGACCESS Conference on Computers and Accessibility, Reno, NV, USA.","DOI":"10.1145\/2982142.2982164"},{"key":"ref_2","doi-asserted-by":"crossref","first-page":"1418","DOI":"10.1109\/TCYB.2013.2265337","article-title":"Discriminative exemplar coding for sign language recognition with Kinect","volume":"43","author":"Sun","year":"2013","journal-title":"IEEE Trans. Cybern."},{"key":"ref_3","doi-asserted-by":"crossref","first-page":"21","DOI":"10.1145\/2735952","article-title":"A real-time hand posture recognition system using deep neural networks","volume":"6","author":"Tang","year":"2015","journal-title":"ACM Trans. Intell. Syst. Technol."},{"key":"ref_4","doi-asserted-by":"crossref","unstructured":"Chuan, C., Regina, E., and Guardino, C. (2014, January 3\u20135). American Sign Language Recognition Using Leap Motion Sensor. Proceedings of the 2014 13th International Conference on Machine Learning and Applications, Detroit, MI, USA.","DOI":"10.1109\/ICMLA.2014.110"},{"key":"ref_5","doi-asserted-by":"crossref","first-page":"1860","DOI":"10.1109\/LSP.2018.2877891","article-title":"Three-Dimensional Sign Language Recognition With Angular Velocity Maps and Connived Feature ResNet","volume":"25","author":"Kumar","year":"2018","journal-title":"IEEE Signal Process. Lett."},{"key":"ref_6","doi-asserted-by":"crossref","unstructured":"Pigou, L., Dieleman, S., Kindermans, P.J., and Schrauwen, B. (2015). Sign language recognition using convolutional neural networks. Lecture Notes in Computer Science, Springer.","DOI":"10.1007\/978-3-319-16178-5_40"},{"key":"ref_7","unstructured":"Dong, C., Leu, M.C., and Yin, Z. (2015, January 7\u201312). American Sign Language alphabet recognition using Microsoft Kinect. Proceedings of the 2015 IEEE Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), Boston, MA, USA."},{"key":"ref_8","doi-asserted-by":"crossref","unstructured":"Huenerfauth, M., and Hanson, V. (2009). Sign language in the interface: Access for deaf signers. Universal Access Handbook, Erlbaum.","DOI":"10.1201\/9781420064995-c38"},{"key":"ref_9","doi-asserted-by":"crossref","unstructured":"Cao, Z., Hidalgo, G., Simon, T., Wei, S.E., and Sheikh, Y. (2018). OpenPose: Realtime Multi-Person 2D Pose Estimation using Part Affinity Fields. arXiv.","DOI":"10.1109\/CVPR.2017.143"},{"key":"ref_10","doi-asserted-by":"crossref","unstructured":"Cao, Z., Simon, T., Wei, S.E., and Sheikh, Y. (2017, January 21\u201326). Realtime multi-person 2d pose estimation using part affinity fields. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.143"},{"key":"ref_11","doi-asserted-by":"crossref","unstructured":"Simon, T., Joo, H., Matthews, I., and Sheikh, Y. (2017, January 21\u201326). Hand keypoint detection in single images using multiview bootstrapping. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.494"},{"key":"ref_12","doi-asserted-by":"crossref","unstructured":"Wei, S.E., Ramakrishna, V., Kanade, T., and Sheikh, Y. (2016, January 27\u201330). Convolutional pose machines. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.511"},{"key":"ref_13","doi-asserted-by":"crossref","first-page":"1061","DOI":"10.1109\/TPAMI.2002.1023803","article-title":"Extraction of 2d motion trajectories and its application to hand gesture recognition","volume":"24","author":"Yang","year":"2002","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_14","doi-asserted-by":"crossref","unstructured":"Kong, W., and Ranganath, S. (2008, January 17\u201319). Automatic hand trajectory segmentation and phoneme transcription for sign language. Proceedings of the 2008 8th IEEE International Conference on Automatic Face & Gesture Recognition, Amsterdam, The Netherlands.","DOI":"10.1109\/AFGR.2008.4813462"},{"key":"ref_15","unstructured":"Forster, J., Schmidt, C., Koller, O., Bellgardt, M., and Ney, H. (2014, January 26\u201331). Extensions of the Sign Language Recognition and Translation Corpus RWTH-PHOENIX-Weather. Proceedings of the LREC, Reykjavik, Iceland."},{"key":"ref_16","doi-asserted-by":"crossref","unstructured":"Koller, O., Zargaran, O., Ney, H., and Bowden, R. (2016, January 19\u201322). Deep Sign: Hybrid CNN-HMM for continuous sign language recognition. Proceedings of the British Machine Vision Conference 2016, York, UK.","DOI":"10.5244\/C.30.136"},{"key":"ref_17","doi-asserted-by":"crossref","first-page":"1311","DOI":"10.1007\/s11263-018-1121-3","article-title":"Deep Sign: Enabling Robust Statistical Continuous Sign Language Recognition via Hybrid CNN-HMMs","volume":"126","author":"Koller","year":"2018","journal-title":"Int. J. Comput. Vis."},{"key":"ref_18","unstructured":"Von Agris, U., and Kraiss, K.F. (2007, January 23\u201325). Towards a video corpus for signer-independent continuous sign language recognition. Proceedings of the Gesture in Human-Computer Interaction and Simulation, Lisbon, Portugal."},{"key":"ref_19","doi-asserted-by":"crossref","unstructured":"Camgoz, N.C., Hadfield, S., Koller, O., and Bowden, R. (2017, January 22\u201329). Subunets: End-to-end hand shape and continuous sign language recognition. Proceedings of the 2017 IEEE International Conference on Computer Vision (ICCV), Venice, Italy.","DOI":"10.1109\/ICCV.2017.332"},{"key":"ref_20","doi-asserted-by":"crossref","unstructured":"Cui, R., Liu, H., and Zhang, C. (2017, January 21\u201326). Recurrent convolutional neural networks for continuous sign language recognition by staged optimization. Proceedings of the 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.175"},{"key":"ref_21","doi-asserted-by":"crossref","unstructured":"Huang, J., Zhou, W., Zhang, Q., Li, H., and Li, W. (2018, January 2\u20137). Video-based sign language recognition without temporal segmentation. Proceedings of the 32nd AAAI Conference on Artificial Intelligence, New Orleans, LA, USA.","DOI":"10.1609\/aaai.v32i1.11903"},{"key":"ref_22","doi-asserted-by":"crossref","unstructured":"Cihan Camgoz, N., Hadfield, S., Koller, O., Ney, H., and Bowden, R. (2018, January 18\u201323). Neural sign language translation. Proceedings of the 2018 IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00812"},{"key":"ref_23","doi-asserted-by":"crossref","first-page":"1880","DOI":"10.1109\/TMM.2018.2889563","article-title":"A deep neural framework for continuous sign language recognition by iterative training","volume":"21","author":"Cui","year":"2019","journal-title":"IEEE Trans. Multimed."},{"key":"ref_24","first-page":"399","article-title":"Weakly Supervised Learning with Multi-Stream CNN-LSTM-HMMs to Discover Sequential Parallelism in Sign Language Videos","volume":"42","author":"Koller","year":"2019","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_25","doi-asserted-by":"crossref","first-page":"91170","DOI":"10.1109\/ACCESS.2020.2993650","article-title":"Continuous Sign Language Recognition Through Cross-Modal Alignment of Video and Text Embeddings in a Joint-Latent Space","volume":"8","author":"Papastratis","year":"2020","journal-title":"IEEE Access"},{"key":"ref_26","doi-asserted-by":"crossref","unstructured":"Ye, Y., Tian, Y., Huenerfauth, M., and Liu, J. (2018, January 18\u201322). Recognizing American Sign Language Gestures from within Continuous Videos. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition Workshops, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPRW.2018.00280"},{"key":"ref_27","unstructured":"Brock, H., and Nakadai, K. (2018, January 7\u201312). Deep JSLC: A Multimodal Corpus Collection for Data-driven Generation of Japanese Sign Language Expressions. Proceedings of the 11th International Conference on Language Resources and Evaluation, Miyazaki, Japan."},{"key":"ref_28","unstructured":"Su, S.F., and Tai, J.H. (2009). Lexical comparison of signs from Taiwan, Chinese, Japanese, and American Sign Languages: Taking iconicity into account. Taiwan Sign Language and Beyond; The Taiwan Institute for the Humanities, National Chung Cheng University."},{"key":"ref_29","doi-asserted-by":"crossref","first-page":"193","DOI":"10.1016\/0147-1767(90)90005-H","article-title":"Apologies: Japanese and American styles","volume":"14","author":"Barnlund","year":"1990","journal-title":"Int. J. Intercult. Relat."},{"key":"ref_30","unstructured":"Sagawa, H., and Takeuchi, M. (2000, January 28\u201330). A method for recognizing a sequence of sign language words represented in a japanese sign language sentence. Proceedings of the Fourth IEEE International Conference on Automatic Face and Gesture Recognition, Washington, DC, USA."},{"key":"ref_31","doi-asserted-by":"crossref","unstructured":"Balayn, A., Brock, H., and Nakadai, K. (2018, January 27\u201331). Data-driven development of Virtual Sign Language Communication Agents. Proceedings of the 2018 27th IEEE International Symposium on Robot and Human Interactive Communication, Nanjing, China.","DOI":"10.1109\/ROMAN.2018.8525717"},{"key":"ref_32","unstructured":"Joze, H.R.V., and Koller, O. (2018). Ms-asl: A large-scale data set and benchmark for understanding american sign language. arXiv."},{"key":"ref_33","doi-asserted-by":"crossref","unstructured":"Farag, I., and Brock, H. (2019, January 12\u201317). Learning Motion Disfluencies for Automatic Sign Language Segmentation. Proceedings of the 2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Brighton, UK.","DOI":"10.1109\/ICASSP.2019.8683523"},{"key":"ref_34","doi-asserted-by":"crossref","first-page":"645","DOI":"10.1109\/LSP.2018.2817179","article-title":"Training CNNs for 3-D sign language recognition with color texture coded joint angular displacement maps","volume":"25","author":"Kumar","year":"2018","journal-title":"IEEE Signal Process. Lett."},{"key":"ref_35","doi-asserted-by":"crossref","first-page":"624","DOI":"10.1109\/LSP.2017.2678539","article-title":"Joint distance maps based action recognition with convolutional neural networks","volume":"24","author":"Li","year":"2017","journal-title":"IEEE Signal Process. Lett."},{"key":"ref_36","doi-asserted-by":"crossref","first-page":"731","DOI":"10.1109\/LSP.2017.2690339","article-title":"SkeletonNet: Mining deep part features for 3-D action recognition","volume":"24","author":"Ke","year":"2017","journal-title":"IEEE Signal Process. Lett."},{"key":"ref_37","unstructured":"McDonald, J., Wolfe, R., Wilbur, R.B., Moncrief, R., Malaia, E., Fujimoto, S., Baowidan, S., and Stec, J. (2016, January 23\u201328). A new tool to facilitate prosodic analysis of motion capture data and a data-driven technique for the improvement of avatar motion. Proceedings of the 7th Workshop on the Representation and Processing of Sign Languages: Corpus Mining Language Resources and Evaluation Conference (LREC), Portoroz, Slovenia."},{"key":"ref_38","doi-asserted-by":"crossref","unstructured":"M\u00fcnzner, S., Schmidt, P., Reiss, A., Hanselmann, M., Stiefelhagen, R., and D\u00fcrichen, R. (2017, January 13\u201315). CNN-based sensor fusion techniques for multimodal human activity recognition. Proceedings of the 2017 ACM International Symposium on Wearable Computers, Maui, HI, USA.","DOI":"10.1145\/3123021.3123046"},{"key":"ref_39","doi-asserted-by":"crossref","unstructured":"Ha, S., Yun, J.M., and Choi, S. (2015, January 9\u201312). Multi-modal convolutional neural networks for activity recognition. Proceedings of the 2015 IEEE International conference on systems, man, and cybernetics, Hong Kong, China.","DOI":"10.1109\/SMC.2015.525"},{"key":"ref_40","doi-asserted-by":"crossref","unstructured":"Lin, H.I., Hsu, M.H., and Chen, W.K. (2014, January 18\u201322). Human hand gesture recognition using a convolution neural network. Proceedings of the 2014 IEEE International Conference on Automation Science and Engineering (CASE), New Taipei, Taiwan.","DOI":"10.1109\/CoASE.2014.6899454"},{"key":"ref_41","doi-asserted-by":"crossref","first-page":"169","DOI":"10.1109\/LSP.2018.2883864","article-title":"S3DRGF: Spatial 3-D Relational Geometric Features for 3-D Sign Language Representation and Recognition","volume":"26","author":"Kumar","year":"2019","journal-title":"IEEE Signal Process. Lett."}],"container-title":["Sensors"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1424-8220\/20\/19\/5621\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,11]],"date-time":"2025-10-11T10:15:38Z","timestamp":1760177738000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1424-8220\/20\/19\/5621"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2020,10,1]]},"references-count":41,"journal-issue":{"issue":"19","published-online":{"date-parts":[[2020,10]]}},"alternative-id":["s20195621"],"URL":"https:\/\/doi.org\/10.3390\/s20195621","relation":{},"ISSN":["1424-8220"],"issn-type":[{"value":"1424-8220","type":"electronic"}],"subject":[],"published":{"date-parts":[[2020,10,1]]}}}