{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,15]],"date-time":"2026-06-15T15:19:11Z","timestamp":1781536751106,"version":"3.54.5"},"reference-count":71,"publisher":"MDPI AG","issue":"5","license":[{"start":{"date-parts":[[2023,3,6]],"date-time":"2023-03-06T00:00:00Z","timestamp":1678060800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"name":"JSPS KAKENHI","award":["JP21H03496"],"award-info":[{"award-number":["JP21H03496"]}]},{"name":"JSPS KAKENHI","award":["JP22K12157"],"award-info":[{"award-number":["JP22K12157"]}]},{"name":"JSPS KAKENHI","award":["JPMJPR1934"],"award-info":[{"award-number":["JPMJPR1934"]}]},{"name":"JST, PRESTO","award":["JP21H03496"],"award-info":[{"award-number":["JP21H03496"]}]},{"name":"JST, PRESTO","award":["JP22K12157"],"award-info":[{"award-number":["JP22K12157"]}]},{"name":"JST, PRESTO","award":["JPMJPR1934"],"award-info":[{"award-number":["JPMJPR1934"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Sensors"],"abstract":"<jats:p>Word-level sign language recognition (WSLR) is the backbone for continuous sign language recognition (CSLR) that infers glosses from sign videos. Finding the relevant gloss from the sign sequence and detecting explicit boundaries of the glosses from sign videos is a persistent challenge. In this paper, we propose a systematic approach for gloss prediction in WLSR using the Sign2Pose Gloss prediction transformer model. The primary goal of this work is to enhance WLSR\u2019s gloss prediction accuracy with reduced time and computational overhead. The proposed approach uses hand-crafted features rather than automated feature extraction, which is computationally expensive and less accurate. A modified key frame extraction technique is proposed that uses histogram difference and Euclidean distance metrics to select and drop redundant frames. To enhance the model\u2019s generalization ability, pose vector augmentation using perspective transformation along with joint angle rotation is performed. Further, for normalization, we employed YOLOv3 (You Only Look Once) to detect the signing space and track the hand gestures of the signers in the frames. The proposed model experiments on WLASL datasets achieved the top 1% recognition accuracy of 80.9% in WLASL100 and 64.21% in WLASL300. The performance of the proposed model surpasses state-of-the-art approaches. The integration of key frame extraction, augmentation, and pose estimation improved the performance of the proposed gloss prediction model by increasing the model\u2019s precision in locating minor variations in their body posture. We observed that introducing YOLOv3 improved gloss prediction accuracy and helped prevent model overfitting. Overall, the proposed model showed 17% improved performance in the WLASL 100 dataset.<\/jats:p>","DOI":"10.3390\/s23052853","type":"journal-article","created":{"date-parts":[[2023,3,6]],"date-time":"2023-03-06T05:18:43Z","timestamp":1678079923000},"page":"2853","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":29,"title":["Sign2Pose: A Pose-Based Approach for Gloss Prediction Using a Transformer Model"],"prefix":"10.3390","volume":"23","author":[{"given":"Jennifer","family":"Eunice","sequence":"first","affiliation":[{"name":"Department of Electronics and Communication Engineering, Karunya Institute of Technology and Sciences, Coimbatore 641114, India"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-3592-6543","authenticated-orcid":false,"given":"Andrew","family":"J","sequence":"additional","affiliation":[{"name":"Computer Science and Engineering, Manipal Institute of Technology, Manipal Academy of Higher Education, Manipal 576104, India"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-2552-6717","authenticated-orcid":false,"given":"Yuichi","family":"Sei","sequence":"additional","affiliation":[{"name":"Department of Informatics, The University of Electro-Communications, Tokyo 182-8585, Japan"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-6091-1880","authenticated-orcid":false,"given":"D. Jude","family":"Hemanth","sequence":"additional","affiliation":[{"name":"Department of Electronics and Communication Engineering, Karunya Institute of Technology and Sciences, Coimbatore 641114, India"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1968","published-online":{"date-parts":[[2023,3,6]]},"reference":[{"key":"ref_1","first-page":"9","article-title":"Automatic Sign Language Finger Spelling Using Convolution Neural Network: Analysis","volume":"117","author":"Dept","year":"2017","journal-title":"Int. J. Pure Appl. Math."},{"key":"ref_2","first-page":"437","article-title":"Deep CNN for Static Indian Sign Language Digits Recognition","volume":"Volume 347","year":"2022","journal-title":"Frontiers in Artificial Intelligence and Applications"},{"key":"ref_3","doi-asserted-by":"crossref","first-page":"432","DOI":"10.1016\/j.dib.2016.02.060","article-title":"Handwritten mathematical symbols dataset","volume":"7","author":"Chajri","year":"2016","journal-title":"Data Br."},{"key":"ref_4","doi-asserted-by":"crossref","unstructured":"Huang, J., Zhou, W., Zhang, Q., Li, H., and Li, W. (2018, January 2\u20137). Video-based sign language recognition without temporal segmentation. Proceedings of the 32nd Proceedings of the AAAI Conference on Artificial Intelligence, New Orleans, LA, USA.","DOI":"10.1609\/aaai.v32i1.11903"},{"key":"ref_5","doi-asserted-by":"crossref","first-page":"821","DOI":"10.18178\/ijmlc.2019.9.6.879","article-title":"Static sign language recognition using deep learning","volume":"9","author":"Tolentino","year":"2019","journal-title":"Int. J. Mach. Learn. Comput."},{"key":"ref_6","doi-asserted-by":"crossref","first-page":"38044","DOI":"10.1109\/ACCESS.2019.2904749","article-title":"Dynamic Sign Language Recognition Based on Video Sequence with BLSTM-3D Residual Networks","volume":"7","author":"Liao","year":"2019","journal-title":"IEEE Access"},{"key":"ref_7","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1016\/j.patrec.2016.12.004","article-title":"Coupled HMM-based Multi-Sensor Data Fusion for Sign Language Recognition","volume":"86","author":"Kumar","year":"2016","journal-title":"Pattern Recognit. Lett."},{"key":"ref_8","doi-asserted-by":"crossref","first-page":"122","DOI":"10.1007\/978-3-030-32523-7_9","article-title":"Hand Sign Language Feature Extraction Using Image Processing","volume":"1070","author":"Chabchoub","year":"2020","journal-title":"Adv. Intell. Syst. Comput."},{"key":"ref_9","unstructured":"Ong, E.J., and Bowden, R. (2004, January 19). A boosted classifier tree for hand shape detection. Proceedings of the Sixth IEEE International Conference on Automatic Face and Gesture Recognition, Seoul, Republic of Korea."},{"key":"ref_10","doi-asserted-by":"crossref","first-page":"70","DOI":"10.1007\/s11263-013-0672-6","article-title":"Automatic and efficient human pose estimation for sign language videos","volume":"110","author":"Charles","year":"2014","journal-title":"Int. J. Comput. Vis."},{"key":"ref_11","doi-asserted-by":"crossref","first-page":"671","DOI":"10.1016\/j.imavis.2014.02.009","article-title":"Non-manual grammatical marker recognition based on multi-scale, spatio-temporal analysis of head pose and facial expressions","volume":"32","author":"Liu","year":"2014","journal-title":"Image Vis. Comput."},{"key":"ref_12","doi-asserted-by":"crossref","first-page":"697","DOI":"10.1007\/978-3-030-58586-0_41","article-title":"Fully Convolutional Networks for Continuous Sign Language Recognition","volume":"Volume 12369 LNCS","author":"Cheng","year":"2020","journal-title":"Lecture Notes in Computer Science"},{"key":"ref_13","doi-asserted-by":"crossref","unstructured":"Koller, O., Ney, H., and Bowden, R. (2016, January 27\u201330). Deep Hand: How to Train a CNN on 1 Million Hand Images When Your Data Is Continuous and Weakly Labelled. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.412"},{"key":"ref_14","doi-asserted-by":"crossref","unstructured":"Koller, O., Zargaran, S., and Ney, H. (2017\u201326, January 21). Resign: Re-aligned end-to-end sequence modelling with deep recurrent CNN-HMMs. Proceedings of the 30th IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.364"},{"key":"ref_15","doi-asserted-by":"crossref","first-page":"28230","DOI":"10.1109\/ACCESS.2019.2901930","article-title":"Gesture Recognition Based on CNN and DCGAN for Calculation and Text Output","volume":"7","author":"Zhang","year":"2019","journal-title":"IEEE Access"},{"key":"ref_16","unstructured":"Rastgoo, R., Kiani, K., and Escalera, S. (2022). Word separation in continuous sign language using isolated signs and post-processing. arXiv."},{"key":"ref_17","doi-asserted-by":"crossref","unstructured":"Guo, D., Zhou, W., Li, H., and Wang, M. (2018, January 2\u20137). Hierarchical LSTM for sign language translation. Proceedings of the 32nd AAAI Conference on Artificial Intelligence, New Orleans, LA, USA.","DOI":"10.1609\/aaai.v32i1.12235"},{"key":"ref_18","doi-asserted-by":"crossref","unstructured":"Agha, R.A.A.R., Sefer, M.N., and Fattah, P. (2018, January 1\u20132). A comprehensive study on sign languages recognition systems using (SVM, KNN, CNN and ANN). Proceedings of the Proceedings of the First International Conference on Data Science, E-learning and Information Systems-DATA\u201918, New York, NY, USA.","DOI":"10.1145\/3279996.3280024"},{"key":"ref_19","doi-asserted-by":"crossref","unstructured":"Rahim, M.A., Islam, M.R., and Shin, J. (2019). Non-touch sign word recognition based on dynamic hand gesture using hybrid segmentation and CNN feature fusion. Appl. Sci., 9.","DOI":"10.3390\/app9183790"},{"key":"ref_20","doi-asserted-by":"crossref","first-page":"5665","DOI":"10.1109\/JBHI.2022.3197331","article-title":"An Attention-based 3D CNN with Multi-scale Integration Block for Alzheimer\u2019 s Disease Classification","volume":"26","author":"Wu","year":"2022","journal-title":"IEEE J. Biomed. Health Inform."},{"key":"ref_21","doi-asserted-by":"crossref","first-page":"399","DOI":"10.1007\/978-3-319-93000-8_45","article-title":"Sign Language Recognition Based on 3D Convolutional Neural Networks","volume":"Volume 10882 LNCS","author":"Neto","year":"2018","journal-title":"Lecture Notes in Computer Science"},{"key":"ref_22","first-page":"3104","article-title":"Sequence to sequence learning with neural networks","volume":"4","author":"Sutskever","year":"2014","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_23","doi-asserted-by":"crossref","unstructured":"Chen, Y., Wei, F., Sun, X., Wu, Z., and Lin, S. (2022, January 19\u201320). A Simple Multi-Modality Transfer Learning Baseline for Sign Language Translation. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, New Orleans, LA, USA.","DOI":"10.1109\/CVPR52688.2022.00506"},{"key":"ref_24","doi-asserted-by":"crossref","unstructured":"Camgoz, N.C., Hadfield, S., Koller, O., Ney, H., and Bowden, R. (2018, January 18\u201323). Neural Sign Language Translation. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00812"},{"key":"ref_25","first-page":"3766","article-title":"Findings of the Association for Computational Linguistics Prior Knowledge and Memory Enriched Transformer for Sign Language Translation","volume":"2022","author":"Jin","year":"2022","journal-title":"Assoc. Comput. Linguist."},{"key":"ref_26","unstructured":"Camgoz, N.C., Koller, O., Hadfield, S., and Bowden, R. (2020, January 13\u201319). Sign language transformers: Joint end-to-end sign language recognition and translation. Proceedings of the Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA."},{"key":"ref_27","unstructured":"Xu, Y., and Seneff, S. (2008, January 21\u201325). Two-Stage Translation: A Combined Linguistic and Statistical Machine Translation Framework. Proceedings of the Conference of the Association for Machine Translation in the Americas, Waikiki, HI, USA."},{"key":"ref_28","unstructured":"Jang, J.Y., Park, H., Shin, S., Shin, S., Yoon, B., and Gweon, G. (2022, January 20\u201325). Automatic Gloss-level Data Augmentation for Sign Language Translation. Proceedings of the 2022 Language Resources and Evaluation Conference, LREC 2022, Marseille, France."},{"key":"ref_29","doi-asserted-by":"crossref","first-page":"263","DOI":"10.1093\/deafed\/enaa038","article-title":"The ASL-LEX 2.0 Project: A Database of Lexical and Phonological Properties for 2,723 Signs in American Sign Language","volume":"26","author":"Sehyr","year":"2021","journal-title":"J. Deaf Stud. Deaf Educ."},{"key":"ref_30","doi-asserted-by":"crossref","first-page":"784","DOI":"10.3758\/s13428-016-0742-0","article-title":"ASL-LEX: A lexical database of American Sign Language","volume":"49","author":"Caselli","year":"2017","journal-title":"Behav. Res. Methods"},{"key":"ref_31","unstructured":"Vaswani, A., Shazeer, N.M., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A.N., Kaiser, L., and Polosukhin, I. (2017). Attention is All you Need. arXiv."},{"key":"ref_32","doi-asserted-by":"crossref","unstructured":"Koller, O., Zargaran, S., Ney, H., and Bowden, R. (2016, January 19\u201322). Deep sign: Hybrid CNN-HMM for continuous sign language recognition. Proceedings of the British Machine Vision Conference 2016, York, UK.","DOI":"10.5244\/C.30.136"},{"key":"ref_33","doi-asserted-by":"crossref","first-page":"1583","DOI":"10.1109\/TPAMI.2016.2537340","article-title":"Deep Dynamic Neural Networks for Multimodal Gesture Segmentation and Recognition","volume":"38","author":"Wu","year":"2016","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_34","doi-asserted-by":"crossref","first-page":"108","DOI":"10.1016\/j.cviu.2015.09.013","article-title":"Continuous sign language recognition: Towards large vocabulary statistical recognition systems handling multiple signers","volume":"141","author":"Koller","year":"2015","journal-title":"Comput. Vis. Image Underst."},{"key":"ref_35","doi-asserted-by":"crossref","first-page":"1880","DOI":"10.1109\/TMM.2018.2889563","article-title":"A Deep Neural Framework for Continuous Sign Language Recognition by Iterative Training","volume":"21","author":"Cui","year":"2019","journal-title":"IEEE Trans. Multimed."},{"key":"ref_36","first-page":"1531","article-title":"Continuous sign language recognition using isolated signs data and deep transfer learning","volume":"1","author":"Sharma","year":"2021","journal-title":"J. Ambient Intell. Humaniz. Comput."},{"key":"ref_37","unstructured":"Vedaldi, A., Bischof, H., Brox, T., and Frahm, J.-M. (2020, January 23\u201328). Stochastic Fine-Grained Labeling of Multi-state Sign Glosses for Continuous Sign Language Recognition. Proceedings of the Computer Vision\u2014ECCV 2020, Glasgow, UK."},{"key":"ref_38","doi-asserted-by":"crossref","unstructured":"Tunga, A., Nuthalapati, S.V., and Wachs, J. (2021, January 3\u20138). Pose-based Sign Language Recognition using GCN and BERT. Proceedings of the IEEE\/CVF Winter Conference on Applications of Computer Vision, Waikoloa, HI, USA.","DOI":"10.1109\/WACVW52041.2021.00008"},{"key":"ref_39","doi-asserted-by":"crossref","unstructured":"Cui, R., Liu, H., and Zhang, C. (2016, January 21\u201326). Recurrent convolutional neural networks for continuous sign language recognition by staged optimization. Proceedings of the 30th IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.175"},{"key":"ref_40","doi-asserted-by":"crossref","first-page":"2662","DOI":"10.1109\/TMM.2021.3087006","article-title":"Conditional Sentence Generation and Cross-Modal Reranking for Sign Language Translation","volume":"24","author":"Zhao","year":"2022","journal-title":"IEEE Trans. Multimed."},{"key":"ref_41","unstructured":"Kim, Y., Kwak, M., Lee, D., Kim, Y., and Baek, H. (2022). Keypoint based Sign Language Translation without Glosses. arXiv."},{"key":"ref_42","doi-asserted-by":"crossref","first-page":"115","DOI":"10.1016\/j.neucom.2022.05.051","article-title":"Full transformer network with masking future for word-level sign language recognition","volume":"500","author":"Du","year":"2022","journal-title":"Neurocomputing"},{"key":"ref_43","unstructured":"Camg\u00f6z, N.C., Koller, O., Hadfield, S., and Bowden, R. (2020). Sign Language Transformers: Joint End-to-end Sign Language Recognition and Translation. arXiv."},{"key":"ref_44","doi-asserted-by":"crossref","unstructured":"Ko, S.K., Kim, C.J., Jung, H., and Cho, C. (2019). Neural sign language translation based on human keypoint estimation. Appl. Sci., 9.","DOI":"10.3390\/app9132683"},{"key":"ref_45","unstructured":"Read, J., and Polytechnique, E. (2017). Better Sign Language Translation with STMC-Transformer. arXiv."},{"key":"ref_46","unstructured":"Walczynska, J. (2022). HandTalk: American Sign Language Recognition by 3D-CNNs. [Ph.D. Thesis, University of Groningen]."},{"key":"ref_47","doi-asserted-by":"crossref","unstructured":"Papastratis, I., Dimitropoulos, K., and Daras, P. (2021). Continuous Sign Language Recognition through a Context-Aware Generative Adversarial Network. Sensors, 21.","DOI":"10.3390\/s21072437"},{"key":"ref_48","doi-asserted-by":"crossref","unstructured":"Bohacek, M., and Hruz, M. (2022, January 4\u20138). Sign Pose-based Transformer for Word-level Sign Language Recognition. Proceedings of the IEEE\/CVF Winter Conference on Applications of Computer Vision (WACV) Workshops, Waikoloa, HI, USA.","DOI":"10.1109\/WACVW54805.2022.00024"},{"key":"ref_49","doi-asserted-by":"crossref","unstructured":"Inan, M., Zhong, Y., Hassan, S., Quandt, L., and Alikhani, M. (2022). Modeling Intensification for Sign Language Generation: A Computational Approach. arXiv.","DOI":"10.18653\/v1\/2022.findings-acl.228"},{"key":"ref_50","unstructured":"Jiang, S., Sun, B., Wang, L., Bai, Y., Li, K., and Fu, Y. (2021). Sign Language Recognition via Skeleton-Aware Multi-Model Ensemble. arXiv."},{"key":"ref_51","first-page":"9735392","article-title":"Key Frame Extraction Method of Music and Dance Video Based on Multicore Learning Feature Fusion","volume":"2022","author":"Yao","year":"2022","journal-title":"Sci. Program."},{"key":"ref_52","doi-asserted-by":"crossref","first-page":"107540","DOI":"10.1016\/j.compeleceng.2021.107540","article-title":"An improved smart key frame extraction algorithm for vehicle target recognition","volume":"97","author":"Wang","year":"2022","journal-title":"Comput. Electr. Eng."},{"key":"ref_53","doi-asserted-by":"crossref","first-page":"1818","DOI":"10.1109\/JAS.2022.105602","article-title":"Structured Sparse Coding With the Group Log-regularizer for Key Frame Extraction","volume":"9","author":"Li","year":"2022","journal-title":"IEEE\/CAA J. Autom. Sin."},{"key":"ref_54","doi-asserted-by":"crossref","unstructured":"Nie, B.X., Xiong, C., and Zhu, S.C. (2015, January 7\u201312). Joint action recognition and pose estimation from video. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Boston, MA, USA.","DOI":"10.1109\/CVPR.2015.7298734"},{"key":"ref_55","doi-asserted-by":"crossref","unstructured":"Gan, S., Yin, Y., Jiang, Z., Xie, L., and Lu, S. (2021, January 20\u201324). Skeleton-Aware Neural Sign Language Translation. Proceedings of the 29th ACM International Conference on Multimedia, Virtual.","DOI":"10.1145\/3474085.3475577"},{"key":"ref_56","unstructured":"Novopoltsev, M., Verkhovtsev, L., Murtazin, R., Milevich, D., and Zemtsova, I. (2023). Fine-tuning of sign language recognition models: A technical report. arXiv."},{"key":"ref_57","unstructured":"Shalev-Arkushin, R., Moryossef, A., and Fried, O. (2022). Ham2Pose: Animating Sign Language Notation into Pose Sequences. arXiv."},{"key":"ref_58","doi-asserted-by":"crossref","unstructured":"Liu, F., Dai, Q., Wang, S., Zhao, L., Shi, X., and Qiao, J. (2020, January 17\u201319). Multi-relational graph convolutional networks for skeleton-based action recognition. Proceedings of the 2020 IEEE Intl Conf on Parallel & Distributed Processing with Applications, Big Data & Cloud Computing, Sustainable Computing & Communications, Social Computing & Networking (ISPA\/BDCloud\/SocialCom\/SustainCom), Exeter, UK.","DOI":"10.1109\/ISPA-BDCloud-SocialCom-SustainCom51426.2020.00085"},{"key":"ref_59","doi-asserted-by":"crossref","unstructured":"De Coster, M., Van Herreweghe, M., and Dambre, J. (2021, January 20\u201325). Isolated sign recognition from RGB video using pose flow and self-attention. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Nashville, TN, USA.","DOI":"10.1109\/CVPRW53098.2021.00383"},{"key":"ref_60","doi-asserted-by":"crossref","unstructured":"Li, D., Opazo, C.R., Yu, X., and Li, H. (2020, January 1\u20135). Word-level deep sign language recognition from video: A new large-scale dataset and methods comparison. Proceedings of the IEEE\/CVF Winter Conference on Applications of Computer Vision, Snowmass Village, CO, USA.","DOI":"10.1109\/WACV45572.2020.9093512"},{"key":"ref_61","doi-asserted-by":"crossref","unstructured":"Madadi, M., Escalera, S., Carruesco, A., Andujar, C., Bar\u00f3, X., and Gonz\u00e0lez, J. (2017\u20133, January 30). Occlusion Aware Hand Pose Recovery from Sequences of Depth Images. Proceedings of the 2017 12th IEEE International Conference on Automatic Face & Gesture Recognition (FG 2017), Washington, DC, USA.","DOI":"10.1109\/FG.2017.37"},{"key":"ref_62","unstructured":"Joze, H.R.V., and Koller, O. (2019, January 9\u201312). MS-ASL: A large-scale data set and benchmark for understanding American sign language. Proceedings of the 30th British Machine Vision Conference 2019, BMVC 2019, Cardiff, UK."},{"key":"ref_63","unstructured":"Kagirov, I., Ivanko, D., Ryumin, D., Axyonov, A., and Karpov, A. (2020, January 11\u201316). TheRuSLan: Database of Russian sign language. Proceedings of the 12th Language Resources and Evaluation Conference, Marseille, France."},{"key":"ref_64","doi-asserted-by":"crossref","first-page":"181340","DOI":"10.1109\/ACCESS.2020.3028072","article-title":"AUTSL: A large scale multi-modal Turkish sign language dataset and baseline methods","volume":"8","author":"Sincan","year":"2020","journal-title":"IEEE Access"},{"key":"ref_65","doi-asserted-by":"crossref","unstructured":"Pishchulin, L., Insafutdinov, E., Tang, S., Andres, B., Andriluka, M., Gehler, P., and Schiele, B. (2016, January 27\u201330). DeepCut: Joint subset partition and labeling for multi person pose estimation. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.533"},{"key":"ref_66","doi-asserted-by":"crossref","first-page":"130105","DOI":"10.1007\/s11432-020-3065-4","article-title":"Deep graph cut network for weakly-supervised semantic segmentation","volume":"64","author":"Feng","year":"2021","journal-title":"Sci. China Inf. Sci."},{"key":"ref_67","doi-asserted-by":"crossref","first-page":"422","DOI":"10.1080\/10095020.2021.1960779","article-title":"VNLSTM-PoseNet: A novel deep ConvNet for real-time 6-DOF camera relocalization in urban streets","volume":"24","author":"Li","year":"2021","journal-title":"Geo-Spatial Inf. Sci."},{"key":"ref_68","doi-asserted-by":"crossref","unstructured":"Kitamura, T., Teshima, H., Thomas, D., and Kawasaki, H. (2022, January 3\u20138). Refining OpenPose with a new sports dataset for robust 2D pose estimation. Proceedings of the IEEE\/CVF Winter Conference on Applications of Computer Vision, Waikoloa, HI, USA.","DOI":"10.1109\/WACVW54805.2022.00074"},{"key":"ref_69","doi-asserted-by":"crossref","unstructured":"Bauer, A. (2013). The Use of Signing Space in a Shared Sign Language of Australia, De Gruyter Mouton.","DOI":"10.1515\/9781614515470"},{"key":"ref_70","doi-asserted-by":"crossref","unstructured":"Senanayaka, S.A.M.A.S., Perera, R.A.D.B.S., Rankothge, W., Usgalhewa, S.S., Hettihewa, H.D., and Abeygunawardhana, P.K.W. (2022, January 1-03). Continuous American Sign Language Recognition Using Computer Vision And Deep Learning Technologies. Proceedings of the 2022 IEEE Region 10 Symposium (TENSYMP), Mumbai, India.","DOI":"10.1109\/TENSYMP54529.2022.9864539"},{"key":"ref_71","doi-asserted-by":"crossref","unstructured":"Maruyama, M., Singh, S., Inoue, K., Roy, P.P., Iwamura, M., and Yoshioka, M. (2021). Word-Level Sign Language Recognition with Multi-Stream Neural Networks Focusing on Local Regions and Skeletal Information. arXiv.","DOI":"10.2139\/ssrn.4263878"}],"container-title":["Sensors"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1424-8220\/23\/5\/2853\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,10]],"date-time":"2025-10-10T18:48:54Z","timestamp":1760122134000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1424-8220\/23\/5\/2853"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,3,6]]},"references-count":71,"journal-issue":{"issue":"5","published-online":{"date-parts":[[2023,3]]}},"alternative-id":["s23052853"],"URL":"https:\/\/doi.org\/10.3390\/s23052853","relation":{},"ISSN":["1424-8220"],"issn-type":[{"value":"1424-8220","type":"electronic"}],"subject":[],"published":{"date-parts":[[2023,3,6]]}}}