{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,12]],"date-time":"2026-06-12T15:58:35Z","timestamp":1781279915811,"version":"3.54.1"},"reference-count":64,"publisher":"Association for Computing Machinery (ACM)","issue":"1","license":[{"start":{"date-parts":[[2024,1,15]],"date-time":"2024-01-15T00:00:00Z","timestamp":1705276800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"name":"Saudi Data and AI Authority (SDAIA) and King Fahd University of Petroleum and Minerals (KFUPM) under the SDAIA-KFUPM Joint Research Center for Artificial Intelligence","award":["JRC-AI-RFP-05"],"award-info":[{"award-number":["JRC-AI-RFP-05"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Asian Low-Resour. Lang. Inf. Process."],"published-print":{"date-parts":[[2024,1,31]]},"abstract":"<jats:p>Pose-based approaches for sign language recognition provide light-weight and fast models that can be adopted in real-time applications. This article presents a framework for isolated Arabic sign language recognition using hand and face keypoints. We employed MediaPipe pose estimator for extracting the keypoints of sign gestures in the video stream. Using the extracted keypoints, three models were proposed for sign language recognition: Long-Term Short Memory, Temporal Convolution Networks, and Transformer-based models. Moreover, we investigated the importance of non-manual features for sign language recognition systems and the obtained results showed that combining hand and face keypoints boosted the recognition accuracy by around 4% compared with only hand keypoints. The proposed models were evaluated on Arabic and Argentinian sign languages. Using the KArSL-100 dataset, the proposed pose-based Transformer achieved the highest accuracy of 99.74% and 68.2% in signer-dependent and -independent modes, respectively. Additionally, the Transformer was evaluated on the LSA64 dataset and obtained an accuracy of 98.25% and 91.09% in signer-dependent and -independent modes, respectively. Consequently, the pose-based Transformer outperformed the state-of-the-art techniques on both datasets using keypoints from the signer\u2019s hands and face.<\/jats:p>","DOI":"10.1145\/3584984","type":"journal-article","created":{"date-parts":[[2023,2,21]],"date-time":"2023-02-21T11:25:42Z","timestamp":1676978742000},"page":"1-19","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":78,"title":["Isolated Arabic Sign Language Recognition Using a Transformer-based Model and Landmark Keypoints"],"prefix":"10.1145","volume":"23","author":[{"ORCID":"https:\/\/orcid.org\/0000-0003-1379-7234","authenticated-orcid":false,"given":"Sarah","family":"Alyami","sequence":"first","affiliation":[{"name":"Information and Computer Science Department, King Fahd University of Petroleum and Minerals, Saudi Arabia; Applied College, Imam Abdulrahman Bin Faisal University, Saudi Arabia"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-7944-5093","authenticated-orcid":false,"given":"Hamzah","family":"Luqman","sequence":"additional","affiliation":[{"name":"Information and Computer Science Department, King Fahd University of Petroleum and Minerals; SDAIA-KFUPM Joint Research Center for Artificial Intelligence, KFUPM, Saudi Arabia"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-1058-0996","authenticated-orcid":false,"given":"Mohammad","family":"Hammoudeh","sequence":"additional","affiliation":[{"name":"Information and Computer Science Department, King FahdUniversity of Petroleum and Minerals, Saudi Arabia"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2024,1,15]]},"reference":[{"key":"e_1_3_2_2_2","unstructured":"WHO. 2021. World Report On Hearing. Retrieved from https:\/\/www.who.int\/publications\/i\/item\/world-report-on-hearing."},{"key":"e_1_3_2_3_2","doi-asserted-by":"publisher","DOI":"10.1145\/3373625.3417023"},{"key":"e_1_3_2_4_2","doi-asserted-by":"publisher","DOI":"10.1007\/s00521-021-06079-3"},{"key":"e_1_3_2_5_2","doi-asserted-by":"publisher","DOI":"10.3390\/a13120310"},{"issue":"2","key":"e_1_3_2_6_2","first-page":"41","article-title":"Differences between American Sign Language (ASL) and British Sign Language (BSL)","volume":"1","author":"Jachova Zora","year":"2008","unstructured":"Zora Jachova, Olivera Kovacheva, and Aleksandra Karovska. 2008. Differences between American Sign Language (ASL) and British Sign Language (BSL). J. Spec. Educ. Rehab. 1, 2 (2008), 41\u201354.","journal-title":"J. Spec. Educ. Rehab."},{"key":"e_1_3_2_7_2","doi-asserted-by":"publisher","DOI":"10.1017\/CBO9780511712203.020"},{"key":"e_1_3_2_8_2","unstructured":"Saudi Press Agency. 2007. Issuance of the Unified Arabic Dictionary for Sign Language. Retrieved from https:\/\/www.spa.gov.sa\/viewstory.php?lang=en&newsid=473792."},{"key":"e_1_3_2_9_2","doi-asserted-by":"publisher","DOI":"10.3390\/electronics10141739"},{"key":"e_1_3_2_10_2","doi-asserted-by":"publisher","DOI":"10.1093\/deafed\/eni007"},{"key":"e_1_3_2_11_2","doi-asserted-by":"publisher","DOI":"10.12816\/0043439"},{"key":"e_1_3_2_12_2","doi-asserted-by":"publisher","DOI":"10.1007\/s11042-020-08961-z"},{"key":"e_1_3_2_13_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.patrec.2021.02.013"},{"key":"e_1_3_2_14_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCDS.2021.3126637"},{"key":"e_1_3_2_15_2","doi-asserted-by":"publisher","DOI":"10.1109\/WACVW54805.2022.00024"},{"key":"e_1_3_2_16_2","unstructured":"Prem Selvaraj Gokul NC Pratyush Kumar and Mitesh Khapra. 2021. OpenHands: Making sign language recognition accessible with pose-based pretrained models across languages. arxiv:2110.05877. Retrieved from http:\/\/arxiv.org\/abs\/2110.05877."},{"key":"e_1_3_2_17_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.143"},{"key":"e_1_3_2_18_2","article-title":"Blazepose: On-device real-time body pose tracking","author":"Bazarevsky Valentin","year":"2020","unstructured":"Valentin Bazarevsky, Ivan Grishchenko, Karthik Raveendran, Tyler Zhu, Fan Zhang, and Matthias Grundmann. 2020. Blazepose: On-device real-time body pose tracking. arXiv:2006.10204. Retrieved from https:\/\/arxiv.org\/abs\/2006.10204.","journal-title":"arXiv:2006.10204"},{"key":"e_1_3_2_19_2","doi-asserted-by":"publisher","DOI":"10.3390\/s21175856"},{"key":"e_1_3_2_20_2","doi-asserted-by":"publisher","DOI":"10.1145\/3453800.3453818"},{"key":"e_1_3_2_21_2","doi-asserted-by":"publisher","DOI":"10.1109\/R10-HTC53172.2021.9641542"},{"key":"e_1_3_2_22_2","unstructured":"Google. 2001. Mediapipe Solutions. Retrieved May 21 2022 from https:\/\/google.github.io\/mediapipe\/solutions\/solutions.html."},{"key":"e_1_3_2_23_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.compeleceng.2021.107395"},{"key":"e_1_3_2_24_2","doi-asserted-by":"publisher","DOI":"10.1145\/3423420"},{"key":"e_1_3_2_25_2","doi-asserted-by":"publisher","DOI":"10.1109\/isaect53699.2021.9668405"},{"key":"e_1_3_2_26_2","doi-asserted-by":"publisher","DOI":"10.1109\/ACCESS.2020.2990699"},{"key":"e_1_3_2_27_2","doi-asserted-by":"publisher","DOI":"10.1504\/IJISTA.2019.101951"},{"key":"e_1_3_2_28_2","doi-asserted-by":"publisher","DOI":"10.14569\/IJACSA.2018.090442"},{"key":"e_1_3_2_29_2","doi-asserted-by":"publisher","DOI":"10.1109\/ACCESS.2021.3069714"},{"key":"e_1_3_2_30_2","first-page":"6018","article-title":"Sign language recognition with transformer networks","author":"Coster Mathieu de","year":"2020","unstructured":"Mathieu de Coster, Mieke van Herreweghe, and Joni Dambre. 2020. Sign language recognition with transformer networks. In Proceedings of the 12th International Conference on Language Resources and Evaluation (LREC\u201920), 6018\u20136024.","journal-title":"Proceedings of the 12th International Conference on Language Resources and Evaluation (LREC\u201920)"},{"key":"e_1_3_2_31_2","doi-asserted-by":"publisher","DOI":"10.1007\/s11042-020-09048-5"},{"key":"e_1_3_2_32_2","doi-asserted-by":"publisher","DOI":"10.1109\/TSMCB.2006.889630"},{"key":"e_1_3_2_33_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-981-10-7566-7_63"},{"key":"e_1_3_2_34_2","doi-asserted-by":"publisher","DOI":"10.1109\/icABCD54961.2022.9856310"},{"key":"e_1_3_2_35_2","volume-title":"Proceedings of the Southern Africa Telecommunication Networks and Applications Conference (SATNAC\u201922)","author":"Marais Marc","year":"2022","unstructured":"Marc Marais, Dane Brown, James Connan, Alden Boby, and Luxolo Lethukuthula Kuhlane. 2022. Investigating signer-independent sign language recognition on the LSA64 dataset. In Proceedings of the Southern Africa Telecommunication Networks and Applications Conference (SATNAC\u201922). Fancourt, George, Western Cape, South Africa."},{"key":"e_1_3_2_36_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-98998-3_29"},{"key":"e_1_3_2_37_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.jvcir.2021.103280"},{"key":"e_1_3_2_38_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.jvcir.2021.103161"},{"key":"e_1_3_2_39_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.csl.2019.101053"},{"key":"e_1_3_2_40_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.procs.2021.08.216"},{"key":"e_1_3_2_41_2","doi-asserted-by":"publisher","DOI":"10.1109\/THMS.2022.3144000"},{"key":"e_1_3_2_42_2","doi-asserted-by":"publisher","DOI":"10.1007\/s11042-022-12051-7"},{"key":"e_1_3_2_43_2","article-title":"MS-ASL: A large-scale data set and benchmark for understanding American sign language","author":"Joze Hamid Reza Vaezi","year":"2020","unstructured":"Hamid Reza Vaezi Joze and Oscar Koller. 2020. MS-ASL: A large-scale data set and benchmark for understanding American sign language. In Proceedings of the 30th British Machine Vision Conference (BMVC\u201919).","journal-title":"Proceedings of the 30th British Machine Vision Conference (BMVC\u201919)"},{"key":"e_1_3_2_44_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.eswa.2020.114403"},{"key":"e_1_3_2_45_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-93000-8_45"},{"key":"e_1_3_2_46_2","doi-asserted-by":"publisher","DOI":"10.1109\/IT-ELA52201.2021.9773404"},{"key":"e_1_3_2_47_2","doi-asserted-by":"publisher","DOI":"10.1109\/WACVW52041.2021.00008"},{"key":"e_1_3_2_48_2","doi-asserted-by":"publisher","DOI":"10.1109\/ECTIDAMTNCON51128.2021.9425711"},{"key":"e_1_3_2_49_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.neunet.2020.01.030"},{"key":"e_1_3_2_50_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPRW53098.2021.00380"},{"key":"e_1_3_2_51_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICPR48806.2021.9412075"},{"key":"e_1_3_2_52_2","doi-asserted-by":"publisher","DOI":"10.1109\/WACV45572.2020.9093512"},{"key":"e_1_3_2_53_2","doi-asserted-by":"publisher","DOI":"10.3390\/s20185151"},{"key":"e_1_3_2_54_2","doi-asserted-by":"publisher","DOI":"10.1145\/3436754"},{"key":"e_1_3_2_55_2","doi-asserted-by":"publisher","DOI":"10.3390\/s21041120"},{"key":"e_1_3_2_56_2","first-page":"794","article-title":"LSA64: An argentinian sign language dataset","author":"Ronchetti Franco","year":"2016","unstructured":"Franco Ronchetti, Facundo Quiroga, and Laura Lanzarini. 2016. LSA64: An argentinian sign language dataset. Congr. Argent. Cienc. Comput. (2016), 794\u2013803.","journal-title":"Congr. Argent. Cienc. Comput."},{"key":"e_1_3_2_57_2","doi-asserted-by":"publisher","DOI":"10.1162\/neco.1997.9.8.1735"},{"key":"e_1_3_2_58_2","doi-asserted-by":"publisher","DOI":"10.1109\/LGRS.2017.2764915"},{"key":"e_1_3_2_59_2","article-title":"An empirical evaluation of generic convolutional and recurrent networks for sequence modeling","author":"Bai Shaojie","year":"2018","unstructured":"Shaojie Bai, J. Zico Kolter, and Vladlen Koltun. 2018. An empirical evaluation of generic convolutional and recurrent networks for sequence modeling. arXiv:1803.01271. Retrieved from https:\/\/arxiv.org\/abs\/1803.01271.","journal-title":"arXiv:1803.01271"},{"key":"e_1_3_2_60_2","article-title":"Short-term temporal convolutional networks for dynamic hand gesture recognition","author":"Zhang Yi","year":"2019","unstructured":"Yi Zhang, Chong Wang, Ye Zheng, Jieyu Zhao, Yuqi Li, and Xijiong Xie. 2019. Short-term temporal convolutional networks for dynamic hand gesture recognition. arXiv:2001.05833. Retrieved from https:\/\/arxiv.org\/abs\/2001.05833.","journal-title":"arXiv:2001.05833"},{"key":"e_1_3_2_61_2","article-title":"Wavenet: A generative model for raw audio","author":"Oord Aaron van den","year":"2016","unstructured":"Aaron van den Oord, Sander Dieleman, Heiga Zen, Karen Simonyan, Oriol Vinyals, Alex Graves, Nal Kalchbrenner, Andrew Senior, and Koray Kavukcuoglu. 2016. Wavenet: A generative model for raw audio. arXiv:1609.03499. Retrieved from https:\/\/arxiv.org\/abs\/1609.03499.","journal-title":"arXiv:1609.03499"},{"key":"e_1_3_2_62_2","article-title":"Attention is all you need","volume":"30","author":"Vaswani Ashish","year":"2017","unstructured":"Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, \u0141ukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. Adv. Neural Inf. Process. Syst. 30 (2017).","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"e_1_3_2_63_2","article-title":"Bert: Pre-training of deep bidirectional transformers for language understanding","author":"Devlin Jacob","year":"2018","unstructured":"Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018. Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv:1810.04805. Retrieved from https:\/\/arxiv.org\/abs\/1810.04805.","journal-title":"arXiv:1810.04805"},{"key":"e_1_3_2_64_2","doi-asserted-by":"publisher","DOI":"10.1145\/3505244"},{"key":"e_1_3_2_65_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-47955-2_28"}],"container-title":["ACM Transactions on Asian and Low-Resource Language Information Processing"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3584984","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3584984","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T16:37:07Z","timestamp":1750178227000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3584984"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,1,15]]},"references-count":64,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2024,1,31]]}},"alternative-id":["10.1145\/3584984"],"URL":"https:\/\/doi.org\/10.1145\/3584984","relation":{},"ISSN":["2375-4699","2375-4702"],"issn-type":[{"value":"2375-4699","type":"print"},{"value":"2375-4702","type":"electronic"}],"subject":[],"published":{"date-parts":[[2024,1,15]]},"assertion":[{"value":"2022-07-28","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2023-02-17","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2024-01-15","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}