{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,1]],"date-time":"2026-06-01T14:28:54Z","timestamp":1780324134699,"version":"3.54.1"},"reference-count":63,"publisher":"Association for Computing Machinery (ACM)","issue":"4","license":[{"start":{"date-parts":[[2024,4,15]],"date-time":"2024-04-15T00:00:00Z","timestamp":1713139200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"name":"National Key Research and Development Program of China","award":["2020AAA0108004"],"award-info":[{"award-number":["2020AAA0108004"]}]},{"name":"Key Research and Development Program of Yunnan Province","award":["202203AA080004"],"award-info":[{"award-number":["202203AA080004"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Asian Low-Resour. Lang. Inf. Process."],"published-print":{"date-parts":[[2024,4,30]]},"abstract":"<jats:p>\n            Named Entity Recognition (NER) is an indispensable component of Natural Language Processing (NLP), which aims to identify and classify entities within text data. While Deep Learning (DL) models have excelled in NER for well-resourced languages such as English, Spanish, and Chinese, they face significant hurdles when dealing with low-resource languages such as Urdu. These challenges stem from the intricate linguistic characteristics of Urdu, including morphological diversity, a context-dependent lexicon, and the scarcity of training data. This study addresses these issues by focusing on Urdu Named Entity Recognition (U-NER) and introducing three key contributions. First, various pre-trained embedding methods are employed, encompassing Word2vec (W2V), GloVe, FastText, Bidirectional Encoder Representations from Transformers (BERT), and Embeddings from language models (ELMo). In particular, fine-tuning is performed on BERT\n            <jats:sub>BASE<\/jats:sub>\n            and ELMo using Urdu Wikipedia and news articles. Second, a novel generative Data Augmentation (DA) technique replaces Named Entities (NEs) with mask tokens, employing pre-trained masked language models to predict masked tokens, effectively expanding the training dataset. Finally, the study introduces a novel hybrid model combining a Transformer Encoder with a Convolutional Neural Network (CNN) to capture the intricate morphology of Urdu. These modules enable the model to handle polysemy, extract short- and long-range dependencies, and enhance learning capacity. Empirical experiments demonstrate that the proposed model, incorporating BERT embeddings and an innovative DA approach, attains the highest F1-score of 93.99%, highlighting its efficacy for the U-NER task.\n          <\/jats:p>","DOI":"10.1145\/3648362","type":"journal-article","created":{"date-parts":[[2024,2,15]],"date-time":"2024-02-15T11:54:53Z","timestamp":1707998093000},"page":"1-38","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":15,"title":["Enriching Urdu NER with BERT Embedding, Data Augmentation, and Hybrid Encoder-CNN Architecture"],"prefix":"10.1145","volume":"23","author":[{"ORCID":"https:\/\/orcid.org\/0000-0001-6223-875X","authenticated-orcid":false,"given":"Anil","family":"Ahmed","sequence":"first","affiliation":[{"name":"Dalian University of Technology, Dalian, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-8860-7805","authenticated-orcid":false,"given":"Degen","family":"Huang","sequence":"additional","affiliation":[{"name":"Dalian University of Technology, Dalian, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-2121-1865","authenticated-orcid":false,"given":"Syed Yasser","family":"Arafat","sequence":"additional","affiliation":[{"name":"Mirpur University of Science and Technology, Mirpur, Pakistan"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0006-4148-5755","authenticated-orcid":false,"given":"Imran","family":"Hameed","sequence":"additional","affiliation":[{"name":"Dalian University of Technology, Dalian, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2024,4,15]]},"reference":[{"key":"e_1_3_2_2_2","first-page":"1638","volume-title":"Proceedings of the 27th International Conference on Computational Linguistics","author":"Akbik Alan","year":"2018","unstructured":"Alan Akbik, Duncan Blythe, and Roland Vollgraf. 2018. Contextual string embeddings for sequence labeling. In Proceedings of the 27th International Conference on Computational Linguistics. 1638\u20131649."},{"key":"e_1_3_2_3_2","first-page":"91","volume-title":"Proceedings of the 4th Workshop on Open-Source Arabic Corpora and Processing Tools, with a Shared Task on Offensive Language Detection","author":"Alharbi Abdullah I.","year":"2020","unstructured":"Abdullah I. Alharbi and Mark Lee. 2020. Combining character and word embeddings for the detection of offensive language in Arabic. In Proceedings of the 4th Workshop on Open-Source Arabic Corpora and Processing Tools, with a Shared Task on Offensive Language Detection. 91\u201396."},{"key":"e_1_3_2_4_2","doi-asserted-by":"crossref","first-page":"416","DOI":"10.18653\/v1\/2023.acl-short.36","article-title":"Split-NER: Named entity recognition via two question-answering-based classifications","author":"Arora Jatin","year":"2023","unstructured":"Jatin Arora and Youngja Park. 2023. Split-NER: Named entity recognition via two question-answering-based classifications. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), 416\u2013426.","journal-title":"Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers)"},{"key":"e_1_3_2_5_2","first-page":"1","article-title":"UHated: Hate speech detection in Urdu language using transfer learning","author":"Arshad Muhammad Umair","year":"2023","unstructured":"Muhammad Umair Arshad, Raza Ali, Mirza Omer Beg, and Waseem Shahzad. 2023. UHated: Hate speech detection in Urdu language using transfer learning. Language Resources and Evaluation (2023), 1\u201320.","journal-title":"Language Resources and Evaluation"},{"key":"e_1_3_2_6_2","doi-asserted-by":"publisher","DOI":"10.1145\/3465221"},{"key":"e_1_3_2_7_2","doi-asserted-by":"publisher","DOI":"10.1080\/13658816.2022.2133125"},{"key":"e_1_3_2_8_2","doi-asserted-by":"publisher","DOI":"10.1162\/tacl_a_00051"},{"key":"e_1_3_2_9_2","first-page":"1877","article-title":"Language models are few-shot learners","volume":"33","author":"Brown Tom","year":"2020","unstructured":"Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D. Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, and Amanda Askell. 2020. Language models are few-shot learners. Advances in Neural Information Processing Systems 33 (2020), 1877\u20131901.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_2_10_2","doi-asserted-by":"publisher","DOI":"10.5555\/2832415.2832421"},{"key":"e_1_3_2_11_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2020.coling-main.343"},{"key":"e_1_3_2_12_2","article-title":"BERT: Pre-training of deep bidirectional transformers for language understanding","volume":"1810","author":"Devlin Jacob","year":"2018","unstructured":"Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018. BERT: Pre-training of deep bidirectional transformers for language understanding. CoRR abs\/1810.04805 (2018). arXiv:1810.04805http:\/\/arxiv.org\/abs\/1810.04805","journal-title":"CoRR"},{"key":"e_1_3_2_13_2","doi-asserted-by":"publisher","DOI":"10.1109\/PROC.1973.9030"},{"key":"e_1_3_2_14_2","doi-asserted-by":"publisher","DOI":"10.2196\/39077"},{"key":"e_1_3_2_15_2","doi-asserted-by":"publisher","DOI":"10.1007\/s41060-021-00305-w"},{"key":"e_1_3_2_16_2","doi-asserted-by":"crossref","first-page":"6645","DOI":"10.1109\/ICASSP.2013.6638947","article-title":"Speech recognition with deep recurrent neural networks","author":"Graves Alex","year":"2013","unstructured":"Alex Graves, Abdel-rahman Mohamed, and Geoffrey Hinton. 2013. Speech recognition with deep recurrent neural networks. 2013 IEEE International Conference on Acoustics, Speech and Signal Processing, 6645\u20136649.","journal-title":"2013 IEEE International Conference on Acoustics, Speech and Signal Processing"},{"key":"e_1_3_2_17_2","first-page":"23","article-title":"Character-aware neural networks for Arabic named entity recognition for social media","author":"Gridach Mourad","year":"2016","unstructured":"Mourad Gridach. 2016. Character-aware neural networks for Arabic named entity recognition for social media. Proceedings of the 6th Workshop on South and Southeast Asian Natural Language Processing (WSSANLP 2016), 23\u201332.","journal-title":"Proceedings of the 6th Workshop on South and Southeast Asian Natural Language Processing (WSSANLP 2016)"},{"key":"e_1_3_2_18_2","article-title":"Message Understanding Conference-6: A brief history","author":"Grishman Ralph","year":"1996","unstructured":"Ralph Grishman and Beth M. Sundheim. 1996. Message Understanding Conference-6: A brief history. COLING 1996 Volume 1: The 16th International Conference on Computational Linguistics.","journal-title":"COLING 1996 Volume 1: The 16th International Conference on Computational Linguistics"},{"key":"e_1_3_2_19_2","doi-asserted-by":"publisher","DOI":"10.1093\/comjnl\/bxac047"},{"key":"e_1_3_2_20_2","article-title":"From text to map: Combing named entity recognition and geographic information systems","author":"Harper Charlie","year":"2020","unstructured":"Charlie Harper and R. Benjamin Gorham. 2020. From text to map: Combing named entity recognition and geographic information systems. Code4Lib Journal (2020). Issue 49.","journal-title":"Code4Lib Journal"},{"key":"e_1_3_2_21_2","article-title":"LSTM can solve hard long time lag problems","volume":"9","author":"Hochreiter Sepp","year":"1996","unstructured":"Sepp Hochreiter and J\u00fcrgen Schmidhuber. 1996. LSTM can solve hard long time lag problems. Advances in Neural Information Processing Systems 9 (1996).","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_2_22_2","first-page":"1","article-title":"DTranNER: Biomedical named entity recognition with deep learning-based label-label transition model","volume":"21","author":"Hong S. K.","year":"2020","unstructured":"S. K. Hong and Jae-Gil Lee. 2020. DTranNER: Biomedical named entity recognition with deep learning-based label-label transition model. BMC Bioinformatics 21 (2020), 1\u201311.","journal-title":"BMC Bioinformatics"},{"key":"e_1_3_2_23_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.eswa.2023.119880"},{"key":"e_1_3_2_24_2","volume-title":"Proceedings of the 6th Workshop on Asian Language Resources","author":"Hussain Sarmad","year":"2008","unstructured":"Sarmad Hussain. 2008. Resources for Urdu language processing. In Proceedings of the 6th Workshop on Asian Language Resources."},{"key":"e_1_3_2_25_2","first-page":"95","volume-title":"Proceedings of the 10th Workshop on Asian Language Resources","author":"Jahangir Faryal","year":"2012","unstructured":"Faryal Jahangir, Waqas Anwar, Usama Ijaz Bajwa, and Xuan Wang. 2012. N-gram and gazetteer list based named entity recognition for Urdu: A scarce resourced language. In Proceedings of the 10th Workshop on Asian Language Resources. 95\u2013104."},{"key":"e_1_3_2_26_2","doi-asserted-by":"publisher","DOI":"10.1145\/3329710"},{"key":"e_1_3_2_27_2","first-page":"1","article-title":"A deep learning approach to building a framework for Urdu POS and NER","author":"Kazi Samreen","year":"2023","unstructured":"Samreen Kazi, Maria Rahim, and Shakeel Khoja. 2023. A deep learning approach to building a framework for Urdu POS and NER. Journal of Intelligent & Fuzzy SystemsPreprint (2023), 1\u201311.","journal-title":"Journal of Intelligent & Fuzzy Systems"},{"key":"e_1_3_2_28_2","first-page":"1","article-title":"Persian automatic text summarization based on named entity recognition","author":"Khademi Mohammad Ebrahim","year":"2020","unstructured":"Mohammad Ebrahim Khademi and Mohammad Fakhredanesh. 2020. Persian automatic text summarization based on named entity recognition. Iranian Journal of Science and Technology, Transactions of Electrical Engineering (2020), 1\u201312.","journal-title":"Iranian Journal of Science and Technology, Transactions of Electrical Engineering"},{"key":"e_1_3_2_29_2","article-title":"Using data augmentation and bidirectional encoder representations from transformers for improving Punjabi named entity recognition","author":"Khalid Hamza","year":"2023","unstructured":"Hamza Khalid, Ghulam Murtaza, and Qaiser Abbas. 2023. Using data augmentation and bidirectional encoder representations from transformers for improving Punjabi named entity recognition. ACM Transactions on Computing Education (2023).","journal-title":"ACM Transactions on Computing Education"},{"key":"e_1_3_2_30_2","doi-asserted-by":"publisher","DOI":"10.4218\/etrij.2018-0553"},{"key":"e_1_3_2_31_2","doi-asserted-by":"publisher","DOI":"10.3390\/app12136391"},{"key":"e_1_3_2_32_2","article-title":"Named entity dataset for Urdu named entity recognition task","volume":"51","author":"Khana Wahab","year":"2016","unstructured":"Wahab Khana, Ali Daudb, Jamal A. Nasira, and Tehmina Amjada. 2016. Named entity dataset for Urdu named entity recognition task. Language & Technology 51 (2016).","journal-title":"Language & Technology"},{"key":"e_1_3_2_33_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.ins.2018.10.006"},{"key":"e_1_3_2_34_2","article-title":"Adam: A method for stochastic optimization","author":"Kingma Diederik","year":"2015","unstructured":"Diederik Kingma and Jimmy Ba. 2015. Adam: A method for stochastic optimization. International Conference on Learning Representations (ICLR).","journal-title":"International Conference on Learning Representations (ICLR)"},{"key":"e_1_3_2_35_2","first-page":"18","article-title":"Data augmentation using pre-trained transformer models","author":"Kumar Varun","year":"2020","unstructured":"Varun Kumar, Ashutosh Choudhary, and Eunah Cho. 2020. Data augmentation using pre-trained transformer models, William M. Campbell, Alex Waibel, Dilek Hakkani-Tur, Timothy J. Hazen, Kevin Kilgour, Eunah Cho, Varun Kumar, and Hadrien Glaude (Eds.). Proceedings of the 2nd Workshop on Life-long Learning for Spoken Language Systems, 18\u201326. https:\/\/aclanthology.org\/2020.lifelongnlp-1.3","journal-title":"Proceedings of the 2nd Workshop on Life-long Learning for Spoken Language Systems"},{"key":"e_1_3_2_36_2","first-page":"271","article-title":"Word vectors, reuse, and replicability: Towards a community repository of large-text resources","author":"Kutuzov Andrei","year":"2017","unstructured":"Andrei Kutuzov, Murhaf Fares, Stephan Oepen, and Erik Velldal. 2017. Word vectors, reuse, and replicability: Towards a community repository of large-text resources. Proceedings of the 58th Conference on Simulation and Modelling, 271\u2013276.","journal-title":"Proceedings of the 58th Conference on Simulation and Modelling"},{"key":"e_1_3_2_37_2","article-title":"ALBERT: A lite BERT for self-supervised learning of language representations","author":"Lan Zhenzhong","year":"2020","unstructured":"Zhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel, Piyush Sharma, and Radu Soricut. 2020. ALBERT: A lite BERT for self-supervised learning of language representations. 8th International Conference on Learning Representations, ICLR 2020, Addis Ababa, Ethiopia, April 26\u201330, 2020. https:\/\/openreview.net\/forum?id=H1eA7AEtvS","journal-title":"8th International Conference on Learning Representations, ICLR 2020, Addis Ababa, Ethiopia, April 26\u201330, 2020"},{"key":"e_1_3_2_38_2","doi-asserted-by":"publisher","DOI":"10.1038\/nature14539"},{"key":"e_1_3_2_39_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.neucom.2021.10.101"},{"key":"e_1_3_2_40_2","article-title":"RoBERTa: A robustly optimized BERT pretraining approach","volume":"1907","author":"Liu Yinhan","year":"2019","unstructured":"Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019. RoBERTa: A robustly optimized BERT pretraining approach. CoRR abs\/1907.11692 (2019). http:\/\/arxiv.org\/abs\/1907.11692","journal-title":"CoRR"},{"key":"e_1_3_2_41_2","doi-asserted-by":"publisher","DOI":"10.1145\/3129290"},{"key":"e_1_3_2_42_2","article-title":"A survey on self-supervised pre-training for sequential transfer learning in neural networks","volume":"2007","author":"Mao Huanru Henry","year":"2020","unstructured":"Huanru Henry Mao. 2020. A survey on self-supervised pre-training for sequential transfer learning in neural networks. CoRR abs\/2007.00800 (2020). arXiv:2007.00800https:\/\/arxiv.org\/abs\/2007.00800","journal-title":"CoRR"},{"key":"e_1_3_2_43_2","article-title":"Efficient estimation of word representations in vector space","author":"Mikolov Tom\u00e1s","year":"2013","unstructured":"Tom\u00e1s Mikolov, Kai Chen, Greg Corrado, and Jeffrey Dean. 2013. Efficient estimation of word representations in vector space, Yoshua Bengio and Yann LeCun (Eds.). 1st International Conference on Learning Representations, ICLR 2013, Scottsdale, Arizona, USA, May 2-4, 2013, Workshop Track Proceedings. http:\/\/arxiv.org\/abs\/1301.3781","journal-title":"1st International Conference on Learning Representations, ICLR 2013, Scottsdale, Arizona, USA, May 2-4, 2013, Workshop Track Proceedings"},{"key":"e_1_3_2_44_2","first-page":"141","article-title":"Fast-paced improvements to named entity handling for neural machine translation","author":"Mota Pedro","year":"2022","unstructured":"Pedro Mota, Vera Cabarr\u00e3o, and Eduardo Farah. 2022. Fast-paced improvements to named entity handling for neural machine translation. Proceedings of the 23rd Annual Conference of the European Association for Machine Translation, 141\u2013149.","journal-title":"Proceedings of the 23rd Annual Conference of the European Association for Machine Translation"},{"key":"e_1_3_2_45_2","doi-asserted-by":"publisher","DOI":"10.1145\/1838751.1838754"},{"issue":"10","key":"e_1_3_2_46_2","doi-asserted-by":"crossref","first-page":"1272","DOI":"10.19026\/rjaset.8.1095","article-title":"Challenges of Urdu named entity recognition: A scarce resourced language","volume":"8","author":"Naz Saeeda","year":"2014","unstructured":"Saeeda Naz, Arif Iqbal Umar, Syed Hamad Shirazi, Sajjad Ahmad Khan, Imtiaz Ahmed, and Akbar Ali Khan. 2014. Challenges of Urdu named entity recognition: A scarce resourced language. Research Journal of Applied Sciences, Engineering and Technology 8, 10 (2014), 1272\u20131278.","journal-title":"Research Journal of Applied Sciences, Engineering and Technology"},{"key":"e_1_3_2_47_2","doi-asserted-by":"publisher","DOI":"10.3115\/v1\/D14-1162"},{"key":"e_1_3_2_48_2","article-title":"Deep contextualized word representations","volume":"1802","author":"Peters Matthew E.","year":"2018","unstructured":"Matthew E. Peters, Mark Neumann, Mohit Iyyer, Matt Gardner, Christopher Clark, Kenton Lee, and Luke Zettlemoyer. 2018. Deep contextualized word representations. CoRR abs\/1802.05365 (2018). http:\/\/arxiv.org\/abs\/1802.05365","journal-title":"CoRR"},{"issue":"8","key":"e_1_3_2_49_2","first-page":"9","article-title":"Language models are unsupervised multitask learners","volume":"1","author":"Radford Alec","year":"2019","unstructured":"Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et\u00a0al. 2019. Language models are unsupervised multitask learners. OpenAI Blog 1, 8 (2019), 9.","journal-title":"OpenAI Blog"},{"key":"e_1_3_2_50_2","doi-asserted-by":"publisher","DOI":"10.5555\/3455716.3455856"},{"key":"e_1_3_2_51_2","doi-asserted-by":"publisher","DOI":"10.5555\/1870457.1870476"},{"key":"e_1_3_2_52_2","volume-title":"Proceedings of the IJCNLP-08 Workshop on Named Entity Recognition for South and South East Asian Languages","author":"Saha Sujan Kumar","year":"2008","unstructured":"Sujan Kumar Saha, Sanjay Chatterji, Sandipan Dandapat, Sudeshna Sarkar, and Pabitra Mitra. 2008. A hybrid named entity recognition system for South and South East Asian languages. In Proceedings of the IJCNLP-08 Workshop on Named Entity Recognition for South and South East Asian Languages."},{"key":"e_1_3_2_53_2","article-title":"Automatic text summarization using document clustering named entity recognition","volume":"13","author":"Senthamizh Selvan R.","year":"2022","unstructured":"Selvan R. Senthamizh and K. Arutchelvan. 2022. Automatic text summarization using document clustering named entity recognition. International Journal of Advanced Computer Science and Applications 13 (2022). Issue 9.","journal-title":"International Journal of Advanced Computer Science and Applications"},{"key":"e_1_3_2_54_2","article-title":"GLU variants improve transformer","volume":"2002","author":"Shazeer Noam","year":"2020","unstructured":"Noam Shazeer. 2020. GLU variants improve transformer. CoRR abs\/2002.05202 (2020). https:\/\/arxiv.org\/abs\/2002.05202","journal-title":"CoRR"},{"key":"e_1_3_2_55_2","first-page":"2507","volume-title":"Proceedings of COLING 2012","author":"Singh UmrinderPal","year":"2012","unstructured":"UmrinderPal Singh, Vishal Goyal, and Gurpreet Singh Lehal. 2012. Named entity recognition system for Urdu. In Proceedings of COLING 2012. 2507\u20132518."},{"key":"e_1_3_2_56_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2021.emnlp-main.160"},{"key":"e_1_3_2_57_2","first-page":"273","article-title":"An overview of named entity recognition","author":"Sun Peng","year":"2018","unstructured":"Peng Sun, Xuezhen Yang, Xiaobing Zhao, and Zhijuan Wang. 2018. An overview of named entity recognition. 2018 International Conference on Asian Language Processing (IALP), 273\u2013278.","journal-title":"2018 International Conference on Asian Language Processing (IALP)"},{"key":"e_1_3_2_58_2","first-page":"32","article-title":"Empirical studies on the NLP techniques for source code data preprocessing","author":"Sun Xiaobing","year":"2014","unstructured":"Xiaobing Sun, Xiangyue Liu, Jiajun Hu, and Junwu Zhu. 2014. Empirical studies on the NLP techniques for source code data preprocessing. Proceedings of the 2014 3rd International Workshop on Evidential Assessment of Software Technologies, 32\u201339.","journal-title":"Proceedings of the 2014 3rd International Workshop on Evidential Assessment of Software Technologies"},{"key":"e_1_3_2_59_2","unstructured":"Hugo Touvron Thibaut Lavril Gautier Izacard Xavier Martinet Marie-Anne Lachaux Timoth\u00e9e Lacroix Baptiste Rozi\u00e8re Naman Goyal Eric Hambro Faisal Azhar Aurelien Rodriguez Armand Joulin Edouard Grave and Guillaume Lample. 2023. LLaMA: Open and Efficient Foundation Language Models. (2023)."},{"key":"e_1_3_2_60_2","first-page":"3","volume-title":"Mexican International Conference on Artificial Intelligence","author":"Ullah Fida","year":"2022","unstructured":"Fida Ullah, Ihsan Ullah, and Olga Kolesnikova. 2022. Urdu named entity recognition with attention Bi-LSTM-CRF model. In Mexican International Conference on Artificial Intelligence. Springer, 3\u201317."},{"key":"e_1_3_2_61_2","article-title":"Attention is all you need","volume":"30","author":"Vaswani Ashish","year":"2017","unstructured":"Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, \u0141ukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. Advances in Neural Information Processing Systems 30 (2017).","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_2_62_2","doi-asserted-by":"publisher","DOI":"10.1007\/s10994-021-06073-9"},{"key":"e_1_3_2_63_2","article-title":"Context-aware attentive multilevel feature fusion for named entity recognition","author":"Yang Zhiwei","year":"2022","unstructured":"Zhiwei Yang, Jing Ma, Hechang Chen, Jiawei Zhang, and Yi Chang. 2022. Context-aware attentive multilevel feature fusion for named entity recognition. IEEE Transactions on Neural Networks and Learning Systems (2022).","journal-title":"IEEE Transactions on Neural Networks and Learning Systems"},{"key":"e_1_3_2_64_2","article-title":"Root mean square layer normalization","volume":"32","author":"Zhang Biao","year":"2019","unstructured":"Biao Zhang and Rico Sennrich. 2019. Root mean square layer normalization. Advances in Neural Information Processing Systems 32 (2019).","journal-title":"Advances in Neural Information Processing Systems"}],"container-title":["ACM Transactions on Asian and Low-Resource Language Information Processing"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3648362","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3648362","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,19]],"date-time":"2025-06-19T00:04:13Z","timestamp":1750291453000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3648362"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,4,15]]},"references-count":63,"journal-issue":{"issue":"4","published-print":{"date-parts":[[2024,4,30]]}},"alternative-id":["10.1145\/3648362"],"URL":"https:\/\/doi.org\/10.1145\/3648362","relation":{},"ISSN":["2375-4699","2375-4702"],"issn-type":[{"value":"2375-4699","type":"print"},{"value":"2375-4702","type":"electronic"}],"subject":[],"published":{"date-parts":[[2024,4,15]]},"assertion":[{"value":"2023-11-01","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2024-02-02","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2024-04-15","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}