{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,1]],"date-time":"2026-06-01T23:36:47Z","timestamp":1780357007535,"version":"3.54.1"},"reference-count":37,"publisher":"Association for Computing Machinery (ACM)","issue":"6","license":[{"start":{"date-parts":[[2023,6,16]],"date-time":"2023-06-16T00:00:00Z","timestamp":1686873600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Asian Low-Resour. Lang. Inf. Process."],"published-print":{"date-parts":[[2023,6,30]]},"abstract":"<jats:p><jats:bold>Named entity recognition (NER)<\/jats:bold>is a task of proper noun identification from natural language text and classification into various types such as location, person, and organization. Due to NER's applications in different<jats:bold>natural language processing (NLP)<\/jats:bold>tasks, numerous NER approaches and benchmark datasets have been proposed. However, developing NER techniques for low-resource languages is still limited due to the absence of substantial training datasets. Punjabi is a classic example of low resource language. Although various researchers have worked on Punjabi, they focused on the Gurmukhi script. To overcome the challenges in developing NER for the Shahmukhi script, we present an improved technique for Punjabi NER for the Shahmukhi script in this paper. We firstly extend the existing dataset by adding new NER classes by leveraging a novel Pool of Words data augmentation strategy. Our extended dataset has 11,31,509 tokens and 1,25,789 labeled entities with more<jats:bold>named entities (NEs)<\/jats:bold>than the older dataset. In the next step, we fine-tuned a transformer model known as<jats:bold>Bidirectional Encoder Representations from Transformers (BERT)<\/jats:bold>for the NER task. We performed experiments using the proposed approach on a new and older dataset version, showing that our method achieved competitive results.<\/jats:p>","DOI":"10.1145\/3595861","type":"journal-article","created":{"date-parts":[[2023,5,4]],"date-time":"2023-05-04T12:26:27Z","timestamp":1683203187000},"page":"1-13","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":8,"title":["Using Data Augmentation and Bidirectional Encoder Representations from Transformers for Improving Punjabi Named Entity Recognition"],"prefix":"10.1145","volume":"22","author":[{"ORCID":"https:\/\/orcid.org\/0000-0003-4097-0178","authenticated-orcid":false,"given":"Hamza","family":"Khalid","sequence":"first","affiliation":[{"name":"Department of Computer Science, University of Engineering and Technology Lahore, Punjab, Pakistan"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-9205-6356","authenticated-orcid":false,"given":"Ghulam","family":"Murtaza","sequence":"additional","affiliation":[{"name":"Department of Computer Science, University of Engineering and Technology Lahore, Punjab, Pakistan"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-9246-5208","authenticated-orcid":false,"given":"Qaiser","family":"Abbas","sequence":"additional","affiliation":[{"name":"Department of Computer Science, University of Engineering and Technology Lahore, Punjab, Pakistan"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2023,6,16]]},"reference":[{"key":"e_1_3_1_2_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-981-16-3346-1_66"},{"key":"e_1_3_1_3_2","doi-asserted-by":"publisher","DOI":"10.1186\/s12859-017-1868-5"},{"issue":"3","key":"e_1_3_1_4_2","first-page":"589","article-title":"Named entity recognition using support vector machine: A language independent approach","volume":"4","author":"Ekbal A.","year":"2010","unstructured":"A. Ekbal and S. Bandyopadhyay. 2010. Named entity recognition using support vector machine: A language independent approach. International Journal of Electrical and Computer Engineering 4, 3 (2010), 589\u2013604.","journal-title":"International Journal of Electrical and Computer Engineering"},{"key":"e_1_3_1_5_2","article-title":"Neural machine translation for low-resource languages: A survey","author":"Ranathunga S.","year":"2021","unstructured":"S. Ranathunga, E.-S. A. Lee, M. P. Skenduli, R. Shekhar, M. Alam, and R. Kaur. 2021. Neural machine translation for low-resource languages: A survey. arXiv preprint arXiv:2106.15115.","journal-title":"arXiv preprint"},{"key":"e_1_3_1_6_2","doi-asserted-by":"publisher","DOI":"10.3390\/sym13050786"},{"key":"e_1_3_1_7_2","doi-asserted-by":"publisher","DOI":"10.1145\/3329710"},{"key":"e_1_3_1_8_2","doi-asserted-by":"publisher","DOI":"10.1145\/3383306"},{"issue":"3","key":"e_1_3_1_9_2","first-page":"28","article-title":"Named entity recognition for Punjabi language text summarization","volume":"33","author":"Gupta V.","year":"2011","unstructured":"V. Gupta and G. S. Lehal. 2011. Named entity recognition for Punjabi language text summarization. International Journal of Computer Applications 33, 3 (2011), 28\u201332.","journal-title":"International Journal of Computer Applications"},{"key":"e_1_3_1_10_2","first-page":"1","article-title":"DeepSpacy-NER: An efficient deep learning model for named entity recognition for Punjabi language","author":"Singh N.","year":"2022","unstructured":"N. Singh, M. Kumar, B. Singh, and J. Singh. 2022. DeepSpacy-NER: An efficient deep learning model for named entity recognition for Punjabi language. Evolving Systems (2022). 1\u201311.","journal-title":"Evolving Systems"},{"key":"e_1_3_1_11_2","doi-asserted-by":"publisher","DOI":"10.5555\/1870457.1870476"},{"key":"e_1_3_1_12_2","first-page":"2507","volume-title":"Proceedings of COLING 2012","author":"Singh U.","year":"2012","unstructured":"U. Singh, V. Goyal, and G. S. Lehal. 2012. Named entity recognition system for Urdu. In Proceedings of COLING 2012. 2507\u20132518."},{"key":"e_1_3_1_13_2","first-page":"95","volume-title":"Proceedings of the 10th Workshop on Asian Language Resources","author":"Jahangir F.","year":"2012","unstructured":"F. Jahangir, W. Anwar, U. I. Bajwa, and X. Wang. 2012. N-gram and gazetteer list based named entity recognition for Urdu: A scarce resourced language. In Proceedings of the 10th Workshop on Asian Language Resources. 95\u2013104."},{"key":"e_1_3_1_14_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-24033-6_28"},{"key":"e_1_3_1_15_2","article-title":"Urdu named entity recognition system using hidden Markov model","author":"Malik M. K.","year":"2017","unstructured":"M. K. Malik and S. M. Sarwar. 2017. Urdu named entity recognition system using hidden Markov model. Pakistan Journal of Engineering and Applied Sciences (2017).","journal-title":"Pakistan Journal of Engineering and Applied Sciences"},{"key":"e_1_3_1_16_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.procs.2016.09.123"},{"key":"e_1_3_1_17_2","doi-asserted-by":"publisher","DOI":"10.1109\/DAS.2016.15"},{"key":"e_1_3_1_18_2","doi-asserted-by":"publisher","DOI":"10.1145\/3129290"},{"key":"e_1_3_1_19_2","doi-asserted-by":"publisher","DOI":"10.4218\/etrij.2018-0553"},{"key":"e_1_3_1_20_2","first-page":"1","article-title":"Telugu named entity recognition using BERT","author":"Gorla S.","year":"2022","unstructured":"S. Gorla, S. S. Tangeda, L. B. M. Neti, and A. Malapati. 2022. Telugu named entity recognition using BERT. International Journal of Data Science and Analytics (2022), 1\u201314.","journal-title":"International Journal of Data Science and Analytics"},{"key":"e_1_3_1_21_2","doi-asserted-by":"publisher","DOI":"10.32604\/cmc.2021.016054"},{"issue":"10","key":"e_1_3_1_22_2","doi-asserted-by":"crossref","first-page":"5731","DOI":"10.17762\/turcomat.v12i10.5387","article-title":"Deep learning strategy to recognize Kannada named entities","volume":"12","author":"Pushpalatha M.","year":"2021","unstructured":"M. Pushpalatha et al. 2021. Deep learning strategy to recognize Kannada named entities. Turkish Journal of Computer and Mathematics Education (TURCOMAT) 12, 10 (2021), 5731\u20135737.","journal-title":"Turkish Journal of Computer and Mathematics Education (TURCOMAT)"},{"key":"e_1_3_1_23_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICEECCOT43722.2018.9001559"},{"key":"e_1_3_1_24_2","doi-asserted-by":"publisher","DOI":"10.3390\/app12136391"},{"key":"e_1_3_1_25_2","article-title":"BERT: Pre-training of deep bidirectional transformers for language understanding","author":"Devlin J.","year":"2018","unstructured":"J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova. 2018. BERT: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805.","journal-title":"arXiv preprint"},{"issue":"2","key":"e_1_3_1_26_2","first-page":"299","article-title":"Common linguistic patterns in Punjabi and Persian Sufi literature","volume":"42","author":"Zafar A.","year":"2022","unstructured":"A. Zafar and A. Jabeen. 2022. Common linguistic patterns in Punjabi and Persian Sufi literature. Pakistan Journal of Social Sciences 42, 2 (2022), 299\u2013306.","journal-title":"Pakistan Journal of Social Sciences"},{"key":"e_1_3_1_27_2","doi-asserted-by":"publisher","DOI":"10.1007\/s10462-016-9482-x"},{"key":"e_1_3_1_28_2","doi-asserted-by":"publisher","DOI":"10.30630\/joiv.6.2.804"},{"key":"e_1_3_1_29_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.knosys.2021.107791"},{"key":"e_1_3_1_30_2","article-title":"ALBERT: A Lite BERT for self-supervised learning of language representations","author":"Lan Z.","year":"2019","unstructured":"Z. Lan, M. Chen, S. Goodman, K. Gimpel, P. Sharma, and R. Soricut. 2019. ALBERT: A Lite BERT for self-supervised learning of language representations. arXiv preprint arXiv:1909.11942.","journal-title":"arXiv preprint"},{"key":"e_1_3_1_31_2","article-title":"RoBERTa: A robustly optimized BERT pre-training approach","author":"Liu Y.","year":"2019","unstructured":"Y. Liu et al. 2019. RoBERTa: A robustly optimized BERT pre-training approach. arXiv preprint arXiv:1907.11692.","journal-title":"arXiv preprint"},{"key":"e_1_3_1_32_2","first-page":"1877","article-title":"Language models are few-shot learners","volume":"33","author":"Brown T.","year":"2020","unstructured":"T. Brown et al. 2020. Language models are few-shot learners. Advances in Neural Information Processing Systems 33 (2020), 1877\u20131901.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_1_33_2","article-title":"XLNet: Generalized autoregressive pre-training for language understanding","volume":"32","author":"Yang Z.","year":"2019","unstructured":"Z. Yang, Z. Dai, Y. Yang, J. Carbonell, R. R. Salakhutdinov, and Q. V. Le. 2019. XLNet: Generalized autoregressive pre-training for language understanding. Advances in Neural Information Processing Systems 32 (2019).","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_1_34_2","doi-asserted-by":"publisher","DOI":"10.1162\/tacl_a_00300"},{"key":"e_1_3_1_35_2","article-title":"TinyBERT: Distilling BERT for natural language understanding","author":"Jiao X.","year":"2019","unstructured":"X. Jiao et al. 2019. TinyBERT: Distilling BERT for natural language understanding. arXiv preprint arXiv:1909.10351.","journal-title":"arXiv preprint"},{"key":"e_1_3_1_36_2","doi-asserted-by":"publisher","DOI":"10.5555\/3379393"},{"key":"e_1_3_1_37_2","article-title":"An analysis of simple data augmentation for named entity recognition","author":"Dai X.","year":"2020","unstructured":"X. Dai and H. Adel. 2020. An analysis of simple data augmentation for named entity recognition. arXiv preprint arXiv:2010.11683.","journal-title":"arXiv preprint"},{"key":"e_1_3_1_38_2","article-title":"A survey on data augmentation for text classification","author":"Bayer M.","year":"2021","unstructured":"M. Bayer, M.-A. Kaufhold, and C. Reuter. 2021. A survey on data augmentation for text classification. ACM Computing Surveys (2021).","journal-title":"ACM Computing Surveys"}],"container-title":["ACM Transactions on Asian and Low-Resource Language Information Processing"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3595861","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3595861","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T22:48:39Z","timestamp":1750286919000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3595861"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,6,16]]},"references-count":37,"journal-issue":{"issue":"6","published-print":{"date-parts":[[2023,6,30]]}},"alternative-id":["10.1145\/3595861"],"URL":"https:\/\/doi.org\/10.1145\/3595861","relation":{},"ISSN":["2375-4699","2375-4702"],"issn-type":[{"value":"2375-4699","type":"print"},{"value":"2375-4702","type":"electronic"}],"subject":[],"published":{"date-parts":[[2023,6,16]]},"assertion":[{"value":"2022-09-15","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2023-04-19","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2023-06-16","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}