{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T04:19:52Z","timestamp":1750220392096,"version":"3.41.0"},"reference-count":21,"publisher":"Association for Computing Machinery (ACM)","issue":"2","license":[{"start":{"date-parts":[[2021,11,18]],"date-time":"2021-11-18T00:00:00Z","timestamp":1637193600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Asian Low-Resour. Lang. Inf. Process."],"published-print":{"date-parts":[[2022,3,31]]},"abstract":"<jats:p>Hindi is the third most-spoken language in the world (615 million speakers) and has the fourth highest native speakers (341 million). It is an inflectionally rich and relatively free word-order language with an immense vocabulary set. Despite being such a celebrated language across the globe, very few Natural Language Processing (NLP) applications and tools have been developed to support it computationally. Moreover, most of the existing ones are not efficient enough due to the lack of semantic information (or contextual knowledge). Hindi grammar is based on Paninian grammar and derives most of its rules from it. Paninian grammar very aggressively highlights the role of karaka theory in free-word order languages. In this article, we present an application that extracts all possible karakas from simple Hindi sentences with an accuracy of 84.2% and an F1 score of 88.5%. We consider features such as Parts of Speech tags, post-position markers (vibhaktis), semantic tags for nouns and syntactic structure to grab the context in different-sized word windows within a sentence. With the help of these features, we built a rule-based inference engine to extract karakas from a sentence. The application takes in a text file with clean (without punctuation) simple Hindi sentences and gives back karaka tagged sentences in a separate text file as output.<\/jats:p>","DOI":"10.1145\/3479155","type":"journal-article","created":{"date-parts":[[2021,11,18]],"date-time":"2021-11-18T20:41:51Z","timestamp":1637268111000},"page":"1-25","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":0,"title":["Development of Automatic Rule-based Semantic Tagger and Karaka Analyzer for Hindi"],"prefix":"10.1145","volume":"21","author":[{"given":"Pragya","family":"Katyayan","sequence":"first","affiliation":[{"name":"Department of Computer Science, Banasthali Vidyapith, Rajasthan, India"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Nisheeth","family":"Joshi","sequence":"additional","affiliation":[{"name":"Department of Computer Science, Banasthali Vidyapith, Rajasthan, India"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2021,11,18]]},"reference":[{"key":"e_1_3_2_2_2","doi-asserted-by":"crossref","unstructured":"Christine Everaert. 2010. Tracing the Boundaries between Hindi and Urdu: Lost and Added in Translation between 20th Century Short Stories Vol. 32 Brill.","DOI":"10.1163\/ej.9789004177314.i-300"},{"key":"e_1_3_2_3_2","doi-asserted-by":"publisher","DOI":"10.1007\/BF02760393"},{"key":"e_1_3_2_4_2","unstructured":"Vineet Chaitanya Rajiv Sangal and Akshar Bharati. 1996. Natural Language Processing: A Paninian Perspective . Prentice-Hall of India."},{"key":"e_1_3_2_5_2","unstructured":"Rajesh Kumar Kiran Raj and Abhinav Yadav. 2013. PoS tagging and CYK Parsing for Indian Languages. Retrieved October 18 2021 from https:\/\/github.com\/rajesh-iiith\/POS-Tagging-and-CYK-Parsing-for-Indian-Languages."},{"key":"e_1_3_2_6_2","unstructured":"Kamal Sathyarthe Ravi Prakash Gupt and Dipti Prakash. 2012. Manak Hindi Vyakaran Evam Rachana\u2014Class 9 and 10 (Course-A) . New Saraswati House India Pvt. Ltd. New Delhi."},{"key":"e_1_3_2_7_2","doi-asserted-by":"publisher","DOI":"10.5120\/21716-4841"},{"key":"e_1_3_2_8_2","doi-asserted-by":"publisher","DOI":"10.3115\/991146.991151"},{"key":"e_1_3_2_9_2","volume-title":"Proceedings of the Workshop on Recent Advances in Dependency Grammar","author":"Pedersen Mark","year":"2004","unstructured":"Mark Pedersen, Domenyk Eades, Samir K. Amin, and Lakshmi Prakash. 2004. Relative clauses in Hindi and Arabic: A Paninian dependency grammar analysis. In Proceedings of the Workshop on Recent Advances in Dependency Grammar. ACL, Geneva (Switzerland), 9\u201316."},{"key":"e_1_3_2_10_2","doi-asserted-by":"publisher","DOI":"10.5555\/1868771.1868782"},{"key":"e_1_3_2_11_2","volume-title":"arXiv preprint","author":"Anuranjana Kaveri","year":"2019","unstructured":"Kaveri Anuranjana, Vijjini Anvesh Rao, and Radhika Mamidi. 2019. Hindi question generation using dependency structures. arXiv preprint arXiv:1906.08570."},{"key":"e_1_3_2_12_2","volume-title":"Proceedings of the 10th International Conference on Language Resources and Evaluation (LREC'16)","author":"Nomani Maaz A.","year":"2016","unstructured":"Maaz A. Nomani and Dipti Misra Sharma. 2016. Towards building semantic role labeler for Indian languages. In Proceedings of the 10th International Conference on Language Resources and Evaluation (LREC'16). ELRA, Portoroz, 4588\u20134595."},{"key":"e_1_3_2_13_2","volume-title":"Proeedings of the 6th International Conference on Natural Language Processing (ICON.08), NLPAI, CDAC Pune (India)","author":"Gupta Mridul","year":"2008","unstructured":"Mridul Gupta, Vineet Yadav, Samar Husain, and Dipti M. Sharma. 2008. A rule based approach for automatic annotation of a Hindi treebank. In Proeedings of the 6th International Conference on Natural Language Processing (ICON.08), NLPAI, CDAC Pune (India). 1\u201310."},{"key":"e_1_3_2_14_2","article-title":"Improving neural machine translation for low-resource Indian languages using rule-based feature extraction","author":"Singh Muskaan","year":"2020","unstructured":"Muskaan Singh, Ravinder Kumar, and Inderveer Chana. 2020. Improving neural machine translation for low-resource Indian languages using rule-based feature extraction. In Neural Computing & Applications, Vol. 33. Springer, 1103\u20131122.","journal-title":"Neural Computing & Applications"},{"key":"e_1_3_2_15_2","doi-asserted-by":"publisher","DOI":"10.5555\/1698381.1698417"},{"key":"e_1_3_2_16_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/D19-5538"},{"key":"e_1_3_2_17_2","doi-asserted-by":"publisher","DOI":"10.4018\/JITR.2015070102"},{"key":"e_1_3_2_18_2","first-page":"224","article-title":"Sanskrit karaka analyzer for machine translation","author":"Mishra Sudhir K.","year":"2007","unstructured":"Sudhir K. Mishra and Girish Nath Jha. 2007. Sanskrit karaka analyzer for machine translation. SPLASH Proceedings of iSTRANS. 224\u2013225.","journal-title":"SPLASH Proceedings of iSTRANS"},{"key":"e_1_3_2_19_2","first-page":"4","article-title":"Karaka analysis of complicated Sanskrit sentences","author":"Mishra Sudhir K.","year":"2017","unstructured":"Sudhir K. Mishra. 2017. Karaka analysis of complicated Sanskrit sentences. Vagarthah: An International Journal of Sanskrit Research I(II). 4\u20137.","journal-title":"Vagarthah: An International Journal of Sanskrit Research"},{"key":"e_1_3_2_20_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-642-00155-0_9"},{"key":"e_1_3_2_21_2","volume-title":"7th International Conference on Natural Language Processing","author":"Palmer Martha","year":"2009","unstructured":"Martha Palmer, Rajesh Bhatt, Bhuvana Narasimhan, Owen Rambow, Dipti Misra Sharma, and Fei Xia. 2009. Hindi syntax: Annotating dependency, lexical predicate-argument structure, and phrase structure. In 7th International Conference on Natural Language Processing. NLPAI, Hyderabad (India), 14\u201317."},{"key":"e_1_3_2_22_2","doi-asserted-by":"publisher","DOI":"10.5555\/555733"}],"container-title":["ACM Transactions on Asian and Low-Resource Language Information Processing"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3479155","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3479155","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T20:18:38Z","timestamp":1750191518000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3479155"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2021,11,18]]},"references-count":21,"journal-issue":{"issue":"2","published-print":{"date-parts":[[2022,3,31]]}},"alternative-id":["10.1145\/3479155"],"URL":"https:\/\/doi.org\/10.1145\/3479155","relation":{},"ISSN":["2375-4699","2375-4702"],"issn-type":[{"type":"print","value":"2375-4699"},{"type":"electronic","value":"2375-4702"}],"subject":[],"published":{"date-parts":[[2021,11,18]]},"assertion":[{"value":"2020-12-01","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2021-08-01","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2021-11-18","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}