{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,2,28]],"date-time":"2026-02-28T15:42:48Z","timestamp":1772293368613,"version":"3.50.1"},"reference-count":27,"publisher":"Association for Computing Machinery (ACM)","issue":"2","license":[{"start":{"date-parts":[[2016,11,4]],"date-time":"2016-11-04T00:00:00Z","timestamp":1478217600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"DOI":"10.13039\/501100004543","name":"China Scholarship Council","doi-asserted-by":"crossref","award":["CSC-201308655118"],"award-info":[{"award-number":["CSC-201308655118"]}],"id":[{"id":"10.13039\/501100004543","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Asian Low-Resour. Lang. Inf. Process."],"published-print":{"date-parts":[[2017,6,30]]},"abstract":"<jats:p>Morphological analysis, which includes analysis of part-of-speech (POS) tagging, stemming, and morpheme segmentation, is one of the key components in natural language processing (NLP), particularly for agglutinative languages. In this article, we investigate the morphological analysis of the Uyghur language, which is the native language of the people in the Xinjiang Uyghur autonomous region of western China. Morphological analysis of Uyghur is challenging primarily because of factors such as (1) ambiguities arising due to the likelihood of association of a multiple number of POS tags with a word stem or a multiple number of functional tags with a word suffix, (2) ambiguous morpheme boundaries, and (3) complex morphopholonogy of the language. Further, the unavailability of a manually annotated training set in the Uyghur language for the purpose of word segmentation makes Uyghur morphological analysis more difficult. In our proposed work, we address these challenges by undertaking a semisupervised approach of learning a Markov model with the help of a manually constructed dictionary of \u201csuffix to tag\u201d mappings in order to predict the most likely tag transitions in the Uyghur morpheme sequence. Due to the linguistic characteristics of Uyghur, we incorporate a prior belief in our model for favoring word segmentations with a lower number of morpheme units. Empirical evaluation of our proposed model shows an accuracy of about 82%. We further improve the effectiveness of the tag transition model with an active learning paradigm. In particular, we manually investigated a subset of words for which the model prediction ambiguity was within the top 20%. Manually incorporating rules to handle these erroneous cases resulted in an overall accuracy of 93.81%.<\/jats:p>","DOI":"10.1145\/2968410","type":"journal-article","created":{"date-parts":[[2016,11,7]],"date-time":"2016-11-07T13:40:30Z","timestamp":1478526030000},"page":"1-23","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":5,"title":["A Semisupervised Tag-Transition-Based Markovian Model for Uyghur Morphology Analysis"],"prefix":"10.1145","volume":"16","author":[{"given":"Eziz","family":"Tursun","sequence":"first","affiliation":[{"name":"Xinjiang Technical Institute of Physics and Chemistry, Chinese Academy of Science, University of Chinese Academy of Science, Institute of Mathematics and Information of Hotan Teachers College, Urumqi, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Debasis","family":"Ganguly","sequence":"additional","affiliation":[{"name":"ADAPT Centre, School of Computing, Dublin City University, Ireland"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Turghun","family":"Osman","sequence":"additional","affiliation":[{"name":"Xinjiang Technical Institute of Physics and Chemistry, Chinese Academy of Science, University of Chinese Academy of Science, Urumqi, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Ya-Ting","family":"Yang","sequence":"additional","affiliation":[{"name":"Xinjiang Technical Institute of Physics and Chemistry, Chinese Academy of Science, Urumqi, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Ghalip","family":"Abdukerim","sequence":"additional","affiliation":[{"name":"Xinjiang Technical Institute of Physics and Chemistry, Chinese Academy of Science, University of Chinese Academy of Science, Urumqi, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Jun-Lin","family":"Zhou","sequence":"additional","affiliation":[{"name":"Xinjiang Branch of Chinese Academy of Science, Urumqi, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Qun","family":"Liu","sequence":"additional","affiliation":[{"name":"ADAPT Centre, School of Computing, Dublin City University, Ireland"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2016,11,4]]},"reference":[{"key":"e_1_2_1_1_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICOSP.2010.5656065"},{"key":"e_1_2_1_2_1","first-page":"3115","article-title":"Directed graph model of Uyghur morphological analysis","volume":"23","author":"Aili Mairehaba","year":"2012","journal-title":"Ruanjian Xuebao\/Journal of Software"},{"key":"e_1_2_1_3_1","volume-title":"Natural Language Processing and Knowledge Engineering","author":"Aisha Batuer","year":"2009"},{"key":"e_1_2_1_4_1","doi-asserted-by":"publisher","DOI":"10.1214\/aoms\/1177699147"},{"key":"e_1_2_1_5_1","doi-asserted-by":"publisher","DOI":"10.1145\/1187415.1187418"},{"key":"e_1_2_1_6_1","volume-title":"Proceedings of the Second Baltic Conference on Human Language Technologies. 107--112","author":"Creutz Mathias","year":"2005"},{"key":"e_1_2_1_7_1","doi-asserted-by":"publisher","DOI":"10.5555\/1557769.1557833"},{"key":"e_1_2_1_8_1","first-page":"3","article-title":"Unsupervised morphological parsing of Bengali","volume":"40","author":"Dasgupta Sajib","year":"2006","journal-title":"Language Resources and Evaluation"},{"key":"e_1_2_1_9_1","doi-asserted-by":"publisher","DOI":"10.1162\/089120101750300490"},{"key":"e_1_2_1_10_1","volume-title":"Proceedings of the 45th Annual Meeting of the Association of Computational Linguistics. Association for Computational Linguistics, Prague, Czech Republic, 744--751","author":"Goldwater Sharon","year":"2007"},{"key":"e_1_2_1_11_1","article-title":"Joint voice harmony restoration and morphological segmentation for morphological analysis","volume":"28","author":"Haibo Zhang","year":"2014","journal-title":"Journal of Chinese Information Processing"},{"key":"e_1_2_1_12_1","unstructured":"Hemdulla Abdurahman Imam. 2011. A Brief Explanatory Dictionary of Modern Uyghur. Xinjiang Ethnic Language Work Committee.  Hemdulla Abdurahman Imam. 2011. A Brief Explanatory Dictionary of Modern Uyghur. Xinjiang Ethnic Language Work Committee."},{"key":"e_1_2_1_13_1","volume-title":"Proceedings of the Morpho Challenge 2010 Workshop. 30--34","author":"Kohonen Oskar","year":"2010"},{"key":"e_1_2_1_15_1","volume-title":"Proc. of EMNLP\u201904","volume":"4","author":"Kudo Taku","year":"2004"},{"key":"e_1_2_1_16_1","unstructured":"Christopher D. Manning and Hinrich Schutze. 1999. Foundations of Statistical Natural Language Processing. MIT Press Cambridge MA London England.   Christopher D. Manning and Hinrich Schutze. 1999. Foundations of Statistical Natural Language Processing. MIT Press Cambridge MA London England."},{"key":"e_1_2_1_17_1","first-page":"2","article-title":"Tagging English text with a probabilistic model","volume":"20","author":"Merialdo Bernard","year":"1994","journal-title":"Computational Linguistics"},{"key":"e_1_2_1_18_1","doi-asserted-by":"publisher","DOI":"10.5555\/1687878.1687894"},{"key":"e_1_2_1_19_1","volume-title":"Machine Learning: A Probabilistic Perspective","author":"Murphy Kevin P.","year":"2012"},{"key":"e_1_2_1_21_1","first-page":"33","article-title":"Rule based analysis of the Uyghur nouns","volume":"19","author":"Orhun Murat","year":"2009","journal-title":"International Journal of Asian Language Processing"},{"key":"e_1_2_1_22_1","unstructured":"Teemu Ruokolainena Oskar Kohonena Sami Virpiojaa and Mikko Kurimob. 2013. Supervised morphological segmentation in a low-resource learning setting using conditional random fields. CoNLL-2013 (2013) 29.  Teemu Ruokolainena Oskar Kohonena Sami Virpiojaa and Mikko Kurimob. 2013. Supervised morphological segmentation in a low-resource learning setting using conditional random fields. CoNLL-2013 (2013) 29."},{"key":"e_1_2_1_23_1","doi-asserted-by":"publisher","DOI":"10.3115\/v1\/E14-4017"},{"key":"e_1_2_1_24_1","volume-title":"Proceedings of the 10th Pacific Asia Conference on Language, Information and Computation. 163--172","author":"Takeuchi Kouichi","year":"1995"},{"key":"e_1_2_1_25_1","unstructured":"Litip Tohti. 2012. Modern Uyghur Reference Grammar. China Social Science Press.  Litip Tohti. 2012. Modern Uyghur Reference Grammar. China Social Science Press."},{"key":"e_1_2_1_26_1","unstructured":"Kh\u0101mit T\u00f6m\u00fcr. 2003. Modern Uyghur Grammar: Morphology. Vol. 3. Y\u0131ld\u0131z.  Kh\u0101mit T\u00f6m\u00fcr. 2003. Modern Uyghur Grammar: Morphology. Vol. 3. Y\u0131ld\u0131z."},{"key":"e_1_2_1_28_1","doi-asserted-by":"publisher","DOI":"10.1109\/18.87000"},{"key":"e_1_2_1_29_1","volume-title":"2nd IEEE International Conference on Computer Science and Information Technology","author":"Wumaier Aishan","year":"2009"},{"key":"e_1_2_1_30_1","unstructured":"Huajian Xue Yong Yang Turghun Osman Xiao Li and Ronghui Zhang. 2011. Uyghur word segmentation using a combination of rules and statistics. Advances in Information Sciences 8 Service Sciences 3 11 (2011).  Huajian Xue Yong Yang Turghun Osman Xiao Li and Ronghui Zhang. 2011. Uyghur word segmentation using a combination of rules and statistics. Advances in Information Sciences 8 Service Sciences 3 11 (2011)."}],"container-title":["ACM Transactions on Asian and Low-Resource Language Information Processing"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/2968410","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/2968410","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T03:40:10Z","timestamp":1750218010000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/2968410"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2016,11,4]]},"references-count":27,"journal-issue":{"issue":"2","published-print":{"date-parts":[[2017,6,30]]}},"alternative-id":["10.1145\/2968410"],"URL":"https:\/\/doi.org\/10.1145\/2968410","relation":{},"ISSN":["2375-4699","2375-4702"],"issn-type":[{"value":"2375-4699","type":"print"},{"value":"2375-4702","type":"electronic"}],"subject":[],"published":{"date-parts":[[2016,11,4]]},"assertion":[{"value":"2015-02-01","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2016-07-01","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2016-11-04","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}