{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,10,14]],"date-time":"2025-10-14T00:40:10Z","timestamp":1760402410581,"version":"build-2065373602"},"reference-count":49,"publisher":"MDPI AG","issue":"5","license":[{"start":{"date-parts":[[2020,5,2]],"date-time":"2020-05-02T00:00:00Z","timestamp":1588377600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"name":"Opening Foundation of the Key Laboratory of Xinjiang Uyghur Autonomous Region of China","award":["2018D04019"],"award-info":[{"award-number":["2018D04019"]}]},{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["61762084, 61662077, 61462083"],"award-info":[{"award-number":["61762084, 61662077, 61462083"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]},{"name":"Scientific Research Program of the State Language Commission of China","award":["ZDI135-54"],"award-info":[{"award-number":["ZDI135-54"]}]},{"name":"National Key Research and Development Project of China","award":["2017YFB1002103"],"award-info":[{"award-number":["2017YFB1002103"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Information"],"abstract":"<jats:p>Pattern matching is widely used in various fields such as information retrieval, natural language processing (NLP), data mining and network security. In Uyghur (a typical agglutinative, low-resource language with complex morphology, spoken by the ethnic Uyghur group in Xinjiang, China), research on pattern matching is also ongoing. Due to the language characteristics, the pattern matching using characters and words as basic units has insufficient performance. There are two problems for pattern matching: (1) vowel weakening and (2) morphological changes caused by suffixes. In view of the above problems, this paper proposes a Boyer\u2013Moore-U (BM-U) algorithm and a retrievable syllable coding format based on the syllable features of the Uyghur language and the improvement of the Boyer\u2013Moore (BM) algorithm. This algorithm uses syllable features to perform pattern matching, which effectively solves the problem of weakening vowels, and it can better match words with stem shape changes. Finally, in the pattern matching experiments based on character-encoded text and syllable-encoded text for vowel-weakened words, the BM-U algorithm precision, recall, F1-measure and accuracy are improved by 4%, 55%, 33%, 25% and 10%, 52%, 38%, 38% compared to the BM algorithm.<\/jats:p>","DOI":"10.3390\/info11050248","type":"journal-article","created":{"date-parts":[[2020,5,4]],"date-time":"2020-05-04T03:29:39Z","timestamp":1588562979000},"page":"248","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":0,"title":["Research on Uyghur Pattern Matching Based on Syllable Features"],"prefix":"10.3390","volume":"11","author":[{"given":"Wayit","family":"Abliz","sequence":"first","affiliation":[{"name":"School of Information Science and Engineering, Xinjiang University, Urumqi 830046, China"},{"name":"Key Laboratory of Multilingual Information Technology in Xinjiang Uygur Autonomous Region, Urumqi 830046, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Maihemuti","family":"Maimaiti","sequence":"additional","affiliation":[{"name":"School of Information Science and Engineering, Xinjiang University, Urumqi 830046, China"},{"name":"Key Laboratory of Multilingual Information Technology in Xinjiang Uygur Autonomous Region, Urumqi 830046, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Hao","family":"Wu","sequence":"additional","affiliation":[{"name":"Key Laboratory of Multilingual Information Technology in Xinjiang Uygur Autonomous Region, Urumqi 830046, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Jiamila","family":"Wushouer","sequence":"additional","affiliation":[{"name":"School of Information Science and Engineering, Xinjiang University, Urumqi 830046, China"},{"name":"Key Laboratory of Multilingual Information Technology in Xinjiang Uygur Autonomous Region, Urumqi 830046, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Kahaerjiang","family":"Abiderexiti","sequence":"additional","affiliation":[{"name":"School of Information Science and Engineering, Xinjiang University, Urumqi 830046, China"},{"name":"Key Laboratory of Multilingual Information Technology in Xinjiang Uygur Autonomous Region, Urumqi 830046, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Tuergen","family":"Yibulayin","sequence":"additional","affiliation":[{"name":"School of Information Science and Engineering, Xinjiang University, Urumqi 830046, China"},{"name":"Key Laboratory of Multilingual Information Technology in Xinjiang Uygur Autonomous Region, Urumqi 830046, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Aishan","family":"Wumaier","sequence":"additional","affiliation":[{"name":"School of Information Science and Engineering, Xinjiang University, Urumqi 830046, China"},{"name":"Key Laboratory of Multilingual Information Technology in Xinjiang Uygur Autonomous Region, Urumqi 830046, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2020,5,2]]},"reference":[{"key":"ref_1","first-page":"1855","article-title":"A fast improved pattern matching algorithm based on BM","volume":"28","author":"Ma","year":"2013","journal-title":"Control Decis."},{"key":"ref_2","doi-asserted-by":"crossref","unstructured":"Zhou, X., Xu, B., Qi, Y., and Li, J. (2008, January 13\u201318). MRSI: A Fast Pattern Matching Algorithm for Anti-virus Applications. Proceedings of the International Conference on Networking, Cancun, Mexico.","DOI":"10.1109\/ICN.2008.119"},{"key":"ref_3","doi-asserted-by":"crossref","first-page":"1414","DOI":"10.4028\/www.scientific.net\/AMR.532-533.1414","article-title":"A Faster Pattern Matching Algorithm for Intrusion Detection","volume":"532","author":"Du","year":"2012","journal-title":"Adv. Mater. Res."},{"key":"ref_4","doi-asserted-by":"crossref","first-page":"162","DOI":"10.1016\/j.eswa.2017.03.026","article-title":"EPMA: Efficient pattern matching algorithm for DNA sequences","volume":"80","author":"Tahir","year":"2017","journal-title":"Expert Syst. Appl."},{"key":"ref_5","doi-asserted-by":"crossref","first-page":"64","DOI":"10.1016\/j.jda.2014.10.003","article-title":"Approximate pattern matching in LZ77-compressed texts","volume":"32","author":"Gagie","year":"2015","journal-title":"J. Discret. Algorithms"},{"key":"ref_6","first-page":"119","article-title":"Study on the Some Key Technology of Improving the Quality of Uyghur Search","volume":"43","author":"Ablez","year":"2013","journal-title":"Math. Pract. Theory"},{"key":"ref_7","first-page":"236","article-title":"Sensitive information filtering algorithm based on Uyghur text information network research","volume":"54","author":"Xue","year":"2018","journal-title":"Comput. Eng. Appl."},{"key":"ref_8","first-page":"188","article-title":"Name recognition in the Uyghur language based on fuzzy matching and syllable-character conversion","volume":"57","author":"Mahmoud","year":"2017","journal-title":"J. Tsinghua Univ."},{"key":"ref_9","first-page":"50","article-title":"An Improved Method for Uyghur Sentence Similarity Computation","volume":"25","author":"Kahaerjiang","year":"2011","journal-title":"J. Chin. Inf. Process."},{"key":"ref_10","doi-asserted-by":"crossref","first-page":"762","DOI":"10.1145\/359842.359859","article-title":"A Fast String Searching Algorithm","volume":"20","author":"Boyer","year":"1977","journal-title":"Commun. Acm"},{"key":"ref_11","unstructured":"Wu, S., and Manber, U. (1994). A Fast Algorithm for Multi-Pattern Searching, University of Arizona. Technical Report TR-94-17."},{"key":"ref_12","first-page":"84","article-title":"A Boyer-Moore Type String Matching Algorithm with Memory and Its Computational Complexity","volume":"35","author":"Xiaohua","year":"2008","journal-title":"J. Hunan Univ. Nat. Sci."},{"key":"ref_13","first-page":"778","article-title":"An Improved Wu-Manber Multi-pattern Matching Algorithm for Chinese Encoding","volume":"36","author":"Yipe","year":"2015","journal-title":"J. Chin. Comput. Syst."},{"key":"ref_14","first-page":"741","article-title":"Syllable based language model for large vocabulary continuous speech recognition of Uyghur","volume":"53","author":"Nurmemet","year":"2013","journal-title":"J. Tsinghua Univ. Sci. Technol."},{"key":"ref_15","first-page":"141","article-title":"Context dependent syllable based speech synthesis system for Uyghur","volume":"47","author":"Mamateli","year":"2011","journal-title":"Comput. Eng. Appl."},{"key":"ref_16","unstructured":"Mahmut, M., and Turgun, I. (2007, January 6\u20138). A Research on Syllable Based Uyghur Text Proofreading System. Proceedings of the the Ninth National Conference on Computational Linguistics, CCL 2007, Dalian, China."},{"key":"ref_17","first-page":"193","article-title":"Acoustic Analysis on Prosodic Feature of CVC Type Syllable in Uyghur Language","volume":"37","author":"Ranagul","year":"2011","journal-title":"Comput. Eng."},{"key":"ref_18","first-page":"119","article-title":"Comparison of Uyghur and Spanish syllables","volume":"8","author":"Isabel","year":"2018","journal-title":"China Natl. Exhib."},{"key":"ref_19","first-page":"143","article-title":"Research on Multiple Pattern Matching Algorithm for Uyghur","volume":"41","author":"Dawut","year":"2015","journal-title":"Comput. Eng."},{"key":"ref_20","doi-asserted-by":"crossref","unstructured":"Tohti, T., Huang, J., Hamdulla, A., and Tan, X. (2019). Text Filtering through Multi-Pattern Matching: A Case Study of Wu\u2013Manber\u2013Uy on the Language of Uyghur. Inf. Int. Interdiscip. J., 10.","DOI":"10.3390\/info10080246"},{"key":"ref_21","doi-asserted-by":"crossref","first-page":"337","DOI":"10.1007\/11880561_28","article-title":"Phrase-Based Pattern Matching in Compressed Text","volume":"Volume 4209","author":"Culpepper","year":"2006","journal-title":"International Symposium on String Processing and Information Retrieval"},{"key":"ref_22","doi-asserted-by":"crossref","first-page":"240","DOI":"10.1016\/j.tcs.2005.11.022","article-title":"Approximate string matching using compressed suffix arrays","volume":"352","author":"Huynh","year":"2006","journal-title":"Theor. Comput. Sci."},{"key":"ref_23","first-page":"139","article-title":"Approximate string matching algorithm based on compressed suffix array","volume":"51","author":"YongKang","year":"2015","journal-title":"Comput. Eng. Appl."},{"key":"ref_24","doi-asserted-by":"crossref","first-page":"3607","DOI":"10.3906\/elk-1601-92","article-title":"A new word-based compression model allowing compressed pattern matching","volume":"25","author":"Carus","year":"2017","journal-title":"Turk. J. Electr. Eng. Comput. Sci."},{"key":"ref_25","doi-asserted-by":"crossref","first-page":"313","DOI":"10.1016\/S1570-8667(03)00032-7","article-title":"Approximate string matching on Ziv-Lempel compressed text","volume":"1","author":"Karkkainen","year":"2003","journal-title":"J. Discret. Algorithms"},{"key":"ref_26","doi-asserted-by":"crossref","first-page":"717","DOI":"10.1016\/j.compbiomed.2004.06.002","article-title":"Assessment of approximate string matching in a biomedical text retrieval problem","volume":"35","author":"Wang","year":"2005","journal-title":"Comput. Biol. Med."},{"key":"ref_27","doi-asserted-by":"crossref","first-page":"1107","DOI":"10.1002\/spe.663","article-title":"LZgrep: A Boyer\u2013Moore string matching tool for Ziv\u2013Lempel compressed text","volume":"35","author":"Navarro","year":"2005","journal-title":"Softw. Pract. Exp."},{"key":"ref_28","first-page":"59","article-title":"Research of BWT-Boyer-Moore Compressed Domain Search Algorithm","volume":"23","author":"Quanzhu","year":"2006","journal-title":"Appl. Res. Comput."},{"key":"ref_29","doi-asserted-by":"crossref","unstructured":"Narupiyakul, L., Thomas, C., Cercone, N., and Sirinaovakul, B. (2004, January 15\u201321). Thai Syllable-Based Information Extraction Using Hidden Markov Models. Proceedings of the Conference on Intelligent Text Processing and Computational Linguistics, CICLing 2004, Seoul, Korea.","DOI":"10.1007\/978-3-540-24630-5_67"},{"key":"ref_30","unstructured":"Hackett, P.G., and Oard, D.W. (October, January 30). Comparison of word-based and syllable-based retrieval for Tibetan (poster session). Proceedings of the Fifth International Workshop on Information Retrieval with Asian Languages, Hong Kong, China."},{"key":"ref_31","doi-asserted-by":"crossref","unstructured":"Oflazer, K., and Kuruoz, I. (1994, January 13\u201315). Tagging and Morphological Disambiguation of Turkish Text. Proceedings of the Conference on Applied Natural Language Processing, Stuttgart, Germany.","DOI":"10.3115\/974358.974391"},{"key":"ref_32","doi-asserted-by":"crossref","first-page":"381","DOI":"10.1023\/A:1020271707826","article-title":"Statistical Morphological Disambiguation for Agglutinative Languages","volume":"36","author":"Hakkanitur","year":"2002","journal-title":"Comput. Humanit."},{"key":"ref_33","doi-asserted-by":"crossref","first-page":"328","DOI":"10.1049\/cje.2016.03.020","article-title":"Agglutinative Language Speech Recognition Using Automatic Allophone Deriving","volume":"25","author":"Xu","year":"2016","journal-title":"Chin. J. Electron."},{"key":"ref_34","doi-asserted-by":"crossref","first-page":"101604","DOI":"10.1016\/j.datak.2017.07.007","article-title":"Constructing a paraphrase database for agglutinative languages","volume":"123","author":"Park","year":"2019","journal-title":"Data Knowl. Eng."},{"key":"ref_35","doi-asserted-by":"crossref","unstructured":"Saimaiti, A., Wang, L., and Yibulayin, T. (2019). Learning Subword Embedding to Improve Uyghur Named-Entity Recognition. Inf. Int. Interdiscip. J., 10.","DOI":"10.3390\/info10040139"},{"key":"ref_36","first-page":"43","article-title":"A Morphological Analysis Based Algorithm for Uyghur Vowel Weakening Identification","volume":"22","author":"Mireguli","year":"2008","journal-title":"J. Chin. Inf. Process."},{"key":"ref_37","first-page":"116","article-title":"Study on the Rule-based Kazakh Word Lemmatization Algorithm","volume":"28","author":"Dawel","year":"2011","journal-title":"J. Xinjiang Univ. Nat. Sci. Ed."},{"key":"ref_38","first-page":"29","article-title":"Research on the Causes of the Weakening and Even Disappearance of Short Vowels in Mongolian","volume":"31","author":"Saren","year":"2005","journal-title":"J. Inn. Mong. Univ. Natl. Soc. Sci."},{"key":"ref_39","doi-asserted-by":"crossref","first-page":"822","DOI":"10.1353\/lan.0.0169","article-title":"Natural and Unnatural Constraints in Hungarian Vowel Harmony","volume":"85","author":"Hayes","year":"2009","journal-title":"Language"},{"key":"ref_40","doi-asserted-by":"crossref","first-page":"859","DOI":"10.1007\/s11049-012-9169-1","article-title":"Information theoretic approaches to phonological structure: The case of Finnish vowel harmony","volume":"30","author":"Goldsmith","year":"2012","journal-title":"Nat. Lang. Linguist. Theory"},{"key":"ref_41","first-page":"27","article-title":"The Weakening and Dropping of Vowels in Mongolian Language","volume":"36","author":"Genxiong","year":"2010","journal-title":"J. Inn. Mong. Univ. Natl. Soc. Sci."},{"key":"ref_42","first-page":"154","article-title":"Experimental Study on the Characteristics of the Weakened Syllables of Jingpo","volume":"5","author":"Qingxia","year":"2014","journal-title":"J. Minzu Univ. China Philos. Soc. Sci. Ed."},{"key":"ref_43","first-page":"51","article-title":"Phonetic and Phonological Vowel Reduction in Russian","volume":"46","author":"Jaworski","year":"2010","journal-title":"Pozn. Stud. Contemp. Linguist."},{"key":"ref_44","first-page":"1","article-title":"Vowel harmony and stem identity","volume":"1","author":"Bakovic","year":"2003","journal-title":"San Diego Linguistic Papers."},{"key":"ref_45","unstructured":"Xinjiang Uygur Autonomous Region National Language Working Committee (1997). Dictionary of Modern Uyghur Literature Language Orthography, Xinjiang People\u2019s Publishing House."},{"key":"ref_46","first-page":"957","article-title":"Modern Uyghur automatic syllable segmentation method and its implementation","volume":"10","author":"Wayit","year":"2015","journal-title":"China Sci."},{"key":"ref_47","doi-asserted-by":"crossref","unstructured":"Abliz, W., Wu, H., Maimaiti, M., Wushouer, J., Abiderexiti, K., Yibulayin, T., and Wumaier, A. (2020). A Syllable-Based Technique for Uyghur Text Compression. Inf. Int. Interdiscip. J., 11.","DOI":"10.3390\/info11030172"},{"key":"ref_48","first-page":"27","article-title":"Rules and Algorithms for Uyghur Affix Variant Collocation","volume":"32","author":"Ainiwaer","year":"2018","journal-title":"J. Chin. Inf. Process."},{"key":"ref_49","first-page":"1","article-title":"A Survey of Central Asian Language Processing","volume":"32","author":"Tuergen","year":"2018","journal-title":"J. Chin. Inf. Process."}],"container-title":["Information"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/2078-2489\/11\/5\/248\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,13]],"date-time":"2025-10-13T13:32:29Z","timestamp":1760362349000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/2078-2489\/11\/5\/248"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2020,5,2]]},"references-count":49,"journal-issue":{"issue":"5","published-online":{"date-parts":[[2020,5]]}},"alternative-id":["info11050248"],"URL":"https:\/\/doi.org\/10.3390\/info11050248","relation":{},"ISSN":["2078-2489"],"issn-type":[{"type":"electronic","value":"2078-2489"}],"subject":[],"published":{"date-parts":[[2020,5,2]]}}}