{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,10,12]],"date-time":"2025-10-12T02:47:01Z","timestamp":1760237221132,"version":"build-2065373602"},"reference-count":32,"publisher":"MDPI AG","issue":"3","license":[{"start":{"date-parts":[[2020,3,23]],"date-time":"2020-03-23T00:00:00Z","timestamp":1584921600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"name":"Opening Foundation of the Key Laboratory of Xinjiang Uyghur Autonomous Region of China","award":["2018D04019"],"award-info":[{"award-number":["2018D04019"]}]},{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["61762084, 61662077, 61462083"],"award-info":[{"award-number":["61762084, 61662077, 61462083"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]},{"name":"the Scientific Research Program of the State Language Commission of China","award":["ZDI135-54"],"award-info":[{"award-number":["ZDI135-54"]}]},{"name":"National Key Research and Development Project of China","award":["2017YFB1002103"],"award-info":[{"award-number":["2017YFB1002103"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Information"],"abstract":"<jats:p>To improve utilization of text storage resources and efficiency of data transmission, we proposed two syllable-based Uyghur text compression coding schemes. First, according to the statistics of syllable coverage of the corpus text, we constructed a 12-bit and 16-bit syllable code tables and added commonly used symbols\u2014such as punctuation marks and ASCII characters\u2014to the code tables. To enable the coding scheme to process Uyghur texts mixed with other language symbols, we introduced a flag code in the compression process to distinguish the Unicode encodings that were not in the code table. The experiments showed that the 12-bit coding scheme had an average compression ratio of 0.3 on Uyghur text less than 4 KB in size and that the 16-bit coding scheme had an average compression ratio of 0.5 on text less than 2 KB in size. Our compression schemes outperformed GZip, BZip2, and the LZW algorithm on short text and could be effectively applied to the compression of Uyghur short text for storage and applications.<\/jats:p>","DOI":"10.3390\/info11030172","type":"journal-article","created":{"date-parts":[[2020,3,24]],"date-time":"2020-03-24T07:16:08Z","timestamp":1585034168000},"page":"172","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":5,"title":["A Syllable-Based Technique for Uyghur Text Compression"],"prefix":"10.3390","volume":"11","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-6903-5311","authenticated-orcid":false,"given":"Wayit","family":"Abliz","sequence":"first","affiliation":[{"name":"School of Information Science and Engineering, Xinjiang University, Urumqi 830046, China"},{"name":"Xinjiang Laboratory of Multi-Language Information Technology, Xinjiang University, Urumqi 830046, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Hao","family":"Wu","sequence":"additional","affiliation":[{"name":"Xinjiang Laboratory of Multi-Language Information Technology, Xinjiang University, Urumqi 830046, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Maihemuti","family":"Maimaiti","sequence":"additional","affiliation":[{"name":"School of Information Science and Engineering, Xinjiang University, Urumqi 830046, China"},{"name":"Xinjiang Laboratory of Multi-Language Information Technology, Xinjiang University, Urumqi 830046, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Jiamila","family":"Wushouer","sequence":"additional","affiliation":[{"name":"School of Information Science and Engineering, Xinjiang University, Urumqi 830046, China"},{"name":"Xinjiang Laboratory of Multi-Language Information Technology, Xinjiang University, Urumqi 830046, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Kahaerjiang","family":"Abiderexiti","sequence":"additional","affiliation":[{"name":"School of Information Science and Engineering, Xinjiang University, Urumqi 830046, China"},{"name":"Xinjiang Laboratory of Multi-Language Information Technology, Xinjiang University, Urumqi 830046, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Tuergen","family":"Yibulayin","sequence":"additional","affiliation":[{"name":"School of Information Science and Engineering, Xinjiang University, Urumqi 830046, China"},{"name":"Xinjiang Laboratory of Multi-Language Information Technology, Xinjiang University, Urumqi 830046, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Aishan","family":"Wumaier","sequence":"additional","affiliation":[{"name":"School of Information Science and Engineering, Xinjiang University, Urumqi 830046, China"},{"name":"Xinjiang Laboratory of Multi-Language Information Technology, Xinjiang University, Urumqi 830046, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2020,3,23]]},"reference":[{"key":"ref_1","unstructured":"David, S., and Le-Nan, W. (2003). Data Compression, Publishing House of Electronics Industry. [2nd ed.]."},{"key":"ref_2","doi-asserted-by":"crossref","first-page":"379","DOI":"10.1002\/j.1538-7305.1948.tb01338.x","article-title":"A mathematical theory of communication","volume":"27","author":"Shannon","year":"1948","journal-title":"Bell Labs Tech. J."},{"key":"ref_3","doi-asserted-by":"crossref","first-page":"91","DOI":"10.1007\/BF02837279","article-title":"A method for the construction of minimum-redundancy codes","volume":"11","author":"Huffman","year":"2006","journal-title":"Resonance"},{"key":"ref_4","doi-asserted-by":"crossref","first-page":"337","DOI":"10.1109\/TIT.1977.1055714","article-title":"A universal algorithm for sequential data compression","volume":"23","author":"Ziv","year":"1977","journal-title":"IEEE Trans. Inf. Theory"},{"key":"ref_5","doi-asserted-by":"crossref","first-page":"530","DOI":"10.1109\/TIT.1978.1055934","article-title":"Compression of individual sequences via variable-rate coding","volume":"24","author":"Ziv","year":"1978","journal-title":"IEEE Trans. Inf. Theory"},{"key":"ref_6","doi-asserted-by":"crossref","first-page":"87","DOI":"10.1177\/016555159502100203","article-title":"A new text compression technique based on language structure","volume":"21","author":"Akman","year":"1995","journal-title":"J. Inf. Sci."},{"key":"ref_7","doi-asserted-by":"crossref","first-page":"185","DOI":"10.1002\/spe.4380190207","article-title":"Word-based text compression","volume":"19","author":"Moffat","year":"1989","journal-title":"Softw. Pract. Exp."},{"key":"ref_8","doi-asserted-by":"crossref","first-page":"1175","DOI":"10.1016\/j.ipm.2004.08.009","article-title":"Word-based text compression using the Burrows-Wheeler transform","volume":"41","author":"Moffat","year":"2005","journal-title":"Inf. Process. Manag."},{"key":"ref_9","doi-asserted-by":"crossref","first-page":"55","DOI":"10.1016\/j.tcs.2014.01.013","article-title":"Note on the greedy parsing optimality for dictionary-based text compression","volume":"525","author":"Crochemore","year":"2014","journal-title":"Theor. Comput. Sci."},{"key":"ref_10","first-page":"497","article-title":"Improving Compression of Short Messages","volume":"6","author":"Bettison","year":"2013","journal-title":"Int. J. Commun. Netw. Syst. Sci."},{"key":"ref_11","doi-asserted-by":"crossref","first-page":"410","DOI":"10.1016\/j.aei.2008.05.001","article-title":"Compression of small text files","volume":"22","author":"Platos","year":"2008","journal-title":"Adv. Eng. Inform."},{"key":"ref_12","doi-asserted-by":"crossref","first-page":"1","DOI":"10.4304\/jcp.1.6.1-10","article-title":"Compression of Short Text on Embedded Systems","volume":"1","author":"Rein","year":"2006","journal-title":"J. Comput."},{"key":"ref_13","doi-asserted-by":"crossref","first-page":"53","DOI":"10.1016\/j.csi.2014.05.005","article-title":"Rapid lossless compression of short text messages","volume":"37","author":"Kalajdzic","year":"2015","journal-title":"Comput. Stand. Interfaces"},{"key":"ref_14","doi-asserted-by":"crossref","first-page":"3089","DOI":"10.1007\/s13369-016-2070-1","article-title":"Syllable-Based Text Compression: A Language Case Study","volume":"41","author":"Adubi","year":"2016","journal-title":"Arab. J. Sci. Eng."},{"key":"ref_15","doi-asserted-by":"crossref","unstructured":"Nguyen, V.H., Nguyen, H.T., Duong, H.N., and Snasel, V. (2016). A Syllable-Based Method for Vietnamese Text Compression. International Conference on Ubiquitous Information Management and Communication, ACM.","DOI":"10.1145\/2857546.2857564"},{"key":"ref_16","first-page":"66","article-title":"A lossless text compression technique using syllable based morphology","volume":"8","author":"Akman","year":"2011","journal-title":"Int. Arab J. Inf. Technol."},{"key":"ref_17","doi-asserted-by":"crossref","unstructured":"Oswald, C., Ajith, K.J., and Avinash, J. (2017). A Graph-Based Frequent Sequence Mining Approach to Text Compression. Mining Intelligence and Knowledge Exploration, Proceedings of the International Conference on Mining Intelligence and Knowledge Exploration, Hyderabad, India, 13\u201315 December 2017, Springer.","DOI":"10.1007\/978-3-319-71928-3_35"},{"key":"ref_18","doi-asserted-by":"crossref","unstructured":"Bharathi, K., Kumar, H., and Fairouz, A. (2018, January 7\u201310). A Plain-Text Incremental Compression (PIC) Technique with Fast Lookup Ability. Proceedings of the 2018 IEEE 36th International Conference on Computer Design (ICCD), Orlando, FL, USA.","DOI":"10.1109\/ICCD.2018.00065"},{"key":"ref_19","first-page":"85","article-title":"Lossless Text Compression Technique Based on Static Dictionary for Unicode Tamil Document","volume":"118","author":"Vijayalakshmi","year":"2018","journal-title":"Int. J. Pure Appl. Math."},{"key":"ref_20","doi-asserted-by":"crossref","unstructured":"Sinaga, A., and Adiwijaya Nugroho, H. (2015, January 27\u201329). Development of word-based text compression algorithm for Indonesian language document. Proceedings of the 2015 3rd International Conference on Information and Communication Technology, ICoICT, Nusa Dua, Indonesia.","DOI":"10.1109\/ICoICT.2015.7231466"},{"key":"ref_21","doi-asserted-by":"crossref","unstructured":"Farhad Mokter, M., Akter, S., Palash Uddin, M., Ibn Afjal, M., Al Mamun, M., and Abu Marjan, M. (2018, January 21\u201322). An Efficient Technique for Representation and Compression of Bengali Text. Proceedings of the 2018 International Conference on Bangla Speech and Language Processing, ICBSLP, Sylhet, Bangladesh.","DOI":"10.1109\/ICBSLP.2018.8554838"},{"key":"ref_22","doi-asserted-by":"crossref","first-page":"618","DOI":"10.1016\/j.procs.2019.08.087","article-title":"Arabic text lossless compression by characters encoding","volume":"155","author":"Hilal","year":"2019","journal-title":"Procedia Comput. Sci."},{"key":"ref_23","first-page":"112","article-title":"Compression algorithm LZW on Chinese text","volume":"50","author":"Chen","year":"2014","journal-title":"Comput. Eng. Appl."},{"key":"ref_24","doi-asserted-by":"crossref","unstructured":"Satoh, N., Morihara, T., Okada, Y., and Yoshida, S. (1997, January 25\u201327). Study of Japanese text compression. Proceedings of the Data Compression Conference, DCC \u201997, Snowbird, UT, USA.","DOI":"10.1109\/DCC.1997.582134"},{"key":"ref_25","first-page":"9","article-title":"Uyghur text compression technique research","volume":"29","author":"Winira","year":"2012","journal-title":"J. Xinjiang Univ."},{"key":"ref_26","doi-asserted-by":"crossref","unstructured":"Maimaiti, M., Wumaier, A., Abiderexiti, K., and Yibulayin, T. (2017). Bidirectional long short-term memory network with a conditional random field layer for Uyghur part-of-speech tagging. Information, 8.","DOI":"10.3390\/info8040157"},{"key":"ref_27","first-page":"47","article-title":"Collaborative Analysis of Uyghur Morphology Based on Character Level","volume":"55","author":"Osman","year":"2019","journal-title":"Acta Sci. Nat. Univ. Pekin."},{"key":"ref_28","first-page":"741","article-title":"Syllable based language model for large vocabulary continuous speech recognition of Uyghur","volume":"53","author":"Nurmemet","year":"2013","journal-title":"J. Tsinghua Univ."},{"key":"ref_29","first-page":"957","article-title":"Modern Uyghur automatic syllable segmentation method and its implementation","volume":"10","author":"Wayit","year":"2015","journal-title":"China Sci."},{"key":"ref_30","unstructured":"(2020, January 10). LZW Data Compression. Available online: http:\/\/dogma.net\/markn\/articles\/lzw\/lzw.htm\/."},{"key":"ref_31","unstructured":"(2020, January 10). GZipStream Class (System.IO.Compression). Available online: https:\/\/docs.microsoft.com\/zh-cn\/dotnet\/api\/system.io.compression.gzipstream?view=netframework-4.8\/."},{"key":"ref_32","first-page":"1","article-title":"A Survey of Central Asian Language Processing","volume":"32","author":"Tuergen","year":"2018","journal-title":"Chin. Inf. Process"}],"container-title":["Information"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/2078-2489\/11\/3\/172\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,11]],"date-time":"2025-10-11T09:10:53Z","timestamp":1760173853000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/2078-2489\/11\/3\/172"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2020,3,23]]},"references-count":32,"journal-issue":{"issue":"3","published-online":{"date-parts":[[2020,3]]}},"alternative-id":["info11030172"],"URL":"https:\/\/doi.org\/10.3390\/info11030172","relation":{},"ISSN":["2078-2489"],"issn-type":[{"type":"electronic","value":"2078-2489"}],"subject":[],"published":{"date-parts":[[2020,3,23]]}}}