{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,3,20]],"date-time":"2026-03-20T20:03:33Z","timestamp":1774037013621,"version":"3.50.1"},"reference-count":42,"publisher":"Association for Computing Machinery (ACM)","issue":"6","license":[{"start":{"date-parts":[[2022,11,30]],"date-time":"2022-11-30T00:00:00Z","timestamp":1669766400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"name":"Ministry of Education - China Mobile Research Foundation","award":["MCM20170206"],"award-info":[{"award-number":["MCM20170206"]}]},{"name":"The Fundamental Research Funds for the Central Universities","award":["lzujbky-2019-kb51 and lzujbky-2018-k12"],"award-info":[{"award-number":["lzujbky-2019-kb51 and lzujbky-2018-k12"]}]},{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"crossref","award":["61402210"],"award-info":[{"award-number":["61402210"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]},{"name":"Major National Project of High Resolution Earth Observation System","award":["30-Y20A34-9010-15\/17"],"award-info":[{"award-number":["30-Y20A34-9010-15\/17"]}]},{"name":"State Grid Corporation of China Science and Technology Project","award":["SGGSKY00WYJS2000062"],"award-info":[{"award-number":["SGGSKY00WYJS2000062"]}]},{"DOI":"10.13039\/501100004602","name":"Program for New Century Excellent Talents in University","doi-asserted-by":"crossref","award":["NCET-12-0250"],"award-info":[{"award-number":["NCET-12-0250"]}],"id":[{"id":"10.13039\/501100004602","id-type":"DOI","asserted-by":"crossref"}]},{"name":"Strategic Priority Research Program of the Chinese Academy of Sciences","award":["XDA03030100"],"award-info":[{"award-number":["XDA03030100"]}]},{"name":"Google Research Awards and Google Faculty Award"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Asian Low-Resour. Lang. Inf. Process."],"published-print":{"date-parts":[[2022,11,30]]},"abstract":"<jats:p>\n            Tibetan is a low-resource language with few existing electronic reference materials. The goal of Tibetan\n            <jats:bold>sentence boundary disambiguation (SBD)<\/jats:bold>\n            is to segment long text into sentences, and it is the foundation for downstream tasks corpora building. This study implemented the Tibetan SBD at the syllable level to avoid\n            <jats:bold>word segmentation (WS)<\/jats:bold>\n            errors affecting the accuracy of SBD. Specifically, the attention mechanism is introduced based on a\n            <jats:bold>recurrent neural network (RNN)<\/jats:bold>\n            to study Tibetan SBD. The primary objective is to determine, using a trained model, whether the shad contained in Tibetan text is the ending of the sentence, and implement experiments on syllable embedding and component embedding to measure the model's performance. The highest accuracy for Tibetan syllable embedding and component embedding is 96.23% and 95.40 %, respectively, and the F1 score reaches 96.23% and 95.37%, respectively. The experimental results demonstrate that the proposed method can achieve better results than the established rule-based and statistical methods without considering various syntactic and\n            <jats:bold>part-of-speech (POS)<\/jats:bold>\n            tagging rules. German and English data from the Europarl corpus and Thai data from the IWSLT2015 corpus are validated to prove the models\u2019 reliability and generalizability. The results demonstrate that this method is efficient not only for low-resource languages but also for high-resource languages. More importantly, we can formally apply the experimental results of this study to the research of downstream tasks, such as machine translation and automatic summarization.\n          <\/jats:p>","DOI":"10.1145\/3527663","type":"journal-article","created":{"date-parts":[[2022,4,1]],"date-time":"2022-04-01T11:41:09Z","timestamp":1648813269000},"page":"1-18","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":2,"title":["Sentence Boundary Disambiguation for Tibetan Based on Attention Mechanism at the Syllable Level"],"prefix":"10.1145","volume":"21","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-6051-0737","authenticated-orcid":false,"given":"Fenfang","family":"Li","sequence":"first","affiliation":[{"name":"School of Information Science and Engineering, Lanzhou University, Lanzhou, Gansu, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-0623-3556","authenticated-orcid":false,"given":"Hui","family":"Lv","sequence":"additional","affiliation":[{"name":"School of Information Science and Engineering, Lanzhou University, Lanzhou, Gansu, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-4511-7152","authenticated-orcid":false,"given":"Duo","family":"La","sequence":"additional","affiliation":[{"name":"Key Lab of China's National Linguistic Information Technology, Northwest University for Nationalities, Lanzhou, Gansu, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-6945-9122","authenticated-orcid":false,"given":"Binbin","family":"Yong","sequence":"additional","affiliation":[{"name":"School of Information Science and Engineering, Lanzhou University, Lanzhou, Gansu, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-8054-5446","authenticated-orcid":false,"given":"Qingguo","family":"Zhou","sequence":"additional","affiliation":[{"name":"School of Information Science and Engineering, Lanzhou University, Lanzhou, Gansu, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2023,2,21]]},"reference":[{"key":"e_1_3_1_2_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCCIS48478.2019.8974495"},{"key":"e_1_3_1_3_1","article-title":"Detection of sentence boundaries and abbreviations in clinical narratives","author":"Markus Kreuzthaler","year":"2015","unstructured":"Markus Kreuzthaler and Stefan Schulz. 2015. Detection of sentence boundaries and abbreviations in clinical narratives. BMC Medical Informatics Decision Making, 2015.","journal-title":"BMC Medical Informatics Decision Making"},{"issue":"6","key":"e_1_3_1_4_1","first-page":"224","article-title":"Dependency parsing of Tibetan compound sentence","volume":"30","year":"2016","unstructured":"Quecairang Hua and Haixing Zhao. 2016. Dependency parsing of Tibetan compound sentence. Journal of Chinese Information Processing 30, 6 (2016), 224--229.","journal-title":"Journal of Chinese Information Processing"},{"issue":"6","key":"e_1_3_1_5_1","first-page":"42","article-title":"Semantic block recognition method for Tibetan sentences","volume":"33","author":"Rou Te","year":"2019","unstructured":"Te Rou, Chajia Se, and Rangjia Cai. 2019. Semantic block recognition method for Tibetan sentences. Journal of Chinese Information Processing 33, 6 (2019), 42--49.","journal-title":"Journal of Chinese Information Processing"},{"issue":"2","key":"e_1_3_1_6_1","first-page":"33","article-title":"Tibetan word segmentation strategy and algorithm based on part-of-speech constraints","volume":"34","author":"Cai Rangzhuoma","year":"2020","unstructured":"Rangzhuoma Cai and Zhijie Cai. 2020. Tibetan word segmentation strategy and algorithm based on part-of-speech constraints. Journal of Chinese Information Processing 34, 2 (2020), 33--37.","journal-title":"Journal of Chinese Information Processing"},{"issue":"4","key":"e_1_3_1_7_1","first-page":"68","article-title":"Tibetan poem generation with attention based encoder-decoder model","volume":"33","year":"2019","unstructured":"Chajia Se, Guocairang Hua, Rangjia Cai, Zhenjiacuo Ci, and Te Rou. 2019. Tibetan poem generation with attention based encoder-decoder model. Journal of Chinese Information Processing 33, 4 (2019), 68--74.","journal-title":"Journal of Chinese Information Processing"},{"issue":"5","key":"e_1_3_1_8_1","first-page":"406","article-title":"Tibetan syllable segmentation based on mixed mode","volume":"48","author":"Cai Rangdangzhi","year":"2019","unstructured":"Rangdangzhi Cai and Quecairang Hua. 2019. Tibetan syllable segmentation based on mixed mode. Journal of Inner Mongolia Normal University (Natural Science Edition) 48, 5 (2019), 406--412.","journal-title":"Journal of Inner Mongolia Normal University (Natural Science Edition)"},{"key":"e_1_3_1_9_1","volume-title":"International Conference on Chinese Information Processing","author":"Tashi Tsering","year":"1988","unstructured":"Tsering Tashi. 1988. The design of a Tibetan spelling checker. International Conference on Chinese Information Processing, 1988."},{"issue":"2","key":"e_1_3_1_10_1","first-page":"67","article-title":"Tibetan interrogative sentences parsing based on PCFG","volume":"33","year":"2019","unstructured":"Mabao Ban, Zhijie Cai, and Mazhaxi La. 2019. Tibetan interrogative sentences parsing based on PCFG. Journal of Chinese Information Processing 33, 2 (2019), 67--74.","journal-title":"Journal of Chinese Information Processing"},{"issue":"5","key":"e_1_3_1_11_1","first-page":"400","article-title":"Tibetan sentence boundary recognition based on mixed strategy","volume":"48","author":"Que Cuozhuoma","year":"2019","unstructured":"Cuozhuoma Que, Quecairang Hua, Rangdangzhi Cai, and Wuji Xia. 2019. Tibetan sentence boundary recognition based on mixed strategy. Journal of Inner Mongolia Normal University (Natural Science Chinese Edition) 48, 5 (2019), 400--405.","journal-title":"Journal of Inner Mongolia Normal University (Natural Science Chinese Edition)"},{"key":"e_1_3_1_12_1","volume-title":"Research on Rule-Based Analysis of Tibetan Syntax","author":"Wan Mecairang","year":"2014","unstructured":"Mecairang Wan. 2014. Research on Rule-Based Analysis of Tibetan Syntax. Qinghai University for Nationalities. 2014."},{"issue":"1","key":"e_1_3_1_13_1","first-page":"524","article-title":"A reappraisal of sentence and token splitting for life sciences documents","volume":"129","author":"Tomanek Katrin","year":"2007","unstructured":"Katrin Tomanek, Joachim Wermter, and Udo Hahn. 2007. A reappraisal of sentence and token splitting for life sciences documents. Studies in Health Technology Informatics 129, 1 (2007), 524--528.","journal-title":"Studies in Health Technology Informatics"},{"key":"e_1_3_1_14_1","first-page":"985","volume-title":"Proceedings of COLING 2012: Posters","author":"Read Jonathon","year":"2012","unstructured":"Jonathon Read, Rebecca Driden, Stephan Oepen, and Lars J\u00f8rgen Solberg. 2012. Sentence boundary detection: A long solved problem? In Proceedings of COLING 2012: Posters, 985--994."},{"key":"e_1_3_1_15_1","doi-asserted-by":"publisher","DOI":"10.3115\/1075434.1075492"},{"issue":"2","key":"e_1_3_1_16_1","first-page":"242","article-title":"Adaptive multilingual sentence boundary disambiguation","volume":"23","author":"Palmer David D.","year":"1997","unstructured":"David D. Palmer and Marti A. Hearst. 1997. Adaptive multilingual sentence boundary disambiguation. Computational Linguistics 23, 2, 242--267.","journal-title":"Computational Linguistics"},{"key":"e_1_3_1_17_1","doi-asserted-by":"publisher","DOI":"10.3115\/974557.974561"},{"key":"e_1_3_1_18_1","first-page":"241","volume-title":"Proceedings of Human Language Technologies: The 2009 Annual Conference of the North American Chapter of the Association for Computational Linguistics","author":"Gillick Daniel","year":"2009","unstructured":"Daniel Gillick. 2009. Sentence boundary detection and the problem with the US. In Proceedings of Human Language Technologies: The 2009 Annual Conference of the North American Chapter of the Association for Computational Linguistics, 241--244."},{"key":"e_1_3_1_19_1","doi-asserted-by":"publisher","DOI":"10.1162\/089120102760275992"},{"key":"e_1_3_1_20_1","doi-asserted-by":"publisher","DOI":"10.5555\/974305.974340"},{"key":"e_1_3_1_21_1","doi-asserted-by":"publisher","DOI":"10.1162\/coli.2006.32.4.485"},{"issue":"2","key":"e_1_3_1_22_1","first-page":"39","article-title":"Researches of speech classification methods based on Tibetan repertoire","volume":"26","author":"Cai Rangjia","year":"2005","unstructured":"Rangjia Cai and Taijia Ji 2005. Researches of speech classification methods based on Tibetan repertoire. Journal of Northwest University for Nationalities 26, 2 (2005), 39--42.","journal-title":"Journal of Northwest University for Nationalities"},{"key":"e_1_3_1_23_1","first-page":"62","article-title":"Research on automatic recognition method of Tibetan sentence boundary","volume":"8","author":"Ren Qingji","year":"2014","unstructured":"Qingji Ren and Jiancairang An. 2014. Research on automatic recognition method of Tibetan sentence boundary. China Computer and Communication. 8 (2014), 62--63.","journal-title":"China Computer and Communication"},{"issue":"6","key":"e_1_3_1_24_1","first-page":"187","article-title":"Research on the automatic identification of Tibetan sentence boundaries with maximum entropy classifier","volume":"34","author":"Cai Zangtai","year":"2012","unstructured":"Zangtai Cai. 2012. Research on the automatic identification of Tibetan sentence boundaries with maximum entropy classifier. Computer Engineering & Science 34, 6 (2012), 187--190.","journal-title":"Computer Engineering & Science"},{"issue":"4","key":"e_1_3_1_25_1","first-page":"39","article-title":"A maximum entropy and rules approach to identifying Tibetan sentence boundaries","volume":"25","author":"Li Xiang","year":"2011","unstructured":"Xiang Li, Zangtai Cai, Jiang Wenbin, Yajuan Lv, and Qun Liu. 2011. A maximum entropy and rules approach to identifying Tibetan sentence boundaries. Journal of Chinese Information Processing 25, 4 (2011), 39--45.","journal-title":"Journal of Chinese Information Processing"},{"key":"e_1_3_1_26_1","volume-title":"Proceedings of National Symposium on Computational Linguistics for Young People. (YWCL'10)","author":"Zhao Weina","year":"2010","unstructured":"Weina Zhao, Huidan Liu, Xin Yu, Jian Wu, and Pu Zang. 2010. The Tibetan sentence boundary identification based on legal texts. In Proceedings of National Symposium on Computational Linguistics for Young People. (YWCL'10)."},{"issue":"2","key":"e_1_3_1_27_1","first-page":"70","article-title":"Method of identification of Tibetan sentence boundary","volume":"27","author":"Ma Weizhen","year":"2012","unstructured":"Weizhen Ma, Mezhaxi Wan, and Zha Nima. 2012. Method of identification of Tibetan sentence boundary. Journal of Tibet University 27, 2 (2012), 70--76.","journal-title":"Journal of Tibet University"},{"issue":"112","key":"e_1_3_1_28_1","first-page":"39","article-title":"Tibetan sentence extraction method based on feature of function words and sentence ending words","volume":"39","author":"Zha Xiji","year":"2018","unstructured":"Xiji Zha and Ba Luo. 2018. Tibetan sentence extraction method based on feature of function words and sentence ending words. Journal of Northwest Minzu University 39, 112 (2018), 39--44.","journal-title":"Journal of Northwest Minzu University"},{"key":"e_1_3_1_29_1","first-page":"288","volume-title":"Proceedings of the 26th International Conference on Computational Linguistics","author":"Hellwig Oliver","year":"2016","unstructured":"Oliver Hellwig. 2016. Detecting sentence boundaries in Sanskrit texts. In Proceedings of the 26th International Conference on Computational Linguistics, COLING 2016. 288--297."},{"key":"e_1_3_1_30_1","doi-asserted-by":"publisher","DOI":"10.1162\/neco.1997.9.8.1735"},{"key":"e_1_3_1_31_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/D17-1035"},{"key":"e_1_3_1_32_1","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v32i1.11941"},{"key":"e_1_3_1_33_1","volume-title":"The 5th International Conference on Learning Representations (ICLR'17)","author":"Han Linzhou","year":"2017","unstructured":"Linzhou Han, Minwei Feng, Cicero Nogueira dos Santos, Mo Yu, Bing Xiang, Bowen Zhou, and Yoshua Bengio. 2017. A structured self-attentive sentence embedding. The 5th International Conference on Learning Representations (ICLR'17)."},{"key":"e_1_3_1_34_1","article-title":"Efficient estimation of word representations in vector space","author":"Mikolov Tomas","year":"2013","unstructured":"Tomas Mikolov, Kai Chen, Greg Corrado, and Jeffrey Dean. Efficient estimation of word representations in vector space. Computer Science, 2013.","journal-title":"Computer Science"},{"key":"e_1_3_1_35_1","unstructured":"Carlos Emiliano Gonz\u00e1lez-Gallardo Juan Manuel and Torres Moreno. 2018. Sentence boundary detection for French with subword-level information vectors and convolutional neural networks 2018."},{"key":"e_1_3_1_36_1","doi-asserted-by":"publisher","DOI":"10.1007\/s11265-017-1289-8"},{"key":"e_1_3_1_37_1","volume-title":"Proceedings of the 19th Empirical Methods in Natural Language Processing (EMNLP'14)","author":"Yoon Kim","year":"2014","unstructured":"Kim Yoon. 2014. Convolutional neural networks for sentence classification. In Proceedings of the 19th Empirical Methods in Natural Language Processing (EMNLP'14)."},{"key":"e_1_3_1_38_1","volume-title":"Proceedings of the 25th International Joint Conference on Artificial Intelligence 2016 (IJCAI'16)","author":"Liu Pengfei","year":"2016","unstructured":"Pengfei Liu, Xipeng Qiu, and Xuanjing Huang. 2016. Recurrent neural network for text classification with multi-task learning. In Proceedings of the 25th International Joint Conference on Artificial Intelligence 2016 (IJCAI'16)."},{"key":"e_1_3_1_39_1","first-page":"427","volume-title":"Proceedings of the 15th Conference of the European Chapter of the Association for Computational Linguistics","author":"Joulin Armand","year":"2016","unstructured":"Armand Joulin, Edouard Grave, Piotr Bojanowski, and Tomas Mikolov. 2016. Bag of tricks for efficient text classification. In Proceedings of the 15th Conference of the European Chapter of the Association for Computational Linguistics 2 (2016), 427--431."},{"key":"e_1_3_1_40_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/P17-1052"},{"key":"e_1_3_1_41_1","doi-asserted-by":"publisher","DOI":"10.5555\/3295222.3295349"},{"key":"e_1_3_1_42_1","first-page":"79","volume-title":"Proceeding of the 10th Machine Translation Summit (MT summit)","author":"Koehn Philipp","year":"2005","unstructured":"Philipp Koehn. 2005. Europarl: A parallel corpus for statistical machine translation. In Proceeding of the 10th Machine Translation Summit (MT summit), 79--86."},{"key":"e_1_3_1_43_1","volume-title":"Proceedings of the 12th International Workshop on Spoken Language Translation (IWSLT'15)","year":"2015","unstructured":"Mauro Cettolo, Jan Niehues, Sebastian St\u00fcker, Luisa Bentivogli, Roldano Cattoni, and Marcello Federico. 2015. The IWSLT 2015 evaluation campaign. In Proceedings of the 12th International Workshop on Spoken Language Translation (IWSLT'15), Da Nang, Vietnam."}],"container-title":["ACM Transactions on Asian and Low-Resource Language Information Processing"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3527663","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3527663","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T20:18:53Z","timestamp":1750191533000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3527663"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,11,30]]},"references-count":42,"journal-issue":{"issue":"6","published-print":{"date-parts":[[2022,11,30]]}},"alternative-id":["10.1145\/3527663"],"URL":"https:\/\/doi.org\/10.1145\/3527663","relation":{},"ISSN":["2375-4699","2375-4702"],"issn-type":[{"value":"2375-4699","type":"print"},{"value":"2375-4702","type":"electronic"}],"subject":[],"published":{"date-parts":[[2022,11,30]]},"assertion":[{"value":"2020-08-09","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2022-03-18","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2023-02-21","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}