{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T04:13:08Z","timestamp":1750219988441,"version":"3.41.0"},"reference-count":33,"publisher":"Association for Computing Machinery (ACM)","issue":"3","license":[{"start":{"date-parts":[[2023,3,10]],"date-time":"2023-03-10T00:00:00Z","timestamp":1678406400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"DOI":"10.13039\/501100001809","name":"National Nature Science Foundation of China","doi-asserted-by":"crossref","award":["61462055, 61562049"],"award-info":[{"award-number":["61462055, 61562049"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Asian Low-Resour. Lang. Inf. Process."],"published-print":{"date-parts":[[2023,3,31]]},"abstract":"<jats:p>The sentiment lexicon is an important tool for natural language processing tasks. In addition to being able to determine the sentiment polarity of words or phrases, it can assist attribute-level, sentence-level, and text-level sentiment analysis tasks. In light of the fact that tagging data and corpora for the Khmer language are scarce, where most resources related to sentiment lexicons are for English, this paper proposes a method for constructing a sentiment lexicon for Khmer based on<jats:bold>Positive-Unlabeled learning (PU Learning)<\/jats:bold>and the label propagation algorithm. Sentiment words are first extracted from a corpus using the Spy technique of PU learning method. The main idea is to purify the set of N-class examples, train the MLP classifier, and continuously delete spy words and increase the number of P-class words in the iterative process. Following this, the sentiment polarity of the candidate words is determined. By considering the problem of determining the sentiment polarity of the candidate words as one of calculating its probability distribution, a small number of labeled sentiment words and candidate words are used to construct a graph model. The contextual information of the candidate words is used to construct a simple supplementary graph model of the set of sentiment words through word co-occurrence and triangulation, where this enhances the correlation between data items. The sentiment polarity of the candidate words is then determined through the label propagation algorithm. The results of experiments show that the proposed method can be used to construct a Khmer sentiment lexicon with a small number of labeled data and a small corpus without requiring excessive manual labeling.<\/jats:p>","DOI":"10.1145\/3564697","type":"journal-article","created":{"date-parts":[[2022,9,29]],"date-time":"2022-09-29T11:45:27Z","timestamp":1664451927000},"page":"1-18","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":0,"title":["Khmer Sentiment Lexicon Based on PU Learning and Label Propagation Algorithm"],"prefix":"10.1145","volume":"22","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-0480-1839","authenticated-orcid":false,"given":"Chao","family":"Li","sequence":"first","affiliation":[{"name":"Faculty of Information Engineering and Automation, Kunming University of Science and Technology, Yunnan Key Laboratory of Artificial Intelligence, Kunming, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-6952-3967","authenticated-orcid":false,"given":"Xin","family":"Yan","sequence":"additional","affiliation":[{"name":"Faculty of Information Engineering and Automation, Kunming University of Science and Technology, Yunnan Key Laboratory of Artificial Intelligence, Kunming, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-4455-3937","authenticated-orcid":false,"given":"Guangyi","family":"Xu","sequence":"additional","affiliation":[{"name":"Yunnan Nantian Electronic Information Industry Co., Ltd., Yunnan, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-6275-5335","authenticated-orcid":false,"given":"Zhongying","family":"Deng","sequence":"additional","affiliation":[{"name":"Faculty of Information Engineering and Automation, Kunming University of Science and Technology, Yunnan, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-6393-5837","authenticated-orcid":false,"given":"Yuanyuan","family":"Mo","sequence":"additional","affiliation":[{"name":"School of Southeast &amp; South Asia Languages and Culture, Yunnan Minzu University, Yunnan, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2023,3,10]]},"reference":[{"key":"e_1_3_1_2_2","doi-asserted-by":"publisher","DOI":"10.5555\/3019323"},{"key":"e_1_3_1_3_2","doi-asserted-by":"publisher","DOI":"10.1145\/1014052.1014073"},{"key":"e_1_3_1_4_2","doi-asserted-by":"publisher","DOI":"10.3115\/976909.979640"},{"key":"e_1_3_1_5_2","doi-asserted-by":"publisher","DOI":"10.5555\/645531.656022"},{"key":"e_1_3_1_6_2","unstructured":"X. Zhu and Z. Ghahramani. 2002. Learning from Labels and Unlabeled Data with Label Propagation [J]. Tech. Rep. Technical Report CMU-CALD-02.107 2002."},{"key":"e_1_3_1_7_2","volume-title":"Proceedings of LREC","author":"Kamps J.","year":"2004","unstructured":"J. Kamps, M. Marx, R. J. Mokken, and M. de Rijke. 2004. Using WordNet to measure semantic orientation of adjectives. In Proceedings of LREC."},{"key":"e_1_3_1_8_2","doi-asserted-by":"publisher","DOI":"10.5555\/2002736.2002852"},{"key":"e_1_3_1_9_2","first-page":"424","volume-title":"Proceedings of the 45th Annual Meeting of the Association for Computational Linguistics.","author":"Esuli A.","year":"2007","unstructured":"A. Esuli and F. Sebastiani. 2007. PageRanking WordNet Synsets: An application to opinion mining. In Proceedings of the 45th Annual Meeting of the Association for Computational Linguistics. Prague, Association for Computational Linguistics, 424\u2212431."},{"key":"e_1_3_1_10_2","doi-asserted-by":"crossref","DOI":"10.1145\/1871437.1871723","article-title":"Construction of a sentimental word dictionary[C]","author":"Dragut E. C.","year":"2010","unstructured":"E. C. Dragut, C. Yu, P. Sistla, et al. 2010. Construction of a sentimental word dictionary[C]. In Proceedings of the 19th ACM Conference on Information and Knowledge Management (CIKM'10). Toronto, Ontario, Canada, (October 26\u201330, 2010), ACM.","journal-title":"Proceedings of the 19th ACM Conference on Information and Knowledge Management (CIKM'10)"},{"key":"e_1_3_1_11_2","doi-asserted-by":"crossref","unstructured":"P. D. Turney. 2002. Thumbs up or thumbs down? Semantic orientation applied to unsupervised classification of reviews[J]. arXiv preprint cs\/0212032 (2002).","DOI":"10.3115\/1073083.1073153"},{"key":"e_1_3_1_12_2","doi-asserted-by":"publisher","DOI":"10.5555\/1610075.1610125"},{"key":"e_1_3_1_13_2","doi-asserted-by":"publisher","DOI":"10.1145\/2481492.2481506"},{"key":"e_1_3_1_14_2","doi-asserted-by":"publisher","DOI":"10.3115\/v1\/P14-1146"},{"key":"e_1_3_1_15_2","first-page":"595","volume-title":"Proceedings of the Conference on Empirical Methods in Natural Language Processing. Conference on Empirical Methods in Natural Language Processing","author":"Hamilton W. L.","year":"2016","unstructured":"W. L. Hamilton, K. Clark, J. Leskovec, et al. 2016. Inducing domain-specific sentiment lexicons from unlabeled corpora[C]. In Proceedings of the Conference on Empirical Methods in Natural Language Processing. Conference on Empirical Methods in Natural Language Processing. NIH Public Access (2016), 595."},{"key":"e_1_3_1_16_2","volume-title":"Proceedings of the 24th International Joint Conference on Artificial Intelligence (IJCAI'15)","author":"Bravo-Marquez F.","year":"2015","unstructured":"F. Bravo-Marquez, E. Frank, and B. Pfahringer. 2015. Positive, negative, or neutral: Learning an expanded opinion lexicon from emoticon-annotated tweets[C]. In Proceedings of the 24th International Joint Conference on Artificial Intelligence (IJCAI'15). AAAI Press."},{"key":"e_1_3_1_17_2","first-page":"219","article-title":"Don't count, predict! An automatic approach to learning sentiment lexicons for short text[C]","volume":"2","author":"Vo D. T.","year":"2016","unstructured":"D. T. Vo and Y. Zhang. 2016. Don't count, predict! An automatic approach to learning sentiment lexicons for short text[C]. In Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers) 2 (2016), 219\u2212224.","journal-title":"Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers)"},{"key":"e_1_3_1_18_2","doi-asserted-by":"publisher","DOI":"10.1109\/TASLP.2019.2933326"},{"key":"e_1_3_1_19_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.dss.2016.04.007"},{"key":"e_1_3_1_20_2","doi-asserted-by":"publisher","DOI":"10.1007\/s10579-016-9364-5"},{"key":"e_1_3_1_21_2","doi-asserted-by":"crossref","first-page":"145","DOI":"10.1109\/SKG49510.2019.00033","volume-title":"2019 15th International Conference on Semantics, Knowledge and Grids (SKG)","author":"Liu H. T.","year":"2019","unstructured":"H. T. Liu, J. C. Zhu, X. Y. Liu, et al. 2019. Expansion of sentiment lexicon based on label propagation[C]. 2019 15th International Conference on Semantics, Knowledge and Grids (SKG). IEEE, 145\u2013152."},{"key":"e_1_3_1_22_2","first-page":"587","article-title":"Learning to classify texts using positive and unlabeled data","volume":"3","author":"Li X.","year":"2003","unstructured":"X. Li and B. Liu. 2003. Learning to classify texts using positive and unlabeled data. In Proceedings of the Eighteenth International Joint Conference on Artifical Intelligence 3 (2003), 587\u2013592","journal-title":"Proceedings of the Eighteenth International Joint Conference on Artifical Intelligence"},{"key":"e_1_3_1_23_2","first-page":"2445","volume-title":"ICML","author":"Hsieh Cho-Jui","year":"2015","unstructured":"Cho-Jui Hsieh, Nagarajan Natarajan, and Inderjit S. Dhillon. 2015. Pu learning for matrix completion. In ICML. 2445\u20132453."},{"key":"e_1_3_1_24_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-642-19460-3"},{"issue":"3","key":"e_1_3_1_25_2","first-page":"639","article-title":"Suggestion sentence classification method based on PU learning","volume":"39","author":"Pu Zhang","year":"2019","unstructured":"Zhang Pu, Liu Chang, and Li Xiao. 2019. Suggestion sentence classification method based on PU learning. Journal of Computer Applications 39, 3 (2019), 639\u2013643.","journal-title":"Journal of Computer Applications"},{"key":"e_1_3_1_26_2","unstructured":"S. Maekawa K. Takeuch and M. Onizuka. 2018. Non-linear Attributed Graph Clustering by Symmetric NMF with PU Learning[J]. (2018)."},{"key":"e_1_3_1_27_2","first-page":"1024","volume-title":"Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1(Long Papers)","author":"Jiang Chao","year":"2018","unstructured":"Chao Jiang, Hsiang-Fu Yu, Cho-Jui Hsieh, and Kai-Wei Chang. 2018. Learning word embeddings for low-resource languages by PU learning. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1(Long Papers). Association for Computational Linguistics, 1024\u20131034."},{"key":"e_1_3_1_28_2","unstructured":"R. Kiryo G. Niu M. C. D. Plessis et al. 2017. Positive-Unlabeled Learning with Non-Negative Risk Estimator [J]. (2017)."},{"key":"e_1_3_1_29_2","volume-title":"Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing","author":"Wang Y.","year":"2017","unstructured":"Y. Wang, Y. Zhang, and B. Liu. 2017. Sentiment lexicon expansion based on neural PU learning, double dictionary lookup, and polarity association[C]. In Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing."},{"key":"e_1_3_1_30_2","first-page":"420","volume-title":"Proceedings of the 25th Pacific Asia Conference on Language, Information and Computation (PACLIC'25)","author":"Yong Ren","year":"2011","unstructured":"Ren Yong, N. Kaji, N. Yoshinaga, et al. 2011. Sentiment classification in resource-scarce languages by using label propagation[C]. In Proceedings of the 25th Pacific Asia Conference on Language, Information and Computation (PACLIC'25). Singapore, (Dec. 16\u201318, 2011), 420\u2013429."},{"key":"e_1_3_1_31_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.knosys.2013.11.009"},{"key":"e_1_3_1_32_2","article-title":"Exploiting social and local contexts propagation for inducing Chinese microblog-specific sentiment lexicons[J]","author":"Zhao C.","year":"2019","unstructured":"C. Zhao, S. Wang, and D. Li. 2019. Exploiting social and local contexts propagation for inducing Chinese microblog-specific sentiment lexicons[J]. Computer Speech & Language (2019).","journal-title":"Computer Speech & Language"},{"issue":"12","key":"e_1_3_1_33_2","first-page":"1506","article-title":"Emotional polarity recognition of new words based on label propagation algorithm[J]","volume":"9","author":"Xudong Hong","year":"2015","unstructured":"Hong Xudong, Yu Zhengtao, Yan Xin, et al. 2015. Emotional polarity recognition of new words based on label propagation algorithm[J]. Journal of Frontiers of Computer Science & Technology 9, 12 (2015), 1506\u20131512.","journal-title":"Journal of Frontiers of Computer Science & Technology"},{"issue":"4","key":"e_1_3_1_34_2","first-page":"110","article-title":"A Khmer word segmentation and part-of-speech tagging method based on cascaded conditional random fields[J]","volume":"30","author":"Huashan Pan","year":"2016","unstructured":"Pan Huashan, Yan Xin, Zhou Feng, et al. 2016. A Khmer word segmentation and part-of-speech tagging method based on cascaded conditional random fields[J]. Journal of Chinese Information Processing 30, 4 (2016), 110\u2013116.","journal-title":"Journal of Chinese Information Processing"}],"container-title":["ACM Transactions on Asian and Low-Resource Language Information Processing"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3564697","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3564697","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T17:49:30Z","timestamp":1750182570000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3564697"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,3,10]]},"references-count":33,"journal-issue":{"issue":"3","published-print":{"date-parts":[[2023,3,31]]}},"alternative-id":["10.1145\/3564697"],"URL":"https:\/\/doi.org\/10.1145\/3564697","relation":{},"ISSN":["2375-4699","2375-4702"],"issn-type":[{"type":"print","value":"2375-4699"},{"type":"electronic","value":"2375-4702"}],"subject":[],"published":{"date-parts":[[2023,3,10]]},"assertion":[{"value":"2021-01-13","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2022-09-14","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2023-03-10","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}