{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,2]],"date-time":"2026-05-02T04:39:33Z","timestamp":1777696773476,"version":"3.51.4"},"reference-count":28,"publisher":"SAGE Publications","issue":"4","license":[{"start":{"date-parts":[[2024,11,1]],"date-time":"2024-11-01T00:00:00Z","timestamp":1730419200000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/journals.sagepub.com\/page\/policies\/text-and-data-mining-license"}],"content-domain":{"domain":["journals.sagepub.com"],"crossmark-restriction":true},"short-container-title":["Intelligent Decision Technologies"],"published-print":{"date-parts":[[2024,11]]},"abstract":"<jats:p>Art is a symbol of people\u2019s thoughts, and among many forms of artistic expression, literature is the most direct one, which can present art directly to people. How to correctly understand language materials in literature is crucial for understanding literary works and realizing their artistic value. Therefore, in order to strengthen the understanding of Korean literature and analyze its core ideas, this article utilizes modern computer technology and improved Term Frequency-Inverse Document Frequency (TF-IDF) algorithm to process the corpus of Korean literature, in order to quickly extract valuable textual information from Korean literature and facilitate reading and understanding. At the same time, a Korean literature corpus processing model was constructed based on deep learning algorithms. This model is based on the Natural Language Processing (NLP) algorithm, selecting Word Frequency Inverse Document Frequency (TF-IDF) as the feature to calculate the feature weight of keywords. By weighting the naive Bayesian algorithm, it achieves the classification and processing of expected text data in Korean literature. The results of multiple experiments show that the classification accuracy of the model exceeds 97.7%, and the classification recall rate is as high as 94.2%, indicating that the model can effectively achieve corpus processing in Korean literature.<\/jats:p>","DOI":"10.3233\/idt-230772","type":"journal-article","created":{"date-parts":[[2024,7,5]],"date-time":"2024-07-05T11:45:38Z","timestamp":1720179938000},"page":"3011-3024","update-policy":"https:\/\/doi.org\/10.1177\/sage-journals-update-policy","source":"Crossref","is-referenced-by-count":2,"title":["Research on Korean literature corpus processing based on computer system improved TF-IDF algorithm"],"prefix":"10.1177","volume":"18","author":[{"given":"Jing","family":"Xue","sequence":"first","affiliation":[{"name":"North China University of Water Resources and Electric Power, Zhengzhou, China"},{"name":"E-mail:"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"179","published-online":{"date-parts":[[2024,11,1]]},"reference":[{"key":"bibr1-IDT-230772","doi-asserted-by":"crossref","unstructured":"ZhangY. Modern Chinese literature as an institution: Canon and literary history. The Columbia companion to modern Chinese literature. Columbia University Press. 2016; 27-37.","DOI":"10.7312\/dent17008-003"},{"key":"bibr2-IDT-230772","unstructured":"JorgensenJ. The origins and development of Korean literature. In Handbook of Korean Literature, edited by Ch\u2019oe Yun. New York: M.E. Sharpe. 1996. pp. 1-23."},{"key":"bibr3-IDT-230772","unstructured":"KangHBKimN. Development of Korean semantic similarity measures using a web corpus. In 2016 3rd International Conference on Biomedical and Bioinformatics Engineering (ICBBE). 2016. pp. 184-187."},{"key":"bibr4-IDT-230772","first-page":"1180","volume":"41","author":"Kim JK","year":"2014","journal-title":"Expert Systems with Applications."},{"key":"bibr5-IDT-230772","first-page":"209","volume":"48","author":"Kim MK","year":"2014","journal-title":"Journal of the Korean Society for Library and Information Science."},{"key":"bibr6-IDT-230772","unstructured":"AlharthiSA. Empirical study of features and unsupervised sentiment analysis techniques for depression detection in social media. European Journal of Computer Science and Information Technology. 2020(5); 8."},{"key":"bibr7-IDT-230772","unstructured":"KimYSeoY. Learning Korean word vector representations with multiple information sources. In Proceedings of the Twenty-Ninth AAAI Conference on Artificial Intelligence. 2015. pp. 3610-3616."},{"key":"bibr8-IDT-230772","first-page":"337","volume":"49","author":"Kim Y","year":"2013","journal-title":"Language Research."},{"key":"bibr9-IDT-230772","first-page":"119","volume":"45","author":"Lee HJ","year":"2016","journal-title":"Journal of Korean Society of Food Science and Nutrition."},{"key":"bibr10-IDT-230772","first-page":"82","volume":"135","author":"Lee HK","year":"2018","journal-title":"Journal of Pragmatics."},{"key":"bibr11-IDT-230772","unstructured":"LeeKHHanSRYoonSA. Korean text classification using text representation and machine learning. In 2014 International Conference on Information and Communication Technology Convergence (ICTC). 2014. pp. 838-842."},{"key":"bibr12-IDT-230772","unstructured":"LeeSWLeeJ. A comparative study of feature selection methods for Korean text classification. In Proceedings of the 15th International Conference on Ubiquitous Computing and Communications and the 2016 International Symposium on Cyberspace and Security. 2016. pp. 555-560."},{"key":"bibr13-IDT-230772","first-page":"101","volume":"6","author":"Lee WH","year":"2015","journal-title":"International Journal of Computational Linguistics and Applications."},{"key":"bibr14-IDT-230772","first-page":"61","volume":"3","author":"Lim HW","year":"2018","journal-title":"Journal of the Korean Society for Information Management."},{"key":"bibr15-IDT-230772","doi-asserted-by":"publisher","DOI":"10.1007\/s13042-022-01695-4"},{"key":"bibr16-IDT-230772","unstructured":"ParkKCChoY. Korean news headline classification using word embeddings. In Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing. 2015. pp. 90-95."},{"key":"bibr17-IDT-230772","doi-asserted-by":"crossref","unstructured":"LongSRuanJZhangW, et al. Textsnake: A flexible representation for detecting text of arbitrary shapes. In Proceedings of the European Conference on Computer Vision (ECCV). 2018. pp. 20-36.","DOI":"10.1007\/978-3-030-01216-8_2"},{"key":"bibr18-IDT-230772","doi-asserted-by":"publisher","DOI":"10.1145\/3065386"},{"key":"bibr19-IDT-230772","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2016.2646371"},{"key":"bibr20-IDT-230772","doi-asserted-by":"publisher","DOI":"10.1109\/TNNLS.2020.2978386"},{"key":"bibr21-IDT-230772","first-page":"212","volume":"30","author":"Sohrab A","year":"2018","journal-title":"Journal of King Saud University-Computer and Information Sciences."},{"key":"bibr22-IDT-230772","doi-asserted-by":"crossref","unstructured":"AmensisaADPatilSAgrawalP. A survey on text document categorization using enhanced sentence vector space model and bi-gram text representation model based on novel fusion techniques. In 2018 2nd International Conference on Inventive Systems and Control (ICISC). IEEE. 2018. pp. 218-225.","DOI":"10.1109\/ICISC.2018.8399067"},{"key":"bibr23-IDT-230772","doi-asserted-by":"crossref","first-page":"320","DOI":"10.1016\/j.jbi.2015.08.008","volume":"57","author":"Roh JS","year":"2015","journal-title":"Journal of Biomedical Informatics."},{"key":"bibr24-IDT-230772","first-page":"23","volume":"11","author":"Seo JB","year":"2013","journal-title":"Journal of the Korean Society of Information Technology."},{"key":"bibr25-IDT-230772","first-page":"283","volume":"49","author":"Shin SW","year":"2013","journal-title":"Language Research."},{"key":"bibr26-IDT-230772","unstructured":"SonYJKimHR. The development of Korean morphological analyzer for e-learning system. In 2014 International Conference on Advanced Communication Technology (ICACT). 2014. pp. 717-722."},{"key":"bibr27-IDT-230772","doi-asserted-by":"crossref","first-page":"121","DOI":"10.9708\/jksci.2015.20.8.121","volume":"20","author":"Sung Y","year":"2015","journal-title":"Journal of the Korean Society of Computer and Information."},{"key":"bibr28-IDT-230772","first-page":"359","volume":"49","author":"Yim SH","year":"2013","journal-title":"Language Research."}],"container-title":["Intelligent Decision Technologies"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/journals.sagepub.com\/doi\/pdf\/10.3233\/IDT-230772","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/journals.sagepub.com\/doi\/full-xml\/10.3233\/IDT-230772","content-type":"application\/xml","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/journals.sagepub.com\/doi\/pdf\/10.3233\/IDT-230772","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,4,29]],"date-time":"2026-04-29T09:20:55Z","timestamp":1777454455000},"score":1,"resource":{"primary":{"URL":"https:\/\/journals.sagepub.com\/doi\/10.3233\/IDT-230772"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,11]]},"references-count":28,"journal-issue":{"issue":"4","published-print":{"date-parts":[[2024,11]]}},"alternative-id":["10.3233\/IDT-230772"],"URL":"https:\/\/doi.org\/10.3233\/idt-230772","relation":{},"ISSN":["1872-4981","1875-8843"],"issn-type":[{"value":"1872-4981","type":"print"},{"value":"1875-8843","type":"electronic"}],"subject":[],"published":{"date-parts":[[2024,11]]}}}