{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,8,5]],"date-time":"2026-08-05T10:43:33Z","timestamp":1785926613174,"version":"3.56.0"},"reference-count":33,"publisher":"SAGE Publications","issue":"3","license":[{"start":{"date-parts":[[2025,4,10]],"date-time":"2025-04-10T00:00:00Z","timestamp":1744243200000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/journals.sagepub.com\/page\/policies\/text-and-data-mining-license"}],"content-domain":{"domain":["journals.sagepub.com"],"crossmark-restriction":true},"short-container-title":["International Journal of Knowledge-Based and Intelligent Engineering Systems"],"published-print":{"date-parts":[[2025,8]]},"abstract":"<jats:title>Abstract<\/jats:title>\n                  <jats:p>In the fast-moving field of Natural Language Processing (NLP), making lexicons is still a method for many text analysis applications. This process of generating lexicons has traditionally used techniques such as semantic matches, word embeddings, and tools like EMPATH. With the arrival of Large Language Models (LLMs) including GPT-3.5, GPT-4 and Mistral 7b 0.1, we have new ways to create lexicons. This study takes a close look at how these older methods stack up against the newer options brought by LLMs. We carried out a detailed analysis, looking at how well different methods could create lexicons, focusing on their precision, scalability, and concluding on how efficiently they can be used in real-world settings. By using standard NLP tasks like document classification, emotion classification and sentiment analysis, this research prove itself on a variety of datasets to test how well the lexicons worked. This discovery, along with others from our study, aims to help professionals and researchers find the best approaches to lexicon creation today, setting the stage for more research in the NLP field.<\/jats:p>","DOI":"10.1177\/13272314251322523","type":"journal-article","created":{"date-parts":[[2025,4,10]],"date-time":"2025-04-10T09:55:51Z","timestamp":1744278951000},"page":"363-374","update-policy":"https:\/\/doi.org\/10.1177\/sage-journals-update-policy","source":"Crossref","is-referenced-by-count":1,"title":["Generating lexicons; are large language models better?"],"prefix":"10.1177","volume":"29","author":[{"ORCID":"https:\/\/orcid.org\/0009-0004-8064-1859","authenticated-orcid":false,"given":"Julien Pierre Edmond","family":"Ghali","sequence":"first","affiliation":[{"name":"Department\u00a0of Computer Science,\u00a0Nagoya Institute of Technology,\u00a0Nagoya,\u00a0Japan"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-0723-4999","authenticated-orcid":false,"given":"Kosuke","family":"Shima","sequence":"additional","affiliation":[{"name":"Department\u00a0of Computer Science,\u00a0Nagoya Institute of Technology,\u00a0Nagoya,\u00a0Japan"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0005-8770-9155","authenticated-orcid":false,"given":"Atsuko","family":"Mutoh","sequence":"additional","affiliation":[{"name":"Department\u00a0of Computer Science,\u00a0Nagoya Institute of Technology,\u00a0Nagoya,\u00a0Japan"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-1735-1862","authenticated-orcid":false,"given":"Koichi","family":"Moriyama","sequence":"additional","affiliation":[{"name":"Department\u00a0of Computer Science,\u00a0Nagoya Institute of Technology,\u00a0Nagoya,\u00a0Japan"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0006-4574-9527","authenticated-orcid":false,"given":"Nobuhiro","family":"Inuzuka","sequence":"additional","affiliation":[{"name":"Department\u00a0of Computer Science,\u00a0Nagoya Institute of Technology,\u00a0Nagoya,\u00a0Japan"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"179","published-online":{"date-parts":[[2025,4,10]]},"reference":[{"key":"e_1_3_3_2_2","unstructured":"Mohammad S Turney P. Emotions evoked by common words and phrases: Using mechanical turk to create an emotion lexicon. In: Proceedings of the NAACL HLT 2010 workshop on computational approaches to analysis and generation of emotion in text 2010 June pp.26\u201334."},{"key":"e_1_3_3_3_2","doi-asserted-by":"crossref","unstructured":"Giunchiglia F Shvaiko P Yatskevich M. S-Match: an algorithm and an implementation of semantic matching. In: European semantic web symposium 2004 May pp.61\u201375. Springer Berlin Heidelberg.","DOI":"10.1007\/978-3-540-25956-5_5"},{"key":"e_1_3_3_4_2","doi-asserted-by":"crossref","unstructured":"Fast E Chen B Bernstein MS. Empath: Understanding topic signals in large-scale text. In: Proceedings of the 2016 CHI conference on human factors in computing systems 2016 May pp.4647\u20134657.","DOI":"10.1145\/2858036.2858535"},{"key":"e_1_3_3_5_2","unstructured":"Le Q Mikolov T. Distributed representations of sentences and documents. In International conference on machine learning 2014 June pp.1188\u20131196. PMLR."},{"key":"e_1_3_3_6_2","doi-asserted-by":"crossref","unstructured":"McInnes L Healy J Melville J. Umap: Uniform manifold approximation and projection for dimension reduction. arXiv preprint arXiv:1802.03426. 2018.","DOI":"10.32614\/CRAN.package.uwot"},{"key":"e_1_3_3_7_2","unstructured":"Devlin J Chang MW Lee K et al. Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805. 2018."},{"key":"e_1_3_3_8_2","first-page":"9","article-title":"Language models are unsupervised multitask learners","volume":"1","author":"Radford A","year":"2019","unstructured":"Radford A, Wu J, Child R, et al. Language models are unsupervised multitask learners. OpenAI Blog 2019; 1: 9.","journal-title":"OpenAI Blog"},{"key":"e_1_3_3_9_2","unstructured":"Jiang AQ Sablayrolles A Mensch A et al. Mistral 7B. arXiv preprint arXiv:2310.06825. 2023."},{"key":"e_1_3_3_10_2","article-title":"Attention is all you need","volume":"30","author":"Vaswani A","year":"2017","unstructured":"Vaswani A, Shazeer N, Parmar N, et al. Attention is all you need. Advances in neural information processing systems 2017; 30: 5998\u20136008.","journal-title":"Advances in neural information processing systems"},{"key":"e_1_3_3_11_2","unstructured":"Newsgroup Dataset. http:\/\/www.cad.zju.edu.cn\/home\/dengcai\/Data\/TextData.html."},{"key":"e_1_3_3_12_2","unstructured":"Wang B Liakata M Zubiaga A et al. Smile: Twitter emotion classification using domain adaptation. In: CEUR Workshop proceedings 2016 Vol. 1619 pp.15\u201321. Sun SITE Central Europe."},{"key":"e_1_3_3_13_2","doi-asserted-by":"crossref","unstructured":"Saravia E Liu HCT Huang YH et al. CARER: Contextualized affect representations for emotion recognition. In: Proceedings of the 2018 conference on empirical methods in natural language processing 2018 pp.3687\u20133697.","DOI":"10.18653\/v1\/D18-1404"},{"key":"e_1_3_3_14_2","first-page":"16","article-title":"Basic emotions","volume":"98","author":"Ekman P","year":"1999","unstructured":"Ekman P. Basic emotions. Handbook of Cognition and Emotion 1999; 98: 16.","journal-title":"Handbook of Cognition and Emotion"},{"key":"e_1_3_3_15_2","unstructured":"Yelp Dataset. URL: https:\/\/www.yelp.com\/dataset\/."},{"key":"e_1_3_3_16_2","unstructured":"Mikolov T Chen K Corrado G et al. Efficient estimation of word representations in vector space. arXiv preprint arXiv:1301.3781. 2013."},{"key":"e_1_3_3_17_2","doi-asserted-by":"crossref","unstructured":"Rumelhart DE Hinton GE Williams RJ. Learning internal representations by error propagation. 1985.","DOI":"10.21236\/ADA164453"},{"key":"e_1_3_3_18_2","doi-asserted-by":"publisher","DOI":"10.1162\/neco.1997.9.8.1735"},{"key":"e_1_3_3_19_2","unstructured":"Touvron H Martin L Stone K et al. Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288. 2023."},{"key":"e_1_3_3_20_2","doi-asserted-by":"crossref","unstructured":"Penedo G Malartic Q Hesslow D et al. The RefinedWeb dataset for Falcon LLM: outperforming curated corpora with web data and web data only. arXiv preprint arXiv:2306.01116. 2023.","DOI":"10.52202\/075280-3464"},{"key":"e_1_3_3_21_2","doi-asserted-by":"crossref","unstructured":"Qadir A Riloff E. Learning emotion indicators from tweets: Hashtags hashtag patterns and phrases. In: Proceedings of the 2014 conference on empirical methods in natural language processing (EMNLP) 2014 October pp.1203\u20131209.","DOI":"10.3115\/v1\/D14-1127"},{"key":"e_1_3_3_22_2","doi-asserted-by":"crossref","unstructured":"Muppidi S Gorripati SK Kishore B. An approach for bibliographic citation sentiment analysis using deep learning\u2019. 1 Jan. 2020: 353\u2013362.","DOI":"10.3233\/KES-200087"},{"key":"e_1_3_3_23_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.jksues.2016.04.002"},{"key":"e_1_3_3_24_2","doi-asserted-by":"crossref","unstructured":"Zhang H Gan W Jiang B. Machine learning and lexicon-based methods for sentiment classification: A survey. In: 2014 11th web information system and application conference 2014 September pp.262\u201326. IEEE.","DOI":"10.1109\/WISA.2014.55"},{"key":"e_1_3_3_25_2","doi-asserted-by":"crossref","unstructured":"Pablos AG Cuadros M Rigau G. A Comparison of Domain-based Word Polarity Estimation using different Word Embeddings. In: Proceedings of the Tenth International Conference on Language Resources and Evaluation (LREC'16) 2016 pp.54\u201360 Portoro\u017e Slovenia. European Language Resources Association (ELRA).","DOI":"10.63317\/24rnrjsuraet"},{"key":"e_1_3_3_26_2","unstructured":"Musto C Semeraro G Polignano M. A Comparison of Lexicon-based Approaches for Sentiment Analysis of Microblog Posts. In: DART@ AI* IA 2014 December pp.59\u201368)."},{"key":"e_1_3_3_27_2","doi-asserted-by":"crossref","unstructured":"Tabak FS Evrim V. Comparison of emotion lexicons. In 2016 HONET-ICT 2016 October pp.154\u2013158. IEEE.","DOI":"10.1109\/HONET.2016.7753440"},{"key":"e_1_3_3_28_2","doi-asserted-by":"publisher","DOI":"10.1037\/0033-295X.114.2.211"},{"key":"e_1_3_3_29_2","unstructured":"Angelov D. Top2vec: Distributed representations of topics. arXiv preprint arXiv:2008.09470. 2020."},{"key":"e_1_3_3_30_2","unstructured":"Naderalvojoud B Bozkir AS Sezer EA. Investigation of term weighting schemes in classification of imbalanced texts. In: Lisbon: Proceedings of European Conference on Data Mining (ECDM) 2014 July pp.15\u20137."},{"key":"e_1_3_3_31_2","unstructured":"Bauer S Clark DD Lehr W. Understanding broadband speed measurements. Tprc. 2010 August."},{"key":"e_1_3_3_32_2","doi-asserted-by":"crossref","unstructured":"Loper E Bird S. Nltk: The natural language toolkit. arXiv preprint cs\/0205028. 2002.","DOI":"10.3115\/1118108.1118117"},{"key":"e_1_3_3_33_2","unstructured":"The Dataset. https:\/\/www.yelp.com\/dataset."},{"key":"e_1_3_3_34_2","first-page":"29","article-title":"Identification of mental disorders through text mining on social Media","volume":"17","author":"Julien G","year":"2024","unstructured":"Julien G, Kosuke S, Koichi M, et al. Identification of mental disorders through text mining on social Media. (TOM) 2024; 17: 29\u201335.","journal-title":"(TOM)"}],"container-title":["International Journal of Knowledge-Based and Intelligent Engineering Systems"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/journals.sagepub.com\/doi\/pdf\/10.1177\/13272314251322523","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/journals.sagepub.com\/doi\/full-xml\/10.1177\/13272314251322523","content-type":"application\/xml","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/journals.sagepub.com\/doi\/pdf\/10.1177\/13272314251322523","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,8,5]],"date-time":"2026-08-05T09:44:04Z","timestamp":1785923044000},"score":1,"resource":{"primary":{"URL":"https:\/\/journals.sagepub.com\/doi\/10.1177\/13272314251322523"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,4,10]]},"references-count":33,"journal-issue":{"issue":"3","published-print":{"date-parts":[[2025,8]]}},"alternative-id":["10.1177\/13272314251322523"],"URL":"https:\/\/doi.org\/10.1177\/13272314251322523","relation":{},"ISSN":["1327-2314","1875-8827"],"issn-type":[{"value":"1327-2314","type":"print"},{"value":"1875-8827","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,4,10]]}}}