{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,2]],"date-time":"2026-06-02T11:37:19Z","timestamp":1780400239230,"version":"3.54.1"},"reference-count":58,"publisher":"SAGE Publications","issue":"2","license":[{"start":{"date-parts":[[2024,3,20]],"date-time":"2024-03-20T00:00:00Z","timestamp":1710892800000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/journals.sagepub.com\/page\/policies\/text-and-data-mining-license"}],"funder":[{"name":"Hunan Key Laboratory for Internet of Things in Electricity","award":["2019TP1016"],"award-info":[{"award-number":["2019TP1016"]}]},{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["72061147004"],"award-info":[{"award-number":["72061147004"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]},{"name":"National Natural Science Foundation of Hunan Province","award":["2021JJ30055"],"award-info":[{"award-number":["2021JJ30055"]}]},{"name":"project about research on key technologies of power knowledge graph","award":["5216A6200037"],"award-info":[{"award-number":["5216A6200037"]}]}],"content-domain":{"domain":["journals.sagepub.com"],"crossmark-restriction":true},"short-container-title":["Journal of Information Science"],"published-print":{"date-parts":[[2025,4]]},"abstract":"<jats:p>Modelling short text is challenging due to the small number of word co-occurrence and insufficient semantic information that affects downstream Natural Language Processing (NLP) tasks, for example, text classification. Gathering information from external sources is expensive and may increase noise. For efficient short text classification without depending on external knowledge sources, we propose Expressive Short text Classification (EStC). EStC consists of a novel document context-aware semantically enriched topic model called the Short text Topic Model (StTM) that captures words, topics and documents semantics in a joint learning framework. In StTM, the probability of predicting a context word involves the topic distribution of word embeddings and the document vector as the global context, which obtains by weighted averaging of word embeddings on the fly simultaneously with the topic distribution of words without requiring an additional inference method for the document embedding. EStC represents documents in an expressive (number of topics\u2009\u00d7\u2009number of word embedding features) embedding space and uses a linear support vector machine (SVM) classifier for their classification. Experimental results demonstrate that EStC outperforms many state-of-the-art language models in short text classification using several publicly available short text data sets.<\/jats:p>","DOI":"10.1177\/01655515241230793","type":"journal-article","created":{"date-parts":[[2024,3,21]],"date-time":"2024-03-21T00:20:51Z","timestamp":1710980451000},"page":"481-498","update-policy":"https:\/\/doi.org\/10.1177\/sage-journals-update-policy","source":"Crossref","is-referenced-by-count":4,"title":["Short text classification using semantically enriched topic model"],"prefix":"10.1177","volume":"51","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-1199-2709","authenticated-orcid":false,"given":"Farid","family":"Uddin","sequence":"first","affiliation":[{"name":"School of Computer Science and Engineering, Central South University, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Yibo","family":"Chen","sequence":"additional","affiliation":[{"name":"Information and Communication Branch, State Grid Hunan Electric Power Company Limited, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-2528-7808","authenticated-orcid":false,"given":"Zuping","family":"Zhang","sequence":"additional","affiliation":[{"name":"School of Computer Science and Engineering, Central South University, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Xin","family":"Huang","sequence":"additional","affiliation":[{"name":"Information and Communication Branch, State Grid Hunan Electric Power Company Limited, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"179","published-online":{"date-parts":[[2024,3,20]]},"reference":[{"key":"e_1_3_3_2_2","first-page":"1","article-title":"Multiple weak supervision for short text classification","volume":"52","author":"Chen LM","year":"2022","unstructured":"Chen LM, Xiu BX, Ding ZY. Multiple weak supervision for short text classification. Appl Intell 2022; 52: 1\u201316.","journal-title":"Appl Intell"},{"key":"e_1_3_3_3_2","doi-asserted-by":"publisher","DOI":"10.1177\/0165551516653082"},{"issue":"1","key":"e_1_3_3_4_2","first-page":"6252","article-title":"Deep short text classification with knowledge powered attention","volume":"33","author":"Chen J","year":"2019","unstructured":"Chen J, Hu Y, Liu J, et al. Deep short text classification with knowledge powered attention. Proc AAAI Conf Artif Intell 2019; 33(1): 6252\u20136259.","journal-title":"Proc AAAI Conf Artif Intell"},{"key":"e_1_3_3_5_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.knosys.2020.105572"},{"key":"e_1_3_3_6_2","doi-asserted-by":"publisher","DOI":"10.1007\/s00521-018-3442-0"},{"key":"e_1_3_3_7_2","first-page":"3597","volume-title":"Proceedings of the 58th annual meeting of the Association for Computational Linguistics","author":"Wu F","unstructured":"Wu F, Qiao Y, Chen JH, et al. Mind: a large-scale dataset for news recommendation. In: Proceedings of the 58th annual meeting of the Association for Computational Linguistics, Online, 5\u201310 July 2020, pp. 3597\u20133606. Stroudsburg, PA: Association for Computational Linguistics."},{"key":"e_1_3_3_8_2","doi-asserted-by":"publisher","DOI":"10.1145\/3388970"},{"key":"e_1_3_3_9_2","doi-asserted-by":"publisher","DOI":"10.1145\/3450352"},{"key":"e_1_3_3_10_2","unstructured":"Zhu Y Zhou X Qiang J et al. Prompt-learning for short text classification https:\/\/arxiv.org\/abs\/2202.11345"},{"key":"e_1_3_3_11_2","first-page":"261","volume-title":"Proceedings of the third ACM international conference on Web search and data mining","author":"Weng J","unstructured":"Weng J, Lim E, Jiang J, et al. Twitterrank: finding topic-sensitive influential twitterers. In: Proceedings of the third ACM international conference on Web search and data mining, New York, 4\u20136 February 2010, pp. 261\u2013270. New York: ACM."},{"key":"e_1_3_3_12_2","doi-asserted-by":"publisher","DOI":"10.1145\/2484028.2484166"},{"key":"e_1_3_3_13_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.neucom.2019.08.080"},{"key":"e_1_3_3_14_2","first-page":"2105","volume-title":"Proceedings of the 22nd ACM SIGKDD international conference on knowledge discovery and data mining","author":"Zuo Y","unstructured":"Zuo Y, Wu J, Zhang H, et al. Topic modeling of short texts: a pseudo-document view. In: Proceedings of the 22nd ACM SIGKDD international conference on knowledge discovery and data mining, San Francisco, CA, 13\u201317 August 2016, pp. 2105\u20132114. New York: ACM."},{"key":"e_1_3_3_15_2","first-page":"91","volume-title":"Proceedings of the 17th international conference on World Wide Web","author":"Phan X","unstructured":"Phan X, Nguyen L, Horiguchi S. Learning to classify short and sparse text & web with hidden topics from large-scale data collections. In: Proceedings of the 17th international conference on World Wide Web, Beijing, China, 21\u201325 April 2008, pp. 91\u2013100. New York: ACM."},{"key":"e_1_3_3_16_2","volume-title":"International conference on learning representations","author":"Zhao H","unstructured":"Zhao H, Phung D, Huynh V, et al. Neural topic model via optimal transport. In: International conference on learning representations, Vienna, Austria, 3\u20137 May 2021."},{"key":"e_1_3_3_17_2","doi-asserted-by":"publisher","DOI":"10.1177\/0165551512448985"},{"key":"e_1_3_3_18_2","first-page":"363","volume-title":"Pacific-Asia conference on knowledge discovery and data mining","author":"Qiang J","unstructured":"Qiang J, Chen P, Wang T, et al. Topic modeling over short texts by incorporating word embeddings. In: Pacific-Asia conference on knowledge discovery and data mining, Jeju, South Korea, 23\u201326 May 2017, pp. 363\u2013374. Cham: Springer."},{"key":"e_1_3_3_19_2","first-page":"1145","volume-title":"Proceedings of the 35th international ACM SIGIR conference on research and development in information retrieval","author":"Sun A","unstructured":"Sun A. Short text classification using very few words. In: Proceedings of the 35th international ACM SIGIR conference on research and development in information retrieval, Portland, OR, 12\u201316 August 2012, pp. 1145\u20131146. New York: ACM."},{"key":"e_1_3_3_20_2","first-page":"3111","volume-title":"Advances in neural information processing systems","author":"Mikolov T","year":"2013","unstructured":"Mikolov T, Sutskever I, Chen K, et al. Distributed representations of words and phrases and their compositionality. In: Burges C, Bottou L, Welling M, et al. (eds) Advances in neural information processing systems. Red Hook, NY: Curran Associates, 2013, pp. 3111\u20133119."},{"key":"e_1_3_3_21_2","first-page":"3111","article-title":"Latent Dirichlet allocation","volume":"3","author":"Blei DM","year":"2003","unstructured":"Blei DM, Ng AY, Jordan MI. Latent Dirichlet allocation. J Mach Learn Res 2003; 3: 3111\u20133119.","journal-title":"J Mach Learn Res"},{"key":"e_1_3_3_22_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.knosys.2016.03.027"},{"key":"e_1_3_3_23_2","unstructured":"Le Q Mikolov T. Distributed representations of sentences and documents. In: Jebara T Xing EP (eds) International conference on machine learning Beijing China. 2014 pp. 1188\u20131196."},{"key":"e_1_3_3_24_2","doi-asserted-by":"publisher","DOI":"10.1177\/0165551520977453"},{"key":"e_1_3_3_25_2","first-page":"375","volume-title":"Proceedings of the 40th international ACM SIGIR conference on research and development in information retrieval","author":"Bei S","unstructured":"Bei S, Wai L, Shoaib J, et al. Jointly learning word embeddings and latent topics. In: Proceedings of the 40th international ACM SIGIR conference on research and development in information retrieval, Tokyo, Japan, 7\u201311 August 2017, pp. 375\u2013384. New York: ACM."},{"key":"e_1_3_3_26_2","volume-title":"International conference on learning representations (ICLR)","author":"Chen M","unstructured":"Chen M. Efficient vector representation for documents through corruption. In: International conference on learning representations (ICLR), Toulon, France, 24\u201326 April 2017."},{"key":"e_1_3_3_27_2","first-page":"873","volume-title":"Proceedings of the 50th annual meeting of the Association for Computational Linguistics","author":"Huang EH","year":"2012","unstructured":"Huang EH, Socher R, Manning CD, et al. Improving word representations via global context and multiple word prototypes. In: Proceedings of the 50th annual meeting of the Association for Computational Linguistics, Jeju Island, South Korea, 8\u201314 July 2012, pp. 873\u2013882. Stroudsburg, PA: Association for Computational Linguistics."},{"key":"e_1_3_3_28_2","first-page":"2418","volume-title":"Twenty-ninth AAAI conference on artificial intelligence","author":"Liu Y","unstructured":"Liu Y, Liu Z, Chua T, et al. Topical word embeddings. In: Twenty-ninth AAAI conference on artificial intelligence, Austin, TX, 25\u201330 January 2015, pp. 2418\u20132424. Menlo Park, CA: AAAI Publication."},{"key":"e_1_3_3_29_2","volume-title":"International conference of learning representations (ICLR)","author":"Arora S","unstructured":"Arora S, Liang Y, Ma T. A simple but tough-to-beat baseline for sentence embeddings. In: International conference of learning representations (ICLR), Toulon, France, 24\u201326 April 2017."},{"key":"e_1_3_3_30_2","unstructured":"Moody CE. Mixing Dirichlet topic models and word embeddings to make lda2vec https:\/\/arxiv.org\/abs\/1605.02019"},{"key":"e_1_3_3_31_2","volume-title":"Thirty-second AAAI conference on artificial intelligence","author":"Chaplot DS","unstructured":"Chaplot DS, Salakhutdinov R. Knowledge-based word sense disambiguation using topic model. In: Thirty-second AAAI conference on artificial intelligence, New Orleans, Lousiana, USA, 2\u20137 February 2018."},{"key":"e_1_3_3_32_2","first-page":"4207","volume-title":"Proceedings of the twenty-sixth international joint conference on artificial intelligence","author":"Guangxu X","unstructured":"Guangxu X, Yaliang L, Wayne XZ, et al. A correlated topic model using word embeddings. In: Proceedings of the twenty-sixth international joint conference on artificial intelligence, Melbourne, VIC, Australia, 19\u201325 August 2017, pp. 4207\u20134213. New York: ACM."},{"key":"e_1_3_3_33_2","first-page":"38","volume-title":"Proceedings of the conference on Empirical Methods in natural Language Processing: system demonstrations","author":"Wolf T","unstructured":"Wolf T, Debut L, Sanh V, et al. Transformers: state-of-the-art natural language processing. In: Proceedings of the conference on Empirical Methods in natural Language Processing: system demonstrations, 16\u201320 November 2020, pp. 38\u201345. Online: EMNLP."},{"key":"e_1_3_3_34_2","first-page":"38","article-title":"Effectiveness of fine-tuned BERT model in classification of helpful and unhelpful online customer reviews","volume":"23","author":"Bilal M","year":"2022","unstructured":"Bilal M, Almazroi AA. Effectiveness of fine-tuned BERT model in classification of helpful and unhelpful online customer reviews. Electron Commer Res 2022; 23: 38\u201345.","journal-title":"Electron Commer Res"},{"key":"e_1_3_3_35_2","doi-asserted-by":"publisher","DOI":"10.1109\/TASLP.2020.3008390"},{"key":"e_1_3_3_36_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2020.acl-main.630"},{"key":"e_1_3_3_37_2","doi-asserted-by":"publisher","DOI":"10.1162\/tacl_a_00326"},{"key":"e_1_3_3_38_2","doi-asserted-by":"publisher","DOI":"10.1109\/TKDE.2019.2892430"},{"key":"e_1_3_3_39_2","doi-asserted-by":"publisher","DOI":"10.1109\/ACCESS.2018.2863260"},{"key":"e_1_3_3_40_2","first-page":"146","volume-title":"IEEE 4th international conference on knowledge innovation and invention (ICKII)","author":"Murakami R","unstructured":"Murakami R, Chakraborty B. Neural topic models for short text using pretrained word embeddings and its application to real data. In: IEEE 4th international conference on knowledge innovation and invention (ICKII), Taichung, 23\u201325 July 2021, pp. 146\u2013150. New York: IEEE."},{"key":"e_1_3_3_41_2","doi-asserted-by":"publisher","DOI":"10.1109\/ACCESS.2020.2973207"},{"key":"e_1_3_3_42_2","doi-asserted-by":"publisher","DOI":"10.1109\/ACCESS.2021.3125768"},{"key":"e_1_3_3_43_2","doi-asserted-by":"publisher","DOI":"10.1007\/s00607-019-00755-y"},{"key":"e_1_3_3_44_2","doi-asserted-by":"publisher","DOI":"10.1177\/0165551515585264"},{"issue":"5","key":"e_1_3_3_45_2","first-page":"7863","article-title":"P-SIF: document embeddings using partition averaging","volume":"34","author":"Gupta V","year":"2020","unstructured":"Gupta V, Saw A, Nokhiz P, et al. P-SIF: document embeddings using partition averaging. Proc AAAI Conf Artif Intell 2020, 34(5):7863\u20137870.","journal-title":"Proc AAAI Conf Artif Intell"},{"key":"e_1_3_3_46_2","unstructured":"Gupta V Karnick H Bansal A et al. Product classification in e-commerce using distributional semantics https:\/\/arxiv.org\/abs\/1606.06083"},{"key":"e_1_3_3_47_2","unstructured":"Gupta V Saw A Nokhiz P et al. Improving document classification with multi-sense embeddings https:\/\/arxiv.org\/abs\/arXiv:1911.07918"},{"key":"e_1_3_3_48_2","unstructured":"Shaoul C. The Westbury lab Wikipedia corpus. Edmonton AB: University of Alberta 2010."},{"key":"e_1_3_3_49_2","first-page":"1445","volume-title":"Proceedings of the 22nd international conference on World Wide Web","author":"Yan X","unstructured":"Yan X, Guo J, Lan Y, et al. A biterm topic model for short texts. In: Proceedings of the 22nd international conference on World Wide Web, Rio de Janeiro, Brazil, 13\u201317 May 2013, pp. 1445\u20131456. New York: ACM."},{"key":"e_1_3_3_50_2","doi-asserted-by":"publisher","DOI":"10.1162\/tacl_a_00140"},{"key":"e_1_3_3_51_2","first-page":"165","volume-title":"Proceedings of the 39th international ACM SIGIR conference on research and development in information retrieval","author":"Li C","unstructured":"Li C, Wang H, Zhang Z, et al. Topic modeling for short texts with auxiliary word embeddings. In: Proceedings of the 39th international ACM SIGIR conference on research and development in information retrieval, Pisa, Italy, 17\u201321 July 2016, pp. 165\u2013174. New York: ACM."},{"key":"e_1_3_3_52_2","doi-asserted-by":"publisher","DOI":"10.1007\/s10115-015-0882-z"},{"key":"e_1_3_3_53_2","first-page":"90","volume-title":"Proceedings of the 50th annual meeting of the Association for Computational Linguistics","volume":"2","author":"Wang S","unstructured":"Wang S, Christopher D. Baselines and bigrams: simple, good sentiment and topic classification. In: Proceedings of the 50th annual meeting of the Association for Computational Linguistics (Vol. 2: Short Papers), Jeju Island, South Korea, 8\u201314 July 2012, pp. 90\u201394. Stroudsburg, PA: Association for Computational Linguistics."},{"key":"e_1_3_3_54_2","volume-title":"Proceedings of the Empirical Methods in Natural Language Processing","author":"Chen Y","year":"2014","unstructured":"Chen Y. Convolutional neural network for sentence classification. In: Proceedings of the Empirical Methods in Natural Language Processing, Berlin, 26 August 2014."},{"key":"e_1_3_3_55_2","volume-title":"IJCAI\u201917","author":"Wang J","unstructured":"Wang J, Wang Z, Zhang D, et al. Combining knowledge with deep convolutional neural networks for short text classification. In: IJCAI\u201917, Melbourne, VIC, Australia, 19\u201325 August, 2017. New York: ACM."},{"key":"e_1_3_3_56_2","first-page":"649","article-title":"Character-level convolutional networks for text classification","volume":"28","author":"Chen M","year":"2015","unstructured":"Chen M. Character-level convolutional networks for text classification. Adv Neural Inf Process Syst 2015; 28: 649\u2013657.","journal-title":"Adv Neural Inf Process Syst"},{"issue":"11","key":"e_1_3_3_57_2","first-page":"2579","article-title":"Visualizing data using t-SNE","volume":"9","author":"Maaten LV","year":"2008","unstructured":"Maaten LV, Hinton G. Visualizing data using t-SNE. J Mach Learn Res 2008; 9(11): 2579\u20132605.","journal-title":"J Mach Learn Res"},{"key":"e_1_3_3_58_2","first-page":"399","volume-title":"Proceedings of the eighth ACM international conference on web search and data mining","author":"R\u00f6der M","unstructured":"R\u00f6der M, Both A, Hinneburg A. Exploring the space of topic coherence measures. In: Proceedings of the eighth ACM international conference on web search and data mining, Shanghai, China, 2\u20136 February 2015, pp. 399\u2013408. New York: ACM."},{"key":"e_1_3_3_59_2","first-page":"31","volume-title":"Proceedings of the biennial GSCL conference","author":"Bouma G","unstructured":"Bouma G. Normalized (pointwise) mutual information in collection extraction. In: Proceedings of the biennial GSCL conference, Potsdam, Germany, 30 September 2009, pp. 31\u201340."}],"container-title":["Journal of Information Science"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/journals.sagepub.com\/doi\/pdf\/10.1177\/01655515241230793","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/journals.sagepub.com\/doi\/full-xml\/10.1177\/01655515241230793","content-type":"application\/xml","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/journals.sagepub.com\/doi\/pdf\/10.1177\/01655515241230793","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,4,29]],"date-time":"2026-04-29T23:10:21Z","timestamp":1777504221000},"score":1,"resource":{"primary":{"URL":"https:\/\/journals.sagepub.com\/doi\/10.1177\/01655515241230793"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,3,20]]},"references-count":58,"journal-issue":{"issue":"2","published-print":{"date-parts":[[2025,4]]}},"alternative-id":["10.1177\/01655515241230793"],"URL":"https:\/\/doi.org\/10.1177\/01655515241230793","relation":{},"ISSN":["0165-5515","1741-6485"],"issn-type":[{"value":"0165-5515","type":"print"},{"value":"1741-6485","type":"electronic"}],"subject":[],"published":{"date-parts":[[2024,3,20]]}}}