{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,8,2]],"date-time":"2025-08-02T16:26:48Z","timestamp":1754152008513,"version":"3.41.2"},"reference-count":52,"publisher":"Association for Computing Machinery (ACM)","issue":"6","content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Knowl. Discov. Data"],"published-print":{"date-parts":[[2025,7,31]]},"abstract":"<jats:p>Topic models are often used to discover latent semantic patterns from document collections. However, existing unsupervised approaches have the following drawbacks: (1) The mined topics may not match user interests; (2) They are prone to extract semantically similar topics and sacrifice diversity; (3) The mined topics often have low interpretability, which does not meet common sense knowledge. To address these limitations, we propose the Distributed Keyword-guided Topic Model (DiskTM) that incorporates Gaussian-distributed keyword prior knowledge into the modeling process to mine user-interested topics. Furthermore, to inject common-sense knowledge and improve the topic\u2019s interpretability, we extend DiskTM and propose the Distributed Keyword-guided Topic Model with Lexical Knowledge (DiskTM-LK). Experimental results on three publicly available text corpora show that our proposed approaches could extract topics that match user interests (keywords). Moreover, DiskTM and DiskTM-LK could also obtain more coherent and diverse topics, outperforming the state-of-the-art baseline approaches.<\/jats:p>","DOI":"10.1145\/3737881","type":"journal-article","created":{"date-parts":[[2025,5,30]],"date-time":"2025-05-30T11:41:36Z","timestamp":1748605296000},"page":"1-19","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":0,"title":["Distributed Keyword-guided Topic Model with Lexical Knowledge Supervision"],"prefix":"10.1145","volume":"19","author":[{"ORCID":"https:\/\/orcid.org\/0000-0001-9350-3667","authenticated-orcid":false,"given":"Rui","family":"Wang","sequence":"first","affiliation":[{"name":"School of Computer Science, Nanjing University of Posts and Telecommunications, Nanjing, China and Key Laboratory of New Generation Artificial Intelligence Technology and Its Interdisciplinary Applications (Southeast University), Ministry of Education, Nanjing, China and Intelligent Interconnected Systems Laboratory of Anhui Province (Hefei University of Technology), Hefei, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0001-9355-9369","authenticated-orcid":false,"given":"Yanan","family":"Wang","sequence":"additional","affiliation":[{"name":"School of Computer Science, Nanjing University of Posts and Telecommunications, Nanjing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0008-4158-8118","authenticated-orcid":false,"given":"Ziang","family":"Li","sequence":"additional","affiliation":[{"name":"School of Computer Science, Nanjing University of Posts and Telecommunications, Nanjing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-9484-3177","authenticated-orcid":false,"given":"Haitao","family":"Cheng","sequence":"additional","affiliation":[{"name":"School of Computer Science, Nanjing University of Posts and Telecommunications, Nanjing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-1888-7001","authenticated-orcid":false,"given":"Guozi","family":"Sun","sequence":"additional","affiliation":[{"name":"School of Computer Science, Nanjing University of Posts and Telecommunications, Nanjing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2025,7,21]]},"reference":[{"key":"e_1_3_2_2_2","first-page":"128","volume-title":"Proceedings of the 19th International Conference on Natural Language Processing (ICON)","author":"Adhya Suman","year":"2022","unstructured":"Suman Adhya, Avishek Lahiri, Debarshi Kumar Sanyal, and Partha Pratim Das. 2022. Improving contextualized topic models with negative sampling. In Proceedings of the 19th International Conference on Natural Language Processing (ICON). New Delhi, India, 128\u2013138."},{"doi-asserted-by":"publisher","key":"e_1_3_2_3_2","DOI":"10.1016\/J.IPM.2022.103146"},{"key":"e_1_3_2_4_2","first-page":"121","volume-title":"Advances in Neural Information Processing Systems 20, Proceedings of the 21st Annual Conference on Neural Information Processing Systems","author":"Blei David M.","year":"2007","unstructured":"David M. Blei and Jon D. McAuliffe. 2007. Supervised topic models. In Advances in Neural Information Processing Systems 20, Proceedings of the 21st Annual Conference on Neural Information Processing Systems. Curran Associates, Inc., 121\u2013128."},{"doi-asserted-by":"publisher","key":"e_1_3_2_5_2","DOI":"10.5555\/944919.944937"},{"key":"e_1_3_2_6_2","first-page":"1","volume-title":"Proceedings of the 7th International Conference on Learning Representations (ICLR \u201919)","author":"Chen Liqun","year":"2019","unstructured":"Liqun Chen, Yizhe Zhang, Ruiyi Zhang, Chenyang Tao, Zhe Gan, Haichao Zhang, Bai Li, Dinghan Shen, Changyou Chen, and Lawrence Carin. 2019. Improving sequence-to-sequence learning via optimal transport. In Proceedings of the 7th International Conference on Learning Representations (ICLR \u201919). OpenReview.net, 1\u201316."},{"key":"e_1_3_2_7_2","first-page":"2071","volume-title":"Proceedings of the 23rd International Joint Conference on Artificial Intelligence (IJCAI \u201913)","author":"Chen Zhiyuan","year":"2013","unstructured":"Zhiyuan Chen, Arjun Mukherjee, Bing Liu, Meichun Hsu, Mal\u00fa Castellanos, and Riddhiman Ghosh. 2013. Leveraging multi-domain prior knowledge in topic models. In Proceedings of the 23rd International Joint Conference on Artificial Intelligence (IJCAI \u201913). IJCAI\/AAAI, 2071\u20132077."},{"doi-asserted-by":"publisher","key":"e_1_3_2_8_2","DOI":"10.18653\/V1\/N19-1423"},{"doi-asserted-by":"publisher","key":"e_1_3_2_9_2","DOI":"10.1162\/TACL_A_00325"},{"doi-asserted-by":"crossref","unstructured":"Shusei Eshima Kosuke Imai and Tomoya Sasaki. 2020. Keyword assisted topic models. arXiv:2004.05964. Retrieved from https:\/\/arxiv.org\/abs\/2004.05964","key":"e_1_3_2_10_2","DOI":"10.32614\/CRAN.package.keyATM"},{"doi-asserted-by":"publisher","key":"e_1_3_2_11_2","DOI":"10.1111\/ajps.12779"},{"doi-asserted-by":"publisher","key":"e_1_3_2_12_2","DOI":"10.1007\/978-90-481-8847-5_10"},{"key":"e_1_3_2_13_2","first-page":"2672","volume-title":"Advances in Neural Information Processing Systems 27: Annual Conference on Neural Information Processing Systems 2014","author":"Goodfellow Ian J.","year":"2014","unstructured":"Ian J. Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron C. Courville, and Yoshua Bengio. 2014. Generative adversarial nets. In Advances in Neural Information Processing Systems 27: Annual Conference on Neural Information Processing Systems 2014, 2672\u20132680."},{"key":"e_1_3_2_14_2","first-page":"723","article-title":"A kernel two-sample test","volume":"13","author":"Gretton Arthur","year":"2012","unstructured":"Arthur Gretton, Karsten M. Borgwardt, Malte J. Rasch, Bernhard Sch\u00f6lkopf, and Alexander J. Smola. 2012. A kernel two-sample test. J. Mach. Learn. Res. 13 (2012), 723\u2013773.","journal-title":"J. Mach. Learn. Res"},{"doi-asserted-by":"publisher","key":"e_1_3_2_15_2","DOI":"10.18653\/v1\/2023.eacl-main.132"},{"doi-asserted-by":"publisher","key":"e_1_3_2_16_2","DOI":"10.1145\/3488560.3498518"},{"doi-asserted-by":"publisher","key":"e_1_3_2_17_2","DOI":"10.18653\/v1\/2020.emnlp-main.725"},{"doi-asserted-by":"publisher","key":"e_1_3_2_18_2","DOI":"10.5555\/2380816.2380844"},{"unstructured":"Diederik P. Kingma and Jimmy Ba. 2014. Adam: A method for stochastic optimization. arXiv:1412.6980. Retrieved from https:\/\/arxiv.org\/abs\/1412.6980","key":"e_1_3_2_19_2"},{"key":"e_1_3_2_20_2","first-page":"1","volume-title":"Proceedings of the 2nd International Conference on Learning Representations (ICLR)","author":"Kingma Diederik P.","year":"2014","unstructured":"Diederik P. Kingma and Max Welling. 2014. Auto-encoding variational bayes. In Proceedings of the 2nd International Conference on Learning Representations (ICLR), 1\u201314."},{"key":"e_1_3_2_21_2","first-page":"375","volume-title":"Advances in Neural Information Processing Systems 15 [Neural Information Processing Systems, NIPS \u201902]","author":"Lafferty John D.","year":"2002","unstructured":"John D. Lafferty and Guy Lebanon. 2002. Information diffusion kernels. In Advances in Neural Information Processing Systems 15 [Neural Information Processing Systems, NIPS \u201902]. MIT Press, 375\u2013382."},{"doi-asserted-by":"publisher","key":"e_1_3_2_22_2","DOI":"10.1016\/B978-1-55860-377-6.50048-7"},{"doi-asserted-by":"publisher","key":"e_1_3_2_23_2","DOI":"10.1016\/J.IPM.2021.102864"},{"doi-asserted-by":"publisher","key":"e_1_3_2_24_2","DOI":"10.1145\/2507157.2507163"},{"doi-asserted-by":"publisher","key":"e_1_3_2_25_2","DOI":"10.1145\/3366423.3380278"},{"key":"e_1_3_2_26_2","first-page":"1727","volume-title":"Proceedings of the 33nd International Conference on Machine Learning (ICML \u201916) (JMLR Workshop and Conference Proceedings)","author":"Miao Yishu","year":"2016","unstructured":"Yishu Miao, Lei Yu, and Phil Blunsom. 2016. Neural variational inference for text processing. In Proceedings of the 33nd International Conference on Machine Learning (ICML \u201916) (JMLR Workshop and Conference Proceedings), 1727\u20131736."},{"key":"e_1_3_2_27_2","first-page":"1","volume-title":"Proceedings of the 1st International Conference on Learning Representations (ICLR \u201913), Workshop Track Proceedings","author":"Mikolov Tom\u00e1s","year":"2013","unstructured":"Tom\u00e1s Mikolov, Kai Chen, Greg Corrado, and Jeffrey Dean. 2013. Efficient estimation of word representations in vector space. In Proceedings of the 1st International Conference on Learning Representations (ICLR \u201913), Workshop Track Proceedings, 1\u201312."},{"doi-asserted-by":"publisher","key":"e_1_3_2_28_2","DOI":"10.18653\/V1\/P19-1640"},{"key":"e_1_3_2_29_2","first-page":"11974","volume-title":"Proceedings of the 34th Advances in Neural Information Processing Systems","author":"Nguyen Thong","year":"2021","unstructured":"Thong Nguyen and Anh Tuan Luu. 2021. Contrastive learning for neural topic model. In Proceedings of the 34th Advances in Neural Information Processing Systems, 11974\u201311986."},{"doi-asserted-by":"publisher","key":"e_1_3_2_30_2","DOI":"10.1016\/J.ESWA.2020.114231"},{"doi-asserted-by":"publisher","key":"e_1_3_2_31_2","DOI":"10.3115\/v1\/D14-1162"},{"doi-asserted-by":"publisher","key":"e_1_3_2_32_2","DOI":"10.1561\/2200000073"},{"doi-asserted-by":"publisher","key":"e_1_3_2_33_2","DOI":"10.3115\/1699510.1699543"},{"doi-asserted-by":"publisher","key":"e_1_3_2_34_2","DOI":"10.1145\/2684822.2685324"},{"doi-asserted-by":"publisher","key":"e_1_3_2_35_2","DOI":"10.1016\/J.KNOSYS.2022.108636"},{"key":"e_1_3_2_36_2","first-page":"1","volume-title":"Proceedings of the 6th International Conference on Learning Representations (ICLR \u201918), Conference Track Proceedings","author":"Tolstikhin Ilya O.","year":"2018","unstructured":"Ilya O. Tolstikhin, Olivier Bousquet, Sylvain Gelly, and Bernhard Sch\u00f6lkopf. 2018. Wasserstein auto-encoders. In Proceedings of the 6th International Conference on Learning Representations (ICLR \u201918), Conference Track Proceedings. OpenReview.net, 1\u201316."},{"doi-asserted-by":"publisher","key":"e_1_3_2_37_2","DOI":"10.1145\/3510003.3510201"},{"issue":"11","key":"e_1_3_2_38_2","first-page":"2579","article-title":"Visualizing data using t-SNE","volume":"9","author":"Van der Maaten Laurens","year":"2008","unstructured":"Laurens Van der Maaten and Geoffrey Hinton. 2008. Visualizing data using t-SNE. J. Mach. Learn. Res. 9, 11 (2008), 2579\u20132605.","journal-title":"J. Mach. Learn. Res"},{"key":"e_1_3_2_39_2","first-page":"1","article-title":"Rethinking LDA: Why priors matter","volume":"22","author":"Wallach Hanna","year":"2009","unstructured":"Hanna Wallach, David Mimno, and Andrew McCallum. 2009. Rethinking LDA: Why priors matter. Adv. Neural Inf. Process. Syst. 22 (2009), 1\u20139.","journal-title":"Adv. Neural Inf. Process. Syst"},{"doi-asserted-by":"publisher","key":"e_1_3_2_40_2","DOI":"10.18653\/V1\/2020.ACL-MAIN.32"},{"doi-asserted-by":"publisher","key":"e_1_3_2_41_2","DOI":"10.1016\/j.ipm.2019.102098"},{"doi-asserted-by":"publisher","key":"e_1_3_2_42_2","DOI":"10.18653\/v1\/D19-1027"},{"key":"e_1_3_2_43_2","first-page":"37335","volume-title":"Proceedings of the International Conference on Machine Learning (ICML \u201923)","author":"Wu Xiaobao","year":"2023","unstructured":"Xiaobao Wu, Xinshuai Dong, Thong Thanh Nguyen, and Anh Tuan Luu. 2023. Effective neural topic modeling with embedding clustering regularization. In Proceedings of the International Conference on Machine Learning (ICML \u201923). PMLR, 37335\u201337357."},{"unstructured":"Bing Xu Naiyan Wang Tianqi Chen and Mu Li. 2015. Empirical evaluation of rectified activations in convolutional network. arXiv:1505.00853. Retrieved from https:\/\/arxiv.org\/abs\/1505.00853","key":"e_1_3_2_44_2"},{"doi-asserted-by":"publisher","key":"e_1_3_2_45_2","DOI":"10.18653\/V1\/2023.FINDINGS-ACL.271"},{"doi-asserted-by":"publisher","key":"e_1_3_2_46_2","DOI":"10.18653\/V1\/D15-1037"},{"key":"e_1_3_2_47_2","volume-title":"Advances in Neural Information Processing Systems","author":"Zhang Xiang","year":"2015","unstructured":"Xiang Zhang, Junbo Zhao, and Yann LeCun. 2015. Character-level convolutional networks for text classification. In Advances in Neural Information Processing Systems. C. Cortes, N. Lawrence, D. Lee, M. Sugiyama, and R. Garnett (Eds.), Vol. 28. Curran Associates, Inc."},{"key":"e_1_3_2_48_2","first-page":"649","volume-title":"Advances in Neural Information Processing Systems 28: Annual Conference on Neural Information Processing Systems 2015","author":"Zhang Xiang","year":"2015","unstructured":"Xiang Zhang, Junbo Jake Zhao, and Yann LeCun. 2015. Character-level convolutional networks for text classification. In Advances in Neural Information Processing Systems 28: Annual Conference on Neural Information Processing Systems 2015, 649\u2013657."},{"doi-asserted-by":"publisher","key":"e_1_3_2_49_2","DOI":"10.1109\/TKDE.2021.3093350"},{"doi-asserted-by":"publisher","key":"e_1_3_2_50_2","DOI":"10.1145\/3397271.3401168"},{"doi-asserted-by":"publisher","key":"e_1_3_2_51_2","DOI":"10.1145\/3539597.3570475"},{"doi-asserted-by":"publisher","key":"e_1_3_2_52_2","DOI":"10.1016\/J.IPM.2022.103215"},{"doi-asserted-by":"publisher","key":"e_1_3_2_53_2","DOI":"10.1016\/J.IPM.2022.103230"}],"container-title":["ACM Transactions on Knowledge Discovery from Data"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3737881","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,7,21]],"date-time":"2025-07-21T11:02:59Z","timestamp":1753095779000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3737881"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,7,21]]},"references-count":52,"journal-issue":{"issue":"6","published-print":{"date-parts":[[2025,7,31]]}},"alternative-id":["10.1145\/3737881"],"URL":"https:\/\/doi.org\/10.1145\/3737881","relation":{},"ISSN":["1556-4681","1556-472X"],"issn-type":[{"type":"print","value":"1556-4681"},{"type":"electronic","value":"1556-472X"}],"subject":[],"published":{"date-parts":[[2025,7,21]]},"assertion":[{"value":"2024-07-31","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-05-15","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-07-21","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}