{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,2,20]],"date-time":"2026-02-20T12:52:21Z","timestamp":1771591941424,"version":"3.50.1"},"reference-count":49,"publisher":"Association for Computing Machinery (ACM)","issue":"2","license":[{"start":{"date-parts":[[2017,8,21]],"date-time":"2017-08-21T00:00:00Z","timestamp":1503273600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"crossref","award":["61502344"],"award-info":[{"award-number":["61502344"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]},{"name":"Singapore Ministry of Education Academic Research Fund Tier 2","award":["MOE2014-T2-2-066"],"award-info":[{"award-number":["MOE2014-T2-2-066"]}]},{"name":"Academic Team Building Plan for Young Scholars from Wuhan University","award":["Whu2016012"],"award-info":[{"award-number":["Whu2016012"]}]},{"name":"Natural Scientific Research Program of Wuhan University","award":["2042017kf0225 and 2042016kf0190"],"award-info":[{"award-number":["2042017kf0225 and 2042016kf0190"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Inf. Syst."],"published-print":{"date-parts":[[2018,4,30]]},"abstract":"<jats:p>Many applications require semantic understanding of short texts, and inferring discriminative and coherent latent topics is a critical and fundamental task in these applications. Conventional topic models largely rely on word co-occurrences to derive topics from a collection of documents. However, due to the length of each document, short texts are much more sparse in terms of word co-occurrences. Recent studies show that the Dirichlet Multinomial Mixture (DMM) model is effective for topic inference over short texts by assuming that each piece of short text is generated by a single topic. However, DMM has two main limitations. First, even though it seems reasonable to assume that each short text has only one topic because of its shortness, the definition of \u201cshortness\u201d is subjective and the length of the short texts is dataset dependent. That is, the single-topic assumption may be too strong for some datasets. To address this limitation, we propose to model the topic number as a Poisson distribution, allowing each short text to be associated with a small number of topics (e.g., one to three topics). This model is named PDMM. Second, DMM (and also PDMM) does not have access to background knowledge (e.g., semantic relations between words) when modeling short texts. When a human being interprets a piece of short text, the understanding is not solely based on its content words, but also their semantic relations. Recent advances in word embeddings offer effective learning of word semantic relations from a large corpus. Such auxiliary word embeddings enable us to address the second limitation. To this end, we propose to promote the semantically related words under the same topic during the sampling process, by using the generalized P\u00f3lya urn (GPU) model. Through the GPU model, background knowledge about word semantic relations learned from millions of external documents can be easily exploited to improve topic modeling for short texts. By directly extending the PDMM model with the GPU model, we propose two more effective topic models for short texts, named GPU-DMM and GPU-PDMM. Through extensive experiments on two real-world short text collections in two languages, we demonstrate that PDMM achieves better topic representations than state-of-the-art models, measured by topic coherence. The learned topic representation leads to better accuracy in a text classification task, as an indirect evaluation. Both GPU-DMM and GPU-PDMM further improve topic coherence and text classification accuracy. GPU-PDMM outperforms GPU-DMM at the price of higher computational costs.<\/jats:p>","DOI":"10.1145\/3091108","type":"journal-article","created":{"date-parts":[[2017,8,24]],"date-time":"2017-08-24T11:49:04Z","timestamp":1503575344000},"page":"1-30","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":108,"title":["Enhancing Topic Modeling for Short Texts with Auxiliary Word Embeddings"],"prefix":"10.1145","volume":"36","author":[{"given":"Chenliang","family":"Li","sequence":"first","affiliation":[{"name":"Wuhan University, Wuhan, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yu","family":"Duan","sequence":"additional","affiliation":[{"name":"Wuhan University, Wuhan, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Haoran","family":"Wang","sequence":"additional","affiliation":[{"name":"Wuhan University, Wuhan, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Zhiqian","family":"Zhang","sequence":"additional","affiliation":[{"name":"Wuhan University"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Aixin","family":"Sun","sequence":"additional","affiliation":[{"name":"Nanyang Technological University, Nanyang Avenue, Singapore"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Zongyang","family":"Ma","sequence":"additional","affiliation":[{"name":"Nanyang Technological University, Nanyang Avenue, Singapore"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2017,8,21]]},"reference":[{"key":"e_1_2_1_1_1","volume-title":"Neural Probabilistic Language Models","author":"Bengio Yoshua","unstructured":"Yoshua Bengio , Holger Schwenk , Jean-S\u00e9bastien Sen\u00e9cal , Fr\u00e9deric Morin , and Jean-Luc Gauvain . 2006. Neural Probabilistic Language Models . Springer . 137--186 pages Yoshua Bengio, Holger Schwenk, Jean-S\u00e9bastien Sen\u00e9cal, Fr\u00e9deric Morin, and Jean-Luc Gauvain. 2006. Neural Probabilistic Language Models. Springer. 137--186 pages"},{"key":"e_1_2_1_2_1","doi-asserted-by":"publisher","DOI":"10.5555\/944919.944937"},{"key":"e_1_2_1_3_1","volume-title":"Blei","author":"Chang Jonathan","year":"2009","unstructured":"Jonathan Chang , Sean Gerrish , Chong Wang , Jordan L. Boyd-Graber , and David M . Blei . 2009 . Reading tea leaves: How humans interpret topic models. In NIPS. Curran Associates , 288--296. Jonathan Chang, Sean Gerrish, Chong Wang, Jordan L. Boyd-Graber, and David M. Blei. 2009. Reading tea leaves: How humans interpret topic models. In NIPS. Curran Associates, 288--296."},{"key":"e_1_2_1_4_1","doi-asserted-by":"publisher","DOI":"10.5555\/2283696.2283700"},{"key":"e_1_2_1_5_1","doi-asserted-by":"publisher","DOI":"10.1145\/2623330.2623622"},{"key":"e_1_2_1_6_1","doi-asserted-by":"publisher","DOI":"10.1145\/2505515.2505519"},{"key":"e_1_2_1_7_1","volume-title":"IJCAI. IJCAI\/AAAI","author":"Chen Zhiyuan","year":"2013","unstructured":"Zhiyuan Chen , Arjun Mukherjee , Bing Liu , Meichun Hsu , Mal\u00fa Castellanos , and Riddhiman Ghosh . 2013 a. Leveraging multi-domain prior knowledge in topic models . In IJCAI. IJCAI\/AAAI , 2071--2077. Zhiyuan Chen, Arjun Mukherjee, Bing Liu, Meichun Hsu, Mal\u00fa Castellanos, and Riddhiman Ghosh. 2013a. Leveraging multi-domain prior knowledge in topic models. In IJCAI. IJCAI\/AAAI, 2071--2077."},{"key":"e_1_2_1_8_1","doi-asserted-by":"publisher","DOI":"10.1109\/TKDE.2014.2313872"},{"key":"e_1_2_1_9_1","doi-asserted-by":"publisher","DOI":"10.1145\/1390156.1390177"},{"key":"e_1_2_1_10_1","doi-asserted-by":"publisher","DOI":"10.5555\/1953048.2078186"},{"key":"e_1_2_1_11_1","volume-title":"Gaussian LDA for topic models with word embeddings","author":"Das Rajarshi","unstructured":"Rajarshi Das , Manzil Zaheer , and Chris Dyer . 2015. Gaussian LDA for topic models with word embeddings . In ACL. Association for Computer Linguistics , 795--804. Rajarshi Das, Manzil Zaheer, and Chris Dyer. 2015. Gaussian LDA for topic models with word embeddings. In ACL. Association for Computer Linguistics, 795--804."},{"key":"e_1_2_1_12_1","doi-asserted-by":"publisher","DOI":"10.1145\/312624.312649"},{"key":"e_1_2_1_13_1","doi-asserted-by":"publisher","DOI":"10.1145\/1964858.1964870"},{"key":"e_1_2_1_14_1","doi-asserted-by":"publisher","DOI":"10.1145\/2020408.2020485"},{"key":"e_1_2_1_15_1","volume-title":"A latent concept topic model for robust topic inference using word embeddings","author":"Hu Weihua","unstructured":"Weihua Hu and Jun\u2019ichi Tsujii . 2016. A latent concept topic model for robust topic inference using word embeddings . In ACL. Association for Computer Linguistics , 380--386. Weihua Hu and Jun\u2019ichi Tsujii. 2016. A latent concept topic model for robust topic inference using word embeddings. In ACL. Association for Computer Linguistics, 380--386."},{"key":"e_1_2_1_16_1","doi-asserted-by":"publisher","DOI":"10.1145\/2063576.2063689"},{"key":"e_1_2_1_17_1","doi-asserted-by":"publisher","DOI":"10.1145\/2806416.2806475"},{"key":"e_1_2_1_18_1","volume-title":"Weinberger","author":"Kusner Matt J.","year":"2015","unstructured":"Matt J. Kusner , Yu Sun , Nicholas I. Kolkin , and Kilian Q . Weinberger . 2015 . From word embeddings to document distances. In ICML. 957--966. Matt J. Kusner, Yu Sun, Nicholas I. Kolkin, and Kilian Q. Weinberger. 2015. From word embeddings to document distances. In ICML. 957--966."},{"key":"e_1_2_1_19_1","doi-asserted-by":"publisher","DOI":"10.1145\/2911451.2911499"},{"key":"e_1_2_1_20_1","doi-asserted-by":"publisher","DOI":"10.1145\/2396761.2396798"},{"key":"e_1_2_1_21_1","volume-title":"Polya urn models","author":"Mahmoud Hosam","unstructured":"Hosam Mahmoud . 2008. Polya urn models . CRC press . Hosam Mahmoud. 2008. Polya urn models. CRC press."},{"key":"e_1_2_1_22_1","doi-asserted-by":"publisher","DOI":"10.1145\/2484028.2484166"},{"key":"e_1_2_1_23_1","unstructured":"Tomas Mikolov Kai Chen Greg Corrada and Jeffrey Dean. 2013. Efficient estimation of word representations in vector space. arXiv preprint arXiv:1301.3781.  Tomas Mikolov Kai Chen Greg Corrada and Jeffrey Dean. 2013. Efficient estimation of word representations in vector space. arXiv preprint arXiv:1301.3781."},{"key":"e_1_2_1_24_1","volume-title":"Optimizing semantic coherence in topic models","author":"Mimno David","unstructured":"David Mimno , Hanna M. Wallach , Edmund Talley , Miriam Leenders , and Andrew McCallum . 2011. Optimizing semantic coherence in topic models . In EMNLP. Association for Computational Linguistics , 262--272. David Mimno, Hanna M. Wallach, Edmund Talley, Miriam Leenders, and Andrew McCallum. 2011. Optimizing semantic coherence in topic models. In EMNLP. Association for Computational Linguistics, 262--272."},{"key":"e_1_2_1_25_1","volume-title":"Hinton","author":"Mnih Andriy","year":"2009","unstructured":"Andriy Mnih and Geoffrey E . Hinton . 2009 . A scalable hierarchical distributed language model. In NIPS. Curran Associates , 1081--1088. Andriy Mnih and Geoffrey E. Hinton. 2009. A scalable hierarchical distributed language model. In NIPS. Curran Associates, 1081--1088."},{"key":"e_1_2_1_26_1","volume-title":"Karl Grieser, and Timothy Baldwin.","author":"Newman David","year":"2010","unstructured":"David Newman , Jey Han Lau , Karl Grieser, and Timothy Baldwin. 2010 . Automatic evaluation of topic coherence. In HLT-NAACL. Association for Computational Linguistics , 100--108. David Newman, Jey Han Lau, Karl Grieser, and Timothy Baldwin. 2010. Automatic evaluation of topic coherence. In HLT-NAACL. Association for Computational Linguistics, 100--108."},{"key":"e_1_2_1_27_1","doi-asserted-by":"publisher","DOI":"10.1162\/tacl_a_00140"},{"key":"e_1_2_1_28_1","doi-asserted-by":"publisher","DOI":"10.1145\/1102351.1102430"},{"key":"e_1_2_1_29_1","doi-asserted-by":"publisher","DOI":"10.1023\/A:1007692713085"},{"key":"e_1_2_1_30_1","volume-title":"Manning","author":"Pennington Jeffrey","year":"2014","unstructured":"Jeffrey Pennington , Richard Socher , and Christopher D . Manning . 2014 . Glove : Global vectors for word representation. In EMNLP. Association for Computational Linguistics , 1532--1543. Jeffrey Pennington, Richard Socher, and Christopher D. Manning. 2014. Glove: Global vectors for word representation. In EMNLP. Association for Computational Linguistics, 1532--1543."},{"key":"e_1_2_1_31_1","doi-asserted-by":"publisher","DOI":"10.1145\/1367497.1367510"},{"key":"e_1_2_1_32_1","doi-asserted-by":"publisher","DOI":"10.1145\/1401890.1401960"},{"key":"e_1_2_1_33_1","volume-title":"Short and sparse text topic modeling via self-aggregation","author":"Quan Xiaojun","unstructured":"Xiaojun Quan , Chunyu Kit , Yong Ge , and Sinno Jialin Pan . 2015. Short and sparse text topic modeling via self-aggregation . In AAAI. AAAI Press , 2270--2276. Xiaojun Quan, Chunyu Kit, Yong Ge, and Sinno Jialin Pan. 2015. Short and sparse text topic modeling via self-aggregation. In AAAI. AAAI Press, 2270--2276."},{"key":"e_1_2_1_34_1","volume-title":"Liebling","author":"Ramage Daniel","year":"2010","unstructured":"Daniel Ramage , Susan T. Dumais , and Daniel J . Liebling . 2010 . Characterizing microblogs with topic models. In ICWSM. Press , 130--137. Daniel Ramage, Susan T. Dumais, and Daniel J. Liebling. 2010. Characterizing microblogs with topic models. In ICWSM. Press, 130--137."},{"key":"e_1_2_1_35_1","volume-title":"Williams","author":"Rumelhart David E.","year":"1988","unstructured":"David E. Rumelhart , Geoffrey E. Hinton , and Ronald J . Williams . 1988 . Neurocomputing : Foundations of research, J. A. Anderson and E. Rosenfeld (Eds .). 696--699. David E. Rumelhart, Geoffrey E. Hinton, and Ronald J. Williams. 1988. Neurocomputing: Foundations of research, J. A. Anderson and E. Rosenfeld (Eds.). 696--699."},{"key":"e_1_2_1_36_1","doi-asserted-by":"publisher","DOI":"10.3115\/v1\/W15-1526"},{"key":"e_1_2_1_37_1","doi-asserted-by":"publisher","DOI":"10.1145\/1835449.1835643"},{"key":"e_1_2_1_38_1","doi-asserted-by":"publisher","DOI":"10.1080\/00268978600100071"},{"key":"e_1_2_1_39_1","doi-asserted-by":"publisher","DOI":"10.1145\/2348283.2348511"},{"key":"e_1_2_1_40_1","doi-asserted-by":"publisher","DOI":"10.1145\/2020408.2020480"},{"key":"e_1_2_1_41_1","doi-asserted-by":"publisher","DOI":"10.1145\/1281192.1281276"},{"key":"e_1_2_1_42_1","doi-asserted-by":"publisher","DOI":"10.1145\/1718487.1718520"},{"key":"e_1_2_1_43_1","doi-asserted-by":"publisher","DOI":"10.1145\/2488388.2488514"},{"key":"e_1_2_1_44_1","doi-asserted-by":"publisher","DOI":"10.1145\/2623330.2623715"},{"key":"e_1_2_1_45_1","volume-title":"Discriminative bi-term topic model for headline-based social news clustering","author":"Yunqing Xia","unstructured":"Xia Yunqing , Tang Nan , Hussain Amir , and Cambria Erik . 2015. Discriminative bi-term topic model for headline-based social news clustering . In AAAI. AAAI Press , 311--316. Xia Yunqing, Tang Nan, Hussain Amir, and Cambria Erik. 2015. Discriminative bi-term topic model for headline-based social news clustering. In AAAI. AAAI Press, 311--316."},{"key":"e_1_2_1_46_1","volume-title":"Comparing twitter and traditional media using topic models","author":"Zhao Wayne Xin","unstructured":"Wayne Xin Zhao , Jing Jiang , Jianshu Weng , Jing He , Ee-Peng Lim , Hongfei Yan , and Xiaoming Li. 2011. Comparing twitter and traditional media using topic models . In ECIR. Springer , 338--349. Wayne Xin Zhao, Jing Jiang, Jianshu Weng, Jing He, Ee-Peng Lim, Hongfei Yan, and Xiaoming Li. 2011. Comparing twitter and traditional media using topic models. In ECIR. Springer, 338--349."},{"key":"e_1_2_1_47_1","doi-asserted-by":"publisher","DOI":"10.1145\/2766462.2767700"},{"key":"e_1_2_1_48_1","doi-asserted-by":"publisher","DOI":"10.1145\/2939672.2939880"},{"key":"e_1_2_1_49_1","doi-asserted-by":"publisher","DOI":"10.1007\/s10115-015-0882-z"}],"container-title":["ACM Transactions on Information Systems"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3091108","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3091108","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T03:37:28Z","timestamp":1750217848000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3091108"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2017,8,21]]},"references-count":49,"journal-issue":{"issue":"2","published-print":{"date-parts":[[2018,4,30]]}},"alternative-id":["10.1145\/3091108"],"URL":"https:\/\/doi.org\/10.1145\/3091108","relation":{},"ISSN":["1046-8188","1558-2868"],"issn-type":[{"value":"1046-8188","type":"print"},{"value":"1558-2868","type":"electronic"}],"subject":[],"published":{"date-parts":[[2017,8,21]]},"assertion":[{"value":"2016-12-01","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2017-04-01","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2017-08-21","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}