{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,1,19]],"date-time":"2026-01-19T06:33:22Z","timestamp":1768804402472,"version":"3.49.0"},"reference-count":32,"publisher":"Wiley","issue":"1","license":[{"start":{"date-parts":[[2018,5,16]],"date-time":"2018-05-16T00:00:00Z","timestamp":1526428800000},"content-version":"vor","delay-in-days":135,"URL":"http:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["61672553"],"award-info":[{"award-number":["61672553"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100002338","name":"Ministry of Education of the People's Republic of China","doi-asserted-by":"publisher","award":["16YJCZH076"],"award-info":[{"award-number":["16YJCZH076"]}],"id":[{"id":"10.13039\/501100002338","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100002338","name":"Ministry of Education of the People's Republic of China","doi-asserted-by":"publisher","award":["2018KF01"],"award-info":[{"award-number":["2018KF01"]}],"id":[{"id":"10.13039\/501100002338","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["onlinelibrary.wiley.com"],"crossmark-restriction":true},"short-container-title":["Complexity"],"published-print":{"date-parts":[[2018,1]]},"abstract":"<jats:p>In the present big data background, how to effectively excavate useful information is the problem that big data is facing now. The purpose of this study is to construct a more effective method of mining interest preferences of users in a particular field in the context of today\u2019s big data. We mainly use a large number of user text data from microblog to study. LDA is an effective method of text mining, but it will not play a very good role in applying LDA directly to a large number of short texts in microblog. In today\u2019s more effective topic modeling project, short texts need to be aggregated into long texts to avoid data sparsity. However, aggregated short texts are mixed with a lot of noise, reducing the accuracy of mining the user\u2019s interest preferences. In this paper, we propose Combining Latent Dirichlet Allocation (CLDA), a new topic model that can learn the potential topics of microblog short texts and long texts simultaneously. The data sparsity of short texts is avoided by aggregating long texts to assist in learning short texts. Short text filtering long text is reused to improve mining accuracy, making long texts and short texts effectively combined. Experimental results in a real microblog data set show that CLDA outperforms many advanced models in mining user interest, and we also confirm that CLDA also has good performance in recommending systems.<\/jats:p>","DOI":"10.1155\/2018\/2503816","type":"journal-article","created":{"date-parts":[[2018,5,16]],"date-time":"2018-05-16T23:55:29Z","timestamp":1526514929000},"update-policy":"https:\/\/doi.org\/10.1002\/crossmark_policy","source":"Crossref","is-referenced-by-count":8,"title":["CLDA: An Effective Topic Model for Mining User Interest Preference under Big Data Background"],"prefix":"10.1155","volume":"2018","author":[{"ORCID":"https:\/\/orcid.org\/0000-0001-6489-1648","authenticated-orcid":false,"given":"Lirong","family":"Qiu","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Jia","family":"Yu","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"311","published-online":{"date-parts":[[2018,5,16]]},"reference":[{"key":"e_1_2_10_1_2","doi-asserted-by":"publisher","DOI":"10.1109\/CC.2014.6911095"},{"key":"e_1_2_10_2_2","doi-asserted-by":"crossref","unstructured":"HofmannT. Probabilistic latent semantic indexing Proceedings of the 22nd Annual International ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR \u203299) 1999 Berkeley Calif USA 50\u201357 https:\/\/doi.org\/10.1145\/312624.312649.","DOI":"10.1145\/312624.312649"},{"key":"e_1_2_10_3_2","doi-asserted-by":"publisher","DOI":"10.1162\/jmlr.2003.3.4-5.993"},{"key":"e_1_2_10_4_2","doi-asserted-by":"crossref","unstructured":"WengJ. LimE. JiangJ. andHeQ. TwitterRank: finding topic-sensitive influential twitterers Proceedings of the 3rd ACM International Conference on Web Search and Data Mining (WSDM \u203210) February 2010 New York NY USA 261\u2013270 https:\/\/doi.org\/10.1145\/1718487.1718520 2-s2.0-77950897279.","DOI":"10.1145\/1718487.1718520"},{"key":"e_1_2_10_5_2","doi-asserted-by":"publisher","DOI":"10.1145\/1964858.1964870"},{"key":"e_1_2_10_6_2","first-page":"391","article-title":"Indexing by latent semantic analysis","volume":"41","author":"Deerwester S.","year":"1990","journal-title":"Journal of the Association for Information Science and Technology"},{"key":"e_1_2_10_7_2","first-page":"391","article-title":"Indexing by latent semantic analysis","volume":"41","author":"Deerwester S.","year":"2010","journal-title":"Journal of the Association for Information Science & Technology"},{"key":"e_1_2_10_8_2","doi-asserted-by":"publisher","DOI":"10.1023\/A:1007617005950"},{"key":"e_1_2_10_9_2","doi-asserted-by":"crossref","unstructured":"XiongH. A location-sentiment-aware recommender system for both home-town and out-of-town users Proceeding of the ACM SIGKDD International Conference on Knowledge Discovery and Data Mining 2017 1135\u20131143.","DOI":"10.1145\/3097983.3098122"},{"key":"e_1_2_10_10_2","doi-asserted-by":"crossref","unstructured":"LianD. ZhaoC. XieX. SunG. ChenE. andRuiY. GeoMF: Joint geographical modeling and matrix factorization for point-of-interest recommendation Proceedings of the 20th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining KDD 2014 August 2014 USA 831\u2013840 2-s2.0-84907022562 https:\/\/doi.org\/10.1145\/2623330.2623638.","DOI":"10.1145\/2623330.2623638"},{"key":"e_1_2_10_11_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.ipm.2017.08.001"},{"key":"e_1_2_10_12_2","doi-asserted-by":"crossref","unstructured":"EfronM. OrganisciakP. andFenlonK. Improving retrieval of short texts through document expansion Proceedings of the 35th Annual ACM SIGIR Conference on Research and Development in Information Retrieval SIGIR 2012 August 2012 usa 911\u2013920 2-s2.0-84866635150 https:\/\/doi.org\/10.1145\/2348283.2348405.","DOI":"10.1145\/2348283.2348405"},{"key":"e_1_2_10_13_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.neucom.2015.11.102"},{"key":"e_1_2_10_14_2","doi-asserted-by":"crossref","unstructured":"SahamiM.andHeilmanT. D. A web-based kernel function for measuring the similarity of short text snippets Proceedings of the International Conference on World Wide Web (WWW \u203206) 2006 Edinburgh Scotland 377\u2013386 https:\/\/doi.org\/10.1145\/1135777.1135834.","DOI":"10.1145\/1135777.1135834"},{"key":"e_1_2_10_15_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.neucom.2015.09.096"},{"key":"e_1_2_10_16_2","doi-asserted-by":"crossref","unstructured":"HuX. SunN. ZhangC. andChuaT. S. Exploiting internal and external semantics for the clustering of short texts using world knowledge Proceedings of the Acm Conference on Information & Knowledge Management 2009 919\u2013928.","DOI":"10.1145\/1645953.1646071"},{"key":"e_1_2_10_17_2","first-page":"1","article-title":"Sentence similarity based on semantic kernels for intelligent text retrieval","author":"Amir S.","year":"2016","journal-title":"Journal of Intelligent Information Systems"},{"key":"e_1_2_10_18_2","doi-asserted-by":"crossref","unstructured":"SchlaeferN. Chu-CarrollJ. NybergE. FanJ. ZadroznyW. andFerrucciD. Statistical source expansion for question answering Proceedings of the 20th ACM Conference on Information and Knowledge Management CIKM\u203211 October 2011 gbr 345\u2013354 2-s2.0-83055181778 https:\/\/doi.org\/10.1145\/2063576.2063632.","DOI":"10.1145\/2063576.2063632"},{"key":"e_1_2_10_19_2","doi-asserted-by":"crossref","unstructured":"DaltonJ. DietzL. andAllanJ. Entity query feature expansion using knowledge base links Proceedings of the 37th International ACM SIGIR Conference on Research and Development in Information Retrieval SIGIR 2014 July 2014 aus 365\u2013374 2-s2.0-84904580855 https:\/\/doi.org\/10.1145\/2600428.2609628.","DOI":"10.1145\/2600428.2609628"},{"key":"e_1_2_10_20_2","doi-asserted-by":"crossref","unstructured":"TangJ. WangY. ZhengK. andMeiQ. End-to-end learning for short text expansion Proceedings of the 23rd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining KDD 2017 August 2017 can 1105\u20131113 2-s2.0-85029066725 https:\/\/doi.org\/10.1145\/3097983.3098166.","DOI":"10.1145\/3097983.3098166"},{"key":"e_1_2_10_21_2","unstructured":"RosenZviM. GriffithsT. SteyversM. andSmythP. The author-topic model for authors and documents Proceedings of the 20th Conference on Uncertainty in Artificial Intelligence (UAI \u203204) 2012 487\u2013494."},{"key":"e_1_2_10_22_2","doi-asserted-by":"crossref","unstructured":"RamageD. HallD. NallapatiR. andManningC. D. Labeled LDA: a supervised topic model for credit attribution in multi-labeled corpora Proceedings of the Conference on Empirical Methods in Natural Language Processing: Volume Association for Computational Linguistics 2009 248\u2013256.","DOI":"10.3115\/1699510.1699543"},{"key":"e_1_2_10_23_2","doi-asserted-by":"crossref","unstructured":"RamageD. DumaisS. andLieblingD. Characterizing microblogs with topic models Proceedings of the International Conference on Weblogs and Social Media (ICWSM \u203210) 2010 Washington Dc Wash USA 130\u2013137 2-s2.0-84890724114.","DOI":"10.1609\/icwsm.v4i1.14026"},{"key":"e_1_2_10_24_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-642-20161-5_34"},{"key":"e_1_2_10_25_2","doi-asserted-by":"publisher","DOI":"10.1109\/TKDE.2010.27"},{"key":"e_1_2_10_26_2","first-page":"52","article-title":"Domain identification for intention posts on online social media","author":"Luong T. L.","year":"2016","journal-title":"Symposium on Information and Communication Technology"},{"key":"e_1_2_10_27_2","first-page":"1260","article-title":"Performance of using LDA for Chinese news text classification","volume":"2015","author":"Wu X.","year":"2015","journal-title":"Electrical and Computer Engineering"},{"key":"e_1_2_10_28_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.cortex.2016.03.020"},{"key":"e_1_2_10_29_2","volume-title":"Improving Short Text Classification Using Public Search Engines. Uncertainty in Knowledge Modelling and Decision Making","author":"Meng W.","year":"2013"},{"key":"e_1_2_10_30_2","doi-asserted-by":"publisher","DOI":"10.1109\/TASLP.2016.2520371"},{"key":"e_1_2_10_31_2","first-page":"1795","article-title":"Topic mining for microblog based on MB-LDA model","volume":"48","author":"Zhang C.","year":"2011","journal-title":"Computer Research and Development"},{"key":"e_1_2_10_32_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.ins.2016.05.031"}],"container-title":["Complexity"],"original-title":[],"language":"en","link":[{"URL":"http:\/\/downloads.hindawi.com\/journals\/complexity\/2018\/2503816.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"http:\/\/downloads.hindawi.com\/journals\/complexity\/2018\/2503816.xml","content-type":"application\/xml","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/onlinelibrary.wiley.com\/doi\/pdf\/10.1155\/2018\/2503816","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,7,4]],"date-time":"2025-07-04T15:44:17Z","timestamp":1751643857000},"score":1,"resource":{"primary":{"URL":"https:\/\/onlinelibrary.wiley.com\/doi\/10.1155\/2018\/2503816"}},"subtitle":[],"editor":[{"given":"Zhihan","family":"Lv","sequence":"additional","affiliation":[],"role":[{"role":"editor","vocabulary":"crossref"}]}],"short-title":[],"issued":{"date-parts":[[2018,1]]},"references-count":32,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2018,1]]}},"alternative-id":["10.1155\/2018\/2503816"],"URL":"https:\/\/doi.org\/10.1155\/2018\/2503816","archive":["Portico"],"relation":{},"ISSN":["1076-2787","1099-0526"],"issn-type":[{"value":"1076-2787","type":"print"},{"value":"1099-0526","type":"electronic"}],"subject":[],"published":{"date-parts":[[2018,1]]},"assertion":[{"value":"2018-03-05","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2018-04-11","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2018-05-16","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}],"article-number":"2503816"}}