{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,3,30]],"date-time":"2026-03-30T14:39:48Z","timestamp":1774881588733,"version":"3.50.1"},"reference-count":31,"publisher":"Springer Science and Business Media LLC","issue":"3","license":[{"start":{"date-parts":[[2021,3,3]],"date-time":"2021-03-03T00:00:00Z","timestamp":1614729600000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2021,3,3]],"date-time":"2021-03-03T00:00:00Z","timestamp":1614729600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"name":"Alma Mater Studiorum - Universit\u00e0 di Bologna"}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Stat Comput"],"published-print":{"date-parts":[[2021,5]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:p>Mixtures of unigrams are one of the simplest and most efficient tools for clustering textual data, as they assume that documents related to the same topic have similar distributions of terms, naturally described by multinomials. When the classification task is particularly challenging, such as when the document-term matrix is high-dimensional and extremely sparse, a more composite representation can provide better insight into the grouping structure. In this work, we developed a deep version of mixtures of unigrams for the unsupervised classification of very short documents with a large number of terms, by allowing for models with further deeper latent layers; the proposal is derived in a Bayesian framework. The behavior of the deep mixtures of unigrams is empirically compared with that of other traditional and state-of-the-art methods, namely<jats:italic>k<\/jats:italic>-means with cosine distance,<jats:italic>k<\/jats:italic>-means with Euclidean distance on data transformed according to semantic analysis, partition around medoids, mixture of Gaussians on semantic-based transformed data, hierarchical clustering according to Ward\u2019s method with cosine dissimilarity, latent Dirichlet allocation, mixtures of unigrams estimated via the EM algorithm, spectral clustering and affinity propagation clustering. The performance is evaluated in terms of both correct classification rate and Adjusted Rand Index. Simulation studies and real data analysis prove that going deep in clustering such data highly improves the classification accuracy.<\/jats:p>","DOI":"10.1007\/s11222-020-09989-9","type":"journal-article","created":{"date-parts":[[2021,3,3]],"date-time":"2021-03-03T08:09:43Z","timestamp":1614758983000},"update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":6,"title":["Deep mixtures of unigrams for uncovering topics in textual data"],"prefix":"10.1007","volume":"31","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-3278-5266","authenticated-orcid":false,"given":"Cinzia","family":"Viroli","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Laura","family":"Anderlucci","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"297","published-online":{"date-parts":[[2021,3,3]]},"reference":[{"issue":"7","key":"9989_CR1","doi-asserted-by":"publisher","first-page":"1298","DOI":"10.1109\/TPAMI.2009.149","volume":"32","author":"J Baek","year":"2010","unstructured":"Baek, J., McLachlan, G., Flack, L.: Mixtures of factor analyzers with common factor loadings: applications to the clustering and visualization of high-dimensional data. IEEE Trans. Pattern Anal. Mach. Intell. 32(7), 1298\u20131309 (2010)","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"9989_CR2","first-page":"993","volume":"3","author":"DM Blei","year":"2003","unstructured":"Blei, D.M., Ng, A.Y., Jordan, M.I.: Latent Dirichlet allocation. J. Mach. Learn. Res. 3, 993\u20131022 (2003)","journal-title":"J. Mach. Learn. Res."},{"key":"9989_CR3","doi-asserted-by":"crossref","unstructured":"Ciarelli, P.M., Oliveira, E.: Agglomeration and elimination of terms for dimensionality reduction. In: Proceedings of the 2009 Ninth International Conference on Intelligent Systems Design and Applications, ISDA \u201909, pp. 547\u2013552. IEEE Computer Society, Washington, DC, USA (2009)","DOI":"10.1109\/ISDA.2009.9"},{"key":"9989_CR4","doi-asserted-by":"crossref","unstructured":"Ciarelli, P.M., Salles, E.O.T., Oliveira, E.: An evolving system based on probabilistic neural network. In: 2010 Eleventh Brazilian Symposium on Neural Networks, pp. 182\u2013187 (2010)","DOI":"10.1109\/SBRN.2010.39"},{"issue":"1","key":"9989_CR5","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1111\/j.2517-6161.1977.tb01600.x","volume":"39","author":"AP Dempster","year":"1977","unstructured":"Dempster, A.P., Laird, N.M., Rubin, D.B.: Maximum likelihood from incomplete data via the EM algorithm. J. Royal Stat. Soc. Ser. B Methodol. 39(1), 1\u201338 (1977)","journal-title":"J. Royal Stat. Soc. Ser. B Methodol."},{"key":"9989_CR6","doi-asserted-by":"publisher","first-page":"611","DOI":"10.1198\/016214502760047131","volume":"97","author":"C Fraley","year":"2002","unstructured":"Fraley, C., Raftery, A.: Model-based clustering, discriminant analysis and density estimation. J. Am. Stat. Assoc. 97, 611\u2013631 (2002)","journal-title":"J. Am. Stat. Assoc."},{"issue":"5814","key":"9989_CR7","doi-asserted-by":"publisher","first-page":"972","DOI":"10.1126\/science.1136800","volume":"315","author":"BJ Frey","year":"2007","unstructured":"Frey, B.J., Dueck, D.: Clustering by passing messages between data points. Science 315(5814), 972\u2013976 (2007)","journal-title":"Science"},{"key":"9989_CR8","doi-asserted-by":"crossref","unstructured":"Greene, D., Cunningham, P.: Practical solutions to the problem of diagonal dominance in kernel document clustering. In: Proceedings of 23rd International Conference on Machine learning (ICML\u201906), pp. 377\u2013384. ACM Press (2006)","DOI":"10.1145\/1143844.1143892"},{"issue":"suppl 1","key":"9989_CR9","doi-asserted-by":"publisher","first-page":"5228","DOI":"10.1073\/pnas.0307752101","volume":"101","author":"TL Griffiths","year":"2004","unstructured":"Griffiths, T.L., Steyvers, M.: Finding scientific topics. Proceed. Natl. Acad. Sci. 101(suppl 1), 5228\u20135235 (2004). https:\/\/doi.org\/10.1073\/pnas.0307752101","journal-title":"Proceed. Natl. Acad. Sci."},{"issue":"1","key":"9989_CR10","doi-asserted-by":"publisher","first-page":"193","DOI":"10.1007\/BF01908075","volume":"2","author":"L Hubert","year":"1985","unstructured":"Hubert, L., Arabie, P.: Comparing partitions. J. Classification 2(1), 193\u2013218 (1985)","journal-title":"J. Classification"},{"key":"9989_CR11","volume-title":"Finding Groups in Data: An Introduction to Cluster Analysis","author":"L Kaufman","year":"2009","unstructured":"Kaufman, L., Rousseeuw, P.J.: Finding Groups in Data: An Introduction to Cluster Analysis, vol. 344. Wiley, Hoboken (2009)"},{"issue":"7553","key":"9989_CR12","doi-asserted-by":"publisher","first-page":"436","DOI":"10.1038\/nature14539","volume":"521","author":"Y LeCun","year":"2015","unstructured":"LeCun, Y., Bengio, Y., Hinton, G.: Deep learning. Nature 521(7553), 436\u2013444 (2015)","journal-title":"Nature"},{"issue":"3","key":"9989_CR13","doi-asserted-by":"publisher","first-page":"547","DOI":"10.1198\/106186005X59586","volume":"14","author":"J Li","year":"2005","unstructured":"Li, J.: Clustering based on a multilayer mixture model. J. Comput. Graph. Stat. 14(3), 547\u2013568 (2005). https:\/\/doi.org\/10.1198\/106186005X59586","journal-title":"J. Comput. Graph. Stat."},{"issue":"2","key":"9989_CR14","doi-asserted-by":"publisher","first-page":"129","DOI":"10.1109\/TIT.1982.1056489","volume":"28","author":"S Lloyd","year":"1982","unstructured":"Lloyd, S.: Least squares quantization in pcm. IEEE Trans. Inf. Theory 28(2), 129\u2013137 (1982)","journal-title":"IEEE Trans. Inf. Theory"},{"issue":"3","key":"9989_CR15","doi-asserted-by":"publisher","first-page":"379","DOI":"10.1016\/S0167-9473(02)00183-4","volume":"41","author":"G McLachlan","year":"2003","unstructured":"McLachlan, G., Peel, D., Bean, R.: Modelling high-dimensional data by mixtures of factor analyzers. Comput. Stat. Data Anal. 41(3), 379\u2013388 (2003)","journal-title":"Comput. Stat. Data Anal."},{"key":"9989_CR16","doi-asserted-by":"publisher","DOI":"10.1002\/0471721182","volume-title":"Finite Mixture Models","author":"GJ McLachlan","year":"2000","unstructured":"McLachlan, G.J., Peel, D.: Finite Mixture Models. Wiley, Hoboken (2000)"},{"issue":"4","key":"9989_CR17","doi-asserted-by":"publisher","first-page":"441","DOI":"10.1177\/1471082X0901000405","volume":"10","author":"A Montanari","year":"2010","unstructured":"Montanari, A., Viroli, C.: Heteroscedastic factor mixture analysis. Stat. Model. 10(4), 441\u2013460 (2010)","journal-title":"Stat. Model."},{"issue":"3","key":"9989_CR18","doi-asserted-by":"publisher","first-page":"274","DOI":"10.1007\/s00357-014-9161-z","volume":"31","author":"F Murtagh","year":"2014","unstructured":"Murtagh, F., Legendre, P.: Ward\u2019s hierarchical agglomerative clustering method: which algorithms implement ward\u2019s criterion? J. Classification 31(3), 274\u2013295 (2014)","journal-title":"J. Classification"},{"key":"9989_CR19","unstructured":"Ng, A.Y., Jordan, M.I., Weiss, Y.: On spectral clustering: analysis and an algorithm. In: Advances in Neural Information Processing Systems, pp. 849\u2013856 (2002)"},{"key":"9989_CR20","doi-asserted-by":"publisher","first-page":"103","DOI":"10.1023\/A:1007692713085","volume":"39","author":"K Nigam","year":"2000","unstructured":"Nigam, K., McCallum, A., Thrun, S., Mitchell, T.: Text classification from labeled and unlabeled documents using EM. Mach. Learn. 39, 103\u2013134 (2000)","journal-title":"Mach. Learn."},{"key":"9989_CR21","first-page":"3518","volume-title":"Advances in Neural Information Processing Systems 27","author":"A van den Oord","year":"2014","unstructured":"van den Oord, A., Schrauwen, B.: Factoring variations in natural images with deep Gaussian mixture models. In: Ghahramani, Z., Welling, M., Cortes, C., Lawrence, N.D., Weinberger, K.Q. (eds.) Advances in Neural Information Processing Systems 27, pp. 3518\u20133526. Curran Associates Inc, New York (2014)"},{"issue":"10","key":"9989_CR22","doi-asserted-by":"publisher","first-page":"5076","DOI":"10.1109\/TIP.2018.2848470","volume":"27","author":"X Peng","year":"2018","unstructured":"Peng, X., Feng, J., Xiao, S., Yau, W.Y., Zhou, J.T., Yang, S.: Structured autoencoders for subspace clustering. IEEE Trans. Image Process. 27(10), 5076\u20135086 (2018)","journal-title":"IEEE Trans. Image Process."},{"issue":"4","key":"9989_CR23","doi-asserted-by":"publisher","first-page":"1053","DOI":"10.1109\/TCYB.2016.2536752","volume":"47","author":"X Peng","year":"2016","unstructured":"Peng, X., Yu, Z., Yi, Z., Tang, H.: Constructing the l2-graph for robust subspace learning and subspace clustering. IEEE Trans. Cybern. 47(4), 1053\u20131066 (2016)","journal-title":"IEEE Trans. Cybern."},{"issue":"1","key":"9989_CR24","first-page":"4635","volume":"17","author":"S Romano","year":"2016","unstructured":"Romano, S., Vinh, N.X., Bailey, J., Verspoor, K.: Adjusting for chance clustering comparison measures. J. Mach. Learn. Res. 17(1), 4635\u20134666 (2016)","journal-title":"J. Mach. Learn. Res."},{"key":"9989_CR25","doi-asserted-by":"publisher","first-page":"85","DOI":"10.1016\/j.neunet.2014.09.003","volume":"61","author":"J Schmidhuber","year":"2015","unstructured":"Schmidhuber, J.: Deep learning in neural networks: an overview. Neural Netw. 61, 85\u2013117 (2015)","journal-title":"Neural Netw."},{"key":"9989_CR26","unstructured":"Tang, Y., Hinton, G.E., Salakhutdinov, R.: Deep mixtures of factor analysers. In: J.\u00a0Langford, J.\u00a0Pineau (eds.) Proceedings of the 29th International Conference on Machine Learning (ICML-12), pp. 505\u2013512. ACM, New York (2012)"},{"issue":"2","key":"9989_CR27","doi-asserted-by":"publisher","first-page":"137","DOI":"10.1016\/j.cmpb.2005.11.007","volume":"81","author":"A Tomovi\u0107","year":"2006","unstructured":"Tomovi\u0107, A., Jani\u010di\u0107, P., Ke\u0161elj, V.: n-gram-based classification and unsupervised hierarchical clustering of genome sequences. Comput. Methods Prog. Biomed. 81(2), 137\u2013153 (2006)","journal-title":"Comput. Methods Prog. Biomed."},{"issue":"3","key":"9989_CR28","doi-asserted-by":"publisher","first-page":"363","DOI":"10.1007\/s00357-010-9063-7","volume":"27","author":"C Viroli","year":"2010","unstructured":"Viroli, C.: Dimensionally reduced model-based clustering through mixtures of factor mixture analyzers. J. Classification 27(3), 363\u2013388 (2010)","journal-title":"J. Classification"},{"issue":"1","key":"9989_CR29","doi-asserted-by":"publisher","first-page":"43","DOI":"10.1007\/s11222-017-9793-z","volume":"29","author":"C Viroli","year":"2019","unstructured":"Viroli, C., McLachlan, G.J.: Deep Gaussian mixture models. Stat. Comput. 29(1), 43\u201351 (2019)","journal-title":"Stat. Comput."},{"issue":"1","key":"9989_CR30","doi-asserted-by":"crossref","first-page":"103","DOI":"10.1080\/07350015.1991.10509832","volume":"9","author":"J Wilson","year":"1991","unstructured":"Wilson, J., Koehler, K.: Hierarchical models for cross-classified overdispersed multinomial data. J. Business Econ. Stat. 9(1), 103\u2013110 (1991)","journal-title":"J. Business Econ. Stat."},{"issue":"12","key":"9989_CR31","doi-asserted-by":"publisher","first-page":"6191","DOI":"10.1109\/TNNLS.2018.2827036","volume":"29","author":"JT Zhou","year":"2018","unstructured":"Zhou, J.T., Zhao, H., Peng, X., Fang, M., Qin, Z., Goh, R.S.M.: Transfer hashing: from shallow to deep. IEEE Trans. Neural Netw. Learn. Syst. 29(12), 6191\u20136201 (2018)","journal-title":"IEEE Trans. Neural Netw. Learn. Syst."}],"container-title":["Statistics and Computing"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s11222-020-09989-9.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s11222-020-09989-9\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s11222-020-09989-9.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2024,8,25]],"date-time":"2024-08-25T06:23:29Z","timestamp":1724567009000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s11222-020-09989-9"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2021,3,3]]},"references-count":31,"journal-issue":{"issue":"3","published-print":{"date-parts":[[2021,5]]}},"alternative-id":["9989"],"URL":"https:\/\/doi.org\/10.1007\/s11222-020-09989-9","relation":{},"ISSN":["0960-3174","1573-1375"],"issn-type":[{"value":"0960-3174","type":"print"},{"value":"1573-1375","type":"electronic"}],"subject":[],"published":{"date-parts":[[2021,3,3]]},"assertion":[{"value":"1 June 2020","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"17 September 2020","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"3 March 2021","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}}],"article-number":"22"}}