{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,8,25]],"date-time":"2026-08-25T07:11:32Z","timestamp":1787641892556,"version":"build-2736575974"},"reference-count":76,"publisher":"Springer Science and Business Media LLC","issue":"3","license":[{"start":{"date-parts":[[2022,1,10]],"date-time":"2022-01-10T00:00:00Z","timestamp":1641772800000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2022,1,10]],"date-time":"2022-01-10T00:00:00Z","timestamp":1641772800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"DOI":"10.13039\/501100001782","name":"University of Melbourne","doi-asserted-by":"crossref","id":[{"id":"10.13039\/501100001782","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Inf Retrieval J"],"published-print":{"date-parts":[[2022,9]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:p>Clustering of the contents of a document corpus is used to create sub-corpora with the intention that they are expected to consist of documents that are related to each other. However, while clustering is used in a variety of ways in document applications such as information retrieval, and a range of methods have been applied to the task, there has been relatively little exploration of how well it works in practice. Indeed, given the high dimensionality of the data it is possible that clustering may not always produce meaningful outcomes. In this paper we use a well-known clustering method to explore a variety of techniques, existing and novel, to measure clustering effectiveness. Results with our new, extrinsic techniques based on relevance judgements or retrieved documents demonstrate that retrieval-based information can be used to assess the quality of clustering, and also show that clustering can succeed to some extent at gathering together similar material. Further, they show that intrinsic clustering techniques that have been shown to be informative in other domains do not work for information retrieval. Whether clustering is sufficiently effective to have a significant impact on practical retrieval is unclear, but as the results show our measurement techniques can effectively distinguish between clustering methods.<\/jats:p>","DOI":"10.1007\/s10791-021-09401-8","type":"journal-article","created":{"date-parts":[[2022,1,10]],"date-time":"2022-01-10T11:03:23Z","timestamp":1641812603000},"page":"239-268","update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":13,"title":["Measurement of clustering effectiveness for document collections"],"prefix":"10.1007","volume":"25","author":[{"given":"Meng","family":"Yuan","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-6622-032X","authenticated-orcid":false,"given":"Justin","family":"Zobel","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Pauline","family":"Lin","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"297","published-online":{"date-parts":[[2022,1,10]]},"reference":[{"key":"9401_CR1","doi-asserted-by":"publisher","unstructured":"Abdelhaq, H., Sengstock, C., & Gertz, M. (2013) Eventweet: online localized event detection from twitter. In Proceedings of VLDB international conference on very large databases (vol\u00a06, pp. 1326\u20131329) https:\/\/doi.org\/10.14778\/2536274.2536307","DOI":"10.14778\/2536274.2536307"},{"key":"9401_CR2","doi-asserted-by":"publisher","unstructured":"Abraham, A., Das, S., & Konar, A. (2006) Document clustering using differential evolution. In IEEE international conference on evolutionary computation (pp. 1784\u20131791), https:\/\/doi.org\/10.1109\/CEC.2006.1688523","DOI":"10.1109\/CEC.2006.1688523"},{"key":"9401_CR3","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-10674-4","volume-title":"Feature selection and enhanced krill herd algorithm for text document clustering","author":"LMQ Abualigah","year":"2019","unstructured":"Abualigah, L. M. Q. (2019). Feature selection and enhanced krill herd algorithm for text document clustering. Springer."},{"issue":"1","key":"9401_CR4","doi-asserted-by":"publisher","first-page":"243","DOI":"10.1016\/j.patcog.2012.07.021","volume":"46","author":"O Arbelaitz","year":"2013","unstructured":"Arbelaitz, O., Gurrutxaga, I., Muguerza, J., P\u00e9rez, J. M., & Perona, I. (2013). An extensive comparative study of cluster validity indices. Pattern Recognition, 46(1), 243\u2013256. https:\/\/doi.org\/10.1016\/j.patcog.2012.07.021","journal-title":"Pattern Recognition"},{"key":"9401_CR5","doi-asserted-by":"publisher","unstructured":"Avrachenkov, K., Dobrynin, V., Nemirovsky, D., Pham, S.K., & Smirnova, E. (2008) Pagerank based clustering of hypertext document collections. In Proceedings of ACM-SIGIR international conference on research and development in information retrieval (pp. 873\u2013874) https:\/\/doi.org\/10.1145\/1390334.1390549","DOI":"10.1145\/1390334.1390549"},{"key":"9401_CR6","doi-asserted-by":"publisher","unstructured":"Becker, H., Naaman, M., & Gravano, L. (2010) Learning similarity metrics for event identification in social media. In Proceedings of ACM international conference on web search and data mining (pp. 291\u2013300) https:\/\/doi.org\/10.1145\/1718487.1718524","DOI":"10.1145\/1718487.1718524"},{"key":"9401_CR7","unstructured":"Ben-David, S., & Ackerman, M. (2008) Measures of clustering quality: a working set of axioms for clustering. In: Advances in neural information processing systems (vol\u00a021, pp. 121\u2013128)"},{"issue":"2","key":"9401_CR8","doi-asserted-by":"publisher","first-page":"191","DOI":"10.1016\/0098-3004(84)90020-7","volume":"10","author":"JC Bezdek","year":"1984","unstructured":"Bezdek, J. C., Ehrlich, R., & Full, W. (1984). Fcm: the fuzzy c-means clustering algorithm. Computers&amp; Geosciences, 10(2), 191\u2013203. https:\/\/doi.org\/10.1016\/0098-3004(84)90020-7","journal-title":"Computers & Geosciences"},{"issue":"6","key":"9401_CR9","doi-asserted-by":"publisher","first-page":"3105","DOI":"10.1016\/j.eswa.2014.11.038","volume":"42","author":"KK Bharti","year":"2015","unstructured":"Bharti, K. K., & Singh, P. K. (2015). Hybrid dimension reduction by integrating feature selection with feature extraction method for text clustering. Expert Systems with Applications, 42(6), 3105\u20133114. https:\/\/doi.org\/10.1016\/j.eswa.2014.11.038","journal-title":"Expert Systems with Applications"},{"issue":"4","key":"9401_CR10","doi-asserted-by":"publisher","first-page":"77","DOI":"10.1145\/2133806.2133826","volume":"55","author":"DM Blei","year":"2012","unstructured":"Blei, D. M. (2012). Probabilistic topic models. Communications of the ACM, 55(4), 77\u201384. https:\/\/doi.org\/10.1145\/2133806.2133826","journal-title":"Communications of the ACM"},{"key":"9401_CR11","doi-asserted-by":"publisher","unstructured":"Blott, S., & Weber, R. (2008) What\u2019s wrong with high-dimensional similarity search. In Proceedings of VLDB international conference on very large databases, Auckland, New Zealand, https:\/\/doi.org\/10.14778\/1453856.1453861","DOI":"10.14778\/1453856.1453861"},{"key":"9401_CR12","doi-asserted-by":"publisher","unstructured":"Bock, H. (2007) Clustering methods: a history of k-Means algorithms (pp. 161\u2013172) Springer Berlin Heidelberg, Berlin, Heidelberg https:\/\/doi.org\/10.1007\/978-3-540-73560-1_15","DOI":"10.1007\/978-3-540-73560-1_15"},{"key":"9401_CR13","doi-asserted-by":"publisher","unstructured":"Broder, A., Garcia-Pueyo, L., Josifovski, V., Vassilvitskii, S., & Venkatesan, S. (2014) Scalable k-means by ranked retrieval. In Proceedings of ACM international conference on web search and data mining, association for computing machinery (pp. 233\u2013242) New York, NY, USA, WSDM \u201914 https:\/\/doi.org\/10.1145\/2556195.2556260","DOI":"10.1145\/2556195.2556260"},{"issue":"6","key":"9401_CR14","doi-asserted-by":"publisher","first-page":"902","DOI":"10.1109\/TKDE.2010.165","volume":"23","author":"D Cai","year":"2010","unstructured":"Cai, D., He, X., & Han, J. (2010). Locally consistent concept factorization for document clustering. IEEE Transactions on Knowledge and Data Engineering, 23(6), 902\u2013913. https:\/\/doi.org\/10.1109\/TKDE.2010.165","journal-title":"IEEE Transactions on Knowledge and Data Engineering"},{"key":"9401_CR15","doi-asserted-by":"publisher","unstructured":"Callan, J.P., Lu, Z., & Croft, W.B. (1995) Searching distributed collections with inference networks. In Proceedings of ACM-SIGIR international conference on research and development in information retrieval, association for computing machinery (pp. 21\u201328) New York, NY, USA, SIGIR \u201995 https:\/\/doi.org\/10.1145\/215206.215328","DOI":"10.1145\/215206.215328"},{"key":"9401_CR16","doi-asserted-by":"publisher","unstructured":"Cleuziou, G. (2008) An extended version of the k-means method for overlapping clustering. In International conference on pattern recognition (pp. 1\u20134) https:\/\/doi.org\/10.1109\/ICPR.2008.4761079","DOI":"10.1109\/ICPR.2008.4761079"},{"key":"9401_CR17","unstructured":"Croft, W.B., Metzler, D., & Strohman, T. (2015) Search engines: information retrieval in practice. Originally published by Pearson"},{"key":"9401_CR18","doi-asserted-by":"publisher","unstructured":"Cutting, D.R., Karger, D.R., Pedersen, J.O., & Tukey, J.W. (1992) Scatter\/gather: A cluster-based approach to browsing large document collections. In Proceedings of ACM-SIGIR international conference on research and development in information retrieval, association for computing machinery (pp. 318\u2013329) New York, NY, USA, SIGIR \u201992, https:\/\/doi.org\/10.1145\/133160.133214","DOI":"10.1145\/133160.133214"},{"issue":"2","key":"9401_CR19","doi-asserted-by":"publisher","first-page":"224","DOI":"10.1109\/TPAMI.1979.4766909","volume":"1","author":"DL Davies","year":"1979","unstructured":"Davies, D. L., & Bouldin, D. W. (1979). A cluster separation measure. IEEE Transactions on Pattern Analysis and Machine Intelligence PAMI-, 1(2), 224\u2013227. https:\/\/doi.org\/10.1109\/TPAMI.1979.4766909","journal-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence PAMI-"},{"key":"9401_CR20","unstructured":"De Vries, C.M., Geva, S., & Trotman, A. (2012) Document clustering evaluation: divergence from a random baseline. https:\/\/arxiv.org\/abs\/1208.5654"},{"issue":"3","key":"9401_CR21","doi-asserted-by":"publisher","first-page":"32","DOI":"10.1080\/01969727308546046","volume":"3","author":"JC Dunn","year":"1973","unstructured":"Dunn, J. C. (1973). A fuzzy relative of the isodata process and its use in detecting compact well-separated clusters. Journal of Cybernetics, 3(3), 32\u201357. https:\/\/doi.org\/10.1080\/01969727308546046","journal-title":"Journal of Cybernetics"},{"key":"9401_CR22","doi-asserted-by":"publisher","unstructured":"Erman, J., Arlitt, M., & Mahanti, A. (2006) Traffic classification using clustering algorithms. In Proceedings of SIGCOMM workshop on mining network data (pp. 281\u2013286) https:\/\/doi.org\/10.1145\/1162678.1162679","DOI":"10.1145\/1162678.1162679"},{"key":"9401_CR23","unstructured":"Ester, M., Kriegel, H. P., Sander, J., & Xu, X. (1996). A density-based algorithm for discovering clusters in large spatial databases with noise. In Proceedings of international conference on knowledge discovery and data mining (pp. 226\u2013231) AAAI Press."},{"key":"9401_CR24","doi-asserted-by":"publisher","unstructured":"Evans, R., Pfahringer, B., & Holmes, G. (2011) Clustering for classification. In 2011 7th international conference on information technology in Asia (pp. 1\u20138) https:\/\/doi.org\/10.1109\/CITA.2011.5998839","DOI":"10.1109\/CITA.2011.5998839"},{"key":"9401_CR25","volume-title":"Cluster analysis","author":"BS Everitt","year":"2009","unstructured":"Everitt, B. S., Landau, S., & Leese, M. (2009). Cluster analysis (4th ed.). Wiley Publishing.","edition":"4"},{"key":"9401_CR26","first-page":"768","volume":"21","author":"EW Forgy","year":"1965","unstructured":"Forgy, E. W. (1965). Cluster analysis of multivariate data\u202f: efficiency versus interpretability of classifications. Biometrics, 21, 768\u2013769.","journal-title":"Biometrics"},{"issue":"2","key":"9401_CR27","doi-asserted-by":"publisher","first-page":"93","DOI":"10.1007\/s10791-011-9173-9","volume":"15","author":"N Fuhr","year":"2012","unstructured":"Fuhr, N., Lechtenfeld, M., Stein, B., & Gollub, T. (2012). The optimum clustering framework: implementing the cluster hypothesis. Information Retrieval, 15(2), 93\u2013115.","journal-title":"Information Retrieval"},{"key":"9401_CR28","doi-asserted-by":"publisher","unstructured":"Fung, B.C., Wang, K., & Ester, M. (2003) Hierarchical document clustering using frequent itemsets. In Proceedings of SIAM international conference on data mining (pp. 59\u201370) SIAM https:\/\/doi.org\/10.1137\/1.9781611972733.6","DOI":"10.1137\/1.9781611972733.6"},{"issue":"5","key":"9401_CR29","doi-asserted-by":"publisher","first-page":"613","DOI":"10.1002\/asi.20548","volume":"58","author":"D Hawking","year":"2007","unstructured":"Hawking, D., & Zobel, J. (2007). Does topic metadata help with web search? Journal of the American Society for Information Science and Technology, 58(5), 613\u2013628. https:\/\/doi.org\/10.1002\/asi.20548","journal-title":"Journal of the American Society for Information Science and Technology"},{"issue":"1","key":"9401_CR30","doi-asserted-by":"publisher","first-page":"258","DOI":"10.1016\/j.csda.2006.11.025","volume":"52","author":"C Hennig","year":"2007","unstructured":"Hennig, C. (2007). Cluster-wise assessment of cluster stability. Computational Statistics&amp; Data Analysis, 52(1), 258\u2013271.","journal-title":"Computational Statistics & Data Analysis"},{"key":"9401_CR31","unstructured":"Ingaramo, D., Rosso, P., & Errecalde, M. (2008) In: CICLing international conference on computational linguistics and intelligent text processing (pp. 555\u2013567), lNCS 4919"},{"key":"9401_CR32","unstructured":"Ingarmo, D., Errecal, M., Cagnina, L., & Rosso, P. (2009) Particle swarm optimization for clustering short-text corpora. In Proceedings of the conference on computational intelligence and bioengineering: essays in memory of Antonina Starita (pp. 3\u201319)"},{"issue":"8","key":"9401_CR33","doi-asserted-by":"publisher","first-page":"651","DOI":"10.1016\/j.patrec.2009.09.011","volume":"31","author":"AK Jain","year":"2010","unstructured":"Jain, A. K. (2010). Data clustering: 50 years beyond k-means. Pattern Recogn Lett, 31(8), 651\u2013666. https:\/\/doi.org\/10.1016\/j.patrec.2009.09.011","journal-title":"Pattern Recogn Lett"},{"issue":"5","key":"9401_CR34","doi-asserted-by":"publisher","first-page":"217","DOI":"10.1016\/0020-0271(71)90051-9","volume":"7","author":"N Jardine","year":"1971","unstructured":"Jardine, N., & van Rijsbergen, C. J. (1971). The use of hierarchic clustering in information retrieval. Information Storage and Retrieval, 7(5), 217\u2013240. https:\/\/doi.org\/10.1016\/0020-0271(71)90051-9","journal-title":"Information Storage and Retrieval"},{"issue":"3","key":"9401_CR35","doi-asserted-by":"publisher","first-page":"241","DOI":"10.1007\/bf02289588","volume":"32","author":"SC Johnson","year":"1967","unstructured":"Johnson, S. C. (1967). Hierarchical clustering schemes. Psychometrika, 32(3), 241\u2013254. https:\/\/doi.org\/10.1007\/bf02289588","journal-title":"Psychometrika"},{"key":"9401_CR36","doi-asserted-by":"publisher","unstructured":"Kulkarni, A., & Callan, J. (2010) Document allocation policies for selective searching of distributed indexes. In Proceedings of CIKM international conference on information and knowledge management, association for computing machinery (pp. 449\u2013458) New York, NY, USA, CIKM \u201910, https:\/\/doi.org\/10.1145\/1871437.1871497","DOI":"10.1145\/1871437.1871497"},{"key":"9401_CR37","doi-asserted-by":"publisher","unstructured":"Kulkarni, A., Tigelaar, A.S., Hiemstra, D., & Callan, J. (2012) Shard ranking and cutoff estimation for topically partitioned collections. In Proceedings of CIKM international conference on information and knowledge management, association for computing machinery (pp. 555\u2013564) New York, NY, USA, CIKM \u201912 https:\/\/doi.org\/10.1145\/2396761.2396833","DOI":"10.1145\/2396761.2396833"},{"key":"9401_CR38","doi-asserted-by":"publisher","unstructured":"Kummamuru, K., Dhawale, A., & Krishnapuram, R. (2003) Fuzzy co-clustering of documents and keywords. In IEEE international conference on fuzzy systems (Vol 2, pp. 772\u2013777) https:\/\/doi.org\/10.1109\/FUZZ.2003.1206527","DOI":"10.1109\/FUZZ.2003.1206527"},{"key":"9401_CR39","doi-asserted-by":"publisher","unstructured":"Larsen, B., & Aone, C. (1999) Fast and effective text mining using linear-time document clustering. In Proceedings of ACM SIGKDD international conference on knowledge discovery and data mining, association for computing machinery (pp. 16\u201322) New York, NY, USA, KDD \u201999, https:\/\/doi.org\/10.1145\/312129.312186","DOI":"10.1145\/312129.312186"},{"key":"9401_CR40","unstructured":"Le, Q., & Mikolov, T. (2014) Distributed representations of sentences and documents. In Proceedings of international conference on machine learning (pp. II\u20131188\u2013II\u20131196) JMLR.org, ICML\u201914"},{"key":"9401_CR41","doi-asserted-by":"publisher","unstructured":"Leuski, A. (2001) Evaluating document clustering for interactive information retrieval. In Proceedings of CIKM international conference on information and knowledge management (pp. 33\u201340), https:\/\/doi.org\/10.1145\/502585.502592","DOI":"10.1145\/502585.502592"},{"key":"9401_CR42","doi-asserted-by":"publisher","unstructured":"Li, C., Sun, A., & Datta, A. (2012) Twevent: Segment-based event detection from tweets. In Proceedings of CIKM international conference on information and knowledge management (pp. 155\u2013164) https:\/\/doi.org\/10.1145\/2396761.2396785","DOI":"10.1145\/2396761.2396785"},{"key":"9401_CR43","doi-asserted-by":"publisher","unstructured":"Liu, L., Kang, J., Yu, J., & Wang, Z. (2005) A comparative study on unsupervised feature selection methods for text clustering. In International conference on natural language processing and knowledge engineering (pp. 597\u2013601) IEEE https:\/\/doi.org\/10.1109\/NLPKE.2005.1598807","DOI":"10.1109\/NLPKE.2005.1598807"},{"key":"9401_CR44","doi-asserted-by":"publisher","unstructured":"Liu, X., & Croft, W.B. (2004) Cluster-based retrieval using language models. In Proceedings of ACM-SIGIR international conference on research and development in information retrieval, association for computing machinery (pp. 186\u2013193) New York, NY, USA, SIGIR \u201904 https:\/\/doi.org\/10.1145\/1008992.1009026","DOI":"10.1145\/1008992.1009026"},{"key":"9401_CR45","doi-asserted-by":"publisher","unstructured":"Liu, Y., Liu, Z., Chua, T., & Sun, M. (2015) Topical word embeddings. In Proceedings of AAAI conference on artificial intelligence (vol\u00a029) https:\/\/doi.org\/10.5555\/2886521.2886657","DOI":"10.5555\/2886521.2886657"},{"issue":"3","key":"9401_CR46","doi-asserted-by":"publisher","first-page":"496","DOI":"10.1007\/s10766-018-0591-9","volume":"48","author":"EL Lydia","year":"2018","unstructured":"Lydia, E. L., Kumar, P. K., Shankar, K., Lakshmanaprabu, S. K., Vidhyavathi, R. M., & Maseleno, A. (2018). Charismatic document clustering through novel k-means non-negative matrix factorization (KNMF) algorithm using key phrase extraction. International Journal of Parallel Programming, 48(3), 496\u2013514. https:\/\/doi.org\/10.1007\/s10766-018-0591-9","journal-title":"International Journal of Parallel Programming"},{"key":"9401_CR47","unstructured":"MacQueen, J.B. (1967) Some methods for classification and analysis of multivariate observations. In: Proceedings of the fifth berkeley symposium on mathematical statistics and probability (Volume 1: Statistics 1:281\u2013297)"},{"key":"9401_CR48","unstructured":"Mikolov, T., Chen, K., Corrado, G., & Dean, J. (2013) Efficient estimation of word representations in vector space. arXiv preprint arXiv:13013781"},{"issue":"3","key":"9401_CR49","doi-asserted-by":"publisher","first-page":"217","DOI":"10.1023\/A:1024016609528","volume":"52","author":"DS Modha","year":"2003","unstructured":"Modha, D. S., & Spangler, W. S. (2003). Feature weighting in k-means clustering. Machine Learning, 52(3), 217\u2013237. https:\/\/doi.org\/10.1023\/A:1024016609528","journal-title":"Machine Learning"},{"key":"9401_CR50","doi-asserted-by":"publisher","unstructured":"Pal, A., & Counts, S. (2011) Identifying topical authorities in microblogs. In Proceedings of ACM international conference on web search and data mining (pp. 45\u201354) https:\/\/doi.org\/10.1145\/1935826.1935843","DOI":"10.1145\/1935826.1935843"},{"key":"9401_CR51","doi-asserted-by":"publisher","unstructured":"Panuccio, A., Bicego, M., & Murino, V. (2002) A hidden Markov Model-based approach to sequential data clustering. In: Joint IAPR international workshops on statistical techniques in pattern recognition (SPR) and structural and syntactic pattern recognition (SSPR) (pp. 734\u2013743) Springer https:\/\/doi.org\/10.1007\/3-540-70659-3_77","DOI":"10.1007\/3-540-70659-3_77"},{"key":"9401_CR52","doi-asserted-by":"publisher","unstructured":"Pfeifer, D., & Leidner, J.L. (2019) Topic grouper: an agglomerative clustering approach to topic modeling. In Proceedings of ECIR european conference on IR research (pp. 590\u2013603) Springer https:\/\/doi.org\/10.1007\/978-3-030-15712-8_38","DOI":"10.1007\/978-3-030-15712-8_38"},{"key":"9401_CR53","doi-asserted-by":"publisher","unstructured":"Ramage, D., DHall, Nallapati, R., & Manning, C.D. (2009) Labeled LDA: a supervised topic model for credit attribution in multi-labeled corpora. In Proceedings of conference on empirical methods in natural language processing (pp. 248\u2013256) https:\/\/doi.org\/10.5555\/1699510.1699543","DOI":"10.5555\/1699510.1699543"},{"key":"9401_CR55","doi-asserted-by":"publisher","first-page":"53","DOI":"10.1016\/0377-0427(87)90125-7","volume":"20","author":"P Rousseeuw","year":"1987","unstructured":"Rousseeuw, P. (1987). Silhouettes: a graphical aid to the interpretation and validation of cluster analysis. Journal of Computational and Applied Mathematics, 20, 53\u201365. https:\/\/doi.org\/10.1016\/0377-0427(87)90125-7","journal-title":"Journal of Computational and Applied Mathematics"},{"key":"9401_CR56","doi-asserted-by":"publisher","unstructured":"Salton, G. (1971) The SMART retrieval system\u2014experiments in automatic document processing. Prentice-Hall, Inc., https:\/\/doi.org\/10.5555\/1102022","DOI":"10.5555\/1102022"},{"key":"9401_CR57","doi-asserted-by":"crossref","unstructured":"Sch\u00fctze, H. (1992) Dimensions of meaning. In Proceedings of the 1992 ACM\/IEEE conference on supercomputing (pp. 787\u2013796) IEEE Computer Society Press, Washington, DC, USA, Supercomputing \u201992","DOI":"10.1109\/SUPERC.1992.236684"},{"key":"9401_CR58","doi-asserted-by":"publisher","unstructured":"Shafiei, M., Wang, S., Zhang, R., Milios, E., Tang, B., Tougas, J., & Spiteri, R. (2007) Document representation and dimension reduction for text clustering. In Proceedings of IEEE international conference on data engineering (pp. 770\u2013779) IEEE https:\/\/doi.org\/10.1109\/ICDEW.2007.4401066","DOI":"10.1109\/ICDEW.2007.4401066"},{"key":"9401_CR59","doi-asserted-by":"crossref","unstructured":"Song, W., & Park, S.C. (2006). Genetic algorithm-based text clustering technique: Automatic evolution of clusters with high efficiency. In Proceedings of WAIMW seventh international conference on web-age information management workshops","DOI":"10.1109\/WAIMW.2006.14"},{"issue":"2","key":"9401_CR60","doi-asserted-by":"publisher","first-page":"50","DOI":"10.1145\/2207243.2207252","volume":"13","author":"N Spirin","year":"2012","unstructured":"Spirin, N., & Han, J. (2012). Survey on web spam detection: principles and algorithms. SIGKDD Explorations, 13(2), 50\u201364. https:\/\/doi.org\/10.1145\/2207243.2207252","journal-title":"SIGKDD Explorations"},{"issue":"4","key":"9401_CR61","doi-asserted-by":"publisher","first-page":"267","DOI":"10.1007\/BF02289263","volume":"18","author":"RL Thorndike","year":"1953","unstructured":"Thorndike, R. L. (1953). Who belongs in the family? Psychometrika, 18(4), 267\u2013276. https:\/\/doi.org\/10.1007\/BF02289263","journal-title":"Psychometrika"},{"key":"9401_CR62","doi-asserted-by":"publisher","unstructured":"Tomasini, C., Emmendorfer, L., Borges, E.N., & Machado, K. (2016) A methodology for selecting the most suitable cluster validation internal indices. In Proceedings of ACM symposium on applied computing (pp. 901\u2013903) https:\/\/doi.org\/10.1145\/2851613.2851885","DOI":"10.1145\/2851613.2851885"},{"key":"9401_CR63","doi-asserted-by":"publisher","unstructured":"Toma\u0161ev, N., & Radovanovi\u0107, M. (2016) Clustering evaluation in high-dimensional data (pp. 71\u2013107) Springer International Publishing https:\/\/doi.org\/10.1007\/978-3-319-24211-8_4","DOI":"10.1007\/978-3-319-24211-8_4"},{"issue":"4","key":"9401_CR64","doi-asserted-by":"publisher","first-page":"559","DOI":"10.1016\/S0306-4573(01)00048-6","volume":"38","author":"A Tombros","year":"2002","unstructured":"Tombros, A., Villa, R., & van Rijsbergen, C. J. (2002). The effectiveness of query-specific hierarchic clustering in information retrieval. Information Processing&amp; Management, 38(4), 559\u2013582. https:\/\/doi.org\/10.1016\/S0306-4573(01)00048-6","journal-title":"Information Processing & Management"},{"key":"9401_CR54","volume-title":"Information Retrieval","author":"CJ van Rijsbergen","year":"1979","unstructured":"van Rijsbergen, C. J. (1979). Information Retrieval. Butterworths."},{"issue":"6","key":"9401_CR65","doi-asserted-by":"publisher","first-page":"465","DOI":"10.1016\/0306-4573(86)90097-X","volume":"22","author":"EM Voorhees","year":"1986","unstructured":"Voorhees, E. M. (1986). Implementing agglomerative hierarchic clustering algorithms for use in document retrieval. Information Processing and Management, 22(6), 465\u2013476. https:\/\/doi.org\/10.1016\/0306-4573(86)90097-X.","journal-title":"Information Processing and Management"},{"key":"9401_CR66","unstructured":"Voorhees, E. M., & Harman, D. K. (Eds.). (2005). TREC experiment and evaluation in information retrieval. The MIT Press."},{"key":"9401_CR67","unstructured":"Weber, R., Schek, H., & Blott, S. (1998) A quantitative analysis and performance study for similarity-search methods in high-dimensional spaces. In Proceedings of VLDB international conferemce on very large databases (vol 98, pp. 194\u2013205)"},{"key":"9401_CR68","doi-asserted-by":"publisher","unstructured":"Wei, X., & Croft, W.B. (2006) LDA-based document models for ad-hoc retrieval. In Proceedings of ACM-SIGIR international conference on research and development in information retrieval (pp. 178\u2013185) https:\/\/doi.org\/10.1145\/1148170.1148204","DOI":"10.1145\/1148170.1148204"},{"issue":"5","key":"9401_CR69","doi-asserted-by":"publisher","first-page":"577","DOI":"10.1016\/0306-45738890027-1","volume":"24","author":"P Willett","year":"1988","unstructured":"Willett, P. (1988). Recent trends in hierarchic document clustering: a critical review. Information Processing&amp; Management, 24(5), 577\u2013597. https:\/\/doi.org\/10.1016\/0306-45738890027-1","journal-title":"Information Processing & Management"},{"key":"9401_CR70","doi-asserted-by":"publisher","unstructured":"Xu, J., & Croft, W.B. (1999) Cluster-based language models for distributed retrieval. In Proceedings of ACM-SIGIR international conference on research and development in information retrieval, association for computing machinery (pp. 254\u2013261) New York, NY, USA, SIGIR \u201999 https:\/\/doi.org\/10.1145\/312624.312687","DOI":"10.1145\/312624.312687"},{"key":"9401_CR71","doi-asserted-by":"publisher","unstructured":"Xu, W., Liu, X., & Gong, Y. (2003) Document clustering based on non-negative matrix factorization. In Proceedings of ACM-SIGIR international conference on research and development in information retrieval (pp. 267\u2013273) https:\/\/doi.org\/10.1145\/860435.860485","DOI":"10.1145\/860435.860485"},{"key":"9401_CR72","doi-asserted-by":"publisher","unstructured":"Yang, K., & Miao, R. (2018) Research on improvement of text processing and clustering algorithms in public opinion early warning system. In International conference on systems and informaticshttps:\/\/doi.org\/10.1109\/ICSAI.2018.8599424","DOI":"10.1109\/ICSAI.2018.8599424"},{"key":"9401_CR73","doi-asserted-by":"publisher","first-page":"152","DOI":"10.1016\/j.knosys.2014.11.028","volume":"75","author":"W Zhang","year":"2015","unstructured":"Zhang, W., Tang, X., & Yoshida, T. (2015). TESC: an approach to TExt classification using semi-supervised clustering. Knowledge-Based Systems, 75, 152\u2013160.","journal-title":"Knowledge-Based Systems"},{"key":"9401_CR74","doi-asserted-by":"publisher","unstructured":"Zobel, J. (1998) How reliable are the results of large-scale information retrieval experiments? In Proceedings of ACM-SIGIR international conference on research and development in information retrieval (pp. 307\u2013314) https:\/\/doi.org\/10.1145\/290941.291014","DOI":"10.1145\/290941.291014"},{"key":"9401_CR75","doi-asserted-by":"publisher","unstructured":"Zobel, J., & Moffat, A. (2006) Inverted files for text search engines. ACM Computing Surveys 38(2), 6\u2013es, https:\/\/doi.org\/10.1145\/1132956.1132959","DOI":"10.1145\/1132956.1132959"},{"issue":"1","key":"9401_CR76","doi-asserted-by":"publisher","first-page":"3","DOI":"10.1145\/1670598.1670600","volume":"43","author":"J Zobel","year":"2009","unstructured":"Zobel, J., Moffat, A., & Park, L. (2009). Against recall: is it persistence, cardinality, density, coverage, or totality? Proceedings of ACM-SIGIR International Conference on Research and Development in Information Retrieval, 43(1), 3\u20138. https:\/\/doi.org\/10.1145\/1670598.1670600","journal-title":"Proceedings of ACM-SIGIR International Conference on Research and Development in Information Retrieval"}],"container-title":["Information Retrieval Journal"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10791-021-09401-8.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s10791-021-09401-8\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10791-021-09401-8.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2022,9,23]],"date-time":"2022-09-23T20:40:38Z","timestamp":1663965638000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s10791-021-09401-8"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,1,10]]},"references-count":76,"journal-issue":{"issue":"3","published-print":{"date-parts":[[2022,9]]}},"alternative-id":["9401"],"URL":"https:\/\/doi.org\/10.1007\/s10791-021-09401-8","relation":{},"ISSN":["1386-4564","1573-7659"],"issn-type":[{"value":"1386-4564","type":"print"},{"value":"1573-7659","type":"electronic"}],"subject":[],"published":{"date-parts":[[2022,1,10]]},"assertion":[{"value":"7 June 2021","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"16 December 2021","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"10 January 2022","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}}]}}