{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,10,12]],"date-time":"2025-10-12T03:03:37Z","timestamp":1760238217484,"version":"build-2065373602"},"reference-count":40,"publisher":"MDPI AG","issue":"8","license":[{"start":{"date-parts":[[2022,8,16]],"date-time":"2022-08-16T00:00:00Z","timestamp":1660608000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Entropy"],"abstract":"<jats:p>We introduce a modern, optimized, and publicly available implementation of the sequential Information Bottleneck clustering algorithm, which strikes a highly competitive balance between clustering quality and speed. We describe a set of optimizations that make the algorithm computation more efficient, particularly for the common case of sparse data representation. The results are substantiated by an extensive evaluation that compares the algorithm to commonly used alternatives, focusing on the practically important use case of text clustering. The evaluation covers a range of publicly available benchmark datasets and a set of clustering setups employing modern word and sentence embeddings obtained by state-of-the-art neural models. The results show that in spite of using the more basic Term-Frequency representation, the proposed implementation provides a highly attractive trade-off between quality and speed that outperforms the alternatives considered. This new release facilitates the use of the algorithm in real-world applications of text clustering.<\/jats:p>","DOI":"10.3390\/e24081132","type":"journal-article","created":{"date-parts":[[2022,8,16]],"date-time":"2022-08-16T23:44:25Z","timestamp":1660693465000},"page":"1132","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":0,"title":["Revisiting Sequential Information Bottleneck: New Implementation and Evaluation"],"prefix":"10.3390","volume":"24","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-5309-267X","authenticated-orcid":false,"given":"Assaf","family":"Toledo","sequence":"first","affiliation":[{"name":"IBM Research AI, Haifa University Campus, Mount Carmel Haifa, Haifa 3498825, Israel"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-9785-3605","authenticated-orcid":false,"given":"Elad","family":"Venezian","sequence":"additional","affiliation":[{"name":"IBM Research AI, Haifa University Campus, Mount Carmel Haifa, Haifa 3498825, Israel"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-5171-8264","authenticated-orcid":false,"given":"Noam","family":"Slonim","sequence":"additional","affiliation":[{"name":"IBM Research AI, Haifa University Campus, Mount Carmel Haifa, Haifa 3498825, Israel"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2022,8,16]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","unstructured":"Aggarwal, C.C., and Zhai, C. (2012). A Survey of Text Clustering Algorithms. Mining Text Data, Springer US.","DOI":"10.1007\/978-1-4614-3223-4"},{"key":"ref_2","unstructured":"Huang, A. (2008, January 6\u20139). Similarity measures for text document clustering. Proceedings of the Sixth New Zealand Computer Science Research Student Conference (NZCSRSC2008), Christchurch, New Zealand."},{"key":"ref_3","doi-asserted-by":"crossref","unstructured":"Abualigah, L.M.Q. (2018). Feature Selection and Enhanced Krill Herd Algorithm for Text Document Clustering, Springer Publishing Company, Incorporated. [1st ed.].","DOI":"10.1007\/978-3-030-10674-4"},{"key":"ref_4","doi-asserted-by":"crossref","unstructured":"Sheng, Q.Z., Stroulia, E., Tata, S., and Bhiri, S. (2016, January 10\u201313). Clustering and Labeling IT Maintenance Tickets. Proceedings of the Service-Oriented Computing, Banff, AB, Canada.","DOI":"10.1007\/978-3-319-46295-0"},{"key":"ref_5","unstructured":"Compton, J.E., and Adams, M.C. (2020). Clustering Support Tickets with Natural Language Processing: K-Means Applied to Transformer Embeddings, Sandia National Lab. (SNL-NM)."},{"key":"ref_6","doi-asserted-by":"crossref","first-page":"170","DOI":"10.1504\/IJWBC.2015.068540","article-title":"A novel method for clustering tweets in Twitter","volume":"11","author":"Poomagal","year":"2015","journal-title":"Int. J. Web Based Commun."},{"key":"ref_7","unstructured":"Rosa, K.D., Shah, R., Lin, B., Gershman, A., and Frederking, R.E. (2011, January 28). Topical Clustering of Tweets. Proceedings of the ACM SIGIR: SWSM, Beijing, China."},{"key":"ref_8","doi-asserted-by":"crossref","first-page":"102034","DOI":"10.1016\/j.ipm.2019.04.002","article-title":"An evaluation of document clustering and topic modelling in two online social networks: Twitter and Reddit","volume":"57","author":"Curiskis","year":"2020","journal-title":"Inf. Process. Manag."},{"key":"ref_9","doi-asserted-by":"crossref","first-page":"129","DOI":"10.1109\/TIT.1982.1056489","article-title":"Least Squares Quantization in PCM","volume":"28","author":"Lloyd","year":"1982","journal-title":"IEEE Trans. Inf. Theor."},{"key":"ref_10","first-page":"2825","article-title":"Scikit-learn: Machine Learning in Python","volume":"12","author":"Pedregosa","year":"2011","journal-title":"J. Mach. Learn. Res."},{"key":"ref_11","unstructured":"Salton, G., and McGill, M.J. (1986). Introduction to Modern Information Retrieval, McGraw-Hill, Inc."},{"key":"ref_12","doi-asserted-by":"crossref","first-page":"34","DOI":"10.1126\/science.153.3731.34","article-title":"Dynamic programming","volume":"153","author":"Bellman","year":"1966","journal-title":"Science"},{"key":"ref_13","unstructured":"Mikolov, T., Chen, K., Corrado, G., and Dean, J. (2013). Efficient Estimation of Word Representations in Vector Space. arXiv."},{"key":"ref_14","first-page":"1421","article-title":"Distributed Representations of Words and Phrases and their Compositionality","volume":"26","author":"Mikolov","year":"2013","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_15","doi-asserted-by":"crossref","unstructured":"Pennington, J., Socher, R., and Manning, C.D. (2014, January 25\u201329). GloVe: Global Vectors for Word Representation. Proceedings of the Empirical Methods in Natural Language Processing (EMNLP), Doha, Qatar.","DOI":"10.3115\/v1\/D14-1162"},{"key":"ref_16","first-page":"1137","article-title":"A Neural Probabilistic Language Model","volume":"3","author":"Bengio","year":"2003","journal-title":"J. Mach. Learn. Res."},{"key":"ref_17","first-page":"3058","article-title":"Attention Is All You Need","volume":"30","author":"Vaswani","year":"2017","journal-title":"CoRR"},{"key":"ref_18","unstructured":"Devlin, J., Chang, M.W., Lee, K., and Toutanova, K. (2019, January 2\u20137). BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), Minneapolis, MN, USA."},{"key":"ref_19","doi-asserted-by":"crossref","unstructured":"Wang, A., Singh, A., Michael, J., Hill, F., Levy, O., and Bowman, S.R. (2018). GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding. arXiv.","DOI":"10.18653\/v1\/W18-5446"},{"key":"ref_20","first-page":"1828","article-title":"SuperGLUE: A Stickier Benchmark for General-Purpose Language Understanding Systems","volume":"32","author":"Wang","year":"2019","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_21","doi-asserted-by":"crossref","unstructured":"Reimers, N., and Gurevych, I. (2019). Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks. arXiv.","DOI":"10.18653\/v1\/D19-1410"},{"key":"ref_22","doi-asserted-by":"crossref","unstructured":"Slonim, N., Friedman, N., and Tishby, N. (2002, January 11\u201315). Unsupervised Document Classification Using Sequential Information Maximization. Proceedings of the 25th Annual International ACM SIGIR Conference on Research and Development in Information Retrieval, SIGIR\u201902, Tampere, Finland.","DOI":"10.1145\/564400.564401"},{"key":"ref_23","unstructured":"Lang, K. (1995, January 9\u201312). NewsWeeder: Learning to Filter Netnews. Proceedings of the 12th International Machine Learning Conference (ML95), Tahoe City, CA, USA."},{"key":"ref_24","doi-asserted-by":"crossref","unstructured":"Connor, R.C.H., Cardillo, F.A., Moss, R., and Rabitti, F. (2013, January 2\u20134). Evaluation of Jensen-Shannon Distance over Sparse Data. Proceedings of the SISAP, A Coruna, Spain.","DOI":"10.1007\/978-3-642-41062-8_16"},{"key":"ref_25","unstructured":"Tishby, N., Pereira, F.C., and Bialek, W. (2000). The information bottleneck method. arXiv."},{"key":"ref_26","unstructured":"Slonim, N. (2002). The Information Bottleneck: Theory and Applications. [Ph.D. Thesis, Hebrew University of Jerusalem]."},{"key":"ref_27","unstructured":"Cover, T.M., and Thomas, J.A. (2006). Elements of Information Theory (Wiley Series in Telecommunications and Signal Processing), Wiley-Interscience."},{"key":"ref_28","unstructured":"Solla, S., Leen, T., and M\u00fcller, K. (December, January 29). Agglomerative Information Bottleneck. Proceedings of the Advances in Neural Information Processing Systems, Denver, CO, USA."},{"key":"ref_29","unstructured":"Zhang, J.A., and Kurkoski, B.M. (November, January 3). Low-complexity quantization of discrete memoryless channels. Proceedings of the 2016 International Symposium on Information Theory and Its Applications (ISITA), Monterey, CA, USA."},{"key":"ref_30","doi-asserted-by":"crossref","first-page":"340","DOI":"10.1109\/34.88569","article-title":"Optimal partitioning for classification and regression trees","volume":"13","author":"Chou","year":"1991","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_31","first-page":"1705","article-title":"Clustering with Bregman Divergences","volume":"6","author":"Banerjee","year":"2005","journal-title":"J. Mach. Learn. Res."},{"key":"ref_32","unstructured":"Kurkoski, B.M. (2017, January 6\u20139). On the relationship between the KL means algorithm and the information bottleneck method. Proceedings of the SCC 2017, 11th International ITG Conference on Systems, Communications and Coding, Hamburg, Germany."},{"key":"ref_33","unstructured":"Slonim, N., Aharoni, E., and Crammer, K. (2013, January 3\u20139). Hartigan\u2019s K-Means versus Lloyd\u2019s K-Means: Is It Time for a Change?. Proceedings of the Twenty-Third International Joint Conference on Artificial Intelligence, IJCAI\u201913, Beijing, China."},{"key":"ref_34","doi-asserted-by":"crossref","unstructured":"Greene, D., and Cunningham, P. (2006, January 25\u201329). Practical Solutions to the Problem of Diagonal Dominance in Kernel Document Clustering. Proceedings of the 23rd International Conference on Machine Learning (ICML\u201906), New York, NY, USA.","DOI":"10.1145\/1143844.1143892"},{"key":"ref_35","first-page":"456","article-title":"Character-level Convolutional Networks for Text Classification","volume":"28","author":"Zhang","year":"2015","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_36","first-page":"2837","article-title":"Information Theoretic Measures for Clusterings Comparison: Variants, Properties, Normalization and Correction for Chance","volume":"11","author":"Vinh","year":"2010","journal-title":"J. Mach. Learn. Res."},{"key":"ref_37","doi-asserted-by":"crossref","first-page":"873","DOI":"10.1016\/j.jmva.2006.11.013","article-title":"Comparing clusterings\u2014An information based distance","volume":"98","year":"2007","journal-title":"J. Multivar. Anal."},{"key":"ref_38","unstructured":"Rosenberg, A., and Hirschberg, J. (2007, January 28\u201330). V-Measure: A Conditional Entropy-Based External Cluster Evaluation Measure. Proceedings of the 2007 Joint Conference on Empirical Methods in Natural Language Processing and Computational Natural Language Learning (EMNLP-CoNLL), Prague, Czech Republic."},{"key":"ref_39","unstructured":"Arthur, D., and Vassilvitskii, S. (2007, January 7\u20139). K-Means++: The Advantages of Careful Seeding. Proceedings of the Eighteenth Annual ACM-SIAM Symposium on Discrete Algorithms, New Orleans, Louisiana."},{"key":"ref_40","doi-asserted-by":"crossref","unstructured":"Spencer, N. (2013). Essentials of Multivariate Data Analysis, Taylor & Francis.","DOI":"10.1201\/b16344"}],"container-title":["Entropy"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1099-4300\/24\/8\/1132\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,11]],"date-time":"2025-10-11T00:09:21Z","timestamp":1760141361000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1099-4300\/24\/8\/1132"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,8,16]]},"references-count":40,"journal-issue":{"issue":"8","published-online":{"date-parts":[[2022,8]]}},"alternative-id":["e24081132"],"URL":"https:\/\/doi.org\/10.3390\/e24081132","relation":{},"ISSN":["1099-4300"],"issn-type":[{"type":"electronic","value":"1099-4300"}],"subject":[],"published":{"date-parts":[[2022,8,16]]}}}