{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,31]],"date-time":"2026-07-31T05:08:53Z","timestamp":1785474533315,"version":"3.56.0"},"reference-count":76,"publisher":"Springer Science and Business Media LLC","issue":"8","license":[{"start":{"date-parts":[[2025,6,2]],"date-time":"2025-06-02T00:00:00Z","timestamp":1748822400000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2025,6,2]],"date-time":"2025-06-02T00:00:00Z","timestamp":1748822400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"DOI":"10.13039\/501100010624","name":"Karabuk University","doi-asserted-by":"crossref","id":[{"id":"10.13039\/501100010624","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["J Supercomput"],"abstract":"<jats:title>Abstract<\/jats:title>\n          <jats:p>Transformer-based Large Language Models (LLMs), which have recently gained popularity, have significantly impacted the various fields of natural language processing. One of them is text clustering, which involves categorizing the huge texts produced by today\u2019s digital world into meaningful groups. LLMs enable text clustering with a more semantic and contextualized approach than traditional methods. One such model is Sentence-BERT (SBERT), which has been modified to detect semantic similarity between texts. Before being clustered, texts need to be transformed into numerical text embeddings. SBERT-based models have shown promise in generating meaningful sentence embeddings. However, they face limitations when dealing with long texts that exceed their maximum token limit. In this context, this study proposes two distinct methods to overcome these limitations and enhance the performance of SBERT models for clustering long text. The proposed methods are combined with various SBERT models, and their combinations are compared to the existing default method on three datasets containing lengthy texts. This study evaluates the impact of these methods on the models and their contributions to clustering performance. The findings indicate that the proposed methods exhibit a higher clustering performance of up to 14% than the default method in text clustering. Additionally, this study provides valuable insights into the text clustering performance of SBERT models, offering practical implications for further research and applications.<\/jats:p>","DOI":"10.1007\/s11227-025-07414-4","type":"journal-article","created":{"date-parts":[[2025,6,2]],"date-time":"2025-06-02T08:39:36Z","timestamp":1748853576000},"update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":11,"title":["Optimizing SBERT for long text clustering: two novel approaches with empirical insights"],"prefix":"10.1007","volume":"81","author":[{"given":"Yasin","family":"Ortakci","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Burak","family":"Borhan","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"297","published-online":{"date-parts":[[2025,6,2]]},"reference":[{"key":"7414_CR1","doi-asserted-by":"publisher","first-page":"28","DOI":"10.1016\/j.procs.2016.04.005","volume":"82","author":"S Al-Anazi","year":"2016","unstructured":"Al-Anazi S, AlMahmoud H, Al-Turaiki I (2016) Finding similar documents using different clustering techniques. Procedia Comput Sci 82:28\u201334","journal-title":"Procedia Comput Sci"},{"key":"7414_CR2","doi-asserted-by":"crossref","unstructured":"Zhang Z, Fang M, Chen L, Namazi-Rad M-R (2022) Is neural topic modelling better than clustering? An empirical study on clustering with contextual embeddings for topics. arXiv:2204.09874","DOI":"10.18653\/v1\/2022.naacl-main.285"},{"key":"7414_CR3","doi-asserted-by":"crossref","unstructured":"Kamalloo E, Zhang X, Ogundepo O, Thakur N, Alfonso-Hermelo D, Rezagholizadeh M, Lin J (2023) Evaluating embedding APIs for information retrieval. arXiv:2305.06300","DOI":"10.18653\/v1\/2023.acl-industry.50"},{"key":"7414_CR4","doi-asserted-by":"publisher","first-page":"235","DOI":"10.1007\/s00521-016-2444-z","volume":"29","author":"X Xie","year":"2018","unstructured":"Xie X, Wang B (2018) Web page recommendation via twofold clustering: considering user behavior and topic relation. Neural Comput Appl 29:235\u2013243","journal-title":"Neural Comput Appl"},{"key":"7414_CR5","doi-asserted-by":"publisher","DOI":"10.1016\/j.eswa.2022.119308","volume":"215","author":"A Ghadimi","year":"2023","unstructured":"Ghadimi A, Beigy H (2023) SGCSumm: an extractive multi-document summarization method based on pre-trained language model, submodularity, and graph convolutional neural networks. Expert Syst Appl 215:119308","journal-title":"Expert Syst Appl"},{"issue":"1","key":"7414_CR6","first-page":"11","volume":"2","author":"L Huang","year":"2021","unstructured":"Huang L, Liu G, Chen T, Yuan H, Shi P, Miao Y (2021) Similarity-based emergency event detection in social media. J Saf Sci Resil 2(1):11\u201319","journal-title":"J Saf Sci Resil"},{"issue":"8","key":"7414_CR7","doi-asserted-by":"publisher","first-page":"5178","DOI":"10.3390\/app13085178","volume":"13","author":"A Moura","year":"2023","unstructured":"Moura A, Lima P, Mendon\u00e7a F, Mostafa SS, Morgado-Dias F (2023) On the use of transformer-based models for intent detection using clustering algorithms. Appl Sci 13(8):5178","journal-title":"Appl Sci"},{"key":"7414_CR8","doi-asserted-by":"crossref","unstructured":"Schluter N (2018) The word analogy testing caveat. In: Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies: Volume 2 (Short Papers). Association for Computational Linguistics, pp 242\u2013246","DOI":"10.18653\/v1\/N18-2039"},{"key":"7414_CR9","doi-asserted-by":"crossref","unstructured":"Yin B, Zhao M, Guo L, Qiao L (2023) Sentence-BERT and k-means based clustering technology for scientific and technical literature. In: 2023 15th International Conference on Computer Research and Development (ICCRD). IEEE, pp 15\u201320","DOI":"10.1109\/ICCRD56364.2023.10080830"},{"key":"7414_CR10","doi-asserted-by":"crossref","unstructured":"Pappagari R, Zelasko P, Villalba J, Carmiel Y, Dehak N (2019) Hierarchical transformers for long document classification. In: 2019 IEEE automatic speech recognition and understanding workshop (ASRU). IEEE, pp 838\u2013844","DOI":"10.1109\/ASRU46091.2019.9003958"},{"key":"7414_CR11","doi-asserted-by":"crossref","unstructured":"Selva Birunda S, Kanniga Devi R (2021) A review on word embedding techniques for text classification. In: Innovative data communication technologies and application: proceedings of ICIDCA 2020. Springer, pp 267\u2013281","DOI":"10.1007\/978-981-15-9651-3_23"},{"issue":"1","key":"7414_CR12","doi-asserted-by":"publisher","first-page":"34","DOI":"10.9790\/0661-16153438","volume":"16","author":"K Soumya George","year":"2014","unstructured":"Soumya George K, Joseph S (2014) Text classification by augmenting bag of words (BOW) representation with co-occurrence feature. IOSR J Comput Eng 16(1):34\u201338","journal-title":"IOSR J Comput Eng"},{"issue":"2","key":"7414_CR13","doi-asserted-by":"publisher","first-page":"88","DOI":"10.17706\/jcp.14.2.88-92","volume":"14","author":"I Mendon\u00e7a","year":"2019","unstructured":"Mendon\u00e7a I, Trouv\u00e9 A, Fukuda A, Murakami KJ, Tsai CF, Hu YH, Wang MC, Liu KE, Yu X, Yuan Y et al (2019) On clustering algorithms: applications in word-embedding documents. J Comput 14(2):88\u201392","journal-title":"J Comput"},{"key":"7414_CR14","doi-asserted-by":"crossref","unstructured":"Singh AK, Shashi M (2019) Vectorization of text documents for identifying unifiable news articles. Int J Adv Comput Sci Appl 10(7)","DOI":"10.14569\/IJACSA.2019.0100742"},{"key":"7414_CR15","doi-asserted-by":"crossref","unstructured":"Sundararaman D, Srinivasan S (2017) Twigraph: discovering and visualizing influential words between Twitter profiles. In: Social Informatics: 9th International Conference, SocInfo 2017, Oxford, UK, September 13\u201315, 2017, Proceedings, Part II 9. Springer, pp 329\u2013346","DOI":"10.1007\/978-3-319-67256-4_26"},{"issue":"4","key":"7414_CR16","doi-asserted-by":"publisher","DOI":"10.1088\/1742-6596\/1601\/4\/042007","volume":"1601","author":"J Tao","year":"2020","unstructured":"Tao J, Jia L, Wan MC, Meng JH (2020) The Text modeling method of Tibetan text combining Word2vec and improved TF-IDF. J Phys Conf Ser 1601(4):042007","journal-title":"J Phys Conf Ser"},{"key":"7414_CR17","doi-asserted-by":"crossref","unstructured":"Lilleberg J, Zhu Y, Zhang Y (2015) Support vector machines and word2vec for text classification with semantic features. In: 2015 IEEE 14th International Conference on Cognitive Informatics and Cognitive Computing (ICCI* CC). IEEE, pp 136\u2013140","DOI":"10.1109\/ICCI-CC.2015.7259377"},{"key":"7414_CR18","doi-asserted-by":"crossref","unstructured":"Omar A (2020) Feature selection in text clustering applications of literary texts: a hybrid of term weighting methods. Int J Adv Comput Sci Appl 11(2)","DOI":"10.14569\/IJACSA.2020.0110214"},{"issue":"6","key":"7414_CR19","doi-asserted-by":"publisher","first-page":"3105","DOI":"10.1016\/j.eswa.2014.11.038","volume":"42","author":"KK Bharti","year":"2015","unstructured":"Bharti KK, Singh PK (2015) Hybrid dimension reduction by integrating feature selection with feature extraction method for text lustering. Expert Syst Appl 42(6):3105\u20133114","journal-title":"Expert Syst Appl"},{"key":"7414_CR20","unstructured":"Mikolov T, Chen K, Corrado G, Dean J (2013) Efficient estimation of word representations in vector space. arXiv:1301.3781"},{"key":"7414_CR21","unstructured":"Mikolov T, Sutskever I, Chen K, Corrado GS, Dean J (2013) Distributed representations of words and phrases and their compositionality. In: Advances in neural information processing systems, 26"},{"key":"7414_CR22","doi-asserted-by":"publisher","first-page":"1091065","DOI":"10.3389\/fpsyg.2022.1091065","volume":"13","author":"C Yu","year":"2022","unstructured":"Yu C, Zhao Y (2022) Mining of risk perception dimensions of Chinese tourists\u2019 outbound tourism based on word vector method. Front Psychol 13:1091065","journal-title":"Front Psychol"},{"key":"7414_CR23","doi-asserted-by":"publisher","first-page":"34499","DOI":"10.1007\/s11042-019-08607-9","volume":"80","author":"J-W Baek","year":"2021","unstructured":"Baek J-W, Chung K-Y (2021) Multimedia recommendation using Word2Vec-based social relationship mining. Multimedia Tools Appl 80:34499\u201334515","journal-title":"Multimedia Tools Appl"},{"key":"7414_CR24","doi-asserted-by":"crossref","unstructured":"Pennington J, Socher R, Manning CD (2014) Glove: global vectors for word representation. In: Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP), pp 1532\u20131543","DOI":"10.3115\/v1\/D14-1162"},{"key":"7414_CR25","doi-asserted-by":"crossref","unstructured":"Haagsma H, Bjerva J (2016) Detecting novel metaphor using selectional preference information. In: Proceedings of the fourth workshop on metaphor in NLP, pp 10\u201317","DOI":"10.18653\/v1\/W16-1102"},{"key":"7414_CR26","doi-asserted-by":"crossref","unstructured":"Jo H, Cinarel C (2019) Delta-training: Simple semi-supervised text classification using pretrained word embeddings. arXiv:1901.07651","DOI":"10.18653\/v1\/D19-1347"},{"key":"7414_CR27","doi-asserted-by":"publisher","first-page":"135","DOI":"10.1162\/tacl_a_00051","volume":"5","author":"P Bojanowski","year":"2017","unstructured":"Bojanowski P, Grave E, Joulin A, Mikolov T (2017) Enriching word vectors with subword information. Trans Assoc Comput Linguist 5:135\u2013146","journal-title":"Trans Assoc Comput Linguist"},{"issue":"1","key":"7414_CR28","doi-asserted-by":"publisher","first-page":"1105","DOI":"10.11591\/ijece.v13i1.pp1105-1112","volume":"13","author":"I Ghozali","year":"2023","unstructured":"Ghozali I, Sungkono KR, Sarno R, Abdullah R (2023) Synonym based feature expansion for Indonesian hate speech detection. Int J Electr Comput Eng IJECE 13(1):1105","journal-title":"Int J Electr Comput Eng IJECE"},{"issue":"4","key":"7414_CR29","doi-asserted-by":"publisher","first-page":"5569","DOI":"10.1007\/s11042-022-13459-x","volume":"82","author":"M Umer","year":"2023","unstructured":"Umer M, Imtiaz Z, Ahmad M, Nappi M, Medaglia C, Choi GS, Mehmood A (2023) Impact of convolutional neural network and FastText embedding on text classification. Multimedia Tools Appl 82(4):5569\u20135585","journal-title":"Multimedia Tools Appl"},{"key":"7414_CR30","doi-asserted-by":"publisher","DOI":"10.1016\/j.infsof.2023.107202","volume":"159","author":"W Alhoshan","year":"2023","unstructured":"Alhoshan W, Ferrari A, Zhao L (2023) Zero-shot learning for requirements classification: An exploratory study. Inf Softw Technol 159:107202","journal-title":"Inf Softw Technol"},{"key":"7414_CR31","doi-asserted-by":"crossref","unstructured":"Howard J, Ruder S (2018) Universal language model fine-tuning for text classification. arXiv:1801.06146","DOI":"10.18653\/v1\/P18-1031"},{"key":"7414_CR32","doi-asserted-by":"publisher","DOI":"10.1016\/j.psychres.2021.114135","volume":"304","author":"J Sarzynska-Wawer","year":"2021","unstructured":"Sarzynska-Wawer J, Wawer A, Pawlak A, Szymanowska J, Stefaniak I, Jarkiewicz M, Okruszek L (2021) Detecting formal thought disorder by deep contextualized word representations. Psychiatry Res 304:114135","journal-title":"Psychiatry Res"},{"key":"7414_CR33","unstructured":"Vaswani A, Shazeer N, Parmar N, Uszkoreit J, Jones L, Gomez AN, Kaiser \u0141, Polosukhin I (2017) Attention is all you need. In: Advances in neural information processing systems, 30"},{"key":"7414_CR34","first-page":"1877","volume":"33","author":"T Brown","year":"2020","unstructured":"Brown T, Mann B, Ryder N, Subbiah M, Kaplan JD, Dhariwal P, Neelakantan A, Shyam P, Sastry G, Askell A et al (2020) Language models are few-shot learners. Adv Neural Inf Process Syst 33:1877\u20131901","journal-title":"Adv Neural Inf Process Syst"},{"key":"7414_CR35","unstructured":"Devlin J, Chang M-W, Lee K, Toutanova K (2018) Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv:1810.04805"},{"key":"7414_CR36","doi-asserted-by":"crossref","unstructured":"Ethayarajh K (2019) How contextual are contextualized word representations? Comparing the geometry of BERT, ELMo, and GPT-2 embeddings. arXiv:1909.00512","DOI":"10.18653\/v1\/D19-1006"},{"key":"7414_CR37","unstructured":"Muennighoff N (2022) Sgpt: Gpt sentence embeddings for semantic earch. arXiv:2202.08904"},{"issue":"1","key":"7414_CR38","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1186\/s40537-022-00564-9","volume":"9","author":"A Subakti","year":"2022","unstructured":"Subakti A, Murfi H, Hariadi N (2022) The performance of BERT as data representation of text clustering. J Big Data 9(1):1\u201321","journal-title":"J Big Data"},{"issue":"8","key":"7414_CR39","doi-asserted-by":"publisher","first-page":"5178","DOI":"10.3390\/app13085178","volume":"13","author":"A Moura","year":"2023","unstructured":"Moura A, Lima P, Mendon\u00e7a F, Mostafa SS, Morgado-Dias F (2023) On the use of transformer-based models for intent detection using clustering algorithms. Appl Sci 13(8):5178","journal-title":"Appl Sci"},{"key":"7414_CR40","doi-asserted-by":"publisher","first-page":"3211","DOI":"10.1007\/s40747-021-00512-9","volume":"7","author":"V Mehta","year":"2021","unstructured":"Mehta V, Bawa S, Singh J (2021) WEClustering: word embeddings based text clustering technique for large datasets. Complex Intell Syst 7:3211\u20133224","journal-title":"Complex Intell Syst"},{"issue":"8","key":"7414_CR41","doi-asserted-by":"publisher","first-page":"10861","DOI":"10.1007\/s11042-022-12155-0","volume":"81","author":"S Hosseini","year":"2022","unstructured":"Hosseini S, Varzaneh ZA (2022) Deep text clustering using stacked AutoEncoder. Multimedia Tools Appl 81(8):10861\u201310881","journal-title":"Multimedia Tools Appl"},{"key":"7414_CR42","doi-asserted-by":"crossref","unstructured":"Li Y, Cai J, Wang J (2020) A text document clustering method based on weighted Bert model. In: 2020 IEEE 4th Information Technology, Networking, Electronic and Automation Control Conference (ITNEC), vol 1. IEEE, pp 1426\u20131430","DOI":"10.1109\/ITNEC48623.2020.9085059"},{"key":"7414_CR43","doi-asserted-by":"crossref","unstructured":"Reimers N, Gurevych I (2019) Sentence-bert: sentence embeddings using siamese bert-networks. arXiv:1908.10084","DOI":"10.18653\/v1\/D19-1410"},{"key":"7414_CR44","doi-asserted-by":"publisher","first-page":"3479","DOI":"10.1016\/j.procs.2023.10.343","volume":"225","author":"C Sorana","year":"2023","unstructured":"Sorana C, Luc D, Julien B et al (2023) Application and evaluation of sentence embedding and clustering methods in the context of concept hierarchy construction. Procedia Comput Sci 225:3479\u20133487","journal-title":"Procedia Comput Sci"},{"key":"7414_CR45","doi-asserted-by":"publisher","DOI":"10.7717\/peerj-cs.2078","volume":"10","author":"K Abdalgader","year":"2024","unstructured":"Abdalgader K, Matroud AA, Hossin K (2024) Experimental study on short-text clustering using transformer-based semantic similarity measure. PeerJ Comput Sci 10:e2078","journal-title":"PeerJ Comput Sci"},{"key":"7414_CR46","doi-asserted-by":"crossref","unstructured":"Kiran N, Ragha L, Ghorpade T (2023) Summarizing students\u2019 text-only answer sheet using SBERT and K-means clustering and evaluating it using semantic search. In: International Conference on Artificial Intelligence and Knowledge Processing, pp 310\u2013323","DOI":"10.1007\/978-3-031-68617-7_23"},{"key":"7414_CR47","volume":"55","author":"Y Ortakci","year":"2024","unstructured":"Ortakci Y (2024) Revolutionary text clustering: investigating transfer learning capacity of SBERT models through pooling techniques. Int J Eng Sci Technol 55:101730","journal-title":"Int J Eng Sci Technol"},{"key":"7414_CR48","unstructured":"Hammami E, Faiz R (2022) Text clustering based on multi-view representations. In: CIRCLE\u201922: Conference of the Information Retrieval Communities in Europe, pp 3178"},{"issue":"21","key":"7414_CR49","doi-asserted-by":"publisher","first-page":"10893","DOI":"10.3390\/app122110893","volume":"12","author":"L Concei\u00e7\u00e3o","year":"2022","unstructured":"Concei\u00e7\u00e3o L, Rodrigues V, Meira J, Marreiros G, Novais P (2022) Supporting argumentation dialogues in group decision support systems: an approach based on dynamic clustering. Appl Sci 12(21):10893","journal-title":"Appl Sci"},{"key":"7414_CR50","doi-asserted-by":"publisher","DOI":"10.1016\/j.neucom.2023.126453","volume":"555","author":"Y Chu","year":"2023","unstructured":"Chu Y, Cao H, Diao Y, Lin H (2023) Refined sbert: representing sentence bert in manifold space. Neurocomputing 555:126453","journal-title":"Neurocomputing"},{"key":"7414_CR51","doi-asserted-by":"crossref","unstructured":"G\u00fcndo\u011fan E, Kaya M (2023) An Article Similarity-Based Approach for Planning Conference Sessions. In: 2023 13th International Conference on Advanced Computer Information Technologies (ACIT). IEEE, pp 591\u2013594","DOI":"10.1109\/ACIT58437.2023.10275336"},{"key":"7414_CR52","doi-asserted-by":"crossref","unstructured":"Yin B, Zhao M, Guo L, Qiao L (2023) Sentence-BERT and k-means based clustering technology for scientific and technical literature. In: 2023 15th International Conference on Computer Research and Development (ICCRD). IEEE, pp 15\u201320","DOI":"10.1109\/ICCRD56364.2023.10080830"},{"key":"7414_CR53","doi-asserted-by":"crossref","unstructured":"Christou L, Bompotas A, Makris C (2024) Document embeddings for long texts from transformers and autoencoders","DOI":"10.21203\/rs.3.rs-5459822\/v1"},{"key":"7414_CR54","doi-asserted-by":"crossref","unstructured":"Guedes GB, da Silva AEA (2024) Classification and clustering of sentence-level embeddings of scientific articles generated by contrastive learning. arXiv:2404.00224","DOI":"10.5121\/csit.2023.131923"},{"issue":"4","key":"7414_CR55","doi-asserted-by":"publisher","first-page":"599","DOI":"10.1007\/s00354-024-00244-7","volume":"42","author":"Z Wang","year":"2024","unstructured":"Wang Z, Zhu Y, Li Y, Qiang J, Yuan Y, Zhang C (2024) Asymmetric short-text clustering via prompt. N Gener Comput 42(4):599\u2013615","journal-title":"N Gener Comput"},{"key":"7414_CR56","unstructured":"Beltagy I, Peters ME, Cohan A (2020) Longformer: the long-document transformer. arXiv:2004.05150"},{"key":"7414_CR57","first-page":"17283","volume":"33","author":"M Zaheer","year":"2020","unstructured":"Zaheer M, Guruganesh G, Dubey KA, Ainslie J, Alberti C, Ontanon S, Pham P, Ravula A, Wang Q, Yang L et al (2020) Big bird: Transformers for longer sequences. J Adv Neural Inf Process Syst 33:17283\u201317297","journal-title":"J Adv Neural Inf Process Syst"},{"key":"7414_CR58","doi-asserted-by":"crossref","unstructured":"Zhu Y, Kiros R, Zemel R, Salakhutdinov R, Urtasun R, Torralba A, Fidler S (2015) Aligning books and movies: Towards story-like visual explanations by watching movies and reading books. In: Proceedings of the IEEE International Conference on Computer Vision, pp 19\u201327","DOI":"10.1109\/ICCV.2015.11"},{"key":"7414_CR59","doi-asserted-by":"crossref","unstructured":"d\u2019Sa AG, Illina I, Fohr D (2020) Bert and fasttext embeddings for automatic detection of toxic speech. In: 2020 International Multi-Conference on:\"Organization of Knowledge and Advanced Technologies\"(OCTA), pp 1\u20135. IEEE","DOI":"10.1109\/OCTA49274.2020.9151853"},{"key":"7414_CR60","doi-asserted-by":"crossref","unstructured":"Ye, Zhihao and Jiang, Gongyao and Liu, Ye and Li, Zhiyong and Yuan, Jin (2020) Document and word representations generated by graph convolutional network and bert for short text classification. In: 24th European Conference on Artificial Intelligence (ECAI 2020), pp 2275\u20132281","DOI":"10.3233\/FAIA200355"},{"key":"7414_CR61","unstructured":"Wu Y, Schuster M, Chen Z, Le QV, Norouzi M, Macherey W, Krikun M, Cao Y, Gao Q, Macherey K et\u00a0al (2016) Google\u2019s neural machine translation system: Bridging the gap between human and machine translation. arXiv:1609.08144"},{"key":"7414_CR62","unstructured":"Liu Y, Ott M, Goyal N, Du J, Joshi M, Chen D, Levy O, Lewis M, Zettlemoyer L, Stoyanov V (2019) Roberta: a robustly optimized bert pretraining approach. arXiv:1907.11692"},{"key":"7414_CR63","doi-asserted-by":"crossref","unstructured":"Sennrich R, Haddow B, Birch A (2015) Neural machine translation of rare words with subword units. arxiv:1508.0790","DOI":"10.18653\/v1\/P16-1162"},{"key":"7414_CR64","unstructured":"Thompson L, Mimno D(2020) Topic modeling with contextualized word representation clusters. arXiv:2010.12626"},{"key":"7414_CR65","unstructured":"Lan Z, Chen M, Goodman S, Gimpel K, Sharma P, Soricut R (2019) Albert: a lite bert for self-supervised learning of language representations. arXiv:1909.11942"},{"key":"7414_CR66","unstructured":"Sanh V, Debut L, Chaumond J, Wolf T (2019) DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter. arXiv:1910.01108"},{"key":"7414_CR67","first-page":"16857","volume":"33","author":"K Song","year":"2020","unstructured":"Song K, Tan X, Qin T, Lu J, Liu T-Y (2020) Mpnet: masked and permuted pre-training for language understanding. Adv Neural Inf Process Syst 33:16857\u201316867","journal-title":"Adv Neural Inf Process Syst"},{"key":"7414_CR68","unstructured":"Cer D, Yang Y, Kong S-y, Hua N, Limtiaco N, John RS, Constant N, Guajardo-Cespedes M, Yuan S, Tar C et\u00a0al (2018) Universal sentence encoder. arXiv:1803.11175"},{"key":"7414_CR69","doi-asserted-by":"crossref","unstructured":"Conneau A, Kiela D, Schwenk H, Barrault L, Bordes A (2017) Supervised learning of universal sentence representations from natural language inference data. arXiv:1705.02364","DOI":"10.18653\/v1\/D17-1070"},{"issue":"8","key":"7414_CR70","doi-asserted-by":"publisher","first-page":"651","DOI":"10.1016\/j.patrec.2009.09.011","volume":"31","author":"AK Jain","year":"2010","unstructured":"Jain AK (2010) Data clustering: 50 years beyond K-means. Pattern Recogn Lett 31(8):651\u2013666","journal-title":"Pattern Recogn Lett"},{"key":"7414_CR71","unstructured":"Rdusseeun LKPJ, Kaufman P (1987) Clustering by means of medoids. In: Proceedings of the Statistical Data Analysis Based on the L1 Norm Conference, Neuchatel, Switzerland, pp 31\u201328"},{"key":"7414_CR72","unstructured":"Zhang X, Zhao J, LeCun Y (2015) Character-level convolutional networks for text classification. In: Advances in neural information processing systems, 28"},{"issue":"1","key":"7414_CR73","doi-asserted-by":"publisher","DOI":"10.1002\/spy2.9","volume":"1","author":"H Ahmed","year":"2018","unstructured":"Ahmed H, Traore I, Saad S (2018) Detecting opinion spams and fake news using text classification. Secur Privacy 1(1):e9","journal-title":"Secur Privacy"},{"key":"7414_CR74","doi-asserted-by":"publisher","first-page":"22","DOI":"10.1016\/j.neunet.2016.12.008","volume":"88","author":"J Xu","year":"2017","unstructured":"Xu J, Xu B, Wang P, Zheng S, Tian G, Zhao J (2017) Self-taught convolutional neural networks for short text clustering. Neural Netw 88:22\u201331","journal-title":"Neural Netw"},{"key":"7414_CR75","doi-asserted-by":"publisher","first-page":"192","DOI":"10.1016\/j.eswa.2019.05.030","volume":"134","author":"R Janani","year":"2019","unstructured":"Janani R, Vijayarani S (2019) Text document clustering using spectral clustering algorithm with particle swarm optimization. Expert Syst Appl 134:192\u2013200","journal-title":"Expert Syst Appl"},{"issue":"8","key":"7414_CR76","doi-asserted-by":"publisher","first-page":"3669","DOI":"10.1109\/TKDE.2020.3028943","volume":"34","author":"R Guan","year":"2020","unstructured":"Guan R, Zhang H, Liang Y, Giunchiglia F, Huang L, Feng X (2020) Deep feature-based text clustering and its explanation. IEEE Trans Knowl Data Eng 34(8):3669\u20133680","journal-title":"IEEE Trans Knowl Data Eng"}],"container-title":["The Journal of Supercomputing"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s11227-025-07414-4.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s11227-025-07414-4\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s11227-025-07414-4.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,2]],"date-time":"2025-06-02T08:39:51Z","timestamp":1748853591000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s11227-025-07414-4"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,6,2]]},"references-count":76,"journal-issue":{"issue":"8","published-online":{"date-parts":[[2025,6]]}},"alternative-id":["7414"],"URL":"https:\/\/doi.org\/10.1007\/s11227-025-07414-4","relation":{},"ISSN":["1573-0484"],"issn-type":[{"value":"1573-0484","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,6,2]]},"assertion":[{"value":"6 May 2025","order":1,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"2 June 2025","order":2,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"The author states that there are no competing financial interests or personal relationships that could have appeared to influence the work reported in this paper.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Conflict of interest"}},{"value":"This study did not involve human or animal subjects, and thus, no ethical approval was required.","order":3,"name":"Ethics","group":{"name":"EthicsHeading","label":"Ethics approval"}}],"article-number":"950"}}