{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,25]],"date-time":"2026-06-25T13:47:06Z","timestamp":1782395226299,"version":"3.54.5"},"reference-count":41,"publisher":"Springer Science and Business Media LLC","issue":"3","license":[{"start":{"date-parts":[[2026,6,25]],"date-time":"2026-06-25T00:00:00Z","timestamp":1782345600000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2026,6,25]],"date-time":"2026-06-25T00:00:00Z","timestamp":1782345600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"name":"Universidade Federal De Campina Grande"}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Software Qual J"],"published-print":{"date-parts":[[2026,9]]},"abstract":"<jats:title>Abstract<\/jats:title>\n                  <jats:p>In software development, bug reports (BRs) are essential for identifying defects, but the volume of reports in large projects makes manual relatedness analysis slow and error-prone. We study machine-learning approaches for predicting BR relatedness under a file-overlap target, File-Change Similarity (FCS). We compare TF-IDF, frozen sentence-level T5 embeddings (without domain fine-tuning), and a hybrid lexical-semantic representation. Our pipeline covers data retrieval, preprocessing, vectorization, normalization, neural-network training, and evaluation. We evaluated 56 models using various modeling strategies. Analysis reveals that using complete vectors as features is more effective than cosine distance. The hybrid approach shows competitive descriptive performance comparable to TF-IDF alone. Fine-tuning on 14 models tested 168 hyperparameter combinations, with Adam and RMSprop optimizers showing best performance. Key contributions include evaluating T5 and TF-IDF performance for BRs, exploring a hybrid approach, and providing a methodological framework for representation comparison. This research offers suggestions for improving efficiency in development and resource allocation. In the context of frozen T5 embeddings, the findings on T5 performance and the comparison with strong TF-IDF baselines drive future research directions. Since the T5 weights were not specifically trained on the bug report domain (frozen), these results serve as a baseline for future fine-tuning experiments.<\/jats:p>","DOI":"10.1007\/s11219-026-09768-1","type":"journal-article","created":{"date-parts":[[2026,6,25]],"date-time":"2026-06-25T12:53:10Z","timestamp":1782391990000},"update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":0,"title":["Lexical and semantic representations for similar bug report detection: A TF-IDF and frozen T5 comparative study"],"prefix":"10.1007","volume":"34","author":[{"given":"Iann","family":"Barbosa","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Jo\u00e3o","family":"Brunet","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Franklin","family":"Ramalho","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"297","published-online":{"date-parts":[[2026,6,25]]},"reference":[{"key":"9768_CR1","doi-asserted-by":"publisher","unstructured":"Bettenburg, N., Just, S., Schr\u00f6ter, A., Weiss, C., Premraj, R., & Zimmermann, T. (2008). What makes a good bug report? In Proceedings of the 16th ACM SIGSOFT International Symposium on Foundations of Software Engineering (FSE 2008) (pp. 308\u2013318). ACM, Atlanta, GA, USA. https:\/\/doi.org\/10.1145\/1453101.1453146","DOI":"10.1145\/1453101.1453146"},{"key":"9768_CR2","doi-asserted-by":"publisher","unstructured":"Brown, T.B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D.M., Wu, J., Winter, C., & Amodei, D. (2020). Language models are few-shot learners. In Advances in Neural Information Processing Systems 33 (NeurIPS 2020) (pp. 1877\u20131901). Curran Associates, Inc., Virtual Conference. https:\/\/doi.org\/10.48550\/arXiv.2005.14165","DOI":"10.48550\/arXiv.2005.14165"},{"key":"9768_CR3","doi-asserted-by":"publisher","unstructured":"Davis, J., & Goadrich, M. (2006). The relationship between precision-recall and ROC curves. In Proceedings of the 23rd International Conference on Machine Learning (ICML 2006) (pp. 233\u2013240). ACM, Pittsburgh, PA, USA. https:\/\/doi.org\/10.1145\/1143844.1143874","DOI":"10.1145\/1143844.1143874"},{"key":"9768_CR4","doi-asserted-by":"publisher","unstructured":"Devlin, J., Chang, M.-W., Lee, K., & Toutanova, K. (2019). BERT: Pre-training of deep bidirectional transformers for language understanding. Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (NAACL-HLT) (pp. 4171\u20134186). Minneapolis, MN, USA: Association for Computational Linguistics. https:\/\/doi.org\/10.18653\/v1\/N19-1423","DOI":"10.18653\/v1\/N19-1423"},{"issue":"3","key":"9768_CR5","doi-asserted-by":"publisher","first-page":"241","DOI":"10.1080\/00401706.1964.10490181","volume":"6","author":"OJ Dunn","year":"1964","unstructured":"Dunn, O. J. (1964). Multiple comparisons using rank sums. Technometrics, 6(3), 241\u2013252. https:\/\/doi.org\/10.1080\/00401706.1964.10490181","journal-title":"Technometrics"},{"issue":"8","key":"9768_CR6","doi-asserted-by":"publisher","first-page":"861","DOI":"10.1016\/j.patrec.2005.10.010","volume":"27","author":"T Fawcett","year":"2006","unstructured":"Fawcett, T. (2006). An introduction to ROC analysis. Pattern Recognition Letters, 27(8), 861\u2013874. https:\/\/doi.org\/10.1016\/j.patrec.2005.10.010","journal-title":"Pattern Recognition Letters"},{"issue":"2","key":"9768_CR7","doi-asserted-by":"publisher","first-page":"113","DOI":"10.3102\/10769986001002113","volume":"1","author":"PA Games","year":"1976","unstructured":"Games, P. A., & Howell, J. F. (1976). Pairwise multiple comparison procedures with unequal n\u2019s and\/or variances: A monte carlo study. Journal of Educational Statistics, 1(2), 113\u2013125. https:\/\/doi.org\/10.3102\/10769986001002113","journal-title":"Journal of Educational Statistics"},{"key":"9768_CR8","unstructured":"Goodfellow, I., Bengio, Y., & Courville, A. (2016). Deep Learning, (p. 800). MIT Press, Cambridge, MA, USA. https:\/\/www.deeplearningbook.org\/"},{"issue":"9","key":"9768_CR9","doi-asserted-by":"publisher","first-page":"1263","DOI":"10.1109\/TKDE.2008.239","volume":"21","author":"H He","year":"2009","unstructured":"He, H., & Garcia, E. A. (2009). Learning from imbalanced data. IEEE Transactions on Knowledge and Data Engineering, 21(9), 1263\u20131284. https:\/\/doi.org\/10.1109\/TKDE.2008.239","journal-title":"IEEE Transactions on Knowledge and Data Engineering"},{"issue":"1","key":"9768_CR10","doi-asserted-by":"publisher","first-page":"5","DOI":"10.1145\/963770.963772","volume":"22","author":"JL Herlocker","year":"2004","unstructured":"Herlocker, J. L., Konstan, J. A., Terveen, L. G., & Riedl, J. T. (2004). Evaluating collaborative filtering recommender systems. ACM Transactions on Information Systems, 22(1), 5\u201353. https:\/\/doi.org\/10.1145\/963770.963772","journal-title":"ACM Transactions on Information Systems"},{"key":"9768_CR11","doi-asserted-by":"publisher","unstructured":"IEEE Computer Society. (1990). IEEE standard glossary of software engineering terminology. Technical Report IEEE Std 610.12-1990, Institute of Electrical and Electronics Engineers. https:\/\/doi.org\/10.1109\/IEEESTD.1990.101064","DOI":"10.1109\/IEEESTD.1990.101064"},{"key":"9768_CR12","unstructured":"Jurafsky, D., & Martin, J. H. (2023). Speech and Language Processing, 3rd edn., (p. 650). Pearson, Boston, MA, USA. https:\/\/web.stanford.edu\/jurafsky\/slp3\/"},{"key":"9768_CR13","first-page":"83","volume":"4","author":"AN Kolmogorov","year":"1933","unstructured":"Kolmogorov, A. N. (1933). Sulla determinazione empirica di una legge di distribuzione. Giornale dell\u2019Istituto Italiano degli Attuari., 4, 83\u201391.","journal-title":"Giornale dell\u2019Istituto Italiano degli Attuari."},{"issue":"260","key":"9768_CR14","doi-asserted-by":"publisher","first-page":"583","DOI":"10.1080\/01621459.1952.10483441","volume":"47","author":"WH Kruskal","year":"1952","unstructured":"Kruskal, W. H., & Wallis, W. A. (1952). Use of ranks in one-criterion variance analysis. Journal of the American Statistical Association, 47(260), 583\u2013621. https:\/\/doi.org\/10.1080\/01621459.1952.10483441","journal-title":"Journal of the American Statistical Association"},{"key":"9768_CR15","doi-asserted-by":"publisher","unstructured":"Lazar, A., Ritchey, S., & Sharif, B. (2014). Improving the accuracy of duplicate bug report detection using textual similarity measures. In Proceedings of the 11th Working Conference on Mining Software Repositories (MSR 2014) (pp. 308\u2013311). ACM, Hyderabad, India. https:\/\/doi.org\/10.1145\/2597073.2597088","DOI":"10.1145\/2597073.2597088"},{"issue":"7553","key":"9768_CR16","doi-asserted-by":"publisher","first-page":"436","DOI":"10.1038\/nature14539","volume":"521","author":"Y LeCun","year":"2015","unstructured":"LeCun, Y., Bengio, Y., & Hinton, G. (2015). Deep learning. Nature, 521(7553), 436\u2013444. https:\/\/doi.org\/10.1038\/nature14539","journal-title":"Nature"},{"key":"9768_CR17","first-page":"278","volume-title":"Contributions to Probability and Statistics: Essays in Honor of Harold Hotelling","author":"H Levene","year":"1960","unstructured":"Levene, H. (1960). Robust tests for equality of variances. In I. Olkin, S. Ghurye, W. Hoeffding, W. Madow, & H. Mann (Eds.), Contributions to Probability and Statistics: Essays in Honor of Harold Hotelling (pp. 278\u2013292). Stanford, CA: Stanford University Press."},{"key":"9768_CR18","doi-asserted-by":"crossref","unstructured":"Manning, C. D., Raghavan, P., & Sch\u00fctze, H. (2008). Introduction to Information Retrieval, (p. 544). Cambridge University Press, Cambridge, UK. https:\/\/nlp.stanford.edu\/IR-book\/","DOI":"10.1017\/CBO9780511809071"},{"key":"9768_CR19","unstructured":"Mikolov, T., Chen, K., Corrado, G., & Dean, J. (2013). Efficient estimation of word representations in vector space. arxiv:1301.3781"},{"key":"9768_CR20","unstructured":"Mikolov, T., Sutskever, I., Chen, K., Corrado, G., & Dean, J. (2013). Distributed representations of words and phrases and their compositionality. arxiv:1310.4546"},{"key":"9768_CR21","doi-asserted-by":"crossref","unstructured":"Muennighoff, N., Tazi, N., Magne, L., & Reimers, N. (2022). MTEB: Massive text embedding benchmark. arxiv:2210.07316","DOI":"10.18653\/v1\/2023.eacl-main.148"},{"key":"9768_CR22","doi-asserted-by":"publisher","unstructured":"Nguyen, A. T., Nguyen, T. T., Nguyen, T. N., Lo, D., & Sun, C. (2012). Duplicate bug report detection with a combination of information retrieval and topic modeling. In 2012 Proceedings of the 27th IEEE\/ACM International Conference on Automated Software Engineering (pp. 70\u201379). https:\/\/doi.org\/10.1145\/2351676.2351687","DOI":"10.1145\/2351676.2351687"},{"key":"9768_CR23","doi-asserted-by":"publisher","unstructured":"Ni, J., \u00c1brego, G. H., Constant, N., Ma, J., Hall, K. B., Cer, D., & Yang, Y. (2021). Sentence-T5: Scalable sentence encoders from pre-trained text-to-text models. https:\/\/doi.org\/10.48550\/arXiv.2210.07316","DOI":"10.48550\/arXiv.2210.07316"},{"key":"9768_CR24","doi-asserted-by":"publisher","unstructured":"Pennington, J., Socher, R., & Manning, C. D. (2014). GloVe: Global vectors for word representation. In Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP) (pp. 1532\u20131543). Association for Computational Linguistics, Doha, Qatar. https:\/\/doi.org\/10.3115\/v1\/D14-1162","DOI":"10.3115\/v1\/D14-1162"},{"issue":"1","key":"9768_CR25","first-page":"37","volume":"2","author":"DMW Powers","year":"2011","unstructured":"Powers, D. M. W. (2011). Evaluation: from precision, recall and F-measure to ROC, informedness, markedness and correlation. Journal of Machine Learning Technologies, 2(1), 37\u201363.","journal-title":"Journal of Machine Learning Technologies"},{"key":"9768_CR26","doi-asserted-by":"publisher","unstructured":"Prechelt, L. (2012). Early stopping\u2014but when? In G. Montavon, G. B. Orr & K. -R. M\u00fcller (eds.), Neural Networks: Tricks of the Trade. Lecture Notes in Computer Science (vol. 7700, pp. 53\u201367). Springer, Berlin, Heidelberg. https:\/\/doi.org\/10.1007\/978-3-642-35289-8_5","DOI":"10.1007\/978-3-642-35289-8_5"},{"issue":"140","key":"9768_CR27","first-page":"1","volume":"21","author":"C Raffel","year":"2020","unstructured":"Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., & Liu, P. J. (2020). Exploring the limits of transfer learning with a unified text-to-text transformer. Journal of Machine Learning Research, 21(140), 1\u201367.","journal-title":"Journal of Machine Learning Research"},{"key":"9768_CR28","first-page":"1136","volume-title":"Artificial Intelligence: A Modern Approach","author":"SJ Russell","year":"2020","unstructured":"Russell, S. J., & Norvig, P. (2020). Artificial Intelligence: A Modern Approach (4th ed., p. 1136). Pearson: Boston, MA, USA.","edition":"4"},{"issue":"3","key":"9768_CR29","doi-asserted-by":"publisher","first-page":"0118432","DOI":"10.1371\/journal.pone.0118432","volume":"10","author":"T Saito","year":"2015","unstructured":"Saito, T., & Rehmsmeier, M. (2015). The precision-recall plot is more informative than the ROC plot when evaluating binary classifiers on imbalanced datasets. PLoS ONE, 10(3), 0118432. https:\/\/doi.org\/10.1371\/journal.pone.0118432","journal-title":"PLoS ONE"},{"key":"9768_CR30","first-page":"448","volume-title":"Introduction to Modern Information Retrieval","author":"G Salton","year":"1983","unstructured":"Salton, G., & McGill, M. J. (1983). Introduction to Modern Information Retrieval (p. 448). New York, NY, USA: McGraw-Hill."},{"issue":"11","key":"9768_CR31","doi-asserted-by":"publisher","first-page":"613","DOI":"10.1145\/361219.361220","volume":"18","author":"G Salton","year":"1975","unstructured":"Salton, G., Wong, A., & Yang, C.-S. (1975). A vector space model for automatic indexing. Communications of the ACM, 18(11), 613\u2013620. https:\/\/doi.org\/10.1145\/361219.361220","journal-title":"Communications of the ACM"},{"issue":"2","key":"9768_CR32","doi-asserted-by":"publisher","first-page":"11","DOI":"10.1109\/MS.2005.34","volume":"22","author":"N Serrano","year":"2005","unstructured":"Serrano, N., & Ciordia, I. (2005). Bugzilla, ITracker, and other bug trackers. IEEE Software, 22(2), 11\u201313. https:\/\/doi.org\/10.1109\/MS.2005.34","journal-title":"IEEE Software"},{"issue":"2","key":"9768_CR33","doi-asserted-by":"publisher","first-page":"279","DOI":"10.1214\/aoms\/1177730256","volume":"19","author":"NV Smirnov","year":"1948","unstructured":"Smirnov, N. V. (1948). Table for estimating the goodness of fit of empirical distributions. The Annals of Mathematical Statistics, 19(2), 279\u2013281. https:\/\/doi.org\/10.1214\/aoms\/1177730256","journal-title":"The Annals of Mathematical Statistics"},{"issue":"4","key":"9768_CR34","doi-asserted-by":"publisher","first-page":"427","DOI":"10.1016\/j.ipm.2009.03.002","volume":"45","author":"M Sokolova","year":"2009","unstructured":"Sokolova, M., & Lapalme, G. (2009). A systematic analysis of performance measures for classification tasks. Information Processing & Management, 45(4), 427\u2013437. https:\/\/doi.org\/10.1016\/j.ipm.2009.03.002","journal-title":"Information Processing & Management"},{"key":"9768_CR35","doi-asserted-by":"publisher","unstructured":"Sun, C., Lo, D., Wang, X., Jiang, J., & Khoo, S. -C. (2010). A discriminative model approach for accurate duplicate bug report retrieval. In 2010 ACM\/IEEE 32nd International Conference on Software Engineering (vol. 1, pp. 45\u201354). https:\/\/doi.org\/10.1145\/1806799.1806811","DOI":"10.1145\/1806799.1806811"},{"key":"9768_CR36","doi-asserted-by":"publisher","unstructured":"Tripathi, A., Dabral, S., & Sureka, A. (2015). On the use of version control system metadata for bug prediction. In Proceedings of the 19th International Conference on Evaluation and Assessment in Software Engineering (EASE 2015), (pp. 1\u201310). ACM, Nanjing, China. https:\/\/doi.org\/10.1145\/2745802.2745803","DOI":"10.1145\/2745802.2745803"},{"issue":"2","key":"9768_CR37","doi-asserted-by":"publisher","first-page":"99","DOI":"10.2307\/3001913","volume":"5","author":"JW Tukey","year":"1949","unstructured":"Tukey, J. W. (1949). Comparing individual means in the analysis of variance. Biometrics, 5(2), 99\u2013114. https:\/\/doi.org\/10.2307\/3001913","journal-title":"Biometrics"},{"key":"9768_CR38","doi-asserted-by":"publisher","first-page":"141","DOI":"10.1613\/jair.2934","volume":"37","author":"PD Turney","year":"2010","unstructured":"Turney, P. D., & Pantel, P. (2010). From frequency to meaning: Vector space models of semantics. Journal of Artificial Intelligence Research, 37, 141\u2013188. https:\/\/doi.org\/10.1613\/jair.2934","journal-title":"Journal of Artificial Intelligence Research"},{"key":"9768_CR39","doi-asserted-by":"publisher","unstructured":"Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L., & Polosukhin, I. (2017). Attention is all you need. In Advances in Neural Information Processing Systems 30 (NIPS 2017) (pp. 5998\u20136008). Curran Associates, Inc., Long Beach, CA, USA. https:\/\/doi.org\/10.48550\/arXiv.1706.03762","DOI":"10.48550\/arXiv.1706.03762"},{"key":"9768_CR40","doi-asserted-by":"publisher","unstructured":"Yang, X., Lo, D., Xia, X., Zhang, L., Sun, J. (2015). Deep learning for just-in-time defect prediction. In Proceedings of the IEEE International Conference on Software Quality, Reliability and Security (QRS 2015) (pp. 17\u201326). IEEE, Vancouver, BC, Canada. https:\/\/doi.org\/10.1109\/QRS.2015.14","DOI":"10.1109\/QRS.2015.14"},{"key":"9768_CR41","doi-asserted-by":"publisher","unstructured":"Zhang, Z., & Sabuncu, M. (2018). Generalized cross entropy loss for training deep neural networks with noisy labels. In Advances in Neural Information Processing Systems 31 (NeurIPS 2018) (pp. 8778\u20138788). Curran Associates, Inc., Montr\u00e9al, QC, Canada. https:\/\/doi.org\/10.48550\/arXiv.1805.07836","DOI":"10.48550\/arXiv.1805.07836"}],"container-title":["Software Quality Journal"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s11219-026-09768-1.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s11219-026-09768-1","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s11219-026-09768-1.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,6,25]],"date-time":"2026-06-25T12:53:19Z","timestamp":1782391999000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s11219-026-09768-1"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,6,25]]},"references-count":41,"journal-issue":{"issue":"3","published-print":{"date-parts":[[2026,9]]}},"alternative-id":["9768"],"URL":"https:\/\/doi.org\/10.1007\/s11219-026-09768-1","relation":{},"ISSN":["0963-9314","1573-1367"],"issn-type":[{"value":"0963-9314","type":"print"},{"value":"1573-1367","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,6,25]]},"assertion":[{"value":"7 December 2025","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"11 June 2026","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"25 June 2026","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"The authors declare that there is no conflict of interest.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Conflicts of interest"}},{"value":"The authors declare no competing interests.","order":3,"name":"Ethics","group":{"name":"EthicsHeading","label":"Competing interests"}}],"article-number":"34"}}