{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,13]],"date-time":"2026-05-13T18:13:56Z","timestamp":1778696036354,"version":"3.51.4"},"reference-count":62,"publisher":"Association for Computing Machinery (ACM)","issue":"6","license":[{"start":{"date-parts":[[2026,5,13]],"date-time":"2026-05-13T00:00:00Z","timestamp":1778630400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/legalcode"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Softw. Eng. Methodol."],"published-print":{"date-parts":[[2026,6,30]]},"abstract":"<jats:p>Machine learning-based methods are increasingly used to optimize build processes and accelerate the integration of software code. These methods leverage large volumes of historical code changes to train models on predicting and preventing issues in the codebase that could delay code integrations and features delivery to end-users. The objective of this study is to examine the impact of handling class noise present in software code changes collected from Continuous Integration (CI) systems on the predictive performance of machine learning models for predicting the execution outcome of CI builds and negative code reviews. In this study, we conduct a series of computational experiments using data from 110 Java open-source projects, examining the effectiveness of two removal-based statistical techniques - Majority Filter (MF) and Consensus Filter (CF) - and two corrective techniques - Domain Knowledge-based (DB) and CleanLab. Our results show that removal-based techniques significantly improve model predictive performance in both build outcome and negative code review prediction tasks. For build outcome prediction, applying MF increased the F1-score from 82% to 97%, and MCC from 0.13 to 0.58. In negative code review predictions, MF improved the F1-score from 17% to 53%, and MCC from \u22120.03 to 0.57. The DB technique was effective primarily in the context of code review comments but less so for build outcome predictions. While CleanLab yielded more consistent predictions, its overall impact on model performance was more moderate compared to removal-based techniques. Additionally, our findings show that hyperparameter tuning, applied independently or in combination with CleanLab, can further improve model performance; however, these gains did not surpass those achieved by removal-based techniques alone. We conclude that applying removal-based techniques to the training data of code changes is necessary to improve the prediction of build outcomes and negative code review comments.<\/jats:p>","DOI":"10.1145\/3764864","type":"journal-article","created":{"date-parts":[[2025,8,28]],"date-time":"2025-08-28T14:57:43Z","timestamp":1756393063000},"page":"1-66","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":0,"title":["The Impact of Class Noise-handling on the Effectiveness of Machine Learning-based Methods for Build Outcome and Code Change Request Predictions"],"prefix":"10.1145","volume":"35","author":[{"ORCID":"https:\/\/orcid.org\/0000-0003-2571-5099","authenticated-orcid":false,"given":"Khaled","family":"Al-Sabbagh","sequence":"first","affiliation":[{"name":"University of Gothenburg, Goteborg, Sweden"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-9052-0864","authenticated-orcid":false,"given":"Miroslaw","family":"Staron","sequence":"additional","affiliation":[{"name":"University of Gothenburg, Goteborg, Sweden"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-1459-2081","authenticated-orcid":false,"given":"Regina","family":"Hebig","sequence":"additional","affiliation":[{"name":"University of Rostock, Rostock, Germany"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2026,5,13]]},"reference":[{"key":"e_1_3_2_2_2","first-page":"106","volume-title":"Proceedings of the 2017 32nd IEEE\/ACM International Conference on Automated Software Engineering (ASE)","author":"Ahmed Toufique","year":"2017","unstructured":"Toufique Ahmed, Amiangshu Bosu, Anindya Iqbal, and Shahram Rahimi. 2017. SentiCR: A customized sentiment analysis tool for code review interactions. In Proceedings of the 2017 32nd IEEE\/ACM International Conference on Automated Software Engineering (ASE). IEEE, 106\u2013111."},{"key":"e_1_3_2_3_2","doi-asserted-by":"publisher","DOI":"10.1145\/3558489.3559070"},{"key":"e_1_3_2_4_2","first-page":"287","volume-title":"Proceedings of the International Conference on Product-Focused Software Process Improvement","author":"Al-Sabbagh Khaled Walid","year":"2020","unstructured":"Khaled Walid Al-Sabbagh, Regina Hebig, and Miroslaw Staron. 2020. The effect of class noise on continuous test case selection: A controlled experiment on industrial data. In Proceedings of the International Conference on Product-Focused Software Process Improvement. Springer, 287\u2013303."},{"key":"e_1_3_2_5_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.jss.2021.111093"},{"key":"e_1_3_2_6_2","doi-asserted-by":"crossref","first-page":"191","DOI":"10.1109\/SEAA51224.2020.00042","volume-title":"Proceedings of the 2020 46th Euromicro Conference on Software Engineering and Advanced Applications (SEAA)","author":"Al-Sabbagh Khaled Walid","year":"2020","unstructured":"Khaled Walid Al-Sabbagh, Miroslaw Staron, Regina Hebig, and Wilhelm Meding. 2020. Improving data quality for regression test selection by reducing annotation noise. In Proceedings of the 2020 46th Euromicro Conference on Software Engineering and Advanced Applications (SEAA). IEEE, 191\u2013194."},{"key":"e_1_3_2_7_2","unstructured":"Tanmay Basu. 2023. Identification of the relevance of comments in codes using bag of words and transformer based models. arXiv:2308.06144. Retrieved from https:\/\/arxiv.org\/abs\/2308.06144"},{"key":"e_1_3_2_8_2","unstructured":"C. Bent\u00e9jac A. Cs\u00f6rgo and G. Mart\u00ednez-Mu\u00f1oz. 1911. A comparative analysis of XGBoost. Retrieved from https:\/\/arxiv.org\/abs\/1911.01914"},{"key":"e_1_3_2_9_2","doi-asserted-by":"publisher","DOI":"10.5555\/2188385.2188395"},{"key":"e_1_3_2_10_2","doi-asserted-by":"publisher","DOI":"10.1613\/jair.606"},{"key":"e_1_3_2_11_2","doi-asserted-by":"publisher","DOI":"10.1007\/s11219-016-9342-6"},{"key":"e_1_3_2_12_2","doi-asserted-by":"publisher","DOI":"10.1186\/s12864-019-6413-7"},{"key":"e_1_3_2_13_2","unstructured":"Francois Chollet. 2015. Keras. Retrieved from https:\/\/github.com\/fchollet\/keras"},{"key":"e_1_3_2_14_2","unstructured":"Jacob Devlin Ming-Wei Chang Kenton Lee and Kristina Toutanova. 2018. BERT: Pre-training of deep bidirectional transformers for language understanding. arXiv:1810.04805. Retrieved from https:\/\/arxiv.org\/abs\/1810.04805"},{"key":"e_1_3_2_15_2","doi-asserted-by":"publisher","DOI":"10.3390\/s22124637"},{"key":"e_1_3_2_16_2","doi-asserted-by":"crossref","unstructured":"Zhangyin Feng Daya Guo Duyu Tang Nan Duan Xiaocheng Feng Ming Gong Linjun Shou Bing Qin Ting Liu Daxin Jiang et al. 2020. CodeBERT: A pre-trained model for programming and natural languages. arXiv:2002.08155. Retrieved from https:\/\/arxiv.org\/abs\/2002.08155","DOI":"10.18653\/v1\/2020.findings-emnlp.139"},{"key":"e_1_3_2_17_2","doi-asserted-by":"publisher","DOI":"10.1109\/32.815326"},{"key":"e_1_3_2_18_2","doi-asserted-by":"publisher","DOI":"10.1145\/3238147.3238171"},{"key":"e_1_3_2_19_2","first-page":"698","volume-title":"Proceedings of the 2019 34th IEEE\/ACM International Conference on Automated Software Engineering (ASE)","author":"Gong Lina","year":"2019","unstructured":"Lina Gong, Shujuan Jiang, Rongcun Wang, and Li Jiang. 2019. Empirical evaluation of the impact of class overlap on software defect prediction. In Proceedings of the 2019 34th IEEE\/ACM International Conference on Automated Software Engineering (ASE). IEEE, 698\u2013709."},{"key":"e_1_3_2_20_2","doi-asserted-by":"publisher","DOI":"10.1109\/TSE.2022.3220740"},{"key":"e_1_3_2_21_2","doi-asserted-by":"publisher","DOI":"10.1007\/s10489-010-0225-4"},{"key":"e_1_3_2_22_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.procs.2019.11.146"},{"key":"e_1_3_2_23_2","doi-asserted-by":"publisher","DOI":"10.1145\/2597073.2597118"},{"key":"e_1_3_2_24_2","doi-asserted-by":"publisher","DOI":"10.1145\/2372225.2372230"},{"key":"e_1_3_2_25_2","doi-asserted-by":"publisher","DOI":"10.5281\/zenodo.4445747"},{"key":"e_1_3_2_26_2","first-page":"531","volume-title":"Proceedings of the 2015 IEEE International Conference on Software Maintenance and Evolution (ICSME)","author":"Jongeling Robbert","year":"2015","unstructured":"Robbert Jongeling, Subhajit Datta, and Alexander Serebrenik. 2015. Choosing your weapons: On sentiment analysis tools for software engineering research. In Proceedings of the 2015 IEEE International Conference on Software Maintenance and Evolution (ICSME). IEEE, 531\u2013535."},{"key":"e_1_3_2_27_2","doi-asserted-by":"publisher","DOI":"10.1007\/s11390-007-9054-2"},{"key":"e_1_3_2_28_2","doi-asserted-by":"publisher","DOI":"10.5555\/1239069.1239072"},{"key":"e_1_3_2_29_2","first-page":"481","volume-title":"Proceedings of the 2011 33rd International Conference on Software Engineering (ICSE)","author":"Kim Sunghun","year":"2011","unstructured":"Sunghun Kim, Hongyu Zhang, Rongxin Wu, and Liang Gong. 2011. Dealing with noise in defect prediction. In Proceedings of the 2011 33rd International Conference on Software Engineering (ICSE). IEEE, 481\u2013490."},{"key":"e_1_3_2_30_2","unstructured":"Diederik P. Kingma and Jimmy Ba. 2014. Adam: A method for stochastic optimization. arXiv:1412.6980. Retrieved from https:\/\/arxiv.org\/abs\/1412.6980"},{"key":"e_1_3_2_31_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.infsof.2020.106257"},{"key":"e_1_3_2_32_2","first-page":"99","volume-title":"Proceedings of the 1st International Symposium on Empirical Software Engineering and Measurement (ESEM \u201907)","author":"Liebchen Gernot","year":"2007","unstructured":"Gernot Liebchen, Bheki Twala, Martin Shepperd, Michelle Cartwright, and Mark Stephens. 2007. Filtering, robust filtering, polishing: Techniques for addressing quality in software data. In Proceedings of the 1st International Symposium on Empirical Software Engineering and Measurement (ESEM \u201907). IEEE, 99\u2013106."},{"key":"e_1_3_2_33_2","volume-title":"Data Cleaning Techniques for Software Engineering Data Sets","author":"Liebchen Gernot Armin","year":"2010","unstructured":"Gernot Armin Liebchen. 2010. Data Cleaning Techniques for Software Engineering Data Sets. Ph.D. Dissertation. School of Information Systems, Computing and Mathematics, Brunel University."},{"key":"e_1_3_2_34_2","doi-asserted-by":"publisher","DOI":"10.1145\/3558489.3559073"},{"key":"e_1_3_2_35_2","doi-asserted-by":"publisher","DOI":"10.1023\/A:1025832930864"},{"key":"e_1_3_2_36_2","doi-asserted-by":"publisher","DOI":"10.1613\/jair.1.12125"},{"key":"e_1_3_2_37_2","doi-asserted-by":"crossref","first-page":"14","DOI":"10.1109\/MALTESQUE.2017.7882011","volume-title":"Proceedings of the 2017 IEEE Workshop on Machine Learning Techniques for Software Quality Evaluation (MaLTeSQuE)","author":"Ochodek Miroslaw","year":"2017","unstructured":"Miroslaw Ochodek, Miroslaw Staron, Dominik Bargowski, Wilhelm Meding, and Regina Hebig. 2017. Using machine learning to design a flexible LOC counter. In Proceedings of the 2017 IEEE Workshop on Machine Learning Techniques for Software Quality Evaluation (MaLTeSQuE). IEEE, 14\u201320."},{"key":"e_1_3_2_38_2","doi-asserted-by":"publisher","DOI":"10.3390\/app11114793"},{"key":"e_1_3_2_39_2","doi-asserted-by":"publisher","DOI":"10.5555\/1953048.2078195"},{"key":"e_1_3_2_40_2","doi-asserted-by":"publisher","DOI":"10.1109\/ACCESS.2019.2905133"},{"key":"e_1_3_2_41_2","doi-asserted-by":"publisher","DOI":"10.1002\/widm.1301"},{"key":"e_1_3_2_42_2","doi-asserted-by":"publisher","DOI":"10.5555\/3455716.3455856"},{"key":"e_1_3_2_43_2","first-page":"215","volume-title":"Proceedings of the 2017 IEEE\/ACM 14th International Conference on Mining Software Repositories (MSR)","author":"Rahman Mohammad Masudur","year":"2017","unstructured":"Mohammad Masudur Rahman, Chanchal K. Roy, and Raula G. Kula. 2017. Predicting usefulness of code review comments using textual features and developer experience. In Proceedings of the 2017 IEEE\/ACM 14th International Conference on Mining Software Repositories (MSR). IEEE, 215\u2013226."},{"key":"e_1_3_2_44_2","first-page":"43","volume-title":"Advances in Computers","author":"Rebours Pierre","year":"2006","unstructured":"Pierre Rebours and Taghi M. Khoshgoftaar. 2006. Quality problem in software measurement data. In Advances in Computers. Marvin V. Zelkowitz (Ed.), Vol. 66. Elsevier, 43\u201377."},{"issue":"2","key":"e_1_3_2_45_2","first-page":"1","article-title":"Sentiment analysis of free\/open source developers: Preliminary findings from a case study","volume":"13","author":"Rousinopoulos Athanasios-Ilias","year":"2014","unstructured":"Athanasios-Ilias Rousinopoulos, Gregorio Robles, and Jes\u00fas M. Gonz\u00e1lez-Barahona. 2014. Sentiment analysis of free\/open source developers: Preliminary findings from a case study. Revista Eletr\u00f4nica de Sistemas de Informa\u00e7\u00e3o 13, 2, Article 6 (2014), 1\u201321.","journal-title":"Revista Eletr\u00f4nica de Sistemas de Informa\u00e7\u00e3o"},{"key":"e_1_3_2_46_2","doi-asserted-by":"publisher","DOI":"10.1145\/3183519.3183525"},{"issue":"12","key":"e_1_3_2_47_2","doi-asserted-by":"crossref","first-page":"4873","DOI":"10.1109\/TSE.2021.3129165","article-title":"Detecting continuous integration skip commits using multi-objective evolutionary search","volume":"48","author":"Saidani Islem","year":"2021","unstructured":"Islem Saidani, Ali Ouni, and Mohamed Wiem Mkaouer. 2021. Detecting continuous integration skip commits using multi-objective evolutionary search. IEEE Transactions on Software Engineering 48, 12 (2021), 4873\u20134891.","journal-title":"IEEE Transactions on Software Engineering"},{"key":"e_1_3_2_48_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.ins.2010.12.016"},{"key":"e_1_3_2_49_2","first-page":"128","volume-title":"Proceedings of the 3rd International Conference on Information Systems, Technology and Management (ICISTM \u201909)","author":"Singh Yogesh","year":"2009","unstructured":"Yogesh Singh, Pradeep Kumar Bhatia, Arvinder Kaur, and Omprakash Sangwan. 2009. Application of neural networks in software engineering: A review. In Proceedings of the 3rd International Conference on Information Systems, Technology and Management (ICISTM \u201909). Springer, 128\u2013137."},{"key":"e_1_3_2_50_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.neucom.2014.10.086"},{"issue":"2","key":"e_1_3_2_51_2","doi-asserted-by":"crossref","first-page":"137","DOI":"10.35882\/jeeemi.v6i2.375","article-title":"Comparative study of various hyperparameter tuning on random Forest classification with SMOTE and feature selection using genetic algorithm in software defect prediction","volume":"6","author":"Suryadi Mulia Kevin","year":"2024","unstructured":"Mulia Kevin Suryadi, Rudy Herteno, Setyo Wahyu Saputro, Mohammad Reza Faisal, and Radityo Adi Nugroho. 2024. Comparative study of various hyperparameter tuning on random Forest classification with SMOTE and feature selection using genetic algorithm in software defect prediction. Journal of Electronics, Electromedical Engineering, and Medical Informatics 6, 2 (2024), 137\u2013147.","journal-title":"Journal of Electronics, Electromedical Engineering, and Medical Informatics"},{"key":"e_1_3_2_52_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICTAI.2004.93"},{"key":"e_1_3_2_53_2","first-page":"239","volume-title":"Proceedings of the 16th International Conference on Machine Learning (ICML \u201999)","author":"Teng Choh-Man","year":"1999","unstructured":"Choh-Man Teng. 1999. Correcting noisy data. In Proceedings of the 16th International Conference on Machine Learning (ICML \u201999). Citeseer, 239\u2013248."},{"key":"e_1_3_2_54_2","doi-asserted-by":"publisher","DOI":"10.1002\/asi.21416"},{"key":"e_1_3_2_55_2","first-page":"478","volume-title":"Proceedings of the 2006 IEEE International Conference on Information Reuse and Integration","author":"Van Hulse Jason","year":"2006","unstructured":"Jason Van Hulse, Taghi M. Khoshgoftaar, Chris Seiffert, and Lili Zhao. 2006. Noise correction using Bayesian multiple imputation. In Proceedings of the 2006 IEEE International Conference on Information Reuse and Integration. IEEE, 478\u2013483."},{"key":"e_1_3_2_56_2","doi-asserted-by":"publisher","DOI":"10.1145\/3475960.3475986"},{"key":"e_1_3_2_57_2","doi-asserted-by":"publisher","DOI":"10.1023\/A:1007626913721"},{"key":"e_1_3_2_58_2","doi-asserted-by":"publisher","DOI":"10.5555\/2349018"},{"key":"e_1_3_2_59_2","doi-asserted-by":"publisher","DOI":"10.1109\/QRS-C.2017.59"},{"key":"e_1_3_2_60_2","first-page":"234","volume-title":"Proceedings of the 2017 14th Web Information Systems and Applications Conference (WISA)","author":"Xia Jing","year":"2017","unstructured":"Jing Xia, Yanhui Li, and Chuanqi Wang. 2017. An empirical study on the cross-project predictability of continuous integration outcomes. In Proceedings of the 2017 14th Web Information Systems and Applications Conference (WISA). IEEE, 234\u2013239."},{"key":"e_1_3_2_61_2","doi-asserted-by":"publisher","DOI":"10.1109\/MIS.2004.1274907"},{"key":"e_1_3_2_62_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICSME52107.2021.00044"},{"key":"e_1_3_2_63_2","doi-asserted-by":"publisher","DOI":"10.1007\/s10462-004-0751-8"}],"container-title":["ACM Transactions on Software Engineering and Methodology"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3764864","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,5,13]],"date-time":"2026-05-13T17:26:41Z","timestamp":1778693201000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3764864"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,5,13]]},"references-count":62,"journal-issue":{"issue":"6","published-print":{"date-parts":[[2026,6,30]]}},"alternative-id":["10.1145\/3764864"],"URL":"https:\/\/doi.org\/10.1145\/3764864","relation":{},"ISSN":["1049-331X","1557-7392"],"issn-type":[{"value":"1049-331X","type":"print"},{"value":"1557-7392","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,5,13]]},"assertion":[{"value":"2023-06-17","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-08-05","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2026-05-13","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}