{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,3,13]],"date-time":"2026-03-13T02:34:11Z","timestamp":1773369251981,"version":"3.50.1"},"reference-count":66,"publisher":"Association for Computing Machinery (ACM)","issue":"4","funder":[{"name":"Nanjing University\u2014China Mobile Communications Group Co., Ltd"},{"name":"National Key Research and Development Program of China","award":["2022YFF0711404"],"award-info":[{"award-number":["2022YFF0711404"]}]},{"name":"Open Project of National Key Laboratory for Novel Software Technology","award":["ZZKT2024A08, KFKT2024B21, and KFKT2024A11"],"award-info":[{"award-number":["ZZKT2024A08, KFKT2024B21, and KFKT2024A11"]}]},{"name":"CCF-Huawei Populus Grove Fund"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Softw. Eng. Methodol."],"published-print":{"date-parts":[[2026,4,30]]},"abstract":"<jats:p>\n                    Self-Admitted Technical Debt (SATD) refers to sub-optimal solutions deliberately introduced to accelerate the software development process, often at the expense of software maintainability and sustainability. Therefore, timely identification and repayment of the SATD is critical for the software system. As exploration deepens, it is found that effectively prioritizing the repayment of SATD with more significant impacts on software quality requires not only identifying SATD but also further classifying it. However, existing SATD identification and classification approaches face the following challenges: (1) SATDs originate from diverse sources. Code comments are a widespread source, but recent research has revealed that SATDs can originate from other sources, such as pull requests, issues, and commit messages. Nonetheless, existing approaches primarily target code comments, lacking the capability to analyze SATDs from other sources effectively. (2) SATDs fall into diverse categories. Nonetheless, existing SATD classification approaches fail to address all SATD categories comprehensively and show inadequate performance. (3) Imbalance of existing SATD datasets. Real-world SATD data are scarce, making dataset collection challenging. Moreover, SATD distribution across different sources is uneven, further complicating the construction of high-quality datasets. To alleviate these challenges, this article presents an SATD identification and classification framework named\n                    <jats:italic toggle=\"yes\">IMPACT<\/jats:italic>\n                    . First, IMPACT employs ChatGPT to construct an augmented dataset. Subsequently, it utilizes a pipeline with two fine-tuned language models of different parameter sizes to identify and classify SATD separately. To evaluate the effectiveness of IMPACT, we compare it with three state-of-the-art SATD classification methods and its two foundation models. Experimental results demonstrate that IMPACT outperforms state-of-the-art methods by a large margin, and even surpasses its foundation model GLM-4-9B-Chat. It achieves the optimal average F1 score of 0.697 on the source of pull requests, the most challenging data source. Moreover, experiments on the cross-project test set show that IMPACT demonstrates strong generalizability on unseen project data.\n                  <\/jats:p>","DOI":"10.1145\/3747180","type":"journal-article","created":{"date-parts":[[2025,7,10]],"date-time":"2025-07-10T13:53:24Z","timestamp":1752155604000},"page":"1-36","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":0,"title":["IMPACT: Identifying and Classifying Multiple Sourced and Categorized Self-Admitted Technical Debts"],"prefix":"10.1145","volume":"35","author":[{"ORCID":"https:\/\/orcid.org\/0009-0004-3975-6734","authenticated-orcid":false,"given":"Qingyuan","family":"Li","sequence":"first","affiliation":[{"name":"National Key Laboratory for Novel Software Technology, Nanjing University, Nanjing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-0974-4685","authenticated-orcid":false,"given":"Zhixin","family":"Yin","sequence":"additional","affiliation":[{"name":"National Key Laboratory for Novel Software Technology, Nanjing University, Nanjing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0001-7550-0320","authenticated-orcid":false,"given":"Yaopeng","family":"Yang","sequence":"additional","affiliation":[{"name":"National Key Laboratory for Novel Software Technology, Nanjing University, Nanjing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-9270-5072","authenticated-orcid":false,"given":"Chuanyi","family":"Li","sequence":"additional","affiliation":[{"name":"National Key Laboratory for Novel Software Technology, Nanjing University, Nanjing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0005-2492-2530","authenticated-orcid":false,"given":"Zongwen","family":"Shen","sequence":"additional","affiliation":[{"name":"National Key Laboratory for Novel Software Technology, Nanjing University, Nanjing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-1773-0942","authenticated-orcid":false,"given":"Jidong","family":"Ge","sequence":"additional","affiliation":[{"name":"National Key Laboratory for Novel Software Technology, Nanjing University, Nanjing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-7426-1594","authenticated-orcid":false,"given":"Wenkang","family":"Zhong","sequence":"additional","affiliation":[{"name":"National Key Laboratory for Novel Software Technology, Nanjing University, Nanjing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0001-1102-9584","authenticated-orcid":false,"given":"Bin","family":"Luo","sequence":"additional","affiliation":[{"name":"National Key Laboratory for Novel Software Technology, Nanjing University, Nanjing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-8237-429X","authenticated-orcid":false,"given":"Vincent","family":"Ng","sequence":"additional","affiliation":[{"name":"Human Language Technology Research Institute, The University of Texas at Dallas, Richardson, Texas, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2026,3,12]]},"reference":[{"key":"e_1_3_2_2_2","unstructured":"Toufique Ahmed Christian Bird Premkumar Devanbu and Saikat Chakraborty. 2024. Studying LLM performance on closed- and open-source data. arXiv:2404.15247. Retrieved from https:\/\/arxiv.org\/abs\/2404.15247"},{"key":"e_1_3_2_3_2","doi-asserted-by":"publisher","DOI":"10.1109\/TechDebt59074.2023.00011"},{"key":"e_1_3_2_4_2","doi-asserted-by":"publisher","DOI":"10.4230\/DagRep.6.4.110"},{"key":"e_1_3_2_5_2","doi-asserted-by":"publisher","DOI":"10.1007\/s10664-021-10081-7"},{"key":"e_1_3_2_6_2","doi-asserted-by":"publisher","DOI":"10.1145\/2901739.2901742"},{"key":"e_1_3_2_7_2","doi-asserted-by":"publisher","DOI":"10.1145\/2901739.2901754"},{"issue":"70","key":"e_1_3_2_8_2","first-page":"1","article-title":"Scaling instruction-finetuned language models","volume":"25","author":"Won Chung Hyung","year":"2024","unstructured":"Hyung Won Chung, Le Hou, Shayne Longpre, Barret Zoph, Yi Tay, William Fedus, Yunxuan Li, Xuezhi Wang, Mostafa Dehghani, Siddhartha Brahma, et al. 2024. Scaling instruction-finetuned language models. Journal of Machine Learning Research (JMLR) 25, 70 (2024), 1\u201353.","journal-title":"Journal of Machine Learning Research (JMLR)"},{"key":"e_1_3_2_9_2","doi-asserted-by":"publisher","DOI":"10.1109\/TSE.2017.2654244"},{"key":"e_1_3_2_10_2","doi-asserted-by":"publisher","unstructured":"Haixing Dai Zhengliang Liu Wenxiong Liao Xiaoke Huang Yihan Cao Zihao Wu Lin Zhao Shaochen Xu Wei Liu Ninghao Liu et al. 2023. Auggpt: Leveraging ChatGPT for text data augmentation. arXiv:2302.13007. DOI: 10.48550\/arXiv.2302.13007","DOI":"10.48550\/arXiv.2302.13007"},{"key":"e_1_3_2_11_2","unstructured":"Ke Dai and Philippe Kruchten. 2017. Detecting technical debt through issue trackers. In Proceedings of the 5th International Workshop on Quantitative Approaches to Software Quality (QuASoQ@ APSEC) 59\u201365."},{"key":"e_1_3_2_12_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/N19-1423"},{"key":"e_1_3_2_13_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2024.acl-long.699"},{"key":"e_1_3_2_14_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2020.findings-emnlp.139"},{"key":"e_1_3_2_15_2","article-title":"Self-guided noise-free data generation for efficient zero-shot learning","author":"Gao Jiahui","year":"2023","unstructured":"Jiahui Gao, Renjie Pi, Lin Yong, Hang Xu, Jiacheng Ye, Zhiyong Wu, Weizhong Zhang, Xiaodan Liang, Zhenguo Li, and Lingpeng Kong. 2023. Self-guided noise-free data generation for efficient zero-shot learning. In Proceedings of the International Conference on Learning Representations (ICLR \u201923).","journal-title":"Proceedings of the International Conference on Learning Representations (ICLR \u201923)"},{"key":"e_1_3_2_16_2","doi-asserted-by":"publisher","unstructured":"Team GLM. 2024. ChatGLM: A family of large language models from GLM-130B to GLM-4 all tools. arXiv:2406.12793. DOI: 10.48550\/arXiv.2406.12793","DOI":"10.48550\/arXiv.2406.12793"},{"key":"e_1_3_2_17_2","doi-asserted-by":"publisher","DOI":"10.1109\/SANER60148.2024.00087"},{"key":"e_1_3_2_18_2","doi-asserted-by":"publisher","DOI":"10.1145\/3447247"},{"key":"e_1_3_2_19_2","doi-asserted-by":"publisher","DOI":"10.13328\/j.cnki.jos.006292"},{"key":"e_1_3_2_20_2","doi-asserted-by":"publisher","DOI":"10.9781\/ijimai.2016.415"},{"key":"e_1_3_2_21_2","doi-asserted-by":"publisher","unstructured":"Dan Hendrycks and Kevin Gimpel. 2016. Gaussian error linear units (GELUs). arXiv:1606.08415. DOI: 10.48550\/arXiv.1606.08415","DOI":"10.48550\/arXiv.1606.08415"},{"key":"e_1_3_2_22_2","doi-asserted-by":"publisher","DOI":"10.3390\/app14219863"},{"key":"e_1_3_2_23_2","doi-asserted-by":"publisher","unstructured":"Dong Huang Qingwen Bu and Heming Cui. 2023. Codecot and beyond: Learning to program and test like a developer. arXiv:2308.08784. DOI: 10.48550\/arXiv.2308.08784","DOI":"10.48550\/arXiv.2308.08784"},{"key":"e_1_3_2_24_2","doi-asserted-by":"publisher","unstructured":"Kai Huang Zhengzi Xu Su Yang Hongyu Sun Xuejun Li Zheng Yan and Yuqing Zhang. 2023. A survey on automated program repair techniques. arXiv:2303.18184. DOI: 10.48550\/arXiv.2303.18184","DOI":"10.48550\/arXiv.2303.18184"},{"key":"e_1_3_2_25_2","doi-asserted-by":"publisher","DOI":"10.1007\/s10664-017-9522-4"},{"key":"e_1_3_2_26_2","doi-asserted-by":"publisher","unstructured":"Nan Jiang Xiaopeng Li Shiqi Wang Qiang Zhou Soneya Binta Hossain Baishakhi Ray Varun Kumar Xiaofei Ma and Anoop Deoras. 2024. Training LLMs to better self-debug and explain code. arXiv:2405.18649. DOI: 10.48550\/arXiv.2405.18649","DOI":"10.48550\/arXiv.2405.18649"},{"key":"e_1_3_2_27_2","doi-asserted-by":"publisher","DOI":"10.1145\/3611643.3613892"},{"key":"e_1_3_2_28_2","doi-asserted-by":"publisher","unstructured":"Ryo Kamoi Yusen Zhang Nan Zhang Jiawei Han and Rui Zhang. 2024. When can LLMs actually correct their own mistakes? A critical survey of self-correction of LLMs. arXiv:2406.01297. DOI: 10.48550\/arXiv.2406.01297","DOI":"10.48550\/arXiv.2406.01297"},{"key":"e_1_3_2_29_2","doi-asserted-by":"publisher","DOI":"10.1109\/SANER53432.2022.00094"},{"key":"e_1_3_2_30_2","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.1911.00172"},{"key":"e_1_3_2_31_2","doi-asserted-by":"publisher","DOI":"10.1002\/spe.3360"},{"key":"e_1_3_2_32_2","doi-asserted-by":"publisher","DOI":"10.1109\/SEAA51224.2020.00083"},{"key":"e_1_3_2_33_2","doi-asserted-by":"publisher","DOI":"10.1007\/s10664-022-10128-3"},{"key":"e_1_3_2_34_2","doi-asserted-by":"publisher","DOI":"10.1007\/s10515-024-00462-9"},{"key":"e_1_3_2_35_2","doi-asserted-by":"publisher","DOI":"10.1145\/3377815.3381377"},{"key":"e_1_3_2_36_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/P19-1441"},{"key":"e_1_3_2_37_2","first-page":"22631","volume-title":"Proceedings of the International Conference on Machine Learning (ICML \u201923)","author":"Longpre Shayne","year":"2023","unstructured":"Shayne Longpre, Le Hou, Tu Vu, Albert Webson, Hyung Won Chung, Yi Tay, Denny Zhou, Quoc V. Le, Barret Zoph, Jason Wei, et al. 2023. The flan collection: Designing data and methods for effective instruction tuning. In Proceedings of the International Conference on Machine Learning (ICML \u201923), 22631\u201322648."},{"key":"e_1_3_2_38_2","doi-asserted-by":"publisher","DOI":"10.1109\/MTD.2015.7332619"},{"key":"e_1_3_2_39_2","doi-asserted-by":"publisher","DOI":"10.1145\/3540250.3549088"},{"key":"e_1_3_2_40_2","doi-asserted-by":"publisher","unstructured":"Seokjin Oh Su Ah Lee and Woohwan Jung. 2023. Data augmentation for neural machine translation using generative language model. arXiv:2307.16833. DOI: 10.48550\/arXiv.2307.16833","DOI":"10.48550\/arXiv.2307.16833"},{"key":"e_1_3_2_41_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICSME.2014.31"},{"key":"e_1_3_2_42_2","doi-asserted-by":"publisher","DOI":"10.1080\/09540091.2022.2067125"},{"issue":"140","key":"e_1_3_2_43_2","first-page":"1","article-title":"Exploring the limits of transfer learning with a unified text-to-text transformer","volume":"21","author":"Raffel Colin","year":"2020","unstructured":"Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2020. Exploring the limits of transfer learning with a unified text-to-text transformer. Journal of Machine Learning Research 21, 140 (2020), 1\u201367.","journal-title":"Journal of Machine Learning Research"},{"key":"e_1_3_2_44_2","doi-asserted-by":"publisher","DOI":"10.1145\/3324916"},{"key":"e_1_3_2_45_2","first-page":"2980","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR \u201917)","author":"Ross T.-Y. L. P. G.","year":"2017","unstructured":"T.-Y. L. P. G. Ross and G. K. H. P. Doll\u00e1r. 2017. Focal loss for dense object detection. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR \u201917), 2980\u20132988."},{"key":"e_1_3_2_46_2","doi-asserted-by":"publisher","DOI":"10.11591\/ijece.v13i2.pp2142-2155"},{"key":"e_1_3_2_47_2","volume-title":"Proceedings of the International Conference on Learning Representations (ICLR \u201917)","author":"Shazeer Noam","year":"2017","unstructured":"Noam Shazeer, Azalia Mirhoseini, Krzysztof Maziarz, Andy Davis, Quoc Le, Geoffrey Hinton, and Jeff Dean. 2017. Outrageously large neural networks: The Sparsely-Gated mixture-of-experts layer. In Proceedings of the International Conference on Learning Representations (ICLR \u201917)."},{"key":"e_1_3_2_48_2","doi-asserted-by":"publisher","DOI":"10.1007\/s10664-024-10548-3"},{"key":"e_1_3_2_49_2","doi-asserted-by":"publisher","DOI":"10.1109\/APSEC65559.2024.00022"},{"key":"e_1_3_2_50_2","doi-asserted-by":"publisher","DOI":"10.1145\/3643991.3644880"},{"key":"e_1_3_2_51_2","doi-asserted-by":"publisher","DOI":"10.1109\/ASE56229.2023.00076"},{"key":"e_1_3_2_52_2","first-page":"24824","volume-title":"Advances in Neural Information Processing Systems (NeurIPS)","volume":"35","author":"Wei Jason","year":"2022","unstructured":"Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V. Le, Denny Zhou, et al. 2022. Chain-of-thought prompting elicits reasoning in large language models. In Advances in Neural Information Processing Systems (NeurIPS), Vol. 35, 24824\u201324837."},{"key":"e_1_3_2_53_2","doi-asserted-by":"publisher","DOI":"10.1145\/3379597.3387459"},{"key":"e_1_3_2_54_2","doi-asserted-by":"publisher","unstructured":"Chunqiu Steven Xia Yifeng Ding and Lingming Zhang. 2023. Revisiting the plastic surgery hypothesis via large language models. arXiv:2303.10494. DOI: 10.48550\/arXiv.2303.10494","DOI":"10.48550\/arXiv.2303.10494"},{"key":"e_1_3_2_55_2","doi-asserted-by":"publisher","DOI":"10.1145\/3597503.3639121"},{"key":"e_1_3_2_56_2","doi-asserted-by":"publisher","DOI":"10.13328\/j.cnki.jos.006981"},{"key":"e_1_3_2_57_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2022.emnlp-main.801"},{"key":"e_1_3_2_58_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.jss.2022.111219"},{"key":"e_1_3_2_59_2","doi-asserted-by":"publisher","unstructured":"Xiao Yu Lei Liu Xing Hu Jacky Wai Keung Jin Liu and Xin Xia. 2024. Fight fire with fire: How much can we trust ChatGPT on source code-related tasks? arXiv:2405.12641. DOI: 10.48550\/arXiv.2405.12641","DOI":"10.48550\/arXiv.2405.12641"},{"key":"e_1_3_2_60_2","doi-asserted-by":"publisher","DOI":"10.1109\/TSE.2020.3031401"},{"key":"e_1_3_2_61_2","doi-asserted-by":"publisher","DOI":"10.1007\/s10664-021-10031-3"},{"key":"e_1_3_2_62_2","doi-asserted-by":"publisher","DOI":"10.1145\/3196398.3196423"},{"key":"e_1_3_2_63_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICME52920.2022.9859706"},{"key":"e_1_3_2_64_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2023.findings-acl.99"},{"key":"e_1_3_2_65_2","doi-asserted-by":"publisher","DOI":"10.1145\/3580305.3599790"},{"key":"e_1_3_2_66_2","doi-asserted-by":"publisher","unstructured":"Yue Zhou Chenlu Guo Xu Wang Yi Chang and Yuan Wu. 2024. A survey on data augmentation in large model era. arXiv:2401.15422. DOI: 10.48550\/arXiv.2401.15422","DOI":"10.48550\/arXiv.2401.15422"},{"key":"e_1_3_2_67_2","doi-asserted-by":"publisher","DOI":"10.1145\/3512345"}],"container-title":["ACM Transactions on Software Engineering and Methodology"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3747180","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,3,12]],"date-time":"2026-03-12T15:04:39Z","timestamp":1773327879000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3747180"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,3,12]]},"references-count":66,"journal-issue":{"issue":"4","published-print":{"date-parts":[[2026,4,30]]}},"alternative-id":["10.1145\/3747180"],"URL":"https:\/\/doi.org\/10.1145\/3747180","relation":{},"ISSN":["1049-331X","1557-7392"],"issn-type":[{"value":"1049-331X","type":"print"},{"value":"1557-7392","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,3,12]]},"assertion":[{"value":"2024-11-26","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-06-24","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2026-03-12","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}