{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,23]],"date-time":"2026-07-23T09:03:54Z","timestamp":1784797434317,"version":"3.55.0"},"reference-count":72,"publisher":"Springer Science and Business Media LLC","issue":"1","license":[{"start":{"date-parts":[[2026,7,23]],"date-time":"2026-07-23T00:00:00Z","timestamp":1784764800000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2026,7,23]],"date-time":"2026-07-23T00:00:00Z","timestamp":1784764800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"DOI":"10.13039\/100013000","name":"Politecnico di Torino","doi-asserted-by":"crossref","id":[{"id":"10.13039\/100013000","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Empir Software Eng"],"published-print":{"date-parts":[[2027,2]]},"abstract":"<jats:title>Abstract<\/jats:title>\n                  <jats:p>\n                    Code comprehension is vital for software development. Still, unreadable code remains a significant issue, costing substantial time and money losses. While tools to identify code exhibiting a low readability exist, take actions to improve such a quality aspect is far from trivial. To support developers in such a task, Vitale et al. introduced at ASE\u201923 an approach using a transformer model (T5) fine-tuned on code commits in which developers explicitly stated their goal to improve code readability. The authors reported that their model is able to generate readability-improving changes being identical to those implemented by developers (exact matches) in 21% of cases and that, even when they differ, they still improve readability in the vast majority of cases (\n                    <jats:inline-formula>\n                      <jats:tex-math>$$\\sim $$<\/jats:tex-math>\n                    <\/jats:inline-formula>\n                    80%). Given the major advances in AI made in the last few years, we questioned whether a fine-tuning for such a task was still needed in the era of Large Language Models (LLMs), thus revisiting the work by Vitale et al. with state-of-the-art models. In doing so, we found out one major issue with the original study design: The training-test splitting was performed randomly (\n                    <jats:italic>i.e.,<\/jats:italic>\n                    readability-improving commits mined from open source projects were randomly split between training and test) rather than by project, significantly inflating the reported performance due to repeating readability-improving commits done within the same project. Indeed, as we will show, even newer and more powerful LLMs fine-tuned for this task can only achieve a\n                    <jats:inline-formula>\n                      <jats:tex-math>$$\\sim $$<\/jats:tex-math>\n                    <\/jats:inline-formula>\n                    1% of exact matches. Once fixed this issue, we shifted our attention on assessing whether fine-tuning on available developer readability improvements is actually needed given the recent advances in general-purpose LLMs. We compare the performance of fine-tuned state-of-the-art LLMs with what they achieve in zero-shot setting, also including in our study commercial LLMs (GPT-4.1). As we will show, quantitative metrics (\n                    <jats:italic>e.g.,<\/jats:italic>\n                    exact matches, CrystalBLEU) tell very little about the readability-improving capabilities of LLMs and, thus, our study is mainly qualitative in nature, with a total of\n                    <jats:inline-formula>\n                      <jats:tex-math>$$\\sim $$<\/jats:tex-math>\n                    <\/jats:inline-formula>\n                    9,500 inspected LLMs\u2019 change recommendations. While all models do improve code readability, our results clearly show that fine-tuning using mined readability-improving commits does not help and, instead, results in sensibly poorer readability recommendations as compared to LLMs used in a zero-shot setting, with GPT-4.1 being the one achieving the best results.\n                  <\/jats:p>","DOI":"10.1007\/s10664-026-10933-0","type":"journal-article","created":{"date-parts":[[2026,7,23]],"date-time":"2026-07-23T08:38:30Z","timestamp":1784795910000},"update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":0,"title":["Revisiting code readability improvement with LLMs: A critical assessment of fine-tuned models"],"prefix":"10.1007","volume":"32","author":[{"ORCID":"https:\/\/orcid.org\/0009-0009-5840-9777","authenticated-orcid":false,"given":"Antonio","family":"Vitale","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Emanuela","family":"Guglielmi","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Valentina","family":"Piantadosi","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Antonio","family":"Mastropaolo","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Gabriele","family":"Bavota","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Rocco","family":"Oliveto","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Simone","family":"Scalabrino","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"297","published-online":{"date-parts":[[2026,7,23]]},"reference":[{"key":"10933_CR1","doi-asserted-by":"crossref","unstructured":"Ahmed T, Devanbu P (2022) Few-shot training llms for project-specific code-summarization. In: Proceedings of the 37th IEEE\/ACM International conference on automated software engineerin, pp 1\u20135","DOI":"10.1145\/3551349.3559555"},{"key":"10933_CR2","doi-asserted-by":"crossref","unstructured":"Allamanis M, Barr ET, Bird C, Sutton C (2014) Learning natural coding conventions. In: Proceedings of the 22nd ACM SIGSOFT international symposium on foundations of software engineering, pp 281\u2013293","DOI":"10.1145\/2635868.2635883"},{"key":"10933_CR3","doi-asserted-by":"publisher","unstructured":"AnonymousVitale A, Guglielmi E, Piantadosi V, Mastropaolo A, Bavota G, Oliveto R, Scalabrino S (2025) Replication package for \u201crevisiting code readability improvements with llms: a replication and critical assessment of fine-tuned models\u201d. https:\/\/doi.org\/10.5281\/zenodo.16022058","DOI":"10.5281\/zenodo.16022058"},{"issue":"1","key":"10933_CR4","doi-asserted-by":"publisher","first-page":"289","DOI":"10.1111\/j.2517-6161.1995.tb02031.x","volume":"57","author":"Y Benjamini","year":"1995","unstructured":"Benjamini Y, Hochberg Y (1995) Controlling the false discovery rate: a practical and powerful approach to multiple testing. J Roy Stat Soc: Ser B (Methodol) 57(1):289\u2013300","journal-title":"J Roy Stat Soc: Ser B (Methodol)"},{"key":"10933_CR5","first-page":"1877","volume":"33","author":"T Brown","year":"2020","unstructured":"Brown T, Mann B, Ryder N, Subbiah M, Kaplan JD, Dhariwal P, Neelakantan A, Shyam P, Sastry G, Askell A et al (2020) Language models are few-shot learners. Adv Neural Inf Process Syst 33:1877\u20131901","journal-title":"Adv Neural Inf Process Syst"},{"key":"10933_CR6","doi-asserted-by":"crossref","unstructured":"Buse RP, Weimer WR (2008) A metric for software readability. In: Proceedings of the 2008 international symposium on Software testing and analysis, pp 121\u2013130","DOI":"10.1145\/1390630.1390647"},{"issue":"4","key":"10933_CR7","doi-asserted-by":"publisher","first-page":"546","DOI":"10.1109\/TSE.2009.70","volume":"36","author":"RP Buse","year":"2009","unstructured":"Buse RP, Weimer WR (2009) Learning a metric for code readability. IEEE Trans Software Eng 36(4):546\u2013558","journal-title":"IEEE Trans Software Eng"},{"key":"10933_CR8","unstructured":"Chen M, Tworek J, Jun H, Yuan Q, Pinto HPDO, Kaplan J, Edwards H, Burda Y, Joseph N, Brockman G (2021) Evaluating large language models trained on code. arXiv:2107.03374"},{"key":"10933_CR9","unstructured":"Crupi G, Tufano R, Bavota G (2026) Improving code generation via small language model-as-a-judge. arXiv:2602.11911"},{"key":"10933_CR10","unstructured":"Dorn J (2012) A general software readability model. MCS Thesis available from (http:\/\/www.cs.virginia.edu\/weimer\/students\/dorn-mcs-paper.pdf) 5:11\u201314"},{"key":"10933_CR11","doi-asserted-by":"crossref","unstructured":"Eghbali A, Pradel M (2022) Crystalbleu: precisely and efficiently measuring the similarity of code. In: Proceedings of the 37th IEEE\/ACM international conference on automated software engineering, pp 1\u201312","DOI":"10.1145\/3551349.3556903"},{"issue":"3","key":"10933_CR12","doi-asserted-by":"publisher","first-page":"493","DOI":"10.1007\/s00355-014-0853-4","volume":"44","author":"E Elkind","year":"2015","unstructured":"Elkind E, Lang J, Saffidine A (2015) Condorcet winning sets. Soc Choice Welfare 44(3):493\u2013517","journal-title":"Soc Choice Welfare"},{"issue":"3","key":"10933_CR13","doi-asserted-by":"publisher","first-page":"17","DOI":"10.1109\/6294.846201","volume":"2","author":"L Erlikh","year":"2000","unstructured":"Erlikh L (2000) Leveraging legacy system dollars for e-business. IT professional 2(3):17\u201323","journal-title":"IT professional"},{"key":"10933_CR14","unstructured":"Fisher RA et\u00a0al (1956) Mathematics of a lady tasting tea. The world of mathematics 3(part 8):1514\u20131521"},{"key":"10933_CR15","doi-asserted-by":"crossref","unstructured":"Fraser G, Arcuri A (2011) Evosuite: automatic test suite generation for object-oriented software. In: Proceedings of the 19th ACM SIGSOFT symposium and the 13th European conference on Foundations of software engineering, pp 416\u2013419","DOI":"10.1145\/2025113.2025179"},{"key":"10933_CR16","unstructured":"Grigorik I (2012) GitHub Archive. https:\/\/www.gharchive.org\/"},{"key":"10933_CR17","unstructured":"Guo D, Zhu Q, Yang D, Xie Z, Dong K, Zhang W, Chen G, Bi X, Wu Y, Li Y (2024a) Deepseek-coder: When the large language model meets programming-the rise of code intelligence. arXiv:2401.14196"},{"key":"10933_CR18","doi-asserted-by":"crossref","unstructured":"Guo Q, Cao J, Xie X, Liu S, Li X, Chen B, Peng X (2024b) Exploring the potential of chatgpt in automated code refinement: An empirical study. In: Proceedings of the 46th IEEE\/ACM international conference on software engineerin, pp 1\u201313","DOI":"10.1145\/3597503.3623306"},{"key":"10933_CR19","doi-asserted-by":"crossref","unstructured":"Huang K, Zhang J, Meng X, Liu Y (2025) Template-guided program repair in the era of large language models. In: ICSE, pp 1895\u20131907","DOI":"10.1109\/ICSE55347.2025.00030"},{"key":"10933_CR20","unstructured":"Hui B, Yang J, Cui Z, Yang J, Liu D, Zhang L, Liu T, Zhang J, Yu B, Lu K et\u00a0al (2024) Qwen2. 5-coder technical report. arXiv preprint arXiv:2409.12186"},{"key":"10933_CR21","doi-asserted-by":"crossref","unstructured":"Improta C, Tufano R, Liguori P, Cotroneo D, Bavota G (2025) Quality in, quality out: Investigating training data\u2019s role in ai code generation. arXiv:2503.11402","DOI":"10.1109\/ICPC66645.2025.00056"},{"key":"10933_CR22","unstructured":"Javaparser: Javaparser. https:\/\/javaparser.org\/"},{"key":"10933_CR23","doi-asserted-by":"crossref","unstructured":"Kavathekar I, Donakanti R, Kumaraguru P, Vaidhyanathan K (2025) Small models, big tasks: An exploratory empirical study on small language models for function calling. In: Proceedings of the 29th international conference on evaluation and assessment in software engineering, pp 1117\u20131126","DOI":"10.1145\/3756681.3757001"},{"key":"10933_CR24","unstructured":"Krishnamoorthy MS, Raghavachari M (2005) Condorcet winner probabilities-a statistical perspective"},{"key":"10933_CR25","unstructured":"Levenshtein VI (1966) Binary codes capable of correcting deletions, insertions, and reversals. Soviet physics doklady, Soviet Union 10:707\u2013710"},{"key":"10933_CR26","doi-asserted-by":"crossref","unstructured":"Li Z, Lu S, Guo D, Duan N, Jannu S, Jenks G, Majumder D, Green J, Svyatkovskiy A, Fu S (2022) Automating code review activities by large-scale pre-training. In: Proceedings of the 30th ACM joint european software engineering conference and symposium on the foundations of software engineering, pp 1035\u20131047","DOI":"10.1145\/3540250.3549081"},{"key":"10933_CR27","doi-asserted-by":"publisher","first-page":"14200","DOI":"10.52202\/079017-0455","volume":"37","author":"J Li","year":"2024","unstructured":"Li J, Fang A, Smyrnis G, Ivgi M, Jordan M, Gadre SY, Bansal H, Guha E, Keh SS, Arora K et al (2024) Datacomp-lm: In search of the next generation of training sets for language models. Adv Neural Inf Process Syst 37:14200\u201314282","journal-title":"Adv Neural Inf Process Syst"},{"key":"10933_CR28","doi-asserted-by":"crossref","unstructured":"Lin B, Scalabrino S, Mocci A, Oliveto R, Bavota G, Lanza M (2017) Investigating the use of code analysis and nlp to promote a consistent usage of identifiers. In: IEEE 17th International Working Conference on Source Code Analysis and Manipulation (SCAM), pp 81\u201390. IEEE","DOI":"10.1109\/SCAM.2017.17"},{"key":"10933_CR29","doi-asserted-by":"crossref","unstructured":"Lin B, Wang S, Wen M, Chen L, Mao X (2024) One size does not fit all: Multi-granularity patch generation for better automated program repair. In: Proceedings of the 33rd ACM SIGSOFT international symposium on software testing and analysis, pp 1554\u20131566","DOI":"10.1145\/3650212.3680381"},{"key":"10933_CR30","doi-asserted-by":"crossref","unstructured":"L\u00f3pez JAH, Chen B, Saad M, Sharma T, Varr\u00f3 D (2024) On inter-dataset code duplication and data leakage in large language models. IEEE Trans Software Eng","DOI":"10.1109\/TSE.2024.3504286"},{"key":"10933_CR31","unstructured":"Loriot B, Madeiral F, Monperrus M (2019) Styler: Learning formatting conventions to repair checkstyle errors. arXiv:1904.01754"},{"key":"10933_CR32","unstructured":"Loshchilov I, Hutter F (2017) Decoupled weight decay regularization. arXiv:1711.05101"},{"key":"10933_CR33","unstructured":"Lozhkov A, Li R, Allal LB, Cassano F, Lamy-Poirier J, Tazi N, Tang A, Pykhtar D, Liu J, Wei Y et\u00a0al (2024) Starcoder 2 and the stack v2: The next generation. arXiv:2402.19173"},{"key":"10933_CR34","doi-asserted-by":"crossref","unstructured":"Mastropaolo A, Aghajani E, Pascarella L, Bavota G (2021) An empirical study on code comment completion. IN: 2021 IEEE International Conference on Software Maintenance and Evolution (ICSME), pp 159\u2013170. IEEE","DOI":"10.1109\/ICSME52107.2021.00021"},{"key":"10933_CR35","doi-asserted-by":"publisher","first-page":"60","DOI":"10.1016\/j.infsof.2018.07.006","volume":"104","author":"Q Mi","year":"2018","unstructured":"Mi Q, Keung J, Xiao Y, Mensah S, Gao Y (2018) Improving code readability classification using convolutional neural networks. Inf Softw Technol 104:60\u201371","journal-title":"Inf Softw Technol"},{"key":"10933_CR36","doi-asserted-by":"publisher","DOI":"10.1016\/j.jss.2022.111454","volume":"193","author":"Q Mi","year":"2022","unstructured":"Mi Q, Hao Y, Ou L, Ma W (2022) Towards using visual, semantic and structural features to improve code readability classification. J Syst Softw 193:111454","journal-title":"J Syst Softw"},{"issue":"4","key":"10933_CR37","doi-asserted-by":"publisher","first-page":"87","DOI":"10.1007\/s10664-023-10319-6","volume":"28","author":"Q Mi","year":"2023","unstructured":"Mi Q, Zhan Y, Weng H, Bao Q, Cui L, Ma W (2023) A graph-based code representation method to improve code readability classification. Empir Softw Eng 28(4):87","journal-title":"Empir Softw Eng"},{"key":"10933_CR38","doi-asserted-by":"crossref","unstructured":"Minelli R, Mocci A, Lanza M (2015) I know what you did last summer-an investigation of how developers spend their time. In: IEEE 23rd international conference on program comprehension, pp 25\u201335. IEEE","DOI":"10.1109\/ICPC.2015.12"},{"key":"10933_CR39","doi-asserted-by":"crossref","unstructured":"Oliveira D, Santos R, de\u00a0Oliveira B, Monperrus M, Castor F, Madeiral F (2024) Understanding code understandability improvements in code reviews. IEEE Trans Softw Eng","DOI":"10.1109\/TSE.2024.3453783"},{"key":"10933_CR40","unstructured":"OpenAI: GPT4.1. https:\/\/platform.openai.com\/docs\/models\/gpt-4.1"},{"key":"10933_CR41","doi-asserted-by":"crossref","unstructured":"Papineni K, Roukos S, Ward T, Zhu WJ (2002) Bleu: a method for automatic evaluation of machine translation. In: Proceedings of the 40th annual meeting of the Association for Computational Linguistics, pp 311\u2013318","DOI":"10.3115\/1073083.1073135"},{"key":"10933_CR42","doi-asserted-by":"publisher","first-page":"30811","DOI":"10.52202\/079017-0970","volume":"37","author":"G Penedo","year":"2024","unstructured":"Penedo G, Kydl\u00ed\u010dek H, Lozhkov A, Mitchell M, Raffel CA, Von Werra L, Wolf T et al (2024) The fineweb datasets: Decanting the web for the finest text data at scale. Adv Neural Inf Process Syst 37:30811\u201330849","journal-title":"Adv Neural Inf Process Syst"},{"issue":"6","key":"10933_CR43","doi-asserted-by":"publisher","first-page":"5374","DOI":"10.1007\/s10664-020-09886-9","volume":"25","author":"V Piantadosi","year":"2020","unstructured":"Piantadosi V, Fierro F, Scalabrino S, Serebrenik A, Oliveto R (2020) How does code readability change during software evolution? Empir Softw Eng 25(6):5374\u20135412","journal-title":"Empir Softw Eng"},{"key":"10933_CR44","doi-asserted-by":"crossref","unstructured":"Posnett D, Hindle A, Devanbu P (2011) A simpler model of software readability. In: Proceedings of the 8th working conference on mining software repositories, pp 73\u201382","DOI":"10.1145\/1985441.1985454"},{"key":"10933_CR45","doi-asserted-by":"crossref","unstructured":"Prechelt L (2002) Early stopping-but when? Neural Networks: Tricks of the trade, pp 55\u201369. Springer","DOI":"10.1007\/3-540-49430-8_3"},{"key":"10933_CR46","doi-asserted-by":"crossref","unstructured":"Prenner JA, Robbes R (2024) Out of context: How important is local context in neural program repair? In: Proceedings of the IEEE\/ACM 46th international conference on software engineering, pp 1\u201313","DOI":"10.1145\/3597503.3639086"},{"key":"10933_CR47","unstructured":"Raffel C, Shazeer N, Roberts A, Lee K, Narang S, Matena M, Zhou Y, Li W, Liu PJ (2019) Exploring the limits of transfer learning with a unified text-to-text transformer. arXiv:1910.10683"},{"key":"10933_CR48","doi-asserted-by":"crossref","unstructured":"Ren H, Zhan M, Wu Z, Zhou A, Pan J, Li H (2025) Reflectioncoder: Learning from reflection sequence for enhanced one-off code generation. In: Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp 9999\u201310020","DOI":"10.18653\/v1\/2025.acl-long.494"},{"key":"10933_CR49","unstructured":"Roziere B, Gehring J, Gloeckle F, Sootla S, Gat I, Tan XE, Adi Y, Liu J, Sauvestre R, Remez T (2023) Code llama: Open foundation models for code. arXiv:2308.12950"},{"key":"10933_CR50","doi-asserted-by":"crossref","unstructured":"Saari DG, Merlin VR (1996) The copeland method: I.: Relationships and the dictionary. Econ Theor 8:51\u201376","DOI":"10.1007\/BF01212012"},{"key":"10933_CR51","doi-asserted-by":"crossref","unstructured":"Scalabrino S, Linares-Vasquez M, Poshyvanyk D, Oliveto R (2016) Improving code readability models with textual features. In: 2016 IEEE 24th International Conference on Program Comprehension (ICPC), pp 1\u201310. IEEE","DOI":"10.1109\/ICPC.2016.7503707"},{"issue":"6","key":"10933_CR52","doi-asserted-by":"publisher","first-page":"e1958","DOI":"10.1002\/smr.1958","volume":"30","author":"S Scalabrino","year":"2018","unstructured":"Scalabrino S, Linares-V\u00e1squez M, Oliveto R, Poshyvanyk D (2018) A comprehensive model for code readability. J Softw Evol Process 30(6):e1958","journal-title":"J Softw Evol Process"},{"key":"10933_CR53","unstructured":"Sergeyuk A, Lvova O, Titov S, Serova A, Bagirov F, Bryksin T (2024a) Assessing consensus of developers\u2019 views on code readability. arXiv:2407.03790"},{"key":"10933_CR54","doi-asserted-by":"crossref","unstructured":"Sergeyuk A, Lvova O, Titov S, Serova A, Bagirov F, Kirillova E, Bryksin T (2024b) Reassessing java code readability models with a human-centered approach. In: Proceedings of the 32nd IEEE\/ACM international conference on program comprehension, pp 225\u2013235","DOI":"10.1145\/3643916.3644435"},{"key":"10933_CR55","doi-asserted-by":"crossref","unstructured":"Shi E, Wang Y, Du L, Chen J, Han S, Zhang H, Zhang D, Sun H (2022a) On the evaluation of neural code summarization. In: Proceedings of the 44th international conference on software engineering, pp 1597\u20131608","DOI":"10.1145\/3510003.3510060"},{"key":"10933_CR56","doi-asserted-by":"crossref","unstructured":"Shi L, Mu F, Chen X, Wang S, Wang J, Yang Y, Li G, Xia X, Wang Q (2022b) Are we building on the rock? on the importance of data preprocessing for code summarization. In: Proceedings of the 30th ACM joint European software engineering conference and symposium on the foundations of software engineering, pp 107\u2013119","DOI":"10.1145\/3540250.3549145"},{"key":"10933_CR57","doi-asserted-by":"crossref","unstructured":"Silva A, Fang S, Monperrus M (2025) Repairllama: Efficient representations and fine-tuned adapters for program repair. IEEE Trans Software Eng","DOI":"10.1109\/TSE.2025.3581062"},{"key":"10933_CR58","doi-asserted-by":"crossref","unstructured":"Sun W, Miao Y, Li Y, Zhang H, Fang C, Liu Y, Deng G, Liu Y, Chen Z (2024) Source code summarization in the era of large language models. arXiv preprint arXiv:2407.07959","DOI":"10.1109\/ICSE55347.2025.00034"},{"key":"10933_CR59","doi-asserted-by":"crossref","unstructured":"Thies A, Roth C (2010) Recommending rename refactorings. In: Proceedings of the 2nd international workshop on recommendation systems for software engineering, pp 1\u20135","DOI":"10.1145\/1808920.1808921"},{"key":"10933_CR60","doi-asserted-by":"crossref","unstructured":"Tufano R, Pascarella L, Bavota G (2023) Automating code-related tasks through transformers: The impact of pre-training. In: 2023 IEEE\/ACM 45th International Conference on Software Engineering (ICSE), pp 2425\u20132437. IEEE","DOI":"10.1109\/ICSE48619.2023.00203"},{"key":"10933_CR61","doi-asserted-by":"crossref","unstructured":"Vitale A, Piantadosi V, Scalabrino S, Oliveto R (2023) Using deep learning to automatically improve code readability. In: 38th IEEE\/ACM International Conference on Automated Software Engineering (ASE), pp 573\u2013584. IEEE","DOI":"10.1109\/ASE56229.2023.00112"},{"key":"10933_CR62","doi-asserted-by":"crossref","unstructured":"Vitale A, Guglielmi E, Oliveto R, Scalabrino S (2025a) Personalized code readability assessment: Are we there yet?. arXiv:2503.07870","DOI":"10.1109\/ICPC66645.2025.00019"},{"key":"10933_CR63","doi-asserted-by":"crossref","unstructured":"Vitale A, Oliveto R, Scalabrino S (2025b) A catalog of data smells for coding tasks. ACM Trans Softw Eng Methodol 34(4):1\u201332","DOI":"10.1145\/3707457"},{"issue":"4","key":"10933_CR64","doi-asserted-by":"publisher","first-page":"911","DOI":"10.1109\/TSE.2024.3368208","volume":"50","author":"J Wang","year":"2024","unstructured":"Wang J, Huang Y, Chen C, Liu Z, Wang S, Wang Q (2024) Software testing with large language models: Survey, landscape, and vision. IEEE Trans Software Eng 50(4):911\u2013936","journal-title":"IEEE Trans Software Eng"},{"key":"10933_CR65","doi-asserted-by":"crossref","unstructured":"WangJ, Xie X, Hu Q, Liu S, Yu J, Kong J, Li Y (2025a) Defects4c: Benchmarking large language model repair capability with c\/c++ bugs. arXiv:2510.11059","DOI":"10.1109\/ASE63991.2025.00029"},{"key":"10933_CR66","doi-asserted-by":"crossref","unstructured":"Wang P, Zhang L, Liu F, Zhu Y, Xu W, Shi L, Lian X, Li M, Shen B, Fu A (2025b) Efficientedit: Accelerating code editing via edit-oriented speculative decoding. arXiv:2506.02780","DOI":"10.1109\/ASE63991.2025.00215"},{"key":"10933_CR67","doi-asserted-by":"crossref","unstructured":"Wang Q, Sun Z, Wang R, Huang T, Jin Z, Li G, Lyu C (2025c) Semguard: Real-time semantic evaluator for correcting llm-generated code. arXiv:2509.24507","DOI":"10.1109\/ASE63991.2025.00160"},{"key":"10933_CR68","unstructured":"Ye Y, Huang Z, Xiao Y, Chern E, Xia S, Liu P (2025) Limo: Less is more for reasoning. arXiv:2502.03387"},{"key":"10933_CR69","doi-asserted-by":"crossref","unstructured":"Zhang J, Nie P, Li JJ, Gligoric M (2023) Multilingual code co-evolution using large language models. In: Proceedings of the 31st ACM joint european software engineering conference and symposium on the foundations of software engineering, pp 695\u2013707","DOI":"10.1145\/3611643.3616350"},{"key":"10933_CR70","unstructured":"Zhu Q, Guo D, Shao Z, Yang D, Wang P, Xu R, Wu Y, Li Y, Gao H, Ma S (2024a) Deepseek-coder-v2: Breaking the barrier of closed-source models in code intelligence. arXiv:2406.11931"},{"key":"10933_CR71","doi-asserted-by":"crossref","unstructured":"Zhu Q, Liang Q, Sun Z, Xiong Y, Zhang L, Cheng S (2024b) Grammart5: Grammar-integrated pretrained encoder-decoder neural model for code. In: Proceedings of the IEEE\/ACM 46th international conference on software engineering, pp 1\u201313","DOI":"10.1145\/3597503.3639125"},{"key":"10933_CR72","doi-asserted-by":"crossref","unstructured":"Zhu T, Li Z, Pan M, Shi C, Zhang T, Pei Y, Li X (2024c) Deep is better? an empirical comparison of information retrieval and deep learning approaches to code summarization. ACM Trans Softw Eng Methodol 33(3):1\u201337","DOI":"10.1145\/3631975"}],"container-title":["Empirical Software Engineering"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10664-026-10933-0.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s10664-026-10933-0","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10664-026-10933-0.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,7,23]],"date-time":"2026-07-23T08:38:46Z","timestamp":1784795926000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s10664-026-10933-0"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,7,23]]},"references-count":72,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2027,2]]}},"alternative-id":["10933"],"URL":"https:\/\/doi.org\/10.1007\/s10664-026-10933-0","relation":{},"ISSN":["1382-3256","1573-7616"],"issn-type":[{"value":"1382-3256","type":"print"},{"value":"1573-7616","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,7,23]]},"assertion":[{"value":"19 November 2025","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"10 July 2026","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"23 July 2026","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"Not applicable.","order":1,"name":"Ethics","label":"Ethical Approval","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"Not applicable.","order":2,"name":"Ethics","label":"Informed Consent","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"The authors have no conflicts of interest to declare that are relevant to the content of this article.","order":3,"name":"Ethics","label":"Conflict of Interest","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"Not applicable.","order":4,"name":"Ethics","label":"Clinical trial number","group":{"name":"EthicsHeading","label":"Declarations"}}],"article-number":"9"}}