{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,17]],"date-time":"2026-06-17T13:54:47Z","timestamp":1781704487095,"version":"3.54.5"},"reference-count":36,"publisher":"Springer Science and Business Media LLC","issue":"3","license":[{"start":{"date-parts":[[2026,3,27]],"date-time":"2026-03-27T00:00:00Z","timestamp":1774569600000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2026,3,27]],"date-time":"2026-03-27T00:00:00Z","timestamp":1774569600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"name":"Alma Mater Studiorum - Universit\u00e0 di Bologna"}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Int J Softw Tools Technol Transfer"],"published-print":{"date-parts":[[2026,6]]},"abstract":"<jats:title>Abstract<\/jats:title>\n                  <jats:p>We present an empirical study on the ability of Large Language Models (LLMs) to understand code by detecting semantically equivalent and inequivalent programs, that is, whether they compute the same result given the same input or not. To probe this, we deliberately perturb the program text by introducing semantics-preserving code transformations, namely copy propagation and constant folding. Using a benchmark of 11 Python functions with both equivalent and non-equivalent variants, we evaluate seven state-of-the-art LLMs (including ChatGPT, Claude, Gemini, and Deep-Seek) under zero-shot prompting, with and without minimal context. Despite strong performance in code generation tasks, the models often fail in this deeper reasoning challenge, misclassifying 41% of equivalent cases without context and 29% with context. Although prompting can improve performance, it does not address the underlying limitations of the models. We argue that improving LLMs themselves, through targeted fine-tuning, contrastive learning on equivalent and nonequivalent implementations, or training on transformation-invariant code, will be necessary for robust semantic understanding. Meanwhile, practitioners can achieve better results by selecting stronger models, carefully engineering prom-pts, or writing code with tools that normalize low-level differences before inference.<\/jats:p>","DOI":"10.1007\/s10009-026-00842-4","type":"journal-article","created":{"date-parts":[[2026,3,27]],"date-time":"2026-03-27T07:44:51Z","timestamp":1774597491000},"page":"329-343","update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":2,"title":["Understanding code semantics: a benchmark study of LLMs"],"prefix":"10.1007","volume":"28","author":[{"given":"Cosimo","family":"Laneve","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Alvise","family":"Span\u00f2","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Dalila","family":"Ressi","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Sabina","family":"Rossi","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Michele","family":"Bugliesi","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"297","published-online":{"date-parts":[[2026,3,27]]},"reference":[{"key":"842_CR1","volume-title":"Compilers: Principles, Techniques, and Tools","author":"A.V. Aho","year":"2006","unstructured":"Aho, A.V., Lam, M.S., Sethi, R., Ullman, J.D.: Compilers: Principles, Techniques, and Tools, 2nd edn. Addison Wesley, Boston, MA, USA (2006)","edition":"2"},{"key":"842_CR2","unstructured":"Amazon website Services: Amazon q developer website, https:\/\/aws.amazon.com\/it\/q\/developer\/, accessed: 2025-02-21"},{"key":"842_CR3","doi-asserted-by":"crossref","unstructured":"Aydin, F., Aysu, A.: Leaking secrets in homomorphic encryption with side-channel attacks. J. Cryptogr. Eng., 1\u201311 (2024)","DOI":"10.21203\/rs.3.rs-3097727\/v1"},{"key":"842_CR4","unstructured":"Cao, Y., Li, S., Liu, Y., Yan, Z., Dai, Y., Yu, P.S., Sun, L.: A comprehensive survey of ai-generated content (aigc): a history of generative ai from gan to chatgpt (2023). arXiv:2303.04226. arXiv preprint"},{"issue":"2","key":"842_CR5","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1145\/3440755","volume":"54","author":"D. Chandrasekaran","year":"2021","unstructured":"Chandrasekaran, D., Mago, V.: Evolution of semantic similarity\u2014a survey. ACM Comput. Surv. 54(2), 1\u201337 (2021)","journal-title":"ACM Comput. Surv."},{"key":"842_CR6","unstructured":"Chatgpt website. https:\/\/github.com\/features\/copilot, accessed: 2025-02-21"},{"key":"842_CR7","doi-asserted-by":"publisher","DOI":"10.1016\/j.jss.2023.111734","volume":"203","author":"A.M. Dakhel","year":"2023","unstructured":"Dakhel, A.M., Majdinasab, V., Nikanjam, A., Khomh, F., Desmarais, M.C., Jiang, Z.M.J.: Github copilot ai pair programmer: asset or liability? J. Syst. Softw. 203, 111734 (2023)","journal-title":"J. Syst. Softw."},{"issue":"1","key":"842_CR8","doi-asserted-by":"publisher","DOI":"10.2196\/53559","volume":"11","author":"J. Davis","year":"2024","unstructured":"Davis, J., Van Bulck, L., Durieux, B.N., Lindvall, C., et al.: The temperature feature of chatgpt: modifying creativity for clinical research. JMIR Human Factors 11(1), e53559 (2024)","journal-title":"JMIR Human Factors"},{"key":"842_CR9","unstructured":"Dolcetti, G., Arceri, V., Iotti, E., Maffeis, S., Cortesi, A., Zaffanella, E.: Helping llms improve code generation using feedback from testing and static analysis (2024). arXiv:2412.14841. arXiv preprint"},{"key":"842_CR10","unstructured":"Dubey, A., Jauhri, A., Pandey, A., Kadian, A., Al-Dahle, A., Letman, A., Mathur, A., Schelten, A., Yang, A., Fan, A., et\u00a0al.: The llama 3 herd of models (2024). arXiv:2407.21783. arXiv preprint"},{"key":"842_CR11","first-page":"829","volume-title":"33rd USENIX Security Symposium (USENIX Security","author":"C. Fang","year":"2024","unstructured":"Fang, C., Miao, N., Srivastav, S., Liu, J., Zhang, R., Fang, R., Tsang, R., Nazari, N., Wang, H., Homayoun, H., et al.: Large language models for code analysis: do {LLMs} really do their job? In: 33rd USENIX Security Symposium (USENIX Security, vol.\u00a024, pp.\u00a0829\u2013846 (2024)"},{"key":"842_CR12","unstructured":"Han, Z., Gao, C., Liu, J., Zhang, J., Zhang, S.Q.: Parameter-efficient fine-tuning for large models: a comprehensive survey (2024). arXiv:2403.14608. arXiv preprint"},{"key":"842_CR13","doi-asserted-by":"publisher","first-page":"500","DOI":"10.1109\/SANER64311.2025.00053","volume-title":"2025 IEEE International Conference on Software Analysis, Evolution and Reengineering (SANER)","author":"P. He","year":"2025","unstructured":"He, P., Wang, S., Chowdhury, S., Chen, T.H.: Evaluating the effectiveness and efficiency of demonstration retrievers in rag for coding tasks. In: 2025 IEEE International Conference on Software Analysis, Evolution and Reengineering (SANER), pp.\u00a0500\u2013510. IEEE, New York (2025)"},{"key":"842_CR14","doi-asserted-by":"publisher","first-page":"526","DOI":"10.1109\/SANER53432.2022.00070","volume-title":"2022 IEEE International Conference on Software Analysis, Evolution and Reengineering (SANER)","author":"J. Henkel","year":"2022","unstructured":"Henkel, J., Ramakrishnan, G., Wang, Z., Albarghouthi, A., Jha, S., Reps, T.: Semantic robustness of models of source code. In: 2022 IEEE International Conference on Software Analysis, Evolution and Reengineering (SANER), pp.\u00a0526\u2013537. IEEE, New York (2022)"},{"key":"842_CR15","doi-asserted-by":"publisher","DOI":"10.1016\/j.neucom.2024.128645","volume":"610","author":"H. Hu","year":"2024","unstructured":"Hu, H., Wang, X., Zhang, Y., Chen, Q., Guan, Q.: A comprehensive survey on contrastive learning. Neurocomputing 610, 128645 (2024)","journal-title":"Neurocomputing"},{"issue":"1","key":"842_CR16","doi-asserted-by":"publisher","first-page":"2","DOI":"10.3390\/technologies9010002","volume":"9","author":"A. Jaiswal","year":"2020","unstructured":"Jaiswal, A., Babu, A.R., Zadeh, M.Z., Banerjee, D., Makedon, F.: A survey on contrastive self-supervised learning. Technologies 9(1), 2 (2020)","journal-title":"Technologies"},{"key":"842_CR17","unstructured":"Jiang, A.Q., Sablayrolles, A., Roux, A., Mensch, A., Savary, B., Bamford, C., Chaplot, D.S., Casas, D.D.L., Hanna, E.B., Bressand, F., et\u00a0al.: Mixtral of experts (2024). arXiv:2401.04088. arXiv preprint"},{"key":"842_CR18","first-page":"22199","volume":"35","author":"T. Kojima","year":"2022","unstructured":"Kojima, T., Gu, S.S., Reid, M., Matsuo, Y., Iwasawa, Y.: Large language models are zero-shot reasoners. Adv. Neural Inf. Process. Syst. 35, 22199\u201322213 (2022)","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"842_CR19","doi-asserted-by":"publisher","DOI":"10.1016\/j.caeai.2024.100210","volume":"6","author":"E. Latif","year":"2024","unstructured":"Latif, E., Zhai, X.: Fine-tuning chatgpt for automatic scoring. Comput. Educ. Artif. Intell. 6, 100210 (2024)","journal-title":"Comput. Educ. Artif. Intell."},{"issue":"2","key":"842_CR20","first-page":"1","volume":"34","author":"J. Li","year":"2025","unstructured":"Li, J., Li, G., Li, Y., Jin, Z.: Structured chain-of-thought prompting for code generation. ACM Trans. Softw. Eng. Methodol. 34(2), 1\u201323 (2025)","journal-title":"ACM Trans. Softw. Eng. Methodol."},{"key":"842_CR21","unstructured":"Liu, F., Liu, Y., Shi, L., Huang, H., Wang, R., Yang, Z., Zhang, L., Li, Z., Ma, Y.: Exploring and evaluating hallucinations in llm-powered code generation (2024). arXiv:2404.00971. arXiv preprint"},{"key":"842_CR22","doi-asserted-by":"publisher","first-page":"111","DOI":"10.1145\/3323771.3323824","volume-title":"Proceedings of the 2019 7th International Conference on Information and Education Technology","author":"W. Maroengsit","year":"2019","unstructured":"Maroengsit, W., Piyakulpinyo, T., Phonyiam, K., Pongnumkul, S., Chaovalit, P., Theeramunkong, T.: A survey on evaluation methods for chatbots. In: Proceedings of the 2019 7th International Conference on Information and Education Technology, pp.\u00a0111\u2013119 (2019)"},{"key":"842_CR23","unstructured":"OpenAI: Openai chatgpt website, https:\/\/openai.com\/blog\/chatgpt, accessed: 2025-02-21"},{"key":"842_CR24","doi-asserted-by":"publisher","DOI":"10.1016\/j.asoc.2022.109803","volume":"132","author":"L.F.A.O. Pellicer","year":"2023","unstructured":"Pellicer, L.F.A.O., Ferreira, T.M., Costa, A.H.R.: Data augmentation techniques in natural language processing. Appl. Soft Comput. 132, 109803 (2023)","journal-title":"Appl. Soft Comput."},{"key":"842_CR25","doi-asserted-by":"publisher","first-page":"36","DOI":"10.1109\/ICST49551.2021.00016","volume-title":"2021 14th IEEE Conference on Software Testing, Verification and Validation (ICST)","author":"M.V. Pour","year":"2021","unstructured":"Pour, M.V., Li, Z., Ma, L., Hemmati, H.: A search-based testing framework for deep neural networks of source code embedding. In: 2021 14th IEEE Conference on Software Testing, Verification and Validation (ICST), pp.\u00a036\u201346. IEEE, New York (2021)"},{"key":"842_CR26","doi-asserted-by":"crossref","unstructured":"Rabin, M.R.I., Alipour, M.A.: Evaluation of generalizability of neural program analyzers under semantic-preserving transformations (2020). arXiv:2004.07313. arXiv preprint","DOI":"10.1016\/j.infsof.2021.106552"},{"key":"842_CR27","first-page":"75","volume-title":"International Conference of the Italian Association for Artificial Intelligence","author":"D. Ressi","year":"2022","unstructured":"Ressi, D., Romanello, R., Piazza, C., Rossi, S.: Neural networks reduction via lumping. In: International Conference of the Italian Association for Artificial Intelligence, pp.\u00a075\u201390. Springer, Berlin (2022)"},{"key":"842_CR28","doi-asserted-by":"crossref","unstructured":"Ressi, D., Romanello, R., Piazza, C., Rossi, S.: Ai-enhanced blockchain technology: a review of advancements and opportunities. J. Netw. Comput. Appl., 103858 (2024)","DOI":"10.1016\/j.jnca.2024.103858"},{"key":"842_CR29","doi-asserted-by":"crossref","unstructured":"Ressi, D., Romanello, R., Rossi, S., Piazza, C.: Compressing neural networks via formal methods. Neural Netw., 106411 (2024)","DOI":"10.1016\/j.neunet.2024.106411"},{"key":"842_CR30","doi-asserted-by":"crossref","unstructured":"Ressi, D., Span\u00f2, A., Benetollo, L., Piazza, C., Bugliesi, M., Rossi, S.: Vulnerability detection in Ethereum smart contracts via machine learning: a qualitative analysis (2024). arXiv:2407.18639. arXiv preprint","DOI":"10.1016\/j.bcra.2025.100390"},{"key":"842_CR31","unstructured":"Sahoo, P., Singh, A.K., Saha, S., Jain, V., Mondal, S., Chadha, A.: A systematic survey of prompt engineering in large language models: Techniques and applications (2024). arXiv preprint. arXiv:2402.07927"},{"key":"842_CR32","unstructured":"Span\u00f2, A.: Github repository. https:\/\/github.com\/alvisespano\/perturb"},{"key":"842_CR33","first-page":"3687","volume-title":"34th USENIX Security Symposium (USENIX Security","author":"J. Spracklen","year":"2025","unstructured":"Spracklen, J., Wijewickrama, R., Sakib, A.N., Maiti, A., Viswanath, B.: We have a package for you! A comprehensive analysis of package hallucinations by code generating {LLMs}. In: 34th USENIX Security Symposium (USENIX Security, vol.\u00a025, pp.\u00a03687\u20133706 (2025)"},{"key":"842_CR34","unstructured":"Team, G., Mesnard, T., Hardin, C., Dadashi, R., Bhupatiraju, S., Pathak, S., Sifre, L., Rivi\u00e8re, M., Kale, M.S., Love, J., et\u00a0al.: Gemma: Open models based on Gemini research and technology (2024). arXiv preprint. arXiv:2403.08295"},{"key":"842_CR35","unstructured":"Urban, C., Min\u00e9, A.: A review of formal methods applied to machine learning (2021). arXiv:2104.02466. arXiv preprint"},{"key":"842_CR36","first-page":"172","volume-title":"Proceedings of the 54th ACM Technical Symposium on Computer Science Education V","author":"M. Wermelinger","year":"2023","unstructured":"Wermelinger, M.: Using github copilot to solve simple programming problems. In: Proceedings of the 54th ACM Technical Symposium on Computer Science Education V, vol.\u00a01, pp.\u00a0172\u2013178 (2023)"}],"container-title":["International Journal on Software Tools for Technology Transfer"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10009-026-00842-4.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s10009-026-00842-4","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10009-026-00842-4.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,6,17]],"date-time":"2026-06-17T13:05:33Z","timestamp":1781701533000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s10009-026-00842-4"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,3,27]]},"references-count":36,"journal-issue":{"issue":"3","published-print":{"date-parts":[[2026,6]]}},"alternative-id":["842"],"URL":"https:\/\/doi.org\/10.1007\/s10009-026-00842-4","relation":{},"ISSN":["1433-2779","1433-2787"],"issn-type":[{"value":"1433-2779","type":"print"},{"value":"1433-2787","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,3,27]]},"assertion":[{"value":"20 February 2026","order":1,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"27 March 2026","order":2,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}}]}}