{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,12,8]],"date-time":"2025-12-08T01:04:20Z","timestamp":1765155860995,"version":"3.46.0"},"reference-count":25,"publisher":"Springer Science and Business Media LLC","issue":"1","license":[{"start":{"date-parts":[[2025,12,8]],"date-time":"2025-12-08T00:00:00Z","timestamp":1765152000000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2025,12,8]],"date-time":"2025-12-08T00:00:00Z","timestamp":1765152000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"DOI":"10.13039\/100018693","name":"HORIZON EUROPE Framework Programme","doi-asserted-by":"publisher","award":["101120393"],"award-info":[{"award-number":["101120393"]}],"id":[{"id":"10.13039\/100018693","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Cybersecurity"],"abstract":"<jats:title>Abstract<\/jats:title>\n                  <jats:p>\n                    Predicting vulnerability in a code element, such as a function or method, often leverages machine or deep learning models to classify whether it is vulnerable or not. Recently, novel solutions exploiting conversational large language models (LLMs) have emerged, which allow the formulation of the task through a prompt containing natural language elements and the input code element, obtaining a natural language as a response. Although initial promising results, there is currently no broad exploration of (i) how the input prompt influences the prediction capabilities and (ii) what characteristics of the model response relate to correct predictions. In this paper, we conduct an empirical investigation into how accurately two popular conversational LLMs, i.e., GPT-3.5 and Llama-2, predict whether a\n                    <jats:sc>Java<\/jats:sc>\n                    method is vulnerable by employing a thorough prompting strategy by (i) adhering to the\n                    <jats:italic>Zero-Shot<\/jats:italic>\n                    and\n                    <jats:italic>Zero-Shot Chain-of-Thought<\/jats:italic>\n                    techniques and (ii) formulating the prediction task in alternative ways via rephrasing. After a manual inspection of the responses generated, we observed that GPT-3.5 displayed more variable F1 scores compared to Llama-2, which was steadier but often gave no direct classification. ZS prompts achieved F1 scores between 0.53 and 0.69, with a tendency of classifying methods positively (i.e., \u2018vulnerable\u2019); conversely, ZS-CoT presents a broader range of scores, varying from 0.35 to 0.72, with often inconsistencies in the results. Then, we phrased the task in their \u201cinverted form\u201d, i.e., asking the LLM to check for the absence of vulnerabilities, which led to worse results for GPT-3.5, while Llama-2 occasionally performed better. The study further suggests that textual metrics provide important information on LLM outputs. Despite this, these metrics are not correlated with actual outcomes, as the models respond consistently with uniform confidence, irrespective of whether the outcome is correct or not. This underscores the need for customized prompt engineering and response analysis strategies to improve the precision and reliability of LLM-based systems for vulnerability prediction. In addition, we applied our study to two state-of-the-art LLMs, validating the broader applicability of our methodology. Finally, we performed an analysis of various textual properties of the model responses, such as response length and readability scores, to further explore the characteristics of the responses given for vulnerability detection tasks.\n                  <\/jats:p>","DOI":"10.1186\/s42400-025-00476-0","type":"journal-article","created":{"date-parts":[[2025,12,8]],"date-time":"2025-12-08T01:01:58Z","timestamp":1765155718000},"update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":0,"title":["Beyond prompting: the role of phrasing tasks in vulnerability prediction for Java"],"prefix":"10.1186","volume":"8","author":[{"ORCID":"https:\/\/orcid.org\/0000-0001-7489-3540","authenticated-orcid":false,"given":"Torge","family":"Hinrichs","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Emanuele","family":"Iannone","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Riccardo","family":"Scandariato","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"297","published-online":{"date-parts":[[2025,12,8]]},"reference":[{"key":"476_CR1","unstructured":"GitHub: Top 50 Programming Languages Globally . https:\/\/innovationgraph.github.com\/global-metrics\/programming-languages#programming-languages-rankings. Online; accessed 20 May 2025 (2025)"},{"key":"476_CR2","unstructured":"Veracode: Annual Report on the State of Software Security. https:\/\/www.veracode.com\/state-software-security-2024-report\/. Online; accessed 20 May 2025 (2024)"},{"key":"476_CR3","doi-asserted-by":"crossref","unstructured":"Ponta S.E, Plate H, Sabetta A, Bezzi M, Dangremont C (2019) A Manually-Curated Dataset of Fixes to Vulnerabilities of Open-Source Software. arXiv. arXiv:1902.02595 [cs]","DOI":"10.1109\/MSR.2019.00064"},{"key":"476_CR4","doi-asserted-by":"publisher","DOI":"10.1145\/3560815","author":"P Liu","year":"2023","unstructured":"Liu P, Yuan W, Fu J, Jiang Z, Hayashi H, Neubig G (2023) Pre-train, prompt, and predict: A systematic survey of prompting methods in natural language processing. ACM Comput Surv. https:\/\/doi.org\/10.1145\/3560815","journal-title":"ACM Comput Surv"},{"key":"476_CR5","unstructured":"Wang X, Wei J, Schuurmans D, Le Q, Chi E, Narang S, Chowdhery A, Zhou D (2023) Self-Consistency Improves Chain of Thought Reasoning in Language Models. arxiv:2203.11171"},{"key":"476_CR6","unstructured":"Ye X, Durrett G (2022) The unreliability of explanations in few-shot prompting for textual reasoning. In: Koyejo, S., Mohamed, S., Agarwal, A., Belgrave, D., Cho, K., Oh, A. (eds.) Advances in Neural Information Processing Systems, vol. 35, pp. 30378\u201330392. Curran Associates, Inc., ??? . https:\/\/proceedings.neurips.cc\/paper_files\/paper\/2022\/file\/c402501846f9fe03e2cac015b3f0e6b1-Paper-Conference.pdf"},{"key":"476_CR7","unstructured":"Wei J, Wang X, Schuurmans D, Bosma M, Ichter B, Xia F, Chi E, Le Q, Zhou D (2023) Chain-of-Thought Prompting Elicits Reasoning in Large Language Models"},{"key":"476_CR8","unstructured":"Kojima T, Gu S.S, Reid M, Matsuo Y, Iwasawa Y. (2023) Large Language Models are Zero-Shot Reasoners"},{"key":"476_CR9","unstructured":"Steenhoek B, Rahman MM, Roy MK, Alam MS, Barr ET, Le W (2024) A Comprehensive Study of the Capabilities of Large Language Models for Vulnerability Detection . https:\/\/arxiv.org\/abs\/2403.17218"},{"key":"476_CR10","doi-asserted-by":"crossref","unstructured":"Boonthum C (2004) iSTART: Paraphrase recognition. In: Proceedings of the ACL Student Research Workshop, pp. 31\u201336. Association for Computational Linguistics, Barcelona, Spain . https:\/\/aclanthology.org\/P04-2006","DOI":"10.3115\/1219079.1219089"},{"key":"476_CR11","doi-asserted-by":"publisher","unstructured":"Dong Q, Wan X, Cao Y (2021) ParaSCI: A large scientific paraphrase dataset for longer paraphrase generation. In: Merlo, P., Tiedemann, J., Tsarfaty, R. (eds.) Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume, pp. 424\u2013434. Association for Computational Linguistics, Online . https:\/\/doi.org\/10.18653\/v1\/2021.eacl-main.33 . https:\/\/aclanthology.org\/2021.eacl-main.33","DOI":"10.18653\/v1\/2021.eacl-main.33"},{"key":"476_CR12","doi-asserted-by":"crossref","unstructured":"Mallinson J, Sennrich R, Lapata M (2017) Paraphrasing revisited with neural machine translation. In: Lapata M, Blunsom P, Koller A. (eds.) Proceedings of the 15th Conference of the European Chapter of the Association for Computational Linguistics: Volume 1, Long Papers, pp. 881\u2013893. Association for Computational Linguistics, Valencia, Spain . https:\/\/aclanthology.org\/E17-1083","DOI":"10.18653\/v1\/E17-1083"},{"key":"476_CR13","unstructured":"Grammarly: How to Paraphrase (Without Plagiarizing a Thing). website. https:\/\/www.grammarly.com\/blog\/summarizing-paraphrasing\/paraphrase\/ (2024)"},{"key":"476_CR14","doi-asserted-by":"publisher","unstructured":"Narayan S, Gardent C, Cohen SB, Shimorina A (2017) Split and rephrase. In: Palmer, M., Hwa, R., Riedel, S. (eds.) Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing, pp. 606\u2013616. Association for Computational Linguistics, Copenhagen, Denmark . https:\/\/doi.org\/10.18653\/v1\/D17-1064 . https:\/\/aclanthology.org\/D17-1064","DOI":"10.18653\/v1\/D17-1064"},{"key":"476_CR15","doi-asserted-by":"publisher","unstructured":"Aladics T, Heged\u0171s P, Ferenc R (2024) A comparative study of commit representations for jit vulnerability prediction. Computers 13(1) https:\/\/doi.org\/10.3390\/computers13010022","DOI":"10.3390\/computers13010022"},{"key":"476_CR16","doi-asserted-by":"publisher","unstructured":"Zhang C, Liu H, Zeng J, Yang K, Li Y, Li H (2024) Prompt-enhanced software vulnerability detection using chatgpt. In: Proceedings of the 2024 IEEE\/ACM 46th International Conference on Software Engineering: Companion Proceedings. ICSE-Companion \u201924, pp. 276\u2013277. Association for Computing Machinery, New York, NY, USA . https:\/\/doi.org\/10.1145\/3639478.3643065","DOI":"10.1145\/3639478.3643065"},{"issue":"1","key":"476_CR17","doi-asserted-by":"publisher","first-page":"37","DOI":"10.1177\/001316446002000104","volume":"20","author":"J Cohen","year":"1960","unstructured":"Cohen J (1960) A coefficient of agreement for nominal scales. Educ Psychol Measur 20(1):37\u201346. https:\/\/doi.org\/10.1177\/001316446002000104","journal-title":"Educ Psychol Measur"},{"key":"476_CR18","doi-asserted-by":"publisher","unstructured":"McHugh M (2012) Interrater reliability: The kappa statistic. Biochemia medica : \u010dasopis Hrvatskoga dru\u0161tva medicinskih biokemi\u010dara \/ HDMB 22(3), 276\u201382 https:\/\/doi.org\/10.11613\/bm.2012.031","DOI":"10.11613\/bm.2012.031"},{"key":"476_CR19","unstructured":"Powers DMW (2011) Evaluation: From precision, recall and f-measure to roc., informedness, markedness & correlation. Journal of Machine Learning Technologies 2(1), 37\u201363"},{"key":"476_CR20","unstructured":"Baeza-Yates R, Ribeiro-Neto B (1999) Modern Information Retrieval. ACM Press Books. ACM Press, ???"},{"issue":"3","key":"476_CR21","doi-asserted-by":"publisher","first-page":"221","DOI":"10.1037\/h0057532","volume":"32","author":"R Flesch","year":"1948","unstructured":"Flesch R (1948) A new readability yardstick. J Appl Psychol 32(3):221\u2013233","journal-title":"J Appl Psychol"},{"key":"476_CR22","unstructured":"Cheshkov A, Zadorozhny P, Levichev R (2023) Evaluation of chatgpt model for vulnerability detection. arXiv preprint arXiv:2304.07232"},{"key":"476_CR23","doi-asserted-by":"crossref","unstructured":"Jensen RIT, Tawosi V, Alamir S (2024) Software vulnerability and functionality assessment using llms. In: 2024 IEEE\/ACM International Workshop on Natural Language-Based Software Engineering (NLBSE), pp. 25\u201328 . IEEE","DOI":"10.1145\/3643787.3648036"},{"key":"476_CR24","doi-asserted-by":"publisher","unstructured":"Yin X, Ni C, Wang S (2024) Multitask-based evaluation of open-source llm on software vulnerability. IEEE Transactions on Software Engineering, 1\u201316 https:\/\/doi.org\/10.1109\/TSE.2024.3470333","DOI":"10.1109\/TSE.2024.3470333"},{"key":"476_CR25","doi-asserted-by":"crossref","unstructured":"Ni C, Yin X, Yang K, Zhao D, Xing Z, Xia X (2023) Distinguishing Look-Alike Innocent and Vulnerable Code by Subtle Semantic Representation Learning and Explanation . arxiv:2308.11237","DOI":"10.1145\/3611643.3616358"}],"container-title":["Cybersecurity"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1186\/s42400-025-00476-0.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1186\/s42400-025-00476-0\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1186\/s42400-025-00476-0.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,12,8]],"date-time":"2025-12-08T01:02:02Z","timestamp":1765155722000},"score":1,"resource":{"primary":{"URL":"https:\/\/cybersecurity.springeropen.com\/articles\/10.1186\/s42400-025-00476-0"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,12,8]]},"references-count":25,"journal-issue":{"issue":"1","published-online":{"date-parts":[[2025,12]]}},"alternative-id":["476"],"URL":"https:\/\/doi.org\/10.1186\/s42400-025-00476-0","relation":{},"ISSN":["2523-3246"],"issn-type":[{"value":"2523-3246","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,12,8]]},"assertion":[{"value":"15 January 2025","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"1 September 2025","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"8 December 2025","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"The authors declare that they have no competing interests.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Competing interest"}}],"article-number":"111"}}