{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,12]],"date-time":"2026-05-12T19:22:17Z","timestamp":1778613737927,"version":"3.51.4"},"reference-count":57,"publisher":"Springer Science and Business Media LLC","issue":"1","license":[{"start":{"date-parts":[[2025,5,15]],"date-time":"2025-05-15T00:00:00Z","timestamp":1747267200000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2025,5,15]],"date-time":"2025-05-15T00:00:00Z","timestamp":1747267200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Cybersecurity"],"abstract":"<jats:title>Abstract<\/jats:title>\n          <jats:p>With the evolution of modern software development paradigms, component reuse, and low-code approaches have emerged as mainstream in software development. However, developers often lack an in-depth understanding of reused code. The inability of components to operate autonomously leads to insufficient testing of software functionalities and security, further exacerbating the contradiction between the increasing complexity of software architectures and the demand for accurate and efficient software automation testing. This, in turn, increases the frequency of software supply chain security incidents. This paper proposes a test-driven generation framework, LLM4TDG, based on large language models (LLMs). By formally defining the constraint dependency graph and converting it into context constraints, LLMs\u2019 ability to understand natural language descriptions such as test requirements and documents is enhanced. Constraint reasoning and backtracking mechanisms are then used to generate test drivers that satisfy the defined constraints automatically. Using the EvalPlus dataset, we evaluate the comprehensive capabilities of LLM4TDG in test case generation using four general-domain LLMs and five code-generation-domain LLMs. The experimental results indicate that our approach significantly enhances LLMs\u2019 ability to comprehend constraints in testing objectives, achieving a 47.62% increase in constraint understanding across 147 testing tasks. Employing LLM4TDG significantly improves the average pass@k metric of all LLMs by 10.41%. The pass@k metric for CodeQwen-chat has improved by up to 18.66%. The metric surpasses the state-of-the-art GPT-4, with a performance of 92.16% on HUMANEVAL and 87.14% on HUMANEVAL+, which enhances the error correction and functional correctness in test-driven code generation. Meanwhile, Our experiments were conducted on a dataset of Python third-party libraries containing malicious behavior in the context of security testing tasks, validating the effectiveness of our method in real-world applications and its generalization capabilities.<\/jats:p>","DOI":"10.1186\/s42400-024-00335-4","type":"journal-article","created":{"date-parts":[[2025,5,15]],"date-time":"2025-05-15T02:01:51Z","timestamp":1747274511000},"update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":6,"title":["LLM4TDG: test-driven generation of large language models based on enhanced constraint reasoning"],"prefix":"10.1186","volume":"8","author":[{"given":"Jingqiang","family":"Liu","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Ruigang","family":"Liang","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Xiaoxi","family":"Zhu","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yue","family":"Zhang","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yuling","family":"Liu","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Qixu","family":"Liu","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"297","published-online":{"date-parts":[[2025,5,15]]},"reference":[{"key":"335_CR1","unstructured":"Aniculaesei A, Vorwald A, Rausch A (2019) Automated generation of requirements-based test cases for an automotive function using the scade toolchain. In: 11th international conference on adaptive and self-adaptive systems and applications (ADAPTIVE 2019), p 69\u201374"},{"key":"335_CR2","doi-asserted-by":"crossref","unstructured":"Aniculaesei A, Howar F, Denecke P, Rausch A (2018) Automated generation of requirements-based test cases for an adaptive cruise control system. In: 2018 IEEE workshop on validation, analysis and evolution of software tests (VST), IEEE, p 11\u201315","DOI":"10.1109\/VST.2018.8327150"},{"key":"335_CR3","doi-asserted-by":"publisher","unstructured":"Austin J, Odena A, Nye M, Bosma M, Michalewski H, Dohan D, Jiang E, Cai C, Terry M, Le Q, et al. (2021) Program synthesis with large language models. arXiv preprint arXiv:2108.07732 . https:\/\/doi.org\/10.48550\/arXiv.2108.07732","DOI":"10.48550\/arXiv.2108.07732"},{"key":"335_CR4","doi-asserted-by":"crossref","unstructured":"Baliga N (2017) A deep dive into a data-driven world of test. In: 2017 IEEE AUTOTESTCON, IEEE p 1\u20138","DOI":"10.1109\/AUTEST.2017.8080478"},{"key":"335_CR5","doi-asserted-by":"crossref","unstructured":"Bang S, Nam S, Chun I, Jhoo HY, Lee J (2022) Smt-based translation validation for machine learning compiler. In: International conference on computer aided verification, Springer, p 386\u2013407","DOI":"10.1007\/978-3-031-13188-2_19"},{"key":"335_CR6","first-page":"1877","volume":"33","author":"T Brown","year":"2020","unstructured":"Brown T, Mann B, Ryder N, Subbiah M, Kaplan JD, Dhariwal P, Neelakantan A, Shyam P, Sastry G et al (2020) Language models are few-shot learners. Adv Neural Inf Process Syst 33:1877\u20131901","journal-title":"Adv Neural Inf Process Syst"},{"key":"335_CR7","unstructured":"Carlos F, Papadakis M, Durelli V, Delamaro EM (2014) Test data generation techniques for mutation testing: a systematic mapping. In: Workshop on experimental software engineering (ESELAW\u201914), p 419\u2013432"},{"key":"335_CR8","doi-asserted-by":"crossref","unstructured":"Cassano F, Gouwar J, Nguyen D, Nguyen S, Phipps-Costin L, Pinckney D, Yee M-H, Zi Y, Anderson CJ, Feldman MQea (2023) Multipl-e: a scalable and polyglot approach to benchmarking neural code generation. IEEE Transactions on Software Engineering","DOI":"10.1109\/TSE.2023.3267446"},{"key":"335_CR9","unstructured":"Chen M, Tworek J, Jun H, Yuan Q, d Pinto O, HP, Kaplan J, Edwards H, Burda Y, Joseph N, et al, GB (2021) Evaluating large language models trained on code. arXiv preprint arXiv:2107.03374"},{"key":"335_CR10","doi-asserted-by":"crossref","unstructured":"Deng Y, Xia CS, Yang C, Zhang SD, Yang S, Zhang L (2023) Large language models are edge-case fuzzers: testing deep learning libraries via fuzzgpt. arXiv preprint arXiv:2304.02014","DOI":"10.1145\/3597926.3598067"},{"key":"335_CR11","doi-asserted-by":"crossref","unstructured":"Di\u00a0Nardo D, Pastore F, Briand L (2015) Generating complex and faulty test data through model-based mutation analysis. In: 2015 IEEE 8th international conference on software testing, verification and validation (ICST), IEEE, p 1\u201310","DOI":"10.1109\/ICST.2015.7102589"},{"key":"335_CR12","doi-asserted-by":"crossref","unstructured":"Divakaran M, Peddinti DT, LLMs S (2024) LLMs for cyber security: new opportunities. 2404.11338","DOI":"10.1109\/MSEC.2024.3504512"},{"key":"335_CR13","unstructured":"Eichenbaum B (1965) The theory of the formal method. vol. 6"},{"issue":"2","key":"335_CR14","doi-asserted-by":"publisher","first-page":"179","DOI":"10.1207\/s15516709cog1402_1","volume":"14","author":"JL Elman","year":"1990","unstructured":"Elman JL (1990) Finding structure in time. Cognitive Sci 14(2):179\u2013211","journal-title":"Cognitive Sci"},{"key":"335_CR15","unstructured":"Fried D, Aghajanyan A, Lin J, Wang S, Wallace E, Shi F, Zhong R, Yih S, Zettlemoyer L, Lewis M (2023) Incoder: a generative model for code infilling and synthesis. In: The eleventh international conference on learning representations"},{"key":"335_CR16","unstructured":"Green CC (1981) Applications of theorem proving to problem solving. Proc. of IJCAI-69"},{"issue":"1","key":"335_CR17","doi-asserted-by":"publisher","first-page":"317","DOI":"10.1145\/1925844.1926423","volume":"46","author":"S Gulwani","year":"2011","unstructured":"Gulwani S (2011) Automating string processing in spreadsheets using input-output examples. Sigplan Not 46(1):317\u2013330","journal-title":"Sigplan Not"},{"issue":"1\u20132","key":"335_CR18","first-page":"1","volume":"4","author":"S Gulwani","year":"2017","unstructured":"Gulwani S, Polozov O, Singh R (2017) Program synthesis. Found Trends Program Lang 4(1\u20132):1\u2013119","journal-title":"Found Trends Program Lang"},{"key":"335_CR19","doi-asserted-by":"crossref","unstructured":"Guo W, Xu Z, Liu C, Huang C, Fang Y, Liu Y (2023) An empirical study of malicious code in PyPI ecosystem. 2309.11021. https:\/\/arxiv.org\/abs\/2309.11021","DOI":"10.1109\/ASE56229.2023.00135"},{"key":"335_CR20","unstructured":"Holler C, Herzig K, Zeller A (2021) Fuzzing with code fragments. In: 21st USENIX security symposium (USENIX Security 12), USENIX Association, Bellevue, WA, p 445\u2013458"},{"issue":"3","key":"335_CR21","doi-asserted-by":"publisher","first-page":"648","DOI":"10.1016\/j.csi.2013.08.017","volume":"36","author":"I-C Hsu","year":"2014","unstructured":"Hsu I-C, Ting D-H, Hsueh N-L (2014) Mda-based visual modeling approach for resources link relationships using uml profile. Comput Stand Interfaces 36(3):648\u2013656","journal-title":"Comput Stand Interfaces"},{"key":"335_CR22","unstructured":"Jimenez CE, Yang J, Wettig A, Yao S, Pei K, Press O, Narasimhan K (2023) Swe-bench: can language models resolve real-world github issues? arXiv preprint arXiv:2310.06770"},{"key":"335_CR23","unstructured":"Kalyan A, Mohta A, Polozov O, Batra D, Jain P, Gulwani S (2018) Neural-guided deductive search for real-time program synthesis from examples. In: International conference on learning representations"},{"key":"335_CR24","doi-asserted-by":"crossref","unstructured":"Kannavara R, Havlicek CJ, Chen Bea (2015) Challenges and opportunities with concolic testing. In: 2015 national aerospace and electronics conference (NAECON), p 374\u2013378","DOI":"10.1109\/NAECON.2015.7443099"},{"issue":"7","key":"335_CR25","doi-asserted-by":"publisher","first-page":"385","DOI":"10.1145\/360248.360252","volume":"19","author":"JC King","year":"1976","unstructured":"King JC (1976) Symbolic execution and program testing. Commun ACM 19(7):385\u2013394","journal-title":"Commun ACM"},{"issue":"6624","key":"335_CR26","doi-asserted-by":"publisher","first-page":"1092","DOI":"10.1126\/science.abq1158","volume":"378","author":"Y Li","year":"2022","unstructured":"Li Y, Choi D, Chung J, Kushman N, Schrittwieser J, Leblond R, Eccles T, Keeling J, Gimeno F, Aea Dal Lago (2022) Competition-level code generation with alphacode. Science 378(6624):1092\u20131097","journal-title":"Science"},{"issue":"4","key":"335_CR27","first-page":"617","volume":"48","author":"X Liu","year":"2011","unstructured":"Liu X, Xu G, Hu L, Fu X, Dong Y (2011) An approach for constraint-based test data generation in mutation testing. J Comput Res Develop 48(4):617\u2013626","journal-title":"J Comput Res Develop"},{"key":"335_CR28","doi-asserted-by":"crossref","unstructured":"L\u00f3pez-Ib\u00e1\u00f1ez M, St\u00fctzle T, Dorigo M (2015) Ant colony optimization: a component-wise overview","DOI":"10.1007\/978-3-319-07153-4_21-1"},{"key":"335_CR29","doi-asserted-by":"crossref","unstructured":"Lyu Y, Xie Y, Chen P, Chen H (2023) Prompt fuzzing for fuzz driver generation. arXiv preprint arXiv:2312.17677","DOI":"10.1145\/3658644.3670396"},{"key":"335_CR31","first-page":"8887","volume":"975","author":"B Minj","year":"2015","unstructured":"Minj B (2015) Generation of test cases based on analysis of simulink stateflow models. Int J Comput Appl 975:8887","journal-title":"Int J Comput Appl"},{"issue":"3","key":"335_CR32","doi-asserted-by":"publisher","first-page":"341","DOI":"10.1007\/s00766-019-00316-x","volume":"24","author":"A Moitra","year":"2019","unstructured":"Moitra A, Siu K, Crapo AW et al (2019) Automating requirements analysis and test case generation. Requir Eng 24(3):341\u2013364","journal-title":"Requir Eng"},{"key":"335_CR33","doi-asserted-by":"crossref","unstructured":"Mussa M, et al, SO (2009) A survey of model-driven testing techniques. In: Ninth international conference on quality software, p 167\u2013172","DOI":"10.1109\/QSIC.2009.30"},{"key":"335_CR34","unstructured":"Nijkamp E, Pang B, Hayashi H, Tu L, Wang H, Zhou Y, Savarese S, Xiong C (2023) Codegen: an open large language model for code with multi-turn program synthesis. In: The eleventh international conference on learning representations"},{"issue":"2","key":"335_CR35","doi-asserted-by":"publisher","first-page":"167","DOI":"10.1002\/(SICI)1097-024X(199902)29:2<167::AID-SPE225>3.0.CO;2-V","volume":"29","author":"AJ Offutt","year":"1999","unstructured":"Offutt AJ, Jin Z, Pan J (1999) The dynamic domain reduction procedure for test data generation. Softw Pract Exp 29(2):167\u2013193","journal-title":"Softw Pract Exp"},{"key":"335_CR36","unstructured":"Openwall: Follow @Openwall on Twitter for new release announcements and other news (2024). https:\/\/www.openwall.com\/lists\/oss-security\/2024\/03\/29\/4. Accessed 05 Apr 2024"},{"key":"335_CR37","doi-asserted-by":"crossref","unstructured":"Papineni K, Roukos S, Ward T, Zhu W-J (2002) Bleu: a method for automatic evaluation of machine translation. In: Proceedings of the 40th annual meeting of the association for computational linguistics, p 311\u2013318","DOI":"10.3115\/1073083.1073135"},{"issue":"4","key":"335_CR38","doi-asserted-by":"publisher","first-page":"263","DOI":"10.1002\/(SICI)1099-1689(199912)9:4<263::AID-STVR190>3.0.CO;2-Y","volume":"9","author":"R Pargas","year":"1999","unstructured":"Pargas R, Harrold MJ, Peck R (1999) Test-data generation using genetic algorithms. Softw Test Verification Reliab 9(4):263\u2013282","journal-title":"Softw Test Verification Reliab"},{"key":"335_CR39","unstructured":"PyPI: radon (2024). https:\/\/pypi.org\/project\/radon. Accessed 5 Sep 2024"},{"key":"335_CR40","unstructured":"Radford A, Narasimhan K, et al (2018) TS: Improving language understanding by generative pre-training. OpenAI"},{"key":"335_CR41","doi-asserted-by":"crossref","unstructured":"Rani S, Dhawan H, Nagpal G, Suri B (2019) Implementing time-bounded automatic test data generation approach based on search-based mutation testing. In: Progress in advanced computing and intelligent engineering: proceedings of ICACIE 2017, Vol 2. Springer, p 113\u2013122","DOI":"10.1007\/978-981-13-0224-4_11"},{"key":"335_CR42","unstructured":"Schneider S (2021) Using model-based testing for creating behaviour-driven tests. PhD thesis, Wien"},{"key":"335_CR43","unstructured":"Sonatype: 8th annual state of the software supply chain (2022). https:\/\/www.sonatype.com\/state-of-the-software-supply-chain\/open-source-dependency-management-trends-and-recommendations. Accessed 16 Feb 2024-02-16"},{"key":"335_CR44","unstructured":"Synopsys: Open Source Security and Risk Analysis 2023. https:\/\/www.gcomtw.com\/mailshot\/Synopsys\/2303BlackDuck\/repossra2023ch.pdf. Accessed 11 Feb 2024"},{"issue":"4","key":"335_CR45","first-page":"287","volume":"404","author":"GH Tzeng","year":"2011","unstructured":"Tzeng GH, Huang JJ (2011) Multiple attributes decision making - methods and applications. Lecture Notes in Economics Mathematical Systems 404(4):287\u2013288","journal-title":"Lecture Notes in Economics Mathematical Systems"},{"key":"335_CR46","doi-asserted-by":"crossref","unstructured":"Wagner F, Schmuki R, Wagner T, Wolstenholme P (2006) Modeling Software with Finite State Machines: A Practical Approach, New York. Auerbach Publications","DOI":"10.1201\/9781420013641"},{"key":"335_CR47","unstructured":"Wang L, Li X, Zheng G (2005) Research on model-driven software testing. Computer Science"},{"key":"335_CR48","doi-asserted-by":"crossref","unstructured":"Whalen MW, Rajan A, Heimdahl MPEea (2006) Coverage metrics for requirements-based testing. In: Proceedings of the 2006 international symposium on software testing and analysis, pp. 25\u201336","DOI":"10.1145\/1146238.1146242"},{"key":"335_CR49","doi-asserted-by":"crossref","unstructured":"Winterer D, Zhang C, Su Z (2020) On the unusual effectiveness of type-aware operator mutations for testing smt solvers. Proceedings of the ACM on programming languages 4(OOPSLA), pp. 1\u201325","DOI":"10.1145\/3428261"},{"key":"335_CR50","doi-asserted-by":"crossref","unstructured":"Xia CS, Paltenghi M, Tian JL, Pradel M, Zhang L (2024) Universal fuzzing via large language models. In: 46th international conference on software engineering (ICSE)","DOI":"10.1145\/3597503.3639121"},{"key":"335_CR51","doi-asserted-by":"crossref","unstructured":"Xia CS, Wei Y, Zhang L (2023) Automated program repair in the era of large pre-trained language models. In: 2023 IEEE\/ACM 45th international conference on software engineering (ICSE), IEEE, pp. 1482\u20131494","DOI":"10.1109\/ICSE48619.2023.00129"},{"key":"335_CR52","doi-asserted-by":"crossref","unstructured":"Xu FF, Alon U, Neubig G, Hellendoorn VJ (2022) A systematic evaluation of large language models of code. In: Proceedings of the 6th ACM SIGPLAN international symposium on machine programming, pp. 1\u201310","DOI":"10.1145\/3520312.3534862"},{"key":"335_CR53","unstructured":"Xuanjing supply chain security intelligence center: Xuanjing supply chain security intelligence center (2024). https:\/\/www.freebuf.com\/news\/398129.html. Accessed 25 Apr 2024"},{"issue":"3","key":"335_CR54","doi-asserted-by":"publisher","first-page":"17","DOI":"10.3969\/j.issn.1672-528X.2014.11.016","volume":"37","author":"B Yang","year":"2014","unstructured":"Yang B, Wu J, Xu L, Bi K, Liu C (2014) A method for modeling software testing requirements and generating test cases. Chinese J Comput 37(3):17. https:\/\/doi.org\/10.3969\/j.issn.1672-528X.2014.11.016","journal-title":"Chinese J Comput"},{"issue":"04","key":"335_CR55","first-page":"828","volume":"27","author":"X Yao","year":"2016","unstructured":"Yao X, Gong D, Li B (2016) Evolutionary generation of path coverage test data with neural networks. J Softw 27(04):828\u2013838","journal-title":"J Softw"},{"key":"335_CR56","doi-asserted-by":"crossref","unstructured":"Yu T, Zhang R, Yang K, Yasunaga M, Wang D, Li Z, Ma J, Li I, Yao Q, Roman Sea (2018) Spider: a large-scale human-labeled dataset for complex and cross-domain semantic parsing and text-to-sql task. In: Proceedings of the 2018 conference on empirical methods in natural language processing","DOI":"10.18653\/v1\/D18-1425"},{"key":"335_CR57","doi-asserted-by":"crossref","unstructured":"Zhang C, Bai M, Zheng Y, Li Y, Xie X, Li Y, Ma W, Sun L, Liu Y (2023) Understanding large language model based fuzz driver generation. arXiv preprint arXiv:2307.12469","DOI":"10.1145\/3650212.3680355"},{"key":"335_CR58","doi-asserted-by":"crossref","unstructured":"Zheng Q, Xia X, Zou X, Dong Y, Wang S, Xue Y, Wang Z, Shen L, Wang A, Li Yea (2023) Codegeex: A pre-trained model for code generation with multilingual evaluations on humaneval-x. arXiv preprint arXiv:2303.17568","DOI":"10.1145\/3580305.3599790"}],"container-title":["Cybersecurity"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1186\/s42400-024-00335-4.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1186\/s42400-024-00335-4\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1186\/s42400-024-00335-4.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,5,15]],"date-time":"2025-05-15T02:02:32Z","timestamp":1747274552000},"score":1,"resource":{"primary":{"URL":"https:\/\/cybersecurity.springeropen.com\/articles\/10.1186\/s42400-024-00335-4"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,5,15]]},"references-count":57,"journal-issue":{"issue":"1","published-online":{"date-parts":[[2025,12]]}},"alternative-id":["335"],"URL":"https:\/\/doi.org\/10.1186\/s42400-024-00335-4","relation":{},"ISSN":["2523-3246"],"issn-type":[{"value":"2523-3246","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,5,15]]},"assertion":[{"value":"24 June 2024","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"4 November 2024","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"15 May 2025","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"The authors have no Conflict of interest to declare that are relevant to the content of this article.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Competing interests"}}],"article-number":"32"}}