{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,7]],"date-time":"2026-07-07T15:15:15Z","timestamp":1783437315683,"version":"3.54.6"},"reference-count":28,"publisher":"Frontiers Media SA","license":[{"start":{"date-parts":[[2025,10,17]],"date-time":"2025-10-17T00:00:00Z","timestamp":1760659200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":["frontiersin.org"],"crossmark-restriction":true},"short-container-title":["Front. Artif. Intell."],"abstract":"<jats:p>Automated programming has become a powerful tool for solving real-world problems. Code generation, in particular, plays a key role in improving developer productivity and reducing the entry barrier to software development. Recent advances in large language models (LLMs) have significantly improved program synthesis, enabling high-quality code generation from natural language. However, LLMs still struggle with complex tasks, especially in understanding problem intent, conducting multi-step reasoning, and producing code that passes all test cases. As task difficulty increases, existing models often fail to devise complete and reliable generation strategies, leading to reduced accuracy and robustness. To address these limitations, we propose Blueprint2Code, an innovative multi-agent framework for code generation. It emulates the human programming workflow through the coordinated interaction of four agents\u2014Previewing, Blueprint, Coding, and Debugging\u2014forming a closed-loop system from task comprehension to planning, implementation, and iterative refinement. Compared to existing methods, Blueprint2Code shows superior performance on complex programming tasks. Extensive experiments on benchmark datasets\u2014HumanEval, MBPP, their extended versions (HumanEval-ET, MBPP-ET), and the APPS competition dataset\u2014demonstrated its effectiveness, achieving strong pass@1 results: HumanEval 96.3%, MBPP 88.4%, HumanEval-ET 86.5%, MBPP-ET 59.4%, and APPS 24.6%. The related code is available at <jats:ext-link>https:\/\/github.com\/MKH99918\/Blueprint2Code<\/jats:ext-link>.<\/jats:p>","DOI":"10.3389\/frai.2025.1660912","type":"journal-article","created":{"date-parts":[[2025,10,17]],"date-time":"2025-10-17T05:29:37Z","timestamp":1760678977000},"update-policy":"https:\/\/doi.org\/10.3389\/crossmark-policy","source":"Crossref","is-referenced-by-count":1,"title":["Blueprint2Code: a multi-agent pipeline for reliable code generation via blueprint planning and repair"],"prefix":"10.3389","volume":"8","author":[{"given":"Kehao","family":"Mao","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Baokun","family":"Hu","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Ruixin","family":"Lin","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Zewen","family":"Li","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Guanyu","family":"Lu","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Zhengyu","family":"Zhang","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1965","published-online":{"date-parts":[[2025,10,17]]},"reference":[{"key":"B1","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv:2303.08774","article-title":"Gpt-4 technical report","author":"Achiam","year":"2023","journal-title":"arXiv [preprint"},{"key":"B2","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2108.07732","article-title":"Program synthesis with large language models","author":"Austin","year":"2021","journal-title":"arXiv [preprint"},{"key":"B3","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.1707.02275","article-title":"A parallel corpus of python functions and documentation strings for automated code documentation and code generation","author":"Barone","year":"2017","journal-title":"arXiv [preprint"},{"key":"B4","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2401.02954","article-title":"Deepseek llm: Scaling open-source language models with longtermism","author":"Bi","year":"2024","journal-title":"arXiv [preprint"},{"key":"B5","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1145\/3641289","article-title":"A survey on evaluation of large language models","volume":"15","author":"Chang","year":"2024","journal-title":"ACM Trans. Intell. Syst. Technol"},{"key":"B6","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2107.03374","article-title":"Evaluating large language models trained on code","author":"Chen","year":"2021","journal-title":"arXiv [preprint"},{"key":"B7","article-title":"\u201cHuman-inspired episodic memory for infinite context LLMs,\u201d","volume-title":"The Thirteenth International Conference on Learning Representations","author":"Fountas","year":"2025"},{"key":"B8","doi-asserted-by":"publisher","first-page":"2629","DOI":"10.1007\/s10439-023-03272-4","article-title":"Prompt engineering with chatgpt: a guide for academic writers","volume":"51","author":"Giray","year":"2023","journal-title":"Ann. Biomed. Eng"},{"key":"B9","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2401.14196","article-title":"Deepseek-coder: When the large language model meets programming-the rise of code intelligence","author":"Guo","year":"2024","journal-title":"arXiv [preprint"},{"key":"B10","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2105.09938","article-title":"Measuring coding challenge competence with apps","author":"Hendrycks","year":"2021","journal-title":"arXiv [preprint"},{"key":"B11","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2308.08784","article-title":"Codecot: tackling code syntax errors in cot reasoning for code generation","author":"Huang","year":"2023","journal-title":"arXiv [preprint"},{"key":"B12","doi-asserted-by":"publisher","first-page":"933","DOI":"10.12785\/ijcds\/140172","article-title":"An efficient framework for software maintenance cost estimation using genetic hybrid algorithm: oops prospective","volume":"14","author":"Islam","year":"2023","journal-title":"Int. J. Comput. Digit. Syst"},{"key":"B13","first-page":"521","article-title":"\u201cCost estimation model using fifth generation language technique for software maintenance project,\u201d","volume-title":"The International Conference on Recent Innovations in Computing","author":"Islam","year":"2022"},{"key":"B14","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2405.11403","article-title":"Mapcoder: multi-agent code generation for competitive problem solving","author":"Islam","year":"2024","journal-title":"arXiv [preprint"},{"key":"B15","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1145\/3672456","article-title":"Self-planning code generation with large language models","volume":"33","author":"Jiang","year":"2024","journal-title":"ACM Trans. Softw. Eng. Methodol"},{"key":"B16","doi-asserted-by":"publisher","first-page":"1593017","DOI":"10.3389\/frai.2025.1593017","article-title":"Multi-agent systems powered by large language models: applications in swarm intelligence","volume":"8","author":"Jimenez-Romero","year":"2025","journal-title":"Front. Artif. Intell"},{"key":"B17","doi-asserted-by":"publisher","author":"Li","year":"2023","DOI":"10.48550\/arXiv.2305.06161"},{"key":"B18","doi-asserted-by":"publisher","first-page":"1092","DOI":"10.1126\/science.abq1158","article-title":"Competition-level code generation with alphacode","volume":"378","author":"Li","year":"2022","journal-title":"Science"},{"key":"B19","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2404.07143","article-title":"Leave no context behind: Efficient infinite context transformers with infini-attention","author":"Munkhdalai","year":"2024","journal-title":"arXiv [preprint"},{"key":"B20","doi-asserted-by":"publisher","first-page":"106","DOI":"10.1145\/3744746","article-title":"A comprehensive overview of large language models","volume":"16","author":"Naveed","year":"2023","journal-title":"ACM Trans. Intell. Syst. Technol"},{"key":"B21","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.1805.04836","article-title":"Building language models for text with named entities","author":"Parvez","year":"2018","journal-title":"arXiv [preprint"},{"key":"B22","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2401.08500","article-title":"Code generation with alphacodium: from prompt engineering to flow engineering","author":"Ridnik","year":"2024","journal-title":"arXiv [preprint"},{"key":"B23","doi-asserted-by":"publisher","first-page":"3464","DOI":"10.1021\/acs.est.3c01106","article-title":"Risks and benefits of large language models for the environment","volume":"57","author":"Rillig","year":"2023","journal-title":"Environ. Sci. Technol"},{"key":"B24","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2308.12950","article-title":"Code llama: open foundation models for code","author":"Roziere","year":"2023","journal-title":"arXiv [preprint"},{"key":"B25","doi-asserted-by":"publisher","first-page":"8634","DOI":"10.48550\/arXiv.2303.11366","article-title":"Reflexion: language agents with verbal reinforcement learning","volume":"36","author":"Shinn","year":"2023","journal-title":"Adv. Neural Inf. Process. Syst"},{"key":"B26","doi-asserted-by":"publisher","first-page":"24824","DOI":"10.5555\/3600270.3602070","article-title":"Chain-of-thought prompting elicits reasoning in large language models","volume":"35","author":"Wei","year":"2022","journal-title":"Adv. Neural Inf. Process. Syst"},{"key":"B27","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2302.11382","article-title":"A prompt pattern catalog to enhance prompt engineering with chatgpt","author":"White","year":"2023","journal-title":"arXiv [preprint"},{"key":"B28","doi-asserted-by":"crossref","first-page":"5673","DOI":"10.1145\/3580305.3599790","article-title":"\u201cCodegeex: a pre-trained model for code generation with multilingual benchmarking on humaneval-x,\u201d","volume-title":"Proceedings of the 29th ACM SIGKDD Conference on Knowledge Discovery and Data Mining","author":"Zheng","year":"2023"}],"container-title":["Frontiers in Artificial Intelligence"],"original-title":[],"link":[{"URL":"https:\/\/www.frontiersin.org\/articles\/10.3389\/frai.2025.1660912\/full","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,17]],"date-time":"2025-10-17T05:29:41Z","timestamp":1760678981000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.frontiersin.org\/articles\/10.3389\/frai.2025.1660912\/full"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,10,17]]},"references-count":28,"alternative-id":["10.3389\/frai.2025.1660912"],"URL":"https:\/\/doi.org\/10.3389\/frai.2025.1660912","relation":{},"ISSN":["2624-8212"],"issn-type":[{"value":"2624-8212","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,10,17]]},"article-number":"1660912"}}