{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,5]],"date-time":"2026-07-05T04:12:19Z","timestamp":1783224739178,"version":"3.54.6"},"reference-count":20,"publisher":"Springer Science and Business Media LLC","issue":"17","license":[{"start":{"date-parts":[[2025,11,17]],"date-time":"2025-11-17T00:00:00Z","timestamp":1763337600000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2025,11,17]],"date-time":"2025-11-17T00:00:00Z","timestamp":1763337600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"DOI":"10.13039\/501100003051","name":"New Energy and Industrial Technology Development Organization","doi-asserted-by":"publisher","award":["JPNP20017"],"award-info":[{"award-number":["JPNP20017"]}],"id":[{"id":"10.13039\/501100003051","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/100007196","name":"Universiti Utara Malaysia","doi-asserted-by":"crossref","id":[{"id":"10.13039\/100007196","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["J Supercomput"],"abstract":"<jats:title>Abstract<\/jats:title>\n                  <jats:p>Large language model (LLM)-based program generation tasks are hindered by high computational demands. These challenges, along with high deployment costs, often pose a barrier to practical applications. To address these, we propose a novel data-augmented multi-LLM model routing approach that classifies prompts based on whether they should be processed on a weak LLM engine or a strong LLM. Experimental results show up to 16 times better efficiency compared to the existing cascaded approaches, while preserving the inference accuracy. Thus, the proposed method optimally allocates prompts across multiple LLMs, reducing computational costs while maintaining inference accuracy.<\/jats:p>","DOI":"10.1007\/s11227-025-08034-8","type":"journal-article","created":{"date-parts":[[2025,11,17]],"date-time":"2025-11-17T12:11:01Z","timestamp":1763381461000},"update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":3,"title":["A data-augmented model routing framework for efficient LLM deployment in edge\u2013cloud environments"],"prefix":"10.1007","volume":"81","author":[{"given":"Muhammad Syafiq Mohd","family":"Pozi","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Yukinori","family":"Sato","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"297","published-online":{"date-parts":[[2025,11,17]]},"reference":[{"key":"8034_CR1","doi-asserted-by":"crossref","unstructured":"Zhang H, Li X, Bing L (2023) Video-llama: an instruction-tuned audio-visual language model for video understanding. In: Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, pp. 543\u2013553","DOI":"10.18653\/v1\/2023.emnlp-demo.49"},{"key":"8034_CR2","doi-asserted-by":"crossref","unstructured":"Jain N, Vaidyanath S, Iyer A, Natarajan N, Parthasarathy S, Rajamani S, Sharma R (2022) Jigsaw: large language models meet program synthesis. In: Proceedings of the 44th International Conference on Software Engineering, ICSE \u201922, p. 1219\u20131231","DOI":"10.1145\/3510003.3510203"},{"key":"8034_CR3","unstructured":"Chen M, Tworek J, Jun H, Yuan Q, de\u00a0Oliveira\u00a0Pinto HP, Kaplan J et\u00a0al (2021) Evaluating large language models trained on code. arXiv:abs\/2107.03374"},{"key":"8034_CR4","doi-asserted-by":"crossref","unstructured":"Majdinasab V, Bishop MJ, Rasheed S, Moradidakhel A, Tahir A, Khomh F (2024) Assessing the security of github copilot\u2019s generated code-a targeted replication study. In: 2024 IEEE International Conference on Software Analysis, Evolution and Reengineering (SANER), pp. 435\u2013444. IEEE","DOI":"10.1109\/SANER60148.2024.00051"},{"key":"8034_CR5","doi-asserted-by":"crossref","unstructured":"Okuda K, Amarasinghe S (2024) Askit: Unified programming interface for programming with large language models. In: Proceedings of the 2024 IEEE\/ACM International Symposium on Code Generation and Optimization, CGO \u201924, p. 41\u201354","DOI":"10.1109\/CGO57630.2024.10444830"},{"key":"8034_CR6","unstructured":"Alliance AR (2024) Ai-ran alliance vision and mission white paper. Techical Report"},{"key":"8034_CR7","unstructured":"Achiam J, Adler S, Agarwal S, Ahmad L, Akkaya I, Aleman FL, Almeida D, Altenschmidt J, Altman S, Anadkat S et\u00a0al (2023) Gpt-4 technical report. arXiv:2303.08774"},{"key":"8034_CR8","unstructured":"Chen L, Zaharia M, Zou J (2023) Frugalgpt: How to use large language models while reducing cost and improving performance. arXiv:2305.05176"},{"key":"8034_CR9","unstructured":"Liu J, Xia CS, Wang Y, Zhang L (2024) Is your code generated by chatgpt really correct? rigorous evaluation of large language models for code generation. Advances in Neural Information Processing Systems. 36"},{"key":"8034_CR10","unstructured":"Chen B, Zhu M, Dolan-Gavitt B, Shafique M, Garg S (2024) Model cascading for code: Reducing inference costs with model cascading for llm based code generation. arXiv:2405.15842"},{"key":"8034_CR11","doi-asserted-by":"crossref","unstructured":"Pozi MSM, Sato Y (2024) A case for deploying dynamic neural network on edge-cloud continuum environment. In: 2024 IEEE International Conference on Edge Computing and Communications (EDGE), pp. 92\u201398. IEEE","DOI":"10.1109\/EDGE62653.2024.00021"},{"key":"8034_CR12","doi-asserted-by":"publisher","unstructured":"Sakota M, Peyrard M, West R (2024) Fly-swat or cannon? Cost-effective language model choice via meta-modeling. In: Proceedings of the 17th ACM International Conference on Web Search and Data Mining, WSDM \u201924, p. 606\u2013615. Association for Computing Machinery, New York, NY, USA. https:\/\/doi.org\/10.1145\/3616855.3635825","DOI":"10.1145\/3616855.3635825"},{"key":"8034_CR13","doi-asserted-by":"crossref","unstructured":"Si C, Shi W, Zhao C, Zettlemoyer L, Boyd-Graber J (2023) Getting more out of mixture of language model reasoning experts. arXiv:2305.14628","DOI":"10.18653\/v1\/2023.findings-emnlp.552"},{"key":"8034_CR14","unstructured":"Liu J, Xia CS, Wang Y, Zhang L (2023) Is your code generated by chatGPT really correct? rigorous evaluation of large language models for code generation. In: Thirty-seventh Conference on Neural Information Processing Systems. https:\/\/openreview.net\/forum?id=1qvx610Cu7"},{"key":"8034_CR15","unstructured":"Guo D, Ren S, Lu S, Feng Z, Tang D, Liu S, Zhou L, Duan N, Svyatkovskiy A, Fu S, Tufano M, Deng SK, Clement C, Drain D, Sundaresan N, Yin J, Jiang D, Zhou M (2021) Graphcodebert: Pre-training code representations with data flow. https:\/\/arxiv.org\/abs\/2009.08366"},{"key":"8034_CR16","doi-asserted-by":"crossref","unstructured":"Zhou S, Alon U, Agarwal S, Neubig G (2023) Codebertscore: Evaluating code generation with pretrained models of code. arXiv:2302.05527","DOI":"10.18653\/v1\/2023.emnlp-main.859"},{"key":"8034_CR17","unstructured":"Liu C, Belkin M (2018) Accelerating sgd with momentum for over-parameterized learning. arXiv:1810.13395"},{"key":"8034_CR18","unstructured":"Feng Z, Hu Y, Yang X, Zhang S et\u00a0al (2020) Codebert: A pre-trained model for programming and natural languages. In: Proceedings of the 43rd International ACM SIGIR Conference on Research and Development in Information Retrieval, pp. 1160\u20131164. ACM"},{"issue":"1","key":"8034_CR19","doi-asserted-by":"publisher","first-page":"101","DOI":"10.1186\/s40537-021-00492-0","volume":"8","author":"C Shorten","year":"2021","unstructured":"Shorten C, Khoshgoftaar TM, Furht B (2021) Text data augmentation for deep learning. J Big Data 8(1):101","journal-title":"J Big Data"},{"key":"8034_CR20","doi-asserted-by":"crossref","unstructured":"Wang Y, Wang W, Joty S, Hoi SC (2021) Codet5: Identifier-aware unified pre-trained encoder-decoder models for code understanding and generation. arXiv:2109.00859","DOI":"10.18653\/v1\/2021.emnlp-main.685"}],"container-title":["The Journal of Supercomputing"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s11227-025-08034-8.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s11227-025-08034-8\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s11227-025-08034-8.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,11,17]],"date-time":"2025-11-17T12:11:06Z","timestamp":1763381466000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s11227-025-08034-8"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,11,17]]},"references-count":20,"journal-issue":{"issue":"17","published-online":{"date-parts":[[2025,11]]}},"alternative-id":["8034"],"URL":"https:\/\/doi.org\/10.1007\/s11227-025-08034-8","relation":{},"ISSN":["1573-0484"],"issn-type":[{"value":"1573-0484","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,11,17]]},"assertion":[{"value":"6 July 2025","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"3 November 2025","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"17 November 2025","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"The authors declare no Conflict of interest.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Conflict of interest"}}],"article-number":"1573"}}