{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,23]],"date-time":"2026-06-23T21:38:16Z","timestamp":1782250696446,"version":"3.54.5"},"reference-count":62,"publisher":"Association for Computing Machinery (ACM)","issue":"4","funder":[{"DOI":"10.13039\/501100001659","name":"Deutsche Forschungsgemeinschaft","doi-asserted-by":"crossref","award":["320435134"],"award-info":[{"award-number":["320435134"]}],"id":[{"id":"10.13039\/501100001659","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Model. Comput. Simul."],"published-print":{"date-parts":[[2025,10,31]]},"abstract":"<jats:p>Formal languages are an integral part of modeling and simulation. They allow the distillation of knowledge into concise simulation models amenable to automatic execution, interpretation, and analysis. However, the arguably most humanly accessible means of expressing models is through natural language, which is not easily interpretable by computers. Here, we evaluate how a Large Language Model (LLM) might be used for formalizing natural language into simulation models. Existing studies only explored using very large LLMs, like the commercial GPT models, without fine-tuning model weights. To close this gap, we show how an open-weights, 7B-parameter Mistral model can be fine-tuned to translate natural language descriptions to reaction network models in a domain-specific language, offering a self-hostable, compute-efficient, and memory efficient alternative. To this end, we develop a synthetic data generator to serve as the basis for fine-tuning and evaluation. Our quantitative evaluation shows that our fine-tuned Mistral model can recover the ground truth simulation model in up to 84.5% of cases. In addition, our small-scale user study demonstrates the model\u2019s practical potential for one-time generation as well as interactive modeling in various domains. While promising, in its current form, the fine-tuned small LLM cannot catch up with large LLMs. We conclude that higher-quality training data are required, and expect future small and open-source LLMs to offer new opportunities.<\/jats:p>","DOI":"10.1145\/3733719","type":"journal-article","created":{"date-parts":[[2025,5,28]],"date-time":"2025-05-28T05:26:17Z","timestamp":1748409977000},"page":"1-27","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":6,"title":["Using (Not-so) Large Language Models to Generate Simulation Models in a Formal DSL: A Study on Reaction Networks"],"prefix":"10.1145","volume":"35","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-4109-3608","authenticated-orcid":false,"given":"Justin Noah","family":"Kreikemeyer","sequence":"first","affiliation":[{"name":"Faculty of Computer Science and Electrical Engineering, Institute for Visual and Analytic Computing, University of Rostock","place":["Rostock, Germany"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0009-9208-4283","authenticated-orcid":false,"given":"Mi\u0142osz","family":"Jankowski","sequence":"additional","affiliation":[{"name":"Faculty of Computer Science and Electrical Engineering, Institute for Visual and Analytic Computing, University of Rostock","place":["Rostock, Germany"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-7447-6667","authenticated-orcid":false,"given":"Pia","family":"Wilsdorf","sequence":"additional","affiliation":[{"name":"Faculty of Computer Science and Electrical Engineering, Institute for Visual and Analytic Computing, University of Rostock","place":["Rostock, Germany"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-5256-4682","authenticated-orcid":false,"given":"Adelinde M.","family":"Uhrmacher","sequence":"additional","affiliation":[{"name":"Faculty of Computer Science and Electrical Engineering, Institute for Visual and Analytic Computing, University of Rostock","place":["Rostock, Germany"]}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2025,9,12]]},"reference":[{"key":"e_1_3_5_2_2","unstructured":"ISO\/IEC 14977:1996(E). 1996. Information Technology \u2013 Syntactic Metalanguage \u2013 Extended BNF. (1996). Retrieved from https:\/\/www.iso.org\/standard\/26153.html"},{"key":"e_1_3_5_3_2","doi-asserted-by":"publisher","DOI":"10.3389\/fsysb.2024.1308292"},{"key":"e_1_3_5_4_2","volume-title":"Proceedings of the 2017 AAAI spring symposium series","author":"Allen James F.","year":"2017","unstructured":"James F. Allen and Choh Man Teng. 2017. Broad coverage, domain-generic deep semantic parsing. In Proceedings of the 2017 AAAI spring symposium series. Association for the Advancement of Artificial Intelligence. Retrieved from https:\/\/cdn.aaai.org\/ocs\/15377\/15377-68204-1-PB.pdf"},{"key":"e_1_3_5_5_2","unstructured":"Jacob Austin Augustus Odena Maxwell Nye Maarten Bosma Henryk Michalewski David Dohan Ellen Jiang Carrie Cai Michael Terry Quoc Le and Charles Sutton. 2021. Program Synthesis with Large Language Models. (2021). arxiv:cs.PL\/2108.07732. Retrieved from https:\/\/arxiv.org\/abs\/2108.07732"},{"key":"e_1_3_5_6_2","doi-asserted-by":"publisher","unstructured":"Nils Baumann Juan Sebastian Diaz Judith Michael Lukas Netz Haron Nqiri Jan Reimer and Bernhard Rumpe. 2024. Combining retrieval-augmented generation and few-shot learning for model synthesis of uncommon DSLs. Modellierung 2024 Satellite Events. DOI:10.18420\/modellierung2024-ws-007","DOI":"10.18420\/modellierung2024-ws-007"},{"key":"e_1_3_5_7_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.jbiotec.2017.06.1200"},{"key":"e_1_3_5_8_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-540-78911-6_2"},{"key":"e_1_3_5_9_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCBBIO.2024.3507781"},{"key":"e_1_3_5_10_2","doi-asserted-by":"publisher","DOI":"10.1145\/3613904.3642377"},{"key":"e_1_3_5_11_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-81-322-3972-7_19"},{"key":"e_1_3_5_12_2","doi-asserted-by":"publisher","DOI":"10.1002\/wsbm.1245"},{"key":"e_1_3_5_13_2","doi-asserted-by":"publisher","DOI":"10.1093\/nar\/gky1049"},{"key":"e_1_3_5_14_2","doi-asserted-by":"publisher","DOI":"10.1177\/003754979506500402"},{"key":"e_1_3_5_15_2","doi-asserted-by":"publisher","unstructured":"Jacob Devlin Ming-Wei Chang Kenton Lee and Kristina Toutanova. 2019. BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies Volume 1 (Long and Short Papers). Minneapolis Minnesota. Association for Computational Linguistics 4171\u20134186. DOI:10.18653\/v1\/N19-1423","DOI":"10.18653\/v1\/N19-1423"},{"key":"e_1_3_5_16_2","unstructured":"OpenAI API Documentation. last accessed 2024-11-12 16:30 UTC+1. GPT-4o. (last accessed 2024-11-12 16:30 UTC+1). Retrieved November 12 2024 from https:\/\/platform.openai.com\/docs\/models#gpt-4o"},{"key":"e_1_3_5_17_2","doi-asserted-by":"publisher","DOI":"10.1145\/3660791"},{"key":"e_1_3_5_18_2","volume-title":"Mathematical Models of Chemical Reactions: Theory and Applications of Deterministic and Stochastic Models","author":"\u00c9rdi P\u00e9ter","year":"1989","unstructured":"P\u00e9ter \u00c9rdi and J\u00e1nos T\u00f3th. 1989. Mathematical Models of Chemical Reactions: Theory and Applications of Deterministic and Stochastic Models. Manchester University Press."},{"key":"e_1_3_5_19_2","unstructured":"DeepSeek-AI Daya Guo Dejian Yang Haowei Zhang Junxiao Song Ruoyu Zhang Runxin Xu Qihao Zhu Shirong Ma Peiyi Wang et al.2025. DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning. (2025). arxiv:cs.CL\/2501.12948. Retrieved from https:\/\/arxiv.org\/abs\/2501.12948"},{"key":"e_1_3_5_20_2","doi-asserted-by":"publisher","unstructured":"Saibo Geng Martin Josifoski Maxime Peyrard and Robert West. 2023. Grammar-constrained decoding for structured NLP tasks without finetuning. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. Singapore. Association for Computational Linguistics 10932\u201310952. DOI:10.18653\/v1\/2023.emnlp-main.674. Source code Retrieved from: https:\/\/github.com\/epfl-dlab\/transformers-CFG","DOI":"10.18653\/v1\/2023.emnlp-main.674"},{"key":"e_1_3_5_21_2","doi-asserted-by":"publisher","DOI":"10.1145\/3664930"},{"key":"e_1_3_5_22_2","doi-asserted-by":"publisher","DOI":"10.1177\/00375497241239360"},{"key":"e_1_3_5_23_2","doi-asserted-by":"publisher","DOI":"10.1057\/s41599-024-03611-3"},{"key":"e_1_3_5_24_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2023.emnlp-main.674"},{"key":"e_1_3_5_25_2","doi-asserted-by":"publisher","unstructured":"P. J. Giabbanelli J. J. Padilla and A. Agrawal. 2024. Broadening access to simulations for end-users via large language models: Challenges and opportunities. 2024 Winter Simulation Conference (WSC). Orlando FL USA. 2535\u20132546. DOI:10.1109\/WSC63780.2024.10838754","DOI":"10.1109\/WSC63780.2024.10838754"},{"key":"e_1_3_5_26_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-642-38874-3_2"},{"key":"e_1_3_5_27_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.coisb.2021.100362"},{"key":"e_1_3_5_28_2","doi-asserted-by":"publisher","DOI":"10.15252\/msb.20177651"},{"key":"e_1_3_5_29_2","doi-asserted-by":"publisher","DOI":"10.1142\/S0129054111008441"},{"key":"e_1_3_5_30_2","doi-asserted-by":"publisher","DOI":"10.1109\/WSC.2007.4419641"},{"key":"e_1_3_5_31_2","unstructured":"Edward J. Hu Yelong Shen Phillip Wallis Zeyuan Allen-Zhu Yuanzhi Li Shean Wang Lu Wang and Weizhu Chen. 2021. LoRA: Low-Rank Adaptation of Large Language Models. (2021). arxiv:cs.CL\/2106.09685. Retrieved from https:\/\/arxiv.org\/abs\/2106.09685"},{"key":"e_1_3_5_32_2","doi-asserted-by":"publisher","DOI":"10.1080\/00207543.2023.2276811"},{"key":"e_1_3_5_33_2","doi-asserted-by":"publisher","DOI":"10.1021\/acs.jpca.0c09316"},{"key":"e_1_3_5_34_2","unstructured":"Albert Q. Jiang Alexandre Sablayrolles Arthur Mensch Chris Bamford Devendra Singh Chaplot Diego de las Casas Florian Bressand Gianna Lengyel Guillaume Lample Lucile Saulnier L\u00e9lio Renard Lavaud Marie-Anne Lachaux Pierre Stock Teven Le Scao Thibaut Lavril Thomas Wang Timoth\u00e9e Lacroix and William El Sayed. 2023. Mistral 7B. (2023). arxiv:cs.CL\/2310.06825. Retrieved from https:\/\/arxiv.org\/abs\/2310.06825"},{"key":"e_1_3_5_35_2","unstructured":"Sathvik Joel Jie J. W. Wu and Fatemeh H. Fard. 2024. A Survey on LLM-based Code Generation for Low-Resource and Domain-Specific Programming Languages. (2024). arxiv:cs.SE\/2410.03981. Retrieved from https:\/\/arxiv.org\/abs\/2410.03981"},{"key":"e_1_3_5_36_2","article-title":"Large language models for pathway curation: A preliminary investigation","author":"Karkera Nikitha","year":"2024","unstructured":"Nikitha Karkera, Nikshita Karkera, Mahanash Kumar, Samik Ghosh, and Sucheendra K. Palaniappan. 2024. Large language models for pathway curation: A preliminary investigation. bioRxiv (2024).","journal-title":"bioRxiv"},{"key":"e_1_3_5_37_2","doi-asserted-by":"publisher","DOI":"10.15252\/msb.20199110"},{"key":"e_1_3_5_38_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-031-26409-2_12"},{"key":"e_1_3_5_39_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-031-71671-3_10"},{"key":"e_1_3_5_40_2","unstructured":"Lisa Lacy. 2024. GPT-4o and Gemini 1.5 Pro: How the New AI Models Compare. (2024). Retrieved from https:\/\/www.cnet.com\/tech\/services-and-software\/gpt-4o-and-gemini-1-5-pro-how-the-new-ai-models-compare\/"},{"key":"e_1_3_5_41_2","first-page":"9459","volume-title":"Proceedings of the Advances in Neural Information Processing Systems","volume":"33","author":"Lewis Patrick","year":"2020","unstructured":"Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich K\u00fcttler, Mike Lewis, Wen-tau Yih, Tim Rockt\u00e4schel, Sebastian Riedel, and Douwe Kiela. 2020. Retrieval-augmented generation for knowledge-intensive NLP tasks. In Proceedings of the Advances in Neural Information Processing Systems, H. Larochelle, M. Ranzato, R. Hadsell, M.F. Balcan, and H. Lin (Eds.), Vol. 33. Curran Associates, Inc., 9459\u20139474. Retrieved from https:\/\/proceedings.neurips.cc\/paper_files\/paper\/2020\/file\/6b493230205f780e1bc26945df7481e5-Paper.pdf"},{"key":"e_1_3_5_42_2","unstructured":"AI @ Meta Llama Team. 2024. The Llama 3 Herd of Models. (2024). arxiv:cs.AI\/2407.21783. Retrieved from https:\/\/arxiv.org\/abs\/2407.21783"},{"key":"e_1_3_5_43_2","unstructured":"Ilya Loshchilov and Frank Hutter. 2019. Decoupled Weight Decay Regularization. (2019). arxiv:cs.LG\/1711.05101. Retrieved from https:\/\/arxiv.org\/abs\/1711.05101"},{"key":"e_1_3_5_44_2","doi-asserted-by":"publisher","DOI":"10.1073\/pnas.6.7.410"},{"key":"e_1_3_5_45_2","doi-asserted-by":"publisher","DOI":"10.3390\/ijms24087296"},{"key":"e_1_3_5_46_2","unstructured":"Julien Martinelli Jeremy Grignard Sylvain Soliman Annabelle Ballesta and Fran\u00e7ois Fages. 2023. Reactmine: A Statistical Search Algorithm for Inferring Chemical Reactions from Time Series Data. (2023). arxiv:q-bio.QM\/2209.03185. Retrieved from https:\/\/arxiv.org\/abs\/2209.03185"},{"key":"e_1_3_5_47_2","doi-asserted-by":"publisher","DOI":"10.1109\/WSC63780.2024.10838967"},{"key":"e_1_3_5_48_2","doi-asserted-by":"publisher","unstructured":"Michael McCloskey and Neal J. Cohen. 1989. Catastrophic interference in connectionist networks: The sequential learning problem. Psychology of Learning and Motivation Vol. 24. Academic Press 109\u2013165. DOI:10.1016\/S0079-7421(08)60536-8","DOI":"10.1016\/S0079-7421(08)60536-8"},{"key":"e_1_3_5_49_2","unstructured":"Tomas Mikolov Kai Chen Greg Corrado and Jeffrey Dean. 2013. Efficient Estimation of Word Representations in Vector Space. (2013). arxiv:cs.CL\/1301.3781. Retrieved from https:\/\/arxiv.org\/abs\/1301.3781"},{"key":"e_1_3_5_50_2","unstructured":"OpenAI. 2024. ChatGPT. Retrieved from https:\/\/chat.openai.com\/. (2024). GPT 3.5."},{"key":"e_1_3_5_51_2","unstructured":"OpenAI. 2024. GPT-4 Technical Report. (2024). arxiv:cs.CL\/2303.08774. Retrieved from https:\/\/arxiv.org\/abs\/2303.08774"},{"key":"e_1_3_5_52_2","unstructured":"OpenAI. 2025. OpenAI Platform Documentation. (2025). Retrieved from https:\/\/platform.openai.com\/docs\/models#gpt-4o"},{"key":"e_1_3_5_53_2","doi-asserted-by":"publisher","DOI":"10.3390\/sym14010119"},{"key":"e_1_3_5_54_2","doi-asserted-by":"publisher","DOI":"10.1109\/WSC.1998.744892"},{"key":"e_1_3_5_55_2","doi-asserted-by":"publisher","DOI":"10.1016\/0360-8352(89)90170-8"},{"key":"e_1_3_5_56_2","doi-asserted-by":"publisher","DOI":"10.1145\/3673226"},{"key":"e_1_3_5_57_2","volume-title":"Proceedings of the Advances in Neural Information Processing Systems","author":"Vaswani A.","year":"2017","unstructured":"A. Vaswani. 2017. Attention is all you need. In Proceedings of the Advances in Neural Information Processing Systems, I. Guyon, U. Von Luxburg, S. Bengio, H. Wallach, R. Fergus, S. Vishwanathan, and R. Garnett (Eds.). Retrieved from https:\/\/proceedings.neurips.cc\/paper_files\/paper\/2017\/file\/3f5ee243547dee91fbd053c1c4a845aa-Paper.pdf"},{"key":"e_1_3_5_58_2","first-page":"31","volume-title":"Proceedings of the Memoria Della Reale Accademia Nazionale dei Lincei","author":"Volterra V.","year":"1926","unstructured":"V. Volterra. 1926. Variazioni e fluttuazioni del numero d\u2019individui in specie animali conviventi. In Proceedings of the Memoria Della Reale Accademia Nazionale dei Lincei. 31\u2013113."},{"key":"e_1_3_5_59_2","first-page":"65030","volume-title":"Proceedings of the Advances in Neural Information Processing Systems","volume":"36","author":"Wang Bailin","year":"2023","unstructured":"Bailin Wang, Zi Wang, Xuezhi Wang, Yuan Cao, Rif A. Saurous, and Yoon Kim. 2023. Grammar prompting for domain-specific language generation with large language models. In Proceedings of the Advances in Neural Information Processing Systems, A. Oh, T. Naumann, A. Globerson, K. Saenko, M. Hardt, and S. Levine (Eds.), Vol. 36. Curran Associates, Inc., 65030\u201365055. Retrieved from https:\/\/proceedings.neurips.cc\/paper_files\/paper\/2023\/file\/cd40d0d65bfebb894ccc9ea822b47fa8-Paper-Conference.pdf"},{"key":"e_1_3_5_60_2","unstructured":"U. Wilensky. 1999. NetLogo. (1999). Retrieved from http:\/\/ccl.northwestern.edu\/netlogo\/"},{"key":"e_1_3_5_61_2","doi-asserted-by":"crossref","unstructured":"Thomas Wolf Lysandre Debut Victor Sanh Julien Chaumond Clement Delangue Anthony Moi Pierric Cistac Tim Rault R\u00e9mi Louf Morgan Funtowicz Joe Davison Sam Shleifer Patrick von Platen Clara Ma Yacine Jernite Julien Plu Canwen Xu Teven Le Scao Sylvain Gugger Mariama Drame Quentin Lhoest and Alexander M. Rush. 2020. HuggingFace\u2019s Transformers: State-of-the-art Natural Language Processing. (2020). arxiv:cs.CL\/1910.03771. Retrieved from https:\/\/arxiv.org\/abs\/1910.03771","DOI":"10.18653\/v1\/2020.emnlp-demos.6"},{"key":"e_1_3_5_62_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.jtbi.2024.111901"},{"key":"e_1_3_5_63_2","unstructured":"Lingling Xu Haoran Xie Si-Zhao Joe Qin Xiaohui Tao and Fu Lee Wang. 2023. Parameter-Efficient Fine-Tuning Methods for Pretrained Language Models: A Critical Review and Assessment. (2023). arxiv:cs.CL\/2312.12148. Retrieved from https:\/\/arxiv.org\/abs\/2312.12148"}],"container-title":["ACM Transactions on Modeling and Computer Simulation"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3733719","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,22]],"date-time":"2025-10-22T13:18:39Z","timestamp":1761139119000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3733719"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,9,12]]},"references-count":62,"journal-issue":{"issue":"4","published-print":{"date-parts":[[2025,10,31]]}},"alternative-id":["10.1145\/3733719"],"URL":"https:\/\/doi.org\/10.1145\/3733719","relation":{},"ISSN":["1049-3301","1558-1195"],"issn-type":[{"value":"1049-3301","type":"print"},{"value":"1558-1195","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,9,12]]},"assertion":[{"value":"2025-04-18","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-04-28","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-09-12","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}