{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,5]],"date-time":"2026-05-05T12:04:16Z","timestamp":1777982656908,"version":"3.51.4"},"reference-count":31,"publisher":"Springer Science and Business Media LLC","issue":"1","license":[{"start":{"date-parts":[[2025,12,18]],"date-time":"2025-12-18T00:00:00Z","timestamp":1766016000000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2025,12,18]],"date-time":"2025-12-18T00:00:00Z","timestamp":1766016000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"DOI":"10.13039\/501100007210","name":"RWTH Aachen University","doi-asserted-by":"crossref","id":[{"id":"10.13039\/501100007210","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Process Sci"],"abstract":"<jats:title>Abstract<\/jats:title>\n                  <jats:p>\n                    Process models are central artefacts in business process management: they drive analysis, automation, simulation, and compliance checking. Yet creating and maintaining high-quality models is labor-intensive and requires expertise in both the domain and formal notations such as Petri nets or BPMN. At the same time, many organisations already document their processes informally through textual work instructions, guidelines, and tickets. Automatically\n                    <jats:italic>generating<\/jats:italic>\n                    formal process models from such natural-language descriptions would therefore accelerate modeling, keep model repositories aligned with documentation, and provide stronger support for downstream uses such as conformance checking and digital twins. In this work, we study\n                    <jats:italic>process model generation<\/jats:italic>\n                    in a narrow, technical sense: given a textual process description, the model must produce a complete, executable process model. This generative capability can then be used both to propose models from scratch and as a building block for process modeling assistance and model completion. Large Language Models (LLMs) pretrained on generic text often struggle with this task, producing syntactically invalid or behaviorally incorrect process models. To address this limitation, we apply Reinforcement Learning (RL) to specialize a pretrained LLM specifically for process model generation. Our RL approach combines automatically verifiable rewards, based on structural checks and behavioral footprints, with universal judgments provided by an LLM-as-a-Judge. We created a dataset of 1312 textual process descriptions with corresponding reference models to support Supervised Fine-Tuning and RL. Experiments demonstrate that RL significantly reduces invalid model generations, improves behavioral correctness, and allows control over model complexity. On the ProMoAI benchmark, the resulting checkpoint approaches the performance of state-of-the-art proprietary models while producing fewer invalid generations.\n                  <\/jats:p>","DOI":"10.1007\/s44311-025-00034-4","type":"journal-article","created":{"date-parts":[[2025,12,18]],"date-time":"2025-12-18T08:36:32Z","timestamp":1766046992000},"update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":0,"title":["Specializing large language models for process modeling via reinforcement learning with verifiable and universal rewards"],"prefix":"10.1007","volume":"2","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-3279-4795","authenticated-orcid":false,"given":"Alessandro","family":"Berti","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Xiaoting","family":"Wang","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-2375-2152","authenticated-orcid":false,"given":"Humam","family":"Kourani","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-0955-6940","authenticated-orcid":false,"given":"Wil M. P.","family":"Van der Aalst","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"297","published-online":{"date-parts":[[2025,12,18]]},"reference":[{"key":"34_CR1","doi-asserted-by":"crossref","unstructured":"Abbasi M, Khadivi M, Ahang M, Lasserre P, Lucet Y, Najjaran H (2025). An innovative data-driven and adaptive reinforcement learning approach for context-aware prescriptive process monitoring. arXiv preprint arXiv:2501.10543","DOI":"10.2139\/ssrn.5125541"},{"key":"34_CR2","unstructured":"Bai Y, Jones A, Ndousse K et al (2022) : training a helpful and harmless assistant with reinforcement learning from human feedback. CoRR abs\/2204.05862"},{"key":"34_CR3","first-page":"610","volume-title":"ICPM workshops. Lecture notes in business information processing","author":"A Berti","year":"2024","unstructured":"Berti A, Kourani H, van der Aalst WMP (2024) Pm-llm-benchmark: evaluating large language models on process mining tasks. In: ICPM workshops. Lecture notes in business information processing, vol 533. Springer, pp 610\u2013623"},{"key":"34_CR4","doi-asserted-by":"publisher","first-page":"100556","DOI":"10.1016\/j.simpa.2023.100556","volume":"17","author":"A Berti","year":"2023","unstructured":"Berti A, van Zelst SJ, Schuster D (2023) Pm4py: a process mining library for python. Softw Impacts 17:100556","journal-title":"Softw Impacts"},{"issue":"15","key":"34_CR5","doi-asserted-by":"publisher","first-page":"6931","DOI":"10.3390\/s23156931","volume":"23","author":"A Bousdekis","year":"2023","unstructured":"Bousdekis A, Kerasiotis A, Kotsias S, Theodoropoulou G, Miaoulis G, Ghazanfarpour D (2023) Modelling and predictive monitoring of business processes under uncertainty with reinforcement learning. Sensors 23(15):6931","journal-title":"Sensors"},{"key":"34_CR6","first-page":"364","volume-title":"CAiSE. Lecture notes in computer science","author":"ZD Bozorgi","year":"2023","unstructured":"Bozorgi ZD, Dumas M, Rosa ML, Polyvyanyy A, Shoush M, Teinemaa I (2023) Learning when to treat business processes: prescriptive process monitoring with causal inference and reinforcement learning. In: CAiSE. Lecture notes in computer science, vol 13901. Springer, pp 364\u2013380"},{"key":"34_CR7","first-page":"197","volume-title":"Business process management workshops. Lecture notes in business information processing","author":"K Brennig","year":"2024","unstructured":"Brennig K, Kaltenpoth S, M\u00fcller O (2024) Straight outta logs: can large language models overcome preprocessing in next event prediction? In: Business process management workshops. Lecture notes in business information processing, vol 534. Springer, pp 197\u2013208"},{"key":"34_CR8","first-page":"55","volume-title":"CAiSE. Lecture notes in computer science","author":"R Cai","year":"2024","unstructured":"Cai R, Zheng C, Wang J, Li D, Wang C, Li B (2024) Reinforcement learning-based streaming process discovery under concept drift. In: CAiSE. Lecture notes in computer science, vol 14663. Springer, pp 55\u201370"},{"key":"34_CR9","unstructured":"DeepSeek-Ai, Guo D, Yang D et al (2025) : deepseek-r1: incentivizing reasoning capability in llms via reinforcement learning. CoRR abs\/2501.12948"},{"key":"34_CR10","first-page":"4571","volume-title":"ACL (1)","author":"S Dou","year":"2024","unstructured":"Dou S, Liu Y, Jia H et al (2024): Stepcoder: improving code generation with reinforcement learning from compiler feedback. In: ACL (1). Association for Computational Linguistics, pp 4571\u20134585"},{"key":"34_CR11","doi-asserted-by":"crossref","unstructured":"El-Aziz EA, Fathalla R, Shaheen M (2023) Deep reinforcement learning for data-efficient weakly supervised business process anomaly detection. J Big Data 10(1)","DOI":"10.1186\/s40537-023-00708-5"},{"key":"34_CR12","unstructured":"Gu J, Jiang X, Shi Z et al (2024): a survey on llm-as-a-judge. CoRR abs\/2411.15594"},{"key":"34_CR13","doi-asserted-by":"publisher","first-page":"102412","DOI":"10.1016\/j.datak.2025.102412","volume":"157","author":"OA Hundogan","year":"2025","unstructured":"Hundogan OA, Verhoef BJ, Theeven P, Reijers HA, Lu X (2025) Reinforcement learning for optimizing responses in care processes. Data Knowl Eng 157:102412","journal-title":"Data Knowl Eng"},{"key":"34_CR14","doi-asserted-by":"crossref","unstructured":"Kourani H, Berti A, Schuster D, van der Aalst WMP (2024a) Evaluating large language models on business process modeling: framework, benchmark, and self-improvement analysis. CoRR abs\/2412.00023","DOI":"10.1007\/s10270-025-01318-w"},{"key":"34_CR15","first-page":"229","volume-title":"BPMDS\/EMMSAD@CAiSE. Lecture notes in business information processing","author":"H Kourani","year":"2024","unstructured":"Kourani H, Berti A, Schuster D, van der Aalst WMP (2024b) Process modeling with large language models. In: BPMDS\/EMMSAD@CAiSE. Lecture notes in business information processing, vol 511. Springer, pp 229\u2013244"},{"key":"34_CR16","first-page":"92","volume-title":"BPM. Lecture notes in computer science","author":"H Kourani","year":"2023","unstructured":"Kourani H, van Zelst SJ (2023) POWL: partially ordered workflow language. In: BPM. Lecture notes in computer science, vol 14159. Springer, pp 92\u2013108"},{"key":"34_CR17","doi-asserted-by":"publisher","first-page":"102493","DOI":"10.1016\/j.is.2024.102493","volume":"128","author":"H Kourani","year":"2025","unstructured":"Kourani H, van Zelst SJ, Schuster D, van der Aalst WMP (2025) Discovering partially ordered workflow models. Inf Syst 128:102493","journal-title":"Inf Syst"},{"key":"34_CR18","unstructured":"Lee H, Phatale S et al H.M.(2024) RLAIF vs. RLHF: scaling reinforcement learning from human feedback with AI feedback. ICML. OpenReview.net"},{"key":"34_CR19","unstructured":"Lee J (2024) Instructpatentgpt: training patent language models to follow instructions with human feedback. CoRR abs\/2406.16897"},{"key":"34_CR20","first-page":"224","volume-title":"AIME (2). Lecture notes in computer science","author":"G Leonardi","year":"2025","unstructured":"Leonardi G, Montani S, Striani M (2025) Exploiting llms for supporting conformance checking on medical processes. In: AIME (2). Lecture notes in computer science, vol 15735. Springer, pp 224\u2013229"},{"key":"34_CR21","doi-asserted-by":"publisher","first-page":"102492","DOI":"10.1016\/j.is.2024.102492","volume":"128","author":"J Middelhuis","year":"2025","unstructured":"Middelhuis J, Bianco RL, Sherzer E, Bukhsh Z, Adan I, Dijkman RM (2025) Learning policies for resource allocation in business processes. Inf Syst 128:102492","journal-title":"Inf Syst"},{"key":"34_CR22","first-page":"44","volume-title":"Business process management workshops. Lecture notes in business information processing","author":"A Norouzifar","year":"2024","unstructured":"Norouzifar A, Kourani H, Dees M, van der Aalst WMP (2024) Bridging domain knowledge and process discovery using large language models. In: Business process management workshops. Lecture notes in business information processing, vol 534. Springer, pp 44\u201356"},{"key":"34_CR23","unstructured":"Ouyang L, Wu J, Jiang X et al (2022) D.A.: training language models to follow instructions with human feedback. NeurIPS"},{"key":"34_CR24","first-page":"9","volume-title":"ICPM","author":"A Rebmann","year":"2024","unstructured":"Rebmann A, Schmidt FD, Glavas G, van der Aa H (2024) Evaluating the ability of llms to solve semantics-aware process mining tasks. In: ICPM. IEEE, pp 9\u201316"},{"key":"34_CR25","unstructured":"Shao Z, Wang P, Zhu Q et al (2024) R.X.: Deepseekmath: pushing the limits of mathematical reasoning in open language models. CoRR abs\/2402.03300"},{"key":"34_CR26","unstructured":"Shoush M, Dumas M (2023) Prescriptive process monitoring under resource constraints: a reinforcement learning approach. CoRR abs\/2307.06564"},{"key":"34_CR27","first-page":"952","volume-title":"IUI","author":"A Szymanski","year":"2025","unstructured":"Szymanski A, Ziems N, Eicher-Miller HA et al (2025) T.J.L.: limitations of the llm-as-a-judge approach for evaluating LLM outputs in expert knowledge tasks. In: IUI. ACM, pp 952\u2013966"},{"key":"34_CR28","first-page":"293","volume-title":"ICPM workshops. Lecture notes in business information processing","author":"J Yuan","year":"2024","unstructured":"Yuan J, Grigori D, van der Aa H (2024) Enhancing predictive process monitoring using semantic information. In: ICPM workshops. Lecture notes in business information processing, vol 533. Springer, pp 293\u2013305"},{"key":"34_CR29","volume-title":"ICML","author":"W Yuan","year":"2024","unstructured":"Yuan W, Pang RY, Cho K et al (2024) X.L.: self-rewarding language models. In: ICML. OpenReview.net"},{"key":"34_CR30","unstructured":"Zhang S, Dong L, Li X et al (2023) S.Z.: instruction tuning for large language models: a survey. CoRR abs\/2308.10792"},{"key":"34_CR31","unstructured":"Zhao Y, Liu H, Yu D et al (2025) S.Y.K.: one token to fool LLM-as-a-Judge. CoRR abs\/2507.08794"}],"container-title":["Process Science"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s44311-025-00034-4.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s44311-025-00034-4\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s44311-025-00034-4.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,12,18]],"date-time":"2025-12-18T08:36:40Z","timestamp":1766047000000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s44311-025-00034-4"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,12,18]]},"references-count":31,"journal-issue":{"issue":"1","published-online":{"date-parts":[[2025,12]]}},"alternative-id":["34"],"URL":"https:\/\/doi.org\/10.1007\/s44311-025-00034-4","relation":{"has-preprint":[{"id-type":"doi","id":"10.36227\/techrxiv.175977593.34948838\/v1","asserted-by":"object"}]},"ISSN":["2948-2178"],"issn-type":[{"value":"2948-2178","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,12,18]]},"assertion":[{"value":"18 September 2025","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"8 December 2025","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"18 December 2025","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"This study does not involve human or animal subjects. Therefore, ethical approval is not applicable.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Ethical approval"}},{"value":"The authors declare no competing interests.","order":3,"name":"Ethics","group":{"name":"EthicsHeading","label":"Competing interests"}}],"article-number":"26"}}