{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,29]],"date-time":"2026-07-29T19:49:38Z","timestamp":1785354578632,"version":"3.55.0"},"reference-count":36,"publisher":"MDPI AG","issue":"7","license":[{"start":{"date-parts":[[2026,7,21]],"date-time":"2026-07-21T00:00:00Z","timestamp":1784592000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Computers"],"abstract":"<jats:p>Local open-weight large language models (LLMs) are increasingly used in privacy-sensitive settings, yet isolated prompts may not reveal whether safety boundaries remain stable during conversation. SafeBoundary-LLM evaluated seven local models across 14 sensitive domains, 84 boundary sets, 672 single-turn prompts, and 84 five-turn escalation conversations; the same models were evaluated separately on XSTest and JBB-Behaviors. Evaluators R1 and R2 independently classified all 12,194 responses, with R2 labels used for primary outcomes and unreconciled labels used for reliability analysis. Exact agreement exceeded 93% in each dataset. In SafeBoundary-LLM, 456 out of 7644 responses (5.97%) were confirmed-or-mixed failures. The multi-turn failure rate was 14.69% versus 0.51% for single-turn prompts, yielding a rate ratio of 28.80 (95% CI [21.20, 43.80]; Holm-adjusted p = 0.0006); boundary collapse occurred only at Turns 4\u20135, and role-play bypass accounted for 299 out of 456 failures. On answer-expected items, over-refusal was 4.34% in XSTest and 17.29% in JBB-Behaviors, whereas unsafe compliance on refusal-expected items was 0.36% and 0.86%, respectively. These findings support an evaluation strategy that includes public single-turn benchmarks, controlled multi-turn escalation, independent human review, and traceable audit records for locally deployed LLMs.<\/jats:p>","DOI":"10.3390\/computers15070460","type":"journal-article","created":{"date-parts":[[2026,7,21]],"date-time":"2026-07-21T08:14:27Z","timestamp":1784621667000},"page":"460","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":1,"title":["SafeBoundary-LLM: Measuring Safety Boundary Stability in Local Open-Weight LLMs Through Single-Turn Baselines and Multi-Turn Escalation"],"prefix":"10.3390","volume":"15","author":[{"ORCID":"https:\/\/orcid.org\/0009-0008-2537-6713","authenticated-orcid":false,"given":"Andreea Alexandra","family":"Anghel","sequence":"first","affiliation":[{"name":"Department of Computer Science and Information Technology, \u201cDun\u0103rea de Jos\u201d University of Galati, \u0218tiin\u021bei St. 2, 800146 Galati, Romania"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-1849-3072","authenticated-orcid":false,"given":"Catalin","family":"Anghel","sequence":"additional","affiliation":[{"name":"Department of Computer Science and Information Technology, \u201cDun\u0103rea de Jos\u201d University of Galati, \u0218tiin\u021bei St. 2, 800146 Galati, Romania"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-1155-5274","authenticated-orcid":false,"given":"Emilia","family":"Pecheanu","sequence":"additional","affiliation":[{"name":"Department of Computer Science and Information Technology, \u201cDun\u0103rea de Jos\u201d University of Galati, \u0218tiin\u021bei St. 2, 800146 Galati, Romania"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0002-5261-2060","authenticated-orcid":false,"given":"Antonio Stefan","family":"Balau","sequence":"additional","affiliation":[{"name":"Department of Computer Science and Information Technology, \u201cDun\u0103rea de Jos\u201d University of Galati, \u0218tiin\u021bei St. 2, 800146 Galati, Romania"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0006-9970-2556","authenticated-orcid":false,"given":"Marian Viorel","family":"Craciun","sequence":"additional","affiliation":[{"name":"Department of Computer Science and Information Technology, \u201cDun\u0103rea de Jos\u201d University of Galati, \u0218tiin\u021bei St. 2, 800146 Galati, Romania"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-0935-4713","authenticated-orcid":false,"given":"Adina","family":"Cocu","sequence":"additional","affiliation":[{"name":"Department of Computer Science and Information Technology, \u201cDun\u0103rea de Jos\u201d University of Galati, \u0218tiin\u021bei St. 2, 800146 Galati, Romania"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0006-6245-3621","authenticated-orcid":false,"given":"Cristian","family":"Sandu","sequence":"additional","affiliation":[{"name":"Department of Computer Science and Information Technology, \u201cDun\u0103rea de Jos\u201d University of Galati, \u0218tiin\u021bei St. 2, 800146 Galati, Romania"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1968","published-online":{"date-parts":[[2026,7,21]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","first-page":"175","DOI":"10.1007\/s10462-024-10824-0","article-title":"A survey of safety and trustworthiness of large language models through the lens of verification and validation","volume":"57","author":"Huang","year":"2024","journal-title":"Artif. Intell. Rev."},{"key":"ref_2","first-page":"1","article-title":"Security and Privacy Challenges of Large Language Models: A Survey","volume":"57","author":"Das","year":"2025","journal-title":"ACM Comput. Surv."},{"key":"ref_3","doi-asserted-by":"crossref","unstructured":"R\u00f6ttger, P., Kirk, H.R., Vidgen, B., Attanasio, G., Bianchi, F., and Hovy, D. (2024). XSTest: A Test Suite for Identifying Exaggerated Safety Behaviours in Large Language Models. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Mexico City, Mexico, 2024, Association for Computational Linguistics.","DOI":"10.18653\/v1\/2024.naacl-long.301"},{"key":"ref_4","doi-asserted-by":"crossref","unstructured":"Chao, P., Debenedetti, E., Robey, A., Andriushchenko, M., Croce, F., Sehwag, V., Dobriban, E., Flammarion, N., Pappas, G.J., and Tram\u00e8r, F. (2024, January 10\u201315). JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models. Proceedings of the Advances in Neural Information Processing Systems, Vancouver, BC, Canada.","DOI":"10.52202\/079017-1745"},{"key":"ref_5","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/3641289","article-title":"A Survey on Evaluation of Large Language Models","volume":"15","author":"Chang","year":"2024","journal-title":"ACM Trans. Intell. Syst. Technol."},{"key":"ref_6","doi-asserted-by":"crossref","unstructured":"Billah, M.M., Hamjaya, H.S., Shiralizade, H., Singh, V., and Inam, R. (2025). Large Language Models\u2019 Trustworthiness in the Light of the EU AI Act\u2014A Systematic Mapping Study. Appl. Sci., 15.","DOI":"10.3390\/app15147640"},{"key":"ref_7","doi-asserted-by":"crossref","unstructured":"Tyndall, E., Wagner, T., Gayheart, C., Some, A., and Langhals, B. (2025). Feasibility Evaluation of Secure Offline Large Language Models with Retrieval-Augmented Generation for CPU-Only Inference. Information, 16.","DOI":"10.3390\/info16090744"},{"key":"ref_8","doi-asserted-by":"crossref","first-page":"382","DOI":"10.1007\/s10462-025-11389-2","article-title":"Safeguarding Large Language Models: A Survey","volume":"58","author":"Dong","year":"2025","journal-title":"Artif. Intell. Rev."},{"key":"ref_9","doi-asserted-by":"crossref","first-page":"100301","DOI":"10.1016\/j.jnlest.2025.100301","article-title":"On large language models safety, security, and privacy: A survey","volume":"23","author":"Zhang","year":"2025","journal-title":"J. Electron. Sci. Technol."},{"key":"ref_10","doi-asserted-by":"crossref","first-page":"130","DOI":"10.1007\/s10994-026-07060-8","article-title":"Survey on LLM Safety: Attacks, Defenses, Alignment, Metrics, and Guardrails","volume":"115","author":"Jalan","year":"2026","journal-title":"Mach. Learn."},{"key":"ref_11","doi-asserted-by":"crossref","unstructured":"Pribisali\u0107, M., and Martin\u010di\u0107-Ip\u0161i\u0107, S. (2026). Security and Privacy of Large Language Models: Threat Taxonomy, Ethical Implications, and Governance. AI, 7.","DOI":"10.3390\/ai7050152"},{"key":"ref_12","doi-asserted-by":"crossref","first-page":"687","DOI":"10.1613\/jair.1.17654","article-title":"Against The Achilles\u2019 Heel: A Survey on Red Teaming for Generative Models","volume":"82","author":"Lin","year":"2025","journal-title":"J. Artif. Intell. Res."},{"key":"ref_13","doi-asserted-by":"crossref","unstructured":"Wang, Y., Li, H., Han, X., Nakov, P., and Baldwin, T. (2024). Do-Not-Answer: Evaluating Safeguards in LLMs. Proceedings of the Findings of the Association for Computational Linguistics: EACL 2024, St. Julian\u2019s, Malta, 2024, Association for Computational Linguistics.","DOI":"10.18653\/v1\/2024.findings-eacl.61"},{"key":"ref_14","unstructured":"Xie, T., Qi, X., Zeng, Y., Huang, Y., Sehwag, U.M., Huang, K., He, L., Wei, B., Li, D., and Sheng, Y. (2025, January 24\u201328). SORRY-Bench: Systematically Evaluating Large Language Model Safety Refusal. Proceedings of the 13th International Conference on Learning Representations, Singapore."},{"key":"ref_15","unstructured":"Cui, J., Chiang, W.-L., Stoica, I., and Hsieh, C.-J. (2025). OR-Bench: An Over-Refusal Benchmark for Large Language Models. Proceedings of the 42nd International Conference on Machine Learning, 2025, PMLR."},{"key":"ref_16","unstructured":"Mazeika, M., Phan, L., Yin, X., Zou, A., Wang, Z., Mu, N., Sakhaee, E., Li, N., Basart, S., and Li, B. (2024). HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal. Proceedings of the 41st International Conference on Machine Learning, Vienna, Austria, 2024, MLResearch Press."},{"key":"ref_17","doi-asserted-by":"crossref","unstructured":"Anghel, C., Anghel, A.A., Pecheanu, E., Cocu, A., and Istrate, A. (2025). Diagnosing Bias and Instability in LLM Evaluation: A Scalable Pairwise Meta-Evaluator. Information, 16.","DOI":"10.3390\/info16080652"},{"key":"ref_18","doi-asserted-by":"crossref","unstructured":"Anghel, C., Craciun, M.V., Cocu, A., Anghel, A.A., Balau, A.S., Istrate, A., and Anghele, A.-D. (2026). EvalHack: Answer-Side Prompt Injection for Probing LLM Exam-Grading Panel Stability. Information, 17.","DOI":"10.3390\/info17030297"},{"key":"ref_19","doi-asserted-by":"crossref","unstructured":"Guo, W., Li, J., Wang, W., Li, Y., He, D., Yu, J., and Zhang, M. (2025). MTSA: Multi-turn Safety Alignment for LLMs through Multi-round Red-teaming. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics, Vienna, Austria, 2025, Association for Computational Linguistics.","DOI":"10.18653\/v1\/2025.acl-long.1282"},{"key":"ref_20","unstructured":"Neo4j Inc. (2026, June 18). Neo4j Graph Database Platform. Available online: https:\/\/neo4j.com\/product\/neo4j-graph-database\/."},{"key":"ref_21","unstructured":"Ollama Inc. (2026, June 19). Ollama API Documentation. Available online: https:\/\/docs.ollama.com\/api\/introduction."},{"key":"ref_22","unstructured":"Alibaba Cloud (2026, June 18). Qwen3-8B Model Card. Available online: https:\/\/huggingface.co\/Qwen\/Qwen3-8B."},{"key":"ref_23","unstructured":"Microsoft (2026, June 24). Phi-4 Model Card. Available online: https:\/\/huggingface.co\/microsoft\/phi-4."},{"key":"ref_24","unstructured":"Google (2026, June 19). Gemma-3-12B-IT Model Card. Available online: https:\/\/huggingface.co\/google\/gemma-3-12b-it."},{"key":"ref_25","unstructured":"Mistral AI (2026, June 19). Mistral Small 3.2 Model Card. Available online: https:\/\/docs.mistral.ai\/models\/model-cards\/mistral-small-3-2-25-06."},{"key":"ref_26","unstructured":"Microsoft (2026, June 19). Phi-4-Mini-Instruct Model Card. Available online: https:\/\/huggingface.co\/microsoft\/Phi-4-mini-instruct."},{"key":"ref_27","unstructured":"Meta-AI. (2026, June 19). Llama-3.2-3B-Instruct Model Card. Available online: https:\/\/huggingface.co\/meta-llama\/Llama-3.2-3B-Instruct."},{"key":"ref_28","unstructured":"(2026, June 19). Meta. llama3.3:70b-Instruct-q4_K_M. Available online: https:\/\/ollama.com\/library\/llama3.3:70b-instruct-q4_K_M."},{"key":"ref_29","doi-asserted-by":"crossref","unstructured":"Anghel, C., Anghel, A.A., Pecheanu, E., Cocu, A., Craciun, M.V., Iacobescu, P., Balau, A.S., and Andrei, C.A. (2025). GraderAssist: A Graph-Based Multi-LLM Framework for Transparent and Reproducible Automated Evaluation. Informatics, 12.","DOI":"10.3390\/informatics12040123"},{"key":"ref_30","doi-asserted-by":"crossref","unstructured":"Jaffal, N.O., Alkhanafseh, M., and Mohaisen, D. (2025). Large Language Models in Cybersecurity: A Survey of Applications, Vulnerabilities, and Defense Techniques. AI, 6.","DOI":"10.3390\/ai6090216"},{"key":"ref_31","doi-asserted-by":"crossref","unstructured":"Lazo Vera, L.E., Jelodar, H., and Razavi-Far, R. (2026). LLM Security and Safety: Insights from Homotopy-Inspired Prompt Obfuscation. AI, 7.","DOI":"10.3390\/ai7030083"},{"key":"ref_32","doi-asserted-by":"crossref","unstructured":"Sokhansanj, B.A. (2025). Local AI Governance: Addressing Model Safety and Policy Challenges Posed by Decentralized AI. AI, 6.","DOI":"10.3390\/ai6070159"},{"key":"ref_33","doi-asserted-by":"crossref","unstructured":"Podpora, M., Baranowski, M., Chopcian, M., Kwasniewicz, L., and Radziewicz, W. (2026). LLM Firewall Using Validator Agent for Prevention Against Prompt Injection Attacks. Appl. Sci., 16.","DOI":"10.3390\/app16010085"},{"key":"ref_34","doi-asserted-by":"crossref","unstructured":"Muhammad, A.E., and Yow, K.-C. (2026). Risk-Based AI Assurance Framework. Information, 17.","DOI":"10.3390\/info17030263"},{"key":"ref_35","doi-asserted-by":"crossref","unstructured":"Anghel, C., Anghel, A.A., Craciun, M.V., Cocu, A., Vulpe, D.-E., Andrei, C.A., Maier, C., Scheau, C., Dragosloveanu, S., and Cergan, R. (2026). GradeAgentOps: A Verification-First Framework for Evidence-Anchored LLM Exam Grading. AI, 7.","DOI":"10.3390\/ai7060198"},{"key":"ref_36","doi-asserted-by":"crossref","unstructured":"Hinov, N., and Ivanova, M. (2026). LLM-Augmented Algorithmic Management: A Governance-Oriented Architecture for Explainable Organizational Decision Systems. AI, 7.","DOI":"10.3390\/ai7030102"}],"container-title":["Computers"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/2073-431X\/15\/7\/460\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,7,21]],"date-time":"2026-07-21T08:28:59Z","timestamp":1784622539000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/2073-431X\/15\/7\/460"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,7,21]]},"references-count":36,"journal-issue":{"issue":"7","published-online":{"date-parts":[[2026,7]]}},"alternative-id":["computers15070460"],"URL":"https:\/\/doi.org\/10.3390\/computers15070460","relation":{},"ISSN":["2073-431X"],"issn-type":[{"value":"2073-431X","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,7,21]]}}}