{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,1,14]],"date-time":"2026-01-14T19:58:02Z","timestamp":1768420682097,"version":"3.49.0"},"reference-count":36,"publisher":"MDPI AG","issue":"1","license":[{"start":{"date-parts":[[2026,1,14]],"date-time":"2026-01-14T00:00:00Z","timestamp":1768348800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Computers"],"abstract":"<jats:p>Large language models (LLMs) are increasingly used as rubric-guided graders for short-answer exams, but their decisions can be unstable across prompts and vulnerable to answer-side prompt injection. In this paper, we study SteadyEval, a guardrailed exam-grading pipeline in which an adversarially trained LoRA filter (SteadyEval-7B-deep) preprocesses student answers to remove answer-side prompt injection, after which the original Mistral-7B-Instruct rubric-guided grader assigns the final score. We build two exam-grading pipelines on top of Mistral-7B-Instruct: a baseline pipeline that scores student answers directly, and a guardrailed pipeline in which a LoRA-based filter (SteadyEval-7B-deep) first removes injection content from the answer and a downstream grader then assigns the final score. Using two rubric-guided short-answer datasets in machine learning and computer networking, we generate grouped families of clean answers and four classes of answer-side attacks, and we evaluate the impact of these attacks on score shifts, attack success rates, stability across prompt variants, and alignment with human graders. On the pooled dataset, answer-side attacks inflate grades in the unguarded baseline by an average of about +1.2 points on a 1\u201310 scale, and substantially increase score dispersion across prompt variants. The guardrailed pipeline largely removes this systematic grade inflation and reduces instability for many items, especially in the machine-learning exam, while keeping mean absolute error with respect to human reference scores in a similar range to the unguarded baseline on clean answers, with a conservative shift in networking that motivates per-course calibration. Chief-panel comparisons further show that the guardrailed pipeline tracks human grading more closely on machine-learning items, but tends to under-score networking answers. These findings are best interpreted as a proof-of-concept guardrail and require per-course validation and calibration before operational use.<\/jats:p>","DOI":"10.3390\/computers15010055","type":"journal-article","created":{"date-parts":[[2026,1,14]],"date-time":"2026-01-14T11:01:14Z","timestamp":1768388474000},"page":"55","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":0,"title":["SteadyEval: Robust LLM Exam Graders via Adversarial Training and Distillation"],"prefix":"10.3390","volume":"15","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-1849-3072","authenticated-orcid":false,"given":"Catalin","family":"Anghel","sequence":"first","affiliation":[{"name":"Department of Computer Science and Information Technology, \u201cDun\u0103rea de Jos\u201d University of Galati, \u0218tiin\u021bei St. 2, 800146 Galati, Romania"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0006-9970-2556","authenticated-orcid":false,"given":"Marian Viorel","family":"Craciun","sequence":"additional","affiliation":[{"name":"Department of Computer Science and Information Technology, \u201cDun\u0103rea de Jos\u201d University of Galati, \u0218tiin\u021bei St. 2, 800146 Galati, Romania"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-0935-4713","authenticated-orcid":false,"given":"Adina","family":"Cocu","sequence":"additional","affiliation":[{"name":"Department of Computer Science and Information Technology, \u201cDun\u0103rea de Jos\u201d University of Galati, \u0218tiin\u021bei St. 2, 800146 Galati, Romania"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0008-2537-6713","authenticated-orcid":false,"given":"Andreea Alexandra","family":"Anghel","sequence":"additional","affiliation":[{"name":"Faculty of Automation, Computer Science, Electrical and Electronic Engineering, \u201cDun\u0103rea de Jos\u201d University of Galati, 800008 Galati, Romania"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-1317-6676","authenticated-orcid":false,"given":"Adrian","family":"Istrate","sequence":"additional","affiliation":[{"name":"Department of Computer Science and Information Technology, \u201cDun\u0103rea de Jos\u201d University of Galati, \u0218tiin\u021bei St. 2, 800146 Galati, Romania"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2026,1,14]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","unstructured":"Emirtekin, E. (2025). Large Language Model-Powered Automated Assessment: A Systematic Review. Appl. Sci., 15.","DOI":"10.3390\/app15105683"},{"key":"ref_2","doi-asserted-by":"crossref","unstructured":"Hackl, V., Krainz, A., and Bock, A. (2023). Is GPT-4 a Reliable Rater? Evaluating Consistency in GPT-4\u2019s Text Ratings. Front. Educ., 8.","DOI":"10.3389\/feduc.2023.1272229"},{"key":"ref_3","doi-asserted-by":"crossref","unstructured":"Anghel, C., Craciun, M.V., Pecheanu, E., Cocu, A., Anghel, A.A., Iacobescu, P., Maier, C., Andrei, C.A., Scheau, C., and Dragosloveanu, S. (2025). CourseEvalAI: Rubric-Guided Framework for Transparent and Consistent Evaluation of Large Language Models. Computers, 14.","DOI":"10.3390\/computers14100431"},{"key":"ref_4","unstructured":"Liang, P., Bommasani, R., Lee, T., Tsipras, D., Soylu, D., Yasunaga, M., Zhang, Y., Narayanan, D., Wu, Y., and Kumar, A. (2022). Holistic Evaluation of Language Models. arXiv."},{"key":"ref_5","doi-asserted-by":"crossref","unstructured":"Anghel, C., Anghel, A.A., Pecheanu, E., Cocu, A., and Istrate, A. (2025). Diagnosing Bias and Instability in LLM Evaluation: A Scalable Pairwise Meta-Evaluator. Information, 16.","DOI":"10.3390\/info16080652"},{"key":"ref_6","doi-asserted-by":"crossref","unstructured":"Li, Z., Peng, B., He, P., and Yan, X. (2024, January 12\u201316). Evaluating the Instruction-Following Robustness of Large Language Models to Prompt Injection. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, Miami, FL, USA.","DOI":"10.18653\/v1\/2024.emnlp-main.33"},{"key":"ref_7","doi-asserted-by":"crossref","first-page":"113","DOI":"10.1007\/s10648-023-09823-4","article-title":"Effects of Rubrics on Academic Performance, Self-Regulated Learning, and self-Efficacy: A Meta-analytic Review","volume":"35","author":"Panadero","year":"2023","journal-title":"Educ. Psychol. Rev."},{"key":"ref_8","doi-asserted-by":"crossref","first-page":"1465","DOI":"10.1007\/s40593-024-00440-y","article-title":"Revealing Rubric Relations: Investigating the Interdependence of a Research Informed and a Machine Learning Based Rubric in Assessing Student Reasoning in Chemistry","volume":"35","author":"Martin","year":"2024","journal-title":"Int. J. Artif. Intell. Educ."},{"key":"ref_9","doi-asserted-by":"crossref","unstructured":"Kim, S., and Oh, D. (2025). Evaluating Creativity: Can LLMs Be Good Evaluators in Creative Writing Tasks?. Appl. Sci., 15.","DOI":"10.3390\/app15062971"},{"key":"ref_10","doi-asserted-by":"crossref","unstructured":"Cisneros-Gonz\u00e1lez, J., Gordo-Herrera, N., Barcia-Santos, I., and S\u00e1nchez-Soriano, J. (2025). JorGPT: Instructor-Aided Grading of Programming Assignments with Large Language Models (LLMs). Future Internet, 17.","DOI":"10.3390\/fi17060265"},{"key":"ref_11","doi-asserted-by":"crossref","unstructured":"Hashemi, H., Eisner, J., Rosset, C., Van Durme, B., and Kedzie, C. (2024, January 11\u201316). LLM-RUBRIC: A Multidimensional, Calibrated Approach to Automated Evaluation of Natural Language Texts. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (ACL), Bangkok, Thailand.","DOI":"10.18653\/v1\/2024.acl-long.745"},{"key":"ref_12","unstructured":"Celikyilmaz, A., Clark, E., and Gao, J. (2021). Evaluation of Text Generation: A Survey. arXiv."},{"key":"ref_13","doi-asserted-by":"crossref","unstructured":"Pan, Y., and Nehm, R.H. (2025). Large Language Model and Traditional Machine Learning Scoring of Evolutionary Explanations: Benefits and Drawbacks. Educ. Sci., 15.","DOI":"10.3390\/educsci15060676"},{"key":"ref_14","unstructured":"Gupta, R. (2024, January 23\u201327). Evaluating LLMs: Beyond Simple Metrics. Proceedings of the INLG Workshop & ACL, Tokyo, Japan. Available online: https:\/\/medium.com\/@ritesh.gupta.ai\/evaluating-llms-beyond-simple-metrics-1e6babbed195."},{"key":"ref_15","unstructured":"(2025, July 31). Toloka.ai. LLM  Evaluation Framework: Principles, Practices, and Tools. Available online: https:\/\/toloka.ai\/blog\/llm-evaluation-framework-principles-practices-and-tools."},{"key":"ref_16","doi-asserted-by":"crossref","unstructured":"Zhu, K., Wang, J., Zhou, J., Wang, Z., Chen, H., Wang, Y., Yang, L., Ye, W., Gong, N.Z., and Zhang, Y. (2024). PromptRobust: Towards Evaluating the Robustness of Large Language Models on Adversarial Prompts. arXiv.","DOI":"10.1145\/3689217.3690621"},{"key":"ref_17","doi-asserted-by":"crossref","unstructured":"Moradi, M., and Samwald, M. (2021, January 7\u201311). Evaluating the Robustness of Neural Language Models to Input Perturbations. Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing (EMNLP 2021), Punta Cana, Dominican Republic.","DOI":"10.18653\/v1\/2021.emnlp-main.117"},{"key":"ref_18","doi-asserted-by":"crossref","unstructured":"Fan, Z., Wang, W., Xing, W., and Zhang, D. (2024, January 12\u201316). SedarEval: Automated Evaluation using Self-Adaptive Rubrics. Proceedings of the Findings of the Association for Computational Linguistics: EMNLP 2024, Miami, FL, USA.","DOI":"10.18653\/v1\/2024.findings-emnlp.984"},{"key":"ref_19","doi-asserted-by":"crossref","unstructured":"Anghel, C., Anghel, A.A., Pecheanu, E., Craciun, M.V., Cocu, A., and Niculita, C. (2025). PEARL: A Rubric-Driven Multi-Metric Framework for LLM Evaluation. Information, 16.","DOI":"10.3390\/info16110926"},{"key":"ref_20","doi-asserted-by":"crossref","unstructured":"Anghel, C., Anghel, A.A., Pecheanu, E., Susnea, I., Cocu, A., and Istrate, A. (2025). Multi-Model Dialectical Evaluation of LLM Reasoning Chains: A Structured Framework with Dual Scoring Agents. Informatics, 12.","DOI":"10.3390\/informatics12030076"},{"key":"ref_21","unstructured":"Chaudhary, M., Gupta, H., Bhat, S., and Varma, V. (2024). Towards Understanding the Robustness of LLM-based Evaluations under Perturbations. arXiv."},{"key":"ref_22","doi-asserted-by":"crossref","first-page":"1143","DOI":"10.1162\/tacl_a_00692","article-title":"How Often Are Errors in Natural Language Reasoning Due to Paraphrastic Shifts Rather than Incorrect Logic?","volume":"12","author":"Talmor","year":"2024","journal-title":"Trans. Assoc. Comput. Linguist."},{"key":"ref_23","doi-asserted-by":"crossref","unstructured":"Milani, A., Franzoni, V., Florindi, E., Omarbekova, A., Bekmanova, G., and Yergesh, B. (2025). When AI Is Fooled: Hidden Risks in LLM-Assisted Grading. Educ. Sci., 15.","DOI":"10.3390\/educsci15111419"},{"key":"ref_24","unstructured":"Ferdaus, M.M., Abdelguerfi, M., Ioup, E., Niles, K., Pathak, K., and Sloan, S. (2024). Towards Trustworthy AI: A Review of Ethical and Robust Large Language Models. arXiv."},{"key":"ref_25","doi-asserted-by":"crossref","unstructured":"Choudhury, A., and Chaudhry, Z. (2024). Large Language Models and User Trust: Consequence of Self-Referential Learning Loop and the Deskilling of Health Care Professionals. J. Med. Internet Res., 26.","DOI":"10.2196\/56764"},{"key":"ref_26","doi-asserted-by":"crossref","unstructured":"Anghel, C., Anghel, A.A., Pecheanu, E., Cocu, A., Craciun, M.V., Iacobescu, P., Balau, A.S., and Andrei, C.A. (2025). GraderAssist: A Graph-Based Multi-LLM Framework for Transparent and Reproducible Automated Evaluation. Informatics, 12.","DOI":"10.3390\/informatics12040123"},{"key":"ref_27","doi-asserted-by":"crossref","unstructured":"Gozzi, M., and Di Maio, F. (2024). Comparative Analysis of Prompt Strategies for Large Language Models: Single-Task vs. Multitask Prompts. Electronics, 13.","DOI":"10.20944\/preprints202410.1334.v1"},{"key":"ref_28","unstructured":"Jiang, A.Q., Sablayrolles, A., Mensch, A., Bamford, C., Chaplot, D.S., de las Casas, D., Bressand, F., Lengyel, G., Lample, G., and Saulnier, L. (2023). Mistral 7B. arXiv."},{"key":"ref_29","unstructured":"Neo4j, I. (2025, December 12). Neo4j Graph Database Platform. Available online: https:\/\/neo4j.com\/product\/neo4j-graph-database\/."},{"key":"ref_30","unstructured":"Hu, E.J., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., and Chen, W. (2021). LoRA: Low-Rank Adaptation of Large Language Models. arXiv."},{"key":"ref_31","unstructured":"Project, O.G.S. (2026, January 08). LLM01:2025 Prompt Injection. Available online: https:\/\/genai.owasp.org\/llmrisk\/llm01-prompt-injection\/."},{"key":"ref_32","unstructured":"Centre, N.C.S. (2026, January 08). Prompt Injection Is Not SQL Injection (It May be Worse), Available online: https:\/\/www.ncsc.gov.uk\/blog-post\/prompt-injection-is-not-sql-injection."},{"key":"ref_33","doi-asserted-by":"crossref","unstructured":"Rababah, B., Wu, T.S., Kwiatkowski, M., Leung, C., and Akcora, C.G. (2024). SoK: Prompt Hacking of Large Language Models. arXiv.","DOI":"10.1109\/BigData62323.2024.10825103"},{"key":"ref_34","doi-asserted-by":"crossref","unstructured":"Zahid, F., Sewwandi, A., Brandon, L., Kumar, V., and Sinha, R. (2025). Securing educational LLMs: A generalised taxonomy of attacks on LLMs and DREAD risk assessment. High-Confid. Comput., 100371.","DOI":"10.1016\/j.hcc.2025.100371"},{"key":"ref_35","doi-asserted-by":"crossref","first-page":"153097","DOI":"10.1016\/j.ijhydene.2025.153097","article-title":"The impact of electrolyser allocation on Great Britain\u2019s electricity transmission system in 2050","volume":"202","author":"Giannelos","year":"2026","journal-title":"Int. J. Hydrogen Energy"},{"key":"ref_36","doi-asserted-by":"crossref","unstructured":"Kaloev, M., and Krastev, G. (2023, January 23\u201325). Comprehensive Review of Benefits from the Use of Sparse Updates Techniques in Reinforcement Learning: Experimental Simulations in Complex Action Space Environments. Proceedings of the 2023 4th International Conference on Communications, Information, Electronic and Energy Systems (CIEES), Plovdiv, Bulgaria.","DOI":"10.1109\/CIEES58940.2023.10378830"}],"container-title":["Computers"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/2073-431X\/15\/1\/55\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,1,14]],"date-time":"2026-01-14T11:14:45Z","timestamp":1768389285000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/2073-431X\/15\/1\/55"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,1,14]]},"references-count":36,"journal-issue":{"issue":"1","published-online":{"date-parts":[[2026,1]]}},"alternative-id":["computers15010055"],"URL":"https:\/\/doi.org\/10.3390\/computers15010055","relation":{},"ISSN":["2073-431X"],"issn-type":[{"value":"2073-431X","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,1,14]]}}}