{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,26]],"date-time":"2026-06-26T05:46:28Z","timestamp":1782452788793,"version":"3.54.5"},"reference-count":57,"publisher":"MDPI AG","issue":"6","license":[{"start":{"date-parts":[[2026,6,18]],"date-time":"2026-06-18T00:00:00Z","timestamp":1781740800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"DOI":"10.13039\/100015539","name":"Australian Government","doi-asserted-by":"crossref","award":["GA221786"],"award-info":[{"award-number":["GA221786"]}],"id":[{"id":"10.13039\/100015539","id-type":"DOI","asserted-by":"crossref"}]},{"award":["GA221786"],"award-info":[{"award-number":["GA221786"]}],"id":[{"id":"https:\/\/ror.org\/0314h5y94","id-type":"ROR","asserted-by":"publisher"}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Computers"],"abstract":"<jats:p>Generative large language models (LLMs) are increasingly used to support writing feedback. However, the pedagogical safety and usefulness of LLM feedback for primary students remains under-evaluated. This study reports an educator-centered evaluation of GPT-4 Turbo for Year 5 narrative and persuasive writing in the context of an established online tutoring program. Using authentic students\u2019 drafts paired with tutor feedback, we generated parallel LLM feedback via rubric-aligned prompting and compared the two feedback sources in a blinded, within-script design. Four experienced English specialists co-designed a six-dimensional rubric (clarity, specificity, helpfulness, feasibility, relevance, and overall effectiveness) and rated tutor versus LLM feedback for each script; their written reflections were analyzed thematically to surface boundary conditions and risk perceptions. Across dimensions, tutor feedback received slightly higher mean ratings, with the clearest descriptive advantage in perceived helpfulness; however, none of the differences remained statistically significant after Holm-Bonferroni correction. LLM feedback was often rated similarly for clarity and feasibility but was frequently characterized as generic, surface-focused, and occasionally misaligned with the student draft, which increased verification effort and posed a risk of misleading learners if used without mediation. Synthesizing ratings and educator reflections, we identify conditions under which LLM feedback is most appropriate as rapid first-pass support for routine structure and surface revision, and least appropriate for developmental judgment and context-sensitive guidance. We translate these findings into design requirements for teacher-in-the-loop primary writing feedback systems, including alignment to explicit pedagogical constructs, editable workflows, and safeguards that reduce unsupported feedback before release to students.<\/jats:p>","DOI":"10.3390\/computers15060393","type":"journal-article","created":{"date-parts":[[2026,6,19]],"date-time":"2026-06-19T13:04:04Z","timestamp":1781874244000},"page":"393","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":0,"title":["Boundary Conditions for LLM-Generated Feedback in Primary Writing: An Educator-Aligned Evaluation and Design Considerations"],"prefix":"10.3390","volume":"15","author":[{"ORCID":"https:\/\/orcid.org\/0009-0003-9299-8444","authenticated-orcid":false,"given":"Dan","family":"Zhang","sequence":"first","affiliation":[{"name":"Faculty of Science, Engineering and Built Environment, School of Information Technology, Deakin University, Burwood, VIC 3125, Australia"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-7354-260X","authenticated-orcid":false,"given":"Thuong","family":"Hoang","sequence":"additional","affiliation":[{"name":"Faculty of Science, Engineering and Built Environment, School of Information Technology, Deakin University, Burwood, VIC 3125, Australia"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-4776-4932","authenticated-orcid":false,"given":"Ye","family":"Zhu","sequence":"additional","affiliation":[{"name":"Faculty of Science, Engineering and Built Environment, School of Information Technology, Deakin University, Burwood, VIC 3125, Australia"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-6600-6937","authenticated-orcid":false,"given":"Rui","family":"Wang","sequence":"additional","affiliation":[{"name":"Commonwealth Scientific and Industrial Research Organisation(CSIRO), Research Way, Clayton, VIC 3168, Australia"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Paula","family":"Crouch","sequence":"additional","affiliation":[{"name":"Kinetic Education, 506 Nepean Hwy, Frankston, VIC 3199, Australia"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-9975-3912","authenticated-orcid":false,"given":"Yi","family":"Wang","sequence":"additional","affiliation":[{"name":"Faculty of Science, Engineering and Built Environment, School of Information Technology, Deakin University, Burwood, VIC 3125, Australia"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1968","published-online":{"date-parts":[[2026,6,18]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","unstructured":"Wei, P., Wang, X., and Dong, H. (2023). The impact of automated writing evaluation on second language writing skills of Chinese EFL learners: A randomized controlled trial. Front. Psychol., 14.","DOI":"10.3389\/fpsyg.2023.1249991"},{"key":"ref_2","doi-asserted-by":"crossref","first-page":"189","DOI":"10.5539\/elt.v14n12p189","article-title":"The impact of using automated writing feedback in ESL\/EFL classroom contexts","volume":"14","author":"Benali","year":"2021","journal-title":"Engl. Lang. Teach."},{"key":"ref_3","doi-asserted-by":"crossref","first-page":"502","DOI":"10.1109\/TLT.2016.2612659","article-title":"Automated essay feedback generation and its impact on revision","volume":"10","author":"Liu","year":"2016","journal-title":"IEEE Trans. Learn. Technol."},{"key":"ref_4","doi-asserted-by":"crossref","first-page":"100752","DOI":"10.1016\/j.asw.2023.100752","article-title":"Collaborating with ChatGPT in argumentative writing classrooms","volume":"57","author":"Su","year":"2023","journal-title":"Assess. Writ."},{"key":"ref_5","doi-asserted-by":"crossref","first-page":"239","DOI":"10.1177\/0047239517697966","article-title":"E-feedback as a scaffolding teaching strategy in the online language classroom","volume":"46","author":"Alharbi","year":"2017","journal-title":"J. Educ. Technol. Syst."},{"key":"ref_6","doi-asserted-by":"crossref","first-page":"1557","DOI":"10.1007\/s11423-021-10004-9","article-title":"Statewide implementation of automated writing evaluation: Analyzing usage and associations with state test performance in grades 4\u201311","volume":"69","author":"Potter","year":"2021","journal-title":"Educ. Technol. Res. Dev."},{"key":"ref_7","doi-asserted-by":"crossref","first-page":"605","DOI":"10.1080\/09588221.2020.1743323","article-title":"Impact of automated writing evaluation on teacher feedback, student revision, and writing improvement","volume":"35","author":"Link","year":"2022","journal-title":"Comput. Assist. Lang. Learn."},{"key":"ref_8","doi-asserted-by":"crossref","unstructured":"Amoozadeh, M., Daniels, D., Nam, D., Kumar, A., Chen, S., Hilton, M., Srinivasa Ragavan, S., and Alipour, M.A. (2024). Trust in generative AI among students: An exploratory study. Proceedings of the 55th ACM Technical Symposium on Computer Science Education V.1, Association for Computing Machinery.","DOI":"10.1145\/3626252.3630842"},{"key":"ref_9","doi-asserted-by":"crossref","unstructured":"Han, A., Zhou, X., Cai, Z., Han, S., Ko, R., Corrigan, S., and Peppler, K.A. (2024). Teachers, parents, and students\u2019 perspectives on integrating generative AI into elementary literacy education. Proceedings of the 2024 CHI Conference on Human Factors in Computing Systems, Association for Computing Machinery.","DOI":"10.1145\/3613904.3642438"},{"key":"ref_10","doi-asserted-by":"crossref","unstructured":"Park, H., and Ahn, D. (2024). The promise and peril of ChatGPT in higher education: Opportunities, challenges, and design implications. Proceedings of the 2024 CHI Conference on Human Factors in Computing Systems, Association for Computing Machinery.","DOI":"10.1145\/3613904.3642785"},{"key":"ref_11","unstructured":"Wang, S., Xu, T., Li, H., Zhang, C., Liang, J., Tang, J., Yu, P.S., and Wen, Q. (2024). Large language models for education: A survey and outlook. arXiv."},{"key":"ref_12","doi-asserted-by":"crossref","first-page":"100299","DOI":"10.1016\/j.caeai.2024.100299","article-title":"Assessing the proficiency of large language models in automatic feedback generation: An evaluation study","volume":"7","author":"Dai","year":"2024","journal-title":"Comput. Educ. Artif. Intell."},{"key":"ref_13","doi-asserted-by":"crossref","first-page":"105360","DOI":"10.1016\/j.compedu.2025.105360","article-title":"Analytics of learner-centered feedback: A large-scale case study in higher education","volume":"237","author":"Aldino","year":"2025","journal-title":"Comput. Educ."},{"key":"ref_14","doi-asserted-by":"crossref","first-page":"100427","DOI":"10.1016\/j.caeai.2025.100427","article-title":"Evaluating the capability of large language models in characterising relational feedback: A comparative analysis of prompting strategies","volume":"8","author":"Dai","year":"2025","journal-title":"Comput. Educ. Artif. Intell."},{"key":"ref_15","doi-asserted-by":"crossref","first-page":"105511","DOI":"10.1016\/j.compedu.2025.105511","article-title":"Leveraging prompt-based LLMs for automated scoring and feedback generation in higher education","volume":"243","author":"AlGhamdi","year":"2025","journal-title":"Comput. Educ."},{"key":"ref_16","unstructured":"Qian, K., Cheng, Y., Guan, R., Dai, W., Jin, F., Yang, K., Nawaz, S., Swiecki, Z., Chen, G., and Yan, L. (2025). Dean of llm tutors: Exploring comprehensive and automated evaluation of llm-generated educational feedback via llm feedback evaluators. arXiv."},{"key":"ref_17","first-page":"1","article-title":"Automated essay scoring with e-rater\u00ae V.2","volume":"4","author":"Attali","year":"2006","journal-title":"J. Technol. Learn. Assess."},{"key":"ref_18","doi-asserted-by":"crossref","first-page":"2495","DOI":"10.1007\/s10462-021-10068-2","article-title":"An automated essay scoring systems: A systematic literature review","volume":"55","author":"Ramesh","year":"2022","journal-title":"Artif. Intell. Rev."},{"key":"ref_19","doi-asserted-by":"crossref","unstructured":"Taghipour, K., and Ng, H.T. (2016). A neural approach to automated essay scoring. Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing, Association for Computational Linguistics.","DOI":"10.18653\/v1\/D16-1193"},{"key":"ref_20","doi-asserted-by":"crossref","unstructured":"Schmaltz, A., Kim, Y., Rush, A.M., and Shieber, S.M. (2016). Sentence-level grammatical error identification as sequence-to-sequence correction. arXiv.","DOI":"10.18653\/v1\/W16-0528"},{"key":"ref_21","doi-asserted-by":"crossref","unstructured":"Dai, W., Lin, J., Jin, H., Li, T., Tsai, Y.S., Ga\u0161evi\u0107, D., and Chen, G. (2023). Can large language models provide feedback to students? A case study on ChatGPT. Proceedings of the 2023 IEEE International Conference on Advanced Learning Technologies (ICALT), IEEE.","DOI":"10.1109\/ICALT58122.2023.00100"},{"key":"ref_22","doi-asserted-by":"crossref","unstructured":"Naismith, B., Mulcaire, P., and Burstein, J. (2023). Automated evaluation of written discourse coherence using GPT-4. Proceedings of the 18th Workshop on Innovative Use of NLP for Building Educational Applications (BEA 2023), Association for Computational Linguistics.","DOI":"10.18653\/v1\/2023.bea-1.32"},{"key":"ref_23","doi-asserted-by":"crossref","unstructured":"Zhang, D., Hoang, T., Zhu, Y., Wang, R., and Crouch, P. (2025). Generating Feedback for School Students Essay with Large Language Models. Proceedings of the International Conference on Knowledge Science, Engineering and Management, Springer.","DOI":"10.1007\/978-981-95-3055-7_25"},{"key":"ref_24","first-page":"3","article-title":"Challenges, benefits and recommendations for using generative artificial intelligence in academic writing: A case of ChatGPT","volume":"7","author":"Costa","year":"2024","journal-title":"Medicon Eng. Themes"},{"key":"ref_25","doi-asserted-by":"crossref","unstructured":"Bender, E.M., Gebru, T., McMillan-Major, A., and Shmitchell, S. (2021). On the dangers of stochastic parrots: Can language models be too big?. Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency, Association for Computing Machinery.","DOI":"10.1145\/3442188.3445922"},{"key":"ref_26","doi-asserted-by":"crossref","first-page":"391","DOI":"10.1162\/tacl_a_00373","article-title":"Summeval: Re-evaluating summarization evaluation","volume":"9","author":"Fabbri","year":"2021","journal-title":"Trans. Assoc. Comput. Linguist."},{"key":"ref_27","doi-asserted-by":"crossref","first-page":"101894","DOI":"10.1016\/j.learninstruc.2024.101894","article-title":"Comparing the quality of human and ChatGPT feedback of students\u2019 writing","volume":"91","author":"Steiss","year":"2024","journal-title":"Learn. Instr."},{"key":"ref_28","doi-asserted-by":"crossref","first-page":"20385","DOI":"10.1007\/s10639-025-13553-1","article-title":"ChatGPT: A reliable assistant for the evaluation of students\u2019 written texts?","volume":"30","author":"Atasoy","year":"2025","journal-title":"Educ. Inf. Technol."},{"key":"ref_29","doi-asserted-by":"crossref","unstructured":"Papineni, K., Roukos, S., Ward, T., and Zhu, W.J. (2002). Bleu: A method for automatic evaluation of machine translation. Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics, Association for Computational Linguistics.","DOI":"10.3115\/1073083.1073135"},{"key":"ref_30","unstructured":"Lin, C.Y. (2004). Rouge: A package for automatic evaluation of summaries. Proceedings of the Text Summarization Branches Out, Association for Computational Linguistics."},{"key":"ref_31","unstructured":"Zhang, T., Kishore, V., Wu, F., Weinberger, K.Q., and Artzi, Y. (2019). Bertscore: Evaluating text generation with bert. arXiv."},{"key":"ref_32","doi-asserted-by":"crossref","first-page":"638","DOI":"10.1007\/s40593-020-00204-4","article-title":"Collaborating with mature English language learners to combine peer and automated feedback: A user-centered approach to designing writing support","volume":"31","author":"Liaqat","year":"2021","journal-title":"Int. J. Artif. Intell. Educ."},{"key":"ref_33","unstructured":"Tan, K., Pang, T., Fan, C., and Yu, S. (2023). Towards applying powerful large AI models in classroom teaching: Opportunities, challenges and prospects. arXiv."},{"key":"ref_34","unstructured":"Jia, Q., Young, M., Xiao, Y., Cui, J., Liu, C., Rashid, P., and Gehringer, E. (2022). Insta-Reviewer: A data-driven approach for generating instant feedback on students\u2019 project reports. Proceedings of the 15th International Conference on Educational Data Mining, International Educational Data Mining Society."},{"key":"ref_35","unstructured":"Xu, B., Bai, Y., Sun, H., Lin, Y., Liu, S., Liang, X., Li, Y., Gao, Y., and Huang, H. (2025). EduBench: A comprehensive benchmarking dataset for evaluating large language models in diverse educational scenarios. arXiv."},{"key":"ref_36","doi-asserted-by":"crossref","first-page":"39","DOI":"10.1186\/s41239-019-0171-0","article-title":"Systematic review of research on artificial intelligence applications in higher education: Where are the educators?","volume":"16","author":"Bond","year":"2019","journal-title":"Int. J. Educ. Technol. High. Educ."},{"key":"ref_37","doi-asserted-by":"crossref","first-page":"102274","DOI":"10.1016\/j.lindif.2023.102274","article-title":"ChatGPT for good? On opportunities and challenges of large language models for education","volume":"103","author":"Kasneci","year":"2023","journal-title":"Learn. Individ. Differ."},{"key":"ref_38","doi-asserted-by":"crossref","unstructured":"Weisz, J.D., He, J., Muller, M., Hoefer, G., Miles, R., and Geyer, W. (2024). Design principles for generative AI applications. Proceedings of the 2024 CHI Conference on Human Factors in Computing Systems, Association for Computing Machinery.","DOI":"10.1145\/3613904.3642466"},{"key":"ref_39","doi-asserted-by":"crossref","unstructured":"Amershi, S., Weld, D., Vorvoreanu, M., Fourney, A., Nushi, B., Collisson, P., Suh, J., Iqbal, S., Bennett, P.N., and Inkpen, K. (2019). Guidelines for human\u2013AI interaction. Proceedings of the 2019 CHI Conference on Human Factors in Computing Systems, Association for Computing Machinery.","DOI":"10.1145\/3290605.3300233"},{"key":"ref_40","doi-asserted-by":"crossref","unstructured":"Gilpin, L.H., Bau, D., Yuan, B.Z., Bajwa, A., Specter, M., and Kagal, L. (2018). Explaining explanations: An overview of interpretability of machine learning. Proceedings of the 2018 IEEE 5th International Conference on Data Science and Advanced Analytics (DSAA), IEEE.","DOI":"10.1109\/DSAA.2018.00018"},{"key":"ref_41","doi-asserted-by":"crossref","unstructured":"Han, J., Yoo, H., Myung, J., Kim, M., Lim, H., Kim, Y., Lee, T.Y., Hong, H., Kim, J., and Ahn, S.Y. (2023). LLM-as-a-tutor in EFL writing education: Focusing on evaluation of student\u2013LLM interaction. arXiv.","DOI":"10.18653\/v1\/2024.customnlp4u-1.21"},{"key":"ref_42","doi-asserted-by":"crossref","unstructured":"Kim, H., Baghestani, S., Yin, S., Karatay, Y., Kurt, S., Beck, J., and Karatay, L. (2024). ChatGPT for writing evaluation: Examining the accuracy and reliability of AI-generated scores compared to human raters. Exploring Artificial Intelligence in Applied Linguistics, Iowa State University Digital Press.","DOI":"10.31274\/isudp.2024.154.06"},{"key":"ref_43","doi-asserted-by":"crossref","unstructured":"Yazici, A., Mejia-Domenzain, P., Frej, J., and K\u00e4ser, T. (2024). GELEX: Generative AI-hybrid system for example-based learning. Proceedings of the Extended Abstracts of the CHI Conference on Human Factors in Computing Systems, Association for Computing Machinery.","DOI":"10.1145\/3613905.3650900"},{"key":"ref_44","doi-asserted-by":"crossref","first-page":"153","DOI":"10.3102\/0034654307313795","article-title":"Focus on formative feedback","volume":"78","author":"Shute","year":"2008","journal-title":"Rev. Educ. Res."},{"key":"ref_45","doi-asserted-by":"crossref","first-page":"199","DOI":"10.1080\/03075070600572090","article-title":"Formative assessment and self-regulated learning: A model and seven principles of good feedback practice","volume":"31","author":"Nicol","year":"2006","journal-title":"Stud. High. Educ."},{"key":"ref_46","doi-asserted-by":"crossref","first-page":"150","DOI":"10.1080\/13562517.2020.1782372","article-title":"Teacher feedback literacy and its interplay with student feedback literacy","volume":"28","author":"Carless","year":"2023","journal-title":"Teach. High. Educ."},{"key":"ref_47","doi-asserted-by":"crossref","first-page":"467","DOI":"10.1007\/s10734-017-0220-3","article-title":"Developing evaluative judgement: Enabling students to make decisions about the quality of work","volume":"76","author":"Tai","year":"2018","journal-title":"High. Educ."},{"key":"ref_48","doi-asserted-by":"crossref","first-page":"81","DOI":"10.3102\/003465430298487","article-title":"The power of feedback","volume":"77","author":"Hattie","year":"2007","journal-title":"Rev. Educ. Res."},{"key":"ref_49","doi-asserted-by":"crossref","first-page":"295","DOI":"10.3102\/00028312043002295","article-title":"Improving the writing, knowledge, and motivation of struggling young writers: Effects of self-regulated strategy development with and without peer support","volume":"43","author":"Harris","year":"2006","journal-title":"Am. Educ. Res. J."},{"key":"ref_50","unstructured":"Graham, S., and Perin, D. (2007). Writing Next: Effective Strategies to Improve Writing of Adolescents in Middle and High Schools, Carnegie Corporation."},{"key":"ref_51","unstructured":"OpenAI (2024, June 15). GPT-4 Technical Report. Available online: https:\/\/openai.com\/research\/gpt-4."},{"key":"ref_52","doi-asserted-by":"crossref","unstructured":"Lee, S., Cai, Y., Meng, D., Wang, Z., and Wu, Y. (2024). Unleashing large language models\u2019 proficiency in zero-shot essay scoring. arXiv.","DOI":"10.18653\/v1\/2024.findings-emnlp.10"},{"key":"ref_53","doi-asserted-by":"crossref","unstructured":"Mannila, L. (2024). Co-designing AI literacy for K\u201312 education. Proceedings of the 19th WiPSCE Conference on Primary and Secondary Computing Education Research, Association for Computing Machinery.","DOI":"10.1145\/3677619.3678716"},{"key":"ref_54","doi-asserted-by":"crossref","unstructured":"Wang, Z., Makarova, V., Li, Z., Kodner, J., and Rambow, O. (2025). LLMs can perform multi-dimensional analytic writing assessments: A case study of L2 graduate-level academic English writing. arXiv.","DOI":"10.18653\/v1\/2025.acl-long.423"},{"key":"ref_55","doi-asserted-by":"crossref","first-page":"653","DOI":"10.1080\/09588221.2018.1428994","article-title":"Automated written corrective feedback: How well can students make use of it?","volume":"31","author":"Ranalli","year":"2018","journal-title":"Comput. Assist. Lang. Learn."},{"key":"ref_56","first-page":"27","article-title":"Co-designing a real-time classroom orchestration tool to support teacher\u2014AI complementarity","volume":"6","author":"Holstein","year":"2019","journal-title":"J. Learn. Anal."},{"key":"ref_57","doi-asserted-by":"crossref","unstructured":"Lin, P., and Van Brummelen, J. (2021). Engaging teachers to co-design integrated AI curriculum for K\u201312 classrooms. Proceedings of the 2021 CHI Conference on Human Factors in Computing Systems, Association for Computing Machinery.","DOI":"10.1145\/3411764.3445377"}],"container-title":["Computers"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/2073-431X\/15\/6\/393\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,6,26]],"date-time":"2026-06-26T04:57:16Z","timestamp":1782449836000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/2073-431X\/15\/6\/393"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,6,18]]},"references-count":57,"journal-issue":{"issue":"6","published-online":{"date-parts":[[2026,6]]}},"alternative-id":["computers15060393"],"URL":"https:\/\/doi.org\/10.3390\/computers15060393","relation":{},"ISSN":["2073-431X"],"issn-type":[{"value":"2073-431X","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,6,18]]}}}