{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,21]],"date-time":"2026-07-21T03:22:01Z","timestamp":1784604121572,"version":"3.55.0"},"reference-count":113,"publisher":"Springer Science and Business Media LLC","issue":"17","license":[{"start":{"date-parts":[[2025,9,8]],"date-time":"2025-09-08T00:00:00Z","timestamp":1757289600000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2025,9,8]],"date-time":"2025-09-08T00:00:00Z","timestamp":1757289600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"DOI":"10.13039\/501100001779","name":"Monash University","doi-asserted-by":"crossref","id":[{"id":"10.13039\/501100001779","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Educ Inf Technol"],"published-print":{"date-parts":[[2025,11]]},"abstract":"<jats:title>Abstract<\/jats:title>\n                  <jats:p>\n                    Grading and feedback provision for students\u2019 open-ended responses are time-consuming and cognitively demanding. Despite the advocacy to leverage Artificial Intelligence (AI) to automate assessment, concerns persist regarding inaccurate AI assessment on student learning and the potential detrimental effects on educators\u2019 assessment practices due to reliance. One alternative is to leverage AI-powered assessment insights as auxiliary information to support educators\u2019 assessment instead of solely entrusting AI for assessment. However, insights published in existing literature typically required excessive cognitive efforts from human graders to interpret and make use of. Generative AI (GenAI) technologies\u2019 capability To produce natural-language assessment explanations could potentially tackle this challenge. To empirically examine the efficacy of such natural-language insights while accounting for concerns in over-reliance on AI, we invited 60 human graders from diverse backgrounds to participate in multiple phases of assessment for secondary students\u2019 short-answer responses. Participants were assigned to one of the three conditions: no AI-powered assessment support; important-word highlights from AI-powered graders as support; and natural-language grading explanations from GenAI-powered graders as support. Mixed-effect regression analyses were adopted to examine the impacts of different AI-powered assessment insights on human assessment practices both in current assessment with AI-powered insights presented and in later assessment without AI-powered insights. Our findings indicate that GenAI-enabled natural-language insights significantly improved educators\u2019 feedback quality compared to educators without AI support (\n                    <jats:inline-formula>\n                      <jats:tex-math>$$\\beta =0.190$$<\/jats:tex-math>\n                    <\/jats:inline-formula>\n                    ,\n                    <jats:inline-formula>\n                      <jats:tex-math>$$p=0.010$$<\/jats:tex-math>\n                    <\/jats:inline-formula>\n                    ). Important-word highlights from traditional AI graders had negligible impact on educators\u2019 feedback quality. Although significant improvements in grading accuracy were not observed, natural-language insights showed greater potential to enhance grading accuracy (\n                    <jats:inline-formula>\n                      <jats:tex-math>$$\\beta =0.556$$<\/jats:tex-math>\n                    <\/jats:inline-formula>\n                    ,\n                    <jats:inline-formula>\n                      <jats:tex-math>$$p=0.093$$<\/jats:tex-math>\n                    <\/jats:inline-formula>\n                    ) than important-word highlights (\n                    <jats:inline-formula>\n                      <jats:tex-math>$$\\beta =-0.060$$<\/jats:tex-math>\n                    <\/jats:inline-formula>\n                    ,\n                    <jats:inline-formula>\n                      <jats:tex-math>$$p=0.848$$<\/jats:tex-math>\n                    <\/jats:inline-formula>\n                    ). Furthermore, educators reported significantly greater satisfaction and willingness to adopt the natural-language insights in practice compared to the important-word insights (\n                    <jats:inline-formula>\n                      <jats:tex-math>$$p&lt;0.05$$<\/jats:tex-math>\n                    <\/jats:inline-formula>\n                    ,\n                    <jats:inline-formula>\n                      <jats:tex-math>$$r \\in [0.352, 0.490]$$<\/jats:tex-math>\n                    <\/jats:inline-formula>\n                    ). Finally, prior exposure to AI-powered insights to some extent fostered educators\u2019 more effective assessment practices in subsequent assessment activities without AI support, though further longitudinal research is needed to establish empirical significance.\n                  <\/jats:p>","DOI":"10.1007\/s10639-025-13741-z","type":"journal-article","created":{"date-parts":[[2025,9,8]],"date-time":"2025-09-08T09:49:28Z","timestamp":1757324968000},"page":"24931-24964","update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":7,"title":["When AI explains in natural language: Unveiling the impact of generative AI explanations on educators\u2019 grading and feedback practices"],"prefix":"10.1007","volume":"30","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-5971-8469","authenticated-orcid":false,"given":"Yuheng","family":"Li","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0004-2866-6063","authenticated-orcid":false,"given":"Zirui","family":"Shan","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-1413-1103","authenticated-orcid":false,"given":"Mladen","family":"Rakovi\u0107","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-6911-3853","authenticated-orcid":false,"given":"Quanlong","family":"Guan","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-9265-1908","authenticated-orcid":false,"given":"Dragan","family":"Ga\u0161evi\u0107","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-8236-3133","authenticated-orcid":false,"given":"Guanliang","family":"Chen","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"297","published-online":{"date-parts":[[2025,9,8]]},"reference":[{"key":"13741_CR1","doi-asserted-by":"publisher","first-page":"63","DOI":"10.1016\/j.jom.2017.06.001","volume":"53","author":"JD Abbey","year":"2017","unstructured":"Abbey, J. D., & Meloy, M. G. (2017). Attention by design: Using attention checks to detect inattentive respondents and improve data quality. Journal of Operations Management, 53, 63\u201370.","journal-title":"Journal of Operations Management"},{"key":"13741_CR2","doi-asserted-by":"crossref","unstructured":"Ackerman, P.L. (2011). Cognitive fatigue: Multidisciplinary perspectives on current research and future applications. American Psychological Association. isbn:9781433808395. http:\/\/www.jstor.org\/stable\/j.ctv1chrtc3","DOI":"10.1037\/12343-000"},{"key":"13741_CR3","doi-asserted-by":"crossref","unstructured":"Ahn, J., Nguyen, H., Campos, F., Young, W. (2021). Transforming everyday information into practical analytics with crowdsourced assessment tasks. Lak21: 11th international learning analytics and knowledge conference (pp. 66\u201376).","DOI":"10.1145\/3448139.3448146"},{"issue":"1","key":"13741_CR4","first-page":"20","volume":"9","author":"P Atchley","year":"2024","unstructured":"Atchley, P., Pannell, H., Wofford, K., Hopkins, M., & Atchley, R. A. (2024). Human and ai collaboration in the higher education environment: opportunities and concerns. Cognitive Research: Principles and Implications, 9(1), 20.","journal-title":"Cognitive Research: Principles and Implications"},{"key":"13741_CR5","unstructured":"Attali, Y., & Burstein, J. (2006). Automated essay scoring with e-rater\u00ae v. 2. The Journal of Technology, Learning and Assessment, 4(3), 31."},{"issue":"1","key":"13741_CR6","first-page":"7","volume":"5","author":"P Black","year":"1998","unstructured":"Black, P., & Wiliam, D. (1998). Assessment and classroom learning. Assessment in Education: principles, policy & practice, 5(1), 7\u201374.","journal-title":"Assessment in Education: principles, policy & practice"},{"key":"13741_CR7","doi-asserted-by":"crossref","unstructured":"Block, R. A., Hancock, P. A., & Zakay, D. (2010). How cognitive load affects duration judgments: A meta-analytic review. Acta Psychologica, 134(3), 330\u2013343.","DOI":"10.1016\/j.actpsy.2010.03.006"},{"key":"13741_CR8","doi-asserted-by":"crossref","unstructured":"Bolotova, V., Blinov, V., Zheng, Y., Croft, W.B., Scholer, F., Sanderson, M. (2020). Do people and neural nets pay attention to the same words: studying eye-tracking data for non-factoid qa evaluation. Proceedings of the 29th acm international conference on information & knowledge management (pp. 85\u201394).","DOI":"10.1145\/3340531.3412043"},{"key":"13741_CR9","doi-asserted-by":"crossref","unstructured":"Bonthu, S., Rama\u00a0Sree, S., Krishna\u00a0Prasad, M. (2021). Automated short answer grading using deep learning: A survey. International cross-domain conference for machine learning and knowledge extraction (pp. 61\u201378).","DOI":"10.1007\/978-3-030-84060-0_5"},{"key":"13741_CR10","doi-asserted-by":"crossref","unstructured":"Boud, D., & Molloy, E. (2013). Feedback in higher and professional education: understanding it and doing it well. Abingdon, Oxon: Routledge.","DOI":"10.4324\/9780203074336"},{"key":"13741_CR11","unstructured":"Brookhart, S. (2012). Grading and learning: Practices that support student achievement. Solution Tree Press. https:\/\/books.google.com.au\/books?id=bmgXBwAAQBAJ"},{"key":"13741_CR12","unstructured":"Brown, S. (2005). Assessment for learning. Learning and teaching in higher education, 1, 81\u201389."},{"key":"13741_CR13","unstructured":"Brown, G., Bull, J., & Pendlebury, M. (1997). Assessing student learning in higher education. London: Routledge."},{"key":"13741_CR14","doi-asserted-by":"crossref","unstructured":"Butler, R. (1988). Enhancing and undermining intrinsic motivation: The effects of task-involving and ego-involving evaluation on interest and performance. British journal of educational psychology,58(1), 1\u201314.","DOI":"10.1111\/j.2044-8279.1988.tb00874.x"},{"issue":"4","key":"13741_CR15","doi-asserted-by":"publisher","first-page":"254","DOI":"10.1080\/19488300.2013.858378","volume":"3","author":"S Cao","year":"2013","unstructured":"Cao, S., & Liu, Y. (2013). Effects of concurrent tasks on diagnostic decision making: An experimental investigation. IIE Transactions on Healthcare Systems Engineering, 3(4), 254\u2013262. https:\/\/doi.org\/10.1080\/19488300.2013.858378","journal-title":"IIE Transactions on Healthcare Systems Engineering"},{"key":"13741_CR16","doi-asserted-by":"crossref","unstructured":"Cavalcanti, A. P., Barbosa, A., Carvalho, R., Freitas, F., Tsai, S., Ga\u0161evi\u0107, D., & Mello, R. F. (2021). Automatic feedback in online learning environments: A systematic literature review. Computers and Education: Artificial Intelligence, 2, Article 100027.","DOI":"10.1016\/j.caeai.2021.100027"},{"key":"13741_CR17","doi-asserted-by":"crossref","unstructured":"Cegin, J., Pecher, B., Simko, J., Srba, I., Bielikova, M., Brusilovsky, P. (2024). Use random selection for now: Investigation of few-shot selection strategies in llm-based text augmentation for classification. arXiv:2410.10756","DOI":"10.18653\/v1\/2025.findings-emnlp.296"},{"key":"13741_CR18","unstructured":"Chamieh, I., Zesch, T., Giebermann, K. (2024). LLMs in short answer scoring: Limitations and promise of zero-shot and few-shot approaches. E.\u00a0Kochmar et\u00a0al. (Eds.),Proceedings of the 19th workshop on innovative use of nlp for building educational applications (bea 2024) (pp. 309\u2013315). Mexico City, Mexico: Association for Computational Linguistics. https:\/\/aclanthology.org\/2024.bea-1.25"},{"key":"13741_CR19","doi-asserted-by":"crossref","unstructured":"Colonna, L. (2023). Teachers in the loop? an analysis of automatic assessment systems under article 22 gdpr. International Data Privacy Law, 14(1), 3\u201318.","DOI":"10.1093\/idpl\/ipad024"},{"key":"13741_CR20","doi-asserted-by":"crossref","unstructured":"Dai, W., Lin, J., Jin, H., Li, T., Tsai, S., Ga\u0161evi\u0107, D., Chen, G. (2023). Can large language models provide feedback to students? a case study on chatgpt. 2023 ieee international conference on advanced learning technologies (icalt) (pp. 323\u2013325).","DOI":"10.1109\/ICALT58122.2023.00100"},{"key":"13741_CR21","doi-asserted-by":"crossref","unstructured":"Dai, W., Tsai, S., Lin, J., Aldino, A., Jin, H., Li, T., & Chen, G. (2024). Assessing the proficiency of large language models in automatic feedback generation: An evaluation study. Computers and Education: Artificial Intelligence, 7, Article 100299.","DOI":"10.1016\/j.caeai.2024.100299"},{"key":"13741_CR22","doi-asserted-by":"crossref","unstructured":"Deeva, G., Bogdanova, D., Serral, E., Snoeck, M., & De Weerdt, J. (2021). A review of automated feedback systems for learners: Classification framework, challenges and opportunities. Computers & Education, 162, Article 104094.","DOI":"10.1016\/j.compedu.2020.104094"},{"key":"13741_CR23","doi-asserted-by":"crossref","unstructured":"DiSabito, D., Hansen, L., Mennella, T., & Rodriguez, J. (2025). Exploring the frontiers of generative ai in assessment: Is there potential for a human-ai partnership? New Directions for Teaching and Learning, 2025(182), 81\u201396.","DOI":"10.1002\/tl.20630"},{"key":"13741_CR24","doi-asserted-by":"crossref","unstructured":"Fan, Y., Tang, L., Le, H., Shen, K., Tan, S., Zhao, Y., & Ga\u0161evi\u0107, D. (2025). Beware of metacognitive laziness: Effects of generative artificial intelligence on learning motivation, processes, and performance. British Journal of Educational Technology, 56(2), 489\u2013530.","DOI":"10.1111\/bjet.13544"},{"key":"13741_CR25","doi-asserted-by":"publisher","unstructured":"Farazouli, A. (2024). Automation and assessment: Exploring ethical issues of automated grading systems from a relational ethics approach. In: Framing futures in postdigital education: Critical concepts for data-driven practices (pp. 209\u2013226). Cham: Springer Nature Switzerland. https:\/\/doi.org\/10.1007\/978-3-031-58622-4_12","DOI":"10.1007\/978-3-031-58622-4_12"},{"key":"13741_CR26","doi-asserted-by":"publisher","DOI":"10.1016\/j.jpubeco.2022.104773","volume":"216","author":"B Ferman","year":"2022","unstructured":"Ferman, B., & Fontes, L. F. (2022). Assessing knowledge or classroom behavior? evidence of teachers\u2019 grading bias. Journal of Public Economics, 216, Article 104773.","journal-title":"Journal of Public Economics"},{"key":"13741_CR27","doi-asserted-by":"publisher","unstructured":"Figueras, C., Rossitto, C., Cerratto\u00a0Pargman, T. (2024). Doing responsibilities with automated grading systems: An empirical multi-stakeholder exploration. Proceedings of the 13th nordic conference on human-computer interaction. New York, NY, USA: Association for Computing Machinery. https:\/\/doi.org\/10.1145\/3679318.3685334","DOI":"10.1145\/3679318.3685334"},{"key":"13741_CR28","doi-asserted-by":"crossref","unstructured":"Filighera, A., Parihar, S., Steuer, T., Meuser, T., Ochs, S. (2022). Your answer is incorrect... would you like to know why? introducing a bilingual short answer feedback dataset. S.\u00a0Muresan, P.\u00a0Nakov, and A.\u00a0Villavicencio (Eds.), Proceedings of the 60th annual meeting of the association for computational linguistics (volume 1: Long papers) (pp. 8577\u20138591). Dublin, Ireland: Association for Computational Linguistics. https:\/\/aclanthology.org\/2022.acl-long.587","DOI":"10.18653\/v1\/2022.acl-long.587"},{"key":"13741_CR29","doi-asserted-by":"publisher","first-page":"1162454","DOI":"10.3389\/frai.2023.1162454","volume":"6","author":"J Fleckenstein","year":"2023","unstructured":"Fleckenstein, J., Liebenow, L. W., & Meyer, J. (2023). Automated feedback and writing: A multi-level meta-analysis of effects on students\u2019 performance. Frontiers in Artificial Intelligence, 6, 1162454.","journal-title":"Frontiers in Artificial Intelligence"},{"key":"13741_CR30","doi-asserted-by":"crossref","unstructured":"Flod\u00e9n, J. (2025). Grading exams using large language models: A comparison between human and ai grading of exams in higher education using chatgpt. British Educational Research Journal, 51(1), 201\u2013224.","DOI":"10.1002\/berj.4069"},{"key":"13741_CR31","unstructured":"Fragiadakis, G., Diou, C., Kousiouris, G., Nikolaidou, M. (2024). Evaluating human-ai collaboration: A review and methodological framework.arXiv:2407.19098"},{"key":"13741_CR32","doi-asserted-by":"crossref","unstructured":"Funayama, H., Sasaki, S., Matsubayashi, Y., Mizumoto, T., Suzuki, J., Mita, M., Inui, K. (2020). Preventing critical scoring errors in short answer scoring with confidence estimation. Proceedings of the 58th annual meeting of the association for computational linguistics: Student research workshop (pp. 237\u2013243).","DOI":"10.18653\/v1\/2020.acl-srw.32"},{"key":"13741_CR33","doi-asserted-by":"crossref","unstructured":"Funayama, H., Sato, T., Matsubayashi, Y., Mizumoto, T., Suzuki, J., Inui, K. (2022). Balancing cost and quality: An exploration of human-in-the-loop frameworks for automated short answer scoring. M.M.\u00a0Rodrigo, N.\u00a0Matsuda, A.I.\u00a0Cristea, and V.\u00a0Dimitrova (Eds.), Artificial intelligence in education (pp. 465\u2013476). Cham: Springer International Publishing.","DOI":"10.1007\/978-3-031-11644-5_38"},{"key":"13741_CR34","doi-asserted-by":"crossref","unstructured":"Gabbay, H., & Cohen, A. (2024). Combining llm-generated and test-based feedback in a mooc for programming. Proceedings of the eleventh acm conference on learning scale (pp. 177\u2013187).","DOI":"10.1145\/3657604.3662040"},{"key":"13741_CR35","doi-asserted-by":"crossref","unstructured":"Galici, R., Kaser, T., Fenu, G., Marras, M. (2023). Do not trust a model because it is confident: Uncovering and characterizing unknown unknowns to student success predictors in online-based learning. Lak23: 13th international learning analytics and knowledge conference (pp. 441\u2013452).","DOI":"10.1145\/3576050.3576148"},{"key":"13741_CR36","doi-asserted-by":"publisher","first-page":"142","DOI":"10.3389\/fpsyg.2015.00142","volume":"6","author":"B Gathmann","year":"2015","unstructured":"Gathmann, B., Schiebener, J., Wolf, O. T., & Brand, M. (2015). Monitoring supports performance in a dual-task paradigm involving a risky decision-making task and a working memory task. Frontiers in Psychology, 6, 142.","journal-title":"Frontiers in Psychology"},{"key":"13741_CR37","doi-asserted-by":"publisher","DOI":"10.1017\/CBO9780511790942","volume-title":"Data analysis using regression and multilevel\/hierarchical models","author":"A Gelman","year":"2006","unstructured":"Gelman, A., & Hill, J. (2006). Data analysis using regression and multilevel\/hierarchical models. Cambridge University Press."},{"key":"13741_CR38","doi-asserted-by":"publisher","unstructured":"Giannakos, M., Azevedo, R., Brusilovsky, P., Cukurova, M., Dimitriadis, Y., Hernandez-Leo, D., & Rienties, B. (2024). The promise and challenges of generative ai in education. Behaviour & Information Technology, 1\u201327,. https:\/\/doi.org\/10.1080\/0144929X.2024.2394886","DOI":"10.1080\/0144929X.2024.2394886"},{"key":"13741_CR39","doi-asserted-by":"publisher","unstructured":"Gnanaprakasam, J., & Lourdusamy, R. (2024). The role of ai in automating grading: Enhancing feedback and efficiency. S.\u00a0Kadry (Ed.), Artificial intelligence and education (Chap.\u00a07). Rijeka: IntechOpen. https:\/\/doi.org\/10.5772\/intechopen.1005025","DOI":"10.5772\/intechopen.1005025"},{"key":"13741_CR40","unstructured":"Guskey, T.R. (2001). Helping standards make the grade. Educational Leadership, 59(1)."},{"key":"13741_CR41","doi-asserted-by":"publisher","first-page":"108190","DOI":"10.1109\/ACCESS.2021.3100890","volume":"9","author":"MG Hahn","year":"2021","unstructured":"Hahn, M. G., Navarro, S. M. B., De La Fuente Valent\u00edn, L., & Burgos, D. (2021). A systematic review of the effects of automatic scoring and automatic feedback in educational settings. IEEE Access, 9, 108190\u2013108198. https:\/\/doi.org\/10.1109\/ACCESS.2021.3100890","journal-title":"IEEE Access"},{"key":"13741_CR42","doi-asserted-by":"crossref","unstructured":"Hall, E., Seyam, M., Dunlap, D. (2024). Exploring explainability and transparency in automated essay scoring systems: A user-centered evaluation. P.\u00a0Zaphiris and A.\u00a0Ioannou (Eds.), Learning and collaboration technologies (pp. 266\u2013282). Cham: Springer Nature Switzerland.","DOI":"10.1007\/978-3-031-61691-4_18"},{"issue":"1","key":"13741_CR43","doi-asserted-by":"publisher","first-page":"81","DOI":"10.3102\/003465430298487","volume":"77","author":"J Hattie","year":"2007","unstructured":"Hattie, J., & Timperley, H. (2007). The power of feedback. Review of Educational Research, 77(1), 81\u2013112.","journal-title":"Review of Educational Research"},{"key":"13741_CR44","doi-asserted-by":"publisher","first-page":"400","DOI":"10.3758\/s13428-015-0578-z","volume":"48","author":"DJ Hauser","year":"2016","unstructured":"Hauser, D. J., & Schwarz, N. (2016). Attentive turkers: Mturk participants perform better on online attention checks than do subject pool participants. Behavior Research Methods, 48, 400\u2013407.","journal-title":"Behavior Research Methods"},{"issue":"7","key":"13741_CR45","doi-asserted-by":"publisher","first-page":"1401","DOI":"10.1080\/07294360.2019.1657807","volume":"38","author":"M Henderson","year":"2019","unstructured":"Henderson, M., Phillips, M., Ryan, T., Boud, D., Dawson, P., Molloy, E., & Mahoney, P. (2019). Conditions that enable effective feedback. Higher Education Research & Development, 38(7), 1401\u20131416.","journal-title":"Higher Education Research & Development"},{"issue":"8","key":"13741_CR46","doi-asserted-by":"publisher","first-page":"1237","DOI":"10.1080\/02602938.2019.1599815","volume":"44","author":"M Henderson","year":"2019","unstructured":"Henderson, M., Ryan, T., & Phillips, M. (2019). The challenges of feedback in higher education. Assessment & Evaluation in Higher Education, 44(8), 1237\u20131252. https:\/\/doi.org\/10.1080\/02602938.2019.1599815","journal-title":"Assessment & Evaluation in Higher Education"},{"key":"13741_CR47","doi-asserted-by":"publisher","unstructured":"Henkel, O., Hills, L., Boxer, A., Roberts, B., Levonian, Z. (2024). Can large language models make the grade? an empirical study evaluating llms ability to mark short answer questions in k-12 education. Proceedings of the eleventh acm conference on learning @ scale (p.300\u2013304). https:\/\/doi.org\/10.1145\/3657604.3664693","DOI":"10.1145\/3657604.3664693"},{"key":"13741_CR48","doi-asserted-by":"publisher","unstructured":"Hoffman, R. R., Mueller, S. T., Klein, G., & Litman, J. (2023). Measures for explainable ai: Explanation goodness, user satisfaction, mental models, curiosity, trust, and human-ai performance. Frontiers in Computer Science, 5\u20132023,. https:\/\/doi.org\/10.3389\/fcomp.2023.1096257","DOI":"10.3389\/fcomp.2023.1096257"},{"key":"13741_CR49","doi-asserted-by":"publisher","unstructured":"Jia, Q., Cui, J., Du, H., Rashid, P., Xi, R., Li, R., Gehringer, E. (2024a). Llm-generated feedback in real classes and beyond: Perspectives from students and instructors. B.\u00a0Paa\u00c3\u0178en and C.D.\u00a0Epp (Eds.), Proceedings of the 17th international conference on educational data mining (pp. 862\u2013867). https:\/\/doi.org\/10.5281\/zenodo.12729974","DOI":"10.5281\/zenodo.12729974"},{"key":"13741_CR50","doi-asserted-by":"publisher","unstructured":"Jia, Q., Cui, J., Xi, R., Liu, C., Rashid, P., Li, R., Gehringer, E. (2024b). On assessing the faithfulness of llm-generated feedback on student assignments. B.\u00a0Paa\u00c3\u0178en and C.D.\u00a0Epp (Eds.), Proceedings of the 17th international conference on educational data mining (pp. 491\u2013499). https:\/\/doi.org\/10.5281\/zenodo.12729868","DOI":"10.5281\/zenodo.12729868"},{"key":"13741_CR51","doi-asserted-by":"publisher","unstructured":"Jiang, L., & Bosch, N. (2024). Short answer scoring with gpt-4. Proceedings of the eleventh acm conference on learning @ scale (p.438\u2013442). https:\/\/doi.org\/10.1145\/3657604.3664685","DOI":"10.1145\/3657604.3664685"},{"key":"13741_CR52","doi-asserted-by":"publisher","unstructured":"Knight, S., Shibani, A., Abel, S., Gibson, A., Ryan, P. (2020). Acawriter: A learning analytics tool for formative feedback on academic writing. Journal of Writing Research, 12(1), 141\u2013186. https:\/\/doi.org\/10.17239\/jowr-2020.12.01.06","DOI":"10.17239\/jowr-2020.12.01.06"},{"key":"13741_CR53","doi-asserted-by":"crossref","unstructured":"Kumar, Y., Aggarwal, S., Mahata, D., Shah, R.R., Kumaraguru, P., Zimmermann, R. (2019). Get it scored using autosas\u2013an automated system for scoring short answers. Proceedings of the aaai conference on artificial intelligence (Vol.\u00a033, pp. 9662\u20139669).","DOI":"10.1609\/aaai.v33i01.33019662"},{"issue":"2","key":"13741_CR54","doi-asserted-by":"publisher","first-page":"264","DOI":"10.1111\/apps.12108","volume":"67","author":"FY Kung","year":"2018","unstructured":"Kung, F. Y., Kwok, N., & Brown, D. J. (2018). Are attention check questions a threat to scale validity? Applied Psychology, 67(2), 264\u2013283.","journal-title":"Applied Psychology"},{"key":"13741_CR55","volume":"6","author":"G-G Lee","year":"2024","unstructured":"Lee, G.-G., Latif, E., Wu, X., Liu, N., & Zhai, X. (2024). Applying large language models and chain-of-thought for automatic scoring. Computers and Education: Artificial Intelligence, 6, Article 100213.","journal-title":"Computers and Education: Artificial Intelligence"},{"key":"13741_CR56","doi-asserted-by":"crossref","unstructured":"Li, J., Zhou, Y., Hao, T. (2024). Effects of the interaction between time-on-task and task load on response lapses. Behavioral Sciences, 14(11), 1086.","DOI":"10.3390\/bs14111086"},{"key":"13741_CR57","doi-asserted-by":"publisher","unstructured":"Li, T.W., Hsu, S., Fowler, M., Zhang, Z., Zilles, C., Karahalios, K. (2023). Am i wrong, or is the autograder wrong? effects of ai grading mistakes on learning. Proceedings of the 2023 acm conference on international computing education research - volume 1 (p.159\u2013176). New York, NY, USA: Association for Computing Machinery. https:\/\/doi.org\/10.1145\/3568813.3600124","DOI":"10.1145\/3568813.3600124"},{"key":"13741_CR58","doi-asserted-by":"publisher","DOI":"10.1016\/j.compedu.2025.105244","volume":"228","author":"Y Li","year":"2025","unstructured":"Li, Y., Rakovi\u0107, M., Srivastava, N., Li, X., Guan, Q., Ga\u0161evi\u0107, D., & Chen, G. (2025). Can ai support human grading? examining machine attention and confidence in short answer scoring. Computers & Education, 228, Article 105244. https:\/\/doi.org\/10.1016\/j.compedu.2025.105244","journal-title":"Computers & Education"},{"issue":"11","key":"13741_CR59","doi-asserted-by":"publisher","first-page":"1086","DOI":"10.3390\/bs14111086","volume":"14","author":"J Li","year":"2024","unstructured":"Li, J., Zhou, Y., & Hao, T. (2024). Effects of the interaction between time-on-task and task load on response lapses. Behavioral sciences, 14(11), 1086.","journal-title":"Behavioral sciences"},{"key":"13741_CR60","unstructured":"Lu, C., & Cutumisu, M. (2021). Integrating deep learning into an automated feedback generation system for automated essay scoring. Proceedings of the 14th international conference on educational data mining (pp. 573\u2013579)."},{"key":"13741_CR61","doi-asserted-by":"crossref","unstructured":"Lun, J., Zhu, J., Tang, Y., Yang, M. (2020). Multiple data augmentation strategies for improving performance on automatic short answer scoring. Proceedings of the aaai conference on artificial intelligence (Vol.\u00a034, pp. 13389\u201313396).","DOI":"10.1609\/aaai.v34i09.7062"},{"issue":"1","key":"13741_CR62","doi-asserted-by":"publisher","first-page":"58","DOI":"10.1007\/s11528-023-00911-4","volume":"68","author":"J Mao","year":"2024","unstructured":"Mao, J., Chen, B., & Liu, J. C. (2024). Generative artificial intelligence in education and its implications for assessment. TechTrends, 68(1), 58\u201366.","journal-title":"TechTrends"},{"key":"13741_CR63","doi-asserted-by":"publisher","unstructured":"McCabe, D., Langer, K.G., Borod, J.C., Bender, H.A. (2011). Practice effects. In: Kreutzer, J.S., DeLuca, J., and Caplan, B. (Eds.), Encyclopedia of clinical neuropsychology (pp. 1988\u20131989). https:\/\/doi.org\/10.1007\/978-0-387-79948-3_1139","DOI":"10.1007\/978-0-387-79948-3_1139"},{"key":"13741_CR64","doi-asserted-by":"publisher","DOI":"10.1016\/j.caeai.2023.100199","volume":"6","author":"J Meyer","year":"2024","unstructured":"Meyer, J., Jansen, T., Schiller, R., Liebenow, L. W., Steinbach, M., Horbach, A., & Fleckenstein, J. (2024). Using llms to bring evidence-based feedback into the classroom: Ai-generated feedback increases secondary students\u2019 text revision, motivation, and positive emotions. Computers and Education: Artificial Intelligence, 6, Article 100199. https:\/\/doi.org\/10.1016\/j.caeai.2023.100199","journal-title":"Computers and Education: Artificial Intelligence"},{"issue":"4","key":"13741_CR65","doi-asserted-by":"publisher","first-page":"527","DOI":"10.1080\/02602938.2019.1667955","volume":"45","author":"E Molloy","year":"2020","unstructured":"Molloy, E., Boud, D., & Henderson, M. (2020). Developing a learning-centred framework for feedback literacy. Assessment & Evaluation in Higher Education, 45(4), 527\u2013540.","journal-title":"Assessment & Evaluation in Higher Education"},{"key":"13741_CR66","doi-asserted-by":"crossref","unstructured":"Nazaretsky, T., Mejia-Domenzain, P., Swamy, V., Frej, J., K\u00e4ser, T. (2024). Ai or human? evaluating student feedback perceptions in higher education. R.\u00a0Ferreira\u00a0Mello, N.\u00a0Rummel, I.\u00a0Jivet, G.\u00a0Pishtari, and J.A.\u00a0Ruip\u00e9rez\u00a0Valiente (Eds.), Technology enhanced learning for inclusive and equitable quality education (pp. 284\u2013298). Cham: Springer Nature Switzerland.","DOI":"10.31219\/osf.io\/6zm83"},{"key":"13741_CR67","doi-asserted-by":"publisher","unstructured":"Nilson, L.B. (2014). Specifications grading: Restoring rigor, motivating students, and saving faculty time (1st Ed.). Routledge. https:\/\/doi.org\/10.4324\/9781003447061","DOI":"10.4324\/9781003447061"},{"key":"13741_CR68","unstructured":"Okoye, I., Bethard, S., Sumner, T. (2013). Cu: Computational assessment of short free text answers-a tool for evaluating students\u2019 understanding. Second joint conference on lexical and computational semantics (* sem), volume 2: Proceedings of the seventh international workshop on semantic evaluation (semeval 2013) (pp. 603\u2013607)."},{"key":"13741_CR69","doi-asserted-by":"crossref","unstructured":"Ormerod, C.M., & Kwako, A. (2024). Automated text scoring in the age of generative ai for the gpu-poor. arXiv:2407.01873","DOI":"10.59863\/OKUU1904"},{"key":"13741_CR70","doi-asserted-by":"crossref","unstructured":"Orrell, J. (2007). Assessment beyond belief: the cognitive process of grading. Balancing dilemmas in assessment and learning in contemporary education (pp. 261\u2013274). Routledge.","DOI":"10.4324\/9780203942185-28"},{"issue":"2","key":"13741_CR71","doi-asserted-by":"publisher","first-page":"387","DOI":"10.1007\/s10579-021-09547-3","volume":"56","author":"U Pad\u00f3","year":"2022","unstructured":"Pad\u00f3, U., & Pad\u00f3, S. (2022). Determinants of grader agreement: an analysis of multiple short answer corpora. Language Resources and Evaluation, 56(2), 387\u2013416.","journal-title":"Language Resources and Evaluation"},{"issue":"2","key":"13741_CR72","doi-asserted-by":"publisher","first-page":"220","DOI":"10.1037\/0033-2909.116.2.220","volume":"116","author":"H Pashler","year":"1994","unstructured":"Pashler, H. (1994). Dual-task interference in simple tasks: data and theory. Psychological Bulletin, 116(2), 220.","journal-title":"Psychological Bulletin"},{"key":"13741_CR73","doi-asserted-by":"publisher","first-page":"217","DOI":"10.1007\/s00146-020-01005-y","volume":"36","author":"MM Peeters","year":"2021","unstructured":"Peeters, M. M., van Diggelen, J., Van Den Bosch, K., Bronkhorst, A., Neerincx, M. A., Schraagen, J. M., & Raaijmakers, S. (2021). Hybrid collective intelligence in a human-ai society. AI & Society, 36, 217\u2013238.","journal-title":"AI & Society"},{"key":"13741_CR74","doi-asserted-by":"crossref","unstructured":"Poulton, A., & Eliens, S. (2021). Explaining transformer-based models for automatic short answer grading. Proceedings of the 5th international conference on digital technology in education (pp. 110\u2013116).","DOI":"10.1145\/3488466.3488479"},{"key":"13741_CR75","unstructured":"Price, M. (2012). Assessment literacy: The foundation for improving student learning. Oxford Centre for Staff and Learning Development. https:\/\/books.google.com.au\/books?id=zSmDMwEACAAJ"},{"issue":"3","key":"13741_CR76","doi-asserted-by":"publisher","first-page":"2495","DOI":"10.1007\/s10462-021-10068-2","volume":"55","author":"D Ramesh","year":"2022","unstructured":"Ramesh, D., & Sanampudi, S. K. (2022). An automated essay scoring systems: a systematic literature review. Artificial Intelligence Review, 55(3), 2495\u20132527.","journal-title":"Artificial Intelligence Review"},{"key":"13741_CR77","doi-asserted-by":"crossref","unstructured":"Ram\u00edrez, J., Baez, M., Casati, F., & Benatallah, B. (2019). Understanding the impact of text highlighting in crowdsourcing tasks. Proceedings of the AAAI Conference on Human Computation and Crowdsourcing, 7(1), 144\u2013152.","DOI":"10.1609\/hcomp.v7i1.5268"},{"key":"13741_CR78","unstructured":"Rosenbaum, J.E. (2001). Beyond college for all: Career paths for the forgotten half. Russell Sage Foundation. http:\/\/www.jstor.org\/stable\/10.7758\/9781610444767"},{"issue":"6","key":"13741_CR79","doi-asserted-by":"publisher","first-page":"894","DOI":"10.1080\/02602938.2020.1828819","volume":"46","author":"T Ryan","year":"2021","unstructured":"Ryan, T., Henderson, M., Ryan, K., & Kennedy, G. (2021). Designing learner-centred text-based feedback: a rapid review and qualitative synthesis. Assessment & Evaluation in Higher Education, 46(6), 894\u2013912.","journal-title":"Assessment & Evaluation in Higher Education"},{"issue":"2","key":"13741_CR80","doi-asserted-by":"publisher","first-page":"119","DOI":"10.1007\/BF00117714","volume":"18","author":"DR Sadler","year":"1989","unstructured":"Sadler, D. R. (1989). Formative assessment and the design of instructional systems. Instructional Science, 18(2), 119\u2013144.","journal-title":"Instructional Science"},{"key":"13741_CR81","doi-asserted-by":"crossref","unstructured":"Sato, T., Funayama, H., Hanawa, K., Inui, K. (2022). Plausibility and faithfulness of feature attribution-based explanations in automated short answer scoring. M.M.\u00a0Rodrigo, N.\u00a0Matsuda, A.I.\u00a0Cristea, and V.\u00a0Dimitrova (Eds.), Artificial intelligence in education (pp. 231\u2013242). Cham: Springer International Publishing.","DOI":"10.1007\/978-3-031-11644-5_19"},{"key":"13741_CR82","doi-asserted-by":"crossref","unstructured":"Sawatzki, J., Schlippe, T., Benner-Wickner, M. (2022). Deep learning techniques for automatic short answer grading: Predicting scores for english and german answers. E.C.K.\u00a0Cheng, R.B.\u00a0Koul, T.\u00a0Wang, and X.\u00a0Yu (Eds.), Artificial intelligence in education: Emerging technologies, models and applications (pp. 65\u201375). Singapore: Springer Nature Singapore.","DOI":"10.1007\/978-981-16-7527-0_5"},{"issue":"1","key":"13741_CR83","doi-asserted-by":"publisher","first-page":"3","DOI":"10.1111\/j.1745-3992.1991.tb00170.x","volume":"10","author":"WD Schafer","year":"1991","unstructured":"Schafer, W. D. (1991). Essential assessment skills in professional education of teachers. Educational Measurement: Issues and Practice, 10(1), 3\u20136.","journal-title":"Educational Measurement: Issues and Practice"},{"issue":"2","key":"13741_CR84","doi-asserted-by":"publisher","first-page":"118","DOI":"10.1080\/00405849309543585","volume":"32","author":"WD Schafer","year":"1993","unstructured":"Schafer, W. D. (1993). Assessment literacy for teachers. Theory into practice, 32(2), 118\u2013126.","journal-title":"Theory into practice"},{"key":"13741_CR85","doi-asserted-by":"crossref","unstructured":"Schlippe, T., Stierstorfer, Q., Koppel, M.t., Libbrecht, P. (2022). Explainability in automatic short answer grading. International conference on artificial intelligence in education technology (pp. 69\u201387).","DOI":"10.1007\/978-981-19-8040-4_5"},{"key":"13741_CR86","doi-asserted-by":"crossref","unstructured":"Schneider, J., Richner, R., Riser, M. (2023). Towards trustworthy autograding of short, multi-lingual, multi-type answers. International Journal of Artificial Intelligence in Education, 33(1), 88\u2013118.","DOI":"10.1007\/s40593-022-00289-z"},{"issue":"1","key":"13741_CR87","doi-asserted-by":"publisher","first-page":"88","DOI":"10.1007\/s40593-022-00289-z","volume":"33","author":"J Schneider","year":"2023","unstructured":"Schneider, J., Richner, R., & Riser, M. (2023). Towards trustworthy autograding of short, multi-lingual, multi-type answers. International Journal of Artificial Intelligence in Education, 33(1), 88\u2013118.","journal-title":"International Journal of Artificial Intelligence in Education"},{"key":"13741_CR88","doi-asserted-by":"crossref","unstructured":"Selwyn, N. (2022). Less work for teacher? the ironies of automated decision-making in schools. Everyday automation (pp. 73\u201386). Routledge.","DOI":"10.4324\/9781003170884-6"},{"issue":"1","key":"13741_CR89","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1007\/s42438-022-00362-9","volume":"5","author":"N Selwyn","year":"2023","unstructured":"Selwyn, N., Hillman, T., Bergviken-Rensfeldt, A., & Perrotta, C. (2023). Making sense of the digital automation of education. Postdigital Science and Education, 5(1), 1\u201314.","journal-title":"Postdigital Science and Education"},{"issue":"7","key":"13741_CR90","doi-asserted-by":"publisher","first-page":"4","DOI":"10.3102\/0013189X029007004","volume":"29","author":"LA Shepard","year":"2000","unstructured":"Shepard, L. A. (2000). The role of assessment in a learning culture. Educational Researcher, 29(7), 4\u201314. https:\/\/doi.org\/10.3102\/0013189X029007004","journal-title":"Educational Researcher"},{"key":"13741_CR91","doi-asserted-by":"publisher","DOI":"10.1016\/j.chb.2024.108386","volume":"160","author":"M Stadler","year":"2024","unstructured":"Stadler, M., Bannert, M., & Sailer, M. (2024). Cognitive ease at a cost: Llms reduce mental effort but compromise depth in student scientific inquiry. Computers in Human Behavior, 160, Article 108386. https:\/\/doi.org\/10.1016\/j.chb.2024.108386","journal-title":"Computers in Human Behavior"},{"key":"13741_CR92","unstructured":"Stahl, M., Biermann, L., Nehring, A., Wachsmuth, H. (2024). Exploring LLM prompting strategies for joint essay scoring and feedback generation. E.\u00a0Kochmar et\u00a0al. (Eds.), Proceedings of the 19th workshop on innovative use of nlp for building educational applications (bea 2024) (pp. 283\u2013298). Mexico City, Mexico: Association for Computational Linguistics. https:\/\/aclanthology.org\/2024.bea-1.23"},{"key":"13741_CR93","doi-asserted-by":"crossref","unstructured":"Steinert, S., Krupp, L., Avila, K.E., Janssen, A.S., Ruf, V., Dzsotjan, D., Joisten, K.,... et al. (2024). Lessons learned from designing an open-source automated feedback system for stem education. Education and Information Technologies, 1\u201342.","DOI":"10.1007\/s10639-024-13025-y"},{"issue":"10","key":"13741_CR94","doi-asserted-by":"publisher","first-page":"736","DOI":"10.1016\/j.tics.2017.06.007","volume":"21","author":"N Stewart","year":"2017","unstructured":"Stewart, N., Chandler, J., & Paolacci, G. (2017). Crowdsourcing samples in cognitive science. Trends in Cognitive Sciences, 21(10), 736\u2013748. https:\/\/doi.org\/10.1016\/j.tics.2017.06.007","journal-title":"Trends in Cognitive Sciences"},{"key":"13741_CR95","unstructured":"Sundararajan, M., Taly, A., Yan, Q. (2017). Axiomatic attribution for deep networks. D.\u00a0Precup and Y.W.\u00a0Teh (Eds.), Proceedings of the 34th international conference on machine learning (Vol.\u00a070, pp. 3319\u20133328). PMLR. https:\/\/proceedings.mlr.press\/v70\/sundararajan17a.html"},{"issue":"2","key":"13741_CR96","doi-asserted-by":"publisher","first-page":"257","DOI":"10.1207\/s15516709cog1202_4","volume":"12","author":"J Sweller","year":"1988","unstructured":"Sweller, J. (1988). Cognitive load during problem solving: Effects on learning. Cognitive Science, 12(2), 257\u2013285.","journal-title":"Cognitive Science"},{"key":"13741_CR97","volume":"3","author":"Z Swiecki","year":"2022","unstructured":"Swiecki, Z., Khosravi, H., Chen, G., Martinez-Maldonado, R., Lodge, J. M., Milligan, S., & Ga\u0161evi\u0107, D. (2022). Assessment in the age of artificial intelligence. Computers and Education: Artificial Intelligence, 3, Article 100075.","journal-title":"Computers and Education: Artificial Intelligence"},{"issue":"14","key":"13741_CR98","doi-asserted-by":"publisher","DOI":"10.1016\/j.heliyon.2024.e34262","volume":"10","author":"X Tang","year":"2024","unstructured":"Tang, X., Chen, H., Lin, D., & Li, K. (2024). Harnessing llms for multi-dimensional writing assessment: Reliability and alignment with human judgments. Heliyon, 10(14), Article e34262. https:\/\/doi.org\/10.1016\/j.heliyon.2024.e34262","journal-title":"Heliyon"},{"key":"13741_CR99","doi-asserted-by":"crossref","unstructured":"To, J., Tan, K., Lim, M. (2023). From error-focused to learner-centred feedback practices: Unpacking the development of teacher feedback literacy. Teaching and Teacher Education, 131, 104185.","DOI":"10.1016\/j.tate.2023.104185"},{"key":"13741_CR100","doi-asserted-by":"publisher","DOI":"10.1016\/j.tate.2023.104185","volume":"131","author":"J To","year":"2023","unstructured":"To, J., Tan, K., & Lim, M. (2023). From error-focused to learner-centred feedback practices: Unpacking the development of teacher feedback literacy. Teaching and Teacher Education, 131, Article 104185.","journal-title":"Teaching and Teacher Education"},{"key":"13741_CR101","doi-asserted-by":"crossref","unstructured":"Vetrivel, S., Vidhyapriya, P., Arun, V. (2024). The role of ai in transforming assessment practices in education. Ai applications and strategies in teacher education (pp. 43\u201370). IGI Global.","DOI":"10.4018\/979-8-3693-5443-8.ch003"},{"issue":"2","key":"13741_CR102","doi-asserted-by":"publisher","first-page":"213","DOI":"10.1080\/02602938.2021.1902467","volume":"47","author":"N Winstone","year":"2022","unstructured":"Winstone, N., Boud, D., Dawson, P., & Heron, M. (2022). From feedback-as-information to feedback-as-process: a linguistic analysis of the feedback literature. Assessment & Evaluation in Higher Education, 47(2), 213\u2013230.","journal-title":"Assessment & Evaluation in Higher Education"},{"key":"13741_CR103","doi-asserted-by":"publisher","unstructured":"Woods, B., Adamson, D., Miel, S., Mayfield, E. (2017). Formative essay feedback using predictive scoring models. Proceedings of the 23rd acm sigkdd international conference on knowledge discovery and data mining (p.2071\u20132080). New York, NY, USA: Association for Computing Machinery. https:\/\/doi.org\/10.1145\/3097983.3098160","DOI":"10.1145\/3097983.3098160"},{"key":"13741_CR104","doi-asserted-by":"crossref","unstructured":"Wu, X., Saraf, P.P., Lee, G-G., Latif, E., Liu, N., Zhai, X. (2024). Unveiling scoring processes: Dissecting the differences between llms and human graders in automatic scoring. arXiv:2407.18328","DOI":"10.1007\/s10758-025-09836-8"},{"key":"13741_CR105","doi-asserted-by":"publisher","DOI":"10.1016\/j.cedpsych.2019.101826","volume":"60","author":"Y Wu","year":"2020","unstructured":"Wu, Y., & Schunn, C. D. (2020). From feedback to revisions: Effects of feedback features and perceptions. Contemporary Educational Psychology, 60, Article 101826. https:\/\/doi.org\/10.1016\/j.cedpsych.2019.101826","journal-title":"Contemporary Educational Psychology"},{"key":"13741_CR106","doi-asserted-by":"crossref","unstructured":"Xiao, C., Ma, W., Song, Q., Xu, S.X., Zhang, K., Wang, Y., Fu, Q. (2024). Human-ai collaborative essay scoring: A dual-process framework with llms.arXiv:2401.06431","DOI":"10.1145\/3706468.3706507"},{"key":"13741_CR107","doi-asserted-by":"publisher","first-page":"125403","DOI":"10.1109\/ACCESS.2021.3110683","volume":"9","author":"J Xue","year":"2021","unstructured":"Xue, J., Tang, X., & Zheng, L. (2021). A hierarchical bert-based transfer learning approach for multi-dimensional essay scoring. IEEE Access, 9, 125403\u2013125415. https:\/\/doi.org\/10.1109\/ACCESS.2021.3110683","journal-title":"IEEE Access"},{"key":"13741_CR108","doi-asserted-by":"crossref","unstructured":"Yan, L., Sha, L., Zhao, L., Li, Y., Martinez-Maldonado, R., Chen, G.. Ga\u0161evi\u0107, D. (2024). Practical and ethical challenges of large language models in education: A systematic scoping review. British Journal of Educational Technology, 55(1), 90\u2013112.","DOI":"10.1111\/bjet.13370"},{"key":"13741_CR109","doi-asserted-by":"publisher","unstructured":"Yan, Z., Zhang, R., Jia, F. (2024). Exploring the potential of large language models as a grading tool for conceptual short-answer questions in introductory physics. Proceedings of the 2024 9th international conference on distance education and learning (p.308\u2013314). https:\/\/doi.org\/10.1145\/3675812.3675837","DOI":"10.1145\/3675812.3675837"},{"issue":"1","key":"13741_CR110","doi-asserted-by":"publisher","first-page":"90","DOI":"10.1111\/bjet.13370","volume":"55","author":"L Yan","year":"2024","unstructured":"Yan, L., Sha, L., Zhao, L., Li, Y., Martinez-Maldonado, R., Chen, G., & Ga\u0161evi\u0107, D. (2024). Practical and ethical challenges of large language models in education: A systematic scoping review. British Journal of Educational Technology, 55(1), 90\u2013112.","journal-title":"British Journal of Educational Technology"},{"key":"13741_CR111","doi-asserted-by":"crossref","unstructured":"Yuan, X., & Zhong, L. (2024). Effects of multitasking and task interruptions on task performance and cognitive load: considering the moderating role of individual resilience. Current Psychology,1\u201311.","DOI":"10.1007\/s12144-024-06094-2"},{"key":"13741_CR112","doi-asserted-by":"crossref","unstructured":"Zeng, Z., Li, X., Gasevic, D., Chen, G. (2022). Do deep neural nets display human-like attention in short answer scoring? Proceedings of the 2022 conference of the north american chapter of the association for computational linguistics: Human language technologies (pp. 191\u2013205).","DOI":"10.18653\/v1\/2022.naacl-main.14"},{"issue":"1","key":"13741_CR113","doi-asserted-by":"publisher","first-page":"20","DOI":"10.1108\/AJIM-12-2021-0385","volume":"75","author":"L Zhao","year":"2023","unstructured":"Zhao, L., Zhang, Y., & Zhang, C. (2023). Does attention mechanism possess the feature of human reading? a perspective of sentiment classification task. Aslib Journal of Information Management, 75(1), 20\u201343.","journal-title":"Aslib Journal of Information Management"}],"container-title":["Education and Information Technologies"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10639-025-13741-z.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s10639-025-13741-z","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10639-025-13741-z.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,1,3]],"date-time":"2026-01-03T08:20:45Z","timestamp":1767428445000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s10639-025-13741-z"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,9,8]]},"references-count":113,"journal-issue":{"issue":"17","published-print":{"date-parts":[[2025,11]]}},"alternative-id":["13741"],"URL":"https:\/\/doi.org\/10.1007\/s10639-025-13741-z","relation":{},"ISSN":["1360-2357","1573-7608"],"issn-type":[{"value":"1360-2357","type":"print"},{"value":"1573-7608","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,9,8]]},"assertion":[{"value":"27 January 2025","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"26 July 2025","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"8 September 2025","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"There is no known conflict of interest for this paper.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Conflicts of interest"}}]}}