{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,10,2]],"date-time":"2026-10-02T09:48:08Z","timestamp":1790934488518,"version":"4.1.0"},"reference-count":62,"publisher":"Springer Science and Business Media LLC","issue":"1","license":[{"start":{"date-parts":[[2023,10,30]],"date-time":"2023-10-30T00:00:00Z","timestamp":1698624000000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2023,10,30]],"date-time":"2023-10-30T00:00:00Z","timestamp":1698624000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"DOI":"10.13039\/100016135","name":"Universit\u00e4t Passau","doi-asserted-by":"crossref","id":[{"id":"10.13039\/100016135","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Sci Rep"],"abstract":"<jats:title>Abstract<\/jats:title>\n                  <jats:p>ChatGPT and similar generative AI models have attracted hundreds of millions of users and have become part of the public discourse. Many believe that such models will disrupt society and lead to significant changes in the education system and information generation. So far, this belief is based on either colloquial evidence or benchmarks from the owners of the models\u2014both lack scientific rigor. We systematically assess the quality of AI-generated content through a large-scale study comparing human-written versus ChatGPT-generated argumentative student essays. We use essays that were rated by a large number of human experts (teachers). We augment the analysis by considering a set of linguistic characteristics of the generated essays. Our results demonstrate that ChatGPT generates essays that are rated higher regarding quality than human-written essays. The writing style of the AI models exhibits linguistic characteristics that are different from those of the human-written essays. Since the technology is readily available, we believe that educators must act immediately. We must re-invent homework and develop teaching concepts that utilize these AI models in the same way as math utilizes the calculator: teach the general concepts first and then use AI tools to free up time for other learning objectives.<\/jats:p>","DOI":"10.1038\/s41598-023-45644-9","type":"journal-article","created":{"date-parts":[[2023,10,30]],"date-time":"2023-10-30T08:02:33Z","timestamp":1698652953000},"update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":320,"title":["A large-scale comparison of human-written versus ChatGPT-generated essays"],"prefix":"10.1038","volume":"13","author":[{"given":"Steffen","family":"Herbold","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Annette","family":"Hautli-Janisz","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Ute","family":"Heuer","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Zlata","family":"Kikteva","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Alexander","family":"Trautsch","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"297","published-online":{"date-parts":[[2023,10,30]]},"reference":[{"key":"45644_CR1","unstructured":"Ouyang, L. et\u00a0al. Training language models to follow instructions with human feedback (2022). arXiv:2203.02155."},{"key":"45644_CR2","unstructured":"Ruby, D. 30+ detailed chatgpt statistics\u2013users & facts (sep 2023). https:\/\/www.demandsage.com\/chatgpt-statistics\/ (2023). Accessed 09 June 2023."},{"key":"45644_CR3","unstructured":"Leahy, S. & Mishra, P. TPACK and the Cambrian explosion of AI. In Society for Information Technology & Teacher Education International Conference, (ed. Langran, E.) 2465\u20132469 (Association for the Advancement of Computing in Education (AACE), 2023)."},{"key":"45644_CR4","unstructured":"Ortiz, S. Need an ai essay writer? here\u2019s how chatgpt (and other chatbots) can help. https:\/\/www.zdnet.com\/article\/how-to-use-chatgpt-to-write-an-essay\/ (2023). Accessed 09 June 2023."},{"key":"45644_CR5","unstructured":"Openai chat interface. https:\/\/chat.openai.com\/. Accessed 09 June 2023."},{"key":"45644_CR6","unstructured":"OpenAI. Gpt-4 technical report (2023). arXiv:2303.08774."},{"key":"45644_CR7","unstructured":"Brown, T.\u00a0B. et\u00a0al. Language models are few-shot learners (2020). arXiv:2005.14165."},{"key":"45644_CR8","unstructured":"Wang, B. Mesh-Transformer-JAX: Model-Parallel Implementation of Transformer Language Model with JAX. https:\/\/github.com\/kingoflolz\/mesh-transformer-jax (2021)."},{"key":"45644_CR9","unstructured":"Wei, J. et\u00a0al. Finetuned language models are zero-shot learners. In International Conference on Learning Representations (2022)."},{"key":"45644_CR10","unstructured":"Taori, R. et\u00a0al. Stanford alpaca: An instruction-following llama model. https:\/\/github.com\/tatsu-lab\/stanford_alpaca (2023)."},{"key":"45644_CR11","doi-asserted-by":"crossref","unstructured":"Cai, Z.\u00a0G., Haslett, D.\u00a0A., Duan, X., Wang, S. & Pickering, M.\u00a0J. Does chatgpt resemble humans in language use? (2023). arXiv:2303.08014.","DOI":"10.31234\/osf.io\/s49qv"},{"key":"45644_CR12","doi-asserted-by":"crossref","unstructured":"Mahowald, K. A discerning several thousand judgments: Gpt-3 rates the article + adjective + numeral + noun construction (2023). arXiv:2301.12564.","DOI":"10.18653\/v1\/2023.eacl-main.20"},{"key":"45644_CR13","unstructured":"Dentella, V., Murphy, E., Marcus, G. & Leivada, E. Testing ai performance on less frequent aspects of language reveals insensitivity to underlying meaning (2023). arXiv:2302.12313."},{"key":"45644_CR14","unstructured":"Guo, B. et\u00a0al. How close is chatgpt to human experts? comparison corpus, evaluation, and detection (2023). arXiv:2301.07597."},{"key":"45644_CR15","unstructured":"Zhao, W. et\u00a0al. Is chatgpt equipped with emotional dialogue capabilities? (2023). arXiv:2304.09582."},{"key":"45644_CR16","doi-asserted-by":"publisher","unstructured":"Keim, D.\u00a0A. & Oelke, D. Literature fingerprinting : A new method for visual literary analysis. In 2007 IEEE Symposium on Visual Analytics Science and Technology, 115\u2013122, https:\/\/doi.org\/10.1109\/VAST.2007.4389004 (IEEE, 2007).","DOI":"10.1109\/VAST.2007.4389004"},{"key":"45644_CR17","doi-asserted-by":"crossref","unstructured":"El-Assady, M. et\u00a0al. Interactive visual analysis of transcribed multi-party discourse. In Proceedings of ACL 2017, System Demonstrations, 49\u201354 (Association for Computational Linguistics, Vancouver, Canada, 2017).","DOI":"10.18653\/v1\/P17-4009"},{"key":"45644_CR18","unstructured":"Mennatallah El-Assady, A. H.-J. & Butt, M. Discourse maps - feature encoding for the analysis of verbatim conversation transcripts. In Visual Analytics for Linguistics, vol. CSLI Lecture Notes, Number 220, 115\u2013147 (Stanford: CSLI Publications, 2020)."},{"key":"45644_CR19","doi-asserted-by":"publisher","unstructured":"Matt\u00a0Foulis, J.\u00a0V. & Reed, C. Dialogical fingerprinting of debaters. In Proceedings of COMMA 2020, 465\u2013466, https:\/\/doi.org\/10.3233\/FAIA200536 (Amsterdam: IOS Press, 2020).","DOI":"10.3233\/FAIA200536"},{"key":"45644_CR20","unstructured":"Matt\u00a0Foulis, J.\u00a0V. & Reed, C. Interactive visualisation of debater identification and characteristics. In Proceedings of the COMMA workshop on Argument Visualisation, COMMA, 1\u20137 (2020)."},{"key":"45644_CR21","unstructured":"Chatzipanagiotidis, S., Giagkou, M. & Meurers, D. Broad linguistic complexity analysis for Greek readability classification. In Proceedings of the 16th Workshop on Innovative Use of NLP for Building Educational Applications, 48\u201358 (Association for Computational Linguistics, Online, 2021)."},{"key":"45644_CR22","unstructured":"Ajili, M., Bonastre, J.-F., Kahn, J., Rossato, S. & Bernard, G. FABIOLE, a speech database for forensic speaker comparison. In Proceedings of the Tenth International Conference on Language Resources and Evaluation (LREC\u201916), 726\u2013733 (European Language Resources Association (ELRA), Portoro\u017e, Slovenia, 2016)."},{"key":"45644_CR23","doi-asserted-by":"publisher","unstructured":"Deutsch, T., Jasbi, M. & Shieber, S. Linguistic features for readability assessment. In Proceedings of the Fifteenth Workshop on Innovative Use of NLP for Building Educational Applications, 1\u201317, https:\/\/doi.org\/10.18653\/v1\/2020.bea-1.1 (Association for Computational Linguistics, Seattle, WA, USA $$\\rightarrow$$ Online, 2020).","DOI":"10.18653\/v1\/2020.bea-1.1"},{"key":"45644_CR24","doi-asserted-by":"publisher","unstructured":"Fiacco, J., Jiang, S., Adamson, D. & Ros\u00e9, C. Toward automatic discourse parsing of student writing motivated by neural interpretation. In Proceedings of the 17th Workshop on Innovative Use of NLP for Building Educational Applications (BEA 2022), 204\u2013215, https:\/\/doi.org\/10.18653\/v1\/2022.bea-1.25 (Association for Computational Linguistics, Seattle, Washington, 2022).","DOI":"10.18653\/v1\/2022.bea-1.25"},{"key":"45644_CR25","doi-asserted-by":"publisher","unstructured":"Weiss, Z., Riemenschneider, A., Schr\u00f6ter, P. & Meurers, D. Computationally modeling the impact of task-appropriate language complexity and accuracy on human grading of German essays. In Proceedings of the Fourteenth Workshop on Innovative Use of NLP for Building Educational Applications, 30\u201345, https:\/\/doi.org\/10.18653\/v1\/W19-4404 (Association for Computational Linguistics, Florence, Italy, 2019).","DOI":"10.18653\/v1\/W19-4404"},{"key":"45644_CR26","doi-asserted-by":"publisher","unstructured":"Yang, F., Dragut, E. & Mukherjee, A. Predicting personal opinion on future events with fingerprints. In Proceedings of the 28th International Conference on Computational Linguistics, 1802\u20131807, https:\/\/doi.org\/10.18653\/v1\/2020.coling-main.162 (International Committee on Computational Linguistics, Barcelona, Spain (Online), 2020).","DOI":"10.18653\/v1\/2020.coling-main.162"},{"key":"45644_CR27","doi-asserted-by":"crossref","unstructured":"Tumarada, K. et\u00a0al. Opinion prediction with user fingerprinting. In Proceedings of the International Conference on Recent Advances in Natural Language Processing (RANLP 2021), 1423\u20131431 (INCOMA Ltd., Held Online, 2021).","DOI":"10.26615\/978-954-452-072-4_159"},{"key":"45644_CR28","doi-asserted-by":"crossref","unstructured":"Rocca, R. & Yarkoni, T. Language as a fingerprint: Self-supervised learning of user encodings using transformers. In Findings of the Association for Computational Linguistics: EMNLP. 1701\u20131714 (Association for Computational Linguistics, Abu Dhabi, United Arab Emirates, 2022).","DOI":"10.18653\/v1\/2022.findings-emnlp.123"},{"key":"45644_CR29","doi-asserted-by":"crossref","unstructured":"Aiyappa, R., An, J., Kwak, H. & Ahn, Y.-Y. Can we trust the evaluation on chatgpt? (2023). arXiv:2303.12767.","DOI":"10.18653\/v1\/2023.trustnlp-1.5"},{"key":"45644_CR30","doi-asserted-by":"crossref","unstructured":"Yeadon, W., Inyang, O.-O., Mizouri, A., Peach, A. & Testrow, C. The death of the short-form physics essay in the coming ai revolution (2022). arXiv:2212.11661.","DOI":"10.1088\/1361-6552\/acc5cf"},{"key":"45644_CR31","doi-asserted-by":"publisher","unstructured":"TURING, A.\u00a0M. I.-COMPUTING MACHINERY AND INTELLIGENCE. Mind LIX, 433\u2013460, https:\/\/doi.org\/10.1093\/mind\/LIX.236.433 (1950). https:\/\/academic.oup.com\/mind\/article-pdf\/LIX\/236\/433\/30123314\/lix-236-433.pdf.","DOI":"10.1093\/mind\/LIX.236.433"},{"key":"45644_CR32","doi-asserted-by":"crossref","unstructured":"Kortemeyer, G. Could an artificial-intelligence agent pass an introductory physics course? (2023). arXiv:2301.12127.","DOI":"10.1103\/PhysRevPhysEducRes.19.010132"},{"key":"45644_CR33","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1371\/journal.pdig.0000198","volume":"2","author":"TH Kung","year":"2023","unstructured":"Kung, T. H. et al. Performance of chatgpt on usmle: Potential for ai-assisted medical education using large language models. PLOS Digital Health 2, 1\u201312. https:\/\/doi.org\/10.1371\/journal.pdig.0000198 (2023).","journal-title":"PLOS Digital Health"},{"key":"45644_CR34","unstructured":"Frieder, S. et\u00a0al. Mathematical capabilities of chatgpt (2023). arXiv:2301.13867."},{"key":"45644_CR35","unstructured":"Yuan, Z., Yuan, H., Tan, C., Wang, W. & Huang, S. How well do large language models perform in arithmetic tasks? (2023). arXiv:2304.02015."},{"key":"45644_CR36","unstructured":"Touvron, H. et\u00a0al. Llama: Open and efficient foundation language models (2023). arXiv:2302.13971."},{"key":"45644_CR37","unstructured":"Chung, H.\u00a0W. et\u00a0al. Scaling instruction-finetuned language models (2022). arXiv:2210.11416."},{"key":"45644_CR38","unstructured":"Workshop, B. et\u00a0al. Bloom: A 176b-parameter open-access multilingual language model (2023). arXiv:2211.05100."},{"key":"45644_CR39","unstructured":"Spencer, S.\u00a0T., Joshi, V. & Mitchell, A. M.\u00a0W. Can ai put gamma-ray astrophysicists out of a job? (2023). arXiv:2303.17853."},{"key":"45644_CR40","doi-asserted-by":"crossref","unstructured":"Cherian, A., Peng, K.-C., Lohit, S., Smith, K. & Tenenbaum, J.\u00a0B. Are deep neural networks smarter than second graders? (2023). arXiv:2212.09993.","DOI":"10.1109\/CVPR52729.2023.01043"},{"key":"45644_CR41","unstructured":"Stab, C. & Gurevych, I. Annotating argument components and relations in persuasive essays. In Proceedings of COLING 2014, the 25th International Conference on Computational Linguistics: Technical Papers, 1501\u20131510 (Dublin City University and Association for Computational Linguistics, Dublin, Ireland, 2014)."},{"key":"45644_CR42","unstructured":"Essay forum. https:\/\/essayforum.com\/. Last-accessed: 2023-09-07."},{"key":"45644_CR43","unstructured":"Common european framework of reference for languages (cefr). https:\/\/www.coe.int\/en\/web\/common-european-framework-reference-languages. Accessed 09 July 2023."},{"key":"45644_CR44","unstructured":"Kmk guidelines for essay assessment. http:\/\/www.kmk-format.de\/material\/Fremdsprachen\/5-3-2_Bewertungsskalen_Schreiben.pdf. Accessed 09 July 2023."},{"key":"45644_CR45","doi-asserted-by":"publisher","first-page":"57","DOI":"10.1177\/0741088309351547","volume":"27","author":"DS McNamara","year":"2010","unstructured":"McNamara, D. S., Crossley, S. A. & McCarthy, P. M. Linguistic features of writing quality. Writ. Commun. 27, 57\u201386 (2010).","journal-title":"Writ. Commun."},{"key":"45644_CR46","doi-asserted-by":"publisher","first-page":"381","DOI":"10.3758\/BRM.42.2.381","volume":"42","author":"PM McCarthy","year":"2010","unstructured":"McCarthy, P. M. & Jarvis, S. Mtld, vocd-d, and hd-d: A validation study of sophisticated approaches to lexical diversity assessment. Behav. Res. Methods 42, 381\u2013392 (2010).","journal-title":"Behav. Res. Methods"},{"key":"45644_CR47","doi-asserted-by":"crossref","unstructured":"Dasgupta, T., Naskar, A., Dey, L. & Saha, R. Augmenting textual qualitative features in deep convolution recurrent neural network for automatic essay scoring. In Proceedings of the 5th Workshop on Natural Language Processing Techniques for Educational Applications, 93\u2013102 (2018).","DOI":"10.18653\/v1\/W18-3713"},{"key":"45644_CR48","doi-asserted-by":"publisher","first-page":"554","DOI":"10.1016\/j.system.2012.10.012","volume":"40","author":"R Koizumi","year":"2012","unstructured":"Koizumi, R. & In\u2019nami, Y. Effects of text length on lexical diversity measures: Using short texts with less than 200 tokens. System 40, 554\u2013564 (2012).","journal-title":"System"},{"key":"45644_CR49","unstructured":"spacy industrial-strength natural language processing in python. https:\/\/spacy.io\/."},{"key":"45644_CR50","unstructured":"Siskou, W., Friedrich, L., Eckhard, S., Espinoza, I. & Hautli-Janisz, A. Measuring plain language in public service encounters. In Proceedings of the 2nd Workshop on Computational Linguistics for Political Text Analysis (CPSS-2022) (Potsdam, Germany, 2022)."},{"key":"45644_CR51","unstructured":"El-Assady, M. & Hautli-Janisz, A. Discourse Maps - Feature Encoding for the Analysis of Verbatim Conversation Transcripts (CSLI lecture notes (CSLI Publications, Center for the Study of Language and Information, 2019)."},{"key":"45644_CR52","unstructured":"Hautli-Janisz, A. et\u00a0al. QT30: A corpus of argument and conflict in broadcast debate. In Proceedings of the Thirteenth Language Resources and Evaluation Conference, 3291\u20133300 (European Language Resources Association, Marseille, France, 2022)."},{"key":"45644_CR53","doi-asserted-by":"publisher","first-page":"91","DOI":"10.1162\/tacl_a_00007","volume":"6","author":"S Somasundaran","year":"2018","unstructured":"Somasundaran, S. et al. Towards evaluating narrative quality in student writing. Trans. Assoc. Comput. Linguist. 6, 91\u2013106 (2018).","journal-title":"Trans. Assoc. Comput. Linguist."},{"key":"45644_CR54","doi-asserted-by":"publisher","unstructured":"Nadeem, F., Nguyen, H., Liu, Y. & Ostendorf, M. Automated essay scoring with discourse-aware neural models. In Proceedings of the Fourteenth Workshop on Innovative Use of NLP for Building Educational Applications, 484\u2013493, https:\/\/doi.org\/10.18653\/v1\/W19-4450 (Association for Computational Linguistics, Florence, Italy, 2019).","DOI":"10.18653\/v1\/W19-4450"},{"key":"45644_CR55","unstructured":"Prasad, R. et\u00a0al. The Penn Discourse TreeBank 2.0. In Proceedings of the Sixth International Conference on Language Resources and Evaluation (LREC\u201908) (European Language Resources Association (ELRA), Marrakech, Morocco, 2008)."},{"key":"45644_CR56","doi-asserted-by":"publisher","first-page":"297","DOI":"10.1007\/bf02310555","volume":"16","author":"LJ Cronbach","year":"1951","unstructured":"Cronbach, L. J. Coefficient alpha and the internal structure of tests. Psychometrika 16, 297\u2013334. https:\/\/doi.org\/10.1007\/bf02310555 (1951).","journal-title":"Psychometrika"},{"key":"45644_CR57","doi-asserted-by":"publisher","first-page":"80","DOI":"10.2307\/3001968","volume":"1","author":"F Wilcoxon","year":"1945","unstructured":"Wilcoxon, F. Individual comparisons by ranking methods. Biom. Bull. 1, 80\u201383 (1945).","journal-title":"Biom. Bull."},{"key":"45644_CR58","first-page":"65","volume":"6","author":"S Holm","year":"1979","unstructured":"Holm, S. A simple sequentially rejective multiple test procedure. Scand. J. Stat. 6, 65\u201370 (1979).","journal-title":"Scand. J. Stat."},{"key":"45644_CR59","doi-asserted-by":"crossref","unstructured":"Cohen, J. Statistical power analysis for the behavioral sciences (Academic press, 2013).","DOI":"10.4324\/9780203771587"},{"key":"45644_CR60","unstructured":"Freedman, D., Pisani, R. & Purves, R. Statistics (international student edition). Pisani, R. Purves, 4th edn. WW Norton & Company, New York (2007)."},{"key":"45644_CR61","unstructured":"Scipy documentation. https:\/\/docs.scipy.org\/doc\/scipy\/reference\/generated\/scipy.stats.pearsonr.html. Accessed 09 June 2023."},{"key":"45644_CR62","doi-asserted-by":"publisher","first-page":"131","DOI":"10.3102\/00346543072002131","volume":"72","author":"M Windschitl","year":"2002","unstructured":"Windschitl, M. Framing constructivism in practice as the negotiation of dilemmas: An analysis of the conceptual, pedagogical, cultural, and political challenges facing teachers. Rev. Educ. Res. 72, 131\u2013175 (2002).","journal-title":"Rev. Educ. Res."}],"container-title":["Scientific Reports"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.nature.com\/articles\/s41598-023-45644-9.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/www.nature.com\/articles\/s41598-023-45644-9","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/www.nature.com\/articles\/s41598-023-45644-9.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2023,10,30]],"date-time":"2023-10-30T08:04:55Z","timestamp":1698653095000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.nature.com\/articles\/s41598-023-45644-9"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,10,30]]},"references-count":62,"journal-issue":{"issue":"1","published-online":{"date-parts":[[2023,12]]}},"alternative-id":["45644"],"URL":"https:\/\/doi.org\/10.1038\/s41598-023-45644-9","relation":{},"ISSN":["2045-2322"],"issn-type":[{"value":"2045-2322","type":"electronic"}],"subject":[],"published":{"date-parts":[[2023,10,30]]},"assertion":[{"value":"1 June 2023","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"22 October 2023","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"30 October 2023","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"The authors declare no competing interests.","order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Competing interests"}}],"article-number":"18617"}}