{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,10]],"date-time":"2026-07-10T16:28:24Z","timestamp":1783700904229,"version":"3.55.0"},"reference-count":120,"publisher":"Association for Computing Machinery (ACM)","issue":"2","license":[{"start":{"date-parts":[[2025,1,22]],"date-time":"2025-01-22T00:00:00Z","timestamp":1737504000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by-nd\/4.0\/"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Inf. Syst."],"published-print":{"date-parts":[[2025,3,31]]},"abstract":"<jats:p>The use of machine learning (ML) models to assess and score textual data has become increasingly pervasive in an array of contexts including natural language processing, information retrieval, search and recommendation, and credibility assessment of online content. A significant disruption at the intersection of ML and text are text-generating large-language models (LLMs) such as generative pre-trained transformers (GPTs). We empirically assess the differences in how ML-based scoring models trained on human content assess the quality of content generated by humans versus GPTs. To do so, we propose an analysis framework that encompasses essay scoring ML models, human- and ML-generated essays, and a statistical model that parsimoniously considers the impact of type of respondent, prompt genre, and the ML model used for assessment model. A rich testbed is utilized that encompasses 18,460 human-generated and GPT-based essays. Results of our benchmark analysis reveal that LLMs and transformer pretrained language models (PLMs) more accurately score human essay quality as compared to CNN\/RNN and feature-based ML methods. Interestingly, we find that LLMs and transformer PLMs tend to score GPT-generated text 10\u201320% higher on average, relative to human-authored documents. Conversely, traditional deep learning and feature-based ML models score human text considerably higher. Further analysis reveals that even though the LLMs and transformer PLMs are exclusively fine-tuned on human text, they more prominently attend to certain tokens appearing only in GPT-generated text, possibly (in part) due to familiarity\/overlap in pre-training. Our framework and results have implications for text classification settings where automated scoring of text is likely to be disrupted by generative AI.<\/jats:p>","DOI":"10.1145\/3702639","type":"journal-article","created":{"date-parts":[[2024,11,5]],"date-time":"2024-11-05T17:43:58Z","timestamp":1730828638000},"page":"1-36","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":27,"title":["When Automated Assessment Meets Automated Content Generation: Examining Text Quality in the Era of GPTs"],"prefix":"10.1145","volume":"43","author":[{"ORCID":"https:\/\/orcid.org\/0009-0001-1641-6428","authenticated-orcid":false,"given":"Marialena","family":"Bevilacqua","sequence":"first","affiliation":[{"name":"University of Notre Dame, Notre Dame, Indiana, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0000-1089-2530","authenticated-orcid":false,"given":"Kezia","family":"Oketch","sequence":"additional","affiliation":[{"name":"University of Notre Dame, Notre Dame, Indiana, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-0827-2257","authenticated-orcid":false,"given":"Ruiyang","family":"Qin","sequence":"additional","affiliation":[{"name":"University of Notre Dame, Notre Dame, Indiana, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0006-0374-5878","authenticated-orcid":false,"given":"Will","family":"Stamey","sequence":"additional","affiliation":[{"name":"University of Notre Dame, Notre Dame, Indiana, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-5119-0073","authenticated-orcid":false,"given":"Xinyuan","family":"Zhang","sequence":"additional","affiliation":[{"name":"University of Notre Dame, Notre Dame, Indiana, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0003-0852-0705","authenticated-orcid":false,"given":"Yi","family":"Gan","sequence":"additional","affiliation":[{"name":"Georgia Tech, Atlanta, Georgia, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-6371-7741","authenticated-orcid":false,"given":"Kai","family":"Yang","sequence":"additional","affiliation":[{"name":"Shenzhen University, Shenzhen, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-7698-7794","authenticated-orcid":false,"given":"Ahmed","family":"Abbasi","sequence":"additional","affiliation":[{"name":"University of Notre Dame, Notre Dame, Indiana, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2025,1,22]]},"reference":[{"key":"e_1_3_2_2_2","doi-asserted-by":"publisher","DOI":"10.1145\/1361684.1361685"},{"key":"e_1_3_2_3_2","doi-asserted-by":"publisher","DOI":"10.17705\/1jais.00849"},{"key":"e_1_3_2_4_2","doi-asserted-by":"publisher","DOI":"10.1109\/TKDE.2010.110"},{"key":"e_1_3_2_5_2","doi-asserted-by":"publisher","DOI":"10.1287\/isre.2024.editorial.v35.n2"},{"key":"e_1_3_2_6_2","doi-asserted-by":"publisher","DOI":"10.1145\/2382438.2382441"},{"issue":"2","key":"e_1_3_2_7_2","article-title":"Text analytics to support sense-making in social media: A language-action perspective","volume":"42","author":"Abbasi Ahmed","year":"2018","unstructured":"Ahmed Abbasi, Yilu Zhou, Shasha Deng, and Pengzhu Zhang. 2018. Text analytics to support sense-making in social media: A language-action perspective. MIS Quarterly 42, 2 (2018).","journal-title":"MIS Quarterly"},{"key":"e_1_3_2_8_2","doi-asserted-by":"publisher","DOI":"10.1145\/3365211"},{"key":"e_1_3_2_9_2","doi-asserted-by":"crossref","unstructured":"Dimitrios Alikaniotis Helen Yannakoudakis and Marek Rei. 2016. Automatic text scoring using neural networks. arXiv:1606.04289. Retrieved from https:\/\/arxiv.org\/abs\/1606.04289","DOI":"10.18653\/v1\/P16-1068"},{"key":"e_1_3_2_10_2","first-page":"229","volume-title":"Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies","volume":"1","author":"Amorim Evelin","year":"2018","unstructured":"Evelin Amorim, Marcia Can\u00e7ado, and Adriano Veloso. 2018. Automated essay scoring in the presence of biased ratings. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Vol. 1 (Long Papers), 229\u2013237."},{"key":"e_1_3_2_11_2","unstructured":"Rohan Anil Andrew M. Dai Orhan Firat Melvin Johnson Dmitry Lepikhin Alexandre Passos Siamak Shakeri Emanuel Taropa Paige Bailey Zhifeng Chen et al. 2023. Palm 2 technical report. arXiv:2305.10403. Retrieved from https:\/\/arxiv.org\/abs\/2305.10403"},{"key":"e_1_3_2_12_2","doi-asserted-by":"publisher","DOI":"10.1145\/45941.214328"},{"key":"e_1_3_2_13_2","doi-asserted-by":"publisher","DOI":"10.1186\/s13054-023-04393-x"},{"key":"e_1_3_2_14_2","first-page":"2200","volume-title":"Proceedings of the 7th International Conference on Language Resources and Evaluation (LREC \u201910)","volume":"10","author":"Baccianella Stefano","year":"2010","unstructured":"Stefano Baccianella, Andrea Esuli, and Fabrizio Sebastiani. 2010. Sentiwordnet 3.0: An enhanced lexical resource for sentiment analysis and opinion mining. In Proceedings of the 7th International Conference on Language Resources and Evaluation (LREC \u201910), Vol. 10, 2200\u20132204."},{"key":"e_1_3_2_15_2","doi-asserted-by":"publisher","unstructured":"Dzmitry Bahdanau Kyunghyun Cho and Yoshua Bengio. 2014. Neural machine translation by jointly learning to align and translate. arXiv:1409.0473. Retrieved from https:\/\/doi.org\/10.48550\/arXiv.1409.0473","DOI":"10.48550\/arXiv.1409.0473"},{"key":"e_1_3_2_16_2","doi-asserted-by":"publisher","DOI":"10.61969\/jai.1337500"},{"key":"e_1_3_2_17_2","first-page":"40","article-title":"Assessing writing in MOOCs: Automated essay scoring and calibrated peer review\u2122","volume":"8","author":"Balfour Stephen P.","year":"2013","unstructured":"Stephen P. Balfour. 2013. Assessing writing in MOOCs: Automated essay scoring and calibrated peer review\u2122. Research & Practice in Assessment 8 (2013), 40\u201348.","journal-title":"Research & Practice in Assessment"},{"issue":"2","key":"e_1_3_2_18_2","doi-asserted-by":"crossref","first-page":"46","DOI":"10.1023\/A:1015570104121","article-title":"Hospitalisations caused by adverse drug reactions (ADR): A meta-analysis of observational studies","volume":"24","author":"Beijer H. J. M.","year":"2002","unstructured":"H. J. M. Beijer and C. J. De Blaey. 2002. Hospitalisations caused by adverse drug reactions (ADR): A meta-analysis of observational studies. Pharmacy World and Science 24, 2 (2002), 46\u201354.","journal-title":"Pharmacy World and Science"},{"issue":"3","key":"e_1_3_2_19_2","first-page":"1433","article-title":"Managing artificial intelligence","volume":"45","author":"Berente Nicholas","year":"2021","unstructured":"Nicholas Berente, Bin Gu, Jan Recker, and Radhika Santhanam. 2021. Managing artificial intelligence. MIS Quarterly 45, 3 (2021), 1433\u20131450.","journal-title":"MIS Quarterly"},{"key":"e_1_3_2_20_2","first-page":"92","volume-title":"Proceedings of the 14th Workshop on Innovative Use of NLP for Building Educational Applications","author":"Berggren Stig Johan","year":"2019","unstructured":"Stig Johan Berggren, Taraka Rama, and Lilja \u00d8vrelid. 2019. Regression or classification? Automated essay scoring for Norwegian. In Proceedings of the 14th Workshop on Innovative Use of NLP for Building Educational Applications, 92\u2013102."},{"key":"e_1_3_2_21_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.caeai.2022.100081"},{"key":"e_1_3_2_22_2","doi-asserted-by":"crossref","first-page":"735","DOI":"10.1109\/CSSE.2008.623","volume-title":"Proceedings of the 2008 International Conference on Computer Science and Software Engineering","volume":"1","author":"Bin Li","year":"2008","unstructured":"Li Bin, Lu Jun, Yao Jian-Min, and Zhu Qiao-Ming. 2008. Automated essay scoring using the KNN algorithm. In Proceedings of the 2008 International Conference on Computer Science and Software Engineering, Vol. 1. IEEE, 735\u2013738."},{"issue":"1","key":"e_1_3_2_23_2","doi-asserted-by":"crossref","first-page":"27","DOI":"10.1080\/08957347.2012.635502","article-title":"Comparison of human and machine scoring of essays: Differences by gender, ethnicity, and country","volume":"25","author":"Bridgeman Brent","year":"2012","unstructured":"Brent Bridgeman, Catherine Trapani, and Yigal Attali. 2012. Comparison of human and machine scoring of essays: Differences by gender, ethnicity, and country. Applied Measurement in Education 25, 1 (2012), 27\u201340.","journal-title":"Applied Measurement in Education"},{"key":"e_1_3_2_24_2","doi-asserted-by":"crossref","DOI":"10.4324\/9781315162058","volume-title":"Assessment of Student Achievement","author":"Brown Gavin T. L.","year":"2017","unstructured":"Gavin T. L. Brown. 2017. Assessment of Student Achievement. Routledge."},{"key":"e_1_3_2_25_2","first-page":"1877","volume-title":"Proceedings of the Advances in Neural Information Processing Systems","volume":"33","author":"Brown Tom","year":"2020","unstructured":"Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020. Language models are few-shot learners. In Proceedings of the Advances in Neural Information Processing Systems, Vol. 33, 1877\u20131901."},{"key":"e_1_3_2_26_2","doi-asserted-by":"publisher","DOI":"10.1145\/185462.185477"},{"key":"e_1_3_2_27_2","doi-asserted-by":"publisher","DOI":"10.1145\/2939672.2939785"},{"key":"e_1_3_2_28_2","doi-asserted-by":"publisher","unstructured":"Yi Chen Rui Wang Haiyun Jiang Shuming Shi and Ruifeng Xu. 2023. Exploring the use of large language models for reference-free text quality evaluation: A preliminary empirical study. arXiv:2304.00723. Retrieved from https:\/\/doi.org\/10.48550\/arXiv.2304.00723","DOI":"10.48550\/arXiv.2304.00723"},{"key":"e_1_3_2_29_2","doi-asserted-by":"publisher","DOI":"10.1145\/3558548"},{"key":"e_1_3_2_30_2","doi-asserted-by":"publisher","unstructured":"Cheng-Han Chiang and Hung-Yi Lee. 2023. Can large language models be an alternative to human evaluations? arXiv:2305.01937. Retrieved from https:\/\/doi.org\/10.48550\/arXiv.2305.01937","DOI":"10.48550\/arXiv.2305.01937"},{"key":"e_1_3_2_31_2","unstructured":"Jonathan H. Choi Kristin E. Hickman Amy Monahan and Daniel Schwarcz. 2023. ChatGPT Goes to Law School. Retrieved from https:\/\/ssrn.com\/abstract=4335905"},{"key":"e_1_3_2_32_2","doi-asserted-by":"crossref","unstructured":"M\u0103d\u0103lina Cozma Andrei M. Butnaru and Radu Tudor Ionescu. 2018. Automated essay scoring with string kernels and word embeddings. arXiv:1804.07954. Retrieved from https:\/\/arxiv.org\/abs\/1804.07954","DOI":"10.18653\/v1\/P18-2080"},{"issue":"3","key":"e_1_3_2_33_2","doi-asserted-by":"crossref","first-page":"415","DOI":"10.17239\/jowr-2020.11.03.01","article-title":"Linguistic features in writing quality and development: An overview","volume":"11","author":"Crossley Scott A.","year":"2020","unstructured":"Scott A. Crossley. 2020. Linguistic features in writing quality and development: An overview. Journal of Writing Research 11, 3 (2020), 415\u2013443.","journal-title":"Journal of Writing Research"},{"key":"e_1_3_2_34_2","doi-asserted-by":"publisher","unstructured":"Sumanth Dathathri Andrea Madotto Janice Lan Jane Hung Eric Frank Piero Molino Jason Yosinski and Rosanne Liu. 2019. Plug and play language models: A simple approach to controlled text generation. arXiv:1912.02164. Retrieved from https:\/\/doi.org\/10.48550\/arXiv.1912.02164","DOI":"10.48550\/arXiv.1912.02164"},{"key":"e_1_3_2_35_2","unstructured":"Jacob Devlin Ming-Wei Chang Kenton Lee and Kristina Toutanova. 2018. Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv:1810.04805. Retrieved from https:\/\/ar5iv.labs.arxiv.org\/html\/1810.04805"},{"issue":"1","key":"e_1_3_2_36_2","first-page":"4","article-title":"An overview of automated scoring of essays","volume":"5","author":"Dikli Semire","year":"2006","unstructured":"Semire Dikli. 2006. An overview of automated scoring of essays. The Journal of Technology, Learning and Assessment 5, 1 (2006), 4\u201328.","journal-title":"The Journal of Technology, Learning and Assessment"},{"key":"e_1_3_2_37_2","doi-asserted-by":"crossref","first-page":"1072","DOI":"10.18653\/v1\/D16-1115","volume-title":"Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing","author":"Dong Fei","year":"2016","unstructured":"Fei Dong and Yue Zhang. 2016. Automatic features for essay scoring\u2013an empirical study. In Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing, 1072\u20131077."},{"key":"e_1_3_2_38_2","doi-asserted-by":"publisher","DOI":"10.5121\/csit.2021.111901"},{"key":"e_1_3_2_39_2","volume-title":"Proceedings of the National Academy of Sciences","volume":"119","author":"Drori Iddo","year":"2022","unstructured":"Iddo Drori, Sarah Zhang, Reece Shuttleworth, Leonard Tang, Albert Lu, Elizabeth Ke, Kevin Liu, Linda Chen, Sunny Tran, Newman Cheng, et al. 2022. A neural network solves, explains, and generates university math problems by program synthesis and few-shot learning at human level. Proceedings of the National Academy of Sciences 119, 32 (2022), e2123433119."},{"key":"e_1_3_2_40_2","doi-asserted-by":"publisher","unstructured":"Hanyu Duan Yixuan Tang Yi Yang Ahmed Abbasi and Kar Yan Tam. 2023. Exploring the relationship between in-context learning and instruction tuning. arXiv:2311.10367. Retrieved from https:\/\/doi.org\/10.48550\/arXiv.2311.10367","DOI":"10.48550\/arXiv.2311.10367"},{"key":"e_1_3_2_41_2","first-page":"14","volume-title":"Proceedings of the International AAAI Conference on Web and Social Media","volume":"7","author":"Farnadi Golnoosh","year":"2013","unstructured":"Golnoosh Farnadi, Susana Zoghbi, Marie-Francine Moens, and Martine De Cock. 2013. Recognising personality traits using facebook status updates. In Proceedings of the International AAAI Conference on Web and Social Media. D. Archambault, E. Kandogan, M. Harrigan, and Organizers (Eds.), Vol. 7, 14\u201318."},{"key":"e_1_3_2_42_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-90-481-8847-5_10"},{"key":"e_1_3_2_43_2","doi-asserted-by":"publisher","unstructured":"Jessica Ficler and Yoav Goldberg. 2017. Controlling linguistic style aspects in neural language generation. arXiv:1707.02633. Retrieved from https:\/\/doi.org\/10.48550\/arXiv.1707.02633","DOI":"10.48550\/arXiv.1707.02633"},{"key":"e_1_3_2_44_2","doi-asserted-by":"publisher","DOI":"10.1287\/isre.2021.1079"},{"key":"e_1_3_2_45_2","unstructured":"Xinyang Geng and Hao Liu. 2023. OpenLLaMA: An Open Reproduction of LLaMA. Retrieved from https:\/\/github.com\/openlm-research\/open_llama"},{"key":"e_1_3_2_46_2","first-page":"1012","volume-title":"Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics","volume":"1","author":"Guo Yue","year":"2022","unstructured":"Yue Guo, Yi Yang, and Ahmed Abbasi. 2022. Auto-debias: Debiasing masked language models with automated biased prompts. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics, Vol. 1 (Long Papers), 1012\u20131023."},{"key":"e_1_3_2_47_2","first-page":"853","volume-title":"Proceedings of the International Conference on Artificial Intelligence and Smart Communication (AISC \u201923)","author":"Gupta Kshitij","year":"2023","unstructured":"Kshitij Gupta. 2023. Data augmentation for automated essay scoring using transformer models. In Proceedings of the International Conference on Artificial Intelligence and Smart Communication (AISC \u201923). IEEE, 853\u2013857."},{"key":"e_1_3_2_48_2","unstructured":"Edward J Hu Yelong Shen Phillip Wallis Zeyuan Allen-Zhu Yuanzhi Li Shean Wang Lu Wang and Weizhu Chen. 2021. Lora: Low-rank adaptation of large language models. arXiv:2106.09685."},{"key":"e_1_3_2_49_2","doi-asserted-by":"publisher","DOI":"10.1109\/JBHI.2024.3352075"},{"key":"e_1_3_2_50_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/P18-1100"},{"key":"e_1_3_2_51_2","doi-asserted-by":"publisher","DOI":"10.1145\/3603374"},{"key":"e_1_3_2_52_2","doi-asserted-by":"publisher","DOI":"10.24963\/ijcai.2019\/879"},{"key":"e_1_3_2_53_2","first-page":"2","volume-title":"Proceedings of the naacL-HLT","volume":"1","author":"Kenton Jacob Devlin Ming-Wei Chang","year":"2019","unstructured":"Jacob Devlin Ming-Wei Chang Kenton and Lee Kristina Toutanova. 2019. BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the naacL-HLT, Vol. 1, 2."},{"key":"e_1_3_2_54_2","doi-asserted-by":"publisher","DOI":"10.3115\/v1\/D14-1181"},{"key":"e_1_3_2_55_2","doi-asserted-by":"publisher","DOI":"10.1007\/s12195-022-00754-8"},{"key":"e_1_3_2_56_2","doi-asserted-by":"crossref","first-page":"899","DOI":"10.25300\/MISQ\/2023\/17381","article-title":"Timely, granular, and actionable: Designing a social listening platform for Public Health 3.0","volume":"48","author":"Kitchens Brent","year":"2024","unstructured":"Brent Kitchens, Jennifer Claggett, and Ahmed Abbasi. 2024. Timely, granular, and actionable: Designing a social listening platform for Public Health 3.0. Management Information Systems Quarterly 48 (2024), 899\u2013930.","journal-title":"Management Information Systems Quarterly"},{"key":"e_1_3_2_57_2","doi-asserted-by":"publisher","DOI":"10.1371\/journal.pdig.0000198"},{"key":"e_1_3_2_58_2","doi-asserted-by":"publisher","DOI":"10.1145\/3641276"},{"key":"e_1_3_2_59_2","first-page":"3598","volume-title":"Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies","author":"Lalor John P.","year":"2022","unstructured":"John P. Lalor, Yi Yang, Kendall Smith, Nicole Forsgren, and Ahmed Abbasi. 2022. Benchmarking intersectional biases in NLP. In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, 3598\u20133609."},{"key":"e_1_3_2_60_2","doi-asserted-by":"publisher","DOI":"10.1145\/3529954"},{"key":"e_1_3_2_61_2","doi-asserted-by":"publisher","DOI":"10.1016\/S2589-7500(23)00019-5"},{"key":"e_1_3_2_62_2","doi-asserted-by":"publisher","unstructured":"Yang Liu Dan Iter Yichong Xu Shuohang Wang Ruochen Xu and Chenguang Zhu. 2023. GPTEVAL: NLG evaluation using GPT-4 with better human alignment. arXiv:2303.16634. Retrieved from https:\/\/doi.org\/10.48550\/arXiv.2303.16634","DOI":"10.48550\/arXiv.2303.16634"},{"key":"e_1_3_2_63_2","doi-asserted-by":"publisher","unstructured":"Yinhan Liu Myle Ott Naman Goyal Jingfei Du Mandar Joshi Danqi Chen Omer Levy Mike Lewis Luke Zettlemoyer and Veselin Stoyanov. 2019. Roberta: A robustly optimized bert pretraining approach. arXiv:1907.11692. Retrieved from https:\/\/doi.org\/10.48550\/arXiv.1907.11692","DOI":"10.48550\/arXiv.1907.11692"},{"key":"e_1_3_2_64_2","doi-asserted-by":"publisher","DOI":"10.1145\/3386253"},{"key":"e_1_3_2_65_2","doi-asserted-by":"publisher","DOI":"10.1108\/LHTN-01-2023-0009"},{"key":"e_1_3_2_66_2","doi-asserted-by":"publisher","DOI":"10.1109\/MIS.2017.23"},{"key":"e_1_3_2_67_2","doi-asserted-by":"publisher","DOI":"10.3115\/v1\/P14-5010"},{"key":"e_1_3_2_68_2","first-page":"26","volume-title":"Proceedings of the Advances in Neural Information Processing Systems","author":"Mikolov Tomas","year":"2013","unstructured":"Tomas Mikolov, Ilya Sutskever, Kai Chen, Greg S. Corrado, and Jeff Dean. 2013. Distributed representations of words and phrases and their compositionality. In Proceedings of the Advances in Neural Information Processing Systems, Vol. 26."},{"key":"e_1_3_2_69_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.rmal.2023.100050"},{"key":"e_1_3_2_70_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.neucom.2021.03.091"},{"key":"e_1_3_2_71_2","doi-asserted-by":"publisher","unstructured":"OpenAI. 2023. GPT-4 technical report. arXiv:2303.08774. Retrieved from https:\/\/doi.org\/10.48550\/arXiv.2303.08774","DOI":"10.48550\/arXiv.2303.08774"},{"issue":"5","key":"e_1_3_2_72_2","first-page":"238","article-title":"The imminence of\u2026 grading essays by computer","volume":"47","author":"Page Ellis B.","year":"1966","unstructured":"Ellis B. Page. 1966. The imminence of\u2026 grading essays by computer. The Phi Delta Kappan 47, 5 (1966), 238\u2013243.","journal-title":"The Phi Delta Kappan"},{"key":"e_1_3_2_73_2","doi-asserted-by":"publisher","DOI":"10.1007\/BF01419938"},{"key":"e_1_3_2_74_2","doi-asserted-by":"crossref","first-page":"149","DOI":"10.1109\/IUCS.2010.5666229","volume-title":"Proceedings of the 2010 4th International Universal Communication Symposium","author":"Peng Xingyuan","year":"2010","unstructured":"Xingyuan Peng, Dengfeng Ke, Zhenbiao Chen, and Bo Xu. 2010. Automated Chinese essay scoring using vector space models. In Proceedings of the 2010 4th International Universal Communication Symposium. IEEE, 149\u2013153."},{"key":"e_1_3_2_75_2","first-page":"135","volume-title":"Linguistic Inquiry and Word Count: LIWC [Computer Software]","author":"Pennebaker James W.","year":"2007","unstructured":"James W. Pennebaker, Roger J. Booth, and Martha E. Francis. 2007. Linguistic Inquiry and Word Count: LIWC [Computer Software]. liwc. net, 135."},{"key":"e_1_3_2_76_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/D15-1049"},{"key":"e_1_3_2_77_2","doi-asserted-by":"publisher","DOI":"10.1080\/09296170802326699"},{"key":"e_1_3_2_78_2","doi-asserted-by":"crossref","first-page":"170","DOI":"10.1109\/ICODSE.2015.7436992","volume-title":"Proceedings of the International Conference on Data and Software Engineering (ICoDSE \u201915)","author":"Pratama Bayu Yudha","year":"2015","unstructured":"Bayu Yudha Pratama and Riyanarto Sarno. 2015. Personality classification based on Twitter text using Naive Bayes, KNN and SVM. In Proceedings of the International Conference on Data and Software Engineering (ICoDSE \u201915). IEEE, 170\u2013174."},{"key":"e_1_3_2_79_2","doi-asserted-by":"publisher","DOI":"10.1145\/3624475"},{"issue":"8","key":"e_1_3_2_80_2","first-page":"9","article-title":"Language models are unsupervised multitask learners","volume":"1","author":"Radford Alec","year":"2019","unstructured":"Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019. Language models are unsupervised multitask learners. OpenAI Blog 1, 8 (2019), 9.","journal-title":"OpenAI Blog"},{"key":"e_1_3_2_81_2","doi-asserted-by":"publisher","DOI":"10.1007\/s10462-021-10068-2"},{"key":"e_1_3_2_82_2","doi-asserted-by":"publisher","unstructured":"Pedro Uria Rodriguez Amir Jafari and Christopher M. Ormerod. 2019. Language models and automated essay scoring. arXiv:1909.09482. Retrieved from https:\/\/doi.org\/10.48550\/arXiv.1909.09482","DOI":"10.48550\/arXiv.1909.09482"},{"key":"e_1_3_2_83_2","doi-asserted-by":"publisher","DOI":"10.1162\/tacl_a_00349"},{"issue":"1","key":"e_1_3_2_84_2","first-page":"26","article-title":"An overview of three approaches to scoring written essays by computer","volume":"7","author":"Rudner Lawrence M.","year":"2000","unstructured":"Lawrence M. Rudner and Phill Gagne. 2000. An overview of three approaches to scoring written essays by computer. Practical Assessment, Research, and Evaluation 7, 1 (2000), 26.","journal-title":"Practical Assessment, Research, and Evaluation"},{"issue":"10","key":"e_1_3_2_85_2","doi-asserted-by":"crossref","first-page":"1207","DOI":"10.1037\/apl0000405","article-title":"Using machine learning to translate applicant work history into predictors of performance and turnover","volume":"104","author":"Sajjadiani Sima","year":"2019","unstructured":"Sima Sajjadiani, Aaron J. Sojourner, John D. Kammeyer-Mueller, and Elton Mykerezi. 2019. Using machine learning to translate applicant work history into predictors of performance and turnover. Journal of Applied Psychology 104, 10 (2019), 1207.","journal-title":"Journal of Applied Psychology"},{"key":"e_1_3_2_86_2","doi-asserted-by":"publisher","DOI":"10.1145\/3494833"},{"key":"e_1_3_2_87_2","first-page":"1","volume-title":"Proceedings of the IEEE International Conference on Engineering, Technology and Education (TALE \u201919)","author":"Salim Yafet","year":"2019","unstructured":"Yafet Salim, Valdi Stevanus, Edwardo Barlian, Azani Cempaka Sari, and Derwin Suhartono. 2019. Automated English digital essay grader using machine learning. In Proceedings of the IEEE International Conference on Engineering, Technology and Education (TALE \u201919). IEEE, 1\u20136."},{"key":"e_1_3_2_88_2","volume-title":"C4. 5: Programs for Machine Learning by J. Ross Quinlan","author":"Salzberg Steven L.","year":"1994","unstructured":"Steven L. Salzberg. 1994. C4. 5: Programs for Machine Learning by J. Ross Quinlan. Morgan Kaufmann Publishers, Inc., 1993."},{"key":"e_1_3_2_89_2","volume-title":"Proceedings of the 8th Italian Conference on Computational Linguistics","author":"Schmalz Veronica Juliana","year":"2021","unstructured":"Veronica Juliana Schmalz and Alessio Brutti. 2021. Automatic assessment of English CEFR levels using BERT Embeddings. In Proceedings of the 8th Italian Conference on Computational Linguistics."},{"key":"e_1_3_2_90_2","doi-asserted-by":"publisher","DOI":"10.1145\/3210753"},{"key":"e_1_3_2_91_2","doi-asserted-by":"publisher","DOI":"10.2196\/48517"},{"key":"e_1_3_2_92_2","doi-asserted-by":"publisher","DOI":"10.1145\/3569929"},{"key":"e_1_3_2_93_2","unstructured":"Mark D. Shermis and Felicia D. Barrera. 2002. Exit assessments: Evaluating writing ability through automated essay scoring. (2002). Retrieved from https:\/\/www.researchgate.net\/publication\/234726160_Assessing_Writing_through_the_Curriculum_with_Automated_Essay_Scoring"},{"key":"e_1_3_2_94_2","doi-asserted-by":"publisher","DOI":"10.4324\/9781410606860"},{"key":"e_1_3_2_95_2","first-page":"2917","volume-title":"Proceedings of the 29th International Conference on Computational Linguistics","author":"Shibata Takumi","year":"2022","unstructured":"Takumi Shibata and Masaki Uto. 2022. Analytic automated essay scoring based on deep neural networks integrating multidimensional item response theory. In Proceedings of the 29th International Conference on Computational Linguistics, 2917\u20132926."},{"key":"e_1_3_2_96_2","doi-asserted-by":"publisher","DOI":"10.1038\/s41586-023-06291-2"},{"key":"e_1_3_2_97_2","doi-asserted-by":"publisher","DOI":"10.1145\/3578931"},{"key":"e_1_3_2_98_2","volume-title":"Proceedings of the 4th International Conference on Language Resources and Evaluation (LREC \u201904)","volume":"4","author":"Strapparava Carlo","year":"2004","unstructured":"Carlo Strapparava and Alessandro Valitutti. 2004. WordNet affect: An affective extension of WordNet.. In Proceedings of the 4th International Conference on Language Resources and Evaluation (LREC \u201904), Vol. 4, 40."},{"key":"e_1_3_2_99_2","first-page":"469","volume-title":"Proceedings of the 20th International Conference on Artificial Intelligence in Education (AIED \u201919)","author":"Sung Chul","year":"2019","unstructured":"Chul Sung, Tejas Indulal Dhamecha, and Nirmal Mukhi. 2019. Improving short answer grading using transformer-based pre-training. In Proceedings of the 20th International Conference on Artificial Intelligence in Education (AIED \u201919). Springer, 469\u2013481."},{"key":"e_1_3_2_100_2","doi-asserted-by":"publisher","DOI":"10.1109\/ACCESS.2018.2876502"},{"key":"e_1_3_2_101_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/D16-1193"},{"key":"e_1_3_2_102_2","unstructured":"Christian Terwiesch. 2023. Would Chat GPT3 Get a Wharton MBA? A Prediction Based on Its Performance in the Operations Management Course. Mack Institute for Innovation Management at the Wharton School University of Pennsylvania."},{"key":"e_1_3_2_103_2","doi-asserted-by":"publisher","DOI":"10.1126\/science.adg7879"},{"key":"e_1_3_2_104_2","doi-asserted-by":"publisher","DOI":"10.1007\/s11409-011-9072-x"},{"key":"e_1_3_2_105_2","doi-asserted-by":"publisher","unstructured":"Hugo Touvron Thibaut Lavril Gautier Izacard Xavier Martinet Marie-Anne Lachaux Timoth\u00e9e Lacroix Baptiste Rozi\u00e8re Naman Goyal Eric Hambro Faisal Azhar et al. 2023. LLaMA: Open and efficient foundation language models. arXiv:2302.13971. Retrieved from https:\/\/doi.org\/10.48550\/arXiv.2302.13971","DOI":"10.48550\/arXiv.2302.13971"},{"key":"e_1_3_2_106_2","doi-asserted-by":"publisher","DOI":"10.1007\/s41237-021-00142-y"},{"key":"e_1_3_2_107_2","doi-asserted-by":"publisher","DOI":"10.1109\/TLT.2022.3145352"},{"key":"e_1_3_2_108_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2020.coling-main.535"},{"key":"e_1_3_2_109_2","doi-asserted-by":"publisher","DOI":"10.28945\/331"},{"key":"e_1_3_2_110_2","volume-title":"Proceedings of the Advances in Neural Information Processing Systems","volume":"30","author":"Vaswani Ashish","year":"2017","unstructured":"Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, \u0141ukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. In Proceedings of the Advances in Neural Information Processing Systems, Vol. 30."},{"key":"e_1_3_2_111_2","doi-asserted-by":"publisher","unstructured":"Jesse Vig. 2019. A multiscale visualization of attention in the transformer model. arXiv:1906.05714. Retrieved from https:\/\/doi.org\/10.48550\/arXiv.1906.05714","DOI":"10.48550\/arXiv.1906.05714"},{"key":"e_1_3_2_112_2","doi-asserted-by":"publisher","DOI":"10.1145\/3530257"},{"key":"e_1_3_2_113_2","first-page":"576","volume-title":"Proceedings of the 18th Workshop on Innovative Use of NLP for Building Educational Applications (BEA \u201923)","author":"Yancey Kevin P","year":"2023","unstructured":"Kevin P Yancey, Geoffrey Laflair, Anthony Verardi, and Jill Burstein. 2023. Rating short L2 essays on the CEFR scale with GPT-4. In Proceedings of the 18th Workshop on Innovative Use of NLP for Building Educational Applications (BEA \u201923), 576\u2013584."},{"key":"e_1_3_2_114_2","doi-asserted-by":"publisher","DOI":"10.1287\/isre.2022.1111"},{"key":"e_1_3_2_115_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.ymssp.2020.106885"},{"key":"e_1_3_2_116_2","doi-asserted-by":"publisher","DOI":"10.1145\/502795.502798"},{"key":"e_1_3_2_117_2","first-page":"383","volume-title":"Proceedings of the IEEE 8th International Conference on Awareness Science and Technology (iCAST \u201917)","author":"Yu Jianguo","year":"2017","unstructured":"Jianguo Yu and Konstantin Markov. 2017. Deep learning based personality recognition from facebook status updates. In Proceedings of the IEEE 8th International Conference on Awareness Science and Technology (iCAST \u201917). IEEE, 383\u2013387."},{"key":"e_1_3_2_118_2","doi-asserted-by":"publisher","DOI":"10.1177\/07356331221127300"},{"key":"e_1_3_2_119_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.birob.2023.100131"},{"key":"e_1_3_2_120_2","doi-asserted-by":"publisher","unstructured":"Jing Zhao Jingya Wang Madhav Sigdel Bopeng Zhang Phuong Hoang Mengshu Liu and Mohammed Korayem. 2021. Embedding-based recommender system for job to candidate matching on scale. arXiv:2107.00221. Retrieved from https:\/\/doi.org\/10.48550\/arXiv.2107.00221","DOI":"10.48550\/arXiv.2107.00221"},{"key":"e_1_3_2_121_2","doi-asserted-by":"publisher","unstructured":"Mingkai Zheng Xiu Su Shan You Fei Wang Chen Qian Chang Xu and Samuel Albanie. 2023. Can GPT-4 perform neural architecture search? arXiv:2304.10970. Retrieved from https:\/\/doi.org\/10.48550\/arXiv.2304.10970","DOI":"10.48550\/arXiv.2304.10970"}],"container-title":["ACM Transactions on Information Systems"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3702639","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3702639","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,19]],"date-time":"2025-06-19T01:18:09Z","timestamp":1750295889000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3702639"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,1,22]]},"references-count":120,"journal-issue":{"issue":"2","published-print":{"date-parts":[[2025,3,31]]}},"alternative-id":["10.1145\/3702639"],"URL":"https:\/\/doi.org\/10.1145\/3702639","relation":{},"ISSN":["1046-8188","1558-2868"],"issn-type":[{"value":"1046-8188","type":"print"},{"value":"1558-2868","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,1,22]]},"assertion":[{"value":"2023-09-25","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2024-10-17","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-01-22","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}