{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,8,18]],"date-time":"2026-08-18T03:17:16Z","timestamp":1787023036997,"version":"build-2736575974"},"reference-count":40,"publisher":"Institute of Electronics, Information and Communications Engineers (IEICE)","issue":"9","content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["IEICE Trans. Inf. &amp; Syst."],"published-print":{"date-parts":[[2025,9,1]]},"DOI":"10.1587\/transinf.2024edp7189","type":"journal-article","created":{"date-parts":[[2025,3,6]],"date-time":"2025-03-06T17:13:03Z","timestamp":1741281183000},"page":"1095-1107","source":"Crossref","is-referenced-by-count":2,"title":["Japanese Essay Scoring with Generative Pre-Trained Language Models Combined with Soft Labels"],"prefix":"10.1587","volume":"E108.D","author":[{"given":"Boago","family":"OKGETHENG","sequence":"first","affiliation":[{"name":"Graduate School of Environmental, Life, Natural Science and Technology, Okayama University"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Koichi","family":"TAKEUCHI","sequence":"additional","affiliation":[{"name":"Graduate School of Environmental, Life, Natural Science and Technology, Okayama University"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"532","reference":[{"key":"1","unstructured":"[1] Y. Attali and J. Burstein, \u201cAutomated essay scoring with e-rater\u00ae v.2,\u201d The Journal of Technology, Learning and Assessment, vol.4, no.3, Feb. 2006."},{"key":"2","doi-asserted-by":"crossref","unstructured":"[2] I. Persing and V. Ng, \u201cModeling prompt adherence in student essays,\u201d Proceedings of the 52nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp.1534-1543, Baltimore, Maryland, Association for Computational Linguistics, June 2014. 10.3115\/v1\/p14-1144","DOI":"10.3115\/v1\/P14-1144"},{"key":"3","doi-asserted-by":"crossref","unstructured":"[3] P. Phandi, K.M.A. Chai, and H.T. Ng, \u201cFlexible domain adaptation for automated essay scoring using correlated linear regression,\u201d Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing, pp.431-439, Lisbon, Portugal, Association for Computational Linguistics, Sept. 2015. 10.18653\/v1\/d15-1049","DOI":"10.18653\/v1\/D15-1049"},{"key":"4","doi-asserted-by":"crossref","unstructured":"[4] L.S. Larkey, \u201cAutomatic essay grading using text categorization techniques,\u201d Proceedings of the 21st Annual International ACM SIGIR Conference on Research and Development in Information Retrieval, SIGIR \u201998, pp.90-95, New York, NY, USA, Association for Computing Machinery, 1998. 10.1145\/290941.290965","DOI":"10.1145\/290941.290965"},{"key":"5","unstructured":"[5] L.M. Rudner and T. Liang, \u201cAutomated essay scoring using bayes\u2019 theorem,\u201d The Journal of Technology, Learning and Assessment, vol.1, no.2, June 2002."},{"key":"6","unstructured":"[6] H. Yannakoudakis, T. Briscoe, and B. Medlock, \u201cA new dataset and method for automatically grading ESOL texts,\u201d Proceedings of the 49th Annual Meeting of the Association for Computational Linguistics: Human Language Technologies, pp.180-189, Portland, Oregon, USA, Association for Computational Linguistics, June 2011."},{"key":"7","doi-asserted-by":"crossref","unstructured":"[7] F. Dong, Y. Zhang, and J. Yang, \u201cAttention-based recurrent convolutional neural network for automatic essay scoring,\u201d R. Levy and L. Specia, editors, Proceedings of the 21st Conference on Computational Natural Language Learning (CoNLL 2017), pp.153-162, Vancouver, Canada, Association for Computational Linguistics, Aug. 2017","DOI":"10.18653\/v1\/K17-1017"},{"key":"8","doi-asserted-by":"publisher","unstructured":"[8] Y. Tay, M. Phan, L.A. Tuan, and S.C. Hui, \u201cSkipflow: Incorporating neural coherence features for end-to-end automatic text scoring,\u201d Proceedings of the AAAI Conference on Artificial Intelligence, vol.32, no.1, 2018. 10.1609\/aaai.v32i1.12045","DOI":"10.1609\/aaai.v32i1.12045"},{"key":"9","unstructured":"[9] J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, \u201cBERT: Pre-training of deep bidirectional transformers for language understanding,\u201d Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pp.4171-4186, Minneapolis, Minnesota, Association for Computational Linguistics, June 2019."},{"key":"10","unstructured":"[10] A. Radford and K. Narasimhan, \u201cImproving language understanding by generative pre-training,\u201d 2018."},{"key":"11","doi-asserted-by":"crossref","unstructured":"[11] M. Uto, Y. Xie, and M. Ueno, \u201cNeural automated essay scoring incorporating handcrafted features,\u201d Proceedings of the 28th International Conference on Computational Linguistics, pp.6077-6088, Barcelona, Spain (Online), International Committee on Computational Linguistics, Dec. 2020. 10.18653\/v1\/2020.coling-main.535","DOI":"10.18653\/v1\/2020.coling-main.535"},{"key":"12","unstructured":"[12] P.U. Rodriguez, A. Jafari, and C.M. Ormerod, \u201cLanguage models and automated essay scoring,\u201d 2019."},{"key":"13","doi-asserted-by":"crossref","unstructured":"[13] E. Mayfield and A.W. Black, \u201cShould you fine-tune BERT for automated essay scoring?,\u201d J. Burstein, E. Kochmar, C. Leacock, N. Madnani, I. Pil\u00e1n, H. Yannakoudakis, and T. Zesch, editors, Proceedings of the Fifteenth Workshop on Innovative Use of NLP for Building Educational Applications, pp.151-162, Seattle, WA, USA \u2192 Online, Association for Computational Linguistics, July 2020. 10.18653\/v1\/2020.bea-1.15","DOI":"10.18653\/v1\/2020.bea-1.15"},{"key":"14","doi-asserted-by":"crossref","unstructured":"[14] R. Yang, J. Cao, Z. Wen, Y. Wu, and X. He, \u201cEnhancing automated essay scoring performance via fine-tuning pre-trained language models with combination of regression and ranking,\u201d Findings of the Association for Computational Linguistics: EMNLP 2020, pp.1560-1569, Online, Association for Computational Linguistics, Nov. 2020. 10.18653\/v1\/2020.findings-emnlp.141","DOI":"10.18653\/v1\/2020.findings-emnlp.141"},{"key":"15","unstructured":"[15] T.B. Brown, B. Mann, N. Ryder, M. Subbiah, J. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, S. Agarwal, A. Herbert-Voss, G. Krueger, T. Henighan, R. Child, A. Ramesh, D.M. Ziegler, J. Wu, C. Winter, C. Hesse, M. Chen, E. Sigler, M. Litwin, S. Gray, B. Chess, J. Clark, C. Berner, S. McCandlish, A. Radford, I. Sutskever, and D. Amodei, \u201cLanguage models are few-shot learners,\u201d 2020."},{"key":"16","unstructured":"[16] OpenAI, J. Achiam, S. Adler, S. Agarwal, L. Ahmad, I. Akkaya, F.L. Aleman, D. Almeida, J. Altenschmidt, S. Altman, S. Anadkat, R. Avila, I. Babuschkin, S. Balaji, V. Balcom, P. Baltescu, H. Bao, M. Bavarian, J. Belgum, I. Bello, J. Berdine, G. Bernadett-Shapiro, C. Berner, L. Bogdonoff, O. Boiko, M. Boyd, A.-L. Brakman, G. Brockman, T. Brooks, M. Brundage, K. Button, T. Cai, R. Campbell, A. Cann, B. Carey, C. Carlson, R. Carmichael, B. Chan, C. Chang, F. Chantzis, D. Chen, S. Chen, R. Chen, J. Chen, M. Chen, B. Chess, C. Cho, C. Chu, H.W. Chung, D. Cummings, J. Currier, Y. Dai, C. Decareaux, T. Degry, N. Deutsch, D. Deville, A. Dhar, D. Dohan, S. Dowling, S. Dunning, A. Ecoffet, A. Eleti, T. Eloundou, D. Farhi, L. Fedus, N. Felix, S.P. Fishman, J. Forte, I. Fulford, L. Gao, E. Georges, C. Gibson, V. Goel, T. Gogineni, G. Goh, R. Gontijo-Lopes, J. Gordon, M. Grafstein, S. Gray, R. Greene, J. Gross, S.S. Gu, Y. Guo, C. Hallacy, J. Han, J. Harris, Y. He, M. Heaton, J. Heidecke, C. Hesse, A. Hickey, W. Hickey, P. Hoeschele, B. Houghton, K. Hsu, S. Hu, X. Hu, J. Huizinga, S. Jain, S. Jain, J. Jang, A. Jiang, R. Jiang, H. Jin, D. Jin, S. Jomoto, B. Jonn, H. Jun, T. Kaftan, \u0141. Kaiser, A. Kamali, I. Kanitscheider, NS. Keskar, T. Khan, L. Kilpatrick, J.W. Kim, C. Kim, Y. Kim, J.H. Kirchner, J. Kiros, M. Knight, D. Kokotajlo, \u0141. Kondraciuk, A. Kondrich, A. Konstantinidis, K. Kosic, G. Krueger, V. Kuo, M. Lampe, I. Lan, T. Lee, J. Leike, J. Leung, D. Levy, C.M. Li, R. Lim, M. Lin, S. Lin, M. Litwin, T. Lopez, R. Lowe, P. Lue, A. Makanju, K. Malfacini, S. Manning, T. Markov, Y. Markovski, B. Martin, K. Mayer, A. Mayne, B. McGrew, S.M. McKinney, C. McLeavey, P. McMillan, J. McNeil, D. Medina, A. Mehta, J. Menick, L. Metz, A. Mishchenko, P. Mishkin, V. Monaco, E. Morikawa, D. Mossing, T. Mu, M. Murati, O. Murk, D. M\u00e9ly, A. Nair, R. Nakano, R. Nayak, A. Neelakantan, R. Ngo, H. Noh, L. Ouyang, C. O\u2019Keefe, J. Pachocki, A. Paino, J. Palermo, A. Pantuliano, G. Parascandolo, J. Parish, E. Parparita, A. Passos, M. Pavlov, A. Peng, A. Perelman, F. de Avila Belbute Peres, M. Petrov, H.P. de Oliveira Pinto, Michael, Pokorny, M. Pokrass, V.H. Pong, T. Powell, A. Power, B. Power, E. Proehl, R. Puri, A. Radford, J. Rae, A. Ramesh, C. Raymond, F. Real, K. Rimbach, C. Ross, B. Rotsted, H. Roussez, N. Ryder, M. Saltarelli, T. Sanders, S. Santurkar, G. Sastry, H. Schmidt, D. Schnurr, J. Schulman, D. Selsam, K. Sheppard, T. Sherbakov, J. Shieh, S. Shoker, P. Shyam, S. Sidor, E. Sigler, M. Simens, J. Sitkin, K. Slama, I. Sohl, B. Sokolowsky, Y. Song, N. Staudacher, F.P. Such, N. Summers, I. Sutskever, J. Tang, N. Tezak, M.B. Thompson, P. Tillet, A. Tootoonchian, E. Tseng, P. Tuggle, N. Turley, J. Tworek, J.F.C. Uribe, A. Vallone, A. Vijayvergiya, C. Voss, C. Wainwright, J.J. Wang, A. Wang, B. Wang, J. Ward, J. Wei, C.J. Weinmann, A. Welihinda, P. Welinder, J. Weng, L. Weng, M. Wiethoff, D. Willner, C. Winter, S. Wolrich, H. Wong, L. Workman, S. Wu, J. Wu, M. Wu, K. Xiao, T. Xu, S. Yoo, K. Yu, Q. Yuan, W. Zaremba, R. Zellers, C. Zhang, M. Zhang, S. Zhao, T. Zheng, J. Zhuang, W. Zhuk, and B. Zoph, Gpt-4 technical report, 2024."},{"key":"17","doi-asserted-by":"publisher","unstructured":"[17] A. Mizumoto and M. Eguchi, \u201cExploring the potential of using an ai language model for automated essay scoring,\u201d Research Methods in Applied Linguistics, vol.2, no.2, p.100050, 2023. 10.1016\/j.rmal.2023.100050","DOI":"10.1016\/j.rmal.2023.100050"},{"key":"18","unstructured":"[18] C. Xiao, W. Ma, S.X. Xu, K. Zhang, Y. Wang, and Q. Fu, \u201cFrom automation to augmentation: Large language models elevating essay scoring landscape,\u201d 2024."},{"key":"19","unstructured":"[19] J. Morimoto and K. Takeuchi, \u201cJapanese Automatic Essay Grading Using R2BERT and ChatGPT-3.5,\u201d IPSJ SIG Technical Report, 2024-IFAT-155, pp.1-5, 2024, (in Japanese)."},{"key":"20","doi-asserted-by":"crossref","unstructured":"[20] K. Taghipour and H.T. Ng, \u201cA neural approach to automated essay scoring,\u201d Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing, pp.1882-1891, Austin, Texas, Association for Computational Linguistics, Nov. 2016. 10.18653\/v1\/d16-1193","DOI":"10.18653\/v1\/D16-1193"},{"key":"21","doi-asserted-by":"crossref","unstructured":"[21] W. Song, K. Zhang, R. Fu, L. Liu, T. Liu, and M. Cheng, \u201cMulti-stage pre-training for automated Chinese essay scoring,\u201d Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pp.6723-6733, Online, Association for Computational Linguistics, Nov. 2020. 10.18653\/v1\/2020.emnlp-main.546","DOI":"10.18653\/v1\/2020.emnlp-main.546"},{"key":"22","unstructured":"[22] R. Hirao, M. Arai, H. Shimanaka, S. Katsumata, and M. Komachi, \u201cAutomated essay scoring system for nonnative Japanese learners,\u201d Proceedings of the Twelfth Language Resources and Evaluation Conference, pp.1250-1257, Marseille, France, European Language Resources Association, May 2020."},{"key":"23","doi-asserted-by":"publisher","unstructured":"[23] R. Ridley, L. He, X.-Y. Dai, S. Huang, and J. Chen, \u201cAutomated cross-prompt scoring of essay traits,\u201d Proc. Conf. AAAI Artif. Intell., vol.35, no.15, pp.13745-13753, May 2021. 10.1609\/aaai.v35i15.17620","DOI":"10.1609\/aaai.v35i15.17620"},{"key":"24","doi-asserted-by":"crossref","unstructured":"[24] R. Diaz and A. Marathe, \u201cSoft labels for ordinal regression,\u201d Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp.4733-4742, June 2019. 10.1109\/cvpr.2019.00487","DOI":"10.1109\/CVPR.2019.00487"},{"key":"25","doi-asserted-by":"crossref","unstructured":"[25] H. Chen and B. He, \u201cAutomated essay scoring by maximizing human-machine agreement,\u201d Proceedings of the 2013 Conference on Empirical Methods in Natural Language Processing, pp.1741-1752, Seattle, Washington, USA, Association for Computational Linguistics, pp.1741-1752, Oct. 2013. 10.18653\/v1\/d13-1180","DOI":"10.18653\/v1\/D13-1180"},{"key":"26","doi-asserted-by":"crossref","unstructured":"[26] Y. Wang, C. Wang, R. Li, and H. Lin, \u201cOn the use of bert for automated essay scoring: Joint learning of multi-scale essay representation,\u201d Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pp.3416-3425, Seattle, United States, Association for Computational Linguistics, July 2022. 10.18653\/v1\/2022.naacl-main.249","DOI":"10.18653\/v1\/2022.naacl-main.249"},{"key":"27","doi-asserted-by":"crossref","unstructured":"[27] A. Obata, T. Tagawa, and Y. Ono, \u201cAssessment of ChatGPT\u2019s validity in scoring essays by foreign language learners of japanese and english,\u201d 2023 15th International Congress on Advanced Applied Informatics Winter (IIAI-AAI-Winter), IEEE, pp.105-110, Dec. 2023. 10.1109\/iiai-aai-winter61682.2023.00028","DOI":"10.1109\/IIAI-AAI-Winter61682.2023.00028"},{"key":"28","unstructured":"[28] E.J. Hu, Y. Shen, P. Wallis, Z. Allen-Zhu, Y. Li, S. Wang, L. Wang, and W. Chen, \u201cLora: Low-rank adaptation of large language models,\u201d 2021."},{"key":"29","unstructured":"[29] Z. Li, X. Li, Y. Liu, H. Xie, J. Li, Fu lee Wang, Q. Li, and X. Zhong, \u201cLabel supervised llama finetuning,\u201d 2023."},{"key":"30","doi-asserted-by":"crossref","unstructured":"[30] X. Ma, L. Wang, N. Yang, F. Wei, and J. Lin, \u201cFine-tuning llama for multi-stage text retrieval,\u201d Proceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval, pp.2421-2425, 2024. 10.1145\/3626772.3657951","DOI":"10.1145\/3626772.3657951"},{"key":"31","unstructured":"[31] L. Ouyang, J. Wu, X. Jiang, D. Almeida, C. Wainwright, P. Mishkin, C. Zhang, S. Agarwal, K. Slama, A. Ray, J. Schulman, J. Hilton, F. Kelton, L. Miller, M. Simens, A. Askell, P. Welinder, P.F. Christiano, J. Leike, and R. Lowe, \u201cTraining language models to follow instructions with human feedback,\u201d S. Koyejo, S. Mohamed, A. Agarwal, D. Belgrave, K. Cho, and A. Oh, editors, Advances in Neural Information Processing Systems, vol.35, pp.27730-27744, Curran Associates, Inc., 2022."},{"key":"32","unstructured":"[32] C. Guo, G. Pleiss, Y. Sun, and K.Q. Weinberger, \u201cOn calibration of modern neural networks,\u201d D. Precup and Y.W. Teh, editors, Proceedings of the 34th International Conference on Machine Learning, vol.70 of Proceedings of Machine Learning Research, pp.1321-1330. PMLR, Aug. 2017."},{"key":"33","unstructured":"[33] R. Hirao, M. Arai, H. Shimanaka, S. Katsumata, and M. Komachi, \u201cAutomated essay scoring system for nonnative japanese learners,\u201d Proceedings of the 12th Conference on Language Resources and Evaluation, pp.1250-1257, 2020."},{"key":"34","doi-asserted-by":"crossref","unstructured":"[34] M. Cozma, A. Butnaru, and R.T. Ionescu, \u201cAutomated essay scoring with string kernels and word embeddings,\u201d Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), pp.503-509, Melbourne, Australia, Association for Computational Linguistics, July 2018. 10.18653\/v1\/p18-2080","DOI":"10.18653\/v1\/P18-2080"},{"key":"35","doi-asserted-by":"crossref","unstructured":"[35] S. Mathias and P. Bhattacharyya, \u201cCan neural networks automatically score essay traits?,\u201d Proceedings of the Fifteenth Workshop on Innovative Use of NLP for Building Educational Applications, pp.85-91, Seattle, WA, USA \u2192 Online, Association for Computational Linguistics, pp.85-91, July 2020. 10.18653\/v1\/2020.bea-1.8","DOI":"10.18653\/v1\/2020.bea-1.8"},{"key":"36","unstructured":"[36] J. Burstein, M. Chodorow, and C. Leacock, \u201cAutomated essay evaluation: The criterion online writing service,\u201d AI Magazine, vol.25, no.3, p.27, Sept. 2004."},{"key":"37","unstructured":"[37] Z. Yang, Z. Dai, Y. Yang, J. Carbonell, R.R. Salakhutdinov, and Q.V. Le, \u201cXlnet: Generalized autoregressive pretraining for language understanding,\u201d H. Wallach, H. Larochelle, A. Beygelzimer, F. d\u2019Alch\u00e9 Buc, E. Fox, and R. Garnett, editors, Advances in Neural Information Processing Systems, vol.32. Curran Associates, Inc., 2019."},{"key":"38","unstructured":"[38] C. Sun, L. Huang, and X. Qiu, \u201cUtilizing BERT for aspect-based sentiment analysis via constructing auxiliary sentence,\u201d Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, vol.1 (Long and Short Papers), pp.380-385, Minneapolis, Minnesota, Association for Computational Linguistics, June 2019."},{"key":"39","doi-asserted-by":"crossref","unstructured":"[39] A. Cohan, I. Beltagy, D. King, B. Dalvi, and D. Weld, \u201cPretrained language models for sequential sentence classification,\u201d Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pp.3693-3699, Hong Kong, China, Association for Computational Linguistics, pp.3693-3699, Nov. 2019. 10.18653\/v1\/d19-1383","DOI":"10.18653\/v1\/D19-1383"},{"key":"40","doi-asserted-by":"publisher","unstructured":"[40] Y. Sun, S. Wang, Y. Li, S. Feng, H. Tian, H. Wu, and H. Wang, \u201cErnie 2.0: A continual pre-training framework for language understanding,\u201d Proceedings of the AAAI Conference on Artificial Intelligence, vol.34, no.5, pp.8968-8975, 2020. 10.1609\/aaai.v34i05.6428","DOI":"10.1609\/aaai.v34i05.6428"}],"container-title":["IEICE Transactions on Information and Systems"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.jstage.jst.go.jp\/article\/transinf\/E108.D\/9\/E108.D_2024EDP7189\/_pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,9,6]],"date-time":"2025-09-06T03:25:09Z","timestamp":1757129109000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.jstage.jst.go.jp\/article\/transinf\/E108.D\/9\/E108.D_2024EDP7189\/_article"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,9,1]]},"references-count":40,"journal-issue":{"issue":"9","published-print":{"date-parts":[[2025]]}},"URL":"https:\/\/doi.org\/10.1587\/transinf.2024edp7189","relation":{},"ISSN":["0916-8532","1745-1361"],"issn-type":[{"value":"0916-8532","type":"print"},{"value":"1745-1361","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,9,1]]},"article-number":"2024EDP7189"}}