{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,4,19]],"date-time":"2026-04-19T06:36:43Z","timestamp":1776580603283,"version":"3.51.2"},"reference-count":54,"publisher":"MDPI AG","issue":"12","license":[{"start":{"date-parts":[[2025,12,4]],"date-time":"2025-12-04T00:00:00Z","timestamp":1764806400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Information"],"abstract":"<jats:p>This study aims to assess the capability of Large Language Models (LLMs), particularly GPT-4o, to evaluate and modify the complexity level of Russian school textbooks. We lay the groundwork for developing scalable, context-aware methods for readability assessment and text simplification in Russian educational materials, areas where traditional formulas often fall short. Using a corpus of 154 textbooks covering various subjects and grade levels, we evaluate the extent to which LLMs accurately predict the appropriate comprehension level of a text and how well they simplify texts by targeted grade reduction. Our evaluation framework employs GPT-4o as a multi-role agent in three distinct experiments. First, we prompt the model to estimate the target comprehension age for each segment and identify five key linguistic or conceptual features underpinning its assessment. Second, we simulate student comprehension by instructing the model to reason step-by-step through whether the text is understandable for a hypothetical student of the given grade. Third, we examine the model\u2019s ability to simplify selected fragments by reducing their complexity by three grade levels. We further measure model perplexity and output token probabilities to probe the prediction confidence and coherence. Results indicate that while LLMs show considerable potential in complexity assessment (i.e., MAE of 1 grade level), they tend to overestimate text difficulty and face challenges in achieving precise simplification levels. Ease of understanding assessments generally align with human expectations, although texts with abstract, technical, or poetic content (e.g., Physics, History, and Literary Russian) pose challenges. Our study concludes that LLMs can substantially complement traditional readability metrics and assist teachers in developing suitable Russian educational materials.<\/jats:p>","DOI":"10.3390\/info16121071","type":"journal-article","created":{"date-parts":[[2025,12,4]],"date-time":"2025-12-04T13:53:30Z","timestamp":1764856410000},"page":"1071","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":1,"title":["Assessing the Readability of Russian Textbooks Using Large Language Models"],"prefix":"10.3390","volume":"16","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-7992-4227","authenticated-orcid":false,"given":"Andrei","family":"Paraschiv","sequence":"first","affiliation":[{"name":"Computer Science and Engineering Department, National University of Science and Technology Politehnica of Bucharest, 313 Splaiul Independentei, 060042 Bucharest, Romania"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-4815-9227","authenticated-orcid":false,"given":"Mihai","family":"Dascalu","sequence":"additional","affiliation":[{"name":"Computer Science and Engineering Department, National University of Science and Technology Politehnica of Bucharest, 313 Splaiul Independentei, 060042 Bucharest, Romania"},{"name":"Academy of Romanian Scientists, Str. Ilfov, Nr. 3, 050044 Bucharest, Romania"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-1885-3039","authenticated-orcid":false,"given":"Marina","family":"Solnyshkina","sequence":"additional","affiliation":[{"name":"Laboratory \u201cMultiDisciplinary Investigations of Text\u201d, Institute of Philology and Intercultural Communication, Kazan (Volga Region) Federal University, 18 Kremlevskaya St., 420008 Kazan, Russia"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2025,12,4]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","first-page":"63","DOI":"10.1007\/s10648-011-9181-8","article-title":"Reconstructing readability: Recent developments and recommendations in the analysis of text difficulty","volume":"24","author":"Benjamin","year":"2012","journal-title":"Educ. Psychol. Rev."},{"key":"ref_2","unstructured":"Solovyev, V.D., Solnyshkina, M.I., Ivanov, V., and Timoshenko, S. (2018). Complexity of Russian academic texts as the function of syntactic parameters. Computational Linguistics and Intelligent Text Processing, Proceedings of the 19th International Conference on Computational Linguistics and Intelligent Text Processing, Hanoi, Vietnam, 18\u201324 March 2018, Springer."},{"key":"ref_3","doi-asserted-by":"crossref","first-page":"100","DOI":"10.1007\/s10958-024-07436-y","article-title":"Readability Formulas for Three Levels of Russian School Textbooks","volume":"285","author":"Solovyev","year":"2024","journal-title":"J. Math. Sci."},{"key":"ref_4","doi-asserted-by":"crossref","unstructured":"McNamara, D.S., Graesser, A.C., McCarthy, P.M., and Cai, Z. (2014). Automated Evaluation of Text and Discourse with Coh-Metrix, Cambridge University Press.","DOI":"10.1017\/CBO9780511894664"},{"key":"ref_5","doi-asserted-by":"crossref","first-page":"221","DOI":"10.1037\/h0057532","article-title":"A new readability yardstick","volume":"32","author":"Flesch","year":"1948","journal-title":"J. Appl. Psychol."},{"key":"ref_6","first-page":"132","article-title":"Readability Formula for Russian Texts: A Modified Version","volume":"11289","author":"Batyrshin","year":"2018","journal-title":"Advances in Computational Intelligence, Proceedings of the 17th Mexican International Conference on Artificial Intelligence, MICAI 2018, Guadalajara, Mexico, 22\u201327 October 2018"},{"key":"ref_7","unstructured":"Achiam, J., Adler, S., Agarwal, S., Ahmad, L., Akkaya, I., Aleman, F.L., Almeida, D., Altenschmidt, J., Altman, S., and Anadkat, S. (2023). Gpt-4 technical report. arXiv."},{"key":"ref_8","unstructured":"Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Mahdaouy, Y.E., Lample, G., Babaei, Y., Bashlykov, N., and Batra, S. (2023). LLaMA 2: Open Foundation and Fine-Tuned Chat Models. arXiv."},{"key":"ref_9","doi-asserted-by":"crossref","unstructured":"Quah, B., Zheng, L., Sng, T.J.H., Yong, C.W., and Islam, I. (2024). Reliability of ChatGPT in automated essay scoring for dental undergraduate examinations. BMC Med. Educ., 24.","DOI":"10.1186\/s12909-024-05881-6"},{"key":"ref_10","unstructured":"Wang, T., Tao, M., Fang, R., Wang, H., Wang, S., Jiang, Y.E., and Zhou, W. (2024). AI PERSONA: Towards Life-long Personalization of LLMs. arXiv."},{"key":"ref_11","doi-asserted-by":"crossref","unstructured":"Khan, I., Chohan, I., and Malik, S.I. (2024, January 12\u201313). Leveraging ChatGPT-4 for Enhanced Education: Personalized Problem-Solving and Consistent Learning. Proceedings of the 2024 2nd International Conference on Computing and Data Analytics (ICCDA), Shinas, Oman.","DOI":"10.1109\/ICCDA64887.2024.10867340"},{"key":"ref_12","doi-asserted-by":"crossref","first-page":"317","DOI":"10.22363\/2687-0088-30171","article-title":"Natural language processing and discourse complexity studies","volume":"26","author":"Solnyshkina","year":"2022","journal-title":"Russ. J. Linguist."},{"key":"ref_13","doi-asserted-by":"crossref","first-page":"862","DOI":"10.1007\/s40593-023-00362-1","article-title":"Text-based question difficulty prediction: A systematic review of automatic approaches","volume":"34","author":"AlKhuzaey","year":"2024","journal-title":"Int. J. Artif. Intell. Educ."},{"key":"ref_14","doi-asserted-by":"crossref","unstructured":"Maddela, M., and Xu, W. (November, January 31). A Word-Complexity Lexicon and A Neural Readability Ranking Model for Lexical Simplification. Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, Brussels, Belgium.","DOI":"10.18653\/v1\/D18-1410"},{"key":"ref_15","doi-asserted-by":"crossref","unstructured":"Schwarm, S.E., and Ostendorf, M. (2005, January 25\u201330). Reading level assessment using support vector machines and statistical language models. Proceedings of the 43rd Annual Meeting of the Association for Computational Linguistics (ACL\u201905), Ann Arbor, MI, USA.","DOI":"10.3115\/1219840.1219905"},{"key":"ref_16","doi-asserted-by":"crossref","unstructured":"Pitler, E., and Nenkova, A. (2008, January 25\u201327). Revisiting readability: A unified framework for predicting text quality. Proceedings of the 2008 Conference on Empirical Methods in Natural Language Processing, Honolulu, HI, USA.","DOI":"10.3115\/1613715.1613742"},{"key":"ref_17","unstructured":"Kincaid, J.P., Fishburne, R.P., Rogers, R.L., and Chissom, B.S. (2025, November 28). Derivation of New Readability Formulas (Automated Readability Index, Fog Count and Flesch Reading Ease Formula) for Navy Enlisted Personnel. Technical Report, Naval Technical Training Command Millington TN Research Branch. Available online: https:\/\/apps.dtic.mil\/sti\/tr\/pdf\/ADA006655.pdf."},{"key":"ref_18","doi-asserted-by":"crossref","unstructured":"Yaneva, V., Ha, L.A., Baldwin, P., and Mee, J. (2020, January 11\u201316). Predicting Item Survival for Multiple Choice Questions in a High-Stakes Medical Exam. Proceedings of the Twelfth Language Resources and Evaluation Conference, Marseille, France.","DOI":"10.18653\/v1\/W19-4402"},{"key":"ref_19","unstructured":"Imperial, J.M. (2021, January 1\u20133). BERT Embeddings for Automatic Readability Assessment. Proceedings of the International Conference on Recent Advances in Natural Language Processing (RANLP 2021), Online."},{"key":"ref_20","doi-asserted-by":"crossref","first-page":"73","DOI":"10.18413\/2313-8912-2023-9-1-0-4","article-title":"Classification of Russian textbooks by grade level and topic using ReaderBench","volume":"9","author":"Paraschiv","year":"2023","journal-title":"Res. Result Theor. Appl. Linguist."},{"key":"ref_21","unstructured":"Al-Onaizan, Y., Bansal, M., and Chen, Y. (2024, January 12\u201316). ReadMe++: Benchmarking Multilingual Language Models for Multi-Domain Readability Assessment. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, EMNLP 2024, Miami, FL, USA."},{"key":"ref_22","first-page":"46","article-title":"Automated Assessment of Text Complexity through the Fusion of AutoML and Psycholinguistic Models","volume":"7","author":"Herianah","year":"2025","journal-title":"Forum Linguist. Stud."},{"key":"ref_23","doi-asserted-by":"crossref","first-page":"426","DOI":"10.22363\/2687-0088-30132","article-title":"Text complexity and linguistic features: Their correlation in English and Russian","volume":"26","author":"Morozov","year":"2022","journal-title":"Russ. J. Linguist."},{"key":"ref_24","unstructured":"Farajidizaji, A., Raina, V., and Gales, M. (2024, January 20\u201325). Is It Possible to Modify Text to a Target Readability Level? An Initial Investigation Using Zero-Shot Large Language Models. Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024), Torino, Italia."},{"key":"ref_25","doi-asserted-by":"crossref","unstructured":"Trott, S., and Rivi\u00e8re, P. (2024, January 15). Measuring and Modifying the Readability of English Texts with GPT-4. Proceedings of the Third Workshop on Text Simplification, Accessibility and Readability (TSAR 2024), Miami, FL, USA.","DOI":"10.18653\/v1\/2024.tsar-1.13"},{"key":"ref_26","doi-asserted-by":"crossref","unstructured":"Huang, C.Y., Wei, J., and Huang, T.H.K. (2024, January 11\u201316). Generating educational materials with different levels of readability using LLMs. Proceedings of the Third Workshop on Intelligent and Interactive Writing Assistants, Honolulu, HI, USA.","DOI":"10.1145\/3690712.3690718"},{"key":"ref_27","unstructured":"Kochmar, E., Bexte, M., Burstein, J., Horbach, A., Laarmann-Quante, R., Tack, A., Yaneva, V., and Yuan, Z. (2024, January 20). Beyond Flesch-Kincaid: Prompt-based Metrics Improve Difficulty Classification of Educational Texts. Proceedings of the 19th Workshop on Innovative Use of NLP for Building Educational Applications, BEA 2024, Mexico City, Mexico."},{"key":"ref_28","unstructured":"Gobara, S., Kamigaito, H., and Watanabe, T. (2024, January 7\u20139). Do LLMs Implicitly Determine the Suitable Text Difficulty for Users?. Proceedings of the 38th Pacific Asia Conference on Language, Information and Computation, Tokyo, Japan."},{"key":"ref_29","unstructured":"Imperial, J.M., and Tayyar Madabushi, H. (2023, January 6). Flesch or Fumble? Evaluating Readability Standard Alignment of Instruction-Tuned Language Models. Proceedings of the Third Workshop on Natural Language Generation, Evaluation, and Metrics (GEM), Singapore."},{"key":"ref_30","unstructured":"Alemi, A.A., and Ginsparg, P. (2015). Text segmentation based on semantic word embeddings. arXiv."},{"key":"ref_31","doi-asserted-by":"crossref","first-page":"457","DOI":"10.1613\/jair.1523","article-title":"Lexrank: Graph-based lexical centrality as salience in text summarization","volume":"22","author":"Erkan","year":"2004","journal-title":"J. Artif. Intell. Res."},{"key":"ref_32","unstructured":"Page, L., Brin, S., Motwani, R., and Winograd, T. (1999). The PageRank Citation Ranking: Bringing Order to the Web, Stanford Infolab. Technical Report."},{"key":"ref_33","doi-asserted-by":"crossref","first-page":"919","DOI":"10.1016\/j.ipm.2003.10.006","article-title":"Centroid-based summarization of multiple documents","volume":"40","author":"Radev","year":"2004","journal-title":"Inf. Process. Manag."},{"key":"ref_34","unstructured":"Li, M., Conrad, F., and Gagnon-Bartsch, J. (2025, January 23\u201326). FastLexRank: Bring Order into Social Media Posts Using Lexical Ranking Algorithm. Proceedings of the ICWSM Workshops: R2CASS 2025\u2014Social Science Meets Web Data, Copenhagen, Denmark."},{"key":"ref_35","doi-asserted-by":"crossref","unstructured":"Olney, A.M. (2022). Assessing Readability by Filling Cloze Items with Transformers. Artificial Intelligence in Education, Proceedings of the 23rd Proceedings of the International Conference on Artificial Intelligence in Education, Durham, UK, 27\u201331 July 2022, Springer.","DOI":"10.1007\/978-3-031-11644-5_25"},{"key":"ref_36","doi-asserted-by":"crossref","unstructured":"Kamalloo, E., Upadhyay, S., and Lin, J. (2024, January 14\u201318). Towards robust qa evaluation via open llms. Proceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval, Washington DC, USA.","DOI":"10.1145\/3626772.3657675"},{"key":"ref_37","doi-asserted-by":"crossref","unstructured":"Benedetto, L., Aradelli, G., Donvito, A., Lucchetti, A., Cappelli, A., and Buttery, P. (2024, January 12\u201316). Using LLMs to simulate students\u2019 responses to exam questions. Proceedings of the Findings of the Association for Computational Linguistics: EMNLP 2024, Miami, FL, USA.","DOI":"10.18653\/v1\/2024.findings-emnlp.663"},{"key":"ref_38","first-page":"202","article-title":"Exploring the Potential of Large Language Models for Estimating the Reading Comprehension Question Difficulty","volume":"Volume 15812","author":"Sottilare","year":"2025","journal-title":"Proceedings of the Adaptive Instructional Systems\u20147th International Conference, AIS 2025, Held as Part of the 27th HCI International Conference, HCII 2025"},{"key":"ref_39","doi-asserted-by":"crossref","first-page":"47","DOI":"10.1007\/s44163-024-00147-y","article-title":"Performance of the pre-trained large language model GPT-4 on automated short answer grading","volume":"4","author":"Kortemeyer","year":"2024","journal-title":"Discov. Artif. Intell."},{"key":"ref_40","unstructured":"Al-Onaizan, Y., Bansal, M., and Chen, Y. (2024, January 12\u201316). Large Language Model as an Assignment Evaluator: Insights, Feedback, and Challenges in a 1000+ Student Course. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, EMNLP 2024, Miami, FL, USA."},{"key":"ref_41","unstructured":"Hu, Y., Huang, Q., Tao, M., Zhang, C., and Feng, Y. (2024, January 11). Can Perplexity Reflect Large Language Model\u2019s Ability in Long Text Understanding?. Proceedings of the Second Tiny Papers Track at ICLR 2024, Tiny Papers @ ICLR 2024, Vienna, Austria."},{"key":"ref_42","first-page":"24824","article-title":"Chain-of-thought prompting elicits reasoning in large language models","volume":"35","author":"Wei","year":"2022","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_43","doi-asserted-by":"crossref","first-page":"501","DOI":"10.22363\/2618-8163-2024-22-4-501-517","article-title":"Approaches and tools for Russian text linguistic profiling","volume":"22","author":"Solnyshkina","year":"2024","journal-title":"Russ. Lang. Stud."},{"key":"ref_44","doi-asserted-by":"crossref","first-page":"321","DOI":"10.1162\/tacl_a_00648","article-title":"Large language models enable few-shot clustering","volume":"12","author":"Viswanathan","year":"2024","journal-title":"Trans. Assoc. Comput. Linguist."},{"key":"ref_45","unstructured":"Tarekegn, A.N. (2024). Large Language Model Enhanced Clustering for News Event Detection. arXiv."},{"key":"ref_46","unstructured":"Grootendorst, M. (2022). BERTopic: Neural topic modeling with a class-based TF-IDF procedure. arXiv."},{"key":"ref_47","doi-asserted-by":"crossref","first-page":"331","DOI":"10.22363\/2618-8163-2021-19-3-331-345","article-title":"Textometr: An online tool for automated complexity level assessment of texts for Russian language learners","volume":"19","author":"Laposhina","year":"2021","journal-title":"Russ. Lang. Stud."},{"key":"ref_48","doi-asserted-by":"crossref","first-page":"540","DOI":"10.22363\/2618-8163-2024-22-4-540-554","article-title":"Russian language textbook as agent of change: From USSR to the new century","volume":"22","author":"Bulina","year":"2024","journal-title":"Russ. Lang. Stud."},{"key":"ref_49","unstructured":"Haim, A., Salinas, A., and Nyarko, J. (2024). What\u2019s in a Name? Auditing Large Language Models for Race and Gender Bias. arXiv."},{"key":"ref_50","unstructured":"Delikoura, I., Fung, Y.R., and Hui, P. (2025). From Superficial Outputs to Superficial Learning: Risks of Large Language Models in Education. arXiv."},{"key":"ref_51","doi-asserted-by":"crossref","unstructured":"Dokic, K., Pisker, B., and Radisic, B. (2025). Mirroring Cultural Dominance: Disclosing Large Language Models Social Values, Attitudes and Stereotypes. Societies, 15.","DOI":"10.3390\/soc15050142"},{"key":"ref_52","doi-asserted-by":"crossref","unstructured":"Dudy, S., Ahmad, I.S., Kitajima, R., and Lapedriza, \u00c0. (2024, January 15\u201318). Analyzing Cultural Representations of Emotions in LLMs Through Mixed Emotion Survey. Proceedings of the 12th International Conference on Affective Computing and Intelligent Interaction, ACII 2024, Glasgow, UK.","DOI":"10.1109\/ACII63134.2024.00044"},{"key":"ref_53","doi-asserted-by":"crossref","unstructured":"Se\u00dfler, K., F\u00fcrstenberg, M., B\u00fchler, B., and Kasneci, E. (2025, January 3\u20137). Can AI grade your essays? A comparative analysis of large language models and teacher ratings in multidimensional essay scoring. Proceedings of the 15th International Learning Analytics and Knowledge Conference, LAK 2025, Dublin, Ireland.","DOI":"10.1145\/3706468.3706527"},{"key":"ref_54","doi-asserted-by":"crossref","unstructured":"Gerlich, M. (2025). AI Tools in Society: Impacts on Cognitive Offloading and the Future of Critical Thinking. Societies, 15.","DOI":"10.2139\/ssrn.5082524"}],"container-title":["Information"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/2078-2489\/16\/12\/1071\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,12,6]],"date-time":"2025-12-06T05:14:46Z","timestamp":1764998086000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/2078-2489\/16\/12\/1071"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,12,4]]},"references-count":54,"journal-issue":{"issue":"12","published-online":{"date-parts":[[2025,12]]}},"alternative-id":["info16121071"],"URL":"https:\/\/doi.org\/10.3390\/info16121071","relation":{},"ISSN":["2078-2489"],"issn-type":[{"value":"2078-2489","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,12,4]]}}}