{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,10]],"date-time":"2026-07-10T00:00:35Z","timestamp":1783641635850,"version":"3.55.0"},"reference-count":37,"publisher":"Springer Science and Business Media LLC","issue":"5","license":[{"start":{"date-parts":[[2024,2,9]],"date-time":"2024-02-09T00:00:00Z","timestamp":1707436800000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2024,2,9]],"date-time":"2024-02-09T00:00:00Z","timestamp":1707436800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"name":"Tomorrow Advancing Life","award":["gift"],"award-info":[{"award-number":["gift"]}]},{"name":"NSF CISE RI","award":["2112926"],"award-info":[{"award-number":["2112926"]}]},{"DOI":"10.13039\/100020670","name":"Stanford Institute for Human-Centered Artificial Intelligence, Stanford University","doi-asserted-by":"publisher","award":["Hoffman Yee grant (to 7 PIs including Landay and Brunskill)"],"award-info":[{"award-number":["Hoffman Yee grant (to 7 PIs including Landay and Brunskill)"]}],"id":[{"id":"10.13039\/100020670","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Mach Learn"],"published-print":{"date-parts":[[2024,5]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:p>Resource limitations make it challenging to provide all students with one of the most effective educational interventions: personalized instruction. Reinforcement learning could be a pivotal tool to decrease the development costs and enhance the effectiveness of intelligent tutoring software, that aims to provide the right support, at the right time, to a student. Here we illustrate that deep reinforcement learning can be used to provide adaptive pedagogical support to students learning about the concept of volume in a narrative storyline software. Using explainable artificial intelligence tools, we extracted interpretable insights about the pedagogical policy learned and demonstrated that the resulting policy had similar performance in a different student population. Most importantly, in both studies, the reinforcement-learning narrative system had the largest benefit for those students with the lowest initial pretest scores, suggesting the opportunity for AI to adapt and provide support for those most in need.<\/jats:p>","DOI":"10.1007\/s10994-023-06423-9","type":"journal-article","created":{"date-parts":[[2024,2,9]],"date-time":"2024-02-09T18:03:01Z","timestamp":1707501781000},"page":"3023-3048","update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":26,"title":["Reinforcement learning tutor better supported lower performers in a math task"],"prefix":"10.1007","volume":"113","author":[{"given":"Sherry","family":"Ruan","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Allen","family":"Nie","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"William","family":"Steenbergen","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Jiayu","family":"He","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"J. Q.","family":"Zhang","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Meng","family":"Guo","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Yao","family":"Liu","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Kyle","family":"Dang Nguyen","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Catherine Y.","family":"Wang","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Rui","family":"Ying","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"James A.","family":"Landay","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-3971-7127","authenticated-orcid":false,"given":"Emma","family":"Brunskill","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"297","published-online":{"date-parts":[[2024,2,9]]},"reference":[{"key":"6423_CR1","doi-asserted-by":"crossref","unstructured":"Bassen, J., Balaji, B., Schaarschmidt, M., Thille, C., Painter, J., Zimmaro, D., Games, A., Fast, E., & Mitchell, J. C. (2020). Reinforcement learning for the adaptive scheduling of educational activities. In CHI, pp. 1\u201312","DOI":"10.1145\/3313831.3376518"},{"issue":"1","key":"6423_CR2","first-page":"1","volume":"9","author":"CR Beal","year":"2010","unstructured":"Beal, C. R., Arroyo, I. M., Cohen, P. R., & Woolf, B. P. (2010). Evaluation of animalwatch: An intelligent tutoring system for arithmetic and fractions. Journal of Interactive Online Learning, 9(1), 1\u201314.","journal-title":"Journal of Interactive Online Learning"},{"key":"6423_CR3","unstructured":"Bird, S., Klein, E., & Loper, E. (2009). Natural language processing with Python: Analyzing text with the natural language toolkit. O\u2019Reilly Media, Inc."},{"key":"6423_CR4","doi-asserted-by":"publisher","first-page":"11","DOI":"10.3389\/fpsyg.2017.00011","volume":"8","author":"E Carey","year":"2017","unstructured":"Carey, E., Hill, F., Devine, A., & Szucs, D. (2017). The modified abbreviated math anxiety scale: A valid and reliable instrument for use with children. Frontiers in Psychology, 8, 11. https:\/\/doi.org\/10.3389\/fpsyg.2017.00011","journal-title":"Frontiers in Psychology"},{"key":"6423_CR5","doi-asserted-by":"publisher","first-page":"11","DOI":"10.3389\/fpsyg.2017.00011","volume":"8","author":"E Carey","year":"2017","unstructured":"Carey, E., Hill, F., Devine, A., & Sz\u0171cs, D. (2017). The modified abbreviated math anxiety scale: A valid and reliable instrument for use with children. Frontiers in Psychology, 8, 11.","journal-title":"Frontiers in Psychology"},{"issue":"1","key":"6423_CR6","doi-asserted-by":"publisher","first-page":"137","DOI":"10.1007\/s11257-010-9093-1","volume":"21","author":"M Chi","year":"2011","unstructured":"Chi, M., VanLehn, K., Litman, D., & Jordan, P. (2011). Empirically evaluating the application of reinforcement learning to the induction of effective and adaptive pedagogical strategies. User Modeling and User-Adapted Interaction, 21(1), 137\u2013180.","journal-title":"User Modeling and User-Adapted Interaction"},{"issue":"11","key":"6423_CR7","doi-asserted-by":"publisher","first-page":"1062","DOI":"10.1126\/sciadv.aay1062","volume":"5","author":"KW Choe","year":"2019","unstructured":"Choe, K. W., Jenifer, J. B., Rozek, C. S., Berman, M. G., & Beilock, S. L. (2019). Calculated avoidance: Math anxiety predicts math avoidance in effort-based decision-making. Science Advances, 5(11), 1062.","journal-title":"Science Advances"},{"key":"6423_CR8","doi-asserted-by":"crossref","unstructured":"Corbett, A. (2001) Cognitive computer tutors: Solving the two-sigma problem. In International Conference on User Modeling, pp. 137\u2013147. Springer","DOI":"10.1007\/3-540-44566-8_14"},{"key":"6423_CR9","doi-asserted-by":"crossref","unstructured":"de Barros, A., & Ganimian, A.J. (2021). Which students benefit from personalized learning? Experimental evidence from a math software in public schools in India","DOI":"10.1080\/19345747.2021.2005203"},{"key":"6423_CR10","unstructured":"Dietz, G., Pease, Z., McNally, B., & Foss, E. (2020). Giggle gauge: a self-report instrument for evaluating children\u2019s engagement with technology. InProceedings of the Interaction Design and Children Conference, pp. 614\u2013623"},{"issue":"4","key":"6423_CR11","doi-asserted-by":"publisher","first-page":"568","DOI":"10.1007\/s40593-019-00187-x","volume":"29","author":"S Doroudi","year":"2019","unstructured":"Doroudi, S., Aleven, V., & Brunskill, E. (2019). Where\u2019s the reward? International Journal of Artificial Intelligence in Education, 29(4), 568\u2013620.","journal-title":"International Journal of Artificial Intelligence in Education"},{"key":"6423_CR12","unstructured":"Facebook: Facebook React. https:\/\/github.com\/facebook\/react. Accessed: 2019-08-20 (2019)"},{"key":"6423_CR13","unstructured":"Hasura: Hasura GraphQL. https:\/\/github.com\/hasura\/graphql-engine. Accessed: 2019-08-20 (2019)"},{"key":"6423_CR14","unstructured":"Hendrycks, D., & Gimpel, K. (2016). Gaussian error linear units (gelus). arXiv preprint arXiv:1606.08415"},{"issue":"1","key":"6423_CR15","first-page":"1334","volume":"17","author":"S Levine","year":"2016","unstructured":"Levine, S., Finn, C., Darrell, T., & Abbeel, P. (2016). End-to-end training of deep visuomotor policies. The Journal of Machine Learning Research, 17(1), 1334\u20131373.","journal-title":"The Journal of Machine Learning Research"},{"key":"6423_CR16","unstructured":"Liu, Y., Swaminathan, A., Agarwal, A., & Brunskill, E. (2020). Off-policy policy gradient with stationary distribution correction. In Uncertainty in Artificial Intelligence, pp. 1180\u20131190. PMLR"},{"key":"6423_CR17","unstructured":"Lundberg, S.M., & Lee, S.-I. (2017). A unified approach to interpreting model predictions. Advances in Neural Information Processing Systems 30"},{"key":"6423_CR18","unstructured":"Mandel, T., Liu, Y.-E., Levine, S., Brunskill, E., & Popovic, Z. (2014). Offline policy evaluation across representations with applications to educational games. In AAMAS, vol. 1077"},{"key":"6423_CR19","unstructured":"Metelli, A.M., Papini, M., Faccio, F., & Restelli, M. (2018). Policy optimization via importance sampling. arXiv preprint arXiv:1809.06098"},{"key":"6423_CR20","unstructured":"Microsoft: Microsoft TypeScript. https:\/\/github.com\/microsoft\/TypeScript. Accessed: 2019-08-20 (2019)"},{"key":"6423_CR21","doi-asserted-by":"crossref","unstructured":"Nickow, A., Oreopoulos, P., & Quan, V. (2020). The impressive effects of tutoring on prek-12 learning: A systematic review and meta-analysis of the experimental evidence. working paper 27476. National Bureau of Economic Research","DOI":"10.3386\/w27476"},{"key":"6423_CR22","first-page":"14810","volume":"35","author":"A Nie","year":"2022","unstructured":"Nie, A., Flet-Berliac, Y., Jordan, D., Steenbergen, W., & Brunskill, E. (2022). Data-efficient pipeline for offline reinforcement learning with limited data. Advances in Neural Information Processing Systems, 35, 14810\u201314823.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"6423_CR23","doi-asserted-by":"publisher","first-page":"687","DOI":"10.1609\/aaai.v33i01.3301687","volume":"33","author":"HW Park","year":"2019","unstructured":"Park, H. W., Grover, I., Spaulding, S., Gomez, L., & Breazeal, C. (2019). A model-free affective reinforcement learning approach to personalization of an autonomous social robot companion for early literacy education. AAAI, 33, 687\u2013694.","journal-title":"AAAI"},{"key":"6423_CR24","doi-asserted-by":"crossref","unstructured":"Pomerleau, D. (1990). Rapidly adapting artificial neural networks for autonomous navigation. NeurIPS 3","DOI":"10.1162\/neco.1991.3.1.88"},{"key":"6423_CR25","unstructured":"Postgres: Postgres. https:\/\/www.postgresql.org\/. Accessed: 2019-08-20 (2019)"},{"key":"6423_CR26","unstructured":"Projects, T.P.: Flask. https:\/\/flask.palletsprojects.com\/. Accessed: 2021-03-03 (2010)"},{"key":"6423_CR27","doi-asserted-by":"crossref","unstructured":"Rowe, J.P., Lester, J.C. (2015). Improving student problem solving in narrative-centered learning environments: A modular reinforcement learning framework. In International Conference on Artificial Intelligence in Education, pp. 419\u2013428. Springer","DOI":"10.1007\/978-3-319-19773-9_42"},{"key":"6423_CR28","doi-asserted-by":"publisher","unstructured":"Ruan, S., He, J., Ying, R., Burkle, J., Hakim, D., Wang, A., Yin, Y., Zhou, L., Xu, Q., AbuHashem, A., Dietz, G., Murnane, E.L., Brunskill, E., & Landay, J.A. (2020). Supporting children\u2019s math learning with feedback-augmented narrative technology. In IDC, pp. 567\u2013580. https:\/\/doi.org\/10.1145\/3392063.3394400.","DOI":"10.1145\/3392063.3394400"},{"key":"6423_CR29","doi-asserted-by":"crossref","unstructured":"Sammut, C., Hurst, S., Kedzier, D., & Michie, D. (1992). Learning to fly. In Machine Learning Proceedings 1992, pp. 385\u2013393. Elsevier.","DOI":"10.1016\/B978-1-55860-247-2.50055-3"},{"key":"6423_CR30","unstructured":"Schaarschmidt, M., Mika, S., Fricke, K., & Yoneki, E. (2019). Rlgraph: Modular computation graphs for deep reinforcement learning. In Proceedings of the 2nd Conference on Systems and Machine Learning (SysML)"},{"key":"6423_CR31","unstructured":"Schulman, J., Wolski, F., Dhariwal, P., Radford, A., & Klimov, O. (2017). Proximal policy optimization algorithms. arXiv preprint arXiv:1707.06347"},{"key":"6423_CR32","doi-asserted-by":"crossref","unstructured":"Shen, S., & Chi, M. (2016). Reinforcement learning: the sooner the better, or the later the better? In UMAP, pp. 37\u201344.","DOI":"10.1145\/2930238.2930247"},{"issue":"6419","key":"6423_CR33","doi-asserted-by":"publisher","first-page":"1140","DOI":"10.1126\/science.aar6404","volume":"362","author":"D Silver","year":"2018","unstructured":"Silver, D., Hubert, T., Schrittwieser, J., Antonoglou, I., Lai, M., Guez, A., Lanctot, M., Sifre, L., Kumaran, D., Graepel, T., et al. (2018). A general reinforcement learning algorithm that masters chess, shogi, and go through self-play. Science, 362(6419), 1140\u20131144.","journal-title":"Science"},{"key":"6423_CR34","unstructured":"Sundararajan, M., Taly, A., & Yan, Q. (2017). Axiomatic attribution for deep networks. In International Conference on Machine Learning, pp. 3319\u20133328. PMLR"},{"issue":"4","key":"6423_CR35","doi-asserted-by":"publisher","first-page":"197","DOI":"10.1080\/00461520.2011.611369","volume":"46","author":"K VanLehn","year":"2011","unstructured":"VanLehn, K. (2011). The relative effectiveness of human tutoring, intelligent tutoring systems, and other tutoring systems. Educational Psychologist, 46(4), 197\u2013221.","journal-title":"Educational Psychologist"},{"issue":"2","key":"6423_CR36","doi-asserted-by":"publisher","first-page":"454","DOI":"10.1007\/s40593-021-00269-9","volume":"32","author":"G Zhou","year":"2022","unstructured":"Zhou, G., Azizsoltani, H., Ausin, M. S., Barnes, T., & Chi, M. (2022). Leveraging granularity: Hierarchical reinforcement learning for pedagogical policy induction. International Journal of Artificial Intelligence in Education, 32(2), 454\u2013500.","journal-title":"International Journal of Artificial Intelligence in Education"},{"key":"6423_CR37","doi-asserted-by":"crossref","unstructured":"Zhou, G., Azizsoltani, H., Ausin, M.S., Barnes, T., & Chi, M. (2019). Hierarchical reinforcement learning for pedagogical policy induction. In Artificial Intelligence in Education: 20th International Conference, AIED 2019, Chicago, IL, USA, June 25\u201329, 2019, Proceedings, Part I 20, pp. 544\u2013556. Springer","DOI":"10.1007\/978-3-030-23204-7_45"}],"container-title":["Machine Learning"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10994-023-06423-9.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s10994-023-06423-9\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10994-023-06423-9.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2024,5,2]],"date-time":"2024-05-02T18:13:47Z","timestamp":1714673627000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s10994-023-06423-9"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,2,9]]},"references-count":37,"journal-issue":{"issue":"5","published-print":{"date-parts":[[2024,5]]}},"alternative-id":["6423"],"URL":"https:\/\/doi.org\/10.1007\/s10994-023-06423-9","relation":{},"ISSN":["0885-6125","1573-0565"],"issn-type":[{"value":"0885-6125","type":"print"},{"value":"1573-0565","type":"electronic"}],"subject":[],"published":{"date-parts":[[2024,2,9]]},"assertion":[{"value":"10 April 2023","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"27 September 2023","order":2,"name":"revised","label":"Revised","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"3 October 2023","order":3,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"9 February 2024","order":4,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"The authors declare no conflict of interest.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Conflict of interest"}},{"value":"The two experimental studies were approved by the Stanford Institutional Review Board.","order":3,"name":"Ethics","group":{"name":"EthicsHeading","label":"Ethical approval"}},{"value":"A copy of the subject consent forms are provided in Figs. , ,  and .","order":4,"name":"Ethics","group":{"name":"EthicsHeading","label":"Consent to participate"}},{"value":"Not relevant.","order":5,"name":"Ethics","group":{"name":"EthicsHeading","label":"Consent for publication"}}]}}