{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,20]],"date-time":"2026-06-20T04:52:45Z","timestamp":1781931165727,"version":"3.54.5"},"reference-count":32,"publisher":"Association for Natural Language Processing","issue":"2","content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Journal of Natural Language Processing"],"published-print":{"date-parts":[[2026]]},"DOI":"10.5715\/jnlp.33.630","type":"journal-article","created":{"date-parts":[[2026,6,14]],"date-time":"2026-06-14T22:11:39Z","timestamp":1781475099000},"page":"630-657","source":"Crossref","is-referenced-by-count":0,"title":["AnaToM: A Dataset Generation Framework for Evaluating Theory of Mind Reasoning Toward Anatomy of Difficulty Through Structurally Controlled Story Generation"],"prefix":"10.5715","volume":"33","author":[{"given":"Jundai","family":"Suzuki","sequence":"first","affiliation":[{"name":"Tokyo Denki University"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Ryoma","family":"Ishigaki","sequence":"additional","affiliation":[{"name":"Tokyo Denki University"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Eisaku","family":"Maeda","sequence":"additional","affiliation":[{"name":"Tokyo Denki University"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"3685","reference":[{"key":"1","doi-asserted-by":"crossref","unstructured":"Baron-Cohen, S., Leslie, A. M., and Frith, U. (1985). \u201cDoes the Autistic Child Have a \u201cTheory of Mind\u201d ?\u201d  <i>Cognition<\/i>, <b>21<\/b> (1), pp. 37\u201346.","DOI":"10.1016\/0010-0277(85)90022-8"},{"key":"2","doi-asserted-by":"crossref","unstructured":"Baron-Cohen, S., O\u2019Riordan, M., Stone, V., Jones, R., and Plaisted, K. (1999). \u201cRecognition of Faux Pas by Normally Developing Children and Children with Asperger Syndrome or High-functioning Autism.\u201d  <i>Journal of Autism and Developmental Disorders<\/i>, <b>29<\/b> (5), pp. 407\u2013418.","DOI":"10.1023\/A:1023035012436"},{"key":"3","doi-asserted-by":"crossref","unstructured":"Beaudoin, C., Leblanc, \u00c9., Gagner, C., and Beauchamp, M. H. (2020). \u201cSystematic Review and Inventory of Theory of Mind Measures for Young Children.\u201d  <i>Frontiers in Psychology<\/i>, <b>10<\/b>.","DOI":"10.3389\/fpsyg.2019.02905"},{"key":"4","doi-asserted-by":"crossref","unstructured":"Chen, Z., Wu, J., Zhou, J., Wen, B., Bi, G., Jiang, G., Cao, Y., Hu, M., Lai, Y., Xiong, Z., and Huang, M. (2024). \u201cToMBench: Benchmarking Theory of Mind in Large Language Models.\u201d  In <i>Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)<\/i>, pp. 15959\u201315983, Bangkok, Thailand. Association for Computational Linguistics.","DOI":"10.18653\/v1\/2024.acl-long.847"},{"key":"5","doi-asserted-by":"crossref","unstructured":"Gandhi, K., Fr\u00e4nken, J.-P., Gerstenberg, T., and Goodman, N. D. (2023). \u201cUnderstanding Social Reasoning in Language Models with Language Models.\u201d  In <i>Proceedings of the 37th International Conference on Neural Information Processing Systems<\/i>, NIPS \u201923, pp. 13518\u201313529, Red Hook, NY, USA. Curran Associates Inc.","DOI":"10.52202\/075280-0595"},{"key":"6","unstructured":"Grant, E., Nematzadeh, A., and Griffiths, T. L. (2017). \u201cHow Can Memory-Augmented Neural Networks Pass a False-Belief Task?\u201d  In <i>Proceedings of the Annual Meeting of the Cognitive Science Society<\/i>, Vol. 39, pp. 427\u2013432."},{"key":"7","unstructured":"Grattafiori, A., Dubey, A., Jauhri, A., Pandey, A., Kadian, A., Al-Dahle, A., Letman, A., Mathur, A., Schelten, A., Vaughan, A., et al. (2024). \u201cThe Llama 3 Herd of Models.\u201d  <i>arXiv preprint arXiv:2407.21783<\/i>. Version 3."},{"key":"8","unstructured":"Isozaki, H., and Katsuno, H. (1996). \u201cA Semantic Characterization of an Algorithm for Estimating Others\u2019 Beliefs from Observation.\u201d  In <i>Proceedings of the 13th National Conference on Artificial Intelligence and 8th Innovative Applications of Artificial Intelligence Conference, AAAI 96, IAAI 96, Portland, Oregon, USA, August 4\u20138, 1996, Volume 1<\/i>, pp. 543\u2013549. AAAI Press \/ The MIT Press."},{"key":"9","doi-asserted-by":"crossref","unstructured":"Kosinski, M. (2024). \u201cEvaluating Large Language Models in Theory of Mind Tasks.\u201d  <i>Proceedings of the National Academy of Sciences<\/i>, <b>121<\/b> (45), p. e2405460121.","DOI":"10.1073\/pnas.2405460121"},{"key":"10","doi-asserted-by":"crossref","unstructured":"Le, M., Boureau, Y.-L., and Nickel, M. (2019). \u201cRevisiting the Evaluation of Theory of Mind through Question Answering.\u201d  In <i>Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP)<\/i>, pp. 5872\u20135877, Hong Kong, China. Association for Computational Linguistics.","DOI":"10.18653\/v1\/D19-1598"},{"key":"11","doi-asserted-by":"crossref","unstructured":"Li, H., Chong, Y., Stepputtis, S., Campbell, J., Hughes, D., Lewis, C., and Sycara, K. (2023). \u201cTheory of Mind for Multi-Agent Collaboration via Large Language Models.\u201d  In <i>Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing<\/i>, pp. 180\u2013192, Singapore. Association for Computational Linguistics.","DOI":"10.18653\/v1\/2023.emnlp-main.13"},{"key":"12","doi-asserted-by":"crossref","unstructured":"Ma, Z., Sansom, J., Peng, R., and Chai, J. (2023). \u201cTowards A Holistic Landscape of Situated Theory of Mind in Large Language Models.\u201d  In <i>Findings of the Association for Computational Linguistics: EMNLP 2023<\/i>, pp. 1011\u20131031, Singapore. Association for Computational Linguistics.","DOI":"10.18653\/v1\/2023.findings-emnlp.72"},{"key":"13","doi-asserted-by":"crossref","unstructured":"Markowska, M., Taghizadeh, M., Soubki, A., Mirroshandel, S., and Rambow, O. (2023). \u201cFinding Common Ground: Annotating and Predicting Common Ground in Spoken Conversations.\u201d  In <i>Findings of the Association for Computational Linguistics: EMNLP 2023<\/i>, pp. 8221\u20138233, Singapore. Association for Computational Linguistics.","DOI":"10.18653\/v1\/2023.findings-emnlp.551"},{"key":"14","unstructured":"Moghaddam, S. R., and Honey, C. J. (2023). \u201cBoosting Theory-of-Mind Performance in Large Language Models via Prompting.\u201d  <i>arXiv preprint arXiv:2304.11490<\/i>. Version 3."},{"key":"15","doi-asserted-by":"crossref","unstructured":"Nematzadeh, A., Burns, K., Grant, E., Gopnik, A., and Griffiths, T. (2018). \u201cEvaluating Theory of Mind in Question Answering.\u201d  In <i>Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing<\/i>, pp. 2392\u20132400, Brussels, Belgium. Association for Computational Linguistics.","DOI":"10.18653\/v1\/D18-1261"},{"key":"16","doi-asserted-by":"crossref","unstructured":"Parmar, M., Patel, N., Varshney, N., Nakamura, M., Luo, M., Mashetty, S., Mitra, A., and Baral, C. (2024). \u201cLogicBench: Towards Systematic Evaluation of Logical Reasoning Ability of Large Language Models.\u201d  In <i>Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)<\/i>, pp. 13679\u201313707, Bangkok, Thailand. Association for Computational Linguistics.","DOI":"10.18653\/v1\/2024.acl-long.739"},{"key":"17","doi-asserted-by":"crossref","unstructured":"Premack, D., and Woodruff, G. (1978). \u201cDoes the Chimpanzee Have a Theory of Mind?\u201d  <i>Behavioral and Brain Sciences<\/i>, <b>1<\/b> (4), pp. 515\u2013526.","DOI":"10.1017\/S0140525X00076512"},{"key":"18","doi-asserted-by":"crossref","unstructured":"Sap, M., Le Bras, R., Fried, D., and Choi, Y. (2022). \u201cNeural Theory-of-Mind? On the Limits of Social Intelligence in Large LMs.\u201d  In <i>Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing<\/i>, pp. 3762\u20133780, Abu Dhabi, United Arab Emirates. Association for Computational Linguistics.","DOI":"10.18653\/v1\/2022.emnlp-main.248"},{"key":"19","doi-asserted-by":"crossref","unstructured":"Sap, M., Rashkin, H., Chen, D., Le Bras, R., and Choi, Y. (2019). \u201cSocial IQa: Commonsense Reasoning about Social Interactions.\u201d  In <i>Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP)<\/i>, pp. 4463\u20134473, Hong Kong, China. Association for Computational Linguistics.","DOI":"10.18653\/v1\/D19-1454"},{"key":"20","unstructured":"Sar\u0131ta\u015f, K., Tez\u00f6ren, K., and Durmazkeser, Y. (2025). \u201cA Systematic Review on the Evaluation of Large Language Models in Theory of Mind Tasks.\u201d  <i>arXiv preprint arXiv:2502.08796<\/i>. Version 1."},{"key":"21","unstructured":"Sclar, M., Dwivedi-Yu, J., Fazel-Zarandi, M., Tsvetkov, Y., Bisk, Y., Choi, Y., and Celikyilmaz, A. (2025). \u201cExplore Theory of Mind: Program-guided Adversarial Data Generation for Theory of Mind Reasoning.\u201d  In <i>International Conference on Representation Learning<\/i>, Vol. 2025, pp. 67635\u201367660."},{"key":"22","doi-asserted-by":"crossref","unstructured":"Shapira, N., Zwirn, G., and Goldberg, Y. (2023). \u201cHow Well Do Large Language Models Perform on Faux Pas Tests?\u201d  In <i>Findings of the Association for Computational Linguistics: ACL 2023<\/i>, pp. 10438\u201310451, Toronto, Canada. Association for Computational Linguistics.","DOI":"10.18653\/v1\/2023.findings-acl.663"},{"key":"23","doi-asserted-by":"crossref","unstructured":"Shinoda, K., Hojo, N., Nishida, K., Mizuno, S., Suzuki, K., Masumura, R., Sugiyama, H., and Saito, K. (2025). \u201cToMATO: Verbalizing the Mental States of Role-Playing LLMs for Benchmarking Theory of Mind.\u201d  In <i>AAAI-25, Sponsored by the Association for the Advancement of Artificial Intelligence, February 25\u2013March 4, 2025, Philadelphia, PA, USA<\/i>, pp. 1520\u20131528. AAAI Press.","DOI":"10.1609\/aaai.v39i2.32143"},{"key":"24","doi-asserted-by":"crossref","unstructured":"Soubki, A., Murzaku, J., Yousefi Jordehi, A., Zeng, P., Markowska, M., Mirroshandel, S. A., and Rambow, O. (2024). \u201cViews Are My Own, but Also Yours: Benchmarking Theory of Mind Using Common Ground.\u201d  In <i>Findings of the Association for Computational Linguistics: ACL 2024<\/i>, pp. 14815\u201314823, Bangkok, Thailand. Association for Computational Linguistics.","DOI":"10.18653\/v1\/2024.findings-acl.880"},{"key":"25","doi-asserted-by":"crossref","unstructured":"Suzuki, J., Ishigaki, R., and Maeda, E. (2025). \u201cAnaToM: A Dataset Generation Framework for Evaluating Theory of Mind Reasoning Toward the Anatomy of Difficulty through Structurally Controlled Story Generation.\u201d  In <i>Proceedings of the 14th International Joint Conference on Natural Language Processing and the 4th Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics<\/i>, pp. 244\u2013257, Mumbai, India. The Asian Federation of Natural Language Processing and The Association for Computational Linguistics.","DOI":"10.18653\/v1\/2025.findings-ijcnlp.14"},{"key":"26","unstructured":"Ullman, T. (2023). \u201cLarge Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks.\u201d  <i>arXiv preprint arXiv:2302.08399<\/i>. Version 5."},{"key":"27","doi-asserted-by":"crossref","unstructured":"Wei, J., Wang, X., Schuurmans, D., Bosma, M., Ichter, B., Xia, F., Chi, E., Le, Q., and Zhou, D. (2023). \u201cChain-of-Thought Prompting Elicits Reasoning in Large Language Models.\u201d  <i>arXiv preprint arXiv:2201.11903<\/i>. Version 6.","DOI":"10.52202\/068431-1800"},{"key":"28","unstructured":"Weston, J., Bordes, A., Chopra, S., Rush, A. M., van Merri\u00ebnboer, B., Joulin, A., and Mikolov, T. (2015). \u201cTowards AI-Complete Question Answering: A Set of Prerequisite Toy Tasks.\u201d  <i>arXiv preprint arXiv:1502.05698<\/i>. Version 10."},{"key":"29","doi-asserted-by":"crossref","unstructured":"Wilf, A., Lee, S., Liang, P. P., and Morency, L.-P. (2024). \u201cThink Twice: Perspective-Taking Improves Large Language Models\u2019 Theory-of-Mind Capabilities.\u201d  In <i>Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)<\/i>, pp. 8292\u20138308, Bangkok, Thailand. Association for Computational Linguistics.","DOI":"10.18653\/v1\/2024.acl-long.451"},{"key":"30","doi-asserted-by":"crossref","unstructured":"Wimmer, H., and Perner, J. (1983). \u201cBeliefs about Beliefs: Representation and Constraining Function of Wrong Beliefs in Young Children\u2019s Understanding of Deception.\u201d  <i>Cognition<\/i>, <b>13<\/b> (1), pp. 103\u2013128.","DOI":"10.1016\/0010-0277(83)90004-5"},{"key":"31","doi-asserted-by":"crossref","unstructured":"Wu, Y., He, Y., Jia, Y., Mihalcea, R., Chen, Y., and Deng, N. (2023). \u201cHi-ToM: A Benchmark for Evaluating Higher-Order Theory of Mind Reasoning in Large Language Models.\u201d  In <i>Findings of the Association for Computational Linguistics: EMNLP 2023<\/i>, pp. 10691\u201310706, Singapore. Association for Computational Linguistics.","DOI":"10.18653\/v1\/2023.findings-emnlp.717"},{"key":"32","doi-asserted-by":"crossref","unstructured":"Xu, H., Zhao, R., Zhu, L., Du, J., and He, Y. (2024). \u201cOpenToM: A Comprehensive Benchmark for Evaluating Theory-of-Mind Reasoning Capabilities of Large Language Models.\u201d  In <i>Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)<\/i>, pp. 8593\u20138623, Bangkok, Thailand. Association for Computational Linguistics.","DOI":"10.18653\/v1\/2024.acl-long.466"}],"container-title":["Journal of Natural Language Processing"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.jstage.jst.go.jp\/article\/jnlp\/33\/2\/33_630\/_pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,6,20]],"date-time":"2026-06-20T04:43:46Z","timestamp":1781930626000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.jstage.jst.go.jp\/article\/jnlp\/33\/2\/33_630\/_article\/-char\/ja\/"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026]]},"references-count":32,"journal-issue":{"issue":"2","published-print":{"date-parts":[[2026]]}},"URL":"https:\/\/doi.org\/10.5715\/jnlp.33.630","relation":{},"ISSN":["1340-7619","2185-8314"],"issn-type":[{"value":"1340-7619","type":"print"},{"value":"2185-8314","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026]]}}}