{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,13]],"date-time":"2026-05-13T17:16:59Z","timestamp":1778692619971,"version":"3.51.4"},"reference-count":33,"publisher":"Springer Science and Business Media LLC","issue":"4","license":[{"start":{"date-parts":[[2026,3,10]],"date-time":"2026-03-10T00:00:00Z","timestamp":1773100800000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2026,3,10]],"date-time":"2026-03-10T00:00:00Z","timestamp":1773100800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"DOI":"10.13039\/501100004462","name":"Consiglio Nazionale Delle Ricerche","doi-asserted-by":"crossref","id":[{"id":"10.13039\/501100004462","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Mach Learn"],"published-print":{"date-parts":[[2026,4]]},"abstract":"<jats:title>Abstract<\/jats:title>\n                  <jats:p>We introduce GSM-Identity, a pipeline to modify existing mathematical reasoning benchmarks by adding extra complexity to the questions while preserving their fundamental meaning. By systematically transforming numerical values in the GSM8K dataset into mathematically equivalent but less obvious expressions, we create a benchmark to measure Large Language Models (LLMs) mathematical understanding. We evaluate LLMs ranging from 7 billions to 72 billions parameters using multiple prompting strategies, including standard, notice-based, and chain-of-thought approaches. We find that Math oriented models can retain most of their performance on GSM8K when evaluated on GSM-Identity, while general purpose models show significant performance degradation. A comparison with human evaluations reveals that models in the 7 billion parameters range perform similar to humans when exposed to the kind of modifications we study, while models with more than 70 billion parameters are more accurate than humans in answering the questions and they are also more resilient to modifications. Our findings highlight GSM-Identity as a valuable tool for distinguishing reasoning from memorization, offering insights into the abilities of LLMs to understand higher level mathematical concepts.<\/jats:p>","DOI":"10.1007\/s10994-026-07029-7","type":"journal-article","created":{"date-parts":[[2026,4,3]],"date-time":"2026-04-03T10:12:42Z","timestamp":1775211162000},"update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":0,"title":["GSM-Identity: Evaluating Mathematical Reasoning in LLMs via Equivalence Transformations"],"prefix":"10.1007","volume":"115","author":[{"given":"Kajal","family":"Negi","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Giovanni","family":"Puccetti","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Andrea","family":"Esuli","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"297","published-online":{"date-parts":[[2026,3,10]]},"reference":[{"key":"7029_CR1","unstructured":"Baevski, A., Zhou, H., Mohamed, A., & Auli, M. (2020). wav2vec 2.0: a framework for self-supervised learning of speech representations. In Proceedings of the 34th international conference on neural information processing systems. Curran Associates."},{"key":"7029_CR2","doi-asserted-by":"crossref","unstructured":"Bender, E. M., Gebru, T., McMillan-Major, A., & Shmitchell, S. (2021). On the dangers of stochastic parrots: Can language models be too big? In Proceedings of the 2021 ACM conference on fairness, accountability, and transparency (pp. 610\u2013623). Association for Computing Machinery.","DOI":"10.1145\/3442188.3445922"},{"key":"7029_CR3","unstructured":"Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., & Amodei, D. (2020). Language models are few-shot learners. In Advances in neural information processing systems (Vol.\u00a033, pp. 1877\u20131901). Curran Associates."},{"key":"7029_CR4","unstructured":"Cobbe, K., Kosaraju, V., Bavarian, M., Chen, M., Jun, H., Kaiser, L., & Schulman, J. (2021). Training verifiers to solve math word problems. arXiv:2110.14168"},{"key":"7029_CR5","unstructured":"DeepSeek-AI (2025). Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning. arXiv:org\/abs\/2501.12948"},{"key":"7029_CR6","doi-asserted-by":"crossref","unstructured":"Dong, Q., Li, L., Dai, D., Zheng, C., Ma, J., Li, R., & Sui, Z. (2024). A survey on in-context learning. In Y.\u00a0Al-Onaizan, M.\u00a0Bansal, & Y.-N.\u00a0Chen (Eds.), Proceedings of EMNLP-2024 (pp. 1107\u20131128). Miami, Florida, USA.","DOI":"10.18653\/v1\/2024.emnlp-main.64"},{"key":"7029_CR7","unstructured":"Grattafiori, A., Dubey, A., Jauhri, A., Pandey, A., Kadian, A., Al-Dahle, A., & Ma, Z. (2024). The llama 3 herd of models. arXiv:org\/abs\/2407.21783"},{"key":"7029_CR8","unstructured":"Gulati, A., Miranda, B., Chen, E., Xia, E., Fronsdal, K., de Moraes\u00a0Dumont, B., & Koyejo, S. (2024). Putnam-AXIOM: A functional and static benchmark for measuring higher level mathematical reasoning. In The 4th workshop on mathematical reasoning and AI at NEURIPS\u201924. https:\/\/openreview.net\/forum?id=YXnwlZe0yf"},{"key":"7029_CR9","unstructured":"Hendrycks, D., Burns, C., Kadavath, S., Arora, A., Basart, S., Tang, E., & Steinhardt, J. (2021). Measuring mathematical problem solving with the MATH dataset. In Thirty-fifth conference on neural information processing systems datasets and benchmarks track (round 2). https:\/\/openreview.net\/forum?id=7Bywt2mQsCe"},{"key":"7029_CR10","unstructured":"Huang, K., Guo, J., Li, Z., Ji, X., Ge, J., Li, W., & Wang, M. (2025). MATH-Perturb: Benchmarking llms\u2019 math reasoning abilities against hard perturbations. Retrieved March 27, 2025, from arXiv:2502.06453 [cs]"},{"key":"7029_CR11","unstructured":"Hui, B., Yang, J., Cui, Z., Yang, J., Liu, D., Zhang, L., & Lin, J. (2024). Qwen2.5-coder technical report. arXiv:org\/abs\/2409.12186"},{"issue":"10","key":"7029_CR12","doi-asserted-by":"publisher","first-page":"1161","DOI":"10.1038\/s42256-023-00729-y","volume":"5","author":"D Hupkes","year":"2023","unstructured":"Hupkes, D., Giulianelli, M., Dankers, V., Artetxe, M., Elazar, Y., Pimentel, T., & Jin, Z. (2023). A taxonomy and review of generalization research in NLP. Nature Machine Intelligence, 5(10), 1161\u20131174.","journal-title":"Nature Machine Intelligence"},{"key":"7029_CR13","doi-asserted-by":"crossref","unstructured":"Kwon, W., Li, Z., Zhuang, S., Sheng, Y., Zheng, L., Yu, C.H., & Stoica, I. (2023). Efficient memory management for large language model serving with pagedattention. In Proceedings of the 29th symposium on operating systems principles (pp. 611\u2013626). Association for Computing Machinery.","DOI":"10.1145\/3600006.3613165"},{"key":"7029_CR14","doi-asserted-by":"crossref","unstructured":"Li, Q., Cui, L., Zhao, X., Kong, L., & Bi, W. (2024). Gsm-plus: A comprehensive benchmark for evaluating the robustness of LLMS as mathematical problem solvers. arXiv:org\/abs\/2402.19255","DOI":"10.18653\/v1\/2024.acl-long.163"},{"key":"7029_CR15","doi-asserted-by":"crossref","unstructured":"Li, Y.A., Han, C., Raghavan, V., Mischler, G., & Mesgarani, N. (2023). Styletts 2: Towards human-level text-to-speech through style diffusion and adversarial training with large speech language models. In Advances in neural information processing systems (Vol.\u00a036, pp. 19594\u201319621). Curran Associates.","DOI":"10.52202\/075280-0860"},{"key":"7029_CR16","doi-asserted-by":"crossref","unstructured":"Lugmayr, A., Danelljan, M., Romero, A., Yu, F., Timofte, R., & Van\u00a0Gool, L. (2022). Repaint: Inpainting using denoising diffusion probabilistic models. In 2022 IEEE\/CVF conference on computer vision and pattern recognition (CVPR) (pp. 11451\u201311461).","DOI":"10.1109\/CVPR52688.2022.01117"},{"issue":"3","key":"7029_CR17","doi-asserted-by":"publisher","first-page":"276","DOI":"10.11613\/BM.2012.031","volume":"22","author":"ML McHugh","year":"2012","unstructured":"McHugh, M. L. (2012). Interrater reliability: The kappa statistic. Biochemia medica, 22(3), 276\u2013282.","journal-title":"Biochemia medica"},{"key":"7029_CR18","unstructured":"Mirzadeh, S. I., Alizadeh, K., Shahrokhi, H., Tuzel, O., Bengio, S., & Farajtabar, M. (2025). GSM-symbolic: Understanding the limitations of mathematical reasoning in large language models. In The 13th international conference on learning representations. https:\/\/openreview.net\/forum?id=AjXkRZIvjB"},{"key":"7029_CR19","unstructured":"Mistral AI Team. (2023). Mistral 7b. arXiv:org\/abs\/2310.06825"},{"key":"7029_CR20","doi-asserted-by":"publisher","unstructured":"Mitchell, M. (2021). Why AI is harder than we think. In Proceedings of the genetic and evolutionary computation conference (p. 3). Association for Computing Machinery. https:\/\/doi.org\/10.1145\/3449639.3465421","DOI":"10.1145\/3449639.3465421"},{"key":"7029_CR21","unstructured":"Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., & Sutskever, I. (2019). Language models are unsupervised multitask learners. https:\/\/api.semanticscholar.org\/CorpusID:160025533"},{"key":"7029_CR22","unstructured":"Ray\u00a0Choudhury, S., Rogers, A., & Augenstein, I. (2022). Machine reading, fast and slow: When do models \u201cunderstand\u201d language? In Proceedings of coling-2022 (pp. 78\u201393). International Committee on Computational Linguistics."},{"key":"7029_CR23","unstructured":"Shah, V., Yu, D., Lyu, K., Park, S., Ke, N. R., Mozer, M. C., & Goyal, A. (2024). AI-assisted generation of difficult math questions. InThe 4th workshop on mathematical reasoning and AI at NEURIPS\u201924. https:\/\/openreview.net\/forum?id=6pUdfsJmd1"},{"key":"7029_CR24","unstructured":"Shi, F., Chen, X., Misra, K., Scales, N., Dohan, D., Chi, E., & Zhou, D. (2023). Large language models can be easily distracted by irrelevant context. Proceedings of the 40th international conference on machine learning. www.JMLR.org."},{"key":"7029_CR25","unstructured":"Srivastava, A., Rastogi, A., Rao, A., Shoeb, A.A.M., Abid, A., Fisch, A., & Wu, Z. (2023). Beyond the imitation game: Quantifying and extrapolating the capabilities of language models. Transactions on Machine Learning Research. https:\/\/openreview.net\/forum?id=uyTL5Bvosj (Featured Certification)"},{"key":"7029_CR26","unstructured":"Team, G., Mesnard, T., Hardin, C., Dadashi, R., Bhupatiraju, S., Pathak, S., & Kenealy, K. (2024). Gemma: Open models based on gemini research and technology. arXiv:org\/abs\/2403.08295"},{"key":"7029_CR27","unstructured":"Touvron, H., Lavril, T., Izacard, G., Martinet, X., Lachaux, M-A., Lacroix, T., & Lample, G. (2023). Llama: Open and efficient foundation language models. arXiv:org\/abs\/2302.13971"},{"key":"7029_CR28","unstructured":"Vendrow, J., Vendrow, E., Beery, S., & Madry, A. (2025). Do large language model benchmarks test reliability?. arXiv:org\/abs\/2502.03461"},{"key":"7029_CR29","unstructured":"Wei, J., Wang, X., Schuurmans, D., Bosma, M., Ichter, B., Xia, F.& Zhou, D. (2022). Chain-of-thought prompting elicits reasoning in large language models. In Proceedings of the 36th international conference on neural information processing systems. Curran Associates Inc."},{"key":"7029_CR30","doi-asserted-by":"crossref","unstructured":"Wolf, T., Debut, L., Sanh, V., Chaumond, J., Delangue, C., Moi, A., & Rush, A. (2020). Transformers: State-of-the-art natural language processing. In Q.\u00a0Liu & D.\u00a0Schlangen (Eds.), Proceedings of EMNLP-2020 (pp. 38\u201345).","DOI":"10.18653\/v1\/2020.emnlp-demos.6"},{"key":"7029_CR31","unstructured":"Xu, R., Wang, Z., Fan, R.-Z., & Liu, P. (2024). Benchmarking benchmark leakage in large language models. arXiv:org\/abs\/2404.18824"},{"key":"7029_CR32","unstructured":"Yang, A., Zhang, B., Hui, B., Gao, B., Yu, B., Li, C., & Zhang, Z. (2024). Qwen2.5-math technical report: Toward mathematical expert model via self-improvement. arxiv:org\/abs\/2409.12122"},{"key":"7029_CR33","unstructured":"Zhou, Y., Liu, H., Chen, Z., Tian, Y., & Chen, B. (2025). GSM-infinite: How do your LLMS behave over infinitely increasing context length and reasoning complexity?. arXiv:org\/abs\/2502.05252"}],"container-title":["Machine Learning"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10994-026-07029-7.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s10994-026-07029-7","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10994-026-07029-7.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,5,13]],"date-time":"2026-05-13T16:36:05Z","timestamp":1778690165000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s10994-026-07029-7"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,3,10]]},"references-count":33,"journal-issue":{"issue":"4","published-print":{"date-parts":[[2026,4]]}},"alternative-id":["7029"],"URL":"https:\/\/doi.org\/10.1007\/s10994-026-07029-7","relation":{},"ISSN":["0885-6125","1573-0565"],"issn-type":[{"value":"0885-6125","type":"print"},{"value":"1573-0565","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,3,10]]},"assertion":[{"value":"3 April 2025","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"10 February 2026","order":2,"name":"revised","label":"Revised","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"10 March 2026","order":3,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"10 March 2026","order":4,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"Authors have no conflict of interest or conflict of interest with respect to the content of this article.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Conflict of interest"}},{"value":"Not applicable.","order":3,"name":"Ethics","group":{"name":"EthicsHeading","label":"Ethical Approval"}},{"value":"Not applicable.","order":4,"name":"Ethics","group":{"name":"EthicsHeading","label":"Consent to Participate"}},{"value":"Not applicable.","order":5,"name":"Ethics","group":{"name":"EthicsHeading","label":"Consent for Publication"}}],"article-number":"88"}}