{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,4,24]],"date-time":"2026-04-24T16:13:17Z","timestamp":1777047197990,"version":"3.51.4"},"reference-count":72,"publisher":"MIT Press","license":[{"start":{"date-parts":[[2026,4,24]],"date-time":"2026-04-24T00:00:00Z","timestamp":1776988800000},"content-version":"vor","delay-in-days":113,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":["direct.mit.edu"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2026,4,10]]},"abstract":"<jats:title>Abstract<\/jats:title>\n                  <jats:p>Large Language Models (LLMs) have great potential to accelerate and support scholarly peer review and are increasingly used as fully automatic review generators (ARGs). However, potential biases and systematic errors may pose significant risks to scientific integrity; understanding the specific capabilities and limitations of state-of-the-art ARGs is essential. We focus on a core reviewing skill that underpins high-quality peer review: detecting faulty research logic. This involves evaluating the internal consistency between a paper\u2019s results, interpretations, and claims. We present a fully automated counterfactual evaluation framework that isolates and tests this skill under controlled conditions. Testing a range of ARG approaches, we find that, contrary to expectation, flaws in research logic have no significant effect on their output reviews. Based on our findings, we derive three actionable recommendations for future work and release our counterfactual dataset and evaluation framework publicly.1<\/jats:p>","DOI":"10.1162\/tacl.a.642","type":"journal-article","created":{"date-parts":[[2026,4,24]],"date-time":"2026-04-24T15:19:42Z","timestamp":1777043982000},"page":"465-488","update-policy":"https:\/\/doi.org\/10.1162\/mitpressjournals.corrections.policy","source":"Crossref","is-referenced-by-count":0,"title":["Automatic Reviewers Fail to Detect Faulty Reasoning in Research Papers: A New Counterfactual Evaluation Framework"],"prefix":"10.1162","volume":"14","author":[{"given":"Nils","family":"Dycke","sequence":"first","affiliation":[{"name":"UKP Lab, Department of Computer Science and National Research Center for Applied Cybersecurity ATHENE Technical University of Darmstadt, Germany"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Iryna","family":"Gurevych","sequence":"additional","affiliation":[{"name":"UKP Lab, Department of Computer Science and National Research Center for Applied Cybersecurity ATHENE Technical University of Darmstadt, Germany"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"281","published-online":{"date-parts":[[2026,4,10]]},"reference":[{"key":"2026042411193802700_bib1","article-title":"Phi-4 technical report","author":"Abdin","year":"2024","journal-title":"arXiv preprint arXiv:2412.08905v1"},{"key":"2026042411193802700_bib2","article-title":"A review on language models as knowledge bases","author":"AlKhamissi","year":"2022","journal-title":"arXiv preprint arXiv:2204.06031v1"},{"key":"2026042411193802700_bib3","unstructured":"Anthropic. 2024. Claude 3.5 sonnet. https:\/\/www.anthropic.com\/news\/claude-3-5-sonnet. Accessed: 2025-11-05."},{"key":"2026042411193802700_bib4","doi-asserted-by":"publisher","DOI":"10.1017\/9781009092265","volume-title":"The Scientific Method: A Guide to Finding Useful Knowledge","author":"Armstrong","year":"2022"},{"key":"2026042411193802700_bib5","unstructured":"Association for the Advancement of Artificial Intelligence. 2025. AAAI launches AI-powered peer review assessment system. https:\/\/aaai.org\/aaai-launches-ai-powered-peer-review-assessment-system\/. Accessed: 2025-07-18."},{"key":"2026042411193802700_bib6","volume-title":"Novum Organum","author":"Bacon","year":"1878"},{"issue":"1","key":"2026042411193802700_bib7","doi-asserted-by":"publisher","first-page":"289","DOI":"10.1111\/j.2517-6161.1995.tb02031.x","article-title":"Controlling the false discovery rate: A practical and powerful approach to multiple testing","volume":"57","author":"Benjamini","year":"1995","journal-title":"Journal of the Royal Statistical Society: Series B (Methodological)"},{"key":"2026042411193802700_bib8","doi-asserted-by":"publisher","first-page":"303","DOI":"10.1086\/266520","article-title":"Communications through limited response questioning","author":"Bennett","year":"1954","journal-title":"Public Opinion Quarterly"},{"key":"2026042411193802700_bib9","doi-asserted-by":"publisher","DOI":"10.3389\/fncom.2011.00056","article-title":"Alternatives to peer review: Novel approaches for research evaluation","volume":"5","author":"Birukou","year":"2011","journal-title":"Frontiers in Computational Neuroscience"},{"issue":"1","key":"2026042411193802700_bib10","doi-asserted-by":"publisher","first-page":"197","DOI":"10.1002\/aris.2011.1440450112","article-title":"Scientific peer review","volume":"45","author":"Bornmann","year":"2011","journal-title":"Annual Review of Information Science and Technology"},{"key":"2026042411193802700_bib11","doi-asserted-by":"publisher","first-page":"9742","DOI":"10.18653\/v1\/2024.findings-acl.580","article-title":"Automated focused feedback generation for scientific writing assistance","volume-title":"Findings of the Association for Computational Linguistics: ACL 2024","author":"Chamoun","year":"2024"},{"key":"2026042411193802700_bib12","doi-asserted-by":"publisher","first-page":"15662","DOI":"10.18653\/v1\/2025.emnlp-main.790","article-title":"TreeReview: A dynamic tree of questions framework for deep and efficient LLM-based scientific peer review","volume-title":"Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing","author":"Chang","year":"2025"},{"key":"2026042411193802700_bib13","doi-asserted-by":"publisher","first-page":"5872","DOI":"10.18653\/v1\/2024.acl-long.320","article-title":"Masked thought: Simply masking partial reasoning steps can improve mathematical reasoning learning of language models","volume-title":"Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)","author":"Chen","year":"2024"},{"key":"2026042411193802700_bib14","article-title":"MARG: Multi-agent review generation for scientific papers","author":"D\u2019Arcy","year":"2024","journal-title":"arXiv preprint arXiv:2401.04259v1"},{"key":"2026042411193802700_bib15","article-title":"Deepseek-v3 technical report","author":"DeepSeek-AI","year":"2024","journal-title":"arXiv preprint arXiv:2412.19437v2"},{"key":"2026042411193802700_bib16","doi-asserted-by":"publisher","first-page":"5081","DOI":"10.18653\/v1\/2024.emnlp-main.292","article-title":"LLMs assist NLP researchers: Critique paper (meta-)reviewing","volume-title":"Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing","author":"Jiangshu","year":"2024"},{"key":"2026042411193802700_bib17","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1162\/COLI.a.19","article-title":"Problem solving through human\u2013AI preference-based cooperation","author":"Dutta","year":"2025","journal-title":"Computational Linguistics"},{"key":"2026042411193802700_bib18","doi-asserted-by":"publisher","first-page":"5049","DOI":"10.18653\/v1\/2023.acl-long.277","article-title":"NLPeer: A unified resource for the computational study of peer review","volume-title":"Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)","author":"Dycke","year":"2023"},{"key":"2026042411193802700_bib19","doi-asserted-by":"publisher","first-page":"187","DOI":"10.18653\/v1\/2023.argmining-1.21","article-title":"Overview of PragTag-2023: Low-resource multi-domain pragmatic tagging of peer reviews","volume-title":"Proceedings of the 10th Workshop on Argument Mining","author":"Dycke","year":"2023"},{"key":"2026042411193802700_bib20","doi-asserted-by":"publisher","first-page":"22687","DOI":"10.18653\/v1\/2025.acl-long.1107","article-title":"STRICTA: Structured reasoning in critical text assessment for peer review and beyond","volume-title":"Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)","author":"Dycke","year":"2025"},{"key":"2026042411193802700_bib21","article-title":"Large language models (LLMs) on tabular data: Prediction, generation, and understanding - a survey","author":"Xi","year":"2024","journal-title":"Transactions on Machine Learning Research"},{"issue":"6","key":"2026042411193802700_bib22","doi-asserted-by":"publisher","first-page":"543","DOI":"10.1016\/0895-4356(90)90158-L","article-title":"High agreement but low kappa: I. the problems of two paradoxes","volume":"43","author":"Feinstein","year":"1990","journal-title":"Journal of Clinical Epidemiology"},{"issue":"3","key":"2026042411193802700_bib23","doi-asserted-by":"publisher","first-page":"1097","DOI":"10.1162\/coli_a_00524","article-title":"Bias and fairness in large language models: A survey","volume":"50","author":"Gallegos","year":"2024","journal-title":"Computational Linguistics"},{"key":"2026042411193802700_bib24","article-title":"Reviewer2: Optimizing review generation through prompt generation","author":"Gao","year":"2024","journal-title":"arXiv preprint arXiv:2402.10886v2"},{"key":"2026042411193802700_bib25","article-title":"Faithful explanations of black-box NLP models using LLM-generated counterfactuals","volume-title":"The Twelfth International Conference on Learning Representations, ICLR 2024, Vienna, Austria, May 7\u201311, 2024","author":"Gat","year":"2024"},{"key":"2026042411193802700_bib26","article-title":"Explaining classifiers with causal concept effect (CaCE)","author":"Goyal","year":"2019","journal-title":"arXiv preprint arXiv: 1907.07165v2"},{"key":"2026042411193802700_bib27","article-title":"RULER: What\u2019s the real context size of your long-context language models?","volume-title":"First Conference on Language Modeling","author":"Hsieh","year":"2024"},{"issue":"5","key":"2026042411193802700_bib28","doi-asserted-by":"publisher","first-page":"2550","DOI":"10.1109\/TKDE.2024.3513320","article-title":"From pixels to insights: A survey on automatic chart understanding in the era of large foundation models","volume":"37","author":"Huang","year":"2025","journal-title":"IEEE Transactions on Knowledge and Data Engineering"},{"key":"2026042411193802700_bib29","doi-asserted-by":"publisher","first-page":"550","DOI":"10.18653\/v1\/2025.naacl-demo.44","article-title":"OpenReviewer: A specialized large language model for generating critical scientific paper reviews","volume-title":"Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (System Demonstrations)","author":"Idahl","year":"2025"},{"issue":"21","key":"2026042411193802700_bib30","doi-asserted-by":"publisher","first-page":"2786","DOI":"10.1001\/jama.287.21.2786","article-title":"Measuring the quality of editorial peer review","volume":"287","author":"Jefferson","year":"2002","journal-title":"JAMA"},{"issue":"12","key":"2026042411193802700_bib31","doi-asserted-by":"publisher","DOI":"10.1145\/3571730","article-title":"Survey of hallucination in natural language generation","volume":"55","author":"Ji","year":"2023","journal-title":"ACM Computing Surveys"},{"key":"2026042411193802700_bib32","article-title":"ReviewEval: An evaluation framework for AI-generated reviews","author":"Kirtani","year":"2025","journal-title":"arXiv preprint arXiv:2502.11736v3"},{"key":"2026042411193802700_bib33","article-title":"What can natural language processing do for peer review?","author":"Kuznetsov","year":"2024","journal-title":"arXiv preprint arXiv:2405.06563v1"},{"key":"2026042411193802700_bib34","article-title":"Aspect-guided multi-level perturbation analysis of large language models in automated peer review","author":"Li","year":"2025","journal-title":"arXiv preprint arXiv:2502.12510v1"},{"key":"2026042411193802700_bib35","doi-asserted-by":"publisher","first-page":"13201","DOI":"10.63317\/3xeiq2g96jjy","article-title":"Prompting large language models for counterfactual generation: An empirical study","volume-title":"Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation, LREC\/COLING 2024, 20\u201325 May, 2024, Torino, Italy","author":"Li","year":"2024"},{"key":"2026042411193802700_bib36","first-page":"29575","article-title":"Monitoring AI-modified content at scale: A case study on the impact of ChatGPT on AI conference peer reviews","volume-title":"Proceedings of the 41st International Conference on Machine Learning","author":"Liang","year":"2024"},{"issue":"8","key":"2026042411193802700_bib37","doi-asserted-by":"publisher","first-page":"AIoa2400196","DOI":"10.1056\/AIoa2400196","article-title":"Can large language models provide useful feedback on research papers? A large-scale empirical analysis","volume":"1","author":"Liang","year":"2024","journal-title":"NEJM AI"},{"key":"2026042411193802700_bib38","doi-asserted-by":"publisher","first-page":"6626","DOI":"10.18653\/v1\/2024.acl-long.358","article-title":"Advancing large language models to capture varied speaking styles and respond properly in spoken conversations","volume-title":"Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)","author":"Lin","year":"2024"},{"issue":"404","key":"2026042411193802700_bib39","doi-asserted-by":"publisher","first-page":"1014","DOI":"10.1080\/01621459.1988.10478693","article-title":"Newton\u2014Raphson and EM algorithms for linear mixed-effects models for repeated-measures data","volume":"83","author":"Lindstrom","year":"1988","journal-title":"Journal of the American Statistical Association"},{"key":"2026042411193802700_bib40","article-title":"ReviewerGPT? An exploratory study on using large language models for paper reviewing","author":"Liu","year":"2023","journal-title":"arXiv preprint arXiv:2306.00622v1"},{"key":"2026042411193802700_bib41","article-title":"AAAR-1.0: Assessing AI\u2019s potential to assist research","volume-title":"Forty-second International Conference on Machine Learning","author":"Lou","year":"2025"},{"key":"2026042411193802700_bib42","doi-asserted-by":"publisher","first-page":"6145","DOI":"10.18653\/v1\/2025.findings-emnlp.326","article-title":"Identifying aspects in peer reviews","volume-title":"Findings of the Association for Computational Linguistics: EMNLP 2025","author":"Sheng","year":"2025"},{"key":"2026042411193802700_bib43","doi-asserted-by":"publisher","first-page":"46534","DOI":"10.52202\/075280-2019","article-title":"Self-refine: Iterative refinement with self-feedback","volume-title":"Advances in Neural Information Processing Systems","author":"Madaan","year":"2023"},{"issue":"2","key":"2026042411193802700_bib44","first-page":"26","article-title":"Is peer review broken? Submissions are up, reviewers are overtaxed, and authors are lodging complaint after complaint about the process at top-tier journals. What\u2019s wrong with peer review?","volume":"20","author":"McCook","year":"2006","journal-title":"The Scientist"},{"key":"2026042411193802700_bib45","doi-asserted-by":"publisher","DOI":"10.21105\/joss.00786","volume-title":"Interpretable Machine Learning","author":"Molnar","year":"2025","edition":"3rd"},{"key":"2026042411193802700_bib46","doi-asserted-by":"publisher","first-page":"6556","DOI":"10.18653\/v1\/2024.acl-long.354","article-title":"A causal approach for counterfactual reasoning in narratives","volume-title":"Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)","author":"Feiteng","year":"2024"},{"key":"2026042411193802700_bib47","article-title":"GPT-4 technical report","author":"OpenAI, Josh Achiam,","year":"2024","journal-title":"arXiv preprint arXiv:2303.08774v6"},{"key":"2026042411193802700_bib48","volume-title":"Causality: Models, Reasoning and Inference","author":"Pearl","year":"2000","edition":"2"},{"key":"2026042411193802700_bib49","volume-title":"Prompt Engineering for Generative AI","author":"Phoenix","year":"2024"},{"key":"2026042411193802700_bib50","doi-asserted-by":"publisher","first-page":"1184","DOI":"10.18653\/v1\/2023.findings-emnlp.84","article-title":"Can ChatGPT assess human personalities? A general evaluation framework","volume-title":"Findings of the Association for Computational Linguistics: EMNLP 2023","author":"Rao","year":"2023"},{"key":"2026042411193802700_bib51","article-title":"A systematic survey of prompt engineering in large language models: Techniques and applications","author":"Sahoo","year":"2024","journal-title":"arXiv preprint arXiv:2402.07927v2"},{"key":"2026042411193802700_bib52","doi-asserted-by":"publisher","first-page":"10776","DOI":"10.18653\/v1\/2023.findings-emnlp.722","article-title":"NLP evaluation in trouble: On the need to measure LLM data contamination for each benchmark","volume-title":"Findings of the Association for Computational Linguistics: EMNLP 2023","author":"Sainz","year":"2023"},{"key":"2026042411193802700_bib53","article-title":"Quantifying language models\u2019 sensitivity to spurious features in prompt design or: How I learned to start worrying about prompt formatting","volume-title":"The Twelfth International Conference on Learning Representations","author":"Sclar","year":"2024"},{"key":"2026042411193802700_bib54","article-title":"Avoiding a tragedy of the commons in the peer review process","author":"Sculley","year":"2018","journal-title":"arXiv preprint arXiv:1901.06246v1"},{"key":"2026042411193802700_bib55","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2025.emnlp-main.1805","article-title":"Mind the blind spots: A focus-level evaluation framework for LLM reviews","author":"Shin","year":"2025","journal-title":"arXiv preprint arXiv:2502.17086v4"},{"key":"2026042411193802700_bib56","article-title":"When AI co-scientists fail: SPOT-a benchmark for automated verification of scientific research","author":"Son","year":"2025","journal-title":"arXiv preprint arXiv:2505.11855v1"},{"key":"2026042411193802700_bib57","doi-asserted-by":"publisher","first-page":"1493","DOI":"10.3115\/1699648.1699696","article-title":"Towards domain-independent argumentative zoning: Evidence from chemistry and computational linguistics","volume-title":"Proceedings of the 2009 Conference on Empirical Methods in Natural Language Processing","author":"Teufel","year":"2009"},{"key":"2026042411193802700_bib58","article-title":"Llama 2: Open foundation and fine-tuned chat models","author":"Touvron","year":"2023","journal-title":"CoRR"},{"key":"2026042411193802700_bib59","article-title":"AI-driven review systems: Evaluating LLMs in scalable and bias-aware academic reviews","author":"Tyser","year":"2024","journal-title":"arXiv preprint arXiv:2408.10365v1"},{"issue":"3","key":"2026042411193802700_bib60","doi-asserted-by":"publisher","first-page":"334","DOI":"10.1002\/leap.1544","article-title":"How to improve scientific peer review: Four schools of thought","volume":"36","author":"Waltman","year":"2023","journal-title":"Learned Publishing"},{"key":"2026042411193802700_bib61","doi-asserted-by":"publisher","first-page":"4798","DOI":"10.18653\/v1\/2024.findings-emnlp.276","article-title":"A survey on natural language counterfactual generation","volume-title":"Findings of the Association for Computational Linguistics: EMNLP 2024","author":"Wang","year":"2024"},{"key":"2026042411193802700_bib62","first-page":"6522","article-title":"Beyond what if: Advancing counterfactual text generation with structural causal modeling","volume-title":"Proceedings of the Thirty-Third International Joint Conference on Artificial Intelligence","author":"Wang","year":"2024"},{"key":"2026042411193802700_bib63","article-title":"CycleResearcher: Improving automated research via automated review","volume-title":"The Thirteenth International Conference on Learning Representations","author":"Weng","year":"2025"},{"key":"2026042411193802700_bib64","doi-asserted-by":"publisher","first-page":"6707","DOI":"10.18653\/v1\/2021.acl-long.523","article-title":"Polyjuice: Generating counterfactuals for explaining, evaluating, and improving models","volume-title":"Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers)","author":"Tongshuang","year":"2021"},{"key":"2026042411193802700_bib65","doi-asserted-by":"publisher","first-page":"6175","DOI":"10.18653\/v1\/2024.findings-acl.369","article-title":"Understanding fine-grained distortions in reports of scientific findings","volume-title":"Findings of the Association for Computational Linguistics: ACL 2024","author":"Wuehrl","year":"2024"},{"key":"2026042411193802700_bib66","article-title":"Benchmarking LLMs\u2019 judgments with no gold standard","volume-title":"The Thirteenth International Conference on Learning Representations, ICLR 2025, Singapore, April 24\u201328, 2025","author":"Shengwei","year":"2025"},{"key":"2026042411193802700_bib67","doi-asserted-by":"publisher","first-page":"10164","DOI":"10.18653\/v1\/2024.findings-emnlp.595","article-title":"Automated peer reviewing in paper SEA: Standardization, evaluation, and analysis","volume-title":"Findings of the Association for Computational Linguistics: EMNLP 2024","author":"Jianxiang","year":"2024"},{"key":"2026042411193802700_bib68","doi-asserted-by":"publisher","first-page":"171","DOI":"10.1613\/jair.1.12862","article-title":"Can we automate scientific reviewing?","volume":"75","author":"Yuan","year":"2022","journal-title":"Journal of Artificial Intelligence Research"},{"key":"2026042411193802700_bib69","article-title":"Reviewing scientific papers for critical problems with reasoning LLMs: Baseline approaches and automatic evaluation","author":"Zhang","year":"2025","journal-title":"arXiv preprint arXiv:2505.23824v2"},{"key":"2026042411193802700_bib70","doi-asserted-by":"publisher","first-page":"46595","DOI":"10.52202\/075280-2020","article-title":"Judging LLM-as-a-judge with MT-bench and chatbot arena","volume-title":"Advances in Neural Information Processing Systems","author":"Zheng","year":"2023"},{"key":"2026042411193802700_bib71","article-title":"Large language models are human-level prompt engineers","volume-title":"The Eleventh International Conference on Learning Representations","author":"Zhou","year":"2023"},{"key":"2026042411193802700_bib72","article-title":"Deepreview: Improving LLM-based paper review with human-like deep thinking process","author":"Zhu","year":"2025","journal-title":"arXiv preprint arXiv: 2503.08569v1"}],"container-title":["Transactions of the Association for Computational Linguistics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/direct.mit.edu\/tacl\/article-pdf\/doi\/10.1162\/TACl.a.642\/2597095\/tacl.a.642.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/direct.mit.edu\/tacl\/article-pdf\/doi\/10.1162\/TACl.a.642\/2597095\/tacl.a.642.pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,4,24]],"date-time":"2026-04-24T15:19:48Z","timestamp":1777043988000},"score":1,"resource":{"primary":{"URL":"https:\/\/direct.mit.edu\/tacl\/article\/doi\/10.1162\/TACl.a.642\/136338\/Automatic-Reviewers-Fail-to-Detect-Faulty"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026]]},"references-count":72,"URL":"https:\/\/doi.org\/10.1162\/tacl.a.642","relation":{},"ISSN":["2307-387X"],"issn-type":[{"value":"2307-387X","type":"electronic"}],"subject":[],"published-other":{"date-parts":[[2026]]},"published":{"date-parts":[[2026]]}}}