{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,27]],"date-time":"2026-07-27T18:50:33Z","timestamp":1785178233452,"version":"3.55.0"},"reference-count":122,"publisher":"MIT Press","license":[{"start":{"date-parts":[[2024,11,7]],"date-time":"2024-11-07T00:00:00Z","timestamp":1730937600000},"content-version":"vor","delay-in-days":311,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":["direct.mit.edu"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2024,11,4]]},"abstract":"<jats:title>Abstract<\/jats:title>\n               <jats:p>Self-correction is an approach to improving responses from large language models (LLMs) by refining the responses using LLMs during inference. Prior work has proposed various self-correction frameworks using different sources of feedback, including self-evaluation and external feedback. However, there is still no consensus on the question of when LLMs can correct their own mistakes, as recent studies also report negative results. In this work, we critically survey broad papers and discuss the conditions required for successful self-correction. We first find that prior studies often do not define their research questions in detail and involve impractical frameworks or unfair evaluations that over-evaluate self-correction. To tackle these issues, we categorize research questions in self-correction research and provide a checklist for designing appropriate experiments. Our critical survey based on the newly categorized research questions shows that (1) no prior work demonstrates successful self-correction with feedback from prompted LLMs, except for studies in tasks that are exceptionally suited for self-correction, (2) self-correction works well in tasks that can use reliable external feedback, and (3) large-scale fine-tuning enables self-correction.<\/jats:p>","DOI":"10.1162\/tacl_a_00713","type":"journal-article","created":{"date-parts":[[2024,11,7]],"date-time":"2024-11-07T20:18:42Z","timestamp":1731010722000},"page":"1417-1440","update-policy":"https:\/\/doi.org\/10.1162\/mitpressjournals.corrections.policy","source":"Crossref","is-referenced-by-count":61,"title":["When Can LLMs <i>Actually<\/i> Correct Their Own Mistakes? A Critical Survey of Self-Correction of LLMs"],"prefix":"10.1162","volume":"12","author":[{"given":"Ryo","family":"Kamoi","sequence":"first","affiliation":[{"name":"Penn State University, USA. ryokamoi@psu.edu"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Yusen","family":"Zhang","sequence":"additional","affiliation":[{"name":"Penn State University, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Nan","family":"Zhang","sequence":"additional","affiliation":[{"name":"Penn State University, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Jiawei","family":"Han","sequence":"additional","affiliation":[{"name":"University of Illinois Urbana-Champaign, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Rui","family":"Zhang","sequence":"additional","affiliation":[{"name":"Penn State University, USA. rmz5227@psu.edu"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"281","published-online":{"date-parts":[[2024,11,4]]},"reference":[{"key":"2024110720183855300_bib1","doi-asserted-by":"publisher","first-page":"7716","DOI":"10.18653\/v1\/2023.acl-long.427","article-title":"RL4F: Generating natural language feedback with reinforcement learning for repairing model outputs","volume-title":"Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)","author":"Akyurek","year":"2023"},{"key":"2024110720183855300_bib2","article-title":"Self-RAG: Learning to retrieve, generate, and critique through self-reflection","volume-title":"The Twelfth International Conference on Learning Representations","author":"Asai","year":"2024"},{"issue":"OOPSLA","key":"2024110720183855300_bib3","doi-asserted-by":"publisher","DOI":"10.1145\/3360585","article-title":"Getafix: Learning to fix bugs automatically","volume":"3","author":"Bader","year":"2019","journal-title":"Proceedings of the ACM on Programming Languages"},{"key":"2024110720183855300_bib4","article-title":"Constitutional AI: Harmlessness from AI feedback","author":"Bai","year":"2022","journal-title":"arXiv preprint arXiv:2212.08073"},{"key":"2024110720183855300_bib5","doi-asserted-by":"publisher","first-page":"5454","DOI":"10.18653\/v1\/2020.acl-main.485","article-title":"Language (technology) is power: A critical survey of \u201cbias\u201d in NLP","volume-title":"Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics","author":"Lin Blodgett","year":"2020"},{"key":"2024110720183855300_bib6","doi-asserted-by":"publisher","first-page":"6251","DOI":"10.18653\/v1\/2020.emnlp-main.506","article-title":"Factual error correction for abstractive summarization models","volume-title":"Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP)","author":"Cao","year":"2020"},{"key":"2024110720183855300_bib7","article-title":"Chateval: Towards better LLM-based evaluators through multi-agent debate","volume-title":"The Twelfth International Conference on Learning Representations","author":"Chan","year":"2024"},{"key":"2024110720183855300_bib8","article-title":"A new era in software security: Towards self-healing software via large language models and formal verification","author":"Charalambous","year":"2023","journal-title":"arXiv preprint arXiv:2305.14752"},{"key":"2024110720183855300_bib9","article-title":"Learning from natural language feedback","author":"Chen","year":"2024","journal-title":"Transactions on Machine Learning Research"},{"key":"2024110720183855300_bib10","article-title":"Codet: Code generation with generated tests","volume-title":"The Eleventh International Conference on Learning Representations","author":"Chen","year":"2023"},{"key":"2024110720183855300_bib11","article-title":"Can LLM-generated misinformation be detected?","volume-title":"The Twelfth International Conference on Learning Representations","author":"Chen","year":"2024"},{"key":"2024110720183855300_bib12","article-title":"Reconcile: Round-table conference improves reasoning via consensus among diverse LLMs","author":"Chen","year":"2024","journal-title":"arXiv preprint arXiv: 2309.13007"},{"key":"2024110720183855300_bib13","article-title":"Iterative translation refinement with large language models","author":"Chen","year":"2023","journal-title":"arXiv preprint arXiv:2306.03856"},{"key":"2024110720183855300_bib14","article-title":"Universal self-consistency for large language models","volume-title":"ICML 2024 Workshop on In-Context Learning","author":"Chen","year":"2024"},{"key":"2024110720183855300_bib15","article-title":"Teaching large language models to self-debug","volume-title":"The Twelfth International Conference on Learning Representations","author":"Chen","year":"2024"},{"issue":"9","key":"2024110720183855300_bib16","doi-asserted-by":"publisher","first-page":"1943","DOI":"10.1109\/TSE.2019.2940179","article-title":"Sequencer: Sequence-to-sequence learning for end-to-end program repair","volume":"47","author":"Chen","year":"2021","journal-title":"IEEE Transactions on Software Engineering"},{"key":"2024110720183855300_bib17","article-title":"Factool: Factuality detection in generative AI \u2013 a tool augmented framework for multi-task and multi-domain scenarios","author":"Chern","year":"2023","journal-title":"arXiv preprint arXiv:2307.13528"},{"key":"2024110720183855300_bib18","doi-asserted-by":"publisher","first-page":"15607","DOI":"10.18653\/v1\/2023.acl-long.870","article-title":"Can large language models be an alternative to human evaluations?","volume-title":"Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)","author":"Chiang","year":"2023"},{"key":"2024110720183855300_bib19","article-title":"Training verifiers to solve math word problems","author":"Cobbe","year":"2021","journal-title":"arXiv preprint arXiv:2110.14168"},{"key":"2024110720183855300_bib20","doi-asserted-by":"publisher","first-page":"12621","DOI":"10.18653\/v1\/2023.emnlp-main.778","article-title":"LM vs LM: Detecting factual errors via cross examination","volume-title":"Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing","author":"Cohen","year":"2023"},{"key":"2024110720183855300_bib21","article-title":"Faithful reasoning using large language models","author":"Creswell","year":"2022","journal-title":"arXiv preprint arXiv:2208.14271"},{"key":"2024110720183855300_bib22","article-title":"Chain-of-verification reduces hallucination in large language models","author":"Dhuliawala","year":"2023","journal-title":"arXiv preprint arXiv:2309.11495"},{"key":"2024110720183855300_bib23","article-title":"Improving factuality and reasoning in language models through multiagent debate","author":"Yilun","year":"2023","journal-title":"arXiv preprint arXiv:2305.14325"},{"key":"2024110720183855300_bib24","doi-asserted-by":"publisher","first-page":"5055","DOI":"10.18653\/v1\/2020.acl-main.454","article-title":"FEQA: A question answering evaluation framework for faithfulness assessment in abstractive summarization","volume-title":"Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics","author":"Durmus","year":"2020"},{"key":"2024110720183855300_bib25","article-title":"Halo: Estimation and reduction of hallucinations in open-source weak large language models","author":"Elaraby","year":"2023","journal-title":"arXiv preprint arXiv:2308 .11764"},{"key":"2024110720183855300_bib26","doi-asserted-by":"publisher","first-page":"11737","DOI":"10.18653\/v1\/2023.acl-long.656","article-title":"From pretraining data to language models to downstream tasks: Tracking the trails of political biases leading to unfair NLP models","volume-title":"Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)","author":"Feng","year":"2023"},{"key":"2024110720183855300_bib27","doi-asserted-by":"publisher","first-page":"1229","DOI":"10.1145\/3611643.3616243","article-title":"Baldur: Whole-proof generation and repair with large language models","volume-title":"Proceedings of the 31st ACM Joint European Software Engineering Conference and Symposium on the Foundations of Software Engineering","author":"First","year":"2023"},{"key":"2024110720183855300_bib28","first-page":"6556","article-title":"GPTScore: Evaluate as you desire","volume-title":"Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers)","author":"Jinlan","year":"2024"},{"key":"2024110720183855300_bib29","doi-asserted-by":"publisher","first-page":"16477","DOI":"10.18653\/v1\/2023.acl-long.910","article-title":"RARR: Researching and revising what language models say, using language models","volume-title":"Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)","author":"Gao","year":"2023"},{"key":"2024110720183855300_bib30","doi-asserted-by":"publisher","first-page":"1173","DOI":"10.18653\/v1\/2023.emnlp-main.75","article-title":"From wrong to right: A recursive approach towards vision-language explanation","volume-title":"Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing","author":"Ge","year":"2023"},{"key":"2024110720183855300_bib31","article-title":"Self-verification improves few-shot clinical information extraction","volume-title":"ICML 3rd Workshop on Interpretable Machine Learning in Healthcare (IMLH)","author":"Gero","year":"2023"},{"key":"2024110720183855300_bib32","article-title":"CRITIC: Large language models can self-correct with tool-interactive critiquing","volume-title":"The Twelfth International Conference on Learning Representations","author":"Gou","year":"2024"},{"key":"2024110720183855300_bib33","article-title":"Reinforced self-training (rest) for language modeling","author":"Gulcehre","year":"2023","journal-title":"arXiv preprint arXiv:2308.08998"},{"issue":"1","key":"2024110720183855300_bib34","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v31i1.10742","article-title":"Deepfix: Fixing common c language errors by deep learning","volume":"31","author":"Gupta","year":"2017","journal-title":"Proceedings of the AAAI Conference on Artificial Intelligence"},{"issue":"16","key":"2024110720183855300_bib35","doi-asserted-by":"publisher","first-page":"18162","DOI":"10.1609\/aaai.v38i16.29774","article-title":"Small language model can self-correct","volume":"38","author":"Han","year":"2024","journal-title":"Proceedings of the AAAI Conference on Artificial Intelligence"},{"key":"2024110720183855300_bib36","doi-asserted-by":"publisher","first-page":"8154","DOI":"10.18653\/v1\/2023.emnlp-main.507","article-title":"Reasoning with language model is planning with world model","volume-title":"Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing","author":"Hao","year":"2023"},{"key":"2024110720183855300_bib37","doi-asserted-by":"publisher","first-page":"900","DOI":"10.18653\/v1\/2024.naacl-long.52","article-title":"A closer look at the self-verification abilities of large language models in logical reasoning","volume-title":"Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers)","author":"Hong","year":"2024"},{"key":"2024110720183855300_bib38","doi-asserted-by":"publisher","first-page":"1051","DOI":"10.18653\/v1\/2023.emnlp-main.67","article-title":"Large language models can self-improve","volume-title":"Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing","author":"Huang","year":"2023"},{"key":"2024110720183855300_bib39","article-title":"Large language models cannot self-correct reasoning yet","volume-title":"The Twelfth International Conference on Learning Representations","author":"Huang","year":"2024"},{"key":"2024110720183855300_bib40","doi-asserted-by":"publisher","first-page":"730","DOI":"10.18653\/v1\/2024.findings-acl.41","article-title":"Do LVLMs understand charts? Analyzing and correcting factual errors in chart captioning","volume-title":"Findings of the Association for Computational Linguistics ACL 2024","author":"Huang","year":"2024"},{"key":"2024110720183855300_bib41","doi-asserted-by":"publisher","first-page":"3670","DOI":"10.18653\/v1\/2022.naacl-main.269","article-title":"FRUIT: Faithfully reflecting updated information in text","volume-title":"Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies","author":"Iv","year":"2022"},{"key":"2024110720183855300_bib42","article-title":"Self-[in]correct: Llms struggle with refining self-generated responses","author":"Jiang","year":"2024","journal-title":"arXiv preprint arXiv:2404.04298"},{"key":"2024110720183855300_bib43","article-title":"Selfevolve: A code evolution framework via large language models","author":"Jiang","year":"2023","journal-title":"arXiv preprint arXiv:2306.02907"},{"key":"2024110720183855300_bib44","doi-asserted-by":"publisher","first-page":"7969","DOI":"10.18653\/v1\/2023.emnlp-main.495","article-title":"Active retrieval augmented generation","volume-title":"Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing","author":"Jiang","year":"2023"},{"key":"2024110720183855300_bib45","doi-asserted-by":"publisher","first-page":"1266","DOI":"10.18653\/v1\/2022.emnlp-main.82","article-title":"Maieutic prompting: Logically consistent reasoning with recursive explanations","volume-title":"Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing","author":"Jung","year":"2022"},{"key":"2024110720183855300_bib46","article-title":"Evaluating LLMs at detecting errors in LLM responses","author":"Kamoi","year":"2024","journal-title":"arXiv preprint arXiv:2404.03602"},{"key":"2024110720183855300_bib47","doi-asserted-by":"publisher","first-page":"4253","DOI":"10.18653\/v1\/2024.findings-naacl.265","article-title":"Guiding large language models to post-edit machine translation with error annotations","volume-title":"Findings of the Association for Computational Linguistics: NAACL 2024","author":"Ki","year":"2024"},{"key":"2024110720183855300_bib48","first-page":"39648","article-title":"Language models can solve computer tasks","volume-title":"Advances in Neural Information Processing Systems","author":"Kim","year":"2023"},{"key":"2024110720183855300_bib49","first-page":"21314","article-title":"Coderl: Mastering code generation through pretrained models and deep reinforcement learning","volume-title":"Advances in Neural Information Processing Systems","author":"Le","year":"2022"},{"key":"2024110720183855300_bib50","doi-asserted-by":"crossref","first-page":"391","DOI":"10.18653\/v1\/2024.naacl-long.23","article-title":"Volcano: Mitigating multimodal hallucination through self-feedback guided revision","volume-title":"Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers)","author":"Lee","year":"2024"},{"key":"2024110720183855300_bib51","article-title":"Confidence matters: Revisiting intrinsic self-correction capabilities of large language models","author":"Li","year":"2024","journal-title":"arXiv preprint arXiv:2402 .12563"},{"key":"2024110720183855300_bib52","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2307.02762","article-title":"Prd: Peer rank and discussion improve large language model based evaluations","author":"Li","year":"2023","journal-title":"arXiv preprint arXiv:2307.02762"},{"key":"2024110720183855300_bib53","doi-asserted-by":"publisher","first-page":"3741","DOI":"10.18653\/v1\/2024.findings-naacl.237","article-title":"When hindsight is not 20\/20: Testing limits on reflective thinking in large language models","volume-title":"Findings of the Association for Computational Linguistics: NAACL 2024","author":"Li","year":"2024"},{"key":"2024110720183855300_bib54","article-title":"Encouraging divergent thinking in large language models through multi-agent debate","author":"Liang","year":"2023","journal-title":"arXiv preprint arXiv:2305.19118"},{"key":"2024110720183855300_bib55","doi-asserted-by":"publisher","first-page":"3291","DOI":"10.18653\/v1\/N19-1333","article-title":"Corpora generation for grammatical error correction","volume-title":"Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers)","author":"Lichtarge","year":"2019"},{"key":"2024110720183855300_bib56","article-title":"Let\u2019s verify step by step","volume-title":"The Twelfth International Conference on Learning Representations","author":"Lightman","year":"2024"},{"key":"2024110720183855300_bib57","article-title":"On the intrinsic self-correction capability of LLMs: Uncertainty and latent concept","author":"Liu","year":"2024","journal-title":"arXiv preprint arXiv:2406.02378"},{"key":"2024110720183855300_bib58","doi-asserted-by":"publisher","first-page":"2511","DOI":"10.18653\/v1\/2023.emnlp-main.153","article-title":"G-eval: NLG evaluation using gpt-4 with better human alignment","volume-title":"Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing","author":"Liu","year":"2023"},{"key":"2024110720183855300_bib59","first-page":"46534","article-title":"Self-refine: Iterative refinement with self-feedback","volume-title":"Advances in Neural Information Processing Systems","author":"Madaan","year":"2023"},{"key":"2024110720183855300_bib60","doi-asserted-by":"publisher","first-page":"9004","DOI":"10.18653\/v1\/2023.emnlp-main.557","article-title":"SelfCheckGPT: Zero-resource black-box hallucination detection for generative large language models","volume-title":"Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing","author":"Manakul","year":"2023"},{"key":"2024110720183855300_bib61","article-title":"Flirt: Feedback loop in-context red teaming","author":"Mehrabi","year":"2023","journal-title":"arXiv preprint arXiv: 2308.04265"},{"key":"2024110720183855300_bib62","first-page":"462","article-title":"Generating training data with language models: Towards zero-shot language understanding","volume-title":"Advances in Neural Information Processing Systems","author":"Meng","year":"2022"},{"key":"2024110720183855300_bib63","doi-asserted-by":"publisher","first-page":"925","DOI":"10.1145\/3338906.3340455","article-title":"Deepdelta: Learning to repair compilation errors","volume-title":"Proceedings of the 2019 27th ACM Joint Meeting on European Software Engineering Conference and Symposium on the Foundations of Software Engineering","author":"Mesbah","year":"2019"},{"key":"2024110720183855300_bib64","article-title":"Selfcheck: Using LLMs to zero-shot check their own step-by-step reasoning","volume-title":"The Twelfth International Conference on Learning Representations","author":"Miao","year":"2024"},{"key":"2024110720183855300_bib65","doi-asserted-by":"publisher","first-page":"12076","DOI":"10.18653\/v1\/2023.emnlp-main.741","article-title":"FActScore: Fine-grained atomic evaluation of factual precision in long form text generation","volume-title":"Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing","author":"Min","year":"2023"},{"key":"2024110720183855300_bib66","article-title":"Fine-grained hallucination detection and editing for language models","author":"Mishra","year":"2024","journal-title":"arXiv preprint arXiv:2401.06855"},{"key":"2024110720183855300_bib67","doi-asserted-by":"publisher","first-page":"6591","DOI":"10.18653\/v1\/2023.emnlp-main.407","article-title":"MAF: Multi-aspect feedback for improving reasoning in large language models","volume-title":"Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing","author":"Nathani","year":"2023"},{"key":"2024110720183855300_bib68","doi-asserted-by":"publisher","first-page":"1","DOI":"10.3115\/v1\/W14-1701","article-title":"The CoNLL-2014 shared task on grammatical error correction","volume-title":"Proceedings of the Eighteenth Conference on Computational Natural Language Learning: Shared Task","author":"Ng","year":"2014"},{"key":"2024110720183855300_bib69","first-page":"26106","article-title":"LEVER: Learning to verify language-to-code generation with execution","volume-title":"Proceedings of the 40th International Conference on Machine Learning","author":"Ni","year":"2023"},{"key":"2024110720183855300_bib70","article-title":"Is self-repair a silver bullet for code generation?","volume-title":"The Twelfth International Conference on Learning Representations","author":"Olausson","year":"2024"},{"key":"2024110720183855300_bib71","doi-asserted-by":"crossref","first-page":"3806","DOI":"10.18653\/v1\/2023.findings-emnlp.248","article-title":"Logic-LM: Empowering large language models with symbolic solvers for faithful logical reasoning","volume-title":"Findings of the Association for Computational Linguistics: EMNLP 2023","author":"Pan","year":"2023"},{"key":"2024110720183855300_bib72","doi-asserted-by":"publisher","first-page":"484","DOI":"10.1162\/tacl_a_00660","article-title":"Automatically correcting large language models: Surveying the landscape of diverse automated correction strategies","volume":"12","author":"Pan","year":"2024","journal-title":"Transactions of the Association for Computational Linguistics"},{"key":"2024110720183855300_bib73","article-title":"Language model self-improvement by reinforcement learning contemplation","volume-title":"The Twelfth International Conference on Learning Representations","author":"Pang","year":"2024"},{"key":"2024110720183855300_bib74","first-page":"1100","article-title":"REFINER: Reasoning feedback on intermediate representations","volume-title":"Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers)","author":"Paul","year":"2024"},{"key":"2024110720183855300_bib75","article-title":"Check your facts and try again: Improving large language models with external knowledge and automated feedback","author":"Peng","year":"2023","journal-title":"arXiv preprint arXiv:2302.12813"},{"key":"2024110720183855300_bib76","article-title":"Llm self defense: By self examination, llms know they are being tricked","author":"Phute","year":"2024","journal-title":"arXiv preprint arXiv:2308.07308"},{"key":"2024110720183855300_bib77","doi-asserted-by":"publisher","first-page":"7957","DOI":"10.18653\/v1\/2023.emnlp-main.494","article-title":"Automatic prompt optimization with \u201cgradient descent\u201d and beam search","volume-title":"Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing","author":"Pryzant","year":"2023"},{"key":"2024110720183855300_bib78","article-title":"Characteristics of harmful text: Towards rigorous benchmarking of language models","volume-title":"Thirty-sixth Conference on Neural Information Processing Systems Datasets and Benchmarks Track","author":"Rauh","year":"2022"},{"key":"2024110720183855300_bib79","doi-asserted-by":"publisher","first-page":"12009","DOI":"10.18653\/v1\/2023.findings-emnlp.804","article-title":"Leveraging GPT-4 for automatic translation post-editing","volume-title":"Findings of the Association for Computational Linguistics: EMNLP 2023","author":"Raunak","year":"2023"},{"key":"2024110720183855300_bib80","doi-asserted-by":"publisher","first-page":"8352","DOI":"10.18653\/v1\/2024.naacl-long.462","article-title":"Branch-solve-merge improves large language model evaluation and generation","volume-title":"Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers)","author":"Saha","year":"2024"},{"key":"2024110720183855300_bib81","article-title":"Self-critiquing models for assisting human evaluators","author":"Saunders","year":"2022","journal-title":"arXiv preprint arXiv:2206.05802"},{"key":"2024110720183855300_bib82","doi-asserted-by":"publisher","first-page":"1408","DOI":"10.1162\/tacl_a_00434","article-title":"Self-diagnosis and self-debiasing: A proposal for reducing corpus-based bias in NLP","volume":"9","author":"Schick","year":"2021","journal-title":"Transactions of the Association for Computational Linguistics"},{"key":"2024110720183855300_bib83","article-title":"PEER: A collaborative language model","volume-title":"The Eleventh International Conference on Learning Representations","author":"Schick","year":"2023"},{"key":"2024110720183855300_bib84","doi-asserted-by":"publisher","first-page":"6594","DOI":"10.18653\/v1\/2021.emnlp-main.529","article-title":"QuestEval: Summarization asks for fact-based evaluation","volume-title":"Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing","author":"Scialom","year":"2021"},{"issue":"05","key":"2024110720183855300_bib85","doi-asserted-by":"publisher","first-page":"8791","DOI":"10.1609\/aaai.v34i05.6406","article-title":"Automatic fact-guided sentence modification","volume":"34","author":"Shah","year":"2020","journal-title":"Proceedings of the AAAI Conference on Artificial Intelligence"},{"key":"2024110720183855300_bib86","doi-asserted-by":"publisher","first-page":"2269","DOI":"10.18653\/v1\/2021.findings-emnlp.195","article-title":"Generate & rank: A multi-task framework for math word problems","volume-title":"Findings of the Association for Computational Linguistics: EMNLP 2021","author":"Shen","year":"2021"},{"key":"2024110720183855300_bib87","doi-asserted-by":"publisher","first-page":"3533","DOI":"10.18653\/v1\/2022.emnlp-main.231","article-title":"Natural language to code translation with execution","volume-title":"Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing","author":"Shi","year":"2022"},{"key":"2024110720183855300_bib88","first-page":"8634","article-title":"Reflexion: Language agents with verbal reinforcement learning","volume-title":"Advances in Neural Information Processing Systems","author":"Shinn","year":"2023"},{"key":"2024110720183855300_bib89","article-title":"GPT-4 doesn\u2019t know it\u2019s wrong: An analysis of iterative prompting for reasoning problems","volume-title":"NeurIPS 2023 Foundation Models for Decision Making Workshop","author":"Stechly","year":"2023"},{"key":"2024110720183855300_bib90","article-title":"Regal: Refactoring programs to discover generalizable abstractions","author":"Stengel-Eskin","year":"2024","journal-title":"arXiv preprint arXiv:2401.16467"},{"key":"2024110720183855300_bib91","doi-asserted-by":"publisher","first-page":"2078","DOI":"10.18653\/v1\/2022.emnlp-main.134","article-title":"Entailer: Answering questions with faithful and truthful chains of reasoning","volume-title":"Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing","author":"Tafjord","year":"2022"},{"key":"2024110720183855300_bib92","doi-asserted-by":"publisher","first-page":"3298","DOI":"10.18653\/v1\/2021.acl-long.256","article-title":"Evidence-based factual error correction","volume-title":"Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers)","author":"Thorne","year":"2021"},{"key":"2024110720183855300_bib93","doi-asserted-by":"publisher","first-page":"13894","DOI":"10.18653\/v1\/2024.findings-acl.826","article-title":"LLMs cannot find reasoning errors, but can correct them given the error location","volume-title":"Findings of the Association for Computational Linguistics ACL 2024","author":"Tyen","year":"2024"},{"key":"2024110720183855300_bib94","article-title":"Solving math word problems with process- and outcome-based feedback","author":"Uesato","year":"2022","journal-title":"arXiv preprint arXiv:2211.14275"},{"key":"2024110720183855300_bib95","article-title":"Investigating the effectiveness of self-critiquing in LLMs solving planning tasks","volume-title":"NeurIPS 2023 Foundation Models for Decision Making Workshop","author":"Valmeekam","year":"2023"},{"key":"2024110720183855300_bib96","article-title":"A stitch in time saves nine: Detecting and mitigating hallucinations of LLMs by validating low-confidence generation","author":"Varshney","year":"2023","journal-title":"arXiv preprint arXiv: 2307.03987"},{"key":"2024110720183855300_bib97","doi-asserted-by":"publisher","first-page":"5008","DOI":"10.18653\/v1\/2020.acl-main.450","article-title":"Asking and answering questions to evaluate the factual consistency of summaries","volume-title":"Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics","author":"Wang","year":"2020"},{"key":"2024110720183855300_bib98","doi-asserted-by":"publisher","first-page":"6106","DOI":"10.18653\/v1\/2024.acl-long.331","article-title":"Rethinking the bounds of LLM reasoning: Are multi-agent discussions the key?","volume-title":"Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)","author":"Wang","year":"2024"},{"key":"2024110720183855300_bib99","article-title":"Self-consistency improves chain of thought reasoning in language models","volume-title":"The Eleventh International Conference on Learning Representations","author":"Wang","year":"2023"},{"key":"2024110720183855300_bib100","article-title":"A theoretical understanding of self-correction through in-context alignment","volume-title":"ICML 2024 Workshop on In-Context Learning","author":"Wang","year":"2024"},{"key":"2024110720183855300_bib101","article-title":"Generating sequences by learning to self-correct","volume-title":"The Eleventh International Conference on Learning Representations","author":"Welleck","year":"2023"},{"key":"2024110720183855300_bib102","doi-asserted-by":"publisher","first-page":"2550","DOI":"10.18653\/v1\/2023.findings-emnlp.167","article-title":"Large language models are better reasoners with self-verification","volume-title":"Findings of the Association for Computational Linguistics: EMNLP 2023","author":"Weng","year":"2023"},{"key":"2024110720183855300_bib103","article-title":"Large language models can self-correct with minimal effort","author":"Zhenyu","year":"2024","journal-title":"arXiv preprint arXiv: 2405.14092"},{"key":"2024110720183855300_bib104","article-title":"Self-evaluation guided beam search for reasoning","volume-title":"Thirty-seventh Conference on Neural Information Processing Systems","author":"Xie","year":"2023"},{"key":"2024110720183855300_bib105","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2023.emnlp-main.365","article-title":"INSTRUCTSCORE: Towards explainable text generation evaluation with automatic feedback","volume-title":"The 2023 Conference on Empirical Methods in Natural Language Processing","author":"Wenda","year":"2023"},{"key":"2024110720183855300_bib106","article-title":"Large language models as optimizers","volume-title":"The Twelfth International Conference on Learning Representations","author":"Yang","year":"2024"},{"key":"2024110720183855300_bib107","doi-asserted-by":"publisher","first-page":"89","DOI":"10.18653\/v1\/2022.emnlp-main.7","article-title":"Generating natural language proofs with verifier-guided search","volume-title":"Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing","author":"Yang","year":"2022"},{"key":"2024110720183855300_bib108","doi-asserted-by":"publisher","first-page":"4393","DOI":"10.18653\/v1\/2022.emnlp-main.296","article-title":"Re3: Generating longer stories with recursive reprompting and revision","volume-title":"Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing","author":"Yang","year":"2022"},{"key":"2024110720183855300_bib109","first-page":"11809","article-title":"Tree of thoughts: Deliberate problem solving with large language models","volume-title":"Advances in Neural Information Processing Systems","author":"Yao","year":"2023"},{"key":"2024110720183855300_bib110","first-page":"10799","article-title":"Graph-based, self-supervised program repair from diagnostic feedback","volume-title":"Proceedings of the 37th International Conference on Machine Learning","author":"Yasunaga","year":"2020"},{"key":"2024110720183855300_bib111","first-page":"11941","article-title":"Break-it-fix-it: Unsupervised learning for program repair","volume-title":"Proceedings of the 38th International Conference on Machine Learning","author":"Yasunaga","year":"2021"},{"key":"2024110720183855300_bib112","article-title":"Selfee: Iterative self-revising LLM empowered by self-feedback generation","author":"Ye","year":"2023"},{"key":"2024110720183855300_bib113","article-title":"Woodpecker: Hallucination correction for multimodal large language models","author":"Yin","year":"2023","journal-title":"arXiv preprint arXiv:2310.16045"},{"key":"2024110720183855300_bib114","article-title":"Improving language models via plug-and-play retrieval feedback","author":"Wenhao","year":"2023","journal-title":"arXiv preprint arXiv: 2305.14002"},{"key":"2024110720183855300_bib115","first-page":"15476","article-title":"Star: Bootstrapping reasoning with reasoning","volume-title":"Advances in Neural Information Processing Systems","author":"Zelikman","year":"2022"},{"key":"2024110720183855300_bib116","article-title":"Exploring collaboration mechanisms for llm agents: A social psychology view","author":"Zhang","year":"2023","journal-title":"arXiv preprint arXiv:2310.02124"},{"key":"2024110720183855300_bib117","doi-asserted-by":"publisher","first-page":"769","DOI":"10.18653\/v1\/2023.acl-long.45","article-title":"Self-edit: Fault-aware code editor for code generation","volume-title":"Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)","author":"Zhang","year":"2023"},{"key":"2024110720183855300_bib118","article-title":"How language model hallucinations can snowball","author":"Zhang","year":"2023","journal-title":"arXiv preprint arXiv:2305.13534"},{"key":"2024110720183855300_bib119","first-page":"41832","article-title":"Coder reviewer reranking for code generation","volume-title":"Proceedings of the 40th International Conference on Machine Learning","author":"Zhang","year":"2023"},{"key":"2024110720183855300_bib120","doi-asserted-by":"publisher","first-page":"15637","DOI":"10.18653\/v1\/2024.findings-acl.924","article-title":"Small language models need strong verifiers to self-correct reasoning","volume-title":"Findings of the Association for Computational Linguistics ACL 2024","author":"Zhang","year":"2024"},{"key":"2024110720183855300_bib121","doi-asserted-by":"publisher","first-page":"5823","DOI":"10.18653\/v1\/2023.acl-long.320","article-title":"Verify-and-edit: A knowledge-enhanced chain-of-thought framework","volume-title":"Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)","author":"Zhao","year":"2023"},{"key":"2024110720183855300_bib122","article-title":"Analyzing and mitigating object hallucination in large vision-language models","volume-title":"The Twelfth International Conference on Learning Representations","author":"Zhou","year":"2024"}],"container-title":["Transactions of the Association for Computational Linguistics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/direct.mit.edu\/tacl\/article-pdf\/doi\/10.1162\/tacl_a_00713\/2478635\/tacl_a_00713.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/direct.mit.edu\/tacl\/article-pdf\/doi\/10.1162\/tacl_a_00713\/2478635\/tacl_a_00713.pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2024,11,7]],"date-time":"2024-11-07T20:19:04Z","timestamp":1731010744000},"score":1,"resource":{"primary":{"URL":"https:\/\/direct.mit.edu\/tacl\/article\/doi\/10.1162\/tacl_a_00713\/125177\/When-Can-LLMs-Actually-Correct-Their-Own-Mistakes"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024]]},"references-count":122,"URL":"https:\/\/doi.org\/10.1162\/tacl_a_00713","relation":{},"ISSN":["2307-387X"],"issn-type":[{"value":"2307-387X","type":"electronic"}],"subject":[],"published-other":{"date-parts":[[2024]]},"published":{"date-parts":[[2024]]}}}