{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,4,6]],"date-time":"2026-04-06T18:08:27Z","timestamp":1775498907365,"version":"3.50.1"},"reference-count":44,"publisher":"MIT Press","license":[{"start":{"date-parts":[[2026,4,6]],"date-time":"2026-04-06T00:00:00Z","timestamp":1775433600000},"content-version":"vor","delay-in-days":95,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":["direct.mit.edu"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2026,4,1]]},"abstract":"<jats:title>Abstract<\/jats:title>\n                  <jats:p>Self-consistency (SC) improves the performance of large language models (LLMs) across various tasks and domains that involve short content. However, does this support its effectiveness for long-context problems? We challenge the assumption that SC\u2019s benefits generalize to long-context settings, where LLMs often struggle with position bias\u2014the systematic over-reliance on specific context regions\u2014which hinders their ability to utilize information effectively from all parts of their context. Through comprehensive experimentation with varying state-of-the-art models, tasks, and SC formulations, we find that SC not only fails to improve but actively degrades performance on long-context tasks. This degradation is driven by persistent position bias, which worsens with longer context lengths and smaller model sizes but remains invariant to prompt format or task type. Unlike short-context tasks, where SC diversifies reasoning paths, long-context SC amplifies positional errors. These comprehensive results provide valuable insight into the limitations of current LLMs in long-context understanding and highlight the need for more sophisticated approaches.<\/jats:p>","DOI":"10.1162\/tacl.a.625","type":"journal-article","created":{"date-parts":[[2026,4,6]],"date-time":"2026-04-06T16:49:59Z","timestamp":1775494199000},"page":"292-317","update-policy":"https:\/\/doi.org\/10.1162\/mitpressjournals.corrections.policy","source":"Crossref","is-referenced-by-count":0,"title":["<i>Self-Consistency Falls Short!<\/i>\n                    The Adverse Effects of Positional Bias on Long-Context Problems"],"prefix":"10.1162","volume":"14","author":[{"given":"Adam","family":"Byerly","sequence":"first","affiliation":[{"name":"Johns Hopkins University, USA. abyerly2@jhu.edu"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Daniel","family":"Khashabi","sequence":"additional","affiliation":[{"name":"Johns Hopkins University, USA. danielk@cs.jhu.edu"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"281","published-online":{"date-parts":[[2026,4,1]]},"reference":[{"key":"2026040612495313400_bib1","doi-asserted-by":"publisher","first-page":"3119","DOI":"10.18653\/v1\/2024.acl-long.172","article-title":"LongBench: A bilingual, multitask benchmark for long context understanding","volume-title":"Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)","author":"Bai","year":"2024"},{"key":"2026040612495313400_bib2","first-page":"1877","article-title":"Language models are few-shot learners","volume-title":"Advances in Neural Information Processing Systems","author":"Brown","year":"2020"},{"key":"2026040612495313400_bib3","article-title":"Universal self-consistency for large language models","volume-title":"ICML 2024 Workshop on In-Context Learning","author":"Chen","year":"2024"},{"key":"2026040612495313400_bib4","doi-asserted-by":"publisher","first-page":"4599","DOI":"10.18653\/v1\/2021.naacl-main.365","article-title":"A dataset of information-seeking questions and answers anchored in research papers","volume-title":"Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies","author":"Dasigi","year":"2021"},{"key":"2026040612495313400_bib5","doi-asserted-by":"publisher","first-page":"391","DOI":"10.1162\/tacl_a_00373","article-title":"Summeval: Re-evaluating summarization evaluation","volume":"9","author":"Fabbri","year":"2021","journal-title":"Transactions of the Association for Computational Linguistics"},{"key":"2026040612495313400_bib6","first-page":"27699","article-title":"Mathematical capabilities of chatgpt","volume-title":"Advances in Neural Information Processing Systems","author":"Frieder","year":"2024"},{"key":"2026040612495313400_bib7","article-title":"The llama 3 herd of models","author":"Grattafiori","year":"2024","journal-title":"arXiv preprint arXiv:2407.21783v3"},{"key":"2026040612495313400_bib8","article-title":"RULER: What\u2019s the real context size of your long-context language models?","volume-title":"First Conference on Language Modeling","author":"Hsieh","year":"2024"},{"key":"2026040612495313400_bib9","doi-asserted-by":"publisher","first-page":"14982","DOI":"10.18653\/v1\/2024.findings-acl.890","article-title":"Found in the middle: Calibrating positional attention bias improves long context utilization","volume-title":"Findings of the Association for Computational Linguistics ACL 2024","author":"Hsieh","year":"2024"},{"key":"2026040612495313400_bib10","doi-asserted-by":"publisher","first-page":"1419","DOI":"10.18653\/v1\/2021.naacl-main.112","article-title":"Efficient attentions for long document summarization","volume-title":"Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies","author":"Huang","year":"2021"},{"key":"2026040612495313400_bib11","article-title":"The consensus game: Language model generation via equilibrium search","volume-title":"Twelfth International Conference on Learning Representations","author":"Jacob","year":"2024"},{"key":"2026040612495313400_bib12","doi-asserted-by":"publisher","first-page":"1658","DOI":"10.18653\/v1\/2024.acl-long.91","article-title":"Longllmlingua: Accelerating and enhancing llms in long context scenarios via prompt compression","volume-title":"Findings of the Association for Computational Linguistics ACL 2024","author":"Jiang","year":"2024"},{"key":"2026040612495313400_bib13","article-title":"Needle in a haystack - pressure testing llms","author":"Kamradt","year":"2023","journal-title":"GitHub"},{"key":"2026040612495313400_bib14","doi-asserted-by":"publisher","first-page":"317","DOI":"10.1162\/tacl_a_00023","article-title":"The NarrativeQA reading comprehension challenge","volume":"6","author":"Ko\u010disky\u0300","year":"2018","journal-title":"Transactions of the Association for Computational Linguistics"},{"key":"2026040612495313400_bib15","doi-asserted-by":"publisher","first-page":"453","DOI":"10.1162\/tacl_a_00276","article-title":"Natural Questions: A benchmark for question answering research","volume":"7","author":"Kwiatkowski","year":"2019","journal-title":"Transactions of the Association for Computational Linguistics"},{"key":"2026040612495313400_bib16","doi-asserted-by":"publisher","DOI":"10.1145\/3600006.3613165","article-title":"Efficient memory management for large language model serving with PagedAttention","volume-title":"Proceedings of the ACM SIGOPS 29th Symposium on Operating Systems Principles","author":"Kwon","year":"2023"},{"key":"2026040612495313400_bib17","article-title":"Can long-context language models subsume retrieval, rag, sql, and more?","author":"Lee","year":"2024","journal-title":"arXiv preprint arXiv:2406.13121v1"},{"key":"2026040612495313400_bib18","doi-asserted-by":"publisher","first-page":"6086","DOI":"10.18653\/v1\/P19-1612","article-title":"Latent retrieval for weakly supervised open domain question answering","volume-title":"Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics","author":"Lee","year":"2019"},{"key":"2026040612495313400_bib19","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2024.findings-acl.230","article-title":"ROUGE: A Package for Automatic Evaluation of Summaries","volume-title":"ACL Workshop on Text Summarization Branches Out","author":"Lin","year":"2004"},{"key":"2026040612495313400_bib20","doi-asserted-by":"crossref","first-page":"3829","DOI":"10.18653\/v1\/2024.findings-acl.230","article-title":"Just ask one more time! Self-agreement improves reasoning of language models in (almost) all scenarios","volume-title":"Findings of the Association for Computational Linguistics ACL 2024","author":"Lin","year":"2024"},{"key":"2026040612495313400_bib21","doi-asserted-by":"publisher","first-page":"157","DOI":"10.1162\/tacl_a_00638","article-title":"Lost in the middle: How language models use long contexts","volume":"12","author":"Liu","year":"2024","journal-title":"Transactions of the Association for Computational Linguistics"},{"key":"2026040612495313400_bib22","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2022.acl-long.556","article-title":"Fantastically ordered prompts and where to find them: Overcoming few-shot prompt order sensitivity","volume-title":"Annual Meeting of the Association for Computational Linguistics (ACL)","author":"Yao","year":"2022"},{"key":"2026040612495313400_bib23","article-title":"GSM-symbolic: Understanding the limitations of mathematical reasoning in large language models","volume-title":"Thirteenth International Conference on Learning Representations","author":"Mirzadeh","year":"2025"},{"key":"2026040612495313400_bib24","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2022.findings-acl.50","article-title":"Reframing instructional prompts to gptk\u2019s language","volume-title":"Annual Meeting of the Association for Computational Linguistics (ACL) - Findings","author":"Mishra","year":"2022"},{"key":"2026040612495313400_bib25","article-title":"NoLiMa: Long-context evaluation beyond literal matching","volume-title":"Forty-second International Conference on Machine Learning","author":"Modarressi","year":"2025"},{"key":"2026040612495313400_bib26","article-title":"Alice in wonderland: Simple tasks showing complete reasoning breakdown in state-of-the-art large language models","author":"Nezhurina","year":"2024","journal-title":"arXiv preprint arXiv:2406.02061v4"},{"key":"2026040612495313400_bib27","unstructured":"OpenAI. 2024. Hello GPT-4o."},{"key":"2026040612495313400_bib28","doi-asserted-by":"publisher","first-page":"5336","DOI":"10.18653\/v1\/2022.naacl-main.391","article-title":"QuALITY: Question answering with long input texts, yes!","volume-title":"Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies","author":"Pang","year":"2022"},{"key":"2026040612495313400_bib29","article-title":"True few-shot learning with language models","volume-title":"Advances in Neural Information Processing Systems (NeurIPS)","author":"Perez","year":"2021"},{"key":"2026040612495313400_bib30","doi-asserted-by":"publisher","first-page":"7977","DOI":"10.18653\/v1\/2023.findings-emnlp.536","article-title":"ZeroSCROLLS: A zero-shot benchmark for long text understanding","volume-title":"Findings of the Association for Computational Linguistics: EMNLP 2023","author":"Shaham","year":"2023"},{"key":"2026040612495313400_bib31","doi-asserted-by":"publisher","first-page":"12007","DOI":"10.18653\/v1\/2022.emnlp-main.823","article-title":"SCROLLS: Standardized comparison over long language sequences","volume-title":"Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing","author":"Shaham","year":"2022"},{"key":"2026040612495313400_bib32","article-title":"Long range arena: A benchmark for efficient transformers","volume-title":"International Conference on Learning Representations","author":"Yi","year":"2021"},{"key":"2026040612495313400_bib33","doi-asserted-by":"publisher","first-page":"539","DOI":"10.1162\/tacl_a_00475","article-title":"MuSiQue: Multihop questions via single-hop question composition","volume":"10","author":"Trivedi","year":"2022","journal-title":"Transactions of the Association for Computational Linguistics"},{"key":"2026040612495313400_bib34","doi-asserted-by":"publisher","first-page":"1139","DOI":"10.18653\/v1\/2022.emnlp-main.75","article-title":"SQuALITY: Building a long-document summarization dataset the hard way","volume-title":"Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing","author":"Wang","year":"2022"},{"key":"2026040612495313400_bib35","doi-asserted-by":"publisher","first-page":"287","DOI":"10.18653\/v1\/2024.acl-short.28","article-title":"Soft self-consistency improves language models agents","volume-title":"Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers)","author":"Wang","year":"2024"},{"key":"2026040612495313400_bib36","article-title":"Self-consistency improves chain of thought reasoning in language models","volume-title":"Eleventh International Conference on Learning Representations","author":"Wang","year":"2023"},{"key":"2026040612495313400_bib37","doi-asserted-by":"publisher","first-page":"108","DOI":"10.18653\/v1\/2023.emnlp-main.8","article-title":"Primacy effect of ChatGPT","volume-title":"Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing","author":"Wang","year":"2023"},{"key":"2026040612495313400_bib38","first-page":"24824","article-title":"Chain-of-Thought prompting elicits reasoning in large language models","volume-title":"Advances in Neural Information Processing Systems","author":"Wei","year":"2022"},{"key":"2026040612495313400_bib39","doi-asserted-by":"publisher","first-page":"1819","DOI":"10.18653\/v1\/2024.naacl-long.102","article-title":"Reasoning or reciting? Exploring the capabilities and limitations of language models through counterfactual tasks","volume-title":"Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers)","author":"Zhaofeng","year":"2024"},{"key":"2026040612495313400_bib40","article-title":"From artificial needles to real haystacks: Improving retrieval capabilities in LLMs by finetuning on synthetic data","volume-title":"Thirteenth International Conference on Learning Representations","author":"Xiong","year":"2025"},{"key":"2026040612495313400_bib41","article-title":"Qwen2.5 technical report","author":"An","year":"2024","journal-title":"arXiv preprint arXiv:2412.15115v2"},{"key":"2026040612495313400_bib42","article-title":"HELMET: How to evaluate long-context language models effectively and thoroughly","author":"Yen","year":"2024","journal-title":"arXiv preprint arXiv:2410.02694v2"},{"key":"2026040612495313400_bib43","article-title":"Large language models are not robust multiple choice selectors","volume-title":"Twelfth International Conference on Learning Representations","author":"Zheng","year":"2023"},{"key":"2026040612495313400_bib44","doi-asserted-by":"publisher","first-page":"5905","DOI":"10.18653\/v1\/2021.naacl-main.472","article-title":"QMSum: A new benchmark for query-based multi-domain meeting summarization","volume-title":"Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies","author":"Zhong","year":"2021"}],"container-title":["Transactions of the Association for Computational Linguistics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/direct.mit.edu\/tacl\/article-pdf\/doi\/10.1162\/TACL.a.625\/2590845\/tacl.a.625.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/direct.mit.edu\/tacl\/article-pdf\/doi\/10.1162\/TACL.a.625\/2590845\/tacl.a.625.pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,4,6]],"date-time":"2026-04-06T16:50:04Z","timestamp":1775494204000},"score":1,"resource":{"primary":{"URL":"https:\/\/direct.mit.edu\/tacl\/article\/doi\/10.1162\/TACL.a.625\/136156\/Self-Consistency-Falls-Short-The-Adverse-Effects"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026]]},"references-count":44,"URL":"https:\/\/doi.org\/10.1162\/tacl.a.625","relation":{},"ISSN":["2307-387X"],"issn-type":[{"value":"2307-387X","type":"electronic"}],"subject":[],"published-other":{"date-parts":[[2026]]},"published":{"date-parts":[[2026]]}}}