{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,6]],"date-time":"2026-05-06T18:17:31Z","timestamp":1778091451292,"version":"3.51.4"},"reference-count":60,"publisher":"MIT Press","license":[{"start":{"date-parts":[[2026,5,6]],"date-time":"2026-05-06T00:00:00Z","timestamp":1778025600000},"content-version":"vor","delay-in-days":125,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":["direct.mit.edu"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2026,5,4]]},"abstract":"<jats:title>Abstract<\/jats:title>\n                  <jats:p>We present a systematic analysis of module-level design choices in GraphRAG, a retrieval-augmented generation framework that integrates structured knowledge graphs into question answering. Focusing on triple extraction, community clustering, and report generation, we evaluate multiple strategies across two knowledge-intensive benchmarks. Our results show that high-quality triple extraction is critical, as the accuracy and coverage of the resulting knowledge graph can become a bottleneck for downstream reasoning. We also find that the granularity of fundamental knowledge units, as determined by community clustering, has a significant impact on downstream performance: Achieving a balance between factual detail and topical coherence within each unit is important to enable precise and comprehensive retrieval and to facilitate effective multi-hop reasoning. In addition, simple template-based reporting outperforms LLM-based summarization in both accuracy and efficiency. These findings provide practical guidance for the structure- aware design of retrieval-augmented systems.<\/jats:p>","DOI":"10.1162\/tacl.a.615","type":"journal-article","created":{"date-parts":[[2026,5,6]],"date-time":"2026-05-06T18:00:58Z","timestamp":1778090458000},"page":"627-655","update-policy":"https:\/\/doi.org\/10.1162\/mitpressjournals.corrections.policy","source":"Crossref","is-referenced-by-count":0,"title":["Dissecting GraphRAG: A Modular Analysis of Knowledge Structuring for Factoid Question Answering"],"prefix":"10.1162","volume":"14","author":[{"given":"Noriki","family":"Nishida","sequence":"first","affiliation":[{"name":"RIKEN, Japan. noriki.nishida@riken.jp"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Rumana Ferdous","family":"Munne","sequence":"additional","affiliation":[{"name":"RIKEN, Japan"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Shanshan","family":"Liu","sequence":"additional","affiliation":[{"name":"RIKEN, Japan"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Narumi","family":"Tokunaga","sequence":"additional","affiliation":[{"name":"RIKEN, Japan"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yuki","family":"Yamagata","sequence":"additional","affiliation":[{"name":"RIKEN, Japan"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Fei","family":"Cheng","sequence":"additional","affiliation":[{"name":"Kyoto University, Japan"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Kouji","family":"Kozaki","sequence":"additional","affiliation":[{"name":"Osaka Electro-Communication University, Japan"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yuji","family":"Matsumoto","sequence":"additional","affiliation":[{"name":"RIKEN, Japan"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"281","published-online":{"date-parts":[[2026,5,4]]},"reference":[{"key":"2026050614004782000_bib1","doi-asserted-by":"publisher","first-page":"318","DOI":"10.18653\/v1\/2024.bionlp-1.24","article-title":"Document-level clinical entity and relation extraction via knowledge base-guided generation","volume-title":"Proceedings of the 23rd Workshop on Biomedical Natural Language Processing","author":"Bhattarai","year":"2024"},{"key":"2026050614004782000_bib2","doi-asserted-by":"publisher","first-page":"154","DOI":"10.1016\/j.websem.2009.07.002","article-title":"Dbpedia - A crystallization point for the web of data","volume":"7","author":"Bizer","year":"2009","journal-title":"Journal Web Semantics"},{"key":"2026050614004782000_bib3","article-title":"Language models are few-shot learners","volume-title":"Proceedings of the 34th International Conference on Neural Information Processing Systems","author":"Brown","year":"2020"},{"key":"2026050614004782000_bib4","article-title":"Communitykg-rag: Leveraging community structures in knowledge graphs for advanced retrieval-augmented generation in fact-checking","author":"Chang","year":"2024"},{"key":"2026050614004782000_bib5","doi-asserted-by":"crossref","DOI":"10.1609\/aaai.v38i16.29728","article-title":"Benchmarking large language models in retrieval-augmented generation","volume-title":"Proceedings of the Thirty-Eighth AAAI Conference on Artificial Intelligence and Thirty-Sixth Conference on Innovative Applications of Artificial Intelligence and Fourteenth Symposium on Educational Advances in Artificial Intelligence","author":"Chen","year":"2024"},{"key":"2026050614004782000_bib6","doi-asserted-by":"publisher","first-page":"15159","DOI":"10.18653\/v1\/2024.emnlp-main.845","article-title":"Dense X retrieval: What retrieval granularity should we use?","volume-title":"Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing","author":"Chen","year":"2024"},{"key":"2026050614004782000_bib7","doi-asserted-by":"publisher","first-page":"719","DOI":"10.1145\/3626772.3657834","article-title":"The power of noise: Redefining retrieval for rag systems","volume-title":"Proceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval","author":"Cuconasu","year":"2024"},{"key":"2026050614004782000_bib8","doi-asserted-by":"publisher","first-page":"257","DOI":"10.1162\/tacl_a_00459","article-title":"Time-aware language models as temporal knowledge bases","volume":"10","author":"Dhingra","year":"2022","journal-title":"Transactions of the Association for Computational Linguistics"},{"key":"2026050614004782000_bib9","article-title":"From local to global: A graph rag approach to query-focused summarization","author":"Edge","year":"2024"},{"key":"2026050614004782000_bib10","doi-asserted-by":"publisher","first-page":"1295","DOI":"10.18653\/v1\/2020.emnlp-main.99","article-title":"Scalable multi-hop relational reasoning for knowledge-aware question answering","volume-title":"Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP)","author":"Feng","year":"2020"},{"key":"2026050614004782000_bib11","article-title":"Retrieval-augmented generation for large language models: A survey","author":"Gao","year":"2024"},{"key":"2026050614004782000_bib12","doi-asserted-by":"publisher","first-page":"3064","DOI":"10.1145\/3539618.3591912","article-title":"Linked-docred - enhancing docred with entity-linking to evaluate end-to-end document- level information extraction pipelines","volume-title":"Proceedings of the 46th International ACM SIGIR Conference on Research and Development in Information Retrieval","author":"Genest","year":"2023"},{"key":"2026050614004782000_bib13","article-title":"The llama 3 herd of models","author":"Grattafiori","year":"2024"},{"key":"2026050614004782000_bib14","article-title":"Rag vs. graphrag: A systematic evaluation and key insights","author":"Han","year":"2025"},{"key":"2026050614004782000_bib15","article-title":"Retrieval-augmented generation with graphs (graphrag)","author":"Han","year":"2025"},{"key":"2026050614004782000_bib16","doi-asserted-by":"publisher","first-page":"25698","DOI":"10.18653\/v1\/2025.findings-acl.1319","article-title":"Reasoning with graphs: Structuring implicit knowledge to enhance LLMs reasoning","volume-title":"Findings of the Association for Computational Linguistics: ACL 2025","author":"Han","year":"2025"},{"key":"2026050614004782000_bib17","doi-asserted-by":"publisher","first-page":"13417","DOI":"10.18653\/v1\/2023.acl-long.750","article-title":"MVP-tuning: Multi-view knowledge retrieval with prompt tuning for commonsense reasoning","volume-title":"Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)","author":"Huang","year":"2023"},{"key":"2026050614004782000_bib18","doi-asserted-by":"publisher","DOI":"10.48550\/ARXIV.2112.09118","article-title":"Unsupervised dense information retrieval with contrastive learning","author":"Izacard","year":"2021"},{"issue":"1","key":"2026050614004782000_bib19","article-title":"Atlas: Few-shot learning with retrieval augmented language models","volume":"24","author":"Izacard","year":"2023","journal-title":"Journal of Machine Learning Research"},{"key":"2026050614004782000_bib20","doi-asserted-by":"publisher","first-page":"163","DOI":"10.18653\/v1\/2024.findings-acl.11","article-title":"Graph chain-of-thought: Augmenting large language models by reasoning on graphs","volume-title":"Findings of the Association for Computational Linguistics: ACL 2024","author":"Jin","year":"2024"},{"key":"2026050614004782000_bib21","doi-asserted-by":"publisher","DOI":"10.52202\/075280-2130","article-title":"Realtime qa: What\u2019s the answer right now?","volume-title":"Proceedings of the 37th International Conference on Neural Information Processing Systems","author":"Kasai","year":"2023"},{"key":"2026050614004782000_bib22","doi-asserted-by":"publisher","first-page":"2526","DOI":"10.18653\/v1\/2021.findings-acl.223","article-title":"JointGT: Graph-text joint representation learning for text generation from knowledge graphs","volume-title":"Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021","author":"Ke","year":"2021"},{"key":"2026050614004782000_bib23","doi-asserted-by":"publisher","first-page":"12894","DOI":"10.18653\/v1\/2025.findings-acl.668","article-title":"BioHopR: A benchmark for multi-hop, multi-answer reasoning in biomedical domain","volume-title":"Findings of the Association for Computational Linguistics: ACL 2025","author":"Kim","year":"2025"},{"key":"2026050614004782000_bib24","article-title":"Large language models are zero-shot reasoners","volume-title":"Proceedings of the 36th International Conference on Neural Information Processing Systems","author":"Kojima","year":"2022"},{"key":"2026050614004782000_bib25","article-title":"Toward optimal search and retrieval for rag","author":"Leto","year":"2024"},{"key":"2026050614004782000_bib26","article-title":"Retrieval-augmented generation for knowledge-intensive nlp tasks","volume-title":"Proceedings of the 34th International Conference on Neural Information Processing Systems","author":"Lewis","year":"2020"},{"key":"2026050614004782000_bib27","doi-asserted-by":"publisher","DOI":"10.1093\/database\/baw068","article-title":"Biocreative V CDR task corpus: A resource for chemical disease relation extraction","volume":"2016","author":"Li","year":"2016","journal-title":"Database: Journal of Biology Databases and Curation"},{"key":"2026050614004782000_bib28","article-title":"Simple is effective: The roles of graphs and large language models in knowledge-graph-based retrieval-augmented generation","author":"Li","year":"2025"},{"key":"2026050614004782000_bib29","article-title":"Unioqa: A unified framework for knowledge graph question answering with large language models","author":"Li","year":"2024"},{"issue":"9","key":"2026050614004782000_bib30","doi-asserted-by":"publisher","DOI":"10.1145\/3560815","article-title":"Pre-train, prompt, and predict: A systematic survey of prompting methods in natural language processing","volume":"55","author":"Liu","year":"2023","journal-title":"ACM Computing Surveys"},{"key":"2026050614004782000_bib31","doi-asserted-by":"publisher","first-page":"5944","DOI":"10.18653\/v1\/2022.naacl-main.435","article-title":"Time waits for no one! Analysis and challenges of temporal misalignment","volume-title":"Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies","author":"Luu","year":"2022"},{"key":"2026050614004782000_bib32","doi-asserted-by":"publisher","first-page":"9802","DOI":"10.18653\/v1\/2023.acl-long.546","article-title":"When not to trust language models: Investigating effectiveness of parametric and non-parametric memories","volume-title":"Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)","author":"Mallen","year":"2023"},{"key":"2026050614004782000_bib33","article-title":"Gpt-4 technical report","author":"OpenAI","year":"2024"},{"key":"2026050614004782000_bib34","article-title":"Document-level in-context few-shot relation extraction via pre-trained language models","author":"Ozyurt","year":"2024"},{"key":"2026050614004782000_bib35","unstructured":"Qwen-Team. 2024. Qwen2.5: A party of foundation models."},{"key":"2026050614004782000_bib36","article-title":"RAPTOR: Recursive abstractive processing for tree-organized retrieval","volume-title":"The Twelfth International Conference on Learning Representations","author":"Sarthi","year":"2024"},{"key":"2026050614004782000_bib37","doi-asserted-by":"publisher","first-page":"6138","DOI":"10.18653\/v1\/2021.emnlp-main.496","article-title":"Simple entity-centric questions challenge dense retrievers","volume-title":"Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing","author":"Sciavolino","year":"2021"},{"key":"2026050614004782000_bib38","doi-asserted-by":"publisher","first-page":"12","DOI":"10.1145\/3673791.3698415","article-title":"Fine tuning vs. retrieval augmented generation for less popular knowledge","author":"Soudani","year":"2024"},{"key":"2026050614004782000_bib39","doi-asserted-by":"publisher","first-page":"5233","DOI":"10.1038\/s41598-019-41695-z","article-title":"From Louvain to Leiden: Guaranteeing well-connected communities","volume":"9","author":"Traag","year":"2019","journal-title":"Scientific Reports"},{"key":"2026050614004782000_bib40","doi-asserted-by":"publisher","first-page":"15566","DOI":"10.18653\/v1\/2023.acl-long.868","article-title":"Revisiting relation extraction in the era of large language models","volume-title":"Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)","author":"Wadhwa","year":"2023"},{"key":"2026050614004782000_bib41","doi-asserted-by":"publisher","first-page":"30553","DOI":"10.18653\/v1\/2025.acl-long.1476","article-title":"Astute RAG: Overcoming imperfect retrieval augmentation and knowledge conflicts for large language models","volume-title":"Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)","author":"Wang","year":"2025"},{"key":"2026050614004782000_bib42","doi-asserted-by":"publisher","first-page":"27152","DOI":"10.18653\/v1\/2025.acl-long.1317","article-title":"Knowledge graph retrieval-augmented generation for LLM-based recommendation","volume-title":"Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)","author":"Wang","year":"2025"},{"key":"2026050614004782000_bib43","doi-asserted-by":"publisher","first-page":"10370","DOI":"10.18653\/v1\/2024.acl-long.558","article-title":"MindMap: Knowledge graph prompting sparks graph of thoughts in large language models","volume-title":"Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)","author":"Wen","year":"2024"},{"key":"2026050614004782000_bib44","article-title":"Medical graph rag: Towards safe medical large language model via graph retrieval-augmented generation","author":"Junde","year":"2024"},{"key":"2026050614004782000_bib45","doi-asserted-by":"publisher","first-page":"6397","DOI":"10.18653\/v1\/2020.emnlp-main.519","article-title":"Scalable zero-shot entity linking with dense entity retrieval","volume-title":"Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP)","author":"Ledell","year":"2020"},{"key":"2026050614004782000_bib46","doi-asserted-by":"publisher","first-page":"4275","DOI":"10.18653\/v1\/2025.naacl-long.216","article-title":"Knowledge-aware query expansion with large language models for textual and relational retrieval","volume-title":"Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers)","author":"Xia","year":"2025"},{"key":"2026050614004782000_bib47","article-title":"Harnessing large language models for knowledge graph question answering via adaptive multi-aspect retrieval-augmentation","author":"Derong","year":"2025"},{"key":"2026050614004782000_bib48","article-title":"Qwen2 technical report","author":"An","year":"2024","journal-title":"arXiv preprint arXiv:2407.10671"},{"key":"2026050614004782000_bib49","article-title":"CRAG - comprehensive RAG benchmark","volume-title":"The Thirty-eight Conference on Neural Information Processing Systems Datasets and Benchmarks Track","author":"Yang","year":"2024"},{"key":"2026050614004782000_bib50","doi-asserted-by":"publisher","first-page":"2369","DOI":"10.18653\/v1\/D18-1259","article-title":"HotpotQA: A dataset for diverse, explainable multi-hop question answering","volume-title":"Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing","author":"Yang","year":"2018"},{"key":"2026050614004782000_bib51","doi-asserted-by":"publisher","first-page":"1","DOI":"10.18653\/v1\/2022.emnlp-main.1","article-title":"Generative knowledge graph construction: A review","volume-title":"Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing","author":"Ye","year":"2022"},{"key":"2026050614004782000_bib52","article-title":"DecAF: Joint decoding of answers and logical forms for question answering over knowledge bases","volume-title":"The Eleventh International Conference on Learning Representations","author":"Donghan","year":"2023"},{"key":"2026050614004782000_bib53","doi-asserted-by":"publisher","first-page":"6470","DOI":"10.18653\/v1\/2020.acl-main.577","article-title":"Named entity recognition as dependency parsing","volume-title":"Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics","author":"Juntao","year":"2020"},{"key":"2026050614004782000_bib54","doi-asserted-by":"publisher","first-page":"1256","DOI":"10.26615\/978-954-452-092-2_133","article-title":"Evaluating generative models for graph-to-text generation","volume-title":"Proceedings of the 14th International Conference on Recent Advances in Natural Language Processing","author":"Yuan","year":"2023"},{"key":"2026050614004782000_bib55","doi-asserted-by":"publisher","first-page":"8693","DOI":"10.18653\/v1\/2023.emnlp-main.537","article-title":"Structure-aware knowledge graph-to-text generation with planning selection and similarity distinction","volume-title":"Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing","author":"Zhao","year":"2023"},{"key":"2026050614004782000_bib56","article-title":"Retrieval augmented generation (rag) and beyond: A comprehensive survey on how to make your llms use external data more wisely","author":"Zhao","year":"2024"},{"issue":"4","key":"2026050614004782000_bib57","doi-asserted-by":"publisher","DOI":"10.1145\/3618295","article-title":"A comprehensive survey on automatic knowledge graph construction","volume":"56","author":"Zhong","year":"2023","journal-title":"ACM Computing Surveys"},{"key":"2026050614004782000_bib58","first-page":"5756","article-title":"Mix-of-granularity: Optimize the chunking granularity for retrieval-augmented generation","volume-title":"Proceedings of the 31st International Conference on Computational Linguistics","author":"Zhong","year":"2025"},{"key":"2026050614004782000_bib59","article-title":"Document-level relation extraction with adaptive thresholding and localized context pooling","author":"Zhou","year":"2020","journal-title":"CoRR"},{"key":"2026050614004782000_bib60","doi-asserted-by":"publisher","first-page":"8912","DOI":"10.18653\/v1\/2025.naacl-long.449","article-title":"Knowledge graph-guided retrieval augmented generation","volume-title":"Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers)","author":"Zhu","year":"2025"}],"container-title":["Transactions of the Association for Computational Linguistics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/direct.mit.edu\/tacl\/article-pdf\/doi\/10.1162\/TACL.a.615\/2598680\/tacl.a.615.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/direct.mit.edu\/tacl\/article-pdf\/doi\/10.1162\/TACL.a.615\/2598680\/tacl.a.615.pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,5,6]],"date-time":"2026-05-06T18:01:05Z","timestamp":1778090465000},"score":1,"resource":{"primary":{"URL":"https:\/\/direct.mit.edu\/tacl\/article\/doi\/10.1162\/TACL.a.615\/136549\/Dissecting-GraphRAG-A-Modular-Analysis-of"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026]]},"references-count":60,"URL":"https:\/\/doi.org\/10.1162\/tacl.a.615","relation":{},"ISSN":["2307-387X"],"issn-type":[{"value":"2307-387X","type":"electronic"}],"subject":[],"published-other":{"date-parts":[[2026]]},"published":{"date-parts":[[2026]]}}}