{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,8,6]],"date-time":"2026-08-06T03:07:02Z","timestamp":1785985622472,"version":"3.56.0"},"reference-count":65,"publisher":"Association for Computing Machinery (ACM)","issue":"2","license":[{"start":{"date-parts":[[2025,1,21]],"date-time":"2025-01-21T00:00:00Z","timestamp":1737417600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"DOI":"10.13039\/501100018537","name":"National Science and Technology Major Project","doi-asserted-by":"crossref","award":["2023ZD0121104"],"award-info":[{"award-number":["2023ZD0121104"]}],"id":[{"id":"10.13039\/501100018537","id-type":"DOI","asserted-by":"crossref"}]},{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"crossref","award":["62222213, 62072423"],"award-info":[{"award-number":["62222213, 62072423"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Inf. Syst."],"published-print":{"date-parts":[[2025,3,31]]},"abstract":"<jats:p>\n            Retrieval-augmented generation (RAG) is a technique that enhances the capabilities of large language models (LLMs) by incorporating external knowledge sources. This method addresses common LLM limitations, including outdated information and the tendency to produce inaccurate \u201challucinated\u201d content. However, evaluating RAG systems is a challenge. Most benchmarks focus primarily on question-answering applications, neglecting other potential scenarios where RAG could be beneficial. Accordingly, in the experiments, these benchmarks often assess only the LLM components of the RAG pipeline or the retriever in knowledge-intensive scenarios, overlooking the impact of external knowledge base construction and the retrieval component on the entire RAG pipeline in non-knowledge-intensive scenarios. To address these issues, this article constructs a large-scale and more comprehensive benchmark and evaluates all the components of RAG systems in various RAG application scenarios. Specifically, we refer to the CRUD actions that describe interactions between users and knowledge bases and also categorize the range of RAG applications into four distinct types\u2014create, read, update, and delete (CRUD). \u201cCreate\u201d refers to scenarios requiring the generation of original, varied content. \u201cRead\u201d involves responding to intricate questions in knowledge-intensive situations. \u201cUpdate\u201d focuses on revising and rectifying inaccuracies or inconsistencies in pre-existing texts. \u201cDelete\u201d pertains to the task of summarizing extensive texts into more concise forms. For each of these CRUD categories, we have developed different datasets to evaluate the performance of RAG systems. We also analyze the effects of various components of the RAG system, such as the retriever, context length, knowledge base construction, and LLM. Finally, we provide useful insights for optimizing the RAG technology for different scenarios. The source code is available at GitHub:\n            <jats:ext-link xmlns:xlink=\"http:\/\/www.w3.org\/1999\/xlink\" ext-link-type=\"url\" xlink:href=\"https:\/\/github.com\/IAAR-Shanghai\/CRUD_RAG\">https:\/\/github.com\/IAAR-Shanghai\/CRUD_RAG<\/jats:ext-link>\n            .\n          <\/jats:p>","DOI":"10.1145\/3701228","type":"journal-article","created":{"date-parts":[[2024,10,19]],"date-time":"2024-10-19T17:58:07Z","timestamp":1729360687000},"page":"1-32","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":89,"title":["CRUD-RAG: A Comprehensive Chinese Benchmark for Retrieval-Augmented Generation of Large Language Models"],"prefix":"10.1145","volume":"43","author":[{"ORCID":"https:\/\/orcid.org\/0009-0001-2628-334X","authenticated-orcid":false,"given":"Yuanjie","family":"Lyu","sequence":"first","affiliation":[{"name":"University of Science and Technology of China, Hefei, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0008-3196-7739","authenticated-orcid":false,"given":"Zhiyu","family":"Li","sequence":"additional","affiliation":[{"name":"Institute for Advanced Algorithms Research (Shanghai), Shanghai, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0009-1862-8959","authenticated-orcid":false,"given":"Simin","family":"Niu","sequence":"additional","affiliation":[{"name":"Renmin University of China, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-1456-2202","authenticated-orcid":false,"given":"Feiyu","family":"Xiong","sequence":"additional","affiliation":[{"name":"Institute for Advanced Algorithms Research (Shanghai), Shanghai, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-7129-0250","authenticated-orcid":false,"given":"Bo","family":"Tang","sequence":"additional","affiliation":[{"name":"Institute for Advanced Algorithms Research (Shanghai), Shanghai, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-7150-6162","authenticated-orcid":false,"given":"Wenjin","family":"Wang","sequence":"additional","affiliation":[{"name":"Institute for Advanced Algorithms Research (Shanghai), Shanghai, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0005-1497-2876","authenticated-orcid":false,"given":"Hao","family":"Wu","sequence":"additional","affiliation":[{"name":"Institute for Advanced Algorithms Research (Shanghai), Shanghai, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0002-5370-3888","authenticated-orcid":false,"given":"Huanyong","family":"Liu","sequence":"additional","affiliation":[{"name":"360 AI Research Institute, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-4246-5386","authenticated-orcid":false,"given":"Tong","family":"Xu","sequence":"additional","affiliation":[{"name":"University of Science and Technology of China, Hefei, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-4835-4102","authenticated-orcid":false,"given":"Enhong","family":"Chen","sequence":"additional","affiliation":[{"name":"University of Science and Technology of China, Hefei, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2025,1,21]]},"reference":[{"key":"e_1_3_1_2_2","doi-asserted-by":"publisher","DOI":"10.1145\/3331184.3331265"},{"key":"e_1_3_1_3_2","first-page":"41","volume-title":"Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics: Tutorial Abstracts (ACL \u201923)","author":"Asai Akari","year":"2023","unstructured":"Akari Asai, Sewon Min, Zexuan Zhong, and Danqi Chen. 2023. Retrieval-based language models and applications. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics: Tutorial Abstracts (ACL \u201923). 41\u201346."},{"key":"e_1_3_1_4_2","doi-asserted-by":"publisher","DOI":"10.1145\/3539618.3591923"},{"key":"e_1_3_1_5_2","unstructured":"Alec Berntson. 2023. Azure AI Search: Outperforming Vector Search with Hybrid Retrieval and Ranking Capabilities. Retrieved from https:\/\/techcommunity.microsoft.com\/t5\/ai-azure-ai-services-blog\/azure-ai-search-outperforming-vector-search-with-hybrid\/ba-p\/3929167"},{"key":"e_1_3_1_6_2","unstructured":"S\u00e9bastien Bubeck Varun Chandrasekaran Ronen Eldan Johannes Gehrke Eric Horvitz Ece Kamar Peter Lee Yin Tat Lee Yuanzhi Li Scott Lundberg et al. 2023. Sparks of artificial general intelligence: Early experiments with gpt-4. arXiv:2303.12712. Retrieved from https:\/\/arxiv.org\/abs\/2303.12712"},{"key":"e_1_3_1_7_2","first-page":"6251","volume-title":"Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP \u201920)","author":"Cao Meng","year":"2020","unstructured":"Meng Cao, Yue Dong, Jiapeng Wu, and Jackie Chi Kit Cheung. 2020. Factual error correction for abstractive summarization models. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP \u201920). 6251\u20136258."},{"key":"e_1_3_1_8_2","unstructured":"Shuyang Cao and Lu Wang. 2024. Verifiable generation with subsentence-level fine-grained citations. arXiv:2406.06125."},{"key":"e_1_3_1_9_2","unstructured":"Jiawei Chen Hongyu Lin Xianpei Han and Le Sun. 2023. Benchmarking large language models in retrieval-augmented generation. arXiv:2309.01431. Retrieved from https:\/\/arxiv.org\/abs\/2309.01431"},{"key":"e_1_3_1_10_2","doi-asserted-by":"publisher","DOI":"10.1145\/3583780.3614905"},{"key":"e_1_3_1_11_2","unstructured":"Xin Cheng Di Luo Xiuying Chen Lemao Liu Dongyan Zhao and Rui Yan. 2023. Lift Yourself Up: Retrieval-augmented text generation with self memory. arXiv:2305.02437. Retrieved from https:\/\/arxiv.org\/abs\/2305.02437"},{"key":"e_1_3_1_12_2","volume-title":"Proceedings of the 11th International Conference on Learning Representations (ICLR \u201923)","author":"Dai Zhuyun","year":"2023","unstructured":"Zhuyun Dai, Vincent Y. Zhao, Ji Ma, Yi Luan, Jianmo Ni, Jing Lu, Anton Bakalov, Kelvin Guu, Keith B. Hall, and Ming-Wei Chang. 2023. Promptagator: Few-shot dense retrieval from 8 examples. In Proceedings of the 11th International Conference on Learning Representations (ICLR \u201923)."},{"key":"e_1_3_1_13_2","doi-asserted-by":"crossref","unstructured":"Esin Durmus He He and Mona Diab. 2020. FEQA: A question answering evaluation framework for faithfulness assessment in abstractive summarization. arXiv:2005.03754. Retrieved from https:\/\/arxiv.org\/abs\/2005.03754","DOI":"10.18653\/v1\/2020.acl-main.454"},{"key":"e_1_3_1_14_2","unstructured":"Shahul Es Jithin James Luis Espinosa-Anke and Steven Schockaert. 2023. RAGAs: Automated evaluation of retrieval augmented generation. arXiv:2309.15217. Retrieved from https:\/\/arxiv.org\/abs\/2309.15217"},{"key":"e_1_3_1_15_2","unstructured":"Joe Ferrara Ethan-Tonic and Oguzhan Mete Ozturk. 2024. The RAG Triad. Retrieved from https:\/\/www.trulens.org\/trulens_eval\/core_concepts_rag_triad\/"},{"key":"e_1_3_1_16_2","doi-asserted-by":"publisher","DOI":"10.1145\/2766462.2767780"},{"key":"e_1_3_1_17_2","first-page":"1762","volume-title":"Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) (ACL \u201923)","author":"Gao Luyu","year":"2023","unstructured":"Luyu Gao, Xueguang Ma, Jimmy Lin, and Jamie Callan. 2023. Precise zero-shot dense retrieval without relevance labels. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) (ACL \u201923), 1762\u20131777."},{"key":"e_1_3_1_18_2","first-page":"6465","volume-title":"Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing (EMNLP \u201923)","author":"Gao Tianyu","year":"2023","unstructured":"Tianyu Gao, Howard Yen, Jiatong Yu, and Danqi Chen. 2023. Enabling large language models to generate text with citations. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing (EMNLP \u201923), 6465\u20136488."},{"key":"e_1_3_1_19_2","unstructured":"Yunfan Gao Yun Xiong Xinyu Gao Kangxiang Jia Jinliu Pan Yuxi Bi Yi Dai Jiawei Sun Qianyu Guo Meng Wang and Haofen Wang. 2023. Retrieval-augmented generation for large language models: A survey. arXiv:2312.10997. Retrieved from https:\/\/arxiv.org\/abs\/2312.10997"},{"key":"e_1_3_1_20_2","unstructured":"Hangfeng He Hongming Zhang and Dan Roth. 2022. Rethinking with retrieval: Faithful large language model inference. arXiv:2301.00303. Retrieved from https:\/\/arxiv.org\/abs\/2301.00303"},{"key":"e_1_3_1_21_2","unstructured":"Ivan Ilin. 2023. Advanced RAG Techniques: An Illustrated Overview. Retrieved from https:\/\/pub.towardsai.net\/advanced-rag-techniques-an-illustrated-overview-04d193d8fec6"},{"key":"e_1_3_1_22_2","unstructured":"Gautier Izacard Patrick S. H. Lewis Maria Lomeli Lucas Hosseini Fabio Petroni Timo Schick Jane Dwivedi-Yu Armand Joulin Sebastian Riedel and Edouard Grave. 2022. Few-shot learning with retrieval augmented language models. arXiv:2208.03299. Retrieved from https:\/\/arxiv.org\/abs\/2208.03299"},{"key":"e_1_3_1_23_2","first-page":"18345","volume-title":"Proceedings of 38th AAAI Conference on Artificial Intelligence (AAAI \u201924), 36th Conference on Innovative Applications of Artificial Intelligence (IAAI \u201924), 14th Symposium on Educational Advances in Artificial Intelligence (EAAI \u201914)","author":"Ji Bin","year":"2024","unstructured":"Bin Ji, Huijun Liu, Mingzhe Du, and See-Kiong Ng. 2024. Chain-of-thought improves text generation with citations in large language models. In Proceedings of 38th AAAI Conference on Artificial Intelligence (AAAI \u201924), 36th Conference on Innovative Applications of Artificial Intelligence (IAAI \u201924), 14th Symposium on Educational Advances in Artificial Intelligence (EAAI \u201914), 18345\u201318353."},{"key":"e_1_3_1_24_2","doi-asserted-by":"crossref","first-page":"7969","DOI":"10.18653\/v1\/2023.emnlp-main.495","volume-title":"Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing (EMNLP \u201923)","author":"Jiang Zhengbao","year":"2023","unstructured":"Zhengbao Jiang, Frank F. Xu, Luyu Gao, Zhiqing Sun, Qian Liu, Jane Dwivedi-Yu, Yiming Yang, Jamie Callan, and Graham Neubig. 2023. Active retrieval augmented generation. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing (EMNLP \u201923), 7969\u20137992."},{"key":"e_1_3_1_25_2","unstructured":"Minki Kang Jin Myung Kwak Jinheon Baek and Sung Ju Hwang. 2023. Knowledge graph-augmented language models for knowledge-grounded dialogue generation. arXiv:2305.18846. Retrieved from https:\/\/arxiv.org\/abs\/2305.18846"},{"key":"e_1_3_1_26_2","first-page":"385","volume-title":"Proceedings of the 1st International Conference on Systems Integration (Systems Integration \u201990)","author":"Kilov Haim","year":"1990","unstructured":"Haim Kilov. 1990. From semantic to object-oriented data modeling. In Proceedings of the 1st International Conference on Systems Integration (Systems Integration \u201990). IEEE, 385\u2013393."},{"key":"e_1_3_1_27_2","doi-asserted-by":"publisher","DOI":"10.1162\/tacl_a_00276"},{"key":"e_1_3_1_28_2","unstructured":"Langchain. 2023. Evaluating RAG Architectures on Benchmark Tasks. Retrieved from https:\/\/langchain-ai.github.io\/langchain-benchmarks\/notebooks\/retrieval\/comparing_techniques.html"},{"key":"e_1_3_1_29_2","doi-asserted-by":"publisher","DOI":"10.5555\/3495724.3496517"},{"key":"e_1_3_1_30_2","unstructured":"Weitao Li Junkai Li Weizhi Ma and Yang Liu. 2024. Citation-enhanced generation for LLM-based chatbots. arXiv:2402.16063. Retrieved from https:\/\/arxiv.org\/abs\/2402.16063"},{"key":"e_1_3_1_31_2","first-page":"408","volume-title":"Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing: Industry Track","author":"Li Xianzhi","year":"2023","unstructured":"Xianzhi Li, Samuel Chan, Xiaodan Zhu, Yulong Pei, Zhiqiang Ma, Xiaomo Liu, and Sameena Shah. 2023. Are ChatGPT and GPT-4 general-purpose solvers for financial text analytics? A study on several typical tasks. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing: Industry Track, 408\u2013422."},{"key":"e_1_3_1_32_2","unstructured":"Xun Liang Shichao Song Simin Niu Zhiyu Li Feiyu Xiong Bo Tang Zhaohui Wy Dawei He Peng Cheng Zhonghao Wang et al. 2023. UHGEval: Benchmarking the hallucination of Chinese large language models via unconstrained generation. arXiv:2311.15296. Retrieved from https:\/\/arxiv.org\/abs\/2311.15296"},{"key":"e_1_3_1_33_2","first-page":"7001","article-title":"Evaluating verifiability in generative search engines","author":"Liu Nelson F.","year":"2023","unstructured":"Nelson F. Liu, Tianyi Zhang, and Percy Liang. 2023. Evaluating verifiability in generative search engines. In Findings of the Association for Computational Linguistics (EMNLP \u201923), 7001\u20137025.","journal-title":"Findings of the Association for Computational Linguistics (EMNLP \u201923)"},{"issue":"5","key":"e_1_3_1_34_2","first-page":"118:1","article-title":"An analysis on matching mechanisms and token pruning for late-interaction models","volume":"42","author":"Liu Qi","year":"2024","unstructured":"Qi Liu, Gang Guo, Jiaxin Mao, Zhicheng Dou, Ji-Rong Wen, Hao Jiang, Xinyu Zhang, and Zhao Cao. 2024. An analysis on matching mechanisms and token pruning for late-interaction models. ACM Transactions on Information Systems 42, 5 (2024), 118:1\u2013118:28.","journal-title":"ACM Transactions on Information Systems"},{"key":"e_1_3_1_35_2","doi-asserted-by":"crossref","unstructured":"Yang Liu and Mirella Lapata. 2019. Hierarchical transformers for multi-document summarization. arXiv:1905.13164. Retrieved from https:\/\/arxiv.org\/abs\/1905.13164","DOI":"10.18653\/v1\/P19-1500"},{"key":"e_1_3_1_36_2","doi-asserted-by":"publisher","DOI":"10.1145\/3511808.3557319"},{"key":"e_1_3_1_37_2","unstructured":"Xinbei Ma Yeyun Gong Pengcheng He Hai Zhao and Nan Duan. 2023. Query rewriting for retrieval-augmented large language models. arXiv:2305.14283. Retrieved from https:\/\/arxiv.org\/abs\/2305.14283"},{"key":"e_1_3_1_38_2","first-page":"10572","article-title":"Large language model is not a good few-shot information extractor, but a good reranker for hard samples!","author":"Ma Yubo","year":"2023","unstructured":"Yubo Ma, Yixin Cao, Yong Hong, and Aixin Sun. 2023. Large language model is not a good few-shot information extractor, but a good reranker for hard samples!. In Findings of the Association for Computational Linguistics (EMNLP \u201923), 10572\u201310601.","journal-title":"Findings of the Association for Computational Linguistics (EMNLP \u201923)"},{"key":"e_1_3_1_39_2","doi-asserted-by":"crossref","unstructured":"Niklas Muennighoff Nouamane Tazi Lo\u00efc Magne and Nils Reimers. 2022. MTEB: Massive text embedding benchmark. arXiv:2210.07316. Retrieved from https:\/\/arxiv.org\/abs\/2210.07316","DOI":"10.18653\/v1\/2023.eacl-main.148"},{"key":"e_1_3_1_40_2","unstructured":"Darren Oberst. 2023. How to Evaluate LLMs for RAG? Retrieved from https:\/\/medium.com\/@darrenoberst\/how-accurate-is-rag-8f0706281fd9"},{"key":"e_1_3_1_41_2","doi-asserted-by":"crossref","unstructured":"Fabio Petroni Aleksandra Piktus Angela Fan Patrick Lewis Majid Yazdani Nicola De Cao James Thorne Yacine Jernite Vladimir Karpukhin Jean Maillard et al. 2020. KILT: A benchmark for knowledge intensive language tasks. arXiv:2009.02252. Retrieved from https:\/\/arxiv.org\/abs\/2009.02252","DOI":"10.18653\/v1\/2021.naacl-main.200"},{"key":"e_1_3_1_42_2","doi-asserted-by":"publisher","DOI":"10.1145\/3397271.3401110"},{"key":"e_1_3_1_43_2","doi-asserted-by":"publisher","DOI":"10.1162\/coli_a_00486"},{"key":"e_1_3_1_44_2","first-page":"1172","volume-title":"Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (NAACL-HLT \u201921)","author":"Raunak Vikas","year":"2021","unstructured":"Vikas Raunak, Arul Menezes, and Marcin Junczys-Dowmunt. 2021. The curious case of hallucinations in neural machine translation. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (NAACL-HLT \u201921), 1172\u20131183."},{"key":"e_1_3_1_45_2","doi-asserted-by":"crossref","unstructured":"Jon Saad-Falcon Omar Khattab Christopher Potts and Matei Zaharia. 2023. ARES: An automated evaluation framework for retrieval-augmented generation systems. arXiv:2311.09476. Retrieved from https:\/\/arxiv.org\/abs\/2311.09476","DOI":"10.18653\/v1\/2024.naacl-long.20"},{"key":"e_1_3_1_46_2","doi-asserted-by":"crossref","unstructured":"Thomas Scialom Paul-Alexis Dray Patrick Gallinari Sylvain Lamprier Benjamin Piwowarski Jacopo Staiano and Alex Wang. 2021. QuestEval: Summarization asks for fact-based evaluation. arXiv:2103.12693. Retrieved from https:\/\/arxiv.org\/abs\/2103.12693","DOI":"10.18653\/v1\/2021.emnlp-main.529"},{"key":"e_1_3_1_47_2","unstructured":"Xinyue Shen Zeyuan Chen Michael Backes and Yang Zhang. 2023. In ChatGPT we trust? Measuring and characterizing the reliability of ChatGPT. arXiv:2304.08979. Retrieved from https:\/\/arxiv.org\/abs\/2304.08979"},{"key":"e_1_3_1_48_2","unstructured":"Weijia Shi Sewon Min Michihiro Yasunaga Minjoon Seo Rich James Mike Lewis Luke Zettlemoyer and Wen-tau Yih. 2023. REPLUG: Retrieval-augmented black-box language models. arXiv:2301.12652. Retrieved from https:\/\/arxiv.org\/abs\/2301.12652"},{"key":"e_1_3_1_49_2","first-page":"191","volume-title":"Proceedings of the 20th International Conference on Control Systems and Computer Science","author":"Truica Ciprian-Octavian","year":"2015","unstructured":"Ciprian-Octavian Truica, Florin Radulescu, Alexandru Boicea, and Ion Bucur. 2015. Performance evaluation for CRUD operations in asynchronously replicated document oriented database. In Proceedings of the 20th International Conference on Control Systems and Computer Science. IEEE, 191\u2013196."},{"key":"e_1_3_1_50_2","doi-asserted-by":"crossref","unstructured":"Alex Wang Kyunghyun Cho and Mike Lewis. 2020. Asking and answering questions to evaluate the factual consistency of summaries. arXiv:2004.04228. Retrieved from https:\/\/arxiv.org\/abs\/2004.04228","DOI":"10.18653\/v1\/2020.acl-main.450"},{"key":"e_1_3_1_51_2","doi-asserted-by":"crossref","first-page":"9414","DOI":"10.18653\/v1\/2023.emnlp-main.585","volume-title":"Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing (EMNLP \u201923)","author":"Wang Liang","year":"2023","unstructured":"Liang Wang, Nan Yang, and Furu Wei. 2023. Query2doc: Query expansion with large language models. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing (EMNLP \u201923), 9414\u20139423."},{"key":"e_1_3_1_52_2","first-page":"24824","volume-title":"Proceedings of the Advances in Neural Information Processing Systems,","volume":"35","author":"Wei Jason","year":"2022","unstructured":"Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V. Le, and Denny Zhou. 2022. Chain-of-thought prompting elicits reasoning in large language models. In Proceedings of the Advances in Neural Information Processing Systems, Vol. 35, 24824\u201324837."},{"key":"e_1_3_1_53_2","doi-asserted-by":"publisher","DOI":"10.1007\/s11704-023-2689-5"},{"key":"e_1_3_1_54_2","unstructured":"Yilin Wen Zifeng Wang and Jimeng Sun. 2023. MindMap: Knowledge graph prompting sparks graph of thoughts in large language models. arXiv:2308.09729. Retrieved from https:\/\/arxiv.org\/abs\/2308.09729"},{"key":"e_1_3_1_55_2","unstructured":"Derong Xu Wei Chen Wenjun Peng Chao Zhang Tong Xu Xiangyu Zhao Xian Wu Yefeng Zheng and Enhong Chen. 2023. Large language models for generative information extraction: A survey. Frontiers of Computer Science. Retrieved from https:\/\/journal.hep.com.cn\/fcs\/EN\/10.1007\/s11704-024-40555-y"},{"key":"e_1_3_1_56_2","unstructured":"Fangyuan Xu Weijia Shi and Eunsol Choi. 2023. RECOMP: Improving retrieval-augmented LMs with compression and selective augmentation. arXiv:2310.04408. Retrieved from https:\/\/arxiv.org\/abs\/2310.04408"},{"key":"e_1_3_1_57_2","unstructured":"Yilong Xu Jinhua Gao Xiaoming Yu Baolong Bi Huawei Shen and Xueqi Cheng. 2024. ALiiCE: Evaluating positional fine-grained citation generation. arXiv:2406.13375. Retrieved from https:\/\/arxiv.org\/abs\/2406.13375"},{"key":"e_1_3_1_58_2","series-title":"Association for Computational Linguistics","doi-asserted-by":"crossref","first-page":"5364","DOI":"10.18653\/v1\/2023.emnlp-main.326","volume-title":"Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing (EMNLP \u201923)","author":"Yang Haoyan","year":"2023","unstructured":"Haoyan Yang, Zhitao Li, Yong Zhang, Jianzong Wang, Ning Cheng, Ming Li, and Jing Xiao. 2023. PRCA: Fitting black-box large language models for retrieval question answering via pluggable reward-driven contextual adapter. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing (EMNLP \u201923). Association for Computational Linguistics, 5364\u20135375."},{"key":"e_1_3_1_59_2","first-page":"7170","volume-title":"Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP \u201920)","author":"Ye Deming","year":"2020","unstructured":"Deming Ye, Yankai Lin, Jiaju Du, Zhenghao Liu, Peng Li, Maosong Sun, and Zhiyuan Liu. 2020. Coreferential reasoning learning for language representation. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP \u201920), 7170\u20137186."},{"key":"e_1_3_1_60_2","doi-asserted-by":"publisher","DOI":"10.1145\/3397271.3401323"},{"key":"e_1_3_1_61_2","unstructured":"Wenhao Yu Hongming Zhang Xiaoman Pan Kaixin Ma Hongwei Wang and Dong Yu. 2023. Chain-of-Note: Enhancing robustness in retrieval-augmented language models. arXiv:2311.09210. Retrieved from https:\/\/arxiv.org\/abs\/2311.09210"},{"key":"e_1_3_1_62_2","doi-asserted-by":"publisher","DOI":"10.1145\/984321.984322"},{"key":"e_1_3_1_63_2","doi-asserted-by":"crossref","first-page":"170","DOI":"10.1145\/3589335.3648314","volume-title":"Companion Proceedings of the ACM on Web Conference 2024","author":"Zhang Chao","year":"2024","unstructured":"Chao Zhang, Shiwei Wu, Haoxin Zhang, Tong Xu, Yan Gao, Yao Hu, and Enhong Chen. 2024. NoteLLM: A retrievable large language model for note recommendation. In Companion Proceedings of the ACM on Web Conference 2024, 170\u2013179."},{"key":"e_1_3_1_64_2","unstructured":"Peitian Zhang Shitao Xiao Zheng Liu Zhicheng Dou and Jian-Yun Nie. 2023. Retrieve anything to augment large language models. arXiv:2310.07554. Retrieved from https:\/\/arxiv.org\/abs\/2310.07554"},{"key":"e_1_3_1_65_2","unstructured":"Weijia Zhang Mohammad Aliannejadi Yifei Yuan Jiahuan Pei Jia-Hong Huang and Evangelos Kanoulas. 2024. Towards fine-grained citation evaluation in generated text: A comparative analysis of faithfulness metrics. arXiv:2406.15264. Retrieved from https:\/\/arxiv.org\/abs\/2406.15264"},{"key":"e_1_3_1_66_2","doi-asserted-by":"publisher","DOI":"10.1145\/3624918.3625329"}],"container-title":["ACM Transactions on Information Systems"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3701228","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3701228","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,19]],"date-time":"2025-06-19T01:17:17Z","timestamp":1750295837000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3701228"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,1,21]]},"references-count":65,"journal-issue":{"issue":"2","published-print":{"date-parts":[[2025,3,31]]}},"alternative-id":["10.1145\/3701228"],"URL":"https:\/\/doi.org\/10.1145\/3701228","relation":{},"ISSN":["1046-8188","1558-2868"],"issn-type":[{"value":"1046-8188","type":"print"},{"value":"1558-2868","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,1,21]]},"assertion":[{"value":"2024-02-03","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2024-10-06","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-01-21","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}