{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,24]],"date-time":"2026-07-24T13:56:09Z","timestamp":1784901369224,"version":"3.55.0"},"publisher-location":"New York, NY, USA","reference-count":83,"publisher":"ACM","license":[{"start":{"date-parts":[[2023,10,21]],"date-time":"2023-10-21T00:00:00Z","timestamp":1697846400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2023,10,21]]},"DOI":"10.1145\/3583780.3614905","type":"proceedings-article","created":{"date-parts":[[2023,10,21]],"date-time":"2023-10-21T07:45:26Z","timestamp":1697874326000},"page":"245-255","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":85,"title":["Hallucination Detection: Robustly Discerning Reliable Answers in Large Language Models"],"prefix":"10.1145","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-4381-486X","authenticated-orcid":false,"given":"Yuyan","family":"Chen","sequence":"first","affiliation":[{"name":"Shanghai Key Laboratory of Data Science &amp; School of Computer Science, Fudan University, Shanghai, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-5821-7267","authenticated-orcid":false,"given":"Qiang","family":"Fu","sequence":"additional","affiliation":[{"name":"Microsoft, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0009-1321-3850","authenticated-orcid":false,"given":"Yichen","family":"Yuan","sequence":"additional","affiliation":[{"name":"Shanghai Key Laboratory of Data Science, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-7688-5381","authenticated-orcid":false,"given":"Zhihao","family":"Wen","sequence":"additional","affiliation":[{"name":"Singapore Management University, Singapore, Singapore"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-5653-1626","authenticated-orcid":false,"given":"Ge","family":"Fan","sequence":"additional","affiliation":[{"name":"Tencent, Shenzhen, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-8755-8941","authenticated-orcid":false,"given":"Dayiheng","family":"Liu","sequence":"additional","affiliation":[{"name":"DAMO Academy, Hangzhou, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-2355-288X","authenticated-orcid":false,"given":"Dongmei","family":"Zhang","sequence":"additional","affiliation":[{"name":"Microsoft, Shanghai, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-8403-9591","authenticated-orcid":false,"given":"Zhixu","family":"Li","sequence":"additional","affiliation":[{"name":"Shanghai Key Laboratory of Data Science &amp; School of Computer Science, Fudan University &amp; Fudan-Aishu Cognitive Intelligence Joint Research Center, Shanghai, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-9230-2799","authenticated-orcid":false,"given":"Yanghua","family":"Xiao","sequence":"additional","affiliation":[{"name":"Shanghai Key Laboratory of Data Science &amp; School of Computer Science, Fudan University &amp; Fudan-Aishu Cognitive Intelligence Joint Research Center, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2023,10,21]]},"reference":[{"key":"e_1_3_2_1_1_1","volume-title":"Can we trust the evaluation on ChatGPT? arXiv preprint arXiv:2303.12767","author":"Aiyappa Rachith","year":"2023","unstructured":"Rachith Aiyappa , Jisun An , Haewoon Kwak , and Yong-Yeol Ahn . 2023. Can we trust the evaluation on ChatGPT? arXiv preprint arXiv:2303.12767 ( 2023 ). Rachith Aiyappa, Jisun An, Haewoon Kwak, and Yong-Yeol Ahn. 2023. Can we trust the evaluation on ChatGPT? arXiv preprint arXiv:2303.12767 (2023)."},{"key":"e_1_3_2_1_2_1","volume-title":"The Internal State of an LLM Knows When its Lying. arXiv preprint arXiv:2304.13734","author":"Azaria Amos","year":"2023","unstructured":"Amos Azaria and Tom Mitchell . 2023. The Internal State of an LLM Knows When its Lying. arXiv preprint arXiv:2304.13734 ( 2023 ). Amos Azaria and Tom Mitchell. 2023. The Internal State of an LLM Knows When its Lying. arXiv preprint arXiv:2304.13734 (2023)."},{"key":"e_1_3_2_1_3_1","doi-asserted-by":"crossref","unstructured":"Yejin Bang Samuel Cahyawijaya Nayeon Lee Wenliang Dai Dan Su Bryan Wilie Holy Lovenia Ziwei Ji Tiezheng Yu Willy Chung etal 2023. A multitask multilingual multimodal evaluation of chatgpt on reasoning hallucination and interactivity. arXiv preprint arXiv:2302.04023 (2023). Yejin Bang Samuel Cahyawijaya Nayeon Lee Wenliang Dai Dan Su Bryan Wilie Holy Lovenia Ziwei Ji Tiezheng Yu Willy Chung et al. 2023. A multitask multilingual multimodal evaluation of chatgpt on reasoning hallucination and interactivity. arXiv preprint arXiv:2302.04023 (2023).","DOI":"10.18653\/v1\/2023.ijcnlp-main.45"},{"key":"e_1_3_2_1_4_1","volume-title":"Scheduled sampling for sequence prediction with recurrent neural networks. Advances in neural information processing systems","author":"Bengio Samy","year":"2015","unstructured":"Samy Bengio , Oriol Vinyals , Navdeep Jaitly , and Noam Shazeer . 2015. Scheduled sampling for sequence prediction with recurrent neural networks. Advances in neural information processing systems , Vol. 28 ( 2015 ). Samy Bengio, Oriol Vinyals, Navdeep Jaitly, and Noam Shazeer. 2015. Scheduled sampling for sequence prediction with recurrent neural networks. Advances in neural information processing systems, Vol. 28 (2015)."},{"key":"e_1_3_2_1_5_1","volume-title":"A categorical archive of ChatGPT failures. arXiv preprint arXiv:2302.03494","author":"Borji Ali","year":"2023","unstructured":"Ali Borji . 2023. A categorical archive of ChatGPT failures. arXiv preprint arXiv:2302.03494 ( 2023 ). Ali Borji. 2023. A categorical archive of ChatGPT failures. arXiv preprint arXiv:2302.03494 (2023)."},{"key":"e_1_3_2_1_6_1","unstructured":"Tom Brown Benjamin Mann Nick Ryder Melanie Subbiah Jared D Kaplan Prafulla Dhariwal Arvind Neelakantan Pranav Shyam Girish Sastry Amanda Askell etal 2020. Language models are few-shot learners. Advances in neural information processing systems Vol. 33 (2020) 1877--1901. Tom Brown Benjamin Mann Nick Ryder Melanie Subbiah Jared D Kaplan Prafulla Dhariwal Arvind Neelakantan Pranav Shyam Girish Sastry Amanda Askell et al. 2020. Language models are few-shot learners. Advances in neural information processing systems Vol. 33 (2020) 1877--1901."},{"key":"e_1_3_2_1_7_1","volume-title":"Could a large language model be conscious? arXiv preprint arXiv:2303.07103","author":"Chalmers David J","year":"2023","unstructured":"David J Chalmers . 2023. Could a large language model be conscious? arXiv preprint arXiv:2303.07103 ( 2023 ). David J Chalmers. 2023. Could a large language model be conscious? arXiv preprint arXiv:2303.07103 (2023)."},{"key":"e_1_3_2_1_8_1","volume-title":"Improving faithfulness in abstractive summarization with contrast candidate generation and selection. arXiv preprint arXiv:2104.09061","author":"Chen Sihao","year":"2021","unstructured":"Sihao Chen , Fan Zhang , Kazoo Sone , and Dan Roth . 2021. Improving faithfulness in abstractive summarization with contrast candidate generation and selection. arXiv preprint arXiv:2104.09061 ( 2021 ). Sihao Chen, Fan Zhang, Kazoo Sone, and Dan Roth. 2021. Improving faithfulness in abstractive summarization with contrast candidate generation and selection. arXiv preprint arXiv:2104.09061 (2021)."},{"key":"e_1_3_2_1_9_1","unstructured":"Yongcong Chen Ting Zeng Xiaoyi Qian Jun Zhang and Xinyue Chen. [n. d.]. Apreliminary STUDY ON THE CAPABILITY BOUNDARY OF LLM AND A NEW IMPLEMENTATION APPROACH FOR AGI. ( [n. d.]). Yongcong Chen Ting Zeng Xiaoyi Qian Jun Zhang and Xinyue Chen. [n. d.]. Apreliminary STUDY ON THE CAPABILITY BOUNDARY OF LLM AND A NEW IMPLEMENTATION APPROACH FOR AGI. ( [n. d.])."},{"key":"e_1_3_2_1_10_1","volume-title":"Can Large Language Models Be an Alternative to Human Evaluations? arXiv preprint arXiv:2305.01937","author":"Chiang Cheng-Han","year":"2023","unstructured":"Cheng-Han Chiang and Hung-yi Lee. 2023. Can Large Language Models Be an Alternative to Human Evaluations? arXiv preprint arXiv:2305.01937 ( 2023 ). Cheng-Han Chiang and Hung-yi Lee. 2023. Can Large Language Models Be an Alternative to Human Evaluations? arXiv preprint arXiv:2305.01937 (2023)."},{"key":"e_1_3_2_1_11_1","volume-title":"QuAC: Question answering in context. arXiv preprint arXiv:1808.07036","author":"Choi Eunsol","year":"2018","unstructured":"Eunsol Choi , He He , Mohit Iyyer , Mark Yatskar , Wen-tau Yih, Yejin Choi , Percy Liang , and Luke Zettlemoyer . 2018. QuAC: Question answering in context. arXiv preprint arXiv:1808.07036 ( 2018 ). Eunsol Choi, He He, Mohit Iyyer, Mark Yatskar, Wen-tau Yih, Yejin Choi, Percy Liang, and Luke Zettlemoyer. 2018. QuAC: Question answering in context. arXiv preprint arXiv:1808.07036 (2018)."},{"key":"e_1_3_2_1_12_1","volume-title":"Manning","author":"Clark Kevin","year":"2020","unstructured":"Kevin Clark , Minh-Thang Luong , Quoc V. Le , and Christopher D . Manning . 2020 . ELECTRA : Pre-training Text Encoders as Discriminators Rather Than Generators . https:\/\/doi.org\/10.48550\/ARXIV.2003.10555 10.48550\/ARXIV.2003.10555 Kevin Clark, Minh-Thang Luong, Quoc V. Le, and Christopher D. Manning. 2020. ELECTRA: Pre-training Text Encoders as Discriminators Rather Than Generators. https:\/\/doi.org\/10.48550\/ARXIV.2003.10555"},{"key":"e_1_3_2_1_13_1","volume-title":"Sentence Similarity Even Better. arXiv preprint arXiv:2212.08597","author":"Dale David","year":"2022","unstructured":"David Dale , Elena Voita , Lo\"ic Barrault, and Marta R Costa-juss\u00e0. 2022. Detecting and Mitigating Hallucinations in Machine Translation: Model Internal Workings Alone Do Well , Sentence Similarity Even Better. arXiv preprint arXiv:2212.08597 ( 2022 ). David Dale, Elena Voita, Lo\"ic Barrault, and Marta R Costa-juss\u00e0. 2022. Detecting and Mitigating Hallucinations in Machine Translation: Model Internal Workings Alone Do Well, Sentence Similarity Even Better. arXiv preprint arXiv:2212.08597 (2022)."},{"key":"e_1_3_2_1_14_1","volume-title":"BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. https:\/\/doi.org\/10.48550\/ARXIV.1810.04805","author":"Devlin Jacob","year":"2018","unstructured":"Jacob Devlin , Ming-Wei Chang , Kenton Lee , and Kristina Toutanova . 2018 . BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. https:\/\/doi.org\/10.48550\/ARXIV.1810.04805 10.48550\/ARXIV.1810.04805 Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. https:\/\/doi.org\/10.48550\/ARXIV.1810.04805"},{"key":"e_1_3_2_1_15_1","volume-title":"Handling divergent reference texts when evaluating table-to-text generation. arXiv preprint arXiv:1906.01081","author":"Dhingra Bhuwan","year":"2019","unstructured":"Bhuwan Dhingra , Manaal Faruqui , Ankur Parikh , Ming-Wei Chang , Dipanjan Das , and William W Cohen . 2019. Handling divergent reference texts when evaluating table-to-text generation. arXiv preprint arXiv:1906.01081 ( 2019 ). Bhuwan Dhingra, Manaal Faruqui, Ankur Parikh, Ming-Wei Chang, Dipanjan Das, and William W Cohen. 2019. Handling divergent reference texts when evaluating table-to-text generation. arXiv preprint arXiv:1906.01081 (2019)."},{"key":"e_1_3_2_1_16_1","volume-title":"Can AI language models replace human participants? Trends in Cognitive Sciences","author":"Dillion Danica","year":"2023","unstructured":"Danica Dillion , Niket Tandon , Yuling Gu , and Kurt Gray . 2023. Can AI language models replace human participants? Trends in Cognitive Sciences ( 2023 ). Danica Dillion, Niket Tandon, Yuling Gu, and Kurt Gray. 2023. Can AI language models replace human participants? Trends in Cognitive Sciences (2023)."},{"key":"e_1_3_2_1_17_1","volume-title":"FEQA: A question answering evaluation framework for faithfulness assessment in abstractive summarization. arXiv preprint arXiv:2005.03754","author":"Durmus Esin","year":"2020","unstructured":"Esin Durmus , He He , and Mona Diab . 2020 . FEQA: A question answering evaluation framework for faithfulness assessment in abstractive summarization. arXiv preprint arXiv:2005.03754 (2020). Esin Durmus, He He, and Mona Diab. 2020. FEQA: A question answering evaluation framework for faithfulness assessment in abstractive summarization. arXiv preprint arXiv:2005.03754 (2020)."},{"key":"e_1_3_2_1_18_1","volume-title":"Evaluating groundedness in dialogue systems: The begin benchmark. arXiv preprint arXiv:2105.00071","author":"Dziri Nouha","year":"2021","unstructured":"Nouha Dziri , Hannah Rashkin , Tal Linzen , and David Reitter . 2021. Evaluating groundedness in dialogue systems: The begin benchmark. arXiv preprint arXiv:2105.00071 ( 2021 ). Nouha Dziri, Hannah Rashkin, Tal Linzen, and David Reitter. 2021. Evaluating groundedness in dialogue systems: The begin benchmark. arXiv preprint arXiv:2105.00071 (2021)."},{"key":"e_1_3_2_1_19_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/P19-1213"},{"key":"e_1_3_2_1_20_1","volume-title":"Controlled hallucinations: Learning to generate faithfully from noisy data. arXiv preprint arXiv:2010.05873","author":"Filippova Katja","year":"2020","unstructured":"Katja Filippova . 2020. Controlled hallucinations: Learning to generate faithfully from noisy data. arXiv preprint arXiv:2010.05873 ( 2020 ). Katja Filippova. 2020. Controlled hallucinations: Learning to generate faithfully from noisy data. arXiv preprint arXiv:2010.05873 (2020)."},{"key":"e_1_3_2_1_21_1","volume-title":"Evaluating factuality in generation with dependency-level entailment. arXiv preprint arXiv:2010.05478","author":"Goyal Tanya","year":"2020","unstructured":"Tanya Goyal and Greg Durrett . 2020. Evaluating factuality in generation with dependency-level entailment. arXiv preprint arXiv:2010.05478 ( 2020 ). Tanya Goyal and Greg Durrett. 2020. Evaluating factuality in generation with dependency-level entailment. arXiv preprint arXiv:2010.05478 (2020)."},{"key":"e_1_3_2_1_22_1","volume-title":"Union: An unreferenced metric for evaluating open-ended story generation. arXiv preprint arXiv:2009.07602","author":"Guan Jian","year":"2020","unstructured":"Jian Guan and Minlie Huang . 2020 . Union: An unreferenced metric for evaluating open-ended story generation. arXiv preprint arXiv:2009.07602 (2020). Jian Guan and Minlie Huang. 2020. Union: An unreferenced metric for evaluating open-ended story generation. arXiv preprint arXiv:2009.07602 (2020)."},{"key":"#cr-split#-e_1_3_2_1_23_1.1","unstructured":"Pengcheng He Xiaodong Liu Jianfeng Gao and Weizhu Chen. 2020. DeBERTa: Decoding-enhanced BERT with Disentangled Attention. https:\/\/doi.org\/10.48550\/ARXIV.2006.03654 10.48550\/ARXIV.2006.03654"},{"key":"#cr-split#-e_1_3_2_1_23_1.2","unstructured":"Pengcheng He Xiaodong Liu Jianfeng Gao and Weizhu Chen. 2020. DeBERTa: Decoding-enhanced BERT with Disentangled Attention. https:\/\/doi.org\/10.48550\/ARXIV.2006.03654"},{"key":"e_1_3_2_1_24_1","volume-title":"Exposure Bias versus Self-Recovery: Are Distortions Really Incremental for Autoregressive Text Generation? arXiv preprint arXiv:1905.10617","author":"He Tianxing","year":"2019","unstructured":"Tianxing He , Jingzhao Zhang , Zhiming Zhou , and James Glass . 2019. Exposure Bias versus Self-Recovery: Are Distortions Really Incremental for Autoregressive Text Generation? arXiv preprint arXiv:1905.10617 ( 2019 ). Tianxing He, Jingzhao Zhang, Zhiming Zhou, and James Glass. 2019. Exposure Bias versus Self-Recovery: Are Distortions Really Incremental for Autoregressive Text Generation? arXiv preprint arXiv:1905.10617 (2019)."},{"key":"e_1_3_2_1_25_1","doi-asserted-by":"crossref","unstructured":"Wei He Kai Liu Jing Liu Yajuan Lyu Shiqi Zhao Xinyan Xiao Yuan Liu Yizhong Wang Hua Wu Qiaoqiao She etal 2017. Dureader: a chinese machine reading comprehension dataset from real-world applications. arXiv preprint arXiv:1711.05073 (2017). Wei He Kai Liu Jing Liu Yajuan Lyu Shiqi Zhao Xinyan Xiao Yuan Liu Yizhong Wang Hua Wu Qiaoqiao She et al. 2017. Dureader: a chinese machine reading comprehension dataset from real-world applications. arXiv preprint arXiv:1711.05073 (2017).","DOI":"10.18653\/v1\/W18-2605"},{"key":"e_1_3_2_1_26_1","volume-title":"TRUE: Re-evaluating factual consistency evaluation. arXiv preprint arXiv:2204.04991","author":"Honovich Or","year":"2022","unstructured":"Or Honovich , Roee Aharoni , Jonathan Herzig , Hagai Taitelbaum , Doron Kukliansy , Vered Cohen , Thomas Scialom , Idan Szpektor , Avinatan Hassidim , and Yossi Matias . 2022 . TRUE: Re-evaluating factual consistency evaluation. arXiv preprint arXiv:2204.04991 (2022). Or Honovich, Roee Aharoni, Jonathan Herzig, Hagai Taitelbaum, Doron Kukliansy, Vered Cohen, Thomas Scialom, Idan Szpektor, Avinatan Hassidim, and Yossi Matias. 2022. TRUE: Re-evaluating factual consistency evaluation. arXiv preprint arXiv:2204.04991 (2022)."},{"key":"e_1_3_2_1_27_1","volume-title":"Evaluating Factual Consistency in Knowledge-Grounded Dialogues via Question Generation and Question Answering. arXiv preprint arXiv:2104.08202","author":"Honovich Or","year":"2021","unstructured":"Or Honovich , Leshem Choshen , Roee Aharoni , Ella Neeman , Idan Szpektor , and Omri Abend . 2021. $Q^2$ : Evaluating Factual Consistency in Knowledge-Grounded Dialogues via Question Generation and Question Answering. arXiv preprint arXiv:2104.08202 ( 2021 ). Or Honovich, Leshem Choshen, Roee Aharoni, Ella Neeman, Idan Szpektor, and Omri Abend. 2021. $Q^2$: Evaluating Factual Consistency in Knowledge-Grounded Dialogues via Question Generation and Question Answering. arXiv preprint arXiv:2104.08202 (2021)."},{"key":"e_1_3_2_1_28_1","volume-title":"The factual inconsistency problem in abstractive text summarization: A survey. arXiv preprint arXiv:2104.14839","author":"Huang Yichong","year":"2021","unstructured":"Yichong Huang , Xiachong Feng , Xiaocheng Feng , and Bing Qin . 2021. The factual inconsistency problem in abstractive text summarization: A survey. arXiv preprint arXiv:2104.14839 ( 2021 ). Yichong Huang, Xiachong Feng, Xiaocheng Feng, and Bing Qin. 2021. The factual inconsistency problem in abstractive text summarization: A survey. arXiv preprint arXiv:2104.14839 (2021)."},{"key":"e_1_3_2_1_29_1","doi-asserted-by":"publisher","DOI":"10.1145\/3571730"},{"key":"e_1_3_2_1_30_1","volume-title":"Graphologue: Exploring Large Language Model Responses with Interactive Diagrams. arXiv preprint arXiv:2305.11473","author":"Jiang Peiling","year":"2023","unstructured":"Peiling Jiang , Jude Rayan , Steven P Dow , and Haijun Xia . 2023 . Graphologue: Exploring Large Language Model Responses with Interactive Diagrams. arXiv preprint arXiv:2305.11473 (2023). Peiling Jiang, Jude Rayan, Steven P Dow, and Haijun Xia. 2023. Graphologue: Exploring Large Language Model Responses with Interactive Diagrams. arXiv preprint arXiv:2305.11473 (2023)."},{"key":"e_1_3_2_1_31_1","doi-asserted-by":"crossref","unstructured":"Douglas Johnson Rachel Goodman J Patrinely Cosby Stone Eli Zimmerman Rebecca Donald Sam Chang Sean Berkowitz Avni Finn Eiman Jahangir etal 2023. Assessing the accuracy and reliability of AI-generated medical responses: an evaluation of the Chat-gpt model. (2023). Douglas Johnson Rachel Goodman J Patrinely Cosby Stone Eli Zimmerman Rebecca Donald Sam Chang Sean Berkowitz Avni Finn Eiman Jahangir et al. 2023. Assessing the accuracy and reliability of AI-generated medical responses: an evaluation of the Chat-gpt model. (2023).","DOI":"10.21203\/rs.3.rs-2566942\/v1"},{"key":"e_1_3_2_1_32_1","volume-title":"Triviaqa: A large scale distantly supervised challenge dataset for reading comprehension. arXiv preprint arXiv:1705.03551","author":"Joshi Mandar","year":"2017","unstructured":"Mandar Joshi , Eunsol Choi , Daniel S Weld , and Luke Zettlemoyer . 2017 . Triviaqa: A large scale distantly supervised challenge dataset for reading comprehension. arXiv preprint arXiv:1705.03551 (2017). Mandar Joshi, Eunsol Choi, Daniel S Weld, and Luke Zettlemoyer. 2017. Triviaqa: A large scale distantly supervised challenge dataset for reading comprehension. arXiv preprint arXiv:1705.03551 (2017)."},{"key":"e_1_3_2_1_33_1","volume-title":"How Secure is Code Generated by ChatGPT? arXiv preprint arXiv:2304.09655","author":"Khoury Rapha\u00ebl","year":"2023","unstructured":"Rapha\u00ebl Khoury , Anderson R Avila , Jacob Brunelle , and Baba Mamadou Camara . 2023. How Secure is Code Generated by ChatGPT? arXiv preprint arXiv:2304.09655 ( 2023 ). Rapha\u00ebl Khoury, Anderson R Avila, Jacob Brunelle, and Baba Mamadou Camara. 2023. How Secure is Code Generated by ChatGPT? arXiv preprint arXiv:2304.09655 (2023)."},{"key":"e_1_3_2_1_34_1","volume-title":"A sliding-window approach to automatic creation of meeting minutes. arXiv preprint arXiv:2104.12324","author":"Koay Jia Jin","year":"2021","unstructured":"Jia Jin Koay , Alexander Roustai , Xiaojin Dai , and Fei Liu . 2021. A sliding-window approach to automatic creation of meeting minutes. arXiv preprint arXiv:2104.12324 ( 2021 ). Jia Jin Koay, Alexander Roustai, Xiaojin Dai, and Fei Liu. 2021. A sliding-window approach to automatic creation of meeting minutes. arXiv preprint arXiv:2104.12324 (2021)."},{"key":"e_1_3_2_1_35_1","doi-asserted-by":"publisher","DOI":"10.1162\/tacl_a_00453"},{"key":"e_1_3_2_1_36_1","first-page":"34586","article-title":"Factuality enhanced language models for open-ended text generation","volume":"35","author":"Lee Nayeon","year":"2022","unstructured":"Nayeon Lee , Wei Ping , Peng Xu , Mostofa Patwary , Pascale N Fung , Mohammad Shoeybi , and Bryan Catanzaro . 2022 . Factuality enhanced language models for open-ended text generation . Advances in Neural Information Processing Systems , Vol. 35 (2022), 34586 -- 34599 . Nayeon Lee, Wei Ping, Peng Xu, Mostofa Patwary, Pascale N Fung, Mohammad Shoeybi, and Bryan Catanzaro. 2022. Factuality enhanced language models for open-ended text generation. Advances in Neural Information Processing Systems, Vol. 35 (2022), 34586--34599.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_2_1_37_1","volume-title":"BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and Comprehension. https:\/\/doi.org\/10.48550\/ARXIV.1910.13461","author":"Lewis Mike","year":"2019","unstructured":"Mike Lewis , Yinhan Liu , Naman Goyal , Marjan Ghazvininejad , Abdelrahman Mohamed , Omer Levy , Ves Stoyanov , and Luke Zettlemoyer . 2019 . BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and Comprehension. https:\/\/doi.org\/10.48550\/ARXIV.1910.13461 10.48550\/ARXIV.1910.13461 Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Ves Stoyanov, and Luke Zettlemoyer. 2019. BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and Comprehension. https:\/\/doi.org\/10.48550\/ARXIV.1910.13461"},{"key":"e_1_3_2_1_38_1","volume-title":"A diversity-promoting objective function for neural conversation models. arXiv preprint arXiv:1510.03055","author":"Li Jiwei","year":"2015","unstructured":"Jiwei Li , Michel Galley , Chris Brockett , Jianfeng Gao , and Bill Dolan . 2015. A diversity-promoting objective function for neural conversation models. arXiv preprint arXiv:1510.03055 ( 2015 ). Jiwei Li, Michel Galley, Chris Brockett, Jianfeng Gao, and Bill Dolan. 2015. A diversity-promoting objective function for neural conversation models. arXiv preprint arXiv:1510.03055 (2015)."},{"key":"e_1_3_2_1_39_1","volume-title":"Rouge: A package for automatic evaluation of summaries. In Text summarization branches out. 74--81.","author":"Lin Chin-Yew","year":"2004","unstructured":"Chin-Yew Lin . 2004 . Rouge: A package for automatic evaluation of summaries. In Text summarization branches out. 74--81. Chin-Yew Lin. 2004. Rouge: A package for automatic evaluation of summaries. In Text summarization branches out. 74--81."},{"key":"e_1_3_2_1_40_1","volume-title":"How not to evaluate your dialogue system: An empirical study of unsupervised evaluation metrics for dialogue response generation. arXiv preprint arXiv:1603.08023","author":"Liu Chia-Wei","year":"2016","unstructured":"Chia-Wei Liu , Ryan Lowe , Iulian V Serban , Michael Noseworthy , Laurent Charlin , and Joelle Pineau . 2016. How not to evaluate your dialogue system: An empirical study of unsupervised evaluation metrics for dialogue response generation. arXiv preprint arXiv:1603.08023 ( 2016 ). Chia-Wei Liu, Ryan Lowe, Iulian V Serban, Michael Noseworthy, Laurent Charlin, and Joelle Pineau. 2016. How not to evaluate your dialogue system: An empirical study of unsupervised evaluation metrics for dialogue response generation. arXiv preprint arXiv:1603.08023 (2016)."},{"key":"e_1_3_2_1_41_1","volume-title":"2023 b. Evaluating the logical reasoning ability of chatgpt and gpt-4. arXiv preprint arXiv:2304.03439","author":"Liu Hanmeng","year":"2023","unstructured":"Hanmeng Liu , Ruoxi Ning , Zhiyang Teng , Jian Liu , Qiji Zhou , and Yue Zhang . 2023 b. Evaluating the logical reasoning ability of chatgpt and gpt-4. arXiv preprint arXiv:2304.03439 ( 2023 ). Hanmeng Liu, Ruoxi Ning, Zhiyang Teng, Jian Liu, Qiji Zhou, and Yue Zhang. 2023 b. Evaluating the logical reasoning ability of chatgpt and gpt-4. arXiv preprint arXiv:2304.03439 (2023)."},{"key":"e_1_3_2_1_42_1","volume-title":"A token-level reference-free hallucination detection benchmark for free-form text generation. arXiv preprint arXiv:2104.08704","author":"Liu Tianyu","year":"2021","unstructured":"Tianyu Liu , Yizhe Zhang , Chris Brockett , Yi Mao , Zhifang Sui , Weizhu Chen , and Bill Dolan . 2021. A token-level reference-free hallucination detection benchmark for free-form text generation. arXiv preprint arXiv:2104.08704 ( 2021 ). Tianyu Liu, Yizhe Zhang, Chris Brockett, Yi Mao, Zhifang Sui, Weizhu Chen, and Bill Dolan. 2021. A token-level reference-free hallucination detection benchmark for free-form text generation. arXiv preprint arXiv:2104.08704 (2021)."},{"key":"e_1_3_2_1_43_1","volume-title":"2023 a. Summary of chatgpt\/gpt-4 research and perspective towards the future of large language models. arXiv preprint arXiv:2304.01852","author":"Liu Yiheng","year":"2023","unstructured":"Yiheng Liu , Tianle Han , Siyuan Ma , Jiayue Zhang , Yuanyuan Yang , Jiaming Tian , Hao He , Antong Li , Mengshen He , Zhengliang Liu , 2023 a. Summary of chatgpt\/gpt-4 research and perspective towards the future of large language models. arXiv preprint arXiv:2304.01852 ( 2023 ). Yiheng Liu, Tianle Han, Siyuan Ma, Jiayue Zhang, Yuanyuan Yang, Jiaming Tian, Hao He, Antong Li, Mengshen He, Zhengliang Liu, et al. 2023 a. Summary of chatgpt\/gpt-4 research and perspective towards the future of large language models. arXiv preprint arXiv:2304.01852 (2023)."},{"key":"#cr-split#-e_1_3_2_1_44_1.1","unstructured":"Yinhan Liu Myle Ott Naman Goyal Jingfei Du Mandar Joshi Danqi Chen Omer Levy Mike Lewis Luke Zettlemoyer and Veselin Stoyanov. 2019. RoBERTa: A Robustly Optimized BERT Pretraining Approach. https:\/\/doi.org\/10.48550\/ARXIV.1907.11692 10.48550\/ARXIV.1907.11692"},{"key":"#cr-split#-e_1_3_2_1_44_1.2","unstructured":"Yinhan Liu Myle Ott Naman Goyal Jingfei Du Mandar Joshi Danqi Chen Omer Levy Mike Lewis Luke Zettlemoyer and Veselin Stoyanov. 2019. RoBERTa: A Robustly Optimized BERT Pretraining Approach. https:\/\/doi.org\/10.48550\/ARXIV.1907.11692"},{"key":"e_1_3_2_1_45_1","volume-title":"Entity-based knowledge conflicts in question answering. arXiv preprint arXiv:2109.05052","author":"Longpre Shayne","year":"2021","unstructured":"Shayne Longpre , Kartik Perisetla , Anthony Chen , Nikhil Ramesh , Chris DuBois , and Sameer Singh . 2021. Entity-based knowledge conflicts in question answering. arXiv preprint arXiv:2109.05052 ( 2021 ). Shayne Longpre, Kartik Perisetla, Anthony Chen, Nikhil Ramesh, Chris DuBois, and Sameer Singh. 2021. Entity-based knowledge conflicts in question answering. arXiv preprint arXiv:2109.05052 (2021)."},{"key":"e_1_3_2_1_46_1","volume-title":"Language models as few-shot learner for task-oriented dialogue systems. arXiv preprint arXiv:2008.06239","author":"Madotto Andrea","year":"2020","unstructured":"Andrea Madotto , Zihan Liu , Zhaojiang Lin , and Pascale Fung . 2020. Language models as few-shot learner for task-oriented dialogue systems. arXiv preprint arXiv:2008.06239 ( 2020 ). Andrea Madotto, Zihan Liu, Zhaojiang Lin, and Pascale Fung. 2020. Language models as few-shot learner for task-oriented dialogue systems. arXiv preprint arXiv:2008.06239 (2020)."},{"key":"e_1_3_2_1_47_1","volume-title":"MS MARCO: A human generated machine reading comprehension dataset. choice","author":"Nguyen Tri","year":"2016","unstructured":"Tri Nguyen , Mir Rosenberg , Xia Song , Jianfeng Gao , Saurabh Tiwary , Rangan Majumder , and Li Deng . 2016 . MS MARCO: A human generated machine reading comprehension dataset. choice , Vol. 2640 (2016), 660. Tri Nguyen, Mir Rosenberg, Xia Song, Jianfeng Gao, Saurabh Tiwary, Rangan Majumder, and Li Deng. 2016. MS MARCO: A human generated machine reading comprehension dataset. choice, Vol. 2640 (2016), 660."},{"key":"e_1_3_2_1_48_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/P19-1256"},{"key":"e_1_3_2_1_49_1","volume-title":"Chatgpt versus traditional question answering for knowledge graphs: Current status and future directions towards knowledge graph chatbots. arXiv preprint arXiv:2302.06466","author":"Omar Reham","year":"2023","unstructured":"Reham Omar , Omij Mangukiya , Panos Kalnis , and Essam Mansour . 2023. Chatgpt versus traditional question answering for knowledge graphs: Current status and future directions towards knowledge graph chatbots. arXiv preprint arXiv:2302.06466 ( 2023 ). Reham Omar, Omij Mangukiya, Panos Kalnis, and Essam Mansour. 2023. Chatgpt versus traditional question answering for knowledge graphs: Current status and future directions towards knowledge graph chatbots. arXiv preprint arXiv:2302.06466 (2023)."},{"key":"e_1_3_2_1_50_1","first-page":"27730","article-title":"Training language models to follow instructions with human feedback","volume":"35","author":"Ouyang Long","year":"2022","unstructured":"Long Ouyang , Jeffrey Wu , Xu Jiang , Diogo Almeida , Carroll Wainwright , Pamela Mishkin , Chong Zhang , Sandhini Agarwal , Katarina Slama , Alex Ray , 2022 . Training language models to follow instructions with human feedback . Advances in Neural Information Processing Systems , Vol. 35 (2022), 27730 -- 27744 . Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al. 2022. Training language models to follow instructions with human feedback. Advances in Neural Information Processing Systems, Vol. 35 (2022), 27730--27744.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_2_1_51_1","volume-title":"Understanding factuality in abstractive summarization with FRANK: A benchmark for factuality metrics. arXiv preprint arXiv:2104.13346","author":"Pagnoni Artidoro","year":"2021","unstructured":"Artidoro Pagnoni , Vidhisha Balachandran , and Yulia Tsvetkov . 2021. Understanding factuality in abstractive summarization with FRANK: A benchmark for factuality metrics. arXiv preprint arXiv:2104.13346 ( 2021 ). Artidoro Pagnoni, Vidhisha Balachandran, and Yulia Tsvetkov. 2021. Understanding factuality in abstractive summarization with FRANK: A benchmark for factuality metrics. arXiv preprint arXiv:2104.13346 (2021)."},{"key":"e_1_3_2_1_52_1","volume-title":"Proceedings of the 40th annual meeting of the Association for Computational Linguistics. 311--318","author":"Papineni Kishore","year":"2002","unstructured":"Kishore Papineni , Salim Roukos , Todd Ward , and Wei-Jing Zhu . 2002 . Bleu: a method for automatic evaluation of machine translation . In Proceedings of the 40th annual meeting of the Association for Computational Linguistics. 311--318 . Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002. Bleu: a method for automatic evaluation of machine translation. In Proceedings of the 40th annual meeting of the Association for Computational Linguistics. 311--318."},{"key":"e_1_3_2_1_53_1","unstructured":"Hyun Jin Park and Changwan Ryu. 2023. Query Augmentation Using Search Engine Results to Improve Answers Generated by Large Language Models. (2023). Hyun Jin Park and Changwan Ryu. 2023. Query Augmentation Using Search Engine Results to Improve Answers Generated by Large Language Models. (2023)."},{"key":"e_1_3_2_1_54_1","unstructured":"Baolin Peng Michel Galley Pengcheng He Hao Cheng Yujia Xie Yu Hu Qiuyuan Huang Lars Liden Zhou Yu Weizhu Chen etal 2023. Check your facts and try again: Improving large language models with external knowledge and automated feedback. arXiv preprint arXiv:2302.12813 (2023). Baolin Peng Michel Galley Pengcheng He Hao Cheng Yujia Xie Yu Hu Qiuyuan Huang Lars Liden Zhou Yu Weizhu Chen et al. 2023. Check your facts and try again: Improving large language models with external knowledge and automated feedback. arXiv preprint arXiv:2302.12813 (2023)."},{"key":"e_1_3_2_1_55_1","volume-title":"Language models as knowledge bases? arXiv preprint arXiv:1909.01066","author":"Petroni Fabio","year":"2019","unstructured":"Fabio Petroni , Tim Rockt\"aschel, Patrick Lewis , Anton Bakhtin , Yuxiang Wu , Alexander H Miller , and Sebastian Riedel . 2019. Language models as knowledge bases? arXiv preprint arXiv:1909.01066 ( 2019 ). Fabio Petroni, Tim Rockt\"aschel, Patrick Lewis, Anton Bakhtin, Yuxiang Wu, Alexander H Miller, and Sebastian Riedel. 2019. Language models as knowledge bases? arXiv preprint arXiv:1909.01066 (2019)."},{"key":"e_1_3_2_1_56_1","volume-title":"100,000 questions for machine comprehension of text. arXiv preprint arXiv:1606.05250","author":"Rajpurkar Pranav","year":"2016","unstructured":"Pranav Rajpurkar , Jian Zhang , Konstantin Lopyrev , and Percy Liang . 2016. Squad : 100,000 questions for machine comprehension of text. arXiv preprint arXiv:1606.05250 ( 2016 ). Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang. 2016. Squad: 100,000 questions for machine comprehension of text. arXiv preprint arXiv:1606.05250 (2016)."},{"key":"e_1_3_2_1_57_1","volume-title":"Sequence level training with recurrent neural networks. arXiv preprint arXiv:1511.06732","author":"Ranzato Marc'Aurelio","year":"2015","unstructured":"Marc'Aurelio Ranzato , Sumit Chopra , Michael Auli , and Wojciech Zaremba . 2015. Sequence level training with recurrent neural networks. arXiv preprint arXiv:1511.06732 ( 2015 ). Marc'Aurelio Ranzato, Sumit Chopra, Michael Auli, and Wojciech Zaremba. 2015. Sequence level training with recurrent neural networks. arXiv preprint arXiv:1511.06732 (2015)."},{"key":"e_1_3_2_1_58_1","volume-title":"Gaurav Singh Tomar, and Dipanjan Das","author":"Rashkin Hannah","year":"2021","unstructured":"Hannah Rashkin , David Reitter , Gaurav Singh Tomar, and Dipanjan Das . 2021 . Increasing faithfulness in knowledge-grounded dialogue with controllable features. arXiv preprint arXiv:2107.06963 (2021). Hannah Rashkin, David Reitter, Gaurav Singh Tomar, and Dipanjan Das. 2021. Increasing faithfulness in knowledge-grounded dialogue with controllable features. arXiv preprint arXiv:2107.06963 (2021)."},{"key":"e_1_3_2_1_59_1","volume-title":"Data-QuestEval: A referenceless metric for data-to-text semantic evaluation. arXiv preprint arXiv:2104.07555","author":"Rebuffel Cl\u00e9ment","year":"2021","unstructured":"Cl\u00e9ment Rebuffel , Thomas Scialom , Laure Soulier , Benjamin Piwowarski , Sylvain Lamprier , Jacopo Staiano , Geoffrey Scoutheeten , and Patrick Gallinari . 2021. Data-QuestEval: A referenceless metric for data-to-text semantic evaluation. arXiv preprint arXiv:2104.07555 ( 2021 ). Cl\u00e9ment Rebuffel, Thomas Scialom, Laure Soulier, Benjamin Piwowarski, Sylvain Lamprier, Jacopo Staiano, Geoffrey Scoutheeten, and Patrick Gallinari. 2021. Data-QuestEval: A referenceless metric for data-to-text semantic evaluation. arXiv preprint arXiv:2104.07555 (2021)."},{"key":"e_1_3_2_1_60_1","doi-asserted-by":"publisher","DOI":"10.1162\/tacl_a_00266"},{"key":"e_1_3_2_1_61_1","volume-title":"How much knowledge can you pack into the parameters of a language model? arXiv preprint arXiv:2002.08910","author":"Roberts Adam","year":"2020","unstructured":"Adam Roberts , Colin Raffel , and Noam Shazeer . 2020. How much knowledge can you pack into the parameters of a language model? arXiv preprint arXiv:2002.08910 ( 2020 ). Adam Roberts, Colin Raffel, and Noam Shazeer. 2020. How much knowledge can you pack into the parameters of a language model? arXiv preprint arXiv:2002.08910 (2020)."},{"key":"e_1_3_2_1_62_1","doi-asserted-by":"crossref","unstructured":"Stephen Roller Emily Dinan Naman Goyal Da Ju Mary Williamson Yinhan Liu Jing Xu Myle Ott Kurt Shuster Eric M Smith etal 2020. Recipes for building an open-domain chatbot. arXiv preprint arXiv:2004.13637 (2020). Stephen Roller Emily Dinan Naman Goyal Da Ju Mary Williamson Yinhan Liu Jing Xu Myle Ott Kurt Shuster Eric M Smith et al. 2020. Recipes for building an open-domain chatbot. arXiv preprint arXiv:2004.13637 (2020).","DOI":"10.18653\/v1\/2021.eacl-main.24"},{"key":"e_1_3_2_1_63_1","volume-title":"Rome was built in 1776: A case study on factual correctness in knowledge-grounded response generation. arXiv preprint arXiv:2110.05456","author":"Santhanam Sashank","year":"2021","unstructured":"Sashank Santhanam , Behnam Hedayatnia , Spandana Gella , Aishwarya Padmakumar , Seokhwan Kim , Yang Liu , and Dilek Hakkani-Tur . 2021. Rome was built in 1776: A case study on factual correctness in knowledge-grounded response generation. arXiv preprint arXiv:2110.05456 ( 2021 ). Sashank Santhanam, Behnam Hedayatnia, Spandana Gella, Aishwarya Padmakumar, Seokhwan Kim, Yang Liu, and Dilek Hakkani-Tur. 2021. Rome was built in 1776: A case study on factual correctness in knowledge-grounded response generation. arXiv preprint arXiv:2110.05456 (2021)."},{"key":"e_1_3_2_1_64_1","volume-title":"Francc ois Yvon, Matthias Gall\u00e9, et al.","author":"Scao Teven Le","year":"2022","unstructured":"Teven Le Scao , Angela Fan , Christopher Akiki , Ellie Pavlick , Suzana Ili\u0107 , Daniel Hesslow , Roman Castagn\u00e9 , Alexandra Sasha Luccioni , Francc ois Yvon, Matthias Gall\u00e9, et al. 2022 . Bloom : A 176b-parameter open-access multilingual language model. arXiv preprint arXiv:2211.05100 (2022). Teven Le Scao, Angela Fan, Christopher Akiki, Ellie Pavlick, Suzana Ili\u0107, Daniel Hesslow, Roman Castagn\u00e9, Alexandra Sasha Luccioni, Francc ois Yvon, Matthias Gall\u00e9, et al. 2022. Bloom: A 176b-parameter open-access multilingual language model. arXiv preprint arXiv:2211.05100 (2022)."},{"key":"e_1_3_2_1_65_1","volume-title":"Questeval: Summarization asks for fact-based evaluation. arXiv preprint arXiv:2103.12693","author":"Scialom Thomas","year":"2021","unstructured":"Thomas Scialom , Paul-Alexis Dray , Patrick Gallinari , Sylvain Lamprier , Benjamin Piwowarski , Jacopo Staiano , and Alex Wang . 2021 . Questeval: Summarization asks for fact-based evaluation. arXiv preprint arXiv:2103.12693 (2021). Thomas Scialom, Paul-Alexis Dray, Patrick Gallinari, Sylvain Lamprier, Benjamin Piwowarski, Jacopo Staiano, and Alex Wang. 2021. Questeval: Summarization asks for fact-based evaluation. arXiv preprint arXiv:2103.12693 (2021)."},{"key":"e_1_3_2_1_66_1","volume-title":"BLEURT: Learning robust metrics for text generation. arXiv preprint arXiv:2004.04696","author":"Sellam Thibault","year":"2020","unstructured":"Thibault Sellam , Dipanjan Das , and Ankur P Parikh . 2020 . BLEURT: Learning robust metrics for text generation. arXiv preprint arXiv:2004.04696 (2020). Thibault Sellam, Dipanjan Das, and Ankur P Parikh. 2020. BLEURT: Learning robust metrics for text generation. arXiv preprint arXiv:2004.04696 (2020)."},{"key":"e_1_3_2_1_67_1","unstructured":"Xinyue Shen Zeyuan Chen Michael Backes and Yang Zhang. 2023. In ChatGPT We Trust? Measuring and Characterizing the Reliability of ChatGPT. arXiv preprint arXiv:2304.08979 (2023). Xinyue Shen Zeyuan Chen Michael Backes and Yang Zhang. 2023. In ChatGPT We Trust? Measuring and Characterizing the Reliability of ChatGPT. arXiv preprint arXiv:2304.08979 (2023)."},{"key":"e_1_3_2_1_68_1","volume-title":"Retrieval augmentation reduces hallucination in conversation. arXiv preprint arXiv:2104.07567","author":"Shuster Kurt","year":"2021","unstructured":"Kurt Shuster , Spencer Poff , Moya Chen , Douwe Kiela , and Jason Weston . 2021. Retrieval augmentation reduces hallucination in conversation. arXiv preprint arXiv:2104.07567 ( 2021 ). Kurt Shuster, Spencer Poff, Moya Chen, Douwe Kiela, and Jason Weston. 2021. Retrieval augmentation reduces hallucination in conversation. arXiv preprint arXiv:2104.07567 (2021)."},{"key":"e_1_3_2_1_69_1","volume-title":"Diversifying dialogue generation with non-conversational text. arXiv preprint arXiv:2005.04346","author":"Su Hui","year":"2020","unstructured":"Hui Su , Xiaoyu Shen , Sanqiang Zhao , Xiao Zhou , Pengwei Hu , Randy Zhong , Cheng Niu , and Jie Zhou . 2020. Diversifying dialogue generation with non-conversational text. arXiv preprint arXiv:2005.04346 ( 2020 ). Hui Su, Xiaoyu Shen, Sanqiang Zhao, Xiao Zhou, Pengwei Hu, Randy Zhong, Cheng Niu, and Jie Zhou. 2020. Diversifying dialogue generation with non-conversational text. arXiv preprint arXiv:2005.04346 (2020)."},{"key":"e_1_3_2_1_70_1","volume-title":"Sticking to the facts: Confident decoding for faithful data-to-text generation. arXiv preprint arXiv:1910.08684","author":"Tian Ran","year":"2019","unstructured":"Ran Tian , Shashi Narayan , Thibault Sellam , and Ankur P Parikh . 2019. Sticking to the facts: Confident decoding for faithful data-to-text generation. arXiv preprint arXiv:1910.08684 ( 2019 ). Ran Tian, Shashi Narayan, Thibault Sellam, and Ankur P Parikh. 2019. Sticking to the facts: Confident decoding for faithful data-to-text generation. arXiv preprint arXiv:1910.08684 (2019)."},{"key":"e_1_3_2_1_71_1","volume-title":"Llama: Open and efficient foundation language models. arXiv preprint arXiv:2302.13971","author":"Touvron Hugo","year":"2023","unstructured":"Hugo Touvron , Thibaut Lavril , Gautier Izacard , Xavier Martinet , Marie-Anne Lachaux , Timoth\u00e9e Lacroix , Baptiste Rozi\u00e8re , Naman Goyal , Eric Hambro , Faisal Azhar , 2023 . Llama: Open and efficient foundation language models. arXiv preprint arXiv:2302.13971 (2023). Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timoth\u00e9e Lacroix, Baptiste Rozi\u00e8re, Naman Goyal, Eric Hambro, Faisal Azhar, et al. 2023. Llama: Open and efficient foundation language models. arXiv preprint arXiv:2302.13971 (2023)."},{"key":"e_1_3_2_1_72_1","volume-title":"Newsqa: A machine comprehension dataset. arXiv preprint arXiv:1611.09830","author":"Trischler Adam","year":"2016","unstructured":"Adam Trischler , Tong Wang , Xingdi Yuan , Justin Harris , Alessandro Sordoni , Philip Bachman , and Kaheer Suleman . 2016 . Newsqa: A machine comprehension dataset. arXiv preprint arXiv:1611.09830 (2016). Adam Trischler, Tong Wang, Xingdi Yuan, Justin Harris, Alessandro Sordoni, Philip Bachman, and Kaheer Suleman. 2016. Newsqa: A machine comprehension dataset. arXiv preprint arXiv:1611.09830 (2016)."},{"key":"e_1_3_2_1_73_1","volume-title":"Asking and answering questions to evaluate the factual consistency of summaries. arXiv preprint arXiv:2004.04228","author":"Wang Alex","year":"2020","unstructured":"Alex Wang , Kyunghyun Cho , and Mike Lewis . 2020a. Asking and answering questions to evaluate the factual consistency of summaries. arXiv preprint arXiv:2004.04228 ( 2020 ). Alex Wang, Kyunghyun Cho, and Mike Lewis. 2020a. Asking and answering questions to evaluate the factual consistency of summaries. arXiv preprint arXiv:2004.04228 (2020)."},{"key":"e_1_3_2_1_74_1","unstructured":"Ben Wang and Aran Komatsuzaki. 2021. GPT-J-6B: A 6 Billion Parameter Autoregressive Language Model. https:\/\/github.com\/kingoflolz\/mesh-transformer-jax. https:\/\/github.com\/kingoflolz\/mesh-transformer-jax Ben Wang and Aran Komatsuzaki. 2021. GPT-J-6B: A 6 Billion Parameter Autoregressive Language Model. https:\/\/github.com\/kingoflolz\/mesh-transformer-jax. https:\/\/github.com\/kingoflolz\/mesh-transformer-jax"},{"key":"e_1_3_2_1_75_1","volume-title":"Towards faithful neural table-to-text generation with content-matching constraints. arXiv preprint arXiv:2005.00969","author":"Wang Zhenyi","year":"2020","unstructured":"Zhenyi Wang , Xiaoyang Wang , Bang An , Dong Yu , and Changyou Chen . 2020b. Towards faithful neural table-to-text generation with content-matching constraints. arXiv preprint arXiv:2005.00969 ( 2020 ). Zhenyi Wang, Xiaoyang Wang, Bang An, Dong Yu, and Changyou Chen. 2020b. Towards faithful neural table-to-text generation with content-matching constraints. arXiv preprint arXiv:2005.00969 (2020)."},{"key":"e_1_3_2_1_76_1","volume-title":"A broad-coverage challenge corpus for sentence understanding through inference. arXiv preprint arXiv:1704.05426","author":"Williams Adina","year":"2017","unstructured":"Adina Williams , Nikita Nangia , and Samuel R Bowman . 2017. A broad-coverage challenge corpus for sentence understanding through inference. arXiv preprint arXiv:1704.05426 ( 2017 ). Adina Williams, Nikita Nangia, and Samuel R Bowman. 2017. A broad-coverage challenge corpus for sentence understanding through inference. arXiv preprint arXiv:1704.05426 (2017)."},{"key":"e_1_3_2_1_77_1","volume-title":"Refining the Responses of LLMs by Themselves. arXiv preprint arXiv:2305.04039","author":"Yan Tianqiang","year":"2023","unstructured":"Tianqiang Yan and Tiansheng Xu. 2023. Refining the Responses of LLMs by Themselves. arXiv preprint arXiv:2305.04039 ( 2023 ). Tianqiang Yan and Tiansheng Xu. 2023. Refining the Responses of LLMs by Themselves. arXiv preprint arXiv:2305.04039 (2023)."},{"key":"e_1_3_2_1_78_1","volume-title":"HotpotQA: A dataset for diverse, explainable multi-hop question answering. arXiv preprint arXiv:1809.09600","author":"Yang Zhilin","year":"2018","unstructured":"Zhilin Yang , Peng Qi , Saizheng Zhang , Yoshua Bengio , William W Cohen , Ruslan Salakhutdinov , and Christopher D Manning . 2018. HotpotQA: A dataset for diverse, explainable multi-hop question answering. arXiv preprint arXiv:1809.09600 ( 2018 ). Zhilin Yang, Peng Qi, Saizheng Zhang, Yoshua Bengio, William W Cohen, Ruslan Salakhutdinov, and Christopher D Manning. 2018. HotpotQA: A dataset for diverse, explainable multi-hop question answering. arXiv preprint arXiv:1809.09600 (2018)."},{"key":"e_1_3_2_1_79_1","unstructured":"Wentao Ye Mingfeng Ou Tianyi Li Xuetao Ma Yifan Yanggong Sai Wu Jie Fu Gang Chen Junbo Zhao etal 2023. Assessing Hidden Risks of LLMs: An Empirical Study on Robustness Consistency and Credibility. arXiv preprint arXiv:2305.10235 (2023). Wentao Ye Mingfeng Ou Tianyi Li Xuetao Ma Yifan Yanggong Sai Wu Jie Fu Gang Chen Junbo Zhao et al. 2023. Assessing Hidden Risks of LLMs: An Empirical Study on Robustness Consistency and Credibility. arXiv preprint arXiv:2305.10235 (2023)."},{"key":"e_1_3_2_1_80_1","volume-title":"Bertscore: Evaluating text generation with bert. arXiv preprint arXiv:1904.09675","author":"Zhang Tianyi","year":"2019","unstructured":"Tianyi Zhang , Varsha Kishore , Felix Wu , Kilian Q Weinberger , and Yoav Artzi . 2019 . Bertscore: Evaluating text generation with bert. arXiv preprint arXiv:1904.09675 (2019). Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q Weinberger, and Yoav Artzi. 2019. Bertscore: Evaluating text generation with bert. arXiv preprint arXiv:1904.09675 (2019)."},{"key":"e_1_3_2_1_81_1","volume-title":"Detecting hallucinated content in conditional neural sequence generation. arXiv preprint arXiv:2011.02593","author":"Zhou Chunting","year":"2020","unstructured":"Chunting Zhou , Graham Neubig , Jiatao Gu , Mona Diab , Paco Guzman , Luke Zettlemoyer , and Marjan Ghazvininejad . 2020. Detecting hallucinated content in conditional neural sequence generation. arXiv preprint arXiv:2011.02593 ( 2020 ). Chunting Zhou, Graham Neubig, Jiatao Gu, Mona Diab, Paco Guzman, Luke Zettlemoyer, and Marjan Ghazvininejad. 2020. Detecting hallucinated content in conditional neural sequence generation. arXiv preprint arXiv:2011.02593 (2020)."}],"event":{"name":"CIKM '23: The 32nd ACM International Conference on Information and Knowledge Management","location":"Birmingham United Kingdom","acronym":"CIKM '23","sponsor":["SIGWEB ACM Special Interest Group on Hypertext, Hypermedia, and Web","SIGIR ACM Special Interest Group on Information Retrieval"]},"container-title":["Proceedings of the 32nd ACM International Conference on Information and Knowledge Management"],"original-title":[],"link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3583780.3614905","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3583780.3614905","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T16:36:43Z","timestamp":1750178203000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3583780.3614905"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,10,21]]},"references-count":83,"alternative-id":["10.1145\/3583780.3614905","10.1145\/3583780"],"URL":"https:\/\/doi.org\/10.1145\/3583780.3614905","relation":{},"subject":[],"published":{"date-parts":[[2023,10,21]]},"assertion":[{"value":"2023-10-21","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}