{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,24]],"date-time":"2026-07-24T09:30:58Z","timestamp":1784885458576,"version":"3.55.0"},"reference-count":255,"publisher":"Association for Computing Machinery (ACM)","issue":"3","funder":[{"name":"NSF CAREER Award","award":["2044149"],"award-info":[{"award-number":["2044149"]}]},{"DOI":"10.13039\/100000006","name":"Office of Naval Research","doi-asserted-by":"crossref","award":["N00014-23-1-2148"],"award-info":[{"award-number":["N00014-23-1-2148"]}],"id":[{"id":"10.13039\/100000006","id-type":"DOI","asserted-by":"crossref"}]},{"name":"Sloan Fellowship"},{"name":"National Science Foundation Graduate Research Fellowship"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Comput. Surv."],"published-print":{"date-parts":[[2026,2,28]]},"abstract":"<jats:p>The remarkable performance of large language models (LLMs) in content generation, coding, and common-sense reasoning has spurred widespread integration into many facets of society. However, integration of LLMs raises valid questions on their reliability and trustworthiness, given their propensity to generate hallucinations: plausible, factually-incorrect responses, which are expressed with striking confidence. Previous work has shown that hallucinations and other non-factual responses generated by LLMs can be detected by examining the uncertainty of the LLM in its response to the pertinent prompt, driving significant research efforts devoted to quantifying the uncertainty of LLMs. This survey seeks to provide an extensive review of existing uncertainty quantification methods for LLMs, identifying their salient features, along with their strengths and weaknesses. We present existing methods within a relevant taxonomy, unifying ostensibly disparate methods to aid understanding of the state-of-the-art. Furthermore, we highlight applications of uncertainty quantification methods for LLMs, spanning chatbot and textual applications to embodied artificial intelligence applications in robotics. We conclude with open research challenges in the uncertainty quantification of LLMs, seeking to motivate future research.<\/jats:p>","DOI":"10.1145\/3744238","type":"journal-article","created":{"date-parts":[[2025,6,7]],"date-time":"2025-06-07T05:14:39Z","timestamp":1749273279000},"page":"1-38","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":39,"title":["A Survey on Uncertainty Quantification of Large Language Models: Taxonomy, Open Research Challenges, and Future Directions"],"prefix":"10.1145","volume":"58","author":[{"ORCID":"https:\/\/orcid.org\/0009-0004-2514-6134","authenticated-orcid":false,"given":"Ola","family":"Shorinwa","sequence":"first","affiliation":[{"name":"Princeton University","place":["Princeton, United States"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-1831-8335","authenticated-orcid":false,"given":"Zhiting","family":"Mei","sequence":"additional","affiliation":[{"name":"Princeton University","place":["Princeton, United States"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-8316-1018","authenticated-orcid":false,"given":"Justin","family":"Lidard","sequence":"additional","affiliation":[{"name":"Princeton University","place":["Princeton, United States"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-5306-2844","authenticated-orcid":false,"given":"Allen Z.","family":"Ren","sequence":"additional","affiliation":[{"name":"Princeton University","place":["Princeton, United States"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0002-2296-7485","authenticated-orcid":false,"given":"Anirudha","family":"Majumdar","sequence":"additional","affiliation":[{"name":"Princeton University","place":["Princeton, United States"]}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2025,9,9]]},"reference":[{"key":"e_1_3_2_2_2","unstructured":"Josh Achiam Steven Adler Sandhini Agarwal Lama Ahmad Ilge Akkaya Florencia Leoni Aleman Diogo Almeida Janko Altenschmidt Sam Altman Shyamal Anadkat et\u00a0al. 2023. Gpt-4 technical report. arXiv:2303.08774. Retrieved from https:\/\/arxiv.org\/abs\/2303.08774"},{"key":"e_1_3_2_3_2","unstructured":"Gustaf Ahdritz Tian Qin Nikhil Vyas Boaz Barak and Benjamin L. Edelman. 2024. Distinguishing the knowable from the unknowable with language models. arXiv:2402.03563. Retrieved from https:\/\/arxiv.org\/abs\/2402.03563"},{"key":"e_1_3_2_4_2","unstructured":"Michael Ahn Anthony Brohan Noah Brown Yevgen Chebotar Omar Cortes Byron David Chelsea Finn Chuyuan Fu Keerthana Gopalakrishnan Karol Hausman et\u00a0al. 2022. Do as i can not as i say: Grounding language in robotic affordances. arXiv:2204.01691. Retrieved from https:\/\/arxiv.org\/abs\/2204.01691"},{"key":"e_1_3_2_5_2","unstructured":"Lukas Aichberger Kajetan Schweighofer Mykyta Ielanskyi and Sepp Hochreiter. 2024. Semantically diverse language generation for uncertainty estimation in language models. arXiv:2406.04306. Retrieved from https:\/\/arxiv.org\/abs\/2406.04306"},{"key":"e_1_3_2_6_2","doi-asserted-by":"crossref","unstructured":"Mohammad Aliannejadi Julia Kiseleva Aleksandr Chuklin Jeffrey Dalton and Mikhail Burtsev. 2021. Building and evaluating open-domain dialogue corpora with clarifying questions. arXiv:2109.05794. Retrieved from https:\/\/arxiv.org\/abs\/2109.05794","DOI":"10.18653\/v1\/2021.emnlp-main.367"},{"issue":"2","key":"e_1_3_2_7_2","article-title":"Artificial hallucinations in ChatGPT: Implications in scientific writing","volume":"15","author":"Alkaissi Hussam","year":"2023","unstructured":"Hussam Alkaissi and Samy I. McFarlane. 2023. Artificial hallucinations in ChatGPT: Implications in scientific writing. Cureus 15, 2 (2023).","journal-title":"Cureus"},{"key":"e_1_3_2_8_2","article-title":"The Claude 3 model family: Opus, Sonnet, Haiku","volume":"1","author":"Anthropic AI","year":"2024","unstructured":"AI Anthropic. 2024. The Claude 3 model family: Opus, Sonnet, Haiku. Claude-3 Model Card 1 (2024).","journal-title":"Claude-3 Model Card"},{"key":"e_1_3_2_9_2","unstructured":"Shuang Ao Stefan Rueger and Advaith Siddharthan. 2024. CSS: Contrastive semantic similarity for uncertainty quantification of LLMs. arXiv:2406.03158. Retrieved from https:\/\/arxiv.org\/abs\/2406.03158"},{"key":"e_1_3_2_10_2","unstructured":"Gabriel Y. Arteaga Thomas B. Sch\u00f6n and Nicolas Pielawski. 2024. Hallucination detection in LLMs: Fast and memory-efficient finetuned models. arXiv:2409.02976. Retrieved from https:\/\/arxiv.org\/abs\/2409.02976"},{"key":"e_1_3_2_11_2","volume-title":"Proceedings of the Medical Imaging with Deep Learning","author":"Ayhan Murat Seckin","year":"2018","unstructured":"Murat Seckin Ayhan and Philipp Berens. 2018. Test-time data augmentation for estimation of heteroscedastic aleatoric uncertainty in deep neural networks. In Proceedings of the Medical Imaging with Deep Learning."},{"key":"e_1_3_2_12_2","doi-asserted-by":"publisher","DOI":"10.1186\/s13054-023-04393-x"},{"key":"e_1_3_2_13_2","doi-asserted-by":"crossref","unstructured":"Amos Azaria and Tom Mitchell. 2023. The internal state of an LLM knows when it\u2019s lying. arXiv:2304.13734. Retrieved from https:\/\/arxiv.org\/abs\/2304.13734","DOI":"10.18653\/v1\/2023.findings-emnlp.68"},{"key":"e_1_3_2_14_2","unstructured":"Yuval Bahat and Gregory Shakhnarovich. 2020. Classification confidence estimation with test-time data-augmentation. arXiv:2006.16705. Retrieved from https:\/\/arxiv.org\/abs\/2006.16705"},{"key":"e_1_3_2_15_2","unstructured":"Zechen Bai Pichao Wang Tianjun Xiao Tong He Zongbo Han Zheng Zhang and Mike Zheng Shou. 2024. Hallucination of multimodal large language models: A survey. arXiv:2404.18930. Retrieved from https:\/\/arxiv.org\/abs\/2404.18930"},{"key":"e_1_3_2_16_2","doi-asserted-by":"crossref","unstructured":"Yavuz Faruk Bakman Duygu Nur Yaldiz Baturalp Buyukates Chenyang Tao Dimitrios Dimitriadis and Salman Avestimehr. 2024. MARS: Meaning-aware response scoring for uncertainty estimation in generative LLMs. arXiv:2402.11756. Retrieved from https:\/\/arxiv.org\/abs\/2402.11756","DOI":"10.18653\/v1\/2024.acl-long.419"},{"key":"e_1_3_2_17_2","unstructured":"Oleksandr Balabanov and Hampus Linander. 2024. Uncertainty quantification in fine-tuned LLMs using LoRA ensembles. arXiv:2402.12264. Retrieved from https:\/\/arxiv.org\/abs\/2402.12264"},{"key":"e_1_3_2_18_2","doi-asserted-by":"publisher","DOI":"10.5555\/3692070.3692180"},{"key":"e_1_3_2_19_2","unstructured":"Evan Becker and Stefano Soatto. 2024. Cycles of thought: Measuring LLM confidence through stable explanations. arXiv:2406.03441. Retrieved from https:\/\/arxiv.org\/abs\/2406.03441"},{"key":"e_1_3_2_20_2","doi-asserted-by":"publisher","DOI":"10.1162\/coli_a_00422"},{"key":"e_1_3_2_21_2","unstructured":"Leonard Bereska and Efstratios Gavves. 2024. Mechanistic interpretability for AI safety\u2013a review. arXiv:2404.14082. Retrieved from https:\/\/arxiv.org\/abs\/2404.14082"},{"key":"e_1_3_2_22_2","unstructured":"Anthony Brohan Noah Brown Justice Carbajal Yevgen Chebotar Xi Chen Krzysztof Choromanski Tianli Ding Danny Driess Avinava Dubey Chelsea Finn et\u00a0al. 2023. Rt-2: Vision-language-action models transfer web knowledge to robotic control. arXiv:2307.15818. Retrieved from https:\/\/arxiv.org\/abs\/2307.15818"},{"key":"e_1_3_2_23_2","doi-asserted-by":"crossref","unstructured":"Anthony Brohan Noah Brown Justice Carbajal Yevgen Chebotar Joseph Dabis Chelsea Finn Keerthana Gopalakrishnan Karol Hausman Alex Herzog Jasmine Hsu et\u00a0al. 2022. Rt-1: Robotics transformer for real-world control at scale. arXiv:2212.06817. Retrieved from https:\/\/arxiv.org\/abs\/2212.06817","DOI":"10.15607\/RSS.2023.XIX.025"},{"key":"e_1_3_2_24_2","unstructured":"Tom B. Brown. 2020. Language models are few-shot learners. arXiv:2005.14165. Retrieved from https:\/\/arxiv.org\/abs\/2005.14165"},{"key":"e_1_3_2_25_2","doi-asserted-by":"publisher","DOI":"10.1145\/1150402.1150464"},{"key":"e_1_3_2_26_2","doi-asserted-by":"crossref","unstructured":"Jannis Bulian Christian Buck Wojciech Gajewski Benjamin Boerschinger and Tal Schuster. 2022. Tomayto tomahto. beyond token-level answer equivalence for question answering evaluation. arXiv:2202.07654. Retrieved from https:\/\/arxiv.org\/abs\/2202.07654","DOI":"10.18653\/v1\/2022.emnlp-main.20"},{"key":"e_1_3_2_27_2","unstructured":"Collin Burns Haotian Ye Dan Klein and Jacob Steinhardt. 2022. Discovering latent knowledge in language models without supervision. arXiv:2212.03827. Retrieved from https:\/\/arxiv.org\/abs\/2212.03827"},{"key":"e_1_3_2_28_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.patrec.2016.01.029"},{"key":"e_1_3_2_29_2","unstructured":"Haw-Shiuan Chang Nanyun Peng Mohit Bansal Anil Ramakrishna and Tagyoung Chung. 2024. REAL sampling: Boosting factuality and diversity of open-ended generation via asymptotic entropy. arXiv:2406.07735. Retrieved from https:\/\/arxiv.org\/abs\/2406.07735"},{"key":"e_1_3_2_30_2","unstructured":"Jiuhai Chen and Jonas Mueller. 2023. Quantifying uncertainty in answers from any language model via intrinsic and extrinsic confidence assessment. arXiv:2308.16175. Retrieved from https:\/\/arxiv.org\/abs\/2308.16175"},{"key":"e_1_3_2_31_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICRA57147.2024.10610676"},{"key":"e_1_3_2_32_2","doi-asserted-by":"publisher","DOI":"10.1145\/3583780.3614905"},{"key":"e_1_3_2_33_2","unstructured":"Yangyi Chen Lifan Yuan Ganqu Cui Zhiyuan Liu and Heng Ji. 2022. A close look into the calibration of pre-trained language models. arXiv:2211.00151. Retrieved from https:\/\/arxiv.org\/abs\/2211.00151"},{"key":"e_1_3_2_34_2","unstructured":"Robert Chew John Bollenbacher Michael Wenger Jessica Speer and Annice Kim. 2023. LLM-assisted content analysis: Using large language models to support deductive coding. arXiv:2306.14924. Retrieved from https:\/\/arxiv.org\/abs\/2306.14924"},{"issue":"3","key":"e_1_3_2_35_2","first-page":"6","article-title":"Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality","volume":"2","author":"Chiang Wei-Lin","year":"2023","unstructured":"Wei-Lin Chiang, Zhuohan Li, Zi Lin, Ying Sheng, Zhanghao Wu, Hao Zhang, Lianmin Zheng, Siyuan Zhuang, Yonghao Zhuang, Joseph E. Gonzalez, et\u00a0al. 2023. Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality. See https:\/\/vicuna. lmsys. org (accessed 14 April 2023) 2, 3 (2023), 6.","journal-title":"See https:\/\/vicuna. lmsys. org (accessed 14 April 2023)"},{"key":"e_1_3_2_36_2","unstructured":"Karl Cobbe Vineet Kosaraju Mohammad Bavarian Mark Chen Heewoo Jun Lukasz Kaiser Matthias Plappert Jerry Tworek Jacob Hilton Reiichiro Nakano et\u00a0al. 2021. Training verifiers to solve math word problems. arXiv:2110.14168. Retrieved from https:\/\/arxiv.org\/abs\/2110.14168"},{"key":"e_1_3_2_37_2","doi-asserted-by":"publisher","DOI":"10.3115\/1119239.1119245"},{"key":"e_1_3_2_38_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.ijmedinf.2017.06.004"},{"key":"e_1_3_2_39_2","unstructured":"Hoagy Cunningham Aidan Ewart Logan Riggs Robert Huben and Lee Sharkey. 2023. Sparse autoencoders find highly interpretable features in language models. arXiv:2309.08600. Retrieved from https:\/\/arxiv.org\/abs\/2309.08600"},{"key":"e_1_3_2_40_2","unstructured":"Longchao Da Tiejin Chen Lu Cheng and Hua Wei. 2024. LLM uncertainty quantification through directional entailment graph and claim level response augmentation. arXiv:2407.00994. Retrieved from https:\/\/arxiv.org\/abs\/2407.00994"},{"key":"e_1_3_2_41_2","first-page":"177","volume-title":"Proceedings of the Machine Learning Challenges Workshop","author":"Dagan Ido","year":"2005","unstructured":"Ido Dagan, Oren Glickman, and Bernardo Magnini. 2005. The pascal recognising textual entailment challenge. In Proceedings of the Machine Learning Challenges Workshop. Springer, 177\u2013190."},{"key":"e_1_3_2_42_2","unstructured":"Shih-Chieh Dai Aiping Xiong and Lun-Wei Ku. 2023. LLM-in-the-loop: Leveraging large language model for thematic analysis. arXiv:2310.15100. Retrieved from https:\/\/arxiv.org\/abs\/2310.15100"},{"key":"e_1_3_2_43_2","article-title":"Augmenting judicial practices with LLMs: Re-thinking LLMs\u2019 uncertainty communication features in light of systemic risks","author":"Delacroix Sylvie","year":"2024","unstructured":"Sylvie Delacroix. 2024. Augmenting judicial practices with LLMs: Re-thinking LLMs\u2019 uncertainty communication features in light of systemic risks. Available at SSRN (2024).","journal-title":"Available at SSRN"},{"key":"e_1_3_2_44_2","doi-asserted-by":"crossref","unstructured":"Shrey Desai and Greg Durrett. 2020. Calibration of pre-trained transformers. arXiv:2003.07892. Retrieved from https:\/\/arxiv.org\/abs\/2003.07892","DOI":"10.18653\/v1\/2020.emnlp-main.21"},{"key":"e_1_3_2_45_2","unstructured":"Gianluca Detommaso Martin Bertran Riccardo Fogliato and Aaron Roth. 2024. Multicalibration for confidence scoring in LLMs. arXiv:2404.04689. Retrieved from https:\/\/arxiv.org\/abs\/2404.04689"},{"key":"e_1_3_2_46_2","unstructured":"Jacob Devlin. 2018. Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv:1810.04805. Retrieved from https:\/\/arxiv.org\/abs\/1810.04805"},{"key":"e_1_3_2_47_2","article-title":"Can an embodied agent find your \u201ccat-shaped mug\u201d? llm-based zero-shot object navigation","author":"Dorbala Vishnu Sashank","year":"2023","unstructured":"Vishnu Sashank Dorbala, James F. Mullen Jr, and Dinesh Manocha. 2023. Can an embodied agent find your \u201ccat-shaped mug\u201d? llm-based zero-shot object navigation. IEEE Robotics and Automation Letters (2023).","journal-title":"IEEE Robotics and Automation Letters"},{"key":"e_1_3_2_48_2","unstructured":"Abhimanyu Dubey Abhinav Jauhri Abhinav Pandey Abhishek Kadian Ahmad Al-Dahle Aiesha Letman Akhil Mathur Alan Schelten Amy Yang Angela Fan et\u00a0al. 2024. The llama 3 herd of models. arXiv:2407.21783. Retrieved from https:\/\/arxiv.org\/abs\/2407.21783"},{"key":"e_1_3_2_49_2","unstructured":"Jacob Dunefsky Philippe Chlenski and Neel Nanda. 2024. Transcoders find interpretable LLM feature circuits. arXiv:2406.11944. Retrieved from https:\/\/arxiv.org\/abs\/2406.11944"},{"key":"e_1_3_2_50_2","first-page":"arXiv\u20132308","article-title":"SONAR: Sentence-level multimodal and language-agnostic representations","author":"Duquenne Paul-Ambroise","year":"2023","unstructured":"Paul-Ambroise Duquenne, Holger Schwenk, and Beno\u00eet Sagot. 2023. SONAR: Sentence-level multimodal and language-agnostic representations. arXiv e-prints (2023), arXiv\u20132308.","journal-title":"arXiv e-prints"},{"key":"e_1_3_2_51_2","doi-asserted-by":"publisher","DOI":"10.1162\/tacl_a_00410"},{"key":"e_1_3_2_52_2","unstructured":"Nelson Elhage Tristan Hume Catherine Olsson Nicholas Schiefer Tom Henighan Shauna Kravec Zac Hatfield-Dodds Robert Lasenby Dawn Drain Carol Chen et\u00a0al. 2022. Toy models of superposition. arXiv:2209.10652. Retrieved from https:\/\/arxiv.org\/abs\/2209.10652"},{"key":"e_1_3_2_53_2","unstructured":"Joshua Engels Isaac Liao Eric J Michaud Wes Gurnee and Max Tegmark. 2024. Not all language model features are linear. arXiv:2405.14860. Retrieved from https:\/\/arxiv.org\/abs\/2405.14860"},{"key":"e_1_3_2_54_2","doi-asserted-by":"crossref","unstructured":"Ekaterina Fadeeva Aleksandr Rubashevskii Artem Shelmanov Sergey Petrakov Haonan Li Hamdy Mubarak Evgenii Tsymbalov Gleb Kuzmin Alexander Panchenko Timothy Baldwin et\u00a0al. 2024. Fact-checking the output of large language models via token-level uncertainty quantification. arXiv:2403.04696. Retrieved from https:\/\/arxiv.org\/abs\/2403.04696","DOI":"10.18653\/v1\/2024.findings-acl.558"},{"key":"e_1_3_2_55_2","unstructured":"Fangxiaoyu Feng Yinfei Yang Daniel Cer Naveen Arivazhagan and Wei Wang. 2020. Language-agnostic BERT sentence embedding. arXiv:2007.01852. Retrieved from https:\/\/arxiv.org\/abs\/2007.01852"},{"key":"e_1_3_2_56_2","doi-asserted-by":"crossref","unstructured":"Shangbin Feng Weijia Shi Yike Wang Wenxuan Ding Vidhisha Balachandran and Yulia Tsvetkov. 2024. Don\u2019t Hallucinate Abstain: Identifying LLM knowledge gaps via multi-LLM collaboration. arXiv:2402.00367. Retrieved from https:\/\/arxiv.org\/abs\/2402.00367","DOI":"10.18653\/v1\/2024.acl-long.786"},{"key":"e_1_3_2_57_2","unstructured":"Javier Ferrando Oscar Obeso Senthooran Rajamanoharan and Neel Nanda. 2024. Do I Know This Entity? Knowledge awareness and hallucinations in language models. arXiv:2411.14257. Retrieved from https:\/\/arxiv.org\/abs\/2411.14257"},{"key":"e_1_3_2_58_2","volume-title":"Proceedings of the 2nd Workshop on Inference in Computational Semantics","author":"Fyodorov Yaroslav","year":"2000","unstructured":"Yaroslav Fyodorov, Yoad Winter, and Nissim Francez. 2000. A natural logic inference system. In Proceedings of the 2nd Workshop on Inference in Computational Semantics."},{"key":"e_1_3_2_59_2","first-page":"1050","volume-title":"Proceedings of the International Conference on Machine Learning","author":"Gal Yarin","year":"2016","unstructured":"Yarin Gal and Zoubin Ghahramani. 2016. Dropout as a bayesian approximation: Representing model uncertainty in deep learning. In Proceedings of the International Conference on Machine Learning. PMLR, 1050\u20131059."},{"key":"e_1_3_2_60_2","article-title":"Concrete dropout","volume":"30","author":"Gal Yarin","year":"2017","unstructured":"Yarin Gal, Jiri Hron, and Alex Kendall. 2017. Concrete dropout. Advances in Neural Information Processing Systems 30 (2017).","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_2_61_2","unstructured":"Leo Gao Tom Dupr\u00e9 la Tour Henk Tillman Gabriel Goh Rajan Troll Alec Radford Ilya Sutskever Jan Leike and Jeffrey Wu. 2024. Scaling and evaluating sparse autoencoders. arXiv:2406.04093. Retrieved from https:\/\/arxiv.org\/abs\/2406.04093"},{"key":"e_1_3_2_62_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2024.naacl-long.366"},{"key":"e_1_3_2_63_2","doi-asserted-by":"publisher","DOI":"10.1162\/tacl_a_00370"},{"key":"e_1_3_2_64_2","doi-asserted-by":"crossref","unstructured":"Mor Geva Roei Schuster Jonathan Berant and Omer Levy. 2020. Transformer feed-forward layers are key-value memories. arXiv:2012.14913. Retrieved from https:\/\/arxiv.org\/abs\/2012.14913","DOI":"10.18653\/v1\/2021.emnlp-main.446"},{"key":"e_1_3_2_65_2","doi-asserted-by":"publisher","DOI":"10.1111\/j.1467-9868.2007.00587.x"},{"key":"e_1_3_2_66_2","doi-asserted-by":"publisher","DOI":"10.1198\/016214506000001437"},{"key":"e_1_3_2_67_2","doi-asserted-by":"crossref","unstructured":"Tobias Groot and Matias Valdenegro-Toro. 2024. Overconfidence is key: Verbalized uncertainty evaluation in large language and vision-language models. arXiv:2405.02917. Retrieved from https:\/\/arxiv.org\/abs\/2405.02917","DOI":"10.18653\/v1\/2024.trustnlp-1.13"},{"key":"e_1_3_2_68_2","first-page":"1321","volume-title":"Proceedings of the International Conference on Machine Learning","author":"Guo Chuan","year":"2017","unstructured":"Chuan Guo, Geoff Pleiss, Yu Sun, and Kilian Q Weinberger. 2017. On calibration of modern neural networks. In Proceedings of the International Conference on Machine Learning. PMLR, 1321\u20131330."},{"key":"e_1_3_2_69_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.neucom.2017.06.052"},{"key":"e_1_3_2_70_2","unstructured":"Wes Gurnee Neel Nanda Matthew Pauly Katherine Harvey Dmitrii Troitskii and Dimitris Bertsimas. 2023. Finding neurons in a haystack: Case studies with sparse probing. arXiv:2305.01610. Retrieved from https:\/\/arxiv.org\/abs\/2305.01610"},{"key":"e_1_3_2_71_2","unstructured":"Jiuzhou Han Wray Buntine and Ehsan Shareghi. 2024. Towards uncertainty-aware language agent. arXiv:2401.14016. Retrieved from https:\/\/arxiv.org\/abs\/2401.14016"},{"key":"e_1_3_2_72_2","doi-asserted-by":"publisher","DOI":"10.1093\/biomet\/57.1.97"},{"key":"e_1_3_2_73_2","unstructured":"Jianfeng He Linlin Yu Shuo Lei Chang-Tien Lu and Feng Chen. 2023. Uncertainty estimation on sequential labeling via uncertainty transmission. arXiv:2311.08726. Retrieved from https:\/\/arxiv.org\/abs\/2311.08726"},{"key":"e_1_3_2_74_2","article-title":"Mitigating hallucinations in LLM using K-means clustering of synonym semantic relevance","author":"He Lin","year":"2024","unstructured":"Lin He and Keqin Li. 2024. Mitigating hallucinations in LLM using K-means clustering of synonym semantic relevance. Authorea Preprints (2024).","journal-title":"Authorea Preprints"},{"key":"e_1_3_2_75_2","unstructured":"Pengcheng He Xiaodong Liu Jianfeng Gao and Weizhu Chen. 2020. Deberta: Decoding-enhanced bert with disentangled attention. arXiv:2006.03654. Retrieved from https:\/\/arxiv.org\/abs\/2006.03654"},{"key":"e_1_3_2_76_2","unstructured":"Dan Hendrycks Steven Basart Saurav Kadavath Mantas Mazeika Akul Arora Ethan Guo Collin Burns Samir Puranik Horace He Dawn Song et\u00a0al. 2021. Measuring coding challenge competence with apps. arXiv:2105.09938. Retrieved from https:\/\/arxiv.org\/abs\/2105.09938"},{"key":"e_1_3_2_77_2","unstructured":"Dan Hendrycks Collin Burns Steven Basart Andy Zou Mantas Mazeika Dawn Song and Jacob Steinhardt. 2020. Measuring massive multitask language understanding. arXiv:2009.03300. Retrieved from https:\/\/arxiv.org\/abs\/2009.03300"},{"key":"e_1_3_2_78_2","unstructured":"Geoffrey Hinton. 2015. Distilling the knowledge in a neural network. arXiv:1503.02531. Retrieved from https:\/\/arxiv.org\/abs\/1503.02531"},{"key":"e_1_3_2_79_2","unstructured":"Bairu Hou Yujian Liu Kaizhi Qian Jacob Andreas Shiyu Chang and Yang Zhang. 2023. Decomposing uncertainty for large language models through input clarification ensembling. arXiv:2311.08718. Retrieved from https:\/\/arxiv.org\/abs\/2311.08718"},{"key":"e_1_3_2_80_2","unstructured":"Edward J. Hu Yelong Shen Phillip Wallis Zeyuan Allen-Zhu Yuanzhi Li Shean Wang Lu Wang and Weizhu Chen. 2021. Lora: Low-rank adaptation of large language models. arXiv:2106.09685. Retrieved from https:\/\/arxiv.org\/abs\/2106.09685"},{"key":"e_1_3_2_81_2","doi-asserted-by":"publisher","DOI":"10.1145\/3589335.3648307"},{"key":"e_1_3_2_82_2","unstructured":"Hsiu-Yuan Huang Yutong Yang Zhaoxi Zhang Sanwoo Lee and Yunfang Wu. 2024. A survey of uncertainty estimation in LLMs: Theory meets practice. arXiv:2410.15326. Retrieved from https:\/\/arxiv.org\/abs\/2410.15326"},{"key":"e_1_3_2_83_2","unstructured":"Lei Huang Weijiang Yu Weitao Ma Weihong Zhong Zhangyin Feng Haotian Wang Qianglong Chen Weihua Peng Xiaocheng Feng Bing Qin et\u00a0al. 2023. A survey on hallucination in large language models: Principles taxonomy challenges and open questions. arXiv:2311.05232. Retrieved from https:\/\/arxiv.org\/abs\/2311.05232"},{"key":"e_1_3_2_84_2","first-page":"677","article-title":"On the importance of gradients for detecting distributional shifts in the wild","volume":"34","author":"Huang Rui","year":"2021","unstructured":"Rui Huang, Andrew Geng, and Yixuan Li. 2021. On the importance of gradients for detecting distributional shifts in the wild. Advances in Neural Information Processing Systems 34 (2021), 677\u2013689.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_2_85_2","unstructured":"Yuheng Huang Jiayang Song Zhijie Wang Shengming Zhao Huaming Chen Felix Juefei-Xu and Lei Ma. 2023. Look before you leap: An exploratory study of uncertainty measurement for large language models. arXiv:2307.10236. Retrieved from https:\/\/arxiv.org\/abs\/2307.10236"},{"key":"e_1_3_2_86_2","unstructured":"Conor Igoe Youngseog Chung Ian Char and Jeff Schneider. 2022. How useful are gradients for ood detection really? arXiv:2205.10439. Retrieved from https:\/\/arxiv.org\/abs\/2205.10439"},{"key":"e_1_3_2_87_2","doi-asserted-by":"publisher","DOI":"10.1145\/3571730"},{"key":"e_1_3_2_88_2","unstructured":"Albert Q. Jiang Alexandre Sablayrolles Arthur Mensch Chris Bamford Devendra Singh Chaplot Diego de las Casas Florian Bressand Gianna Lengyel Guillaume Lample Lucile Saulnier et\u00a0al. 2023. Mistral 7B. arXiv:2310.06825. Retrieved from https:\/\/arxiv.org\/abs\/2310.06825"},{"key":"e_1_3_2_89_2","unstructured":"Mingjian Jiang Yangjun Ruan Sicong Huang Saifei Liao Silviu Pitis Roger Baker Grosse and Jimmy Ba. 2023. Calibrating language models via augmented prompt ensembles. (2023)."},{"key":"e_1_3_2_90_2","unstructured":"Mingjian Jiang Yangjun Ruan Prasanna Sattigeri Salim Roukos and Tatsunori Hashimoto. 2024. Graph-based uncertainty metrics for long-form language model outputs. arXiv:2410.20783. Retrieved from https:\/\/arxiv.org\/abs\/2410.20783"},{"key":"e_1_3_2_91_2","unstructured":"Daniel D. Johnson Daniel Tarlow David Duvenaud and Chris J. Maddison. 2024. Experts don\u2019t cheat: Learning what you don\u2019t know by predicting pairs. arXiv:2402.08733. Retrieved from https:\/\/arxiv.org\/abs\/2402.08733"},{"key":"e_1_3_2_92_2","doi-asserted-by":"crossref","unstructured":"Mandar Joshi Eunsol Choi Daniel S. Weld and Luke Zettlemoyer. 2017. Triviaqa: A large scale distantly supervised challenge dataset for reading comprehension. arXiv:1705.03551. Retrieved from https:\/\/arxiv.org\/abs\/1705.03551","DOI":"10.18653\/v1\/P17-1147"},{"key":"e_1_3_2_93_2","doi-asserted-by":"publisher","DOI":"10.1109\/MCI.2022.3155327"},{"key":"e_1_3_2_94_2","unstructured":"Jaehun Jung Faeze Brahman and Yejin Choi. 2024. Trust or escalate: LLM judges with provable guarantees for human agreement. arXiv:2407.18370. Retrieved from https:\/\/arxiv.org\/abs\/2407.18370"},{"key":"e_1_3_2_95_2","unstructured":"Saurav Kadavath Tom Conerly Amanda Askell Tom Henighan Dawn Drain Ethan Perez Nicholas Schiefer Zac Hatfield-Dodds Nova DasSarma Eli Tran-Johnson et\u00a0al. 2022. Language models (mostly) know what they know. arXiv:2207.05221. Retrieved from https:\/\/arxiv.org\/abs\/2207.05221"},{"key":"e_1_3_2_96_2","doi-asserted-by":"publisher","DOI":"10.1162\/tacl_a_00670"},{"key":"e_1_3_2_97_2","unstructured":"Shyam Sundar Kannan Vishnunandan L. N. Venkatesh and Byung-Cheol Min. 2023. Smart-llm: Smart multi-agent robot task planning using large language models. arXiv:2309.10062. Retrieved from https:\/\/arxiv.org\/abs\/2309.10062"},{"key":"e_1_3_2_98_2","unstructured":"Sanyam Kapoor Nate Gruver Manley Roberts Katherine Collins Arka Pal Umang Bhatt Adrian Weller Samuel Dooley Micah Goldblum and Andrew Gordon Wilson. 2024. Large language models must be taught to know what they don\u2019t know. arXiv:2406.08391. Retrieved from https:\/\/arxiv.org\/abs\/2406.08391"},{"key":"e_1_3_2_99_2","doi-asserted-by":"publisher","DOI":"10.1098\/rsta.2023.0254"},{"key":"e_1_3_2_100_2","unstructured":"Geoff Keeling and Winnie Street. 2024. On the attribution of confidence to large language models. arXiv:2407.08388. Retrieved from https:\/\/arxiv.org\/abs\/2407.08388"},{"key":"e_1_3_2_101_2","unstructured":"Moo Jin Kim Karl Pertsch Siddharth Karamcheti Ted Xiao Ashwin Balakrishna Suraj Nair Rafael Rafailov Ethan Foster Grace Lam Pannag Sanketi et\u00a0al. 2024. OpenVLA: An open-source vision-language-action model. arXiv:2406.09246. Retrieved from https:\/\/arxiv.org\/abs\/2406.09246"},{"key":"e_1_3_2_102_2","doi-asserted-by":"publisher","DOI":"10.1145\/3630106.3658941"},{"key":"e_1_3_2_103_2","first-page":"41","volume-title":"Proceedings of the 1st Workshop on Uncertainty-Aware NLP","author":"Kolagar Zahra","year":"2024","unstructured":"Zahra Kolagar and Alessandra Zarcone. 2024. Aligning uncertainty: Leveraging LLMs to analyze uncertainty transfer in text summarization. In Proceedings of the 1st Workshop on Uncertainty-Aware NLP. 41\u201361."},{"key":"e_1_3_2_104_2","doi-asserted-by":"crossref","unstructured":"Lingkai Kong Haoming Jiang Yuchen Zhuang Jie Lyu Tuo Zhao and Chao Zhang. 2020. Calibrated language model fine-tuning for in-and out-of-distribution data. arXiv:2010.11506. Retrieved from https:\/\/arxiv.org\/abs\/2010.11506","DOI":"10.18653\/v1\/2020.emnlp-main.102"},{"key":"e_1_3_2_105_2","unstructured":"Jannik Kossen Jiatong Han Muhammed Razzak Lisa Schut Shreshth Malik and Yarin Gal. 2024. Semantic entropy probes: Robust and cheap hallucination detection in llms. arXiv:2406.15927. Retrieved from https:\/\/arxiv.org\/abs\/2406.15927"},{"key":"e_1_3_2_106_2","first-page":"1","volume-title":"Proceedings of the Workshop on Multimodal, Multilingual Natural Language Generation and Multilingual WebNLG Challenge","author":"Krause Lea","year":"2023","unstructured":"Lea Krause, Wondimagegnhue Tufa, Selene B\u00e1ez Santamar\u00eda, Angel Daza, Urja Khurana, and Piek Vossen. 2023. Confidently wrong: Exploring the calibration and expression of (Un) certainty of large language models in a multilingual setting. In Proceedings of the Workshop on Multimodal, Multilingual Natural Language Generation and Multilingual WebNLG Challenge. 1\u20139."},{"key":"e_1_3_2_107_2","unstructured":"Lorenz Kuhn Yarin Gal and Sebastian Farquhar. 2023. Semantic uncertainty: Linguistic invariances for uncertainty estimation in natural language generation. arXiv:2302.09664. Retrieved from https:\/\/arxiv.org\/abs\/2302.09664"},{"key":"e_1_3_2_108_2","unstructured":"Bhawesh Kumar Charlie Lu Gauri Gupta Anil Palepu David Bellamy Ramesh Raskar and Andrew Beam. 2023. Conformal prediction with large language models for multi-choice question answering. arXiv:2305.18404. Retrieved from https:\/\/arxiv.org\/abs\/2305.18404"},{"key":"e_1_3_2_109_2","doi-asserted-by":"crossref","unstructured":"Guokun Lai Qizhe Xie Hanxiao Liu Yiming Yang and Eduard Hovy. 2017. Race: Large-scale reading comprehension dataset from examinations. arXiv:1704.04683. Retrieved from https:\/\/arxiv.org\/abs\/1704.04683","DOI":"10.18653\/v1\/D17-1082"},{"key":"e_1_3_2_110_2","article-title":"Simple and scalable predictive uncertainty estimation using deep ensembles","volume":"30","author":"Lakshminarayanan Balaji","year":"2017","unstructured":"Balaji Lakshminarayanan, Alexander Pritzel, and Charles Blundell. 2017. Simple and scalable predictive uncertainty estimation using deep ensembles. Advances in Neural Information Processing Systems 30 (2017).","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_2_111_2","article-title":"Generating text from structured data with application to the biography domain","author":"Lebret R\u00e9mi","year":"2016","unstructured":"R\u00e9mi Lebret, David Grangier, and Michael Auli. 2016. Generating text from structured data with application to the biography domain. CoRR, abs\/1603.07771 (2016).","journal-title":"CoRR, abs\/1603.07771"},{"key":"e_1_3_2_112_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.ins.2023.119356"},{"key":"e_1_3_2_113_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICIP40778.2020.9190679"},{"key":"e_1_3_2_114_2","unstructured":"Katherine Lee Orhan Firat Ashish Agarwal Clara Fannjiang and David Sussillo. 2018. Hallucinations in neural machine translation. (2018)."},{"key":"e_1_3_2_115_2","volume-title":"Proceedings of the ICLR 2024 Workshop on Reliable and Responsible Foundation Models","author":"Li Chengzu","year":"2024","unstructured":"Chengzu Li, Han Zhou, Goran Glava\u0161, Anna Korhonen, and Ivan Vuli\u0107. 2024. Can large language models achieve calibration with in-context learning?. In Proceedings of the ICLR 2024 Workshop on Reliable and Responsible Foundation Models."},{"key":"e_1_3_2_116_2","unstructured":"Junyi Li Xiaoxue Cheng Wayne Xin Zhao Jian-Yun Nie and Ji-Rong Wen. 2023. Halueval: A large-scale hallucination evaluation benchmark for large language models. arXiv:2305.11747. Retrieved from https:\/\/arxiv.org\/abs\/2305.11747"},{"key":"e_1_3_2_117_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2024.naacl-long.210"},{"key":"e_1_3_2_118_2","unstructured":"Kaiqu Liang Zixu Zhang and Jaime Fern\u00e1ndez Fisac. 2024. Introspective planning: Guiding language-enabled agents to refine their own uncertainty. arXiv:2402.06529. Retrieved from https:\/\/arxiv.org\/abs\/2402.06529"},{"key":"e_1_3_2_119_2","unstructured":"Tom Lieberum Matthew Rahtz J\u00e1nos Kram\u00e1r Neel Nanda Geoffrey Irving Rohin Shah and Vladimir Mikulik. 2023. Does circuit analysis interpretability scale? Evidence from multiple choice capabilities in chinchilla. arXiv:2307.09458. Retrieved from https:\/\/arxiv.org\/abs\/2307.09458"},{"key":"e_1_3_2_120_2","doi-asserted-by":"crossref","unstructured":"Tom Lieberum Senthooran Rajamanoharan Arthur Conmy Lewis Smith Nicolas Sonnerat Vikrant Varma J\u00e1nos Kram\u00e1r Anca Dragan Rohin Shah and Neel Nanda. 2024. Gemma scope: Open sparse autoencoders everywhere all at once on gemma 2. arXiv:2408.05147. Retrieved from https:\/\/arxiv.org\/abs\/2408.05147","DOI":"10.18653\/v1\/2024.blackboxnlp-1.19"},{"key":"e_1_3_2_121_2","first-page":"74","volume-title":"Proceedings of the Text Summarization Branches Out","author":"Lin Chin-Yew","year":"2004","unstructured":"Chin-Yew Lin. 2004. Rouge: A package for automatic evaluation of summaries. In Proceedings of the Text Summarization Branches Out. 74\u201381."},{"key":"e_1_3_2_122_2","unstructured":"Stephanie Lin Jacob Hilton and Owain Evans. 2021. Truthfulqa: Measuring how models mimic human falsehoods. arXiv:2109.07958. Retrieved from https:\/\/arxiv.org\/abs\/2109.07958"},{"key":"e_1_3_2_123_2","unstructured":"Stephanie Lin Jacob Hilton and Owain Evans. 2022. Teaching models to express their uncertainty in words. arXiv:2205.14334. Retrieved from https:\/\/arxiv.org\/abs\/2205.14334"},{"key":"e_1_3_2_124_2","unstructured":"Zhen Lin Shubhendu Trivedi and Jimeng Sun. 2023. Generating with confidence: Uncertainty quantification for black-box large language models. arXiv:2305.19187. Retrieved from https:\/\/arxiv.org\/abs\/2305.19187"},{"key":"e_1_3_2_125_2","doi-asserted-by":"crossref","unstructured":"Chen Ling Xujiang Zhao Wei Cheng Yanchi Liu Yiyou Sun Xuchao Zhang Mika Oishi Takao Osaki Katsushi Matsuda Jie Ji et\u00a0al. 2024. Uncertainty decomposition and quantification for in-context learning of large language models. arXiv:2402.10189. Retrieved from https:\/\/arxiv.org\/abs\/2402.10189","DOI":"10.18653\/v1\/2024.naacl-long.184"},{"key":"e_1_3_2_126_2","doi-asserted-by":"crossref","unstructured":"Alisa Liu Zhaofeng Wu Julian Michael Alane Suhr Peter West Alexander Koller Swabha Swayamdipta Noah A. Smith and Yejin Choi. 2023. We\u2019re afraid language models aren\u2019t modeling ambiguity. arXiv:2304.14399. Retrieved from https:\/\/arxiv.org\/abs\/2304.14399","DOI":"10.18653\/v1\/2023.emnlp-main.51"},{"key":"e_1_3_2_127_2","unstructured":"Hongfu Liu Hengguan Huang Hao Wang Xiangming Gu and Ye Wang. 2024. On calibration of LLM-based guard models for reliable content moderation. arXiv:2410.10414. Retrieved from https:\/\/arxiv.org\/abs\/2410.10414"},{"key":"e_1_3_2_128_2","article-title":"Visual instruction tuning","volume":"36","author":"Liu Haotian","year":"2024","unstructured":"Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. 2024. Visual instruction tuning. Advances in Neural Information Processing Systems 36 (2024).","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_2_129_2","unstructured":"Hanchao Liu Wenyuan Xue Yifei Chen Dapeng Chen Xiutian Zhao Ke Wang Liping Hou Rongjun Li and Wei Peng. 2024. A survey on hallucination in large vision-language models. arXiv:2402.00253. Retrieved from https:\/\/arxiv.org\/abs\/2402.00253"},{"key":"e_1_3_2_130_2","unstructured":"Linyu Liu Yu Pan Xiaocheng Li and Guanting Chen. 2024. Uncertainty estimation and quantification for LLMs: A simple supervised approach. arXiv:2404.15993. Retrieved from https:\/\/arxiv.org\/abs\/2404.15993"},{"key":"e_1_3_2_131_2","unstructured":"Terrance Liu and Zhiwei Steven Wu. 2024. Multi-group uncertainty quantification for long-form text generation. arXiv:2407.21057. Retrieved from https:\/\/arxiv.org\/abs\/2407.21057"},{"key":"e_1_3_2_132_2","volume-title":"Proceedings of the 12th International Conference on Learning Representations","author":"Liu Xin","year":"2024","unstructured":"Xin Liu, Muhammad Khalifa, and Lu Wang. 2024. LitCab: Lightweight language model calibration over short-and long-form responses. In Proceedings of the 12th International Conference on Learning Representations."},{"key":"e_1_3_2_133_2","unstructured":"Yinhan Liu. 2019. Roberta: A robustly optimized bert pretraining approach. arXiv:1907.11692. Retrieved from https:\/\/arxiv.org\/abs\/1907.11692"},{"key":"e_1_3_2_134_2","unstructured":"Yuxuan Liu Tianchi Yang Shaohan Huang Zihan Zhang Haizhen Huang Furu Wei Weiwei Deng Feng Sun and Qi Zhang. 2023. Calibrating llm-based evaluator. arXiv:2309.13308. Retrieved from https:\/\/arxiv.org\/abs\/2309.13308"},{"key":"e_1_3_2_135_2","unstructured":"Yang Liu Yuanshun Yao Jean-Francois Ton Xiaoying Zhang Ruocheng Guo Hao Cheng Yegor Klochkov Muhammad Faaiz Taufiq and Hang Li. 2023. Trustworthy LLMs: A survey and guideline for evaluating large language models\u2019 alignment. arXiv:2308.05374. Retrieved from https:\/\/arxiv.org\/abs\/2308.05374"},{"key":"e_1_3_2_136_2","doi-asserted-by":"publisher","DOI":"10.1109\/LRA.2020.2974682"},{"key":"e_1_3_2_137_2","unstructured":"Qing Lyu Kumar Shridhar Chaitanya Malaviya Li Zhang Yanai Elazar Niket Tandon Marianna Apidianaki Mrinmaya Sachan and Chris Callison-Burch. 2024. Calibrating large language models with sample consistency. arXiv:2402.13904. Retrieved from https:\/\/arxiv.org\/abs\/2402.13904"},{"key":"e_1_3_2_138_2","doi-asserted-by":"publisher","DOI":"10.3115\/1599081.1599147"},{"key":"e_1_3_2_139_2","doi-asserted-by":"crossref","unstructured":"Mat\u00e9o Mahaut Laura Aina Paula Czarnowska Momchil Hardalov Thomas M\u00fcller and Llu\u00eds M\u00e0rquez. 2024. Factual confidence of LLMs: On reliability and robustness of current estimators. arXiv:2406.13415. Retrieved from https:\/\/arxiv.org\/abs\/2406.13415","DOI":"10.18653\/v1\/2024.acl-long.250"},{"key":"e_1_3_2_140_2","unstructured":"Andrey Malinin and Mark Gales. 2020. Uncertainty estimation in autoregressive structured prediction. arXiv:2002.07650. Retrieved from https:\/\/arxiv.org\/abs\/2002.07650"},{"key":"e_1_3_2_141_2","first-page":"269","volume-title":"Proceedings of the Conformal and Probabilistic Prediction and Applications","author":"Maltoudoglou Lysimachos","year":"2020","unstructured":"Lysimachos Maltoudoglou, Andreas Paisios, and Harris Papadopoulos. 2020. BERT-based conformal predictor for sentiment analysis. In Proceedings of the Conformal and Probabilistic Prediction and Applications. PMLR, 269\u2013284."},{"key":"e_1_3_2_142_2","doi-asserted-by":"crossref","unstructured":"Potsawee Manakul Adian Liusie and Mark JF Gales. 2023. Selfcheckgpt: Zero-resource black-box hallucination detection for generative large language models. arXiv:2303.08896. Retrieved from https:\/\/arxiv.org\/abs\/2303.08896","DOI":"10.18653\/v1\/2023.emnlp-main.557"},{"key":"e_1_3_2_143_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICRA57147.2024.10610855"},{"key":"e_1_3_2_144_2","doi-asserted-by":"crossref","unstructured":"Xin Mao Feng-Lin Li Huimin Xu Wei Zhang and Anh Tuan Luu. 2024. Don\u2019t forget your reward values: Language model alignment via value-based calibration. arXiv:2402.16030. Retrieved from https:\/\/arxiv.org\/abs\/2402.16030","DOI":"10.18653\/v1\/2024.emnlp-main.976"},{"key":"e_1_3_2_145_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.knosys.2022.109265"},{"key":"e_1_3_2_146_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2008.78"},{"key":"e_1_3_2_147_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v36i9.21243"},{"key":"e_1_3_2_148_2","doi-asserted-by":"publisher","DOI":"10.1162\/tacl_a_00494"},{"key":"e_1_3_2_149_2","doi-asserted-by":"crossref","unstructured":"Sewon Min Julian Michael Hannaneh Hajishirzi and Luke Zettlemoyer. 2020. AmbigQA: Answering ambiguous open-domain questions. arXiv:2004.10645. Retrieved from https:\/\/arxiv.org\/abs\/2004.10645","DOI":"10.18653\/v1\/2020.emnlp-main.466"},{"key":"e_1_3_2_150_2","unstructured":"Shervin Minaee Tomas Mikolov Narjes Nikzad Meysam Chenaghlu Richard Socher Xavier Amatriain and Jianfeng Gao. 2024. Large language models: A survey. arXiv:2402.06196. Retrieved from https:\/\/arxiv.org\/abs\/2402.06196"},{"key":"e_1_3_2_151_2","unstructured":"Christopher Mohri and Tatsunori Hashimoto. 2024. Language models with conformal factuality guarantees. arXiv:2402.10978. Retrieved from https:\/\/arxiv.org\/abs\/2402.10978"},{"key":"e_1_3_2_152_2","volume-title":"Proceedings of the 3rd Workshop on Inference in Computational Semantics","author":"Monz Christof","year":"2001","unstructured":"Christof Monz and Maarten de Rijke. 2001. Light-weight entailment checking for computational semantics. In Proceedings of the 3rd Workshop on Inference in Computational Semantics."},{"key":"e_1_3_2_153_2","unstructured":"James F. Mullen Jr and Dinesh Manocha. 2024. Towards robots that know when they need help: Affordance-based uncertainty for large language model planners. arXiv:2403.13198. Retrieved from https:\/\/arxiv.org\/abs\/2403.13198"},{"key":"e_1_3_2_154_2","volume-title":"Proceedings of the AAAI Conference on Artificial Intelligence","author":"Naeini Mahdi Pakdaman","year":"2015","unstructured":"Mahdi Pakdaman Naeini, Gregory Cooper, and Milos Hauskrecht. 2015. Obtaining well calibrated probabilities using bayesian binning. In Proceedings of the AAAI Conference on Artificial Intelligence."},{"key":"e_1_3_2_155_2","unstructured":"Neel Nanda Lawrence Chan Tom Lieberum Jess Smith and Jacob Steinhardt. 2023. Progress measures for grokking via mechanistic interpretability. arXiv:2301.05217. Retrieved from https:\/\/arxiv.org\/abs\/2301.05217"},{"key":"e_1_3_2_156_2","unstructured":"Shiyu Ni Keping Bi Lulu Yu and Jiafeng Guo. 2024. Are large language models more honest in their probabilistic or verbalized confidence? arXiv:2408.09773. Retrieved from https:\/\/arxiv.org\/abs\/2408.09773"},{"key":"e_1_3_2_157_2","doi-asserted-by":"publisher","DOI":"10.1145\/1102351.1102430"},{"key":"e_1_3_2_158_2","unstructured":"Alexander Nikitin Jannik Kossen Yarin Gal and Pekka Marttinen. 2024. Kernel language entropy: Fine-grained uncertainty quantification for LLMs from semantic similarities. arXiv:2405.20003. Retrieved from https:\/\/arxiv.org\/abs\/2405.20003"},{"key":"e_1_3_2_159_2","unstructured":"Ruijia Niu Dongxia Wu Rose Yu and Yi-An Ma. 2024. Functional-level uncertainty quantification for calibrated fine-tuning on LLMs. arXiv:2410.06431. Retrieved from https:\/\/arxiv.org\/abs\/2410.06431"},{"key":"e_1_3_2_160_2","volume-title":"Proceedings of the CVPR Workshops","author":"Nixon Jeremy","year":"2019","unstructured":"Jeremy Nixon, Michael W. Dusenberry, Linchuan Zhang, Ghassen Jerfel, and Dustin Tran. 2019. Measuring calibration in deep learning. In Proceedings of the CVPR Workshops."},{"key":"e_1_3_2_161_2","unstructured":"Ian Osband Seyed Mohammad Asghari Benjamin Van Roy Nat McAleese John Aslanides and Geoffrey Irving. 2022. Fine-tuning language models via epistemic neural networks. arXiv:2211.01568. Retrieved from https:\/\/arxiv.org\/abs\/2211.01568"},{"key":"e_1_3_2_162_2","first-page":"2795","article-title":"Epistemic neural networks","volume":"36","author":"Osband Ian","year":"2023","unstructured":"Ian Osband, Zheng Wen, Seyed Mohammad Asghari, Vikranth Dwaracherla, Morteza Ibrahimi, Xiuyuan Lu, and Benjamin Van Roy. 2023. Epistemic neural networks. Advances in Neural Information Processing Systems 36 (2023), 2795\u20132823.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_2_163_2","unstructured":"Lorenzo Pacchiardi Alex J. Chan S\u00f6ren Mindermann Ilan Moscovitz Alexa Y. Pan Yarin Gal Owain Evans and Jan Brauner. 2023. How to catch an ai liar: Lie detection in black-box llms by asking unrelated questions. arXiv:2309.15840. Retrieved from https:\/\/arxiv.org\/abs\/2309.15840"},{"key":"e_1_3_2_164_2","unstructured":"Alina Petukhova Joao P. Matos-Carvalho and Nuno Fachada. 2024. Text clustering with LLM embeddings. arXiv:2403.15112. Retrieved from https:\/\/arxiv.org\/abs\/2403.15112"},{"key":"e_1_3_2_165_2","first-page":"1341","volume-title":"Proceedings of the 51st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)","author":"Pilehvar Mohammad Taher","year":"2013","unstructured":"Mohammad Taher Pilehvar, David Jurgens, and Roberto Navigli. 2013. Align, disambiguate and walk: A unified approach for measuring semantic similarity. In Proceedings of the 51st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 1341\u20131351."},{"issue":"3","key":"e_1_3_2_166_2","first-page":"61","article-title":"Probabilistic outputs for support vector machines and comparisons to regularized likelihood methods","volume":"10","author":"Platt John","year":"1999","unstructured":"John Platt. 1999. Probabilistic outputs for support vector machines and comparisons to regularized likelihood methods. Advances in Large Margin Classifiers 10, 3 (1999), 61\u201374.","journal-title":"Advances in Large Margin Classifiers"},{"key":"e_1_3_2_167_2","unstructured":"Konstantin Posch Jan Steinbrener and J\u00fcrgen Pilz. 2019. Variational inference to measure model uncertainty in deep neural networks. arXiv:1902.10189. Retrieved from https:\/\/arxiv.org\/abs\/1902.10189"},{"key":"e_1_3_2_168_2","unstructured":"Xin Qiu and Risto Miikkulainen. 2024. Semantic density: Uncertainty quantification in semantic space for large language models. arXiv:2405.13845. Retrieved from https:\/\/arxiv.org\/abs\/2405.13845"},{"key":"e_1_3_2_169_2","unstructured":"Victor Quach Adam Fisch Tal Schuster Adam Yala Jae Ho Sohn Tommi S. Jaakkola and Regina Barzilay. 2023. Conformal language modeling. arXiv:2306.10193. Retrieved from https:\/\/arxiv.org\/abs\/2306.10193"},{"key":"e_1_3_2_170_2","doi-asserted-by":"publisher","DOI":"10.1109\/INISTA49547.2020.9194665"},{"key":"e_1_3_2_171_2","unstructured":"Alec Radford and Karthik Narasimhan. 2018. Improving language understanding by generative pre-training."},{"key":"e_1_3_2_172_2","first-page":"20063","article-title":"Uncertainty quantification and deep ensembles","volume":"34","year":"2021","unstructured":"Rahul Rahaman and Alexandre H. Thiery. 2021. Uncertainty quantification and deep ensembles. Advances in Neural Information Processing Systems 34 (2021), 20063\u201320075.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_2_173_2","unstructured":"Daking Rai Yilun Zhou Shi Feng Abulhair Saparov and Ziyu Yao. 2024. A practical review of mechanistic interpretability for transformer-based language models. arXiv:2407.02646. Retrieved from https:\/\/arxiv.org\/abs\/2407.02646"},{"key":"e_1_3_2_174_2","unstructured":"Vipula Rawte Amit Sheth and Amitava Das. 2023. A survey of hallucination in large foundation models. arXiv:2309.05922. Retrieved from https:\/\/arxiv.org\/abs\/2309.05922"},{"key":"e_1_3_2_175_2","doi-asserted-by":"publisher","DOI":"10.1162\/tacl_a_00266"},{"key":"e_1_3_2_176_2","doi-asserted-by":"crossref","unstructured":"N. Reimers. 2019. Sentence-BERT: Sentence embeddings using siamese BERT-networks. arXiv:1908.10084. Retrieved from https:\/\/arxiv.org\/abs\/1908.10084","DOI":"10.18653\/v1\/D19-1410"},{"key":"e_1_3_2_177_2","unstructured":"David Rein Betty Li Hou Asa Cooper Stickland Jackson Petty Richard Yuanzhe Pang Julien Dirani Julian Michael and Samuel R. Bowman. 2023. Gpqa: A graduate-level google-proof q&a benchmark. arXiv:2311.12022. Retrieved from https:\/\/arxiv.org\/abs\/2311.12022"},{"key":"e_1_3_2_178_2","unstructured":"Allen Z. Ren Jaden Clark Anushri Dixit Masha Itkina Anirudha Majumdar and Dorsa Sadigh. 2024. Explore until Confident: Efficient exploration for embodied question answering. arXiv:2403.15941. Retrieved from https:\/\/arxiv.org\/abs\/2403.15941"},{"key":"e_1_3_2_179_2","unstructured":"Allen Z. Ren Anushri Dixit Alexandra Bodrova Sumeet Singh Stephen Tu Noah Brown Peng Xu Leila Takayama Fei Xia Jake Varley et\u00a0al. 2023. Robots that ask for help: Uncertainty alignment for large language model planners. arXiv:2307.01928. Retrieved from https:\/\/arxiv.org\/abs\/2307.01928"},{"key":"e_1_3_2_180_2","series-title":"Proceedings of Machine Learning Research","first-page":"49","volume-title":"Proceedings on \u201dI Can\u2019t Believe It\u2019s Not Better: Failure Modes in the Age of Foundation Models\u201d at NeurIPS 2023 Workshops","volume":"239","author":"Ren Jie","year":"2023","unstructured":"Jie Ren, Yao Zhao, Tu Vu, Peter J. Liu, and Balaji Lakshminarayanan. 2023. Self-evaluation improves selective generation in large language models. In Proceedings on \u201dI Can\u2019t Believe It\u2019s Not Better: Failure Modes in the Age of Foundation Models\u201d at NeurIPS 2023 Workshops(Proceedings of Machine Learning Research, Vol. 239).Javier Antor\u00e1n, Arno Blaas, Kelly Buchanan, Fan Feng, Vincent Fortuin, Sahra Ghalebikesabi, Andreas Kriegler, Ian Mason, David Rohde, Francisco J. R. Ruiz, et al. (Eds.), PMLR, 49\u201364."},{"key":"e_1_3_2_181_2","unstructured":"Pouria Rouzrokh Shahriar Faghani Cooper U Gamble Moein Shariatnia and Bradley J. Erickson. 2024. CONFLARE: CONFormal LArge language model REtrieval. arXiv:2404.04287. Retrieved from https:\/\/arxiv.org\/abs\/2404.04287"},{"key":"e_1_3_2_182_2","doi-asserted-by":"publisher","DOI":"10.1080\/01621459.2017.1395341"},{"key":"e_1_3_2_183_2","article-title":"Cxplain: Causal explanations for model interpretation under uncertainty","volume":"32","author":"Schwab Patrick","year":"2019","unstructured":"Patrick Schwab and Walter Karlen. 2019. Cxplain: Causal explanations for model interpretation under uncertainty. Advances in Neural Information Processing Systems 32 (2019).","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_2_184_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v35i15.17623"},{"issue":"3","key":"e_1_3_2_185_2","article-title":"A tutorial on conformal prediction.","volume":"9","author":"Shafer Glenn","year":"2008","unstructured":"Glenn Shafer and Vladimir Vovk. 2008. A tutorial on conformal prediction. Journal of Machine Learning Research 9, 3 (2008), 371\u2013412.","journal-title":"Journal of Machine Learning Research"},{"key":"e_1_3_2_186_2","first-page":"492","volume-title":"Proceedings of the Conference on Robot Learning","author":"Shah Dhruv","year":"2023","unstructured":"Dhruv Shah, B\u0142a\u017cej Osi\u0144ski, Brian Ichter, and Sergey Levine. 2023. Lm-nav: Robotic navigation with large pre-trained models of language, vision, and action. In Proceedings of the Conference on Robot Learning. PMLR, 492\u2013504."},{"key":"e_1_3_2_187_2","unstructured":"Eric Michael Smith Diana Gonzalez-Rico Emily Dinan and Y.-Lan Boureau. 2020. Controlling style in generated dialogue. arXiv:2009.10855. Retrieved from https:\/\/arxiv.org\/abs\/2009.10855"},{"key":"e_1_3_2_188_2","unstructured":"Claudio Spiess David Gros Kunal Suresh Pai Michael Pradel Md Rafiqul Islam Rabin Amin Alipour Susmit Jha Prem Devanbu and Toufique Ahmed. 2024. Calibration and correctness of language models for code. arXiv:2402.02047. Retrieved from https:\/\/arxiv.org\/abs\/2402.02047"},{"key":"e_1_3_2_189_2","first-page":"35","volume-title":"Proceedings of the 1st Workshop on Uncertainty-Aware NLP","author":"Steindl Sebastian","year":"2024","unstructured":"Sebastian Steindl, Ulrich Sch\u00e4fer, Bernd Ludwig, and Patrick Levi. 2024. Linguistic obfuscation attacks and large language model uncertainty. In Proceedings of the 1st Workshop on Uncertainty-Aware NLP. 35\u201340."},{"key":"e_1_3_2_190_2","unstructured":"Elias Stengel-Eskin Peter Hase and Mohit Bansal. 2024. LACIE: Listener-aware finetuning for confidence calibration in large language models. arXiv:2405.21028. Retrieved from https:\/\/arxiv.org\/abs\/2405.21028"},{"key":"e_1_3_2_191_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.neucom.2023.127063"},{"key":"e_1_3_2_192_2","unstructured":"Jiayuan Su Jing Luo Hongwei Wang and Lu Cheng. 2024. Api is enough: Conformal prediction for large language models without logit-access. arXiv:2403.01216. Retrieved from https:\/\/arxiv.org\/abs\/2403.01216"},{"key":"e_1_3_2_193_2","unstructured":"Xingpeng Sun Yiran Zhang Xindi Tang Amrit Singh Bedi and Aniket Bera. 2024. TrustNavGPT: Modeling uncertainty to improve trustworthiness of audio-guided LLM-based robot navigation. arXiv:2408.01867. Retrieved from https:\/\/arxiv.org\/abs\/2408.01867"},{"key":"e_1_3_2_194_2","article-title":"ReDeEP: Detecting hallucination in retrieval-augmented generation via mechanistic interpretability","author":"Sun Zhongxiang","year":"2024","unstructured":"Zhongxiang Sun, Xiaoxue Zang, Kai Zheng, Yang Song, Jun Xu, Xiao Zhang, Weijie Yu, and Han Li. 2024. ReDeEP: Detecting hallucination in retrieval-augmented generation via mechanistic interpretability. arXiv:2410.11414. Retrieved from https:\/\/arxiv.org\/abs\/2410.11414","journal-title":"a"},{"key":"e_1_3_2_195_2","doi-asserted-by":"publisher","DOI":"10.1177\/16094069241231168"},{"key":"e_1_3_2_196_2","unstructured":"Alex Tamkin Kunal Handa Avash Shrestha and Noah Goodman. 2022. Task ambiguity in humans and language models. arXiv:2212.10711. Retrieved from https:\/\/arxiv.org\/abs\/2212.10711"},{"key":"e_1_3_2_197_2","unstructured":"Alex Tamkin Mohammad Taufeeque and Noah D. Goodman. 2023. Codebook features: Sparse and discrete interpretability for neural networks. arXiv:2310.17230. Retrieved from https:\/\/arxiv.org\/abs\/2310.17230"},{"key":"e_1_3_2_198_2","unstructured":"Zhisheng Tang Ke Shen and Mayank Kejriwal. 2024. An evaluation of estimative uncertainty in large language models. arXiv:2405.15185. Retrieved from https:\/\/arxiv.org\/abs\/2405.15185"},{"key":"e_1_3_2_199_2","first-page":"1072","volume-title":"Proceedings of the International Conference on Artificial Intelligence and Statistics","author":"Tanneru Sree Harsha","year":"2024","unstructured":"Sree Harsha Tanneru, Chirag Agarwal, and Himabindu Lakkaraju. 2024. Quantifying uncertainty in natural language explanations of large language models. In Proceedings of the International Conference on Artificial Intelligence and Statistics. PMLR, 1072\u20131080."},{"key":"e_1_3_2_200_2","unstructured":"Shuchang Tao Liuyi Yao Hanxing Ding Yuexiang Xie Qi Cao Fei Sun Jinyang Gao Huawei Shen and Bolin Ding. 2024. When to trust LLMs: Aligning confidence with response quality. arXiv:2404.17287. Retrieved from https:\/\/arxiv.org\/abs\/2404.17287"},{"key":"e_1_3_2_201_2","unstructured":"Adly Templeton Tom Conerly Jonathan Marcus Jack Lindsey Trenton Bricken Brian Chen Adam Pearce Craig Citro Emmanuel Ameisen Andy Jones et\u00a0al. 2024. Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnet. Retrieved from https:\/\/transformer-circuits.pub\/2024\/scaling-monosemanticity\/index.html"},{"key":"e_1_3_2_202_2","volume-title":"Elements of Information Theory","author":"Thomas MTCAJ","year":"2006","unstructured":"MTCAJ Thomas and A. Thomas Joy. 2006. Elements of Information Theory. Wiley-Interscience."},{"key":"e_1_3_2_203_2","doi-asserted-by":"crossref","unstructured":"James Thorne Andreas Vlachos Christos Christodoulopoulos and Arpit Mittal. 2018. FEVER: A large-scale dataset for fact extraction and VERification. arXiv:1803.05355. Retrieved from https:\/\/arxiv.org\/abs\/1803.05355","DOI":"10.18653\/v1\/W18-5501"},{"key":"e_1_3_2_204_2","unstructured":"Katherine Tian Eric Mitchell Allan Zhou Archit Sharma Rafael Rafailov Huaxiu Yao Chelsea Finn and Christopher D. Manning. 2023. Just ask for calibration: Strategies for eliciting calibrated confidence scores from language models fine-tuned with human feedback. arXiv:2305.14975. Retrieved from https:\/\/arxiv.org\/abs\/2305.14975"},{"key":"e_1_3_2_205_2","unstructured":"Christian Tomani Kamalika Chaudhuri Ivan Evtimov Daniel Cremers and Mark Ibrahim. 2024. Uncertainty-based abstention in LLMs improves safety and reduces hallucinations. arXiv:2404.10960. Retrieved from https:\/\/arxiv.org\/abs\/2404.10960"},{"key":"e_1_3_2_206_2","unstructured":"S. M. Tonmoy S. M. Zaman Vinija Jain Anku Rani Vipula Rawte Aman Chadha and Amitava Das. 2024. A comprehensive survey of hallucination mitigation techniques in large language models. arXiv:2401.01313. Retrieved from https:\/\/arxiv.org\/abs\/2401.01313"},{"key":"e_1_3_2_207_2","unstructured":"Hugo Touvron Louis Martin Kevin Stone Peter Albert Amjad Almahairi Yasmine Babaei Nikolay Bashlykov Soumya Batra Prajjwal Bhargava Shruti Bhosale et\u00a0al. 2023. Llama 2: Open foundation and fine-tuned chat models. arXiv:2307.09288. Retrieved from https:\/\/arxiv.org\/abs\/2307.09288"},{"key":"e_1_3_2_208_2","unstructured":"Yao-Hung Hubert Tsai Walter Talbott and Jian Zhang. 2024. Efficient non-parametric uncertainty quantification for black-box large language models and decision planning. arXiv:2402.00251. Retrieved from https:\/\/arxiv.org\/abs\/2402.00251"},{"key":"e_1_3_2_209_2","unstructured":"Dennis Ulmer Martin Gubri Hwaran Lee Sangdoo Yun and Seong Joon Oh. 2024. Calibrating large language models using their generations only. arXiv:2403.05973. Retrieved from https:\/\/arxiv.org\/abs\/2403.05973"},{"key":"e_1_3_2_210_2","doi-asserted-by":"crossref","unstructured":"Roman Vashurin Ekaterina Fadeeva Artem Vazhentsev Akim Tsvigun Daniil Vasilev Rui Xing Abdelrahman Boda Sadallah Lyudmila Rvanova Sergey Petrakov Alexander Panchenko et\u00a0al. 2024. Benchmarking uncertainty quantification methods for large language models with LM-polygraph. arXiv:2406.15627. Retrieved from https:\/\/arxiv.org\/abs\/2406.15627","DOI":"10.1162\/tacl_a_00737"},{"key":"e_1_3_2_211_2","article-title":"Attention is all you need","author":"Vaswani A.","year":"2017","unstructured":"A. Vaswani. 2017. Attention is all you need. Advances in Neural Information Processing Systems 30 (2017).","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_2_212_2","unstructured":"Artem Vazhentsev Ekaterina Fadeeva Rui Xing Alexander Panchenko Preslav Nakov Timothy Baldwin Maxim Panov and Artem Shelmanov. 2024. Unconditional truthfulness: Learning conditional dependency for uncertainty quantification of large language models. arXiv:2408.10692. Retrieved from https:\/\/arxiv.org\/abs\/2408.10692"},{"key":"e_1_3_2_213_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-01237-3_34"},{"key":"e_1_3_2_214_2","first-page":"11052","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Wang Hanjing","year":"2024","unstructured":"Hanjing Wang and Qiang Ji. 2024. Epistemic uncertainty quantification for pre-trained neural networks. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition. 11052\u201311061."},{"key":"e_1_3_2_215_2","unstructured":"Jun Wang Guocheng He and Yiannis Kantaros. 2024. Safe task planning for language-instructed multi-robot systems using conformal prediction. arXiv:2402.15368. Retrieved from https:\/\/arxiv.org\/abs\/2402.15368"},{"key":"e_1_3_2_216_2","unstructured":"J. Wang Jiaming Tong Kai Liang Tan Yevgeniy Vorobeychik and Yiannis Kantaros. 2023. Conformal temporal logic planning using large language models: Knowing when to do what and when to ask for help. arXiv:2309.10092. Retrieved from https:\/\/arxiv.org\/abs\/2309.10092"},{"key":"e_1_3_2_217_2","unstructured":"Xi Wang Laurence Aitchison and Maja Rudolph. 2023. LoRA ensembles for large language model fine-tuning. arXiv:2310.00035. Retrieved from https:\/\/arxiv.org\/abs\/2310.00035"},{"key":"e_1_3_2_218_2","unstructured":"Yiming Wang Pei Zhang Baosong Yang Derek F Wong and Rui Wang. 2024. Latent space chain-of-embedding enables output-free LLM self-evaluation. arXiv:2410.13640. Retrieved from https:\/\/arxiv.org\/abs\/2410.13640"},{"key":"e_1_3_2_219_2","unstructured":"Yu-Hsiang Wang Andrew Bai Che-Ping Tsai and Cho-Jui Hsieh. 2024. CLUE: Concept-level uncertainty estimation for large language models. arXiv:2409.03021. Retrieved from https:\/\/arxiv.org\/abs\/2409.03021"},{"key":"e_1_3_2_220_2","doi-asserted-by":"crossref","unstructured":"Zhiyuan Wang Jinhao Duan Lu Cheng Yue Zhang Qingni Wang Hengtao Shen Xiaofeng Zhu Xiaoshuang Shi and Kaidi Xu. 2024. ConU: Conformal uncertainty in large language models with correctness coverage guarantees. arXiv:2407.00499. Retrieved from https:\/\/arxiv.org\/abs\/2407.00499","DOI":"10.18653\/v1\/2024.findings-emnlp.404"},{"key":"e_1_3_2_221_2","first-page":"24824","article-title":"Chain-of-thought prompting elicits reasoning in large language models","volume":"35","author":"Wei Jason","year":"2022","unstructured":"Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V. Le, and Denny Zhou. 2022. Chain-of-thought prompting elicits reasoning in large language models. Advances in Neural Information Processing Systems 35 (2022), 24824\u201324837.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_2_222_2","unstructured":"Adina Williams Nikita Nangia and Samuel R. Bowman. 2017. A broad-coverage challenge corpus for sentence understanding through inference. arXiv:1704.05426. Retrieved from https:\/\/arxiv.org\/abs\/1704.05426"},{"key":"e_1_3_2_223_2","first-page":"3376","volume-title":"Proceedings of the International Conference on Artificial Intelligence and Statistics","author":"Wu Luhuan","year":"2024","unstructured":"Luhuan Wu and Sinead A. Williamson. 2024. Posterior uncertainty quantification in neural networks using data augmentation. In Proceedings of the International Conference on Artificial Intelligence and Statistics. PMLR, 3376\u20133384."},{"key":"e_1_3_2_224_2","unstructured":"Yijun Xiao and William Yang Wang. 2021. On hallucination and predictive uncertainty in conditional language generation. arXiv:2103.15025. Retrieved from https:\/\/arxiv.org\/abs\/2103.15025"},{"key":"e_1_3_2_225_2","doi-asserted-by":"publisher","DOI":"10.1145\/3581754.3584136"},{"key":"e_1_3_2_226_2","unstructured":"Miao Xiong Zhiyuan Hu Xinyang Lu Yifei Li Jie Fu Junxian He and Bryan Hooi. 2023. Can llms express their uncertainty? An empirical evaluation of confidence elicitation in llms. arXiv:2306.13063. Retrieved from https:\/\/arxiv.org\/abs\/2306.13063"},{"key":"e_1_3_2_227_2","unstructured":"Tianyang Xu Shujin Wu Shizhe Diao Xiaoze Liu Xingyao Wang Yangyi Chen and Jing Gao. 2024. SaySelf: Teaching LLMs to express confidence with self-reflective rationales. arXiv:2405.20974. Retrieved from https:\/\/arxiv.org\/abs\/2405.20974"},{"key":"e_1_3_2_228_2","unstructured":"Ziwei Xu Sanjay Jain and Mohan Kankanhalli. 2024. Hallucination is inevitable: An innate limitation of large language models. arXiv:2401.11817. Retrieved from https:\/\/arxiv.org\/abs\/2401.11817"},{"key":"e_1_3_2_229_2","unstructured":"Yasin Abbasi Yadkori Ilja Kuzborskij Andr\u00e1s Gy\u00f6rgy and Csaba Szepesv\u00e1ri. 2024. To believe or not to believe your LLM. arXiv:2406.02543. Retrieved from https:\/\/arxiv.org\/abs\/2406.02543"},{"key":"e_1_3_2_230_2","unstructured":"Adam X. Yang Maxime Robeyns Xi Wang and Laurence Aitchison. 2024. Bayesian low-rank adaptation for large language models. arXiv:2308.13111. Retrieved from https:\/\/arxiv.org\/abs\/2308.13111"},{"key":"e_1_3_2_231_2","unstructured":"Haoyan Yang Yixuan Wang Xingyin Xu Hanyuan Zhang and Yirong Bian. 2024. Can we trust LLMs? Mitigate overconfidence bias in LLMs through knowledge transfer. arXiv:2405.16856. Retrieved from https:\/\/arxiv.org\/abs\/2405.16856"},{"key":"e_1_3_2_232_2","unstructured":"Yuqing Yang Ethan Chern Xipeng Qiu Graham Neubig and Pengfei Liu. 2023. Alignment for honesty. arXiv:2312.07000. Retrieved from https:\/\/arxiv.org\/abs\/2312.07000"},{"key":"e_1_3_2_233_2","unstructured":"Yuchen Yang Houqiang Li Yanfeng Wang and Yu Wang. 2023. Improving the reliability of large language models by leveraging uncertainty-aware in-context learning. arXiv:2310.04782. Retrieved from https:\/\/arxiv.org\/abs\/2310.04782"},{"key":"e_1_3_2_234_2","doi-asserted-by":"crossref","unstructured":"Zhilin Yang Peng Qi Saizheng Zhang Yoshua Bengio William W. Cohen Ruslan Salakhutdinov and Christopher D. Manning. 2018. HotpotQA: A dataset for diverse explainable multi-hop question answering. arXiv:1809.09600. Retrieved from https:\/\/arxiv.org\/abs\/1809.09600","DOI":"10.18653\/v1\/D18-1259"},{"key":"e_1_3_2_235_2","unstructured":"Fanghua Ye Mingming Yang Jianhui Pang Longyue Wang Derek F. Wong Emine Yilmaz Shuming Shi and Zhaopeng Tu. 2024. Benchmarking llms via uncertainty quantification. arXiv:2401.12794. Retrieved from https:\/\/arxiv.org\/abs\/2401.12794"},{"key":"e_1_3_2_236_2","doi-asserted-by":"crossref","unstructured":"Gal Yona Roee Aharoni and Mor Geva. 2024. Can large language models faithfully express their intrinsic uncertainty in words? arXiv:2405.16908. Retrieved from https:\/\/arxiv.org\/abs\/2405.16908","DOI":"10.18653\/v1\/2024.emnlp-main.443"},{"key":"e_1_3_2_237_2","unstructured":"Lei Yu Meng Cao Jackie Chi Kit Cheung and Yue Dong. 2024. Mechanisms of non-factual hallucinations in language models. arXiv:2403.18167. Retrieved from https:\/\/arxiv.org\/abs\/2403.18167"},{"key":"e_1_3_2_238_2","first-page":"27263","article-title":"Bartscore: Evaluating generated text as text generation","volume":"34","author":"Yuan Weizhe","year":"2021","unstructured":"Weizhe Yuan, Graham Neubig, and Pengfei Liu. 2021. Bartscore: Evaluating generated text as text generation. Advances in Neural Information Processing Systems 34 (2021), 27263\u201327277.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_2_239_2","doi-asserted-by":"crossref","unstructured":"Zeyu Yun Yubei Chen Bruno A. Olshausen and Yann LeCun. 2021. Transformer visualization via dictionary learning: Contextualized embedding as a linear superposition of transformer factors. arXiv:2103.15949. Retrieved from https:\/\/arxiv.org\/abs\/2103.15949","DOI":"10.18653\/v1\/2021.deelio-1.1"},{"key":"e_1_3_2_240_2","first-page":"609","volume-title":"Proceedings of the Icml","author":"Zadrozny Bianca","year":"2001","unstructured":"Bianca Zadrozny and Charles Elkan. 2001. Obtaining calibrated probability estimates from decision trees and naive bayesian classifiers. In Proceedings of the Icml. 609\u2013616."},{"key":"e_1_3_2_241_2","doi-asserted-by":"publisher","DOI":"10.1145\/775047.775151"},{"key":"e_1_3_2_242_2","doi-asserted-by":"crossref","unstructured":"Rowan Zellers Ari Holtzman Yonatan Bisk Ali Farhadi and Yejin Choi. 2019. Hellaswag: Can a machine really finish your sentence? arXiv:1905.07830. Retrieved from https:\/\/arxiv.org\/abs\/1905.07830","DOI":"10.18653\/v1\/P19-1472"},{"key":"e_1_3_2_243_2","unstructured":"Qingcheng Zeng Mingyu Jin Qinkai Yu Zhenting Wang Wenyue Hua Zihao Zhou Guangyan Sun Yanda Meng Shiqing Ma Qifan Wang et\u00a0al. 2024. Uncertainty is fragile: Manipulating uncertainty in large language models. arXiv:2407.11282. Retrieved from https:\/\/arxiv.org\/abs\/2407.11282"},{"key":"e_1_3_2_244_2","doi-asserted-by":"crossref","unstructured":"Caiqi Zhang Fangyu Liu Marco Basaldella and Nigel Collier. 2024. LUQ: Long-text uncertainty quantification for LLMs. arXiv:2403.20279. Retrieved from https:\/\/arxiv.org\/abs\/2403.20279","DOI":"10.18653\/v1\/2024.emnlp-main.299"},{"key":"e_1_3_2_245_2","unstructured":"Tianyi Zhang Varsha Kishore Felix Wu Kilian Q. Weinberger and Yoav Artzi. 2019. Bertscore: Evaluating text generation with bert. arXiv:1904.09675. Retrieved from https:\/\/arxiv.org\/abs\/1904.09675"},{"key":"e_1_3_2_246_2","doi-asserted-by":"crossref","unstructured":"Tianhang Zhang Lin Qiu Qipeng Guo Cheng Deng Yue Zhang Zheng Zhang Chenghu Zhou Xinbing Wang and Luoyi Fu. 2023. Enhancing uncertainty-based hallucination detection with stronger focus. arXiv:2311.13230. Retrieved from https:\/\/arxiv.org\/abs\/2311.13230","DOI":"10.18653\/v1\/2023.emnlp-main.58"},{"key":"e_1_3_2_247_2","doi-asserted-by":"crossref","unstructured":"Yuwei Zhang Zihan Wang and Jingbo Shang. 2023. Clusterllm: Large language models as a guide for text clustering. arXiv:2305.14871. Retrieved from https:\/\/arxiv.org\/abs\/2305.14871","DOI":"10.18653\/v1\/2023.emnlp-main.858"},{"key":"e_1_3_2_248_2","doi-asserted-by":"publisher","DOI":"10.1145\/3639372"},{"key":"e_1_3_2_249_2","unstructured":"Qiwei Zhao Xujiang Zhao Yanchi Liu Wei Cheng Yiyou Sun Mika Oishi Takao Osaki Katsushi Matsuda Huaxiu Yao and Haifeng Chen. 2024. SAUP: Situation awareness uncertainty propagation on LLM agent. arXiv:2412.01033. Retrieved from https:\/\/arxiv.org\/abs\/2412.01033"},{"key":"e_1_3_2_250_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2024.acl-long.566"},{"key":"e_1_3_2_251_2","doi-asserted-by":"crossref","unstructured":"Xinran Zhao Hongming Zhang Xiaoman Pan Wenlin Yao Dong Yu Tongshuang Wu and Jianshu Chen. 2024. Fact-and-reflection (FaR) improves confidence calibration of large language models. arXiv:2402.17124. Retrieved from https:\/\/arxiv.org\/abs\/2402.17124","DOI":"10.18653\/v1\/2024.findings-acl.515"},{"key":"e_1_3_2_252_2","volume-title":"Proceedings of the 11th International Conference on Learning Representations","author":"Zhao Yao","year":"2022","unstructured":"Yao Zhao, Mikhail Khalman, Rishabh Joshi, Shashi Narayan, Mohammad Saleh, and Peter J. Liu. 2022. Calibrating sequence likelihood improves conditional language generation. In Proceedings of the 11th International Conference on Learning Representations."},{"key":"e_1_3_2_253_2","first-page":"12697","volume-title":"Proceedings of the International Conference on Machine Learning","author":"Zhao Zihao","year":"2021","unstructured":"Zihao Zhao, Eric Wallace, Shi Feng, Dan Klein, and Sameer Singh. 2021. Calibrate before use: Improving few-shot performance of language models. In Proceedings of the International Conference on Machine Learning. PMLR, 12697\u201312706."},{"key":"e_1_3_2_254_2","unstructured":"Zhi Zheng Qian Feng Hang Li Alois Knoll and Jianxiang Feng. 2024. Evaluating uncertainty-based failure detection for closed-loop LLM planners. arXiv:2406.00430. Retrieved from https:\/\/arxiv.org\/abs\/2406.00430"},{"key":"e_1_3_2_255_2","unstructured":"Chiwei Zhu Benfeng Xu Quan Wang Yongdong Zhang and Zhendong Mao. 2023. On the calibration of large language models and alignment. arXiv:2311.13240. Retrieved from https:\/\/arxiv.org\/abs\/2311.13240"},{"key":"e_1_3_2_256_2","article-title":"Scale alone does not improve mechanistic interpretability in vision models","volume":"36","author":"Zimmermann Roland S.","year":"2024","unstructured":"Roland S. Zimmermann, Thomas Klein, and Wieland Brendel. 2024. Scale alone does not improve mechanistic interpretability in vision models. Advances in Neural Information Processing Systems 36 (2024).","journal-title":"Advances in Neural Information Processing Systems"}],"container-title":["ACM Computing Surveys"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3744238","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,9,9]],"date-time":"2025-09-09T14:32:58Z","timestamp":1757428378000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3744238"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,9,9]]},"references-count":255,"journal-issue":{"issue":"3","published-print":{"date-parts":[[2026,2,28]]}},"alternative-id":["10.1145\/3744238"],"URL":"https:\/\/doi.org\/10.1145\/3744238","relation":{},"ISSN":["0360-0300","1557-7341"],"issn-type":[{"value":"0360-0300","type":"print"},{"value":"1557-7341","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,9,9]]},"assertion":[{"value":"2024-12-07","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-05-28","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-09-09","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}