{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,7]],"date-time":"2026-07-07T04:47:05Z","timestamp":1783399625586,"version":"3.54.6"},"reference-count":70,"publisher":"Association for Computing Machinery (ACM)","license":[{"start":{"date-parts":[[2024,10,29]],"date-time":"2024-10-29T00:00:00Z","timestamp":1730160000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Comput.-Hum. Interact."],"abstract":"<jats:p>Large-scale generative models have enabled the development of AI-powered code completion tools to assist programmers in writing code. Like all AI-powered tools, these code completion tools are not always accurate and can introduce bugs or even security vulnerabilities into code if not properly detected and corrected by a human programmer. One technique that has been proposed and implemented to help programmers locate potential errors is to highlight uncertain tokens. However, little is known about the effectiveness of this technique. Through a mixed-methods study with 30 programmers, we compare three conditions: providing the AI system's code completion alone, highlighting tokens with the lowest likelihood of being generated by the underlying generative model, and highlighting tokens with the highest predicted likelihood of being edited by a programmer. We find that highlighting tokens with the highest predicted likelihood of being edited leads to faster task completion and more targeted edits, and is subjectively preferred by study participants. In contrast, highlighting tokens according to their probability of being generated does not provide any benefit over the baseline with no highlighting. We further explore the design space of how to convey uncertainty in AI-powered code completion tools and find that programmers prefer highlights that are granular, informative, interpretable, and not overwhelming. This work contributes to building an understanding of what uncertainty means for generative models and how to convey it effectively.<\/jats:p>","DOI":"10.1145\/3702320","type":"journal-article","created":{"date-parts":[[2024,10,29]],"date-time":"2024-10-29T20:49:18Z","timestamp":1730234958000},"update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":12,"title":["Generation Probabilities Are Not Enough: Uncertainty Highlighting in AI Code Completions"],"prefix":"10.1145","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-6649-6905","authenticated-orcid":false,"given":"Helena","family":"Vasconcelos","sequence":"first","affiliation":[{"name":"Stanford University, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-7741-3861","authenticated-orcid":false,"given":"Gagan","family":"Bansal","sequence":"additional","affiliation":[{"name":"Microsoft Research, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-4986-7794","authenticated-orcid":false,"given":"Adam","family":"Fourney","sequence":"additional","affiliation":[{"name":"Microsoft Research, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-4543-7196","authenticated-orcid":false,"given":"Q. Vera","family":"Liao","sequence":"additional","affiliation":[{"name":"Microsoft Research, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-7807-2018","authenticated-orcid":false,"given":"Jennifer Wortman","family":"Vaughan","sequence":"additional","affiliation":[{"name":"Microsoft Research, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2024,10,29]]},"reference":[{"key":"e_1_2_1_1_1","volume-title":"Proceedings of the 37th IEEE\/ACM International Conference on Automated Software Engineering. 1\u20135.","author":"Madi Naser Al","year":"2022","unstructured":"Naser Al Madi. 2022. How readable is model-generated code? examining readability and visual inspection of github copilot. In Proceedings of the 37th IEEE\/ACM International Conference on Automated Software Engineering. 1\u20135."},{"key":"e_1_2_1_2_1","volume-title":"Retrieved","author":"Services Amazon Web","year":"2022","unstructured":"Amazon Web Services. 2022. ML-powered coding companion - Amazon CodeWhisperer. Retrieved September, 2022 from https:\/\/aws.amazon.com\/codewhisperer\/"},{"key":"e_1_2_1_3_1","unstructured":"Julia Angwin Jeff Larson Surya Mattu and Lauren Kirchner. 2016. Machine bias: There's software across the country to predict future criminals and it's biased against blacks. (2016)."},{"key":"e_1_2_1_4_1","volume-title":"Proceedings of the 2021 CHI Conference on Human Factors in Computing Systems. 1\u201316","author":"Bansal Gagan","year":"2021","unstructured":"Gagan Bansal, Tongshuang Wu, Joyce Zhou, Raymond Fok, Besmira Nushi, Ece Kamar, Marco Tulio Ribeiro, and Daniel Weld. 2021. Does the whole exceed its parts? the effect of ai explanations on complementary team performance. In Proceedings of the 2021 CHI Conference on Human Factors in Computing Systems. 1\u201316."},{"key":"e_1_2_1_5_1","volume-title":"Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency (FAccT). 610\u2013623","author":"Bender Emily M.","year":"2021","unstructured":"Emily M. Bender, Timnit Gebru, Angelina McMillan-Major, and Shmargaret Shmitchell. 2021. On the Dangers of Stochastic Parrots: Can Language Models Be Too Big?. In Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency (FAccT). 610\u2013623. https:\/\/doi.org\/10.1145\/3442188.3445922"},{"key":"e_1_2_1_6_1","first-page":"1137","article-title":"A neural probabilistic language model","author":"Bengio Yoshua","year":"2003","unstructured":"Yoshua Bengio, R\u00e9jean Ducharme, Pascal Vincent, and Christian Jauvin. 2003. A neural probabilistic language model. Journal of Machine Learning Research 3, Feb (2003), 1137\u20131155.","journal-title":"Journal of Machine Learning Research 3"},{"key":"e_1_2_1_7_1","volume-title":"Proceedings of the 2021 AAAI\/ACM Conference on AI, Ethics, and Society. 401\u2013413","author":"Bhatt Umang","year":"2021","unstructured":"Umang Bhatt, Javier Antor\u00e1n, Yunfeng Zhang, Q Vera Liao, Prasanna Sattigeri, Riccardo Fogliato, Gabrielle Melan\u00e7on, Ranganath Krishnan, Jason Stanley, Omesh Tickoo, et\u00a0al. 2021. Uncertainty as a form of transparency: Measuring, communicating, and using uncertainty. In Proceedings of the 2021 AAAI\/ACM Conference on AI, Ethics, and Society. 401\u2013413."},{"key":"e_1_2_1_8_1","volume-title":"et\u00a0al","author":"Brown Tom","year":"2020","unstructured":"Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et\u00a0al. 2020. Language models are few-shot learners. Advances in neural information processing systems 33 (2020), 1877\u20131901."},{"key":"e_1_2_1_9_1","volume-title":"Proceedings of the ACM on Human-computer Interaction 3, CSCW","author":"Cai Carrie J","year":"2019","unstructured":"Carrie J Cai, Samantha Winter, David Steiner, Lauren Wilcox, and Michael Terry. 2019. \u201d Hello AI\u201d: Uncovering the Onboarding Needs of Medical Practitioners for Human-AI Collaborative Decision-Making. Proceedings of the ACM on Human-computer Interaction 3, CSCW (2019), 1\u201324."},{"key":"e_1_2_1_10_1","volume-title":"Extended Abstracts of the 2021 CHI Conference on Human Factors in Computing Systems","author":"Cai Carrie J.","year":"2021","unstructured":"Carrie J. Cai, Samantha Winter, David F. Steiner, Lauren Wilcox, and Michael Terry. 2021. Onboarding Materials as Cross-functional Boundary Objects for Developing AI Assistants. Extended Abstracts of the 2021 CHI Conference on Human Factors in Computing Systems (2021)."},{"key":"e_1_2_1_11_1","volume-title":"Help me write a poem: Instruction Tuning as a Vehicle for Collaborative Poetry Writing. arXiv preprint arXiv:2210.13669","author":"Chakrabarty Tuhin","year":"2022","unstructured":"Tuhin Chakrabarty, Vishakh Padmakumar, and He He. 2022. Help me write a poem: Instruction Tuning as a Vehicle for Collaborative Poetry Writing. arXiv preprint arXiv:2210.13669 (2022)."},{"key":"e_1_2_1_12_1","volume-title":"23rd International Conference on Intelligent User Interfaces. 329\u2013340","author":"Clark Elizabeth","year":"2018","unstructured":"Elizabeth Clark, Anne Spencer Ross, Chenhao Tan, Yangfeng Ji, and Noah A Smith. 2018. Creative writing with a machine in the loop: Case studies on slogans and stories. In 23rd International Conference on Intelligent User Interfaces. 329\u2013340."},{"key":"e_1_2_1_13_1","volume-title":"Retrieved","year":"2022","unstructured":"DeepMind. 2022. AlphaCode. Retrieved September, 2022 from https:\/\/alphacode.deepmind.com\/"},{"key":"e_1_2_1_14_1","volume-title":"Communicating uncertainty using words and numbers. Trends in Cognitive Sciences","author":"Dhami Mandeep K","year":"2022","unstructured":"Mandeep K Dhami and David R Mandel. 2022. Communicating uncertainty using words and numbers. Trends in Cognitive Sciences (2022)."},{"key":"e_1_2_1_15_1","doi-asserted-by":"crossref","first-page":"20","DOI":"10.1145\/3454122.3454124","article-title":"The SPACE of Developer Productivity: There's more to it than you think","volume":"19","author":"Forsgren Nicole","year":"2021","unstructured":"Nicole Forsgren, Margaret-Anne Storey, Chandra Maddila, Thomas Zimmermann, Brian Houck, and Jenna Butler. 2021. The SPACE of Developer Productivity: There's more to it than you think. Queue 19, 1 (2021), 20\u201348.","journal-title":"Queue"},{"key":"e_1_2_1_16_1","volume-title":"Retrieved","year":"2022","unstructured":"GitHub. 2022. GitHub Copilot - Your AI pair programmer. Retrieved September, 2022 from https:\/\/github.com\/features\/copilot\/"},{"key":"e_1_2_1_17_1","doi-asserted-by":"crossref","unstructured":"Ana Valeria Gonzalez Gagan Bansal Angela Fan Yashar Mehdad Robin Jia and Srini Iyer. 2021. Do Explanations Help Users Detect Errors in Open-Domain QA? An Evaluation of Spoken vs. Visual Explanations. In Findings of ACL.","DOI":"10.18653\/v1\/2021.findings-acl.95"},{"key":"e_1_2_1_18_1","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/3479562","article-title":"Algorithmic Risk Assessments Can Alter Human Decision-Making Processes in High-Stakes Government Contexts","volume":"5","author":"Green Ben","year":"2020","unstructured":"Ben Green and Yiling Chen. 2020. Algorithmic Risk Assessments Can Alter Human Decision-Making Processes in High-Stakes Government Contexts. Proceedings of the ACM on Human-Computer Interaction 5 (2020), 1 \u2013 33.","journal-title":"Proceedings of the ACM on Human-Computer Interaction"},{"key":"e_1_2_1_19_1","volume-title":"International conference on machine learning. PMLR, 1321\u20131330","author":"Guo Chuan","year":"2017","unstructured":"Chuan Guo, Geoff Pleiss, Yu Sun, and Kilian Q Weinberger. 2017. On calibration of modern neural networks. In International conference on machine learning. PMLR, 1321\u20131330."},{"key":"e_1_2_1_20_1","volume-title":"Proceedings of the human factors and ergonomics society annual meeting","volume":"50","author":"Hart Sandra G","year":"2006","unstructured":"Sandra G Hart. 2006. NASA-task load index (NASA-TLX); 20 years later. In Proceedings of the human factors and ergonomics society annual meeting, Vol. 50. Sage publications Sage CA: Los Angeles, CA, 904\u2013908."},{"key":"e_1_2_1_21_1","volume-title":"Companion of the 2017 ACM Conference on Computer Supported Cooperative Work and Social Computing. 195\u2013198","author":"Hayashi Yugo","year":"2017","unstructured":"Yugo Hayashi and Kosuke Wakabayashi. 2017. Can AI become reliable source to support human decision making in a court scene?. In Companion of the 2017 ACM Conference on Computer Supported Cooperative Work and Social Computing. 195\u2013198."},{"key":"e_1_2_1_22_1","volume-title":"Thomas H. McCoy, Roy H. Perlis, Finale Doshi-Velez, and Krzysztof Z Gajos.","author":"Jacobs Maia L.","year":"2021","unstructured":"Maia L. Jacobs, Melanie Fernandes Pradier, Thomas H. McCoy, Roy H. Perlis, Finale Doshi-Velez, and Krzysztof Z Gajos. 2021. How machine-learning recommendations influence clinician treatment selections: the example of antidepressant selection. Translational Psychiatry 11 (2021)."},{"key":"e_1_2_1_23_1","doi-asserted-by":"crossref","unstructured":"Kevin Jesse Toufique Ahmed Premkumar T. Devanbu and Emily Morgan. 2023. Large Language Models and Simple Stupid Bugs. arXiv:2303.11455 [cs.SE]","DOI":"10.1109\/MSR59073.2023.00082"},{"key":"e_1_2_1_24_1","doi-asserted-by":"publisher","DOI":"10.1145\/3571730"},{"key":"e_1_2_1_25_1","doi-asserted-by":"crossref","first-page":"962","DOI":"10.1162\/tacl_a_00407","article-title":"How can we know when language models know? on the calibration of language models for question answering","volume":"9","author":"Jiang Zhengbao","year":"2021","unstructured":"Zhengbao Jiang, Jun Araki, Haibo Ding, and Graham Neubig. 2021. How can we know when language models know? on the calibration of language models for question answering. Transactions of the Association for Computational Linguistics 9 (2021), 962\u2013977.","journal-title":"Transactions of the Association for Computational Linguistics"},{"key":"e_1_2_1_26_1","volume-title":"arXiv preprint arXiv:2303.00732","author":"Johnson Daniel D","year":"2023","unstructured":"Daniel D Johnson, Daniel Tarlow, and Christian Walder. 2023. RU-SURE? Uncertainty-Aware Code Suggestions By Maximizing Utility Across Random User Intents. arXiv preprint arXiv:2303.00732 (2023)."},{"key":"e_1_2_1_27_1","volume-title":"et\u00a0al","author":"Kadavath Saurav","year":"2022","unstructured":"Saurav Kadavath, Tom Conerly, Amanda Askell, Tom Henighan, Dawn Drain, Ethan Perez, Nicholas Schiefer, Zac Hatfield-Dodds, Nova DasSarma, Eli Tran-Johnson, et\u00a0al. 2022. Language models (mostly) know what they know. arXiv preprint arXiv:2207.05221 (2022)."},{"key":"e_1_2_1_28_1","unstructured":"Eirini Kalliamvakou. 2022. Research: quantifying GitHub Copilot's impact on developer productivity and happiness. https:\/\/github.blog\/2022-09-07-research-quantifying-github-copilots-impact-on-developer-productivity-and-happiness\/"},{"key":"e_1_2_1_29_1","unstructured":"Adam Khakhar Stephen Mell and Osbert Bastani. 2023. PAC Prediction Sets for Large Language Models of Code. arXiv:2302.08703 [cs.LG]"},{"key":"e_1_2_1_30_1","volume-title":"Proceedings of the 2019 CHI Conference on Human Factors in Computing Systems","author":"Kocielnik Rafal","year":"2019","unstructured":"Rafal Kocielnik, Saleema Amershi, and Paul N. Bennett. 2019. Will You Accept an Imperfect AI?: Exploring Designs for Adjusting End-user Expectations of AI Systems. Proceedings of the 2019 CHI Conference on Human Factors in Computing Systems (2019)."},{"key":"e_1_2_1_31_1","unstructured":"Maia Kotelanski Robert Gallo Ashwin Nayak and Thomas Savage. 2023. Methods to Estimate Large Language Model Confidence. arXiv:2312.03733 [cs.CL]"},{"key":"e_1_2_1_32_1","volume-title":"Semantic uncertainty: Linguistic invariances for uncertainty estimation in natural language generation. arXiv preprint arXiv:2302.09664","author":"Kuhn Lorenz","year":"2023","unstructured":"Lorenz Kuhn, Yarin Gal, and Sebastian Farquhar. 2023. Semantic uncertainty: Linguistic invariances for uncertainty estimation in natural language generation. arXiv preprint arXiv:2302.09664 (2023)."},{"key":"e_1_2_1_33_1","volume-title":"Proceedings of the 2020 CHI Conference on Human Factors in Computing Systems","author":"Lai Vivian","year":"2020","unstructured":"Vivian Lai, Han Liu, and Chenhao Tan. 2020. \u201dWhy is \u2019Chicago\u2019 deceptive?\u201d Towards Building Model-Driven Tutorials for Humans. Proceedings of the 2020 CHI Conference on Human Factors in Computing Systems (2020)."},{"key":"e_1_2_1_34_1","volume-title":"Proceedings of the SIGCHI Conference on Human Factors in Computing Systems. 1\u20134.","author":"Lank Edward","year":"2010","unstructured":"Edward Lank, Ryan Stedman, and Michael Terry. 2010. Estimating residual error rate in recognized handwritten documents using artificial error injection. In Proceedings of the SIGCHI Conference on Human Factors in Computing Systems. 1\u20134."},{"key":"e_1_2_1_35_1","unstructured":"LeetCode. 2015. The world's leading online programming learning platform. https:\/\/leetcode.com\/"},{"key":"e_1_2_1_36_1","volume-title":"Achmadnoer Sukma Wicaksana, Marise Ph Born, and Cornelius J K\u00f6nig.","author":"Liem Cynthia CS","year":"2018","unstructured":"Cynthia CS Liem, Markus Langer, Andrew Demetriou, Annemarie MF Hiemstra, Achmadnoer Sukma Wicaksana, Marise Ph Born, and Cornelius J K\u00f6nig. 2018. Psychology Meets Machine Learning: Interdisciplinary Perspectives on Algorithmic Job Candidate Screening. In Explainable and Interpretable Models in Computer Vision and Machine Learning. Springer, 197\u2013253."},{"key":"e_1_2_1_37_1","doi-asserted-by":"crossref","first-page":"1501","DOI":"10.1007\/s11145-017-9734-4","article-title":"Effects of spell checkers on English as a second language students\u2019 incidental spelling learning: a cognitive load perspective","volume":"30","author":"Lin Po-Han","year":"2017","unstructured":"Po-Han Lin, Tzu-Chien Liu, and Fred Paas. 2017. Effects of spell checkers on English as a second language students\u2019 incidental spelling learning: a cognitive load perspective. Reading and Writing 30 (2017), 1501\u20131525.","journal-title":"Reading and Writing"},{"key":"e_1_2_1_38_1","volume-title":"Teaching models to express their uncertainty in words. arXiv preprint arXiv:2205.14334","author":"Lin Stephanie","year":"2022","unstructured":"Stephanie Lin, Jacob Hilton, and Owain Evans. 2022. Teaching models to express their uncertainty in words. arXiv preprint arXiv:2205.14334 (2022)."},{"key":"e_1_2_1_39_1","unstructured":"Genglin Liu Xingyao Wang Lifan Yuan Yangyi Chen and Hao Peng. 2023. Prudent Silence or Foolish Babble? Examining Large Language Models\u2019 Responses to the Unknown. arXiv:2311.09731 [cs.CL]"},{"key":"e_1_2_1_40_1","volume-title":"Proceedings of the 35th Annual ACM Symposium on User Interface Software and Technology","author":"Liu Vivian","year":"2022","unstructured":"Vivian Liu, Han Qiao, and Lydia B. Chilton. 2022. Opal: Multimodal Image Generation for News Illustration. Proceedings of the 35th Annual ACM Symposium on User Interface Software and Technology (2022)."},{"key":"e_1_2_1_41_1","volume-title":"Proceedings of the 2020 CHI Conference on Human Factors in Computing Systems","author":"Louie Ryan","year":"2020","unstructured":"Ryan Louie, Andy Coenen, Cheng-Zhi Anna Huang, Michael Terry, and Carrie J. Cai. 2020. Novice-AI Music Co-Creation via AI-Steering Tools for Deep Generative Models. Proceedings of the 2020 CHI Conference on Human Factors in Computing Systems (2020)."},{"key":"e_1_2_1_42_1","volume-title":"Shu-Fang Newman, Jerry Kim, et\u00a0al.","author":"Lundberg Scott M","year":"2018","unstructured":"Scott M Lundberg, Bala Nair, Monica S Vavilala, Mayumi Horibe, Michael J Eisses, Trevor Adams, David E Liston, Daniel King-Wai Low, Shu-Fang Newman, Jerry Kim, et\u00a0al. 2018. Explainable machine-learning predictions for the prevention of hypoxaemia during surgery. Nature biomedical engineering 2, 10 (2018), 749\u2013760."},{"key":"e_1_2_1_43_1","volume-title":"Proceedings of Machine Translation Summit XVII","volume":"243","author":"Martindale Marianna","year":"2019","unstructured":"Marianna Martindale, Marine Carpuat, Kevin Duh, and Paul McNamee. 2019. Identifying fluently inadequate output in neural and statistical machine translation. In Proceedings of Machine Translation Summit XVII Volume 1: Research Track. 233\u2013243."},{"key":"e_1_2_1_44_1","doi-asserted-by":"crossref","first-page":"857","DOI":"10.1162\/tacl_a_00494","article-title":"Reducing conversational agents\u2019 overconfidence through linguistic calibration","volume":"10","author":"Mielke Sabrina J","year":"2022","unstructured":"Sabrina J Mielke, Arthur Szlam, Emily Dinan, and Y-Lan Boureau. 2022. Reducing conversational agents\u2019 overconfidence through linguistic calibration. Transactions of the Association for Computational Linguistics 10 (2022), 857\u2013872.","journal-title":"Transactions of the Association for Computational Linguistics"},{"key":"e_1_2_1_45_1","volume-title":"Reading Between the Lines: Modeling User Behavior and Costs in AI-Assisted Programming. arXiv preprint arXiv:2210.14306","author":"Mozannar Hussein","year":"2022","unstructured":"Hussein Mozannar, Gagan Bansal, Adam Fourney, and Eric Horvitz. 2022. Reading Between the Lines: Modeling User Behavior and Costs in AI-Assisted Programming. arXiv preprint arXiv:2210.14306 (2022)."},{"key":"e_1_2_1_46_1","volume-title":"Sontag","author":"Mozannar Hussein","year":"2021","unstructured":"Hussein Mozannar, Arvindmani Satyanarayan, and David A. Sontag. 2021. Teaching Humans When To Defer to a Classifier via Examplars. In AAAI."},{"key":"e_1_2_1_47_1","volume-title":"Proceedings of the 22nd international conference on Machine learning. 625\u2013632","author":"Niculescu-Mizil Alexandru","year":"2005","unstructured":"Alexandru Niculescu-Mizil and Rich Caruana. 2005. Predicting good probabilities with supervised learning. In Proceedings of the 22nd international conference on Machine learning. 625\u2013632."},{"key":"e_1_2_1_48_1","unstructured":"OpenAI. 2015. https:\/\/beta.openai.com\/playground"},{"key":"e_1_2_1_49_1","volume-title":"Complacency and bias in human use of automation: An attentional integration. Human factors 52, 3","author":"Parasuraman Raja","year":"2010","unstructured":"Raja Parasuraman and Dietrich H Manzey. 2010. Complacency and bias in human use of automation: An attentional integration. Human factors 52, 3 (2010), 381\u2013410."},{"key":"e_1_2_1_50_1","volume-title":"2022 IEEE Symposium on Security and Privacy (SP). 754\u2013768","author":"Pearce Hammond","year":"2022","unstructured":"Hammond Pearce, Baleegh Ahmad, Benjamin Tan, Brendan Dolan-Gavitt, and Ramesh Karri. 2022. Asleep at the Keyboard? Assessing the Security of GitHub Copilot's Code Contributions. In 2022 IEEE Symposium on Security and Privacy (SP). 754\u2013768. https:\/\/doi.org\/10.1109\/SP46214.2022.9833571"},{"key":"e_1_2_1_51_1","volume-title":"Do users write more insecure code with AI assistants? arXiv preprint arXiv:2211.03622","author":"Perry Neil","year":"2022","unstructured":"Neil Perry, Megha Srivastava, Deepak Kumar, and Dan Boneh. 2022. Do users write more insecure code with AI assistants? arXiv preprint arXiv:2211.03622 (2022)."},{"key":"e_1_2_1_52_1","volume-title":"Ernst","author":"Pudari Rohith","year":"2023","unstructured":"Rohith Pudari and Neil A. Ernst. 2023. From Copilot to Pilot: Towards AI Supported Software Development. arXiv:2303.04142 [cs.SE]"},{"key":"e_1_2_1_53_1","unstructured":"Alec Radford Jeffrey Wu Rewon Child David Luan Dario Amodei and Ilya Sutskever. 2019. Language models are unsupervised multitask learners. OpenAI white paper."},{"key":"e_1_2_1_54_1","volume-title":"Sruti Srinivasa Ragavan, and Ben Zorn","author":"Sarkar Advait","year":"2022","unstructured":"Advait Sarkar, Andrew D Gordon, Carina Negreanu, Christian Poelitz, Sruti Srinivasa Ragavan, and Ben Zorn. 2022. What is it like to program with artificial intelligence? arXiv preprint arXiv:2208.06213 (2022)."},{"key":"e_1_2_1_55_1","unstructured":"William Saunders Catherine Yeh Jeff Wu Steven Bills Long Ouyang Jonathan Ward and Jan Leike. 2022. Self-critiquing models for assisting human evaluators. arXiv:2206.05802 [cs.CL]"},{"key":"e_1_2_1_56_1","unstructured":"Vaishnavi Shrivastava Percy Liang and Ananya Kumar. 2023. Llamas Know What GPTs Don\u2019t Show: Surrogate Models for Confidence Estimation. arXiv:2311.08877 [cs.CL]"},{"key":"e_1_2_1_57_1","unstructured":"Aniket Kumar Singh Suman Devkota Bishal Lamichhane Uttam Dhakal and Chandra Dhakal. 2023. The Confidence-Competence Gap in Large Language Models: A Cognitive Study. arXiv:2309.16145 [cs.CL]"},{"key":"e_1_2_1_58_1","volume-title":"27th International Conference on Intelligent User Interfaces","author":"Sun Jiao","unstructured":"Jiao Sun, Q. Vera Liao, Michael Muller, Mayank Agarwal, Stephanie Houde, Kartik Talamadupula, and Justin D. Weisz. 2022. Investigating Explainability of Generative AI for Code through Scenario-Based Design. In 27th International Conference on Intelligent User Interfaces (Helsinki, Finland) (IUI \u201922). Association for Computing Machinery, New York, NY, USA, 212\u2013228. https:\/\/doi.org\/10.1145\/3490099.3511119"},{"key":"e_1_2_1_59_1","unstructured":"Sree Harsha Tanneru Chirag Agarwal and Himabindu Lakkaraju. 2023. Quantifying Uncertainty in Natural Language Explanations of Large Language Models. arXiv:2311.03533 [cs.CL]"},{"key":"e_1_2_1_60_1","volume-title":"CHI Conference on Human Factors in Computing Systems Extended Abstracts. 1\u20137.","author":"Vaithilingam Priyan","year":"2022","unstructured":"Priyan Vaithilingam, Tianyi Zhang, and Elena L Glassman. 2022. Expectation vs. Experience: Evaluating the Usability of Code Generation Tools Powered by Large Language Models. In CHI Conference on Human Factors in Computing Systems Extended Abstracts. 1\u20137."},{"key":"e_1_2_1_61_1","volume-title":"Sander Van Der Linden","author":"Van der Bles Anne Marthe","year":"2019","unstructured":"Anne Marthe Van der Bles, Sander Van Der Linden, Alexandra LJ Freeman, James Mitchell, Ana B Galvao, Lisa Zaval, and David J Spiegelhalter. 2019. Communicating uncertainty about facts, numbers and science. Royal Society open science 6, 5 (2019), 181870."},{"key":"e_1_2_1_62_1","doi-asserted-by":"publisher","unstructured":"Helena Vasconcelos Matthew J\u00f6rke Madeleine Grunde-McLaughlin Tobias Gerstenberg Michael Bernstein and Ranjay Krishna. 2022. Explanations Can Reduce Overreliance on AI Systems During Decision-Making. https:\/\/doi.org\/10.48550\/ARXIV.2212.06823","DOI":"10.48550\/ARXIV.2212.06823"},{"key":"e_1_2_1_63_1","volume-title":"Advances in Neural Information Processing Systems","volume":"30","author":"Vaswani Ashish","year":"2017","unstructured":"Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, \u0141ukasz Kaiser, and Illia Polosukhin. 2017. Attention is All you Need. In Advances in Neural Information Processing Systems, Vol. 30."},{"key":"e_1_2_1_64_1","doi-asserted-by":"crossref","first-page":"103456","DOI":"10.1016\/j.artint.2021.103456","article-title":"Show or suppress? Managing input uncertainty in machine learning model explanations","volume":"294","author":"Wang Danding","year":"2021","unstructured":"Danding Wang, Wencan Zhang, and Brian Y Lim. 2021. Show or suppress? Managing input uncertainty in machine learning model explanations. Artificial Intelligence 294 (2021), 103456.","journal-title":"Artificial Intelligence"},{"key":"e_1_2_1_65_1","volume-title":"Complacency and automation bias in the use of imperfect automation. Human factors 57, 5","author":"Wickens Christopher D","year":"2015","unstructured":"Christopher D Wickens, Benjamin A Clegg, Alex Z Vieane, and Angelia L Sebok. 2015. Complacency and automation bias in the use of imperfect automation. Human factors 57, 5 (2015), 728\u2013739."},{"key":"e_1_2_1_66_1","unstructured":"Qingyun Wu Gagan Bansal Jieyu Zhang Yiran Wu Shaokun Zhang Erkang Zhu Beibin Li Li Jiang Xiaoyun Zhang and Chi Wang. 2023. AutoGen: Enabling Next-Gen LLM Applications via Multi-Agent Conversation Framework. arXiv:2308.08155 [cs.AI]"},{"key":"e_1_2_1_67_1","volume-title":"Proceedings of the 2019 CHI Conference on Human Factors in Computing Systems","author":"Yin Ming","year":"2019","unstructured":"Ming Yin, Jennifer Wortman Vaughan, and Hanna M. Wallach. 2019. Understanding the Effect of Accuracy on Trust in Machine Learning Models. Proceedings of the 2019 CHI Conference on Human Factors in Computing Systems (2019)."},{"key":"e_1_2_1_68_1","volume-title":"Proceedings of the 2020 Conference on Fairness, Accountability, and Transparency","author":"Zhang Yunfeng","year":"2020","unstructured":"Yunfeng Zhang, Qingzi Vera Liao, and Rachel K. E. Bellamy. 2020. Effect of confidence and explanation on accuracy and trust calibration in AI-assisted decision making. Proceedings of the 2020 Conference on Fairness, Accountability, and Transparency (2020)."},{"key":"e_1_2_1_69_1","volume-title":"Navigating the grey area: Expressions of overconfidence and uncertainty in language models. arXiv preprint arXiv:2302.13439","author":"Zhou Kaitlyn","year":"2023","unstructured":"Kaitlyn Zhou, Dan Jurafsky, and Tatsunori Hashimoto. 2023. Navigating the grey area: Expressions of overconfidence and uncertainty in language models. arXiv preprint arXiv:2302.13439 (2023)."},{"key":"e_1_2_1_70_1","volume-title":"Proceedings of the 6th ACM SIGPLAN International Symposium on Machine Programming","author":"Ziegler Albert","year":"2022","unstructured":"Albert Ziegler, Eirini Kalliamvakou, X. Alice Li, Andrew Rice, Devon Rifkin, Shawn Simister, Ganesh Sittampalam, and Edward Aftandilian. 2022. Productivity Assessment of Neural Code Completion. In Proceedings of the 6th ACM SIGPLAN International Symposium on Machine Programming (San Diego, CA, USA) (MAPS 2022). Association for Computing Machinery, New York, NY, USA, 21\u201329. https:\/\/doi.org\/10.1145\/3520312.3534864"}],"container-title":["ACM Transactions on Computer-Human Interaction"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3702320","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3702320","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,19]],"date-time":"2025-06-19T01:18:08Z","timestamp":1750295888000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3702320"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,10,29]]},"references-count":70,"alternative-id":["10.1145\/3702320"],"URL":"https:\/\/doi.org\/10.1145\/3702320","relation":{},"ISSN":["1073-0516","1557-7325"],"issn-type":[{"value":"1073-0516","type":"print"},{"value":"1557-7325","type":"electronic"}],"subject":[],"published":{"date-parts":[[2024,10,29]]},"assertion":[{"value":"2023-12-14","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2024-07-09","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2024-10-29","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}],"article-number":"3702320"}}