{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,29]],"date-time":"2026-06-29T19:48:50Z","timestamp":1782762530335,"version":"3.54.5"},"publisher-location":"New York, NY, USA","reference-count":158,"publisher":"ACM","license":[{"start":{"date-parts":[[2026,6,25]],"date-time":"2026-06-25T00:00:00Z","timestamp":1782345600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/legalcode"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2026,6,25]]},"DOI":"10.1145\/3805689.3812402","type":"proceedings-article","created":{"date-parts":[[2026,6,29]],"date-time":"2026-06-29T17:52:08Z","timestamp":1782755528000},"page":"6438-6465","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":0,"title":["Human oversight of agentic systems in practice: Examining the oversight work, challenges, and heuristics of developers using software agents"],"prefix":"10.1145","author":[{"ORCID":"https:\/\/orcid.org\/0009-0003-7725-5888","authenticated-orcid":false,"given":"Shipi","family":"Dhanorkar","sequence":"first","affiliation":[{"name":"Microsoft, Redmond, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-7921-3820","authenticated-orcid":false,"given":"Samir","family":"Passi","sequence":"additional","affiliation":[{"name":"Microsoft, Redmond, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-3322-3548","authenticated-orcid":false,"given":"Mihaela","family":"Vorvoreanu","sequence":"additional","affiliation":[{"name":"Microsoft, Redmond, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2026,6,25]]},"reference":[{"key":"e_1_3_2_1_1_1","volume-title":"Article 14: Human Oversight | EU Artificial Intelligence Act. https:\/\/artificialintelligenceact.eu\/article\/14\/ [Online","author":"Artificial Intelligence Act E.U.","year":"2025","unstructured":"E.U. Artificial Intelligence Act. 2024. Article 14: Human Oversight | EU Artificial Intelligence Act. https:\/\/artificialintelligenceact.eu\/article\/14\/ [Online; accessed 2025-08-01]."},{"key":"e_1_3_2_1_2_1","volume-title":"Forty-second International Conference on Machine Learning. https:\/\/openreview.net\/forum?id=b0jYs6JOZu","author":"Ahmed Toufique","year":"2025","unstructured":"Toufique Ahmed, Jatin Ganhotra, Rangeet Pan, Avraham Shinnar, Saurabh Sinha, and Martin Hirzel. 2025. Otter: Generating Tests from Issues to Validate SWE Patches. In Forty-second International Conference on Machine Learning. https:\/\/openreview.net\/forum?id=b0jYs6JOZu"},{"key":"e_1_3_2_1_3_1","doi-asserted-by":"publisher","DOI":"10.1145\/3597503.3639133"},{"key":"e_1_3_2_1_4_1","volume-title":"Anthropic Economic Index report: Uneven geographic and enterprise AI adoption \u2014 anthropic.com. https:\/\/www.anthropic.com\/research\/anthropic-economic-index-september-2025-report. [Online","year":"2026","unstructured":"Anthropic. 2025. Anthropic Economic Index report: Uneven geographic and enterprise AI adoption \u2014 anthropic.com. https:\/\/www.anthropic.com\/research\/anthropic-economic-index-september-2025-report. [Online; accessed 12-01-2026]."},{"key":"e_1_3_2_1_5_1","unstructured":"Anthropic. 2025. Disrupting the first reported AI-orchestrated cyber espionage campaign. https:\/\/assets.anthropic.com\/m\/ec212e6566a0d47\/original\/Disrupting-the-first-reported-AI-orchestrated-cyber-espionage-campaign.pdf [Online; accessed 2026-01-09]."},{"key":"e_1_3_2_1_6_1","volume-title":"Magentic Marketplace: An Open-Source Environment for Studying Agentic Markets. arXiv:2510.25779 [cs.MA] https:\/\/arxiv.org\/abs\/2510.25779","author":"Bansal Gagan","year":"2025","unstructured":"Gagan Bansal, Wenyue Hua, Zezhou Huang, Adam Fourney, Amanda Swearngin, Will Epperson, Tyler Payne, Jake M. Hofman, Brendan Lucier, Chinmay Singh, Markus Mobius, Akshay Nambi, Archana Yadav, Kevin Gao, David M. Rothschild, Aleksandrs Slivkins, Daniel G. Goldstein, Hussein Mozannar, Nicole Immorlica, Maya Murad, Matthew Vogel, Subbarao Kambhampati, Eric Horvitz, and Saleema Amershi. 2025. Magentic Marketplace: An Open-Source Environment for Studying Agentic Markets. arXiv:2510.25779 [cs.MA] https:\/\/arxiv.org\/abs\/2510.25779"},{"key":"e_1_3_2_1_7_1","volume-title":"Saleema Amershi, Eric Horvitz, Adam Fourney, Hussein Mozannar, Victor Dibia, and Daniel S. Weld.","author":"Bansal Gagan","year":"2024","unstructured":"Gagan Bansal, Jennifer Wortman Vaughan, Saleema Amershi, Eric Horvitz, Adam Fourney, Hussein Mozannar, Victor Dibia, and Daniel S. Weld. 2024. Challenges in Human-Agent Communication. arXiv:2412.10380 [cs.HC] https:\/\/arxiv.org\/abs\/2412.10380"},{"key":"e_1_3_2_1_8_1","volume-title":"Executive order on the safe, secure, and trustworthy development and use of artificial intelligence. Presidential Actions","author":"Biden Joseph R","year":"2023","unstructured":"Joseph R Biden. 2023. Executive order on the safe, secure, and trustworthy development and use of artificial intelligence. Presidential Actions (2023)."},{"key":"e_1_3_2_1_9_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICSE55347.2025.00157"},{"key":"e_1_3_2_1_10_1","unstructured":"Islem Bouzenia and Michael Pradel. 2025. Understanding Software Engineering Agents: A Study of Thought-Action-Result Trajectories. arXiv:2506.18824 [cs.SE] https:\/\/arxiv.org\/abs\/2506.18824"},{"key":"e_1_3_2_1_11_1","doi-asserted-by":"publisher","DOI":"10.1145\/3708359.3712071"},{"key":"e_1_3_2_1_12_1","doi-asserted-by":"publisher","DOI":"10.1145\/3449287"},{"key":"e_1_3_2_1_13_1","doi-asserted-by":"publisher","DOI":"10.1007\/s11023-022-09591-0"},{"key":"e_1_3_2_1_14_1","doi-asserted-by":"publisher","DOI":"10.52202\/079017-3421"},{"key":"e_1_3_2_1_15_1","doi-asserted-by":"publisher","DOI":"10.1145\/3630106.3659037"},{"key":"e_1_3_2_1_16_1","doi-asserted-by":"publisher","unstructured":"Luciano Cavalcante Siebert Maria Luce Lupetti Evgeni Aizenberg Niek Beckers Arkady Zgonnikov Herman Veluwenkamp David Abbink Elisa Giaccardi Geert-Jan Houben Catholijn M. Jonker Jeroen van den Hoven Deborah Forster and Reginald L. Lagendijk. 2023. Meaningful human control: actionable properties for AI system development. AI and Ethics 3 1 (01 Feb 2023) 241\u2013255. doi:10.1007\/s43681-022-00167-3","DOI":"10.1007\/s43681-022-00167-3"},{"key":"e_1_3_2_1_17_1","unstructured":"Mert Cemri Melissa Z. Pan Shuyi Yang Lakshya A. Agrawal Bhavya Chopra Rishabh Tiwari Kurt Keutzer Aditya Parameswaran Dan Klein Kannan Ramchandran Matei Zaharia Joseph E. Gonzalez and Ion Stoica. 2025. Why Do Multi-Agent LLM Systems Fail? arXiv:2503.13657 [cs.AI]"},{"key":"e_1_3_2_1_18_1","doi-asserted-by":"publisher","DOI":"10.1145\/3630106.3658948"},{"key":"e_1_3_2_1_19_1","doi-asserted-by":"publisher","DOI":"10.1145\/3593013.3594033"},{"key":"e_1_3_2_1_20_1","doi-asserted-by":"publisher","DOI":"10.1145\/3593013.3593970"},{"key":"e_1_3_2_1_21_1","unstructured":"Mark Chen Jerry Tworek Heewoo Jun Qiming Yuan Henrique Ponde de Oliveira Pinto Jared Kaplan Harri Edwards Yuri Burda Nicholas Joseph Greg Brockman Alex Ray Raul Puri Gretchen Krueger Michael Petrov Heidy Khlaaf Girish Sastry Pamela Mishkin Brooke Chan Scott Gray Nick Ryder Mikhail Pavlov Alethea Power Lukasz Kaiser Mohammad Bavarian Clemens Winter Philippe Tillet Felipe Petroski Such Dave Cummings Matthias Plappert Fotios Chantzis Elizabeth Barnes Ariel Herbert-Voss William Hebgen Guss Alex Nichol Alex Paino Nikolas Tezak Jie Tang Igor Babuschkin Suchir Balaji Shantanu Jain William Saunders Christopher Hesse Andrew N. Carr Jan Leike Josh Achiam Vedant Misra Evan Morikawa Alec Radford Matthew Knight Miles Brundage Mira Murati Katie Mayer Peter Welinder Bob McGrew Dario Amodei Sam McCandlish Ilya Sutskever and Wojciech Zaremba. 2021. Evaluating Large Language Models Trained on Code. arXiv:2107.03374 [cs.LG] https:\/\/arxiv.org\/abs\/2107.03374"},{"key":"e_1_3_2_1_22_1","doi-asserted-by":"crossref","unstructured":"Valerie Chen Ameet Talwalkar Robert Brennan and Graham Neubig. 2025. Code with Me or for Me? How Increasing AI Automation Transforms Developer Workflows. arXiv:2507.08149 [cs.SE] https:\/\/arxiv.org\/abs\/2507.08149","DOI":"10.1145\/3772318.3790850"},{"key":"e_1_3_2_1_23_1","unstructured":"Zhi Chen Wei Ma and Lingxiao Jiang. 2025. Beyond Final Code: A Process-Oriented Error Analysis of Software Development Agents in Real-World GitHub Scenarios. arXiv:2503.12374 [cs.SE] https:\/\/arxiv.org\/abs\/2503.12374"},{"key":"e_1_3_2_1_24_1","doi-asserted-by":"publisher","DOI":"10.48550\/ARXIV.2502.01821"},{"key":"e_1_3_2_1_25_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-981-97-8440-0_75-1"},{"key":"e_1_3_2_1_26_1","doi-asserted-by":"publisher","DOI":"10.7551\/mitpress\/10113.003.0007"},{"key":"e_1_3_2_1_27_1","doi-asserted-by":"publisher","DOI":"10.1145\/3593013.3594104"},{"key":"e_1_3_2_1_28_1","volume-title":"Code Execution Through Deception: Gemini AI CLI Hijack. https:\/\/tracebit.com\/blog\/code-exec-deception-gemini-ai-cli-hijack [Online","author":"Cox Sam","year":"2026","unstructured":"Sam Cox. 2025. Code Execution Through Deception: Gemini AI CLI Hijack. https:\/\/tracebit.com\/blog\/code-exec-deception-gemini-ai-cli-hijack [Online; accessed 2026-01-09]."},{"key":"e_1_3_2_1_29_1","doi-asserted-by":"publisher","DOI":"10.5555\/1387649.1387650"},{"key":"e_1_3_2_1_30_1","volume-title":"Tian Xia, Huanzhi Mao, Nicholas Thumiger, Aditya Desai, Ion Stoica, Ana Klimovic, Graham Neubig, and Joseph E. Gonzalez.","author":"Cuadron Alejandro","year":"2025","unstructured":"Alejandro Cuadron, Dacheng Li, Wenjie Ma, Xingyao Wang, Yichuan Wang, Siyuan Zhuang, Shu Liu, Luis Gaspar Schroeder, Tian Xia, Huanzhi Mao, Nicholas Thumiger, Aditya Desai, Ion Stoica, Ana Klimovic, Graham Neubig, and Joseph E. Gonzalez. 2025. The Danger of Overthinking: Examining the Reasoning-Action Dilemma in Agentic Tasks. arXiv:2502.08235 [cs.AI] https:\/\/arxiv.org\/abs\/2502.08235"},{"key":"e_1_3_2_1_31_1","doi-asserted-by":"publisher","DOI":"10.2139\/ssrn.5277762"},{"key":"e_1_3_2_1_32_1","volume-title":"MARG: Multi-Agent Review Generation for Scientific Papers. arXiv:2401.04259 [cs.CL] https:\/\/arxiv.org\/abs\/2401.04259","author":"D'Arcy Mike","year":"2024","unstructured":"Mike D'Arcy, Tom Hope, Larry Birnbaum, and Doug Downey. 2024. MARG: Multi-Agent Review Generation for Scientific Papers. arXiv:2401.04259 [cs.CL] https:\/\/arxiv.org\/abs\/2401.04259"},{"key":"e_1_3_2_1_33_1","doi-asserted-by":"publisher","DOI":"10.1145\/3531146.3533123"},{"key":"e_1_3_2_1_34_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICSE.2012.6227187"},{"key":"e_1_3_2_1_35_1","doi-asserted-by":"publisher","DOI":"10.1145\/3637396"},{"key":"e_1_3_2_1_36_1","doi-asserted-by":"publisher","DOI":"10.1145\/3772318.3791081"},{"key":"e_1_3_2_1_37_1","volume-title":"Towards Translating Real-World Code with LLMs: A Study of Translating to Rust. CoRR","author":"Eniser Hasan Ferit","year":"2024","unstructured":"Hasan Ferit Eniser, Hanliang Zhang, Cristina David, Meng Wang, Maria Christakis, Brandon Paulsen, Joey Dodds, and Daniel Kroening. 2024. Towards Translating Real-World Code with LLMs: A Study of Translating to Rust. CoRR (2024)."},{"key":"e_1_3_2_1_38_1","doi-asserted-by":"publisher","DOI":"10.1080\/17579961.2023.2245683"},{"key":"e_1_3_2_1_39_1","doi-asserted-by":"publisher","DOI":"10.1145\/3706598.3713581"},{"key":"e_1_3_2_1_40_1","doi-asserted-by":"crossref","unstructured":"Michael Feffer Anusha Sinha Wesley H. Deng Zachary C. Lipton and Hoda Heidari. 2025. Red-Teaming for Generative AI: Silver Bullet or Security Theater? AAAI Press 421\u2013437.","DOI":"10.1609\/aies.v7i1.31647"},{"key":"e_1_3_2_1_41_1","doi-asserted-by":"publisher","DOI":"10.30574\/wjarr.2023.20.2.2194"},{"key":"e_1_3_2_1_42_1","doi-asserted-by":"publisher","DOI":"10.4324\/9780203996867"},{"key":"e_1_3_2_1_43_1","volume-title":"Byron Lee, Tiago R D Costa, Jos\u00e9 R Penad\u00e9s, Gary Peltz, Yunhan Xu, Annalisa Pawlosky, Alan Karthikesalingam, and Vivek Natarajan.","author":"Gottweis Juraj","year":"2025","unstructured":"Juraj Gottweis, Wei-Hung Weng, Alexander Daryin, Tao Tu, Anil Palepu, Petar Sirkovic, Artiom Myaskovsky, Felix Weissenberger, Keran Rong, Ryutaro Tanno, Khaled Saab, Dan Popovici, Jacob Blum, Fan Zhang, Katherine Chou, Avinatan Hassidim, Burak Gokturk, Amin Vahdat, Pushmeet Kohli, Yossi Matias, Andrew Carroll, Kavita Kulkarni, Nenad Tomasev, Yuan Guan, Vikram Dhillon, Eeshit Dhaval Vaishnav, Byron Lee, Tiago R D Costa, Jos\u00e9 R Penad\u00e9s, Gary Peltz, Yunhan Xu, Annalisa Pawlosky, Alan Karthikesalingam, and Vivek Natarajan. 2025. Towards an AI co-scientist. arXiv:2502.18864 [cs.AI] https:\/\/arxiv.org\/abs\/2502.18864"},{"key":"e_1_3_2_1_44_1","doi-asserted-by":"crossref","unstructured":"Navita Goyal Eleftheria Briakou Amanda Liu Connor Baumler Claire Bonial Jeffrey Micher Clare R. Voss Marine Carpuat and Hal Daum\u00e9 III. 2023. What Else Do I Need to Know? The Effect of Background Information on Users' Reliance on QA Systems. arXiv:2305.14331 [cs.CL] https:\/\/arxiv.org\/abs\/2305.14331","DOI":"10.18653\/v1\/2023.emnlp-main.201"},{"key":"e_1_3_2_1_45_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.clsr.2022.105681"},{"key":"e_1_3_2_1_46_1","doi-asserted-by":"publisher","DOI":"10.1145\/3359152"},{"key":"e_1_3_2_1_47_1","doi-asserted-by":"publisher","DOI":"10.52202\/079017-3369"},{"key":"e_1_3_2_1_48_1","doi-asserted-by":"publisher","DOI":"10.1109\/SaTML59370.2024.00040"},{"key":"e_1_3_2_1_49_1","unstructured":"Dan Hendrycks Steven Basart Saurav Kadavath Mantas Mazeika Akul Arora Ethan Guo Collin Burns Samir Puranik Horace He Dawn Song and Jacob Steinhardt. 2021. Measuring Coding Challenge Competence With APPS. In Thirty-fifth Conference on Neural Information Processing Systems Datasets and Benchmarks Track (Round 2). https:\/\/openreview.net\/forum?id=sD93GOzH3i5"},{"key":"e_1_3_2_1_50_1","doi-asserted-by":"publisher","DOI":"10.1145\/3593013.3594028"},{"key":"e_1_3_2_1_51_1","volume-title":"The ETTO Principle: Efficiency-Thoroughness Trade-Off: Why Things That Go Right Sometimes Go Wrong","author":"Hollnagel Erik","unstructured":"Erik Hollnagel. 2009. The ETTO Principle: Efficiency-Thoroughness Trade-Off: Why Things That Go Right Sometimes Go Wrong. CRC press."},{"key":"e_1_3_2_1_52_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.nbt.2024.12.003"},{"key":"e_1_3_2_1_53_1","volume-title":"Hassan","author":"Horikawa Kosei","year":"2025","unstructured":"Kosei Horikawa, Hao Li, Yutaro Kashiwa, Bram Adams, Hajimu Iida, and Ahmed E. Hassan. 2025. Agentic Refactoring: An Empirical Study of AI Coding Agents. arXiv:2511.04824 [[cs.SE](http:\/\/cs.se\/)] https:\/\/arxiv.org\/abs\/2511.04824"},{"key":"e_1_3_2_1_54_1","doi-asserted-by":"publisher","DOI":"10.1145\/3695988"},{"key":"e_1_3_2_1_55_1","unstructured":"Kung-Hsiang Huang Akshara Prabhakar Sidharth Dhawan Yixin Mao Huan Wang Silvio Savarese Caiming Xiong Philippe Laban and Chien-Sheng Wu. 2025. CRMArena: Understanding the Capacity of LLM Agents to Perform Professional CRM Tasks in Realistic Environments. arXiv:2411.02305 [cs.CL] https:\/\/arxiv.org\/abs\/2411.02305"},{"key":"e_1_3_2_1_56_1","doi-asserted-by":"publisher","DOI":"10.1145\/3703155"},{"key":"e_1_3_2_1_57_1","volume-title":"Professional Software Developers Don't Vibe","author":"Huang Ruanqianqian","year":"2025","unstructured":"Ruanqianqian Huang, Avery Reyna, Sorin Lerner, Haijun Xia, and Brian Hempel. 2025. Professional Software Developers Don't Vibe, They Control: AI Agent Use for Coding in 2025. arXiv:2512.14012 [cs.SE] https:\/\/arxiv.org\/abs\/2512.14012"},{"key":"e_1_3_2_1_58_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/D18-1192"},{"key":"e_1_3_2_1_59_1","volume-title":"The Twelfth International Conference on Learning Representations. https:\/\/openreview.net\/forum?id=VTF8yNQM66","author":"Jimenez Carlos E","year":"2024","unstructured":"Carlos E Jimenez, John Yang, Alexander Wettig, Shunyu Yao, Kexin Pei, Ofir Press, and Karthik R Narasimhan. 2024. SWE-bench: Can Language Models Resolve Real-world Github Issues?. In The Twelfth International Conference on Learning Representations. https:\/\/openreview.net\/forum?id=VTF8yNQM66"},{"key":"e_1_3_2_1_60_1","doi-asserted-by":"publisher","DOI":"10.1145\/3531146.3533135"},{"key":"e_1_3_2_1_61_1","doi-asserted-by":"publisher","DOI":"10.1145\/3630106.3658941"},{"key":"e_1_3_2_1_62_1","doi-asserted-by":"publisher","DOI":"10.1145\/356802.356806"},{"key":"e_1_3_2_1_63_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.chb.2024.108352"},{"key":"e_1_3_2_1_64_1","volume-title":"Human control over automation: EU policy and AI ethics. European journal of legal studies 12","author":"Riikka KOULU.","year":"2020","unstructured":"Riikka KOULU. 2020. Human control over automation: EU policy and AI ethics. European journal of legal studies 12 (2020), 9\u201346."},{"key":"e_1_3_2_1_65_1","doi-asserted-by":"publisher","DOI":"10.1177\/1023263X20978649"},{"key":"e_1_3_2_1_66_1","volume-title":"Specification gaming: the flip side of AI ingenuity. https:\/\/deepmind.google\/discover\/blog\/specification-gaming-the-flip-side-of-ai-ingenuity\/ [Online","author":"Krakovna Victoria","year":"2025","unstructured":"Victoria Krakovna, Jonathan Uesato, Vladimir Mikulik, Matthew Rahtz, Tom Everitt, Ramana Kumar, Zac Kenton, Jan Leike, and Shane Legg. 2020. Specification gaming: the flip side of AI ingenuity. https:\/\/deepmind.google\/discover\/blog\/specification-gaming-the-flip-side-of-ai-ingenuity\/ [Online; accessed 2025-09-01]."},{"key":"e_1_3_2_1_67_1","doi-asserted-by":"crossref","unstructured":"Aayush Kumar Yasharth Bajpai Sumit Gulwani Gustavo Soares and Emerson Murphy-Hill. 2025. Why AI Agents Still Need You: Findings from Developer-Agent Collaborations in the Wild. arXiv:2506.12347 [cs.SE] https:\/\/arxiv.org\/abs\/2506.12347","DOI":"10.1109\/ASE63991.2025.00043"},{"key":"e_1_3_2_1_68_1","doi-asserted-by":"publisher","unstructured":"Kyriakos Kyriakou and Jahna Otterbacher. 2023. In humans we trust. Discover Artificial Intelligence 3 1 (12 Dec 2023) 44. doi:10.1007\/s44163-023-00092-2","DOI":"10.1007\/s44163-023-00092-2"},{"key":"e_1_3_2_1_69_1","doi-asserted-by":"publisher","DOI":"10.1145\/3593013.3594087"},{"key":"e_1_3_2_1_70_1","doi-asserted-by":"publisher","DOI":"10.1007\/s11023-024-09701-0"},{"key":"e_1_3_2_1_71_1","doi-asserted-by":"publisher","DOI":"10.1145\/1134285.1134355"},{"key":"e_1_3_2_1_72_1","doi-asserted-by":"publisher","DOI":"10.1145\/3610218"},{"key":"e_1_3_2_1_73_1","doi-asserted-by":"publisher","DOI":"10.1145\/3630106.3662681"},{"key":"e_1_3_2_1_74_1","unstructured":"Shanchao Liang Spandan Garg and Roshanak Zilouchian Moghaddam. 2025. The SWE-Bench Illusion: When State-of-the-Art LLMs Remember Instead of Reason. arXiv:2506.12286 [[cs.AI](http:\/\/cs.ai\/)] https:\/\/arxiv.org\/abs\/2506.12286"},{"key":"e_1_3_2_1_75_1","doi-asserted-by":"publisher","DOI":"10.1145\/3531146.3534628"},{"key":"e_1_3_2_1_76_1","doi-asserted-by":"publisher","DOI":"10.1145\/3708519"},{"key":"e_1_3_2_1_77_1","volume-title":"Saurav Sahay, Giuseppe Raffa, and Lama Nachman.","author":"Manuvinakurike Ramesh","year":"2025","unstructured":"Ramesh Manuvinakurike, Emanuel Moss, Elizabeth Anne Watkins, Saurav Sahay, Giuseppe Raffa, and Lama Nachman. 2025. Thoughts without Thinking: Reconsidering the Explanatory Value of Chain-of-Thought Reasoning in LLMs through Agentic Pipelines. arXiv:2505.00875 [cs.AI] https:\/\/arxiv.org\/abs\/2505.00875"},{"key":"e_1_3_2_1_78_1","volume-title":"Marvin: The AI-Native Customer Feedback Repository \u2014 heymarvin.com. https:\/\/heymarvin.com\/. [Online","year":"2026","unstructured":"Marvin. [n.d.]. Marvin: The AI-Native Customer Feedback Repository \u2014 heymarvin.com. https:\/\/heymarvin.com\/. [Online; accessed 12-01-2026]."},{"key":"e_1_3_2_1_79_1","unstructured":"Tula Masterman Sandi Besen Mason Sawtell and Alex Chao. 2024. The Landscape of Emerging AI Agent Architectures for Reasoning Planning and Tool Calling: A Survey. arXiv:2404.11584 [cs.AI] https:\/\/arxiv.org\/abs\/2404.11584"},{"key":"e_1_3_2_1_80_1","unstructured":"Cecily Mauran. 2025. https:\/\/mashable.com\/article\/google-gemini-deletes-users-code"},{"key":"e_1_3_2_1_81_1","unstructured":"METR. [n.d.]. Details about METR's evaluation of OpenAI GPT-5 \/ METR's Autonomy Evaluation Resources. https:\/\/evaluations.metr.org\/gpt-5-report\/?ref=bounded-regret.ghost.io#we-do-find-evidence-of-significant-situational-awareness-though-it-is-not-robust-and-often-gets-things-wrong"},{"key":"e_1_3_2_1_82_1","unstructured":"Microsoft. 2024. Taxonomy of Failure Mode in Agentic AI Systems. Technical Report. Microsoft. https:\/\/cdn-dynmedia-1.microsoft.com\/is\/content\/microsoftcorp\/microsoft\/final\/en-us\/microsoft-brand\/documents\/Taxonomy-of-Failure-Mode-in-Agentic-AI-Systems-Whitepaper.pdf"},{"key":"e_1_3_2_1_83_1","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2406.12952"},{"key":"e_1_3_2_1_84_1","doi-asserted-by":"publisher","DOI":"10.1145\/3706599.3706693"},{"key":"e_1_3_2_1_85_1","volume-title":"Webgpt: Browser-assisted question-answering with human feedback. arXiv preprint arXiv:2112.09332","author":"Nakano Reiichiro","year":"2021","unstructured":"Reiichiro Nakano, Jacob Hilton, Suchir Balaji, Jeff Wu, Long Ouyang, Christina Kim, Christopher Hesse, Shantanu Jain, Vineet Kosaraju, William Saunders, et al. 2021. Webgpt: Browser-assisted question-answering with human feedback. arXiv preprint arXiv:2112.09332 (2021)."},{"key":"e_1_3_2_1_86_1","doi-asserted-by":"publisher","DOI":"10.1145\/3715275.3732089"},{"key":"e_1_3_2_1_87_1","unstructured":"Joe Needham Giles Edkins Govind Pimpale Henning Bartsch and Marius Hobbhahn. 2025. Large Language Models Often Know When They Are Being Evaluated. arXiv:2505.23836 [cs.CL] https:\/\/arxiv.org\/abs\/2505.23836"},{"key":"e_1_3_2_1_88_1","volume-title":"Faultline: Automated proof-of-vulnerability generation using llm agents. arXiv preprint arXiv:2507.15241","author":"Nitin Vikram","year":"2025","unstructured":"Vikram Nitin, Baishakhi Ray, and Roshanak Zilouchian Moghaddam. 2025. Faultline: Automated proof-of-vulnerability generation using llm agents. arXiv preprint arXiv:2507.15241 (2025)."},{"key":"e_1_3_2_1_89_1","volume-title":"AI-powered coding tool wiped out a software company's database in \u2018catastrophic failure","author":"Nolan Beatrice","year":"2025","unstructured":"Beatrice Nolan. 2025. AI-powered coding tool wiped out a software company's database in \u2018catastrophic failure\u2019 | Fortune. https:\/\/fortune.com\/2025\/07\/23\/ai-coding-tool-replit-wiped-database-called-it-a-catastrophic-failure\/ [Online; accessed 2026-01-09]."},{"key":"e_1_3_2_1_90_1","unstructured":"Alexander Novikov Ng\u00e2n V\u0169 Marvin Eisenberger Emilien Dupont Po-Sen Huang Adam Zsolt Wagner Sergey Shirobokov Borislav Kozlovskii Francisco J. R. Ruiz Abbas Mehrabian M. Pawan Kumar Abigail See Swarat Chaudhuri George Holland Alex Davies Sebastian Nowozin Pushmeet Kohli and Matej Balog. 2025. AlphaEvolve: A coding agent for scientific and algorithmic discovery. arXiv:2506.13131 [cs.AI] https:\/\/arxiv.org\/abs\/2506.13131"},{"key":"e_1_3_2_1_91_1","doi-asserted-by":"publisher","DOI":"10.1056\/AIp2400979"},{"key":"e_1_3_2_1_92_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICSE-SEIP66354.2025.00060"},{"key":"e_1_3_2_1_93_1","doi-asserted-by":"publisher","DOI":"10.1145\/3586183.3606763"},{"key":"e_1_3_2_1_94_1","doi-asserted-by":"publisher","DOI":"10.2139\/ssrn.5529058"},{"key":"e_1_3_2_1_95_1","unstructured":"Samir Passi Shipi Dhanorkar and Mihaela Vorvoreanu. 2024. Appropriate reliance on Generative AI: Research synthesis. Technical Report MSR-TR-2024-7. Microsoft. https:\/\/www.microsoft.com\/en-us\/research\/publication\/appropriate-reliance-on-generative-ai-research-synthesis\/"},{"key":"e_1_3_2_1_96_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-981-97-8440-0_98-1"},{"key":"e_1_3_2_1_97_1","unstructured":"Samir Passi and Mihaela Vorvoreanu. 2022. Overreliance on AI: Literature Review. Technical Report MSR-TR-2022-12. Microsoft. https:\/\/www.microsoft.com\/en-us\/research\/publication\/overreliance-on-ai-literature-review\/"},{"key":"e_1_3_2_1_98_1","volume-title":"Qualitative Research & Evaluation Methods: Integrating Theory and Practice","author":"Patton M.Q.","unstructured":"M.Q. Patton. 2014. Qualitative Research & Evaluation Methods: Integrating Theory and Practice. SAGE Publications. https:\/\/books.google.com\/books?id=ovAkBQAAQBAJ"},{"key":"e_1_3_2_1_99_1","doi-asserted-by":"publisher","DOI":"10.1145\/3610721"},{"key":"e_1_3_2_1_100_1","unstructured":"Sida Peng Eirini Kalliamvakou Peter Cihon and Mert Demirer. 2023. The Impact of AI on Developer Productivity: Evidence from GitHub Copilot. arXiv:2302.06590 [[cs.SE](http:\/\/cs.se\/)] https:\/\/arxiv.org\/abs\/2302.06590"},{"key":"e_1_3_2_1_101_1","doi-asserted-by":"publisher","DOI":"10.1037\/0033-295X.106.4.643"},{"key":"e_1_3_2_1_102_1","unstructured":"Inioluwa Deborah Raji Emily Denton Emily M. Bender Alex Hanna and Amandalynne Paullada. 2021. AI and the Everything in the Whole Wide World Benchmark. In Thirty-fifth Conference on Neural Information Processing Systems Datasets and Benchmarks Track (Round 2). https:\/\/openreview.net\/forum?id=j6NxpQbREA1"},{"key":"e_1_3_2_1_103_1","doi-asserted-by":"publisher","DOI":"10.1145\/3687299"},{"key":"e_1_3_2_1_104_1","doi-asserted-by":"publisher","DOI":"10.1145\/3630106.3659032"},{"key":"e_1_3_2_1_105_1","doi-asserted-by":"publisher","DOI":"10.3389\/frobt.2018.00015"},{"key":"e_1_3_2_1_106_1","doi-asserted-by":"publisher","DOI":"10.2139\/ssrn.5713646"},{"key":"e_1_3_2_1_107_1","doi-asserted-by":"publisher","DOI":"10.1145\/3581641.3584066"},{"key":"e_1_3_2_1_108_1","doi-asserted-by":"publisher","DOI":"10.52202\/075280-2997"},{"key":"e_1_3_2_1_109_1","doi-asserted-by":"publisher","DOI":"10.1145\/3630106.3659031"},{"key":"e_1_3_2_1_110_1","doi-asserted-by":"publisher","DOI":"10.1145\/3531146.3533127"},{"key":"e_1_3_2_1_111_1","volume-title":"The phenomenology of the social world","author":"Schutz Alfred","unstructured":"Alfred Schutz. 1967. The phenomenology of the social world. Northwestern university press."},{"key":"e_1_3_2_1_112_1","doi-asserted-by":"publisher","DOI":"10.1145\/3637865"},{"key":"e_1_3_2_1_113_1","doi-asserted-by":"publisher","DOI":"10.1145\/3531146.3533189"},{"key":"e_1_3_2_1_114_1","volume-title":"Robinson","author":"Shavit Yonadav","year":"2023","unstructured":"Yonadav Shavit, Sandhini Agarwal, Miles Brundage, Steven Adler, Cullen O'Keefe, Rosie Campbell, Teddy Lee, Pamela Mishkin, Tyna Eloundou, Alan Hickey, Katarina Slama, Lama Ahmad, Paul McMillan, Alex Beutel, Alexandre Passos, and David G. Robinson. 2023. Practices for Governing Agentic AI Systems. Technical Report. OpenAI. https:\/\/cdn.openai.com\/papers\/practices-for-governing-agentic-ai-systems.pdf"},{"key":"e_1_3_2_1_115_1","doi-asserted-by":"publisher","DOI":"10.52202\/075280-1657"},{"key":"e_1_3_2_1_116_1","doi-asserted-by":"publisher","DOI":"10.52202\/075280-0377"},{"key":"e_1_3_2_1_117_1","doi-asserted-by":"publisher","DOI":"10.1080\/10447318.2020.1741118"},{"key":"e_1_3_2_1_118_1","doi-asserted-by":"crossref","unstructured":"Ben Shneiderman. 2022. Human-centered AI. Oxford University Press.","DOI":"10.1093\/oso\/9780192845290.001.0001"},{"key":"e_1_3_2_1_119_1","doi-asserted-by":"crossref","unstructured":"Parshin Shojaee*\u2020 Iman Mirzadeh* Keivan Alizadeh Maxwell Horton Samy Bengio and Mehrdad Farajtabar. 2025. The Illusion of Thinking: Understanding the Strengths and Limitations of Reasoning Models via the Lens of Problem Complexity. https:\/\/ml-site.cdn-apple.com\/papers\/the-illusion-of-thinking.pdf","DOI":"10.70777\/si.v2i6.15919"},{"key":"e_1_3_2_1_120_1","doi-asserted-by":"crossref","unstructured":"Pradyumna Shome Sashreek Krishnan and Sauvik Das. 2025. Why Johnny Can't Use Agents: Industry Aspirations vs. User Realities with AI Agent Software. arXiv:2509.14528 [cs.HC] https:\/\/arxiv.org\/abs\/2509.14528","DOI":"10.1145\/3786335.3813140"},{"key":"e_1_3_2_1_121_1","first-page":"161","article-title":"Theories of Bounded Rationality","volume":"22","author":"HA","year":"1972","unstructured":"HA SIMON. 1972. Theories of Bounded Rationality. Decision and Organization 22 (1972), 161\u2013176.","journal-title":"Decision and Organization"},{"key":"e_1_3_2_1_122_1","doi-asserted-by":"publisher","DOI":"10.1037\/h0042769"},{"key":"e_1_3_2_1_123_1","doi-asserted-by":"crossref","unstructured":"Hao Song Yiming Shen Wenxuan Luo Leixin Guo Ting Chen Jiashui Wang Beibei Li Xiaosong Zhang and Jiachi Chen. 2025. Beyond the Protocol: Unveiling Attack Vectors in the Model Context Protocol (MCP) Ecosystem. arXiv:2506.02040 [cs.CR] https:\/\/arxiv.org\/abs\/2506.02040","DOI":"10.1109\/TSE.2026.3694876"},{"key":"e_1_3_2_1_124_1","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2406.10279"},{"key":"e_1_3_2_1_125_1","volume-title":"The Thirteenth International Conference on Learning Representations, ICLR 2025","author":"Stechly Kaya","year":"2025","unstructured":"Kaya Stechly, Karthik Valmeekam, and Subbarao Kambhampati. 2025. On the self-verification limitations of large language models on reasoning and planning tasks. In The Thirteenth International Conference on Learning Representations, ICLR 2025, Singapore, April 24-28, 2025. OpenReview.net. https:\/\/openreview.net\/forum?id=4O0v4s3IzY"},{"key":"e_1_3_2_1_126_1","doi-asserted-by":"publisher","DOI":"10.1145\/2675133.2675298"},{"key":"e_1_3_2_1_127_1","doi-asserted-by":"publisher","DOI":"10.1145\/3630106.3659051"},{"key":"e_1_3_2_1_128_1","volume-title":"How People Use AI Agents \u2014 perplexity.ai. https:\/\/www.perplexity.ai\/hub\/blog\/how-people-use-ai-agents. [Online","author":"Team Perplexity","year":"2026","unstructured":"Perplexity Team. 2025. How People Use AI Agents \u2014 perplexity.ai. https:\/\/www.perplexity.ai\/hub\/blog\/how-people-use-ai-agents. [Online; accessed 12-01-2026]."},{"key":"e_1_3_2_1_129_1","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v39i24.34717"},{"key":"e_1_3_2_1_130_1","doi-asserted-by":"publisher","DOI":"10.1145\/3696630.3730565"},{"key":"e_1_3_2_1_131_1","doi-asserted-by":"publisher","DOI":"10.1515\/9783110718508-006"},{"key":"e_1_3_2_1_132_1","doi-asserted-by":"publisher","DOI":"10.1007\/s43681-024-00489-4"},{"key":"e_1_3_2_1_133_1","doi-asserted-by":"publisher","DOI":"10.1145\/3630106.3658985"},{"key":"e_1_3_2_1_134_1","doi-asserted-by":"publisher","DOI":"10.1126\/science.185.4157.1124"},{"key":"e_1_3_2_1_135_1","doi-asserted-by":"publisher","DOI":"10.1145\/3715275.3732157"},{"key":"e_1_3_2_1_136_1","doi-asserted-by":"publisher","DOI":"10.1145\/3630106.3658984"},{"key":"e_1_3_2_1_137_1","doi-asserted-by":"publisher","DOI":"10.5555\/3692070.3694124"},{"key":"e_1_3_2_1_138_1","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2407.16741"},{"key":"e_1_3_2_1_139_1","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2503.15223"},{"key":"e_1_3_2_1_140_1","doi-asserted-by":"publisher","DOI":"10.1007\/s10515-025-00544-2"},{"key":"e_1_3_2_1_141_1","volume-title":"Hassan","author":"Watanabe Miku","year":"2025","unstructured":"Miku Watanabe, Hao Li, Yutaro Kashiwa, Brittany Reid, Hajimu Iida, and Ahmed E. Hassan. 2025. On the Use of Agentic Coding: An Empirical Study of Pull Requests on GitHub. arXiv:2509.14745 [[cs.SE](http:\/\/cs.se\/)] https:\/\/arxiv.org\/abs\/2509.14745"},{"key":"e_1_3_2_1_142_1","doi-asserted-by":"publisher","DOI":"10.1145\/3531146.3533088"},{"key":"e_1_3_2_1_143_1","doi-asserted-by":"publisher","DOI":"10.3386\/w33662"},{"key":"e_1_3_2_1_144_1","doi-asserted-by":"publisher","DOI":"10.1162\/TACL.a.25"},{"key":"e_1_3_2_1_145_1","doi-asserted-by":"publisher","DOI":"10.1145\/3180155.3182538"},{"key":"e_1_3_2_1_146_1","doi-asserted-by":"publisher","DOI":"10.1145\/3654777.3676374"},{"key":"e_1_3_2_1_147_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICSE-Companion66252.2025.00071"},{"key":"e_1_3_2_1_148_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2025.acl-industry.89"},{"key":"e_1_3_2_1_149_1","doi-asserted-by":"publisher","DOI":"10.52202\/079017-1601"},{"key":"e_1_3_2_1_150_1","doi-asserted-by":"publisher","DOI":"10.1145\/3597503.3639074"},{"key":"e_1_3_2_1_151_1","volume-title":"ReAct: Synergizing Reasoning and Acting in Language Models. In The Eleventh International Conference on Learning Representations. https:\/\/openreview.net\/forum?id=WE_vluYUL-X","author":"Yao Shunyu","year":"2023","unstructured":"Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik R Narasimhan, and Yuan Cao. 2023. ReAct: Synergizing Reasoning and Acting in Language Models. In The Eleventh International Conference on Learning Representations. https:\/\/openreview.net\/forum?id=WE_vluYUL-X"},{"key":"e_1_3_2_1_152_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2024.acl-long.737"},{"key":"e_1_3_2_1_153_1","doi-asserted-by":"publisher","DOI":"10.1145\/3650212.3680384"},{"key":"e_1_3_2_1_154_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2024.findings-acl.762"},{"key":"e_1_3_2_1_155_1","unstructured":"Yuxuan Zhu Tengjun Jin Yada Pruksachatkun Andy Zhang Shu Liu Sasha Cui Sayash Kapoor Shayne Longpre Kevin Meng Rebecca Weiss Fazl Barez Rahul Gupta Jwala Dhamala Jacob Merizian Mario Giulianelli Harry Coppock Cozmin Ududec Jasjeet Sekhon Jacob Steinhardt Antony Kellermann Sarah Schwettmann Matei Zaharia Ion Stoica Percy Liang and Daniel Kang. 2025. Establishing Best Practices for Building Rigorous Agentic Benchmarks. arXiv:2507.02825 [cs.AI] https:\/\/arxiv.org\/abs\/2507.02825"},{"key":"e_1_3_2_1_156_1","volume-title":"Agent-as-a-Judge: Evaluate Agents with Agents. In Forty-second International Conference on Machine Learning.","author":"Zhuge Mingchen","year":"2025","unstructured":"Mingchen Zhuge, Changsheng Zhao, Dylan R Ashley, Wenyi Wang, Dmitrii Khizbullin, Yunyang Xiong, Zechun Liu, Ernie Chang, Raghuraman Krishnamoorthi, Yuandong Tian, et al. 2025. Agent-as-a-Judge: Evaluate Agents with Agents. In Forty-second International Conference on Machine Learning."},{"key":"e_1_3_2_1_157_1","unstructured":"Terry Yue Zhuo Vu Minh Chien Jenny Chim Han Hu Wenhao Yu Ratnadira Widyasari Imam Nur Bani Yusuf Haolan Zhan Junda He Indraneil Paul Simon Brunner Chen GONG James Hoang Armel Randy Zebaze Xiaoheng Hong Wen-Ding Li Jean Kaddour Ming Xu Zhihan Zhang Prateek Yadav Naman Jain Alex Gu Zhoujun Cheng Jiawei Liu Qian Liu Zijian Wang David Lo Binyuan Hui Niklas Muennighoff Daniel Fried Xiaoning Du Harm de Vries and Leandro Von Werra. 2025. BigCodeBench: Benchmarking Code Generation with Diverse Function Calls and Complex Instructions. In The Thirteenth International Conference on Learning Representations. https:\/\/openreview.net\/forum?id=YrycTjllL0"},{"key":"e_1_3_2_1_158_1","doi-asserted-by":"publisher","DOI":"10.1145\/3715275.3732086"}],"event":{"name":"FAccT '26: The 2026 ACM Conference on Fairness, Accountability, and Transparency","location":"Montreal QC Canada","acronym":"FAccT '26","sponsor":["ACM\/SIG"]},"container-title":["Proceedings of the 2026 ACM Conference on Fairness, Accountability, and Transparency"],"original-title":[],"link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3805689.3812402","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,6,29]],"date-time":"2026-06-29T19:20:34Z","timestamp":1782760834000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3805689.3812402"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,6,25]]},"references-count":158,"alternative-id":["10.1145\/3805689.3812402","10.1145\/3805689"],"URL":"https:\/\/doi.org\/10.1145\/3805689.3812402","relation":{},"subject":[],"published":{"date-parts":[[2026,6,25]]},"assertion":[{"value":"2026-06-25","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}