{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,13]],"date-time":"2026-07-13T15:06:26Z","timestamp":1783955186080,"version":"3.55.0"},"reference-count":26,"publisher":"Association for Computing Machinery (ACM)","issue":"1","content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["SIGOPS Oper. Syst. Rev."],"published-print":{"date-parts":[[2026,7,13]]},"abstract":"<jats:p>Large language models (LLMs) have shown exceptional capabilities in code understanding and generation. However, they still face significant challenges in analyzing and debugging code. Most existing works rely on a single model, which often struggles to detect and fix bugs in complex program structures and semantic logic. This paper presents DoTA, a novel debugging framework that enhances LLMs' debugging capabilities through multiagent collaboration. Our key innovation is two-fold. First, we enhance code understanding through automated hierarchical documentation analysis, enabling more effective bug detection and localization based on comprehensive program context. Second, we leverage a delta-of-thoughts process where multi LLM agents analyze different aspects of program correctness and iteratively contribute complementary insights to identify bugs. Our experiments show that combining different LLM based agents with enriched documentation context significantly improves LLMs' debugging capabilities. DoTA has been extensively evaluated on the Debug- Bench dataset of 4,253 debugging instances. It achieves an average improvement of 15% in bug detection accuracy across languages compared to GPT-3.5 with task background prompting. The framework shows particular strength in handling complex logical errors (+18.3%) and multiple bugs (+18.9%). On open-source models, DoTA enhances bug detection capabilities by 10.4\u201313.5%.<\/jats:p>","DOI":"10.1145\/3830422.3830424","type":"journal-article","created":{"date-parts":[[2026,7,13]],"date-time":"2026-07-13T14:20:01Z","timestamp":1783952401000},"page":"9-20","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":0,"title":["DOTA: Intelligent Debugging with\n                    <u>D<\/u>\n                    elta\n                    <u>o<\/u>\n                    f\n                    <u>T<\/u>\n                    houghts\n                    <u>A<\/u>\n                    gents"],"prefix":"10.1145","volume":"60","author":[{"given":"Rui","family":"Yang","sequence":"first","affiliation":[{"name":"University of California Riverside, Riverside, CA, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Rajiv","family":"Gupta","sequence":"additional","affiliation":[{"name":"University of California Riverside, Riverside, CA, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Qian","family":"Zhang","sequence":"additional","affiliation":[{"name":"University of California Riverside, Riverside, CA, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Ashish","family":"Kundu","sequence":"additional","affiliation":[{"name":"Cisco, San Jose, CA, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2026,7,13]]},"reference":[{"key":"e_1_2_1_1_1","volume-title":"GPT-4 technical report. arXiv preprint arXiv:2303.08774","author":"AI.","year":"2023","unstructured":"OpenAI. GPT-4 technical report. arXiv preprint arXiv:2303.08774, 2023."},{"key":"e_1_2_1_2_1","volume-title":"Evaluating large language models trained on code. In arXiv preprint arXiv:2107.03374","author":"Chen Mark","year":"2021","unstructured":"Mark Chen, Jerry Tworek, Heewoo Jun, et al. Evaluating large language models trained on code. In arXiv preprint arXiv:2107.03374, 2021."},{"key":"e_1_2_1_3_1","volume-title":"et al. Code Llama: Open foundation models for code. arXiv preprint arXiv:2308.12950","author":"Rozi'ere Baptiste","year":"2023","unstructured":"Baptiste Rozi'ere, Jonas Gehring, Fabian Gloeckle, et al. Code Llama: Open foundation models for code. arXiv preprint arXiv:2308.12950, 2023."},{"key":"e_1_2_1_4_1","volume-title":"International Conference on Learning Representations (ICLR)","author":"Jimenez Carlos E.","year":"2024","unstructured":"Carlos E. Jimenez, John Yang, Alexander Wettig, et al. SWE-bench: Can language models resolve real-world GitHub issues? In International Conference on Learning Representations (ICLR), 2024."},{"key":"e_1_2_1_5_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2021.emnlp-main.685"},{"key":"e_1_2_1_6_1","volume-title":"Advances in NeurIPS Datasets and Benchmarks Track","author":"Lu Shuai","year":"2021","unstructured":"Shuai Lu, Daya Guo, Shuo Ren, et al. CodeXGLUE: A machine learning benchmark dataset for code understanding and generation. In Advances in NeurIPS Datasets and Benchmarks Track, 2021."},{"key":"e_1_2_1_7_1","doi-asserted-by":"publisher","DOI":"10.1109\/TSE.2016.2521368"},{"key":"e_1_2_1_8_1","volume-title":"The Thirteenth International Conference on Learning Representations (ICLR)","author":"McAleese Nat","year":"2025","unstructured":"Nat McAleese, Rai Rai, et al. Llm critics help catch llm bugs. In The Thirteenth International Conference on Learning Representations (ICLR), 2025."},{"key":"e_1_2_1_9_1","doi-asserted-by":"publisher","DOI":"10.52202\/075280-2023"},{"key":"e_1_2_1_10_1","doi-asserted-by":"publisher","DOI":"10.1145\/3597503.3639174"},{"key":"e_1_2_1_11_1","volume-title":"Agentcoder: Multi-agentbased code generation with iterative testing and optimisation. arXiv preprint arXiv:2312.13010","author":"Huang Dong","year":"2023","unstructured":"Dong Huang, Qingwen Bu, Jie M. Zhang, Michael Luck, and Heming Cui. Agentcoder: Multi-agentbased code generation with iterative testing and optimisation. arXiv preprint arXiv:2312.13010, 2023."},{"key":"e_1_2_1_12_1","doi-asserted-by":"publisher","DOI":"10.1007\/s10664-024-10594-x"},{"key":"e_1_2_1_13_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2024.findings-acl.247"},{"key":"e_1_2_1_14_1","doi-asserted-by":"publisher","DOI":"10.1126\/science.abq1158"},{"key":"e_1_2_1_15_1","volume-title":"Yang Zi, et al. StarCoder: May the source be with you! Trans. on Machine Learning Research (TMLR)","author":"Li Raymond","year":"2023","unstructured":"Raymond Li, Loubna Ben Allal, Yang Zi, et al. StarCoder: May the source be with you! Trans. on Machine Learning Research (TMLR), 2023."},{"key":"e_1_2_1_16_1","first-page":"5521","volume-title":"Proc. of the ACM SIGKDD Conference on Knowledge Discovery and Data Mining","author":"Zheng Qinkai","year":"2023","unstructured":"Qinkai Zheng, Xiao Xia, et al. CodeGeeX: A pretrained model for code generation with multilingual evaluations on HumanEval-X. In Proc. of the ACM SIGKDD Conference on Knowledge Discovery and Data Mining, pages 5521-5531, 2023."},{"key":"e_1_2_1_17_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2025.naacl-long.570"},{"key":"e_1_2_1_18_1","volume-title":"Debugging with GDB: The GNU Source-Level Debugger","author":"Stallman Richard","year":"2000","unstructured":"Richard Stallman, Roland Pesch, and Stan Shebs. Debugging with GDB: The GNU Source-Level Debugger. Free Software Foundation, 2000."},{"key":"e_1_2_1_19_1","first-page":"89","volume-title":"Proc. of the ACM SIGPLAN Conference on Programming Language Design and Implementation (PLDI)","author":"Nethercote Nicholas","year":"2007","unstructured":"Nicholas Nethercote and Julian Seward. Valgrind: A framework for heavyweight dynamic binary instrumentation. In Proc. of the ACM SIGPLAN Conference on Programming Language Design and Implementation (PLDI), pages 89-100, 2007."},{"key":"e_1_2_1_20_1","doi-asserted-by":"publisher","DOI":"10.1145\/3182657"},{"key":"e_1_2_1_21_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICPC.2009.5090029"},{"key":"e_1_2_1_22_1","doi-asserted-by":"publisher","DOI":"10.1145\/3395363.3397369"},{"key":"e_1_2_1_23_1","first-page":"735","volume-title":"Proc. of the IEEE\/ACM International Conference on Software Engineering (ICSE)","author":"Drain Dawn","year":"2021","unstructured":"Dawn Drain, Colin B Clement, Guillermo Serrato, and Neel Sundaresan. Deepdebug: fixing Python bugs using stack traces, backtranslation, and code skeletons. In Proc. of the IEEE\/ACM International Conference on Software Engineering (ICSE), pages 735-747, 2021."},{"key":"e_1_2_1_24_1","doi-asserted-by":"publisher","DOI":"10.1093\/comjnl\/27.2.97"},{"key":"e_1_2_1_25_1","first-page":"15174","volume-title":"Proc. of the Annual Meeting of the Association for Computational Linguistics","author":"Qian Chen","year":"2024","unstructured":"Chen Qian, Wei Liu, Hongzhang Liu, Nuo Chen, Yufan Dang, Jiahao Li, Cheng Yang, Weize Chen, Yusheng Su, Xin Cong, Juyuan Xu, Dahai Li, Zhiyuan Liu, and Maosong Sun. ChatDev: Communicative agents for software development. In Proc. of the Annual Meeting of the Association for Computational Linguistics, pages 15174-15186, August 2024."},{"key":"e_1_2_1_26_1","volume-title":"The Twelfth International Conference on Learning Representations (ICLR)","author":"Hong Sirui","year":"2024","unstructured":"Sirui Hong, Xiawu Zheng, et al. MetaGPT: Meta programming for multi-agent collaborative framework. In The Twelfth International Conference on Learning Representations (ICLR), 2024."}],"container-title":["ACM SIGOPS Operating Systems Review"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3830422.3830424","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,7,13]],"date-time":"2026-07-13T14:20:47Z","timestamp":1783952447000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3830422.3830424"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,7,13]]},"references-count":26,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2026,7,13]]}},"alternative-id":["10.1145\/3830422.3830424"],"URL":"https:\/\/doi.org\/10.1145\/3830422.3830424","relation":{},"ISSN":["0163-5980"],"issn-type":[{"value":"0163-5980","type":"print"}],"subject":[],"published":{"date-parts":[[2026,7,13]]},"assertion":[{"value":"2026-07-13","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}