{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,30]],"date-time":"2026-06-30T19:07:00Z","timestamp":1782846420090,"version":"3.54.5"},"reference-count":51,"publisher":"Association for Computing Machinery (ACM)","issue":"FSE","license":[{"start":{"date-parts":[[2026,6,30]],"date-time":"2026-06-30T00:00:00Z","timestamp":1782777600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/legalcode"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Proc. ACM Softw. Eng."],"published-print":{"date-parts":[[2026,6,30]]},"abstract":"<jats:p>Automated Code Review (ACR) is crucial for software quality, yet existing benchmarks often fail to reflect real-world complexities, hindering the evaluation of modern Large Language Models (LLMs). Current benchmarks frequently focus on fine-grained code units, lack complete project context, and use inadequate evaluation metrics. To address these limitations, we introduce SWR-Bench, a new benchmark comprising 1000 manually verified Pull Requests (PRs) from GitHub, offering PR-centric review with full project context. SWR-Bench employs an objective LLM-based evaluation method that aligns strongly with human judgment (\u223c90% agreement) by verifying if issues from a structured ground truth are covered in generated reviews. Our systematic evaluation of mainstream ACR tools and LLMs on SWR-Bench reveals that current systems underperform, and ACR tools are more adept at detecting functional errors. Subsequently, we propose and validate a simple multi-review aggregation strategy that significantly boosts ACR performance, increasing F1 scores by up to 43.67%. Our contributions include the SWR-Bench benchmark, its objective evaluation method, a comprehensive study of current ACR capabilities, and an effective enhancement approach, offering valuable insights for advancing ACR research.<\/jats:p>","DOI":"10.1145\/3808144","type":"journal-article","created":{"date-parts":[[2026,6,30]],"date-time":"2026-06-30T17:06:14Z","timestamp":1782839174000},"page":"3093-3115","source":"Crossref","is-referenced-by-count":0,"title":["SWR-Bench: Assessing LLM Performance in Real-World Code Review Comment Generation"],"prefix":"10.1145","volume":"3","author":[{"ORCID":"https:\/\/orcid.org\/0009-0009-8422-4522","authenticated-orcid":false,"given":"Zhengran","family":"Zeng","sequence":"first","affiliation":[{"name":"Peking University, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0001-2028-670X","authenticated-orcid":false,"given":"Ruikai","family":"Shi","sequence":"additional","affiliation":[{"name":"Peking University, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0003-1843-8545","authenticated-orcid":false,"given":"Keke","family":"Han","sequence":"additional","affiliation":[{"name":"Peking University, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0005-0492-2234","authenticated-orcid":false,"given":"Yixin","family":"Li","sequence":"additional","affiliation":[{"name":"Peking University, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0007-0856-9049","authenticated-orcid":false,"given":"Kaicheng","family":"Sun","sequence":"additional","affiliation":[{"name":"Northwestern Polytechnical University, Xian, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0007-9969-8259","authenticated-orcid":false,"given":"Yidong","family":"Wang","sequence":"additional","affiliation":[{"name":"Peking University, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0000-8256-8588","authenticated-orcid":false,"given":"Zhuohao","family":"Yu","sequence":"additional","affiliation":[{"name":"Peking University, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-1756-7746","authenticated-orcid":false,"given":"Rui","family":"Xie","sequence":"additional","affiliation":[{"name":"Peking University, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-9331-4716","authenticated-orcid":false,"given":"Wei","family":"Ye","sequence":"additional","affiliation":[{"name":"Peking University, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-8576-2674","authenticated-orcid":false,"given":"Shikun","family":"Zhang","sequence":"additional","affiliation":[{"name":"Peking University, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2026,6,30]]},"reference":[{"key":"e_1_2_1_1_1","unstructured":"SWR-Bench: A Benchmarking Suite for Serverless Cold Start Time Reduction. https:\/\/github.com\/ZZR0\/ SWRench. Accessed: 2025-05-23.."},{"key":"e_1_2_1_2_1","doi-asserted-by":"publisher","DOI":"10.1145\/3585004"},{"key":"e_1_2_1_3_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICSM.2015.7332454"},{"key":"e_1_2_1_4_1","doi-asserted-by":"publisher","DOI":"10.1145\/2597073.2597082"},{"key":"e_1_2_1_5_1","doi-asserted-by":"publisher","DOI":"10.48550\/ARXIV.2503.13657"},{"key":"e_1_2_1_6_1","doi-asserted-by":"publisher","DOI":"10.1016\/J.SCICO.2021.102652"},{"key":"e_1_2_1_7_1","doi-asserted-by":"publisher","DOI":"10.1016\/J.SCICO.2021.102652"},{"key":"e_1_2_1_8_1","doi-asserted-by":"publisher","DOI":"10.48550\/ARXIV.2503.16858"},{"key":"e_1_2_1_9_1","doi-asserted-by":"publisher","DOI":"10.48550\/ARXIV.2412.18531"},{"key":"e_1_2_1_10_1","doi-asserted-by":"publisher","DOI":"10.48550\/ARXIV.2409.15152"},{"key":"e_1_2_1_11_1","doi-asserted-by":"publisher","DOI":"10.1007\/S10664-022-10205-7"},{"key":"e_1_2_1_12_1","unstructured":"Google. 2025. Gemini 2.5: Our most intelligent ai model. https:\/\/blog.google\/technology\/google-deepmind\/geminimodel-thinking-updates-march-2025\/#gemini-2-5-thinking."},{"key":"e_1_2_1_13_1","doi-asserted-by":"publisher","DOI":"10.48550\/ARXIV.2411.15594"},{"key":"e_1_2_1_14_1","volume-title":"LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code. In The Thirteenth International Conference on Learning Representations, ICLR 2025","author":"Jain Naman","year":"2025","unstructured":"Naman Jain, King Han, Alex Gu, Wen-Ding Li, Fanjia Yan, Tianjun Zhang, Sida Wang, Armando Solar-Lezama, Koushik Sen, and Ion Stoica. 2025. LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code. In The Thirteenth International Conference on Learning Representations, ICLR 2025, Singapore, April 24-28, 2025. OpenReview.net. https:\/\/openreview.net\/forum?id=chfJJYC3iL"},{"key":"e_1_2_1_15_1","doi-asserted-by":"publisher","DOI":"10.48550\/ARXIV.2502.06633"},{"key":"e_1_2_1_16_1","doi-asserted-by":"publisher","DOI":"10.48550\/ARXIV.2502.06633"},{"key":"e_1_2_1_17_1","doi-asserted-by":"publisher","DOI":"10.48550\/ARXIV.2501.05176"},{"key":"e_1_2_1_18_1","volume-title":"The Twelfth International Conference on Learning Representations, ICLR 2024","author":"Jimenez Carlos E.","year":"2024","unstructured":"Carlos E. Jimenez, John Yang, Alexander Wettig, Shunyu Yao, Kexin Pei, Ofir Press, and Karthik R. Narasimhan. 2024. SWE-bench: Can Language Models Resolve Real-world Github Issues?. In The Twelfth International Conference on Learning Representations, ICLR 2024, Vienna, Austria, May 7-11, 2024. OpenReview.net. https:\/\/openreview.net\/forum? id=VTF8yNQM66"},{"key":"e_1_2_1_19_1","doi-asserted-by":"publisher","DOI":"10.1145\/2610384.2628055"},{"key":"e_1_2_1_20_1","volume-title":"Advances in Neural Information Processing Systems 32: Annual Conference on Neural Information Processing Systems 2019","author":"Kulal Sumith","year":"2019","unstructured":"Sumith Kulal, Panupong Pasupat, Kartik Chandra, Mina Lee, Oded Padon, Alex Aiken, and Percy Liang. 2019. SPoC: Search-based Pseudocode to Code. In Advances in Neural Information Processing Systems 32: Annual Conference on Neural Information Processing Systems 2019, NeurIPS 2019, December 8-14, 2019, Vancouver, BC, Canada, Hanna M. Wallach, Hugo Larochelle, Alina Beygelzimer, Florence d'Alch\u00e9-Buc, Emily B. Fox, and Roman Garnett (Eds.). 11883-11894. https:\/\/proceedings.neurips.cc\/paper\/2019\/hash\/7298332f04ac004a0ca44cc69ecf6f6b-Abstract.html"},{"key":"e_1_2_1_21_1","doi-asserted-by":"publisher","DOI":"10.48550\/ARXIV.2412.15676"},{"key":"e_1_2_1_22_1","doi-asserted-by":"publisher","DOI":"10.1145\/3540250.3549081"},{"key":"e_1_2_1_23_1","volume-title":"ROUGE: A Package for Automatic Evaluation of Summaries. In Text Summarization Branches Out","author":"Lin Chin-Yew","year":"2004","unstructured":"Chin-Yew Lin. 2004. ROUGE: A Package for Automatic Evaluation of Summaries. In Text Summarization Branches Out. Association for Computational Linguistics, Barcelona, Spain, 74-81. https:\/\/aclanthology.org\/W04-1013\/"},{"key":"e_1_2_1_24_1","doi-asserted-by":"publisher","DOI":"10.48550\/ARXIV.2409.02977"},{"key":"e_1_2_1_25_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-031-90900-9_3"},{"key":"e_1_2_1_26_1","doi-asserted-by":"publisher","DOI":"10.1109\/ISSRE59848.2023.00026"},{"key":"e_1_2_1_27_1","doi-asserted-by":"publisher","DOI":"10.3115\/1073083.1073135"},{"key":"e_1_2_1_28_1","unstructured":"Qodana AI. 2025. PR-Agent: AI-Powered Pull Request Agent. https:\/\/github.com\/qodo-ai\/pr-agent"},{"key":"e_1_2_1_29_1","doi-asserted-by":"publisher","DOI":"10.48550\/ARXIV.2404.18496"},{"key":"e_1_2_1_30_1","doi-asserted-by":"publisher","DOI":"10.1016\/J.JSS.2023.111729"},{"key":"e_1_2_1_31_1","doi-asserted-by":"publisher","DOI":"10.1145\/3643775"},{"key":"e_1_2_1_32_1","doi-asserted-by":"publisher","DOI":"10.48550\/ARXIV.2406"},{"key":"e_1_2_1_33_1","unstructured":"SonarSource. 2006. SonarQube. https:\/\/www.sonarsource.com\/products\/sonarqube\/"},{"key":"e_1_2_1_34_1","doi-asserted-by":"publisher","DOI":"10.48550\/ARXIV.2501.15134"},{"key":"e_1_2_1_35_1","volume-title":"Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, EMNLP 2024","author":"Tang Xunzhu","year":"2024","unstructured":"Xunzhu Tang, Kisub Kim, Yewei Song, Cedric Lothritz, Bei Li, Saad Ezzini, Haoye Tian, Jacques Klein, and Tegawend\u00e9 F. Bissyand\u00e9. 2024. CodeAgent: Autonomous Communicative Agents for Code Review. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, EMNLP 2024, Miami, FL, USA, November 12-16, 2024, Yaser Al-Onaizan, Mohit Bansal, and Yun-Nung Chen (Eds.). Association for Computational Linguistics, 11279-11313. https:\/\/aclanthology.org\/2024.emnlp-main.632"},{"key":"e_1_2_1_36_1","doi-asserted-by":"publisher","DOI":"10.1145\/3597503.3639194"},{"key":"e_1_2_1_37_1","doi-asserted-by":"publisher","DOI":"10.48550\/ARXIV.2406.12624"},{"key":"e_1_2_1_38_1","doi-asserted-by":"publisher","DOI":"10.1145\/3510003.3510067"},{"key":"e_1_2_1_39_1","doi-asserted-by":"publisher","DOI":"10.48550\/ARXIV.2302.13971"},{"key":"e_1_2_1_40_1","doi-asserted-by":"publisher","DOI":"10.1145\/3510003.3510621"},{"key":"e_1_2_1_41_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICSE43902.2021.00027"},{"key":"e_1_2_1_42_1","doi-asserted-by":"publisher","DOI":"10.18653\/V1\/2021.EMNLP-MAIN.685"},{"key":"e_1_2_1_43_1","volume-title":"PandaLM: An Automatic Evaluation Benchmark for LLM Instruction Tuning Optimization. In The Twelfth International Conference on Learning Representations, ICLR 2024","author":"Wang Yidong","year":"2024","unstructured":"Yidong Wang, Zhuohao Yu, Wenjin Yao, Zhengran Zeng, Linyi Yang, Cunxiang Wang, Hao Chen, Chaoya Jiang, Rui Xie, Jindong Wang, Xing Xie, Wei Ye, Shikun Zhang, and Yue Zhang. 2024. PandaLM: An Automatic Evaluation Benchmark for LLM Instruction Tuning Optimization. In The Twelfth International Conference on Learning Representations, ICLR 2024, Vienna, Austria, May 7-11, 2024. OpenReview.net. https:\/\/openreview.net\/forum?id=5Nn2BLV7SB"},{"key":"e_1_2_1_44_1","unstructured":"Wikipedia. 2025. Cohen's kappa. https:\/\/en.wikipedia.org\/wiki\/Cohen's_kappa. Accessed: 2025-09-11."},{"key":"e_1_2_1_45_1","unstructured":"Wikipedia. 2025. Stratified sampling. https:\/\/en.wikipedia.org\/wiki\/Stratified_sampling. Accessed: 2025-09-11."},{"key":"e_1_2_1_46_1","doi-asserted-by":"publisher","DOI":"10.48550\/ARXIV.2407.01489"},{"key":"e_1_2_1_47_1","volume-title":"SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering. In Advances in Neural Information Processing Systems 38: Annual Conference on Neural Information Processing Systems 2024","author":"Yang John","year":"2024","unstructured":"John Yang, Carlos E. Jimenez, Alexander Wettig, Kilian Lieret, Shunyu Yao, Karthik Narasimhan, and Ofir Press. 2024. SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering. In Advances in Neural Information Processing Systems 38: Annual Conference on Neural Information Processing Systems 2024, NeurIPS 2024, Vancouver, BC, Canada, December 10 -15, 2024, Amir Globersons, Lester Mackey, Danielle Belgrave, Angela Fan, Ulrich Paquet, Jakub M. Tomczak, and Cheng Zhang (Eds.). http:\/\/papers.nips.cc\/paper_files\/paper\/2024\/hash\/ 5a7c947568c1b1328ccc5230172e1e7c-Abstract-Conference.html"},{"key":"e_1_2_1_48_1","doi-asserted-by":"publisher","DOI":"10.48550\/ARXIV.2405.18216"},{"key":"e_1_2_1_49_1","doi-asserted-by":"publisher","DOI":"10.1145\/3650212.3652115"},{"key":"e_1_2_1_50_1","doi-asserted-by":"publisher","DOI":"10.1145\/3460319.3464819"},{"key":"e_1_2_1_51_1","volume-title":"The Thirteenth International Conference on Learning Representations, ICLR 2025","author":"Zhu Lianghui","year":"2025","unstructured":"Lianghui Zhu, Xinggang Wang, and Xinlong Wang. 2025. JudgeLM: Fine-tuned Large Language Models are Scalable Judges. In The Thirteenth International Conference on Learning Representations, ICLR 2025, Singapore, April 24-28, 2025. OpenReview.net. https:\/\/openreview.net\/forum?id=xsELpEPn4A"}],"container-title":["Proceedings of the ACM on Software Engineering"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3808144","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,6,30]],"date-time":"2026-06-30T18:14:55Z","timestamp":1782843295000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3808144"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,6,30]]},"references-count":51,"journal-issue":{"issue":"FSE","published-print":{"date-parts":[[2026,6,30]]}},"alternative-id":["10.1145\/3808144"],"URL":"https:\/\/doi.org\/10.1145\/3808144","relation":{},"ISSN":["2994-970X"],"issn-type":[{"value":"2994-970X","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,6,30]]}}}