{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,14]],"date-time":"2026-07-14T23:51:30Z","timestamp":1784073090269,"version":"3.55.0"},"reference-count":80,"publisher":"Association for Computing Machinery (ACM)","issue":"5","license":[{"start":{"date-parts":[[2025,5,24]],"date-time":"2025-05-24T00:00:00Z","timestamp":1748044800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Softw. Eng. Methodol."],"published-print":{"date-parts":[[2025,6,30]]},"abstract":"<jats:p>Large language models (LLMs) have become instrumental in advancing software engineering (SE) tasks, showcasing their efficacy in code understanding and beyond. AI code models have demonstrated their value not only in code generation but also in defect detection, enhancing security measures and improving overall software quality. They are emerging as crucial tools for both software development and maintaining software quality. Like traditional SE tools, open source collaboration is key in realizing the excellent products. However, with AI models, the essential need is in data. The collaboration of these AI-based SE models hinges on maximizing the sources of high-quality data. However, data, especially of high quality, often hold commercial or sensitive value, making them less accessible for open source AI-based SE projects. This reality presents a significant barrier to the development and enhancement of AI-based SE tools within the SE community. Therefore, researchers need to find solutions for enabling open source AI-based SE models to tap into resources by different organizations. Addressing this challenge, our position article investigates one solution to facilitate access to diverse organizational resources for open source AI models, ensuring that privacy and commercial sensitivities are respected. We introduce a governance framework centered on federated learning (FL), designed to foster the joint development and maintenance of open source AI code models while safeguarding data privacy and security. Additionally, we present guidelines for developers on AI-based SE tool collaboration, covering data requirements, model architecture, updating strategies, and version control. Given the significant influence of data characteristics on FL, our research examines the effect of code data heterogeneity on FL performance. We consider six different scenarios of data distributions and include four code models. We also include four most common FL algorithms. Our experimental findings highlight the potential for employing FL in the collaborative development and maintenance of AI-based SE models. We also discuss the key issues to be addressed in the co-construction process and future research directions.<\/jats:p>","DOI":"10.1145\/3708529","type":"journal-article","created":{"date-parts":[[2024,12,18]],"date-time":"2024-12-18T12:05:27Z","timestamp":1734523527000},"page":"1-24","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":12,"title":["Open Source AI-based SE Tools: Opportunities and Challenges of Collaborative Software Learning"],"prefix":"10.1145","volume":"34","author":[{"ORCID":"https:\/\/orcid.org\/0009-0003-4868-2805","authenticated-orcid":false,"given":"Zhihao","family":"Lin","sequence":"first","affiliation":[{"name":"Beihang University, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-0044-466X","authenticated-orcid":false,"given":"Wei","family":"Ma","sequence":"additional","affiliation":[{"name":"Singapore Management University, Singapore, Singapore"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-3246-6935","authenticated-orcid":false,"given":"Tao","family":"Lin","sequence":"additional","affiliation":[{"name":"Westlake University, Hangzhou, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-8953-0782","authenticated-orcid":false,"given":"Yaowen","family":"Zheng","sequence":"additional","affiliation":[{"name":"Nanyang Technological University, Singapore, Singapore"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-9879-0260","authenticated-orcid":false,"given":"Jingquan","family":"Ge","sequence":"additional","affiliation":[{"name":"Nanyang Technological University, Singapore, Singapore"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-0481-5341","authenticated-orcid":false,"given":"Jun","family":"Wang","sequence":"additional","affiliation":[{"name":"University of Luxembourg, Esch-sur-Alzette, Luxembourg, Luxembourg"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-4052-475X","authenticated-orcid":false,"given":"Jacques","family":"Klein","sequence":"additional","affiliation":[{"name":"University of Luxembourg, Esch-sur-Alzette, Luxembourg, Luxembourg"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-7270-9869","authenticated-orcid":false,"given":"Tegawend\u00e9 F.","family":"Bissyand\u00e9","sequence":"additional","affiliation":[{"name":"University of Luxembourg, Esch-sur-Alzette, Luxembourg, Luxembourg"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-7300-9215","authenticated-orcid":false,"given":"Yang","family":"Liu","sequence":"additional","affiliation":[{"name":"Nanyang Technological University, Singapore, Singapore"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-2990-1614","authenticated-orcid":false,"given":"Li","family":"Li","sequence":"additional","affiliation":[{"name":"Beihang University, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2025,5,24]]},"reference":[{"key":"e_1_3_2_2_2","unstructured":"ChatGPT. 2022-11. ChatGpt: Optimizing Language Models for Dialogue. Retrieved from https:\/\/chat.openai.com"},{"key":"e_1_3_2_3_2","unstructured":"Hugging Face. 2024. Huggingface: The AI Community Building the Future. Retrieved from https:\/\/huggingface.co\/"},{"key":"e_1_3_2_4_2","unstructured":"Pekka Abrahamsson Outi Salo Jussi Ronkainen and Juhani Warsta. 2017. Agile software development methods: Review and analysis. arXiv:1709.08439. Retrieved from https:\/\/arxiv.org\/abs\/1709.08439"},{"key":"e_1_3_2_5_2","doi-asserted-by":"crossref","unstructured":"Sadi Alawadi Khalid Alkharabsheh Fahed Alkhabbas Victor Kebande Feras M Awaysheh and Fabio Palomba. 2023. Fedcsd: A federated learning based approach for code-smell detection. arXiv:2306.00038. Retrieved from https:\/\/arxiv.org\/abs\/2306.00038","DOI":"10.1109\/ACCESS.2024.3380167"},{"key":"e_1_3_2_6_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICSE-SEIP.2019.00042"},{"key":"e_1_3_2_7_2","doi-asserted-by":"publisher","DOI":"10.1109\/CEC.2008.4630793"},{"key":"e_1_3_2_8_2","doi-asserted-by":"publisher","DOI":"10.1145\/1150402.1150464"},{"key":"e_1_3_2_9_2","unstructured":"Mark Chen Jerry Tworek Heewoo Jun Qiming Yuan Henrique Ponde de Oliveira Pinto Jared Kaplan Harri Edwards Yuri Burda Nicholas Joseph Greg Brockman et al. 2021. Evaluating large language models trained on code. arXiv:2107.03374. Retrieved from https:\/\/arxiv.org\/abs\/2107.03374"},{"key":"e_1_3_2_10_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.00159"},{"key":"e_1_3_2_11_2","unstructured":"Olivia Choudhury Aris Gkoulalas-Divanis Theodoros Salonidis Issa Sylla Yoonyoung Park Grace Hsu and Amar Das. 2019. Differential privacy-enabled federated learning for sensitive health data. arXiv:1910.02578. Retrieved from https:\/\/arxiv.org\/abs\/1910.02578"},{"key":"e_1_3_2_12_2","doi-asserted-by":"publisher","DOI":"10.1109\/82.959866"},{"key":"e_1_3_2_13_2","doi-asserted-by":"publisher","DOI":"10.1145\/3561048"},{"key":"e_1_3_2_14_2","doi-asserted-by":"publisher","DOI":"10.1007\/11787006_1"},{"key":"e_1_3_2_15_2","doi-asserted-by":"publisher","DOI":"10.1145\/3412569.3412579"},{"key":"e_1_3_2_16_2","unstructured":"Angela Fan Beliz Gokkaya Mark Harman Mitya Lyubarskiy Shubho Sengupta Shin Yoo and Jie M. Zhang. 2023. Large language models for software engineering: Survey and open problems. arXiv:2310.03533. Retrieved from https:\/\/arxiv.org\/abs\/2310.03533"},{"issue":"2","key":"e_1_3_2_17_2","first-page":"333","article-title":"CausaLM: Causal model explanation through counterfactual language models","volume":"47","author":"Feder Amir","year":"2021","unstructured":"Amir Feder, Nadav Oved, Uri Shalit, and Roi Reichart. 2021. CausaLM: Causal model explanation through counterfactual language models. Computational Linguistics 47, 2 (2021), 333\u2013386.","journal-title":"Computational Linguistics"},{"key":"e_1_3_2_18_2","doi-asserted-by":"crossref","unstructured":"Zhangyin Feng Daya Guo Duyu Tang Nan Duan Xiaocheng Feng Ming Gong Linjun Shou Bing Qin Ting Liu Daxin Jiang et al. 2020. Codebert: A pre-trained model for programming and natural languages. arXiv:2002.08155. Retrieved from https:\/\/arxiv.org\/abs\/2002.08155","DOI":"10.18653\/v1\/2020.findings-emnlp.139"},{"key":"e_1_3_2_19_2","doi-asserted-by":"publisher","DOI":"10.1007\/s11263-021-01453-z"},{"key":"e_1_3_2_20_2","unstructured":"Daya Guo Shuo Ren Shuai Lu Zhangyin Feng Duyu Tang Shujie Liu Long Zhou Nan Duan Alexey Svyatkovskiy Shengyu Fu et al. 2020. Graphcodebert: Pre-training code representations with data flow. arXiv:2009.08366. Retrieved from https:\/\/arxiv.org\/abs\/2009.08366"},{"key":"e_1_3_2_21_2","unstructured":"Sirui Hong Xiawu Zheng Jonathan Chen Yuheng Cheng Jinlin Wang Ceyao Zhang Zili Wang Steven Ka Shing Yau Zijuan Lin Liyang Zhou et al. 2023. Metagpt: Meta programming for multi-agent collaborative framework. arXiv:2308.00352. Retrieved from https:\/\/arxiv.org\/abs\/2308.00352"},{"key":"e_1_3_2_22_2","doi-asserted-by":"publisher","DOI":"10.1145\/3695988"},{"key":"e_1_3_2_23_2","first-page":"2790","volume-title":"International Conference on Machine Learning","author":"Houlsby Neil","year":"2019","unstructured":"Neil Houlsby, Andrei Giurgiu, Stanislaw Jastrzebski, Bruna Morrone, Quentin De Laroussilhe, Andrea Gesmundo, Mona Attariyan, and Sylvain Gelly. 2019. Parameter-efficient transfer learning for NLP. In International Conference on Machine Learning. PMLR, 2790\u20132799."},{"key":"e_1_3_2_24_2","unstructured":"Edward J. Hu Yelong Shen Phillip Wallis Zeyuan Allen-Zhu Yuanzhi Li Shean Wang Lu Wang and Weizhu Chen. 2021. Lora: Low-rank adaptation of large language models. arXiv:2106.09685. Retrieved from https:\/\/arxiv.org\/abs\/2106.09685"},{"key":"e_1_3_2_25_2","unstructured":"Hamel Husain Ho-Hsiang Wu Tiferet Gazit Miltiadis Allamanis and Marc Brockschmidt. 2019. Codesearchnet challenge: Evaluating the state of semantic code search. arXiv:1909.09436. Retrieved from https:\/\/arxiv.org\/abs\/1909.09436"},{"key":"e_1_3_2_26_2","unstructured":"Johannes Rude Jensen Victor von Wachter and Omri Ross. 2021. How decentralized is the governance of blockchain-based finance: Empirical evidence from four governance token distributions. arXiv:2102.10096. Retrieved from https:\/\/arxiv.org\/abs\/2102.10096"},{"key":"e_1_3_2_27_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-031-19827-4_41"},{"key":"e_1_3_2_28_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICSE43902.2021.00107"},{"key":"e_1_3_2_29_2","unstructured":"Jared Kaplan Sam McCandlish Tom Henighan Tom B. Brown Benjamin Chess Rewon Child Scott Gray Alec Radford Jeffrey Wu and Dario Amodei. 2020. Scaling laws for neural language models. arXiv:2001.08361. Retrieved from https:\/\/arxiv.org\/abs\/2001.08361"},{"key":"e_1_3_2_30_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-031-19827-4_36"},{"key":"e_1_3_2_31_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICC42927.2021.9500877"},{"key":"e_1_3_2_32_2","unstructured":"Jakub Konecn\u1ef3 H. Brendan McMahan Felix X. Yu Peter Richt\u00e1rik Ananda Theertha Suresh and Dave Bacon. 2016. Federated learning: Strategies for improving communication efficiency. arXiv:1610.05492. Retrieved from https:\/\/arxiv.org\/abs\/1610.05492"},{"key":"e_1_3_2_33_2","unstructured":"Raymond Li Loubna Ben Allal Yangtian Zi Niklas Muennighoff Denis Kocetkov Chenghao Mou Marc Marone Christopher Akiki Jia Li Jenny Chim et al. 2023. Starcoder: May the source be with you! arXiv:2305.06161. Retrieved from https:\/\/arxiv.org\/abs\/2305.06161"},{"key":"e_1_3_2_34_2","first-page":"429","volume-title":"Proceedings of Machine Learning and Systems","volume":"2","author":"Li Tian","year":"2020","unstructured":"Tian Li, Anit Kumar Sahu, Manzil Zaheer, Maziar Sanjabi, Ameet Talwalkar, and Virginia Smith. 2020. Federated optimization in heterogeneous networks. Proceedings of Machine Learning and Systems 2 (2020), 429\u2013450."},{"key":"e_1_3_2_35_2","doi-asserted-by":"publisher","DOI":"10.1109\/MNET.011.2000263"},{"key":"e_1_3_2_36_2","unstructured":"Zhen Li Deqing Zou Shouhuai Xu Xinyu Ou Hai Jin Sujuan Wang Zhijun Deng and Yuyi Zhong. 2018. Vuldeepecker: A deep learning-based system for vulnerability detection. arXiv:1801.01681. Retrieved from https:\/\/arxiv.org\/abs\/1801.01681"},{"key":"e_1_3_2_37_2","unstructured":"Shangqing Liu Yanzhou Li Xiaofei Xie and Yang Liu. 2022. Commitbart: A large pre-trained model for github commits. arXiv:2208.08100. Retrieved from https:\/\/arxiv.org\/abs\/2208.08100"},{"key":"e_1_3_2_38_2","unstructured":"Shuai Lu Daya Guo Shuo Ren Junjie Huang Alexey Svyatkovskiy Ambrosio Blanco Colin Clement Dawn Drain Daxin Jiang Duyu Tang et al. 2021. Codexglue: A machine learning benchmark dataset for code understanding and generation. arXiv:2102.04664. Retrieved from https:\/\/arxiv.org\/abs\/2102.04664"},{"key":"e_1_3_2_39_2","unstructured":"Ziyang Luo Can Xu Pu Zhao Qingfeng Sun Xiubo Geng Wenxiang Hu Chongyang Tao Jing Ma Qingwei Lin and Daxin Jiang. 2023. Wizardcoder: Empowering code large language models with evol-instruct. arXiv:2306.08568. Retrieved from https:\/\/arxiv.org\/abs\/2306.08568"},{"key":"e_1_3_2_40_2","unstructured":"Wei Ma Shangqing Liu Zhihao Lin Wenhan Wang Qiang Hu Ye Liu Cen Zhang Liming Nie Li Li and Yang Liu. 2023. LMs: Understanding code syntax and semantics for code analysis. arXiv:2305.12138. Retrieved from https:\/\/arxiv.org\/abs\/2305.12138"},{"key":"e_1_3_2_41_2","doi-asserted-by":"publisher","DOI":"10.1145\/3664606"},{"key":"e_1_3_2_42_2","first-page":"1273","volume-title":"20th International Conference on Artificial Intelligence and Statistics","author":"McMahan Brendan","year":"2017","unstructured":"Brendan McMahan, Eider Moore, Daniel Ramage, Seth Hampson, and Blaise Aguera y Arcas. 2017. Communication-efficient learning of deep networks from decentralized data. In 20th International Conference on Artificial Intelligence and Statistics. PMLR, 1273\u20131282."},{"key":"e_1_3_2_43_2","unstructured":"H. Brendan McMahan Eider Moore Daniel Ramage and Blaise Ag\u00fcera y Arcas. 2016. Federated learning of deep networks using model averaging. arXiv:1602.05629. Retrieved from https:\/\/arxiv.org\/abs\/1602.05629"},{"key":"e_1_3_2_44_2","doi-asserted-by":"publisher","DOI":"10.1109\/SP.2019.00065"},{"key":"e_1_3_2_45_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.imavis.2011.02.002"},{"key":"e_1_3_2_46_2","doi-asserted-by":"publisher","DOI":"10.1002\/widm.1356"},{"key":"e_1_3_2_47_2","unstructured":"Open-Source AI Models. 2024. Open-Source AI Models. Retrieved from https:\/\/github.com\/mathieu0905\/collaborative_software_learning"},{"key":"e_1_3_2_48_2","doi-asserted-by":"publisher","DOI":"10.1109\/TKDE.2009.191"},{"key":"e_1_3_2_49_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPDS.2005.101"},{"key":"e_1_3_2_50_2","doi-asserted-by":"publisher","DOI":"10.1145\/3485128"},{"key":"e_1_3_2_51_2","unstructured":"Baptiste Roziere Jonas Gehring Fabian Gloeckle Sten Sootla Itai Gat Xiaoqing Ellen Tan Yossi Adi Jingyu Liu Tal Remez J\u00e9r\u00e9my Rapin et al. 2023. Code llama: Open foundation models for code. arXiv:2308.12950. Retrieved from https:\/\/arxiv.org\/abs\/2308.12950"},{"issue":"5","key":"e_1_3_2_52_2","first-page":"1","article-title":"Collaborative machine learning without centralized training data for federated learning","volume":"5","author":"Satish Snehal","year":"2022","unstructured":"Snehal Satish, Geeta Sandeep Nadella, Karthik Meduri, and Hari Gonaygunta. 2022. Collaborative machine learning without centralized training data for federated learning. International Machine Learning Journal and Computer Engineering 5, 5 (2022), 1\u201314.","journal-title":"International Machine Learning Journal and Computer Engineering"},{"key":"e_1_3_2_53_2","doi-asserted-by":"publisher","DOI":"10.1109\/ACCESS.2020.2970495"},{"key":"e_1_3_2_54_2","doi-asserted-by":"crossref","unstructured":"Koustuv Sinha Robin Jia Dieuwke Hupkes Joelle Pineau Adina Williams and Douwe Kiela. 2021. Masked language modeling and the distributional hypothesis: Order word matters pre-training for little. arXiv:2104.06644. Retrieved from https:\/\/arxiv.org\/abs\/2104.06644","DOI":"10.18653\/v1\/2021.emnlp-main.230"},{"key":"e_1_3_2_55_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v34i09.7123"},{"issue":"1","key":"e_1_3_2_56_2","first-page":"61","article-title":"Contextual learning: Innovative approach towards the development of students \\({}^{\\prime}\\)  scientific attitude and natural Science performance","volume":"14","author":"Suryawati Evi","year":"2017","unstructured":"Evi Suryawati and Kamisah Osman. 2017. Contextual learning: Innovative approach towards the development of students \\({}^{\\prime}\\) scientific attitude and natural Science performance. Eurasia Journal of Mathematics, Science and Technology Education 14, 1 (2017), 61\u201376.","journal-title":"Eurasia Journal of Mathematics, Science and Technology Education"},{"key":"e_1_3_2_57_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.compag.2018.03.032"},{"key":"e_1_3_2_58_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.rser.2012.03.014"},{"key":"e_1_3_2_59_2","doi-asserted-by":"publisher","DOI":"10.1007\/s43681-021-00043-6"},{"key":"e_1_3_2_60_2","doi-asserted-by":"publisher","DOI":"10.1109\/TSE.2022.3178469"},{"key":"e_1_3_2_61_2","doi-asserted-by":"publisher","DOI":"10.1109\/TSE.2024.3368208"},{"key":"e_1_3_2_62_2","doi-asserted-by":"publisher","DOI":"10.12691\/acis-3-1-3"},{"key":"e_1_3_2_63_2","doi-asserted-by":"crossref","unstructured":"Yue Wang Hung Le Akhilesh Deepak Gotmare Nghi D. Q. Bui Junnan Li and Steven C. H. Hoi. 2023. Codet5+: Open code large language models for code understanding and generation. arXiv:2305.07922. Retrieved from https:\/\/arxiv.org\/abs\/2305.07922","DOI":"10.18653\/v1\/2023.emnlp-main.68"},{"key":"e_1_3_2_64_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2021.emnlp-main.685"},{"key":"e_1_3_2_65_2","unstructured":"Zhilin Wang and Qin Hu. 2021. Blockchain-based federated learning: A comprehensive survey. arXiv:2110.02182. Retrieved from https:\/\/arxiv.org\/abs\/2110.02182"},{"key":"e_1_3_2_66_2","doi-asserted-by":"publisher","DOI":"10.1007\/s10602-005-2234-6"},{"key":"e_1_3_2_67_2","first-page":"795","volume-title":"Proceedings of Machine Learning and Systems","volume":"4","author":"Wu Carole-Jean","year":"2022","unstructured":"Carole-Jean Wu, Ramya Raghavendra, Udit Gupta, Bilge Acun, Newsha Ardalani, Kiwan Maeng, Gloria Chang, Fiona Aga, Jinshi Huang, Charles Bai, et al. 2022. Sustainable AI: Environmental implications, challenges and opportunities. Proceedings of Machine Learning and Systems 4 (2022), 795\u2013813."},{"key":"e_1_3_2_68_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.00046"},{"key":"e_1_3_2_69_2","unstructured":"Miao Xiong Zhiyuan Hu Xinyang Lu Yifei Li Jie Fu Junxian He and Bryan Hooi. 2023. Can llms express their uncertainty? An empirical evaluation of confidence elicitation in llms. arXiv:2306.13063. Retrieved from https:\/\/arxiv.org\/abs\/2306.13063"},{"key":"e_1_3_2_70_2","doi-asserted-by":"publisher","DOI":"10.1145\/3520312.3534862"},{"key":"e_1_3_2_71_2","unstructured":"Hanxiang Xu Wei Ma Ting Zhou Yanjie Zhao Kai Chen Qiang Hu Yang Liu and Haoyu Wang. 2024. A code knowledge graph-enhanced system for LLM-based fuzz driver generation. arXiv:2411.11532. Retrieved from https:\/\/arxiv.org\/abs\/2411.11532"},{"key":"e_1_3_2_72_2","unstructured":"John Yang Carlos E. Jimenez Alexander Wettig Shunyu Yao Karthik Narasimhan and Ofir Press. 2024. SWE-agent: Agent computer interfaces enable software engineering language models. arXiv.2405.15793. Retrieved from https:\/\/arxiv.org\/abs\/2405.15793"},{"key":"e_1_3_2_73_2","doi-asserted-by":"publisher","DOI":"10.1145\/3298981"},{"key":"e_1_3_2_74_2","doi-asserted-by":"publisher","DOI":"10.1109\/TSE.2023.3347898"},{"key":"e_1_3_2_75_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.hcc.2024.100211"},{"key":"e_1_3_2_76_2","first-page":"5650","volume-title":"International Conference on Machine Learning","author":"Yin Dong","year":"2018","unstructured":"Dong Yin, Yudong Chen, Ramchandran Kannan, and Peter Bartlett. 2018. Byzantine-robust distributed learning: Towards optimal statistical rates. In International Conference on Machine Learning. PMLR, 5650\u20135659."},{"key":"e_1_3_2_77_2","doi-asserted-by":"publisher","DOI":"10.1145\/3660783"},{"key":"e_1_3_2_78_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.knosys.2021.106775"},{"key":"e_1_3_2_79_2","doi-asserted-by":"publisher","DOI":"10.1145\/3650212.3680355"},{"key":"e_1_3_2_80_2","first-page":"1","article-title":"Devign: Effective vulnerability identification by learning comprehensive program semantics via graph neural networks","volume":"32","author":"Zhou Yaqin","year":"2019","unstructured":"Yaqin Zhou, Shangqing Liu, Jingkai Siow, Xiaoning Du, and Yang Liu. 2019. Devign: Effective vulnerability identification by learning comprehensive program semantics via graph neural networks. Advances in Neural Information Processing Systems 32 (2019), 1\u201312.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_2_81_2","first-page":"465","article-title":"Federated learning on non-IID data: A survey","author":"Zhu Hangyu","year":"2021","unstructured":"Hangyu Zhu, Jinjin Xu, Shiqing Liu, and Yaochu Jin. 2021. Federated learning on non-IID data: A survey. Neurocomputing 465 (2021), 371\u2013390.","journal-title":"Neurocomputing"}],"container-title":["ACM Transactions on Software Engineering and Methodology"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3708529","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3708529","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,19]],"date-time":"2025-06-19T01:17:46Z","timestamp":1750295866000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3708529"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,5,24]]},"references-count":80,"journal-issue":{"issue":"5","published-print":{"date-parts":[[2025,6,30]]}},"alternative-id":["10.1145\/3708529"],"URL":"https:\/\/doi.org\/10.1145\/3708529","relation":{},"ISSN":["1049-331X","1557-7392"],"issn-type":[{"value":"1049-331X","type":"print"},{"value":"1557-7392","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,5,24]]},"assertion":[{"value":"2024-04-06","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2024-11-27","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-05-24","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}