{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,2]],"date-time":"2026-07-02T04:41:52Z","timestamp":1782967312404,"version":"3.54.5"},"reference-count":45,"publisher":"Oxford University Press (OUP)","issue":"3","license":[{"start":{"date-parts":[[2021,12,18]],"date-time":"2021-12-18T00:00:00Z","timestamp":1639785600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/academic.oup.com\/journals\/pages\/open_access\/funder_policies\/chorus\/standard_publication_model"}],"funder":[{"name":"Basic Research Project of Shenzhen","award":["JCYJ20190806142601687"],"award-info":[{"award-number":["JCYJ20190806142601687"]}]},{"name":"Basic Research Project of Shenzhen","award":["JCYJ20190806143418198"],"award-info":[{"award-number":["JCYJ20190806143418198"]}]},{"name":"Basic Research Project of Shenzhen","award":["JCYJ20200109113405927"],"award-info":[{"award-number":["JCYJ20200109113405927"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":[],"published-print":{"date-parts":[[2023,3,15]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:p>Federated learning (FL), a variant of distributed learning (DL), supports the training of a shared model without accessing private data from different sources. Despite its benefits with regard to privacy preservation, FL\u2019s distributed nature and privacy constraints make it vulnerable to data poisoning attacks. Existing defenses, primarily designed for DL, are typically not well adapted to FL. In this paper, we study such attacks and defenses. In doing so, we start from the perspective of DL and then give consideration to a real-world FL scenario, with the aim being to explore the requisites of a desirable defense in FL. Our study shows that (i) the batch size used in each training round affects the effectiveness of defenses in DL, (ii) the defenses investigated are somewhat effective and moderately influenced by batch size in FL settings and (iii) the non-IID data makes it more difficult to defend against data poisoning attacks in FL. Based on the findings, we discuss the key challenges and possible directions in defending against such attacks in FL. In addition, we propose detect and suppress the potential outliers(DSPO), a defense against data poisoning attacks in FL scenarios. Our results show that DSPO outperforms other defenses in several cases.<\/jats:p>","DOI":"10.1093\/comjnl\/bxab192","type":"journal-article","created":{"date-parts":[[2021,11,18]],"date-time":"2021-11-18T12:09:31Z","timestamp":1637237371000},"page":"711-726","source":"Crossref","is-referenced-by-count":15,"title":["Defending Against Data Poisoning Attacks: From Distributed Learning to Federated Learning"],"prefix":"10.1093","volume":"66","author":[{"given":"Yuchen","family":"Tian","sequence":"first","affiliation":[{"name":"College of Computer Science and Technology , Harbin Institute of Technology (Shenzhen), Shenzhen 518055, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Weizhe","family":"Zhang","sequence":"additional","affiliation":[{"name":"College of Computer Science and Technology , Harbin Institute of Technology (Shenzhen), Shenzhen 518055, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Andrew","family":"Simpson","sequence":"additional","affiliation":[{"name":"Department of Computer Science , University of Oxford, Oxford OX1 3QD, UK"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Yang","family":"Liu","sequence":"additional","affiliation":[{"name":"College of Computer Science and Technology , Harbin Institute of Technology (Shenzhen), Shenzhen 518055, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Zoe Lin","family":"Jiang","sequence":"additional","affiliation":[{"name":"College of Computer Science and Technology , Harbin Institute of Technology (Shenzhen), Shenzhen 518055, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"286","published-online":{"date-parts":[[2021,12,18]]},"reference":[{"key":"2023031708554200700_","doi-asserted-by":"crossref","first-page":"248","DOI":"10.1093\/comjnl\/bxx059","article-title":"Multi-objective optimization techniques for task scheduling problem in distributed systems","volume":"61","author":"Sarathambekai","year":"2018","journal-title":"Comput. J."},{"key":"2023031708554200700_","first-page":"102582","article-title":"Computational intelligence intrusion detection techniques in mobile cloud computing environments: Review, taxonomy, and open research issues","volume":"55","author":"Shamshirband","year":"2020","journal-title":"J. Inform. Secur. Appl."},{"key":"2023031708554200700_","doi-asserted-by":"crossref","first-page":"2802","DOI":"10.1109\/TPDS.2020.3003307","article-title":"Distributed training of deep learning models: A taxonomic perspective","volume":"31","author":"Langer","year":"2020","journal-title":"IEEE Trans. Parallel Distrib. Syst."},{"key":"2023031708554200700_","first-page":"1273","article-title":"Communication-Efficient Learning of Deep Networks from Decentralized Data","volume-title":"Proc. 20th Int. Conf. Artificial Intelligence and Statistics","author":"Mcmahan","year":"2017"},{"key":"2023031708554200700_","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/3339474","article-title":"Comparison and modelling of country-level microblog user and activity in cyber-physical-social systems using Weibo and Twitter data","volume":"10","author":"Yang","year":"2019","journal-title":"ACM Trans. Intell. Syst. Technol."},{"key":"2023031708554200700_","doi-asserted-by":"crossref","first-page":"242","DOI":"10.1109\/MNET.001.1900506","article-title":"On safeguarding privacy and security in the framework of federated learning","volume":"34","author":"Ma","year":"2020","journal-title":"IEEE Netw."},{"key":"2023031708554200700_","doi-asserted-by":"crossref","first-page":"8861","DOI":"10.1109\/ICASSP40776.2020.9054676","article-title":"On the Byzantine Robustness of Clustered Federated Learning","volume-title":"2020 IEEE Int. Conf. Acoustics, Speech and Signal Processing (ICASSP)","author":"Sattler","year":"2020"},{"key":"2023031708554200700_","first-page":"168","article-title":"Poisoning Attacks in Federated Learning: An Evaluation on Traffic Sign Classification","volume-title":"Proc. Tenth ACM Conf. Data and Application Security and Privacy","author":"Nuding","year":"2020"},{"key":"2023031708554200700_","doi-asserted-by":"crossref","first-page":"27","DOI":"10.1145\/3128572.3140451","article-title":"Towards Poisoning of Deep Learning Algorithms with Back-gradient Optimization","volume-title":"Proc. 10th ACM Workshop on Artificial Intelligence and Security","author":"Mu\u00f1oz-Gonz\u00e1lez","year":"2017"},{"key":"2023031708554200700_","first-page":"119","article-title":"Machine Learning with Adversaries: Byzantine Tolerant Gradient Descent","volume-title":"Advances in Neural Information Processing Systems","author":"Blanchard","year":"2017"},{"key":"2023031708554200700_","doi-asserted-by":"crossref","first-page":"103","DOI":"10.1145\/3128572.3140450","article-title":"Mitigating Poisoning Attacks on Machine Learning Models: A Data Provenance based Approach","volume-title":"Proc. 10th ACM Workshop on Artificial Intelligence and Security","author":"Baracaldo","year":"2017"},{"key":"2023031708554200700_","first-page":"3518","article-title":"Certified Defenses for Data Poisoning Attacks","volume-title":"Advances in Neural Information Processing Systems","author":"Steinhardt","year":"2017"},{"key":"2023031708554200700_","first-page":"6893","article-title":"Zeno: Distributed Stochastic Gradient Descent with Suspicion-Based Fault-Tolerance","volume-title":"36th Int. Conf. Machine Learning","author":"Xie","year":"2019"},{"key":"2023031708554200700_","first-page":"1641","article-title":"Justinian\u2019s GAAvernor: Robust Distributed Learning with Gradient Aggregation Agent","volume-title":"29th USENIX Security Symposium (USENIX Security 20)","author":"Pan","year":"2020"},{"key":"2023031708554200700_","first-page":"6112","article-title":"Robust Learning from Untrusted Sources","volume-title":"36th Int. Conf. Machine Learning (ICML 2019)","author":"Konstantinov","year":"2019"},{"key":"2023031708554200700_","first-page":"5650","article-title":"Byzantine-Robust Distributed Learning: Towards Optimal Statistical Rates","volume-title":"35th Int. Conf. Machine Learning (ICML 2018)","author":"Yin","year":"2018"},{"key":"2023031708554200700_","first-page":"1","article-title":"DBA: Distributed Backdoor Attacks against Federated Learning","volume-title":"Int. Conf. Learning Representations","author":"Xie","year":"2020"},{"key":"2023031708554200700_","first-page":"2938","article-title":"How to Backdoor Federated Learning","volume-title":"Proc. 23rd Int. Conf. Artificial Intelligence and Statistics","author":"Bagdasaryan","year":"2020"},{"key":"2023031708554200700_","first-page":"1807","article-title":"Poisoning Attacks against Support Vector Machines","volume-title":"Proc. 29th Int. Conf. Machine Learning (ICML 2012)","author":"Biggio","year":"2012"},{"key":"2023031708554200700_","doi-asserted-by":"crossref","first-page":"19","DOI":"10.1109\/SP.2018.00057","article-title":"Manipulating Machine Learning: Poisoning Attacks and Countermeasures for Regression Learning","volume-title":"2018 IEEE Symposium on Security and Privacy (SP)","author":"Jagielski","year":"2018"},{"key":"2023031708554200700_","first-page":"381","article-title":"Poisoning Attacks to Graph-Based Recommender Systems","volume-title":"Proc. 34th Annual Computer Security Applications Conf.","author":"Fang","year":"2018"},{"key":"2023031708554200700_","first-page":"1596","article-title":"Sever: A Robust Meta-Algorithm for Stochastic Optimization","volume-title":"Proc. 36th Int. Conf. Machine Learning","author":"Diakonikolas","year":"2019"},{"key":"2023031708554200700_","first-page":"6103","article-title":"Poison Frogs! Targeted Clean-Label Poisoning Attacks on Neural Networks","volume-title":"Advances in Neural Information Processing Systems","author":"Shafahi","year":"2018"},{"key":"2023031708554200700_","first-page":"8230","article-title":"Certified Robustness to Label-Flipping Attacks via Randomized Smoothing","volume-title":"Proc. 37th Int. Conf. Machine Learning","author":"Rosenfeld","year":"2020"},{"key":"2023031708554200700_","doi-asserted-by":"crossref","first-page":"5","DOI":"10.1007\/978-3-030-13453-2_1","article-title":"Label Sanitization against Label Flipping Poisoning Attacks","volume-title":"ECML PKDD 2018 Workshops: Nemesis 2018, UrbReas 2018, SoGood 2018, IWAISe 2018, and Green Data Mining 2018","author":"Paudice","year":"2019"},{"key":"2023031708554200700_","doi-asserted-by":"crossref","DOI":"10.14722\/ndss.2018.23291","article-title":"Trojaning Attack on Neural Networks","volume-title":"Proc. 2018 Network and Distributed System Security Symposium (NDSS)","author":"Liu","year":"2018"},{"key":"2023031708554200700_","doi-asserted-by":"crossref","first-page":"481","DOI":"10.1117\/12.2520275","article-title":"Model Poisoning Attacks against Distributed Machine Learning Systems","volume-title":"Artificial Intelligence and Machine Learning for Multi-Domain Operations Applications","author":"Tomsett","year":"2019"},{"key":"2023031708554200700_","first-page":"634","article-title":"Analyzing Federated Learning through an Adversarial Lens","volume-title":"36th Int. Conf. Machine Learning (ICML 2019)","author":"Bhagoji","year":"2019"},{"key":"2023031708554200700_","first-page":"261","article-title":"Fall of Empires: Breaking Byzantine-Tolerant SGD by Inner Product Manipulation","volume-title":"Proc. 35th Uncertainty in Artificial Intelligence Conf.","author":"Xie","year":"2020"},{"key":"2023031708554200700_","first-page":"1605","article-title":"Local Model Poisoning Attacks to Byzantine-Robust Federated Learning","volume-title":"29th USENIX Security Symposium (USENIX Security 20)","author":"Fang","year":"2020"},{"key":"2023031708554200700_","first-page":"5739","article-title":"Learning with Bad Training Data via Iterative Trimmed Loss Minimization","volume-title":"Proc. 36th Int. Conf. Machine Learning","author":"Shen","year":"2019"},{"key":"2023031708554200700_","first-page":"1","article-title":"Shielding collaborative learning: Mitigating poisoning attacks through client-side detection","volume":"1","author":"Zhao","year":"2020","journal-title":"IEEE Trans. Depend. Secure Comput."},{"key":"2023031708554200700_","first-page":"4613","article-title":"Byzantine Stochastic Gradient Descent","volume-title":"Advances in Neural Information Processing Systems","author":"Alistarh","year":"2018"},{"key":"2023031708554200700_","first-page":"301","article-title":"The Limitations of Federated Learning in Sybil Settings","volume-title":"23rd Int. Symposium on Research in Attacks, Intrusions and Defenses (RAID 2020)","author":"Fung","year":"2020"},{"key":"2023031708554200700_","first-page":"583","article-title":"Scaling Distributed Machine Learning with the Parameter Server","volume-title":"11th USENIX Symposium on Operating Systems Design and Implementation (OSDI 14)","author":"Li","year":"2014"},{"key":"2023031708554200700_","doi-asserted-by":"crossref","first-page":"480","DOI":"10.1007\/978-3-030-58951-6_24","article-title":"Data Poisoning Attacks Against Federated Learning Systems","volume-title":"Computer Security (ESORICS 2020): 25th European Symposium on Research in Computer Security, ESORICS 2020, Part I","author":"Tolpegin","year":"2020"},{"key":"2023031708554200700_","article-title":"On Large-Batch Training for Deep Learning: Generalization Gap and Sharp Minima","volume-title":"Int. Conf. Learning Representations 2017","author":"Keskar","year":"2017"},{"key":"2023031708554200700_","first-page":"4348","article-title":"Taming the Noisy Gradient: Train Deep Neural Networks with Small Batch Sizes","volume-title":"Proc. Twenty-Eighth Int. Joint Conf. Artificial Intelligence (IJCAI-19)","author":"Zhang","year":"2019"},{"key":"2023031708554200700_","first-page":"5674","article-title":"The Hidden Vulnerability of Distributed Learning in Byzantium","volume-title":"35th Int. Conf. Machine Learning (ICML 2018)","author":"Mhamdi","year":"2018"},{"key":"2023031708554200700_","first-page":"4387","article-title":"The Non-IID Data Quagmire of Decentralized Machine Learning","volume-title":"Proc. 37th Int. Conf. Machine Learning","author":"Hsieh","year":"2020"},{"key":"2023031708554200700_","article-title":"On the Convergence of FedAvg on Non-IID Data","volume-title":"Int. Conf. Learning Representations 2020","author":"Xiang","year":"2020"},{"key":"2023031708554200700_","article-title":"Can You Really Backdoor Federated Learning?","volume-title":"2nd Int. Workshop on Federated Learning for Data Privacy and Confidentiality in NeurIPS 2019","author":"Sun","year":"2019"},{"key":"2023031708554200700_","doi-asserted-by":"crossref","DOI":"10.14722\/ndss.2021.24498","article-title":"Manipulating the Byzantine: Optimizing Model Poisoning Attacks and Defenses for Federated Learning","volume-title":"Proc. 2021 Network and Distributed System Security Symposium","author":"Shejwalkar","year":"2021"},{"key":"2023031708554200700_","first-page":"1532","article-title":"Distributions of angles in random packing on spheres","volume":"14","author":"Cai","year":"2013","journal-title":"J. Mach. Learn. Res."},{"key":"2023031708554200700_","first-page":"4605","article-title":"Learning Spread-Out Local Feature Descriptors","volume-title":"Proc. IEEE Int. Conf. Computer Vision","author":"Zhang","year":"2017"}],"container-title":["The Computer Journal"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/academic.oup.com\/comjnl\/article-pdf\/66\/3\/711\/49530973\/bxab192.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/academic.oup.com\/comjnl\/article-pdf\/66\/3\/711\/49530973\/bxab192.pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2023,11,12]],"date-time":"2023-11-12T20:49:35Z","timestamp":1699822175000},"score":1,"resource":{"primary":{"URL":"https:\/\/academic.oup.com\/comjnl\/article\/66\/3\/711\/6469154"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2021,12,18]]},"references-count":45,"journal-issue":{"issue":"3","published-online":{"date-parts":[[2021,12,18]]},"published-print":{"date-parts":[[2023,3,15]]}},"URL":"https:\/\/doi.org\/10.1093\/comjnl\/bxab192","relation":{},"ISSN":["0010-4620","1460-2067"],"issn-type":[{"value":"0010-4620","type":"print"},{"value":"1460-2067","type":"electronic"}],"subject":[],"published-other":{"date-parts":[[2023,3]]},"published":{"date-parts":[[2021,12,18]]}}}