{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,4,7]],"date-time":"2026-04-07T03:46:25Z","timestamp":1775533585133,"version":"3.50.1"},"reference-count":39,"publisher":"MDPI AG","issue":"3","license":[{"start":{"date-parts":[[2020,3,11]],"date-time":"2020-03-11T00:00:00Z","timestamp":1583884800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Entropy"],"abstract":"<jats:p>Federated learning is a decentralized topology of deep learning, that trains a shared model through data distributed among each client (like mobile phones, wearable devices), in order to ensure data privacy by avoiding raw data exposed in data center (server). After each client computes a new model parameter by stochastic gradient descent (SGD) based on their own local data, these locally-computed parameters will be aggregated to generate an updated global model. Many current state-of-the-art studies aggregate different client-computed parameters by averaging them, but none theoretically explains why averaging parameters is a good approach. In this paper, we treat each client computed parameter as a random vector because of the stochastic properties of SGD, and estimate mutual information between two client computed parameters at different training phases using two methods in two learning tasks. The results confirm the correlation between different clients and show an increasing trend of mutual information with training iteration. However, when we further compute the distance between client computed parameters, we find that parameters are getting more correlated while not getting closer. This phenomenon suggests that averaging parameters may not be the optimum way of aggregating trained parameters.<\/jats:p>","DOI":"10.3390\/e22030314","type":"journal-article","created":{"date-parts":[[2020,3,12]],"date-time":"2020-03-12T04:13:57Z","timestamp":1583986437000},"page":"314","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":41,"title":["Averaging Is Probably Not the Optimum Way of Aggregating Parameters in Federated Learning"],"prefix":"10.3390","volume":"22","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-7313-4689","authenticated-orcid":false,"given":"Peng","family":"Xiao","sequence":"first","affiliation":[{"name":"Department of Computer Science and Technology, Tongji University, Shanghai 201804, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Samuel","family":"Cheng","sequence":"additional","affiliation":[{"name":"The School of Electrical and Computer Engineering, University of Oklahoma, Tulsa, OK 73019, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-1075-2420","authenticated-orcid":false,"given":"Vladimir","family":"Stankovic","sequence":"additional","affiliation":[{"name":"Department of Electronic and Electrical engineering, University of Strathclyde, Glasgow G1 1XW, UK"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Dejan","family":"Vukobratovic","sequence":"additional","affiliation":[{"name":"Faculty of Technical Sciences, University of Novi Sad, 21000 Novi Sad, Serbia"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2020,3,11]]},"reference":[{"key":"ref_1","first-page":"1","article-title":"Smartphone ownership and internet usage continues to climb in emerging economies","volume":"22","author":"Poushter","year":"2016","journal-title":"Pew Res. Center"},{"key":"ref_2","doi-asserted-by":"crossref","first-page":"4","DOI":"10.4258\/hir.2017.23.1.4","article-title":"Wearable devices in medical internet of things: Scientific research and commercially available devices","volume":"23","author":"Haghi","year":"2017","journal-title":"Healthc. Inform. Res."},{"key":"ref_3","doi-asserted-by":"crossref","unstructured":"Taylor, R., Baron, D., and Schmidt, D. (2015, January 21\u201323). The world in 2025-predictions for the next ten years. Proceedings of the 2015 10th International Microsystems, Packaging, Assembly and Circuits Technology Conference (IMPACT), Taipei, Taiwan.","DOI":"10.1109\/IMPACT.2015.7365193"},{"key":"ref_4","doi-asserted-by":"crossref","unstructured":"Bonomi, F., Milito, R., Zhu, J., and Addepalli, S. (2012, January 17). Fog computing and its role in the internet of things. Proceedings of the First Edition of the MCC Workshop on Mobile Cloud Computing, Helsinki, Finland.","DOI":"10.1145\/2342509.2342513"},{"key":"ref_5","doi-asserted-by":"crossref","first-page":"37","DOI":"10.1145\/2831347.2831354","article-title":"Edge-centric computing: Vision and challenges","volume":"45","author":"Montresor","year":"2015","journal-title":"ACM SIGCOMM Comput. Commun. Rev."},{"key":"ref_6","unstructured":"House, W. (2012). Consumer Data Privacy in a Networked World: A Framework for Protecting Privacy and Promoting Innovation in the Global Digital Economy, White House."},{"key":"ref_7","unstructured":"McMahan, H.B., Moore, E., Ramage, D., and Hampson, S. (2016). Communication-efficient learning of deep networks from decentralized data. arXiv."},{"key":"ref_8","unstructured":"Dean, J., Corrado, G., Monga, R., Chen, K., Devin, M., Mao, M., Ranzato, M., Senior, A., Tucker, P., and Yang, K. (2012, January 3\u20136). Large scale distributed deep networks. Proceedings of the Advances in Neural Information Processing Systems, Lake Tahoe, NV, USA."},{"key":"ref_9","unstructured":"Hard, A., Rao, K., Mathews, R., Ramaswamy, S., Beaufays, F., Augenstein, S., Eichner, H., Kiddon, C., and Ramage, D. (2018). Federated learning for mobile keyboard prediction. arXiv."},{"key":"ref_10","unstructured":"Ramaswamy, S., Mathews, R., Rao, K., and Beaufays, F. (2019). Federated learning for emoji prediction in a mobile keyboard. arXiv."},{"key":"ref_11","unstructured":"Huang, L., Yin, Y., Fu, Z., Zhang, S., Deng, H., and Liu, D. (2018). LoAdaBoost: Loss-Based AdaBoost Federated Machine Learning on medical Data. arXiv."},{"key":"ref_12","doi-asserted-by":"crossref","unstructured":"Samarakoon, S., Bennis, M., Saad, W., and Debbah, M. (2018, January 9\u201313). Federated learning for ultra-reliable low-latency V2V communications. Proceedings of the 2018 IEEE Global Communications Conference (GLOBECOM), Abu Dhabi, UAE.","DOI":"10.1109\/GLOCOM.2018.8647927"},{"key":"ref_13","unstructured":"Bottou, L. (2010, January 22\u201327). Large-scale machine learning with stochastic gradient descent. Proceedings of the 19th International Conference on Computational Statistics, Paris, France."},{"key":"ref_14","doi-asserted-by":"crossref","unstructured":"Sattler, F., Wiedemann, S., M\u00fcller, K.R., and Samek, W. (2019). Robust and communication-efficient federated learning from non-iid data. IEEE Trans. Neural Netw. Learn. Syst.","DOI":"10.1109\/TNNLS.2019.2944481"},{"key":"ref_15","unstructured":"McMahan, H.B., Ramage, D., Talwar, K., and Zhang, L. (May, January 30). Learning differentially private recurrent language models. Proceedings of the International Conference on Learning Representations, Vancouver, BC, Canada."},{"key":"ref_16","unstructured":"Agarwal, N., Suresh, A.T., Yu, F.X.X., Kumar, S., and McMahan, B. (2018, January 3\u20138). cpSGD: Communication-efficient and differentially- private distributed SGD. Proceedings of the Advances in Neural Information Processing Systems, Montreal, QC, Canada."},{"key":"ref_17","unstructured":"He, L., Bian, A., and Jaggi, M. (2018, January 3\u20138). Cola: Decentralized linear learning. Proceedings of the Advances in Neural Information Processing Systems, Montreal, QC, Canada."},{"key":"ref_18","unstructured":"Woodworth, B.E., Wang, J., Smith, A., McMahan, B., and Srebro, N. (2018, January 3\u20138). Graph oracle models, lower bounds, and gaps for parallel stochastic optimization. Proceedings of the Advances in Neural Information Processing Systems, Montreal, QC, Canada."},{"key":"ref_19","unstructured":"Eichner, H., Koren, T., Mcmahan, B., Srebro, N., and Talwar, K. (2019, January 10\u201315). Semi-Cyclic Stochastic Gradient Descent. Proceedings of the International Conference on Machine Learning, Long Beach, CA, USA."},{"key":"ref_20","unstructured":"Cover, T.M., and Thomas, J.A. (2012). Elements of Information Theory, John Wiley & Sons."},{"key":"ref_21","doi-asserted-by":"crossref","first-page":"066138","DOI":"10.1103\/PhysRevE.69.066138","article-title":"Estimating mutual information","volume":"69","author":"Kraskov","year":"2004","journal-title":"Phys. Rev. E"},{"key":"ref_22","doi-asserted-by":"crossref","unstructured":"Ross, B.C. (2014). Mutual information between discrete and continuous data sets. PLoS ONE, 9.","DOI":"10.1371\/journal.pone.0087357"},{"key":"ref_23","unstructured":"Krizhevsky, A., Sutskever, I., and Hinton, G.E. (2012, January 3\u20136). Imagenet classification with deep convolutional neural networks. Proceedings of the Advances in Neural Information Processing Systems, Lake Tahoe, NV, USA."},{"key":"ref_24","doi-asserted-by":"crossref","unstructured":"Mikolov, T., Karafi\u00e1t, M., Burget, L., \u010cernock\u1ef3, J., and Khudanpur, S. (2010, January 26\u201330). Recurrent neural network based language model. Proceedings of the Eleventh Annual Conference of the International Speech Communication Association, Chiba, Japan.","DOI":"10.21437\/Interspeech.2010-343"},{"key":"ref_25","doi-asserted-by":"crossref","first-page":"227","DOI":"10.1016\/0146-664X(80)90054-4","article-title":"Euclidean distance mapping","volume":"14","author":"Danielsson","year":"1980","journal-title":"Comput. Gr. Image Process."},{"key":"ref_26","unstructured":"Krause, E.F. (1986). Taxicab Geometry: An Adventure in Non-Euclidean Geometry, Courier Corporation."},{"key":"ref_27","doi-asserted-by":"crossref","unstructured":"Cantrell, C.D. (2000). Modern Mathematical Methods for Physicists and Engineers, Cambridge University Press.","DOI":"10.1017\/9780511811487"},{"key":"ref_28","doi-asserted-by":"crossref","unstructured":"Tishby, N., and Zaslavsky, N. (May, January 26). Deep learning and the information bottleneck principle. Proceedings of the 2015 IEEE Information Theory Workshop (ITW), Jerusalem, Israel.","DOI":"10.1109\/ITW.2015.7133169"},{"key":"ref_29","unstructured":"Shwartz-Ziv, R., and Tishby, N. (2017). Opening the black box of deep neural networks via information. arXiv."},{"key":"ref_30","unstructured":"Saxe, A.M., Bansal, Y., Dapello, J., Advani, M., Kolchinsky, A., Tracey, B.D., and Cox, D.D. (May, January 30). On the information bottleneck theory of deep learning. Proceedings of the International Conference on Learning Representations, Vancouver, BC, Canada."},{"key":"ref_31","doi-asserted-by":"crossref","unstructured":"Noshad, M., Zeng, Y., and Hero, A.O. (2019, January 12\u201317). Scalable mutual information estimation using dependence graphs. Proceedings of the ICASSP 2019-2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Brighton, UK.","DOI":"10.1109\/ICASSP.2019.8683351"},{"key":"ref_32","doi-asserted-by":"crossref","unstructured":"Kolchinsky, A., Tracey, B.D., and Wolpert, D.H. (2019). Nonlinear information bottleneck. Entropy, 21.","DOI":"10.3390\/e21121181"},{"key":"ref_33","unstructured":"Wickstr\u00f8m, K., L\u00f8kse, S., Kampffmeyer, M., Yu, S., Principe, J., and Jenssen, R. (2019). Information Plane Analysis of Deep Neural Networks via Matrix-Based Renyi\u2019s Entropy and Tensor Kernels. arXiv."},{"key":"ref_34","unstructured":"Adilova, L., Rosenzweig, J., and Kamp, M. (2019). Information-Theoretic Perspective of Federated Learning. arXiv."},{"key":"ref_35","doi-asserted-by":"crossref","first-page":"379","DOI":"10.1002\/j.1538-7305.1948.tb01338.x","article-title":"A mathematical theory of communication","volume":"27","author":"Shannon","year":"1948","journal-title":"Bell Syst. Tech. J."},{"key":"ref_36","doi-asserted-by":"crossref","first-page":"141","DOI":"10.1109\/MSP.2012.2211477","article-title":"The MNIST database of handwritten digit images for machine learning research [best of the web]","volume":"29","author":"Deng","year":"2012","journal-title":"IEEE Signal Process. Mag."},{"key":"ref_37","unstructured":"Kingma, D.P., and Ba, J. (2014). Adam: A method for stochastic optimization. arXiv."},{"key":"ref_38","unstructured":"Gal, Y., and Ghahramani, Z. (2016, January 19\u201324). Dropout as a bayesian approximation: Representing model uncertainty in deep learning. Proceedings of The International Conference on Machine Learning, New York, NY, USA."},{"key":"ref_39","doi-asserted-by":"crossref","first-page":"2278","DOI":"10.1109\/5.726791","article-title":"Gradient-based learning applied to document recognition","volume":"86","author":"LeCun","year":"1998","journal-title":"Proc. IEEE"}],"container-title":["Entropy"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1099-4300\/22\/3\/314\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,11]],"date-time":"2025-10-11T09:05:57Z","timestamp":1760173557000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1099-4300\/22\/3\/314"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2020,3,11]]},"references-count":39,"journal-issue":{"issue":"3","published-online":{"date-parts":[[2020,3]]}},"alternative-id":["e22030314"],"URL":"https:\/\/doi.org\/10.3390\/e22030314","relation":{},"ISSN":["1099-4300"],"issn-type":[{"value":"1099-4300","type":"electronic"}],"subject":[],"published":{"date-parts":[[2020,3,11]]}}}