{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T04:14:44Z","timestamp":1750220084294,"version":"3.41.0"},"reference-count":50,"publisher":"Association for Computing Machinery (ACM)","issue":"6","license":[{"start":{"date-parts":[[2022,7,30]],"date-time":"2022-07-30T00:00:00Z","timestamp":1659139200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Knowl. Discov. Data"],"published-print":{"date-parts":[[2022,12,31]]},"abstract":"<jats:p>\n            In the era of big data, data are usually distributed across numerous connected computing and storage units (i.e., nodes or workers). Under such an environment, many machine learning problems can be reformulated as a\n            <jats:italic>consensus optimization<\/jats:italic>\n            problem, which consists of one objective and constraint terms splitting into N parts (each corresponds to a node). Such a problem can be solved efficiently in a distributed manner via Alternating Direction Method of Multipliers (\n            <jats:sans-serif>ADMM<\/jats:sans-serif>\n            ). However, existing consensus optimization frameworks assume that every node has the same\n            <jats:italic>quality of information (QoI)<\/jats:italic>\n            , i.e., the data from all the nodes are equally informative for the estimation of global model parameters. As a consequence, they may lead to inaccurate estimates in the presence of nodes with low QoI. To overcome this challenge, in this article, we propose a novel consensus optimization framework for distributed machine-learning that incorporates the crucial metric, QoI. Theoretically, we prove that the convergence rate of the proposed framework is linear to the number of iterations, but has a tighter upper bound compared with\n            <jats:sans-serif>ADMM<\/jats:sans-serif>\n            . Experimentally, we show that the proposed framework is more efficient and effective than existing\n            <jats:sans-serif>ADMM<\/jats:sans-serif>\n            -based solutions on both synthetic and real-world datasets due to its faster convergence rate and higher accuracy.\n          <\/jats:p>","DOI":"10.1145\/3522591","type":"journal-article","created":{"date-parts":[[2022,3,15]],"date-time":"2022-03-15T21:30:32Z","timestamp":1647379832000},"page":"1-28","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":0,"title":["Toward Quality of Information Aware Distributed Machine Learning"],"prefix":"10.1145","volume":"16","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-6981-8842","authenticated-orcid":false,"given":"Houping","family":"Xiao","sequence":"first","affiliation":[{"name":"Georgia State University, Atlanta, GA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Shiyu","family":"Wang","sequence":"additional","affiliation":[{"name":"University of Geogria, Athen"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2022,7,30]]},"reference":[{"key":"e_1_3_3_2_2","unstructured":"Naman Agarwal Ananda Theertha Suresh Felix Yu Sanjiv Kumar and H. Brendan McMahan. 2018. cpSGD: Communication-efficient and differentially-private distributed SGD. In Proceedings of the Advances in Neural Information Processing Systems . 7575\u20137586."},{"key":"e_1_3_3_3_2","first-page":"5973","volume-title":"Proceedings of the Advances in Neural Information Processing Systems","author":"Alistarh Dan","year":"2018","unstructured":"Dan Alistarh, Torsten Hoefler, Mikael Johansson, Nikola Konstantinov, Sarit Khirirat, and C\u00e9dric Renggli. 2018. The convergence of sparsified gradient methods. In Proceedings of the Advances in Neural Information Processing Systems. 5973\u20135983."},{"key":"e_1_3_3_4_2","first-page":"436","volume-title":"Proceedings of the International Conference on Machine Learning","author":"Aviv Rotem Zamir","year":"2021","unstructured":"Rotem Zamir Aviv, Ido Hakimi, Assaf Schuster, and Kfir Yehuda Levy. 2021. Asynchronous distributed learning: Adapting to gradient delays without prior knowledge. In Proceedings of the International Conference on Machine Learning. PMLR, 436\u2013445."},{"key":"e_1_3_3_5_2","unstructured":"Debraj Basu Deepesh Data Can Karakus and Suhas Diggavi. 2019. Qsparse-local-SGD: Distributed SGD with quantization sparsification and local computations. In Proceedings of the Advances in Neural Information Processing Systems . 14695\u201314706."},{"key":"e_1_3_3_6_2","doi-asserted-by":"publisher","DOI":"10.5555\/993483"},{"key":"e_1_3_3_7_2","doi-asserted-by":"publisher","DOI":"10.1561\/2200000016"},{"key":"e_1_3_3_8_2","doi-asserted-by":"publisher","DOI":"10.5555\/3122009.3122055"},{"key":"e_1_3_3_9_2","doi-asserted-by":"publisher","DOI":"10.1007\/s10107-014-0826-5"},{"key":"e_1_3_3_10_2","article-title":"Gossipgrad: Scalable deep learning using gossip communication based asynchronous gradient descent","author":"Daily Jeff","year":"2018","unstructured":"Jeff Daily, Abhinav Vishnu, Charles Siegel, Thomas Warfel, and Vinay Amatya. 2018. Gossipgrad: Scalable deep learning using gossip communication based asynchronous gradient descent. arXiv:1803.05880. Retrieved from https:\/\/arxiv.org\/abs\/1803.05880.","journal-title":"arXiv:1803.05880"},{"key":"e_1_3_3_11_2","first-page":"1","article-title":"Parallel multi-block ADMM with O (1\/k) convergence","author":"Deng Wei","year":"2014","unstructured":"Wei Deng, Ming-Jun Lai, Zhimin Peng, and Wotao Yin. 2014. Parallel multi-block ADMM with O (1\/k) convergence. Journal of Scientific Computing 71, 2 (2014), 1\u201325.","journal-title":"Journal of Scientific Computing"},{"key":"e_1_3_3_12_2","doi-asserted-by":"publisher","DOI":"10.1007\/BF01581204"},{"key":"e_1_3_3_13_2","doi-asserted-by":"publisher","DOI":"10.1109\/TAC.2016.2525015"},{"key":"e_1_3_3_14_2","first-page":"2197","volume-title":"Proceedings of the International Conference on Artificial Intelligence and Statistics","author":"Gandikota Venkata","year":"2021","unstructured":"Venkata Gandikota, Daniel Kane, Raj Kumar Maity, and Arya Mazumdar. 2021. vqsgd: Vector quantized stochastic gradient descent. In Proceedings of the International Conference on Artificial Intelligence and Statistics. PMLR, 2197\u20132205."},{"key":"e_1_3_3_15_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2014.2362759"},{"key":"e_1_3_3_16_2","article-title":"Stochastic distributed learning with gradient quantization and variance reduction","author":"Horv\u00e1th Samuel","year":"2019","unstructured":"Samuel Horv\u00e1th, Dmitry Kovalev, Konstantin Mishchenko, Sebastian Stich, and Peter Richt\u00e1rik. 2019. Stochastic distributed learning with gradient quantization and variance reduction. arXiv:1904.05115. Retrieved from https:\/\/arxiv.org\/abs\/1904.05115.","journal-title":"arXiv:1904.05115"},{"key":"e_1_3_3_17_2","doi-asserted-by":"publisher","DOI":"10.1145\/2770876"},{"key":"e_1_3_3_18_2","doi-asserted-by":"publisher","DOI":"10.1145\/3417337"},{"key":"e_1_3_3_19_2","doi-asserted-by":"publisher","DOI":"10.1109\/TIFS.2019.2931068"},{"key":"e_1_3_3_20_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPDS.2017.2712630"},{"key":"e_1_3_3_21_2","doi-asserted-by":"publisher","DOI":"10.1145\/2746285.2746310"},{"key":"e_1_3_3_22_2","first-page":"1905","volume-title":"Proceedings of the Uncertainty in Artificial Intelligence","author":"Kasiviswanathan Shiva Prasad","year":"2021","unstructured":"Shiva Prasad Kasiviswanathan. 2021. SGD with low-dimensional gradients with applications to private and distributed learning. In Proceedings of the Uncertainty in Artificial Intelligence. PMLR, 1905\u20131915."},{"key":"e_1_3_3_23_2","doi-asserted-by":"publisher","DOI":"10.23919\/WiOpt52861.2021.9589802"},{"key":"e_1_3_3_24_2","first-page":"3043","volume-title":"Proceedings of the International Conference on Machine Learning","author":"Lian Xiangru","year":"2018","unstructured":"Xiangru Lian, Wei Zhang, Ce Zhang, and Ji Liu. 2018. Asynchronous decentralized parallel stochastic gradient descent. In Proceedings of the International Conference on Machine Learning. PMLR, 3043\u20133052."},{"key":"e_1_3_3_25_2","doi-asserted-by":"publisher","DOI":"10.1007\/s00365-017-9379-1"},{"key":"e_1_3_3_26_2","doi-asserted-by":"publisher","DOI":"10.1137\/140971178"},{"key":"e_1_3_3_27_2","doi-asserted-by":"publisher","DOI":"10.1145\/3230668"},{"key":"e_1_3_3_28_2","first-page":"1273","volume-title":"Proceedings of the Artificial Intelligence and Statistics","author":"McMahan Brendan","year":"2017","unstructured":"Brendan McMahan, Eider Moore, Daniel Ramage, Seth Hampson, and Blaise Aguera y Arcas. 2017. Communication-efficient learning of deep networks from decentralized data. In Proceedings of the Artificial Intelligence and Statistics. PMLR, 1273\u20131282."},{"key":"e_1_3_3_29_2","unstructured":"Brendan McMahan and Matthew Streeter. 2014. Delay-tolerant algorithms for asynchronous distributed online learning. In Proceedings of the 27th International Conference on Neural Information Processing Systems - Volume 2 ."},{"key":"e_1_3_3_30_2","first-page":"343","volume-title":"Proceedings of the International Conference on Machine Learning","author":"Nishihara Robert","year":"2015","unstructured":"Robert Nishihara, Laurent Lessard, Benjamin Recht, Andrew Packard, and Michael I. Jordan. 2015. A general analysis of the convergence of ADMM. In Proceedings of the International Conference on Machine Learning. 343\u2013352."},{"key":"e_1_3_3_31_2","first-page":"1310","volume-title":"Proceedings of the International Conference on Machine Learning","author":"Pascanu Razvan","year":"2013","unstructured":"Razvan Pascanu, Tomas Mikolov, and Yoshua Bengio. 2013. On the difficulty of training recurrent neural networks. In Proceedings of the International Conference on Machine Learning. PMLR, 1310\u20131318."},{"key":"e_1_3_3_32_2","first-page":"693","volume-title":"Proceedings of the Advances in Neural Information Processing Systems","author":"Recht Benjamin","year":"2011","unstructured":"Benjamin Recht, Christopher Re, Stephen Wright, and Feng Niu. 2011. Hogwild: A lock-free approach to parallelizing stochastic gradient descent. In Proceedings of the Advances in Neural Information Processing Systems. 693\u2013701."},{"key":"e_1_3_3_33_2","doi-asserted-by":"publisher","DOI":"10.21437\/Interspeech.2014-274"},{"key":"e_1_3_3_34_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICDCS.2019.00220"},{"key":"e_1_3_3_35_2","doi-asserted-by":"publisher","DOI":"10.1109\/TSP.2014.2304432"},{"key":"e_1_3_3_36_2","doi-asserted-by":"publisher","DOI":"10.1145\/3278607"},{"key":"e_1_3_3_37_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCI.2021.3094062"},{"key":"e_1_3_3_38_2","doi-asserted-by":"crossref","unstructured":"Christian Szegedy Wei Liu Yangqing Jia Pierre Sermanet Scott Reed Dragomir Anguelov Dumitru Erhan Vincent Vanhoucke and Andrew Rabinovich. 2015. Going deeper with convolutions. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition . 1\u20139.","DOI":"10.1109\/CVPR.2015.7298594"},{"key":"e_1_3_3_39_2","first-page":"2722","volume-title":"Proceedings of the International Conference on Machine Learning","author":"Taylor Gavin","year":"2016","unstructured":"Gavin Taylor, Ryan Burmeister, Zheng Xu, Bharat Singh, Ankit Patel, and Tom Goldstein. 2016. Training neural networks without gradients: A scalable admm approach. In Proceedings of the International Conference on Machine Learning. PMLR, 2722\u20132731."},{"key":"e_1_3_3_40_2","doi-asserted-by":"publisher","DOI":"10.1145\/3292500.3330936"},{"key":"e_1_3_3_41_2","doi-asserted-by":"publisher","DOI":"10.1145\/3451884"},{"key":"e_1_3_3_42_2","doi-asserted-by":"publisher","DOI":"10.1007\/s10915-018-0757-z"},{"key":"e_1_3_3_43_2","first-page":"5325","volume-title":"Proceedings of the International Conference on Machine Learning","author":"Wu Jiaxiang","year":"2018","unstructured":"Jiaxiang Wu, Weidong Huang, Junzhou Huang, and Tong Zhang. 2018. Error compensated quantized SGD and its applications to large-scale distributed optimization. In Proceedings of the International Conference on Machine Learning. PMLR, 5325\u20135333."},{"key":"e_1_3_3_44_2","doi-asserted-by":"publisher","DOI":"10.1145\/2939672.2939831"},{"key":"e_1_3_3_45_2","doi-asserted-by":"publisher","DOI":"10.1145\/2939672.2939816"},{"key":"e_1_3_3_46_2","doi-asserted-by":"publisher","DOI":"10.1109\/TII.2019.2937513"},{"key":"e_1_3_3_47_2","doi-asserted-by":"publisher","DOI":"10.1145\/3434768"},{"key":"e_1_3_3_48_2","first-page":"1701","volume-title":"Proceedings of the International Conference on Machine Learning","author":"Zhang Ruiliang","year":"2014","unstructured":"Ruiliang Zhang and James Kwok. 2014. Asynchronous distributed ADMM for consensus optimization. In Proceedings of the International Conference on Machine Learning. 1701\u20131709."},{"key":"e_1_3_3_49_2","doi-asserted-by":"publisher","DOI":"10.1145\/3464976"},{"key":"e_1_3_3_50_2","doi-asserted-by":"publisher","DOI":"10.23919\/WiOpt52861.2021.9589660"},{"key":"e_1_3_3_51_2","first-page":"4120","volume-title":"Proceedings of the International Conference on Machine Learning","author":"Zheng Shuxin","year":"2017","unstructured":"Shuxin Zheng, Qi Meng, Taifeng Wang, Wei Chen, Nenghai Yu, Zhi-Ming Ma, and Tie-Yan Liu. 2017. Asynchronous stochastic gradient descent with delay compensation. In Proceedings of the International Conference on Machine Learning. PMLR, 4120\u20134129."}],"container-title":["ACM Transactions on Knowledge Discovery from Data"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3522591","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3522591","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T18:09:33Z","timestamp":1750183773000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3522591"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,7,30]]},"references-count":50,"journal-issue":{"issue":"6","published-print":{"date-parts":[[2022,12,31]]}},"alternative-id":["10.1145\/3522591"],"URL":"https:\/\/doi.org\/10.1145\/3522591","relation":{},"ISSN":["1556-4681","1556-472X"],"issn-type":[{"type":"print","value":"1556-4681"},{"type":"electronic","value":"1556-472X"}],"subject":[],"published":{"date-parts":[[2022,7,30]]},"assertion":[{"value":"2021-08-01","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2022-02-01","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2022-07-30","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}