{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,2,6]],"date-time":"2026-02-06T22:51:22Z","timestamp":1770418282357,"version":"3.49.0"},"reference-count":24,"publisher":"Association for Computing Machinery (ACM)","issue":"3","license":[{"start":{"date-parts":[[2019,9,30]],"date-time":"2019-09-30T00:00:00Z","timestamp":1569801600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"DOI":"10.13039\/100000001","name":"NSF","doi-asserted-by":"publisher","award":["CCF-1525024,IIS-1633215,IIS-1546151"],"award-info":[{"award-number":["CCF-1525024,IIS-1633215,IIS-1546151"]}],"id":[{"id":"10.13039\/100000001","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Parallel Comput."],"published-print":{"date-parts":[[2019,9,30]]},"abstract":"<jats:p>Recent years have witnessed an increasing popularity of algorithm design for distributed data, largely due to the fact that massive datasets are often collected and stored in different locations. In the distributed setting, communication typically dominates the query processing time. Thus, it becomes crucial to design communication-efficient algorithms for queries on distributed data. Simultaneously, it has been widely recognized that partial optimizations, where we are allowed to disregard a small part of the data, provide us significantly better solutions. The motivation for disregarded points often arises from noise and other phenomena that are pervasive in large data scenarios.<\/jats:p>\n          <jats:p>\n            In this article, we focus on partial clustering problems,\n            <jats:italic>k<\/jats:italic>\n            -center,\n            <jats:italic>k<\/jats:italic>\n            -median, and\n            <jats:italic>k<\/jats:italic>\n            -means objectives in the distributed model, and provide algorithms with communication sublinear of the input size. As a consequence, we develop the first algorithms for the partial\n            <jats:italic>k<\/jats:italic>\n            -median and means objectives that run in subquadratic running time. We also initiate the study of distributed algorithms for clustering uncertain data, where each data point can possibly fall into multiple locations under certain probability distribution.\n          <\/jats:p>","DOI":"10.1145\/3322808","type":"journal-article","created":{"date-parts":[[2019,10,15]],"date-time":"2019-10-15T16:35:58Z","timestamp":1571157358000},"page":"1-20","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":10,"title":["Distributed Partial Clustering"],"prefix":"10.1145","volume":"6","author":[{"given":"Sudipto","family":"Guha","sequence":"first","affiliation":[{"name":"University of Pennsylvania, United States"}]},{"given":"Yi","family":"Li","sequence":"additional","affiliation":[{"name":"Nanyang Technological University, Singapore"}]},{"given":"Qin","family":"Zhang","sequence":"additional","affiliation":[{"name":"Indiana University Bloomington, United States"}]}],"member":"320","published-online":{"date-parts":[[2019,10,15]]},"reference":[{"key":"e_1_2_1_1_1","volume-title":"Data Clustering: Algorithms and Applications","author":"Aggarwal Charu C."},{"key":"e_1_2_1_2_1","volume-title":"Proceedings of the NIPS. 1995--2003","author":"Balcan Maria-Florina","year":"2013"},{"key":"e_1_2_1_3_1","doi-asserted-by":"publisher","DOI":"10.1109\/SFFCS.1999.814609"},{"key":"e_1_2_1_4_1","volume-title":"Proceedings of the SODA. 642--651","author":"Charikar Moses","year":"2001"},{"key":"e_1_2_1_5_1","volume-title":"Proceedings of the NIPS.","author":"Chen Jiecao","year":"2016"},{"key":"e_1_2_1_6_1","volume-title":"Proceedings of the SODA. 826--835","author":"Chen Ke","year":"2008"},{"key":"e_1_2_1_7_1","doi-asserted-by":"publisher","DOI":"10.1145\/2746539.2746569"},{"key":"e_1_2_1_8_1","doi-asserted-by":"publisher","DOI":"10.1145\/1376916.1376944"},{"key":"e_1_2_1_9_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICDE.2007.368962"},{"key":"e_1_2_1_10_1","volume-title":"Proceedings of the (OSDI\u201904)","author":"Dean Jeffrey","year":"2004"},{"key":"e_1_2_1_11_1","doi-asserted-by":"publisher","DOI":"10.1137\/0215052"},{"key":"e_1_2_1_12_1","doi-asserted-by":"publisher","DOI":"10.1145\/2020408.2020515"},{"key":"e_1_2_1_13_1","volume-title":"Schulman","author":"Feldman Dan","year":"2012"},{"key":"e_1_2_1_14_1","doi-asserted-by":"publisher","DOI":"10.1016\/0304-3975(85)90224-5"},{"key":"e_1_2_1_15_1","doi-asserted-by":"publisher","DOI":"10.1109\/TKDE.2003.1198387"},{"key":"e_1_2_1_16_1","doi-asserted-by":"publisher","DOI":"10.1145\/1559795.1559836"},{"key":"e_1_2_1_17_1","doi-asserted-by":"publisher","DOI":"10.1145\/2755573.2755607"},{"key":"e_1_2_1_18_1","doi-asserted-by":"publisher","DOI":"10.1145\/375827.375845"},{"key":"e_1_2_1_19_1","volume-title":"Sarpatwar","author":"Khuller Samir","year":"2014"},{"key":"e_1_2_1_20_1","volume-title":"Woodruff","author":"Liang Yingyu","year":"2014"},{"key":"e_1_2_1_21_1","volume-title":"Proceedings of the NIPS. 1063--1071","author":"Malkomes Gustavo","year":"2015"},{"key":"e_1_2_1_22_1","volume-title":"Probabilistic Databases","author":"Suciu Dan","edition":"1"},{"key":"e_1_2_1_23_1","doi-asserted-by":"publisher","DOI":"10.1006\/jcss.1997.1547"},{"key":"e_1_2_1_24_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.tcs.2015.08.017"}],"container-title":["ACM Transactions on Parallel Computing"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3322808","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3322808","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3322808","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T01:02:26Z","timestamp":1750208546000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3322808"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2019,9,30]]},"references-count":24,"journal-issue":{"issue":"3","published-print":{"date-parts":[[2019,9,30]]}},"alternative-id":["10.1145\/3322808"],"URL":"https:\/\/doi.org\/10.1145\/3322808","relation":{},"ISSN":["2329-4949","2329-4957"],"issn-type":[{"value":"2329-4949","type":"print"},{"value":"2329-4957","type":"electronic"}],"subject":[],"published":{"date-parts":[[2019,9,30]]},"assertion":[{"value":"2017-10-01","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2018-12-01","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2019-10-15","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}