{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,10,16]],"date-time":"2025-10-16T06:31:16Z","timestamp":1760596276980,"version":"3.41.0"},"reference-count":37,"publisher":"Association for Computing Machinery (ACM)","issue":"1","license":[{"start":{"date-parts":[[2016,7,20]],"date-time":"2016-07-20T00:00:00Z","timestamp":1468972800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"DOI":"10.13039\/501100002858","name":"China Postdoctoral Science Foundation","doi-asserted-by":"crossref","award":["2014M552344 and 2015M580786"],"award-info":[{"award-number":["2014M552344 and 2015M580786"]}],"id":[{"id":"10.13039\/501100002858","id-type":"DOI","asserted-by":"crossref"}]},{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"crossref","award":["61403062 and 61433014"],"award-info":[{"award-number":["61403062 and 61433014"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]},{"DOI":"10.13039\/501100012226","name":"Fundamental Research Funds for the Central Universities","doi-asserted-by":"crossref","award":["ZYGX2014J053"],"award-info":[{"award-number":["ZYGX2014J053"]}],"id":[{"id":"10.13039\/501100012226","id-type":"DOI","asserted-by":"crossref"}]},{"name":"Science-Technology Foundation for Young Scientist of SiChuan Province","award":["2016JQ0007"],"award-info":[{"award-number":["2016JQ0007"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Knowl. Discov. Data"],"published-print":{"date-parts":[[2017,2,28]]},"abstract":"<jats:p>Clustering very large datasets while preserving cluster quality remains a challenging data-mining task to date. In this paper, we propose an effective scalable clustering algorithm for large datasets that builds upon the concept of synchronization. Inherited from the powerful concept of synchronization, the proposed algorithm, CIPA (Clustering by Iterative Partitioning and Point Attractor Representations), is capable of handling very large datasets by iteratively partitioning them into thousands of subsets and clustering each subset separately. Using dynamic clustering by synchronization, each subset is then represented by a set of point attractors and outliers. Finally, CIPA identifies the cluster structure of the original dataset by clustering the newly generated dataset consisting of points attractors and outliers from all subsets. We demonstrate that our new scalable clustering approach has several attractive benefits: (a) CIPA faithfully captures the cluster structure of the original data by performing clustering on each separate data iteratively instead of using any sampling or statistical summarization technique. (b) It allows clustering very large datasets efficiently with high cluster quality. (c) CIPA is parallelizable and also suitable for distributed data. Extensive experiments demonstrate the effectiveness and efficiency of our approach.<\/jats:p>","DOI":"10.1145\/2934688","type":"journal-article","created":{"date-parts":[[2016,7,21]],"date-time":"2016-07-21T15:13:24Z","timestamp":1469114004000},"page":"1-23","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":18,"title":["Scalable Clustering by Iterative Partitioning and Point Attractor Representation"],"prefix":"10.1145","volume":"11","author":[{"given":"Junming","family":"Shao","sequence":"first","affiliation":[{"name":"University of Electronic Science and Technology of China, Chengdu, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Qinli","family":"Yang","sequence":"additional","affiliation":[{"name":"University of Electronic Science and Technology of China, Chengdu, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Hoang-Vu","family":"Dang","sequence":"additional","affiliation":[{"name":"Johannes Gutenberg Universit\u00e4t Mainz, Mainz, Germany"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Bertil","family":"Schmidt","sequence":"additional","affiliation":[{"name":"Johannes Gutenberg Universit\u00e4t Mainz, Mainz, Germany"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Stefan","family":"Kramer","sequence":"additional","affiliation":[{"name":"Johannes Gutenberg Universit\u00e4t Mainz, Mainz, Germany"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2016,7,20]]},"reference":[{"doi-asserted-by":"publisher","key":"e_1_2_1_1_1","DOI":"10.1103\/RevModPhys.77.137"},{"doi-asserted-by":"publisher","key":"e_1_2_1_2_1","DOI":"10.1007\/978-3-642-40047-6_83"},{"key":"e_1_2_1_3_1","volume-title":"Sevcik","author":"Andritsos Periklis","year":"2004","unstructured":"Periklis Andritsos , Panayiotis Tsaparas , Ren\u00e9e J. Miller , and Kenneth C . Sevcik . 2004 . Limbo : Scalable clustering of categorical data. In Advances in Database Technology-EDBT 2004. Springer , 123--146. Periklis Andritsos, Panayiotis Tsaparas, Ren\u00e9e J. Miller, and Kenneth C. Sevcik. 2004. Limbo: Scalable clustering of categorical data. In Advances in Database Technology-EDBT 2004. Springer, 123--146."},{"doi-asserted-by":"publisher","key":"e_1_2_1_4_1","DOI":"10.1016\/j.physrep.2008.09.002"},{"doi-asserted-by":"publisher","key":"e_1_2_1_5_1","DOI":"10.14778\/2180912.2180915"},{"doi-asserted-by":"publisher","key":"e_1_2_1_6_1","DOI":"10.1007\/s10618-006-0040-z"},{"doi-asserted-by":"publisher","key":"e_1_2_1_7_1","DOI":"10.1145\/1645953.1646038"},{"doi-asserted-by":"publisher","key":"e_1_2_1_8_1","DOI":"10.1145\/1835804.1835879"},{"unstructured":"Paul S. Bradley Usama M. Fayyad Cory Reina and others. 1998. Scaling clustering algorithms to large databases. In KDD. ACM 9--15.  Paul S. Bradley Usama M. Fayyad Cory Reina and others. 1998. Scaling clustering algorithms to large databases. In KDD. ACM 9--15.","key":"e_1_2_1_9_1"},{"doi-asserted-by":"publisher","key":"e_1_2_1_10_1","DOI":"10.1145\/376284.375672"},{"doi-asserted-by":"publisher","key":"e_1_2_1_11_1","DOI":"10.1007\/11775300_32"},{"key":"e_1_2_1_12_1","volume-title":"Nonlinear Phenom. 273, 19","author":"Smet Dirk Aeyels Filip De","year":"2008","unstructured":"Filip De Smet Dirk Aeyels . 2008. A mathematical model for the dynamics of clustering. Phys. D , Nonlinear Phenom. 273, 19 ( 2008 ), 2517C2530. Filip De Smet Dirk Aeyels. 2008. A mathematical model for the dynamics of clustering. Phys. D, Nonlinear Phenom. 273, 19 (2008), 2517C2530."},{"volume-title":"Pervasive Computing and the Networked World","author":"Fu Xiufen","unstructured":"Xiufen Fu , Yaguang Wang , Yanna Ge , Peiwen Chen , and Shaohua Teng . 2014. Research and application of DBSCAN algorithm based on Hadoop platform . In Pervasive Computing and the Networked World . Springer , 73--87. Xiufen Fu, Yaguang Wang, Yanna Ge, Peiwen Chen, and Shaohua Teng. 2014. Research and application of DBSCAN algorithm based on Hadoop platform. In Pervasive Computing and the Networked World. Springer, 73--87.","key":"e_1_2_1_13_1"},{"doi-asserted-by":"publisher","key":"e_1_2_1_14_1","DOI":"10.1145\/276304.276312"},{"doi-asserted-by":"publisher","key":"e_1_2_1_15_1","DOI":"10.1109\/TFUZZ.2012.2201485"},{"doi-asserted-by":"publisher","key":"e_1_2_1_16_1","DOI":"10.1063\/1.4747710"},{"doi-asserted-by":"publisher","key":"e_1_2_1_17_1","DOI":"10.1016\/j.knosys.2012.11.015"},{"doi-asserted-by":"publisher","key":"e_1_2_1_18_1","DOI":"10.1109\/TPAMI.2002.1017616"},{"key":"e_1_2_1_19_1","volume-title":"Rousseeuw","author":"Kaufman Leonard","year":"2009","unstructured":"Leonard Kaufman and Peter J . Rousseeuw . 2009 . Finding Groups in Data : An Introduction to Cluster Analysis. Vol. 344 . John Wiley & Sons . Leonard Kaufman and Peter J. Rousseeuw. 2009. Finding Groups in Data: An Introduction to Cluster Analysis. Vol. 344. John Wiley & Sons."},{"key":"e_1_2_1_20_1","volume-title":"Cheol Soo Bae, and Hong Joon Tcha","author":"Kim Chang Sik","year":"2008","unstructured":"Chang Sik Kim , Cheol Soo Bae, and Hong Joon Tcha . 2008 . A phase synchronization clustering algorithm for identifying interesting groups of genes from cell cycle expression data. BMC Bioinformat . 9, 56 (2008). Chang Sik Kim, Cheol Soo Bae, and Hong Joon Tcha. 2008. A phase synchronization clustering algorithm for identifying interesting groups of genes from cell cycle expression data. BMC Bioinformat. 9, 56 (2008)."},{"doi-asserted-by":"publisher","key":"e_1_2_1_21_1","DOI":"10.1016\/j.is.2013.11.002"},{"doi-asserted-by":"publisher","key":"e_1_2_1_22_1","DOI":"10.5555\/1876037.1876051"},{"key":"e_1_2_1_23_1","volume-title":"Int. J. Mach. Learn. Cybern.","author":"Ludwig Simone A.","year":"2015","unstructured":"Simone A. Ludwig . 2015. MapReduce-based fuzzy c-means clustering algorithm: Implementation and scalability . Int. J. Mach. Learn. Cybern. ( 2015 ), 1--12. Simone A. Ludwig. 2015. MapReduce-based fuzzy c-means clustering algorithm: Implementation and scalability. Int. J. Mach. Learn. Cybern. (2015), 1--12."},{"volume-title":"IEEE International Conference on Data Mining. IEEE, 290--297","author":"Boriana","unstructured":"Boriana L. Milenova and Marcos M. Campos. 2002. O-cluster. 2002. Scalable clustering of large high dimensional data sets . In IEEE International Conference on Data Mining. IEEE, 290--297 . Boriana L. Milenova and Marcos M. Campos. 2002. O-cluster. 2002. Scalable clustering of large high dimensional data sets. In IEEE International Conference on Data Mining. IEEE, 290--297.","key":"e_1_2_1_24_1"},{"doi-asserted-by":"publisher","key":"e_1_2_1_25_1","DOI":"10.1145\/1099554.1099592"},{"doi-asserted-by":"publisher","key":"e_1_2_1_26_1","DOI":"10.1080\/01621459.1971.10482356"},{"doi-asserted-by":"publisher","key":"e_1_2_1_27_1","DOI":"10.1145\/2623330.2623609"},{"volume-title":"Machine Learning and Knowledge Discovery in Databases","author":"Shao Junming","unstructured":"Junming Shao , Christian B\u00f6hm , Qinli Yang , and Claudia Plant . 2010. Synchronization based outlier detection . In Machine Learning and Knowledge Discovery in Databases . Springer , 245--260. Junming Shao, Christian B\u00f6hm, Qinli Yang, and Claudia Plant. 2010. Synchronization based outlier detection. In Machine Learning and Knowledge Discovery in Databases. Springer, 245--260.","key":"e_1_2_1_28_1"},{"doi-asserted-by":"publisher","key":"e_1_2_1_29_1","DOI":"10.1109\/TKDE.2012.32"},{"volume-title":"Advances in Knowledge Discovery and Data Mining","author":"Shao Junming","unstructured":"Junming Shao , Xiao He , Qinli Yang , Claudia Plant , and Christian B\u00f6hm . 2013b. Robust synchronization-based graph clustering . In Advances in Knowledge Discovery and Data Mining . Springer , 249--260. Junming Shao, Xiao He, Qinli Yang, Claudia Plant, and Christian B\u00f6hm. 2013b. Robust synchronization-based graph clustering. In Advances in Knowledge Discovery and Data Mining. Springer, 249--260.","key":"e_1_2_1_30_1"},{"doi-asserted-by":"publisher","key":"e_1_2_1_31_1","DOI":"10.1109\/ICDM.2011.50"},{"doi-asserted-by":"publisher","key":"e_1_2_1_32_1","DOI":"10.1162\/153244303321897735"},{"doi-asserted-by":"publisher","key":"e_1_2_1_33_1","DOI":"10.1109\/HiPC.2011.6152713"},{"doi-asserted-by":"publisher","key":"e_1_2_1_34_1","DOI":"10.1109\/TKDE.2013.178"},{"doi-asserted-by":"publisher","key":"e_1_2_1_35_1","DOI":"10.1145\/233269.233324"},{"doi-asserted-by":"publisher","key":"e_1_2_1_36_1","DOI":"10.1007\/978-3-642-10665-1_71"},{"doi-asserted-by":"publisher","key":"e_1_2_1_37_1","DOI":"10.1145\/584792.584877"}],"container-title":["ACM Transactions on Knowledge Discovery from Data"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/2934688","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/2934688","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T03:39:47Z","timestamp":1750217987000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/2934688"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2016,7,20]]},"references-count":37,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2017,2,28]]}},"alternative-id":["10.1145\/2934688"],"URL":"https:\/\/doi.org\/10.1145\/2934688","relation":{},"ISSN":["1556-4681","1556-472X"],"issn-type":[{"type":"print","value":"1556-4681"},{"type":"electronic","value":"1556-472X"}],"subject":[],"published":{"date-parts":[[2016,7,20]]},"assertion":[{"value":"2015-01-01","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2016-04-01","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2016-07-20","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}