{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,11,18]],"date-time":"2025-11-18T09:41:32Z","timestamp":1763458892154,"version":"3.45.0"},"reference-count":56,"publisher":"Association for Computing Machinery (ACM)","issue":"3","license":[{"start":{"date-parts":[[2018,3,21]],"date-time":"2018-03-21T00:00:00Z","timestamp":1521590400000},"content-version":"vor","delay-in-days":365,"URL":"http:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"DOI":"10.13039\/501100013285","name":"Program for Professor of Special Appointment (Eastern Scholar) at Shanghai Institutions of Higher Learning","doi-asserted-by":"crossref","id":[{"id":"10.13039\/501100013285","id-type":"DOI","asserted-by":"crossref"}]},{"DOI":"10.13039\/100000001","name":"National Science Foundation","doi-asserted-by":"publisher","award":["IIP-1444949"],"award-info":[{"award-number":["IIP-1444949"]}],"id":[{"id":"10.13039\/100000001","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Knowl. Discov. Data"],"published-print":{"date-parts":[[2017,8,31]]},"abstract":"<jats:p>\n                    Graphs are popularly used to represent objects with shared dependency relationships. To date, all existing graph clustering algorithms consider each node as a single attribute or a set of independent attributes, without realizing that content inside each node may also have complex structures. In this article, we formulate a new networked graph clustering task where a network contains a set of inter-connected (or networked) super-nodes, each of which is a single-attribute graph. The new super-node representation is applicable to many real-world applications, such as a citation network where each node denotes a paper whose content can be described as a graph, and citation relationships between papers form a networked graph (i.e., a super-graph). Networked graph clustering aims to find similar node groups, each of which contains nodes with similar content and structure information. The main challenge is to properly calculate the similarity between super-nodes for clustering. To solve the problem, we propose to characterize node similarity by integrating\n                    <jats:italic toggle=\"yes\">structure<\/jats:italic>\n                    and\n                    <jats:italic toggle=\"yes\">content<\/jats:italic>\n                    information of each super-node. To measure node content similarity, we use cosine distance by considering overlapped attributes between two super-nodes. To measure structure similarity, we propose an Attributed Random Walk Kernel (ARWK) to calculate the similarity between super-nodes. Detailed node content analysis is also included to build relationships between super-nodes with shared internal structure information, so the structure similarity can be calculated in a precise way. By integrating the structure similarity and content similarity as one matrix, the spectral clustering is used to achieve networked graph clustering. Our method enjoys sound theoretical properties, including bounded similarities and better structure similarity assessment than traditional graph clustering methods. Experiments on real-world applications demonstrate that our method significantly outperforms baseline approaches.\n                  <\/jats:p>","DOI":"10.1145\/2996197","type":"journal-article","created":{"date-parts":[[2017,3,23]],"date-time":"2017-03-23T12:19:44Z","timestamp":1490271584000},"page":"1-29","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":11,"title":["Combining Structured Node Content and Topology Information for Networked Graph Clustering"],"prefix":"10.1145","volume":"11","author":[{"given":"Ting","family":"Guo","sequence":"first","affiliation":[{"name":"University of Technology, Sydney, Australia"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Jia","family":"Wu","sequence":"additional","affiliation":[{"name":"University of Technology, Sydney, Australia"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-4129-9611","authenticated-orcid":false,"given":"Xingquan","family":"Zhu","sequence":"additional","affiliation":[{"name":"Florida Atlantic University, Fudan University, Boca Raton, FL"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Chengqi","family":"Zhang","sequence":"additional","affiliation":[{"name":"University of Technology, Sydney, Australia"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2017,3,21]]},"reference":[{"key":"e_1_2_1_1_1","doi-asserted-by":"publisher","DOI":"10.1137\/1.9781611972801.42"},{"key":"e_1_2_1_2_1","first-page":"3","article-title":"Clustering gene expression patterns","volume":"6","author":"Amir Ben-Dor","year":"1999","unstructured":"Ben-Dor Amir, Ron Shamir, and Zohar Yakhini. 1999. Clustering gene expression patterns. Journal of Computational Biology 6, 3--4 (1999), 281--297.","journal-title":"Journal of Computational Biology"},{"key":"e_1_2_1_3_1","doi-asserted-by":"publisher","DOI":"10.1145\/1148170.1148254"},{"key":"e_1_2_1_4_1","doi-asserted-by":"publisher","unstructured":"Christopher M. Bishop (Ed.). 2006. Pattern Recognition and Machine Learning vol. 1. Springer New York NY.","DOI":"10.5555\/1162264"},{"key":"e_1_2_1_5_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-540-39658-1_52"},{"key":"e_1_2_1_6_1","doi-asserted-by":"publisher","DOI":"10.1007\/BF02579448"},{"key":"e_1_2_1_7_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2010.231"},{"key":"e_1_2_1_8_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICDM.2008.57"},{"key":"e_1_2_1_9_1","doi-asserted-by":"publisher","DOI":"10.5555\/645475.654016"},{"key":"e_1_2_1_10_1","volume-title":"Proceedings of the 3rd IEEE International Conference on Data Mining, Workshop on Clustering Large Data Sets.","author":"Carrasco J. J. M.","year":"2003","unstructured":"J. J. M. Carrasco, D. C. Fain, and L. Zhukov K. J. Lang. 2003. Clustering of bipartite advertiser-keyword graph. In Proceedings of the 3rd IEEE International Conference on Data Mining, Workshop on Clustering Large Data Sets."},{"key":"e_1_2_1_11_1","doi-asserted-by":"publisher","DOI":"10.5555\/2900423.2900472"},{"key":"e_1_2_1_12_1","doi-asserted-by":"publisher","DOI":"10.1109\/TIT.2014.2346205"},{"key":"e_1_2_1_13_1","doi-asserted-by":"publisher","DOI":"10.1145\/1921632.1921638"},{"key":"e_1_2_1_14_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-642-37453-1_19"},{"volume-title":"Proceedings of the 26th International Conference on Machine Learning.","author":"Costa F.","key":"e_1_2_1_15_1","unstructured":"F. Costa and K. De Grave. 2010. Fast neighborhood subgraph pairwise distance kernel. In Proceedings of the 26th International Conference on Machine Learning."},{"key":"e_1_2_1_16_1","first-page":"P09008","article-title":"Comparing community structure identification","volume":"9","author":"Danon Leon","year":"2005","unstructured":"Leon Danon, Albert Diaz-Guilera, Jordi Duch, and Alex Arenas. 2005. Comparing community structure identification. Theory and Experiment 9 (2005), P09008.","journal-title":"Theory and Experiment"},{"key":"e_1_2_1_17_1","doi-asserted-by":"publisher","DOI":"10.1145\/1081870.1081948"},{"key":"e_1_2_1_18_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2008.277"},{"key":"e_1_2_1_19_1","doi-asserted-by":"publisher","DOI":"10.1088\/1742-5468\/2004\/10\/P10012"},{"key":"e_1_2_1_20_1","doi-asserted-by":"publisher","DOI":"10.1080\/15427951.2004.10129093"},{"key":"e_1_2_1_21_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.physrep.2009.11.002"},{"key":"e_1_2_1_22_1","doi-asserted-by":"publisher","DOI":"10.1103\/PhysRevE.70.056104"},{"key":"e_1_2_1_23_1","volume-title":"Proceedings of Conference on Learning Theory (COLT\u201903)","author":"G\u00e4rtner Thomas","year":"2003","unstructured":"Thomas G\u00e4rtner, Peter A. Flach, and Stefan Wrobel. 2003. On graph kernels: Hardness rand efficient alternatives. In Proceedings of Conference on Learning Theory (COLT\u201903). 129--143."},{"key":"e_1_2_1_24_1","doi-asserted-by":"publisher","DOI":"10.1109\/MIC.2012.141"},{"key":"e_1_2_1_25_1","doi-asserted-by":"publisher","DOI":"10.1073\/pnas.122653799"},{"key":"e_1_2_1_26_1","doi-asserted-by":"publisher","DOI":"10.1145\/2505515.2505614"},{"key":"e_1_2_1_27_1","volume-title":"Proceedings of the 5th Berkeley Symposium on Mathematical Statistics and Probability. 281--297","author":"MacQueen J.","year":"1967","unstructured":"J. MacQueen. 1967. Some methods for classification and analysis of multivariate observations. In Proceedings of the 5th Berkeley Symposium on Mathematical Statistics and Probability. 281--297."},{"key":"e_1_2_1_28_1","doi-asserted-by":"publisher","DOI":"10.1137\/1.9781611972825.71"},{"key":"e_1_2_1_29_1","doi-asserted-by":"publisher","DOI":"10.5555\/645496.658027"},{"key":"e_1_2_1_30_1","doi-asserted-by":"publisher","DOI":"10.5808\/GI.2017.15.3.82"},{"key":"e_1_2_1_31_1","first-page":"556","article-title":"Algorithms for non-negative matrix factorization","volume":"13","author":"Lee Daniel D.","year":"2001","unstructured":"Daniel D. Lee and H. Sebastian Seung. 2001. Algorithms for non-negative matrix factorization. Advances in Neural Information Processing Systems 13 (2001), 556--562.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_2_1_32_1","doi-asserted-by":"publisher","DOI":"10.1109\/TASE.2010.2094608"},{"key":"e_1_2_1_33_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.physrep.2013.08.002"},{"key":"e_1_2_1_34_1","doi-asserted-by":"publisher","DOI":"10.1103\/PhysRevE.69.026113"},{"key":"e_1_2_1_35_1","doi-asserted-by":"publisher","DOI":"10.5555\/2980539.2980649"},{"key":"e_1_2_1_36_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.cose.2013.09.006"},{"key":"e_1_2_1_37_1","doi-asserted-by":"publisher","DOI":"10.1145\/1273496.1273595"},{"key":"e_1_2_1_38_1","doi-asserted-by":"publisher","DOI":"10.5555\/1855255"},{"key":"e_1_2_1_39_1","doi-asserted-by":"publisher","DOI":"10.1007\/s10994-008-5089-z"},{"key":"e_1_2_1_40_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.cosrev.2007.05.001"},{"key":"e_1_2_1_41_1","doi-asserted-by":"publisher","DOI":"10.5555\/2034161.2034179"},{"key":"e_1_2_1_42_1","doi-asserted-by":"publisher","DOI":"10.1145\/2339530.2339614"},{"key":"e_1_2_1_43_1","doi-asserted-by":"publisher","DOI":"10.1109\/34.868688"},{"key":"e_1_2_1_44_1","doi-asserted-by":"publisher","DOI":"10.1109\/TCBB.2009.80"},{"key":"e_1_2_1_45_1","doi-asserted-by":"publisher","DOI":"10.1162\/153244303321897735"},{"key":"e_1_2_1_46_1","doi-asserted-by":"publisher","DOI":"10.1145\/1376616.1376675"},{"key":"e_1_2_1_47_1","doi-asserted-by":"publisher","DOI":"10.1137\/S009753970241096X"},{"key":"e_1_2_1_48_1","doi-asserted-by":"publisher","DOI":"10.1007\/s11222-007-9033-z"},{"key":"e_1_2_1_49_1","doi-asserted-by":"publisher","DOI":"10.5555\/2283516.2283654"},{"key":"e_1_2_1_50_1","doi-asserted-by":"publisher","DOI":"10.1145\/1281192.1281280"},{"key":"e_1_2_1_51_1","doi-asserted-by":"publisher","DOI":"10.1145\/2629616"},{"key":"e_1_2_1_52_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICDM.2013.167"},{"key":"e_1_2_1_53_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.jmb.2013.11.009"},{"key":"e_1_2_1_54_1","doi-asserted-by":"publisher","DOI":"10.1145\/1081870.1081910"},{"key":"e_1_2_1_55_1","doi-asserted-by":"publisher","DOI":"10.14778\/1687627.1687709"},{"key":"e_1_2_1_56_1","doi-asserted-by":"publisher","DOI":"10.4103\/0973-1482.119351"}],"container-title":["ACM Transactions on Knowledge Discovery from Data"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/2996197","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/2996197","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/2996197","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,11,18]],"date-time":"2025-11-18T09:37:40Z","timestamp":1763458660000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/2996197"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2017,3,21]]},"references-count":56,"journal-issue":{"issue":"3","published-print":{"date-parts":[[2017,8,31]]}},"alternative-id":["10.1145\/2996197"],"URL":"https:\/\/doi.org\/10.1145\/2996197","relation":{},"ISSN":["1556-4681","1556-472X"],"issn-type":[{"type":"print","value":"1556-4681"},{"type":"electronic","value":"1556-472X"}],"subject":[],"published":{"date-parts":[[2017,3,21]]},"assertion":[{"value":"2015-02-01","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2016-09-01","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2017-03-21","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}