{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,2,21]],"date-time":"2026-02-21T18:52:35Z","timestamp":1771699955469,"version":"3.50.1"},"reference-count":42,"publisher":"Association for Computing Machinery (ACM)","issue":"1","license":[{"start":{"date-parts":[[2016,7,20]],"date-time":"2016-07-20T00:00:00Z","timestamp":1468972800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"crossref","award":["61272374 and 61300190"],"award-info":[{"award-number":["61272374 and 61300190"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Knowl. Discov. Data"],"published-print":{"date-parts":[[2017,2,28]]},"abstract":"<jats:p>Sampling is the key aspect for Nystr\u00f6m extension based spectral clustering. Traditional sampling schemes select the set of landmark points on a whole and focus on how to lower the matrix approximation error. However, the matrix approximation error does not have direct impact on the clustering performance. In this article, we propose a sampling framework from an incremental perspective, i.e., the landmark points are selected one by one, and each next point to be sampled is determined by previously selected landmark points. Incremental sampling builds explicit relationships among landmark points; thus, they work together well and provide a theoretical guarantee on the clustering performance. We provide two novel analysis methods and propose two schemes for selecting-the-next-one of the framework. The first scheme is based on clusterability analysis, which provides a better guarantee on clustering performance than schemes based on matrix approximation error analysis. The second scheme is based on loss analysis, which provides maximized predictive ability of the landmark points on the (implicit) labels of the unsampled points. Experimental results on a wide range of benchmark datasets demonstrate the superiorities of our proposed incremental sampling schemes over existing sampling schemes.<\/jats:p>","DOI":"10.1145\/2934693","type":"journal-article","created":{"date-parts":[[2016,7,21]],"date-time":"2016-07-21T15:13:24Z","timestamp":1469114004000},"page":"1-25","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":8,"title":["Sampling for Nystr\u00f6m Extension-Based Spectral Clustering"],"prefix":"10.1145","volume":"11","author":[{"given":"Xianchao","family":"Zhang","sequence":"first","affiliation":[{"name":"Dalian University of Technology, Dalian, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Linlin","family":"Zong","sequence":"additional","affiliation":[{"name":"Dalian University of Technology, Dalian, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Quanzeng","family":"You","sequence":"additional","affiliation":[{"name":"University of Rochester, Rochester, NY"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Xing","family":"Yong","sequence":"additional","affiliation":[{"name":"Mi Inc."}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2016,7,20]]},"reference":[{"key":"e_1_2_1_1_1","first-page":"7084","article-title":"Proteome survey reveals modularity of the yeast cell machinery","volume":"440","author":"Anne-Claude Gavin","year":"2006","unstructured":"Gavin Anne-Claude , Aloy Patrick , Grandi Paola , Krause Roland , Boesche Markus , Marzioch Martina , Rau Christina , Jensen Lars Juhl , Bastuck Sonja , and Dmpelfeld Birgit . 2006 . Proteome survey reveals modularity of the yeast cell machinery . Nature 440 , 7084 (Mar. 2006), 631--636. Gavin Anne-Claude, Aloy Patrick, Grandi Paola, Krause Roland, Boesche Markus, Marzioch Martina, Rau Christina, Jensen Lars Juhl, Bastuck Sonja, and Dmpelfeld Birgit. 2006. Proteome survey reveals modularity of the yeast cell machinery. Nature 440, 7084 (Mar. 2006), 631--636.","journal-title":"Nature"},{"key":"e_1_2_1_2_1","first-page":"6","article-title":"Spectral methods in machine learning and new strategies for very large datasets","volume":"51","author":"Belabbas Mohamed-Ali","year":"2009","unstructured":"Mohamed-Ali Belabbas and Patrick J. Wolfe . 2009 . Spectral methods in machine learning and new strategies for very large datasets . Proc. Nat. Acad. Sci. USA 51 , 6 (Jan. 2009), 369--374. Mohamed-Ali Belabbas and Patrick J. Wolfe. 2009. Spectral methods in machine learning and new strategies for very large datasets. Proc. Nat. Acad. Sci. USA 51, 6 (Jan. 2009), 369--374.","journal-title":"Proc. Nat. Acad. Sci. USA"},{"key":"e_1_2_1_3_1","doi-asserted-by":"publisher","DOI":"10.1186\/1471-2105-7-488"},{"key":"e_1_2_1_4_1","doi-asserted-by":"crossref","first-page":"8","DOI":"10.1109\/TCYB.2014.2358564","article-title":"Large scale spectral clustering via landmark-based sparse representation","volume":"45","author":"Cai Deng","year":"2015","unstructured":"Deng Cai and Xinlei Chen . 2015 . Large scale spectral clustering via landmark-based sparse representation . IEEE Trans. Cybern. 45 , 8 (Aug. 2015), 1669--1680. Deng Cai and Xinlei Chen. 2015. Large scale spectral clustering via landmark-based sparse representation. IEEE Trans. Cybern. 45, 8 (Aug. 2015), 1669--1680.","journal-title":"IEEE Trans. Cybern."},{"key":"e_1_2_1_5_1","first-page":"4","article-title":"Spectral analysis of protein--protein interactions in Drosophila melanogaster","volume":"71","author":"Christel Kamp","year":"2005","unstructured":"Kamp Christel and Christensen Kim . 2005 . Spectral analysis of protein--protein interactions in Drosophila melanogaster . Phys. Rev. E 71 , 4 (Apr. 2005), 100--119. Kamp Christel and Christensen Kim. 2005. Spectral analysis of protein--protein interactions in Drosophila melanogaster. Phys. Rev. E 71, 4 (Apr. 2005), 100--119.","journal-title":"Phys. Rev. E"},{"key":"e_1_2_1_6_1","doi-asserted-by":"publisher","DOI":"10.1074\/mcp.M600381-MCP200"},{"key":"e_1_2_1_7_1","doi-asserted-by":"publisher","DOI":"10.1109\/CAMSAP.2015.7383728"},{"key":"e_1_2_1_8_1","doi-asserted-by":"publisher","DOI":"10.1137\/0707001"},{"key":"e_1_2_1_9_1","first-page":"2037","article-title":"Spectral clustering algorithm based on adaptive Nystr\u00f6m sampling for big data analysis","volume":"25","author":"Ding Shi Fei","year":"2014","unstructured":"Shi Fei Ding , Hong Jie Jia , and Zhong Zhi Shi . 2014 . Spectral clustering algorithm based on adaptive Nystr\u00f6m sampling for big data analysis . J. Softw. 25 , 9 (2014), 2037 -- 2049 . Shi Fei Ding, Hong Jie Jia, and Zhong Zhi Shi. 2014. Spectral clustering algorithm based on adaptive Nystr\u00f6m sampling for big data analysis. J. Softw. 25, 9 (2014), 2037--2049.","journal-title":"J. Softw."},{"key":"e_1_2_1_10_1","volume-title":"Mahoney","author":"Drineas Petros","year":"2005","unstructured":"Petros Drineas and Michael W . Mahoney . 2005 . On the Nystr\u00f6m method for approximating a gram matrix for improved kernel-based learning. J. Mach. Learn. Res . 6 (Dec. 2005), 2153--2175. Petros Drineas and Michael W. Mahoney. 2005. On the Nystr\u00f6m method for approximating a gram matrix for improved kernel-based learning. J. Mach. Learn. Res. 6 (Dec. 2005), 2153--2175."},{"key":"e_1_2_1_11_1","volume-title":"Proceedings of the Fourteenth International Conference on Artificial Intelligence and Statistics. JMLR","author":"Farahat Ahmed K.","year":"2011","unstructured":"Ahmed K. Farahat , Ali Ghodsi , and Mohamed Kamel . 2011 . A novel greedy algorithm for Nystr\u00f6m approximation . In Proceedings of the Fourteenth International Conference on Artificial Intelligence and Statistics. JMLR , Fort Lauderdale, USA, 269--277. Ahmed K. Farahat, Ali Ghodsi, and Mohamed Kamel. 2011. A novel greedy algorithm for Nystr\u00f6m approximation. In Proceedings of the Fourteenth International Conference on Artificial Intelligence and Statistics. JMLR, Fort Lauderdale, USA, 269--277."},{"key":"e_1_2_1_12_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2004.1262185"},{"key":"e_1_2_1_13_1","volume-title":"Proceedings of the 30th International Conference on Machine Learning. IMLS","author":"Gittens Alex","year":"2013","unstructured":"Alex Gittens and Michael Mahoney . 2013 . Revisiting the Nystr\u00f6m for improved large-scale machine learning . In Proceedings of the 30th International Conference on Machine Learning. IMLS , Atlanta, GA, USA, 567--575. Alex Gittens and Michael Mahoney. 2013. Revisiting the Nystr\u00f6m for improved large-scale machine learning. In Proceedings of the 30th International Conference on Machine Learning. IMLS, Atlanta, GA, USA, 567--575."},{"key":"e_1_2_1_14_1","volume-title":"Performance analysis of spectral clustering on compressed, incomplete and inaccurate measurements. Comput. Res. Repository","author":"Hunter Blake","year":"2010","unstructured":"Blake Hunter and Thomas Strohmer . 2010. Performance analysis of spectral clustering on compressed, incomplete and inaccurate measurements. Comput. Res. Repository ( 2010 ), arXiv: abs\/1011.0997. Blake Hunter and Thomas Strohmer. 2010. Performance analysis of spectral clustering on compressed, incomplete and inaccurate measurements. Comput. Res. Repository (2010), arXiv: abs\/1011.0997."},{"key":"e_1_2_1_15_1","volume-title":"Advances in Knowledge Discovery and Data Mining","author":"Kang Ying","unstructured":"Ying Kang , Bo Yu , Weiping Wang , and Dan Meng . 2015. Spectral clustering for large-scale social networks via a pre-coarsening sampling based Nystr\u00f6m method . In Advances in Knowledge Discovery and Data Mining . Springer , Ho Chi Minh City, Vietnam, 106--118. Ying Kang, Bo Yu, Weiping Wang, and Dan Meng. 2015. Spectral clustering for large-scale social networks via a pre-coarsening sampling based Nystr\u00f6m method. In Advances in Knowledge Discovery and Data Mining. Springer, Ho Chi Minh City, Vietnam, 106--118."},{"key":"e_1_2_1_16_1","first-page":"9","article-title":"Diffusion model based spectral clustering for protein-protein interaction networks","volume":"5","author":"Kentaro Inoue","year":"2010","unstructured":"Inoue Kentaro , Li Weijiang , and Kurata Hiroyuki . 2010 . Diffusion model based spectral clustering for protein-protein interaction networks . PLoS One 5 , 9 (Sep. 2010), 2771--2794. Inoue Kentaro, Li Weijiang, and Kurata Hiroyuki. 2010. Diffusion model based spectral clustering for protein-protein interaction networks. PLoS One 5, 9 (Sep. 2010), 2771--2794.","journal-title":"PLoS One"},{"key":"e_1_2_1_17_1","volume-title":"Proceedings of 15th International Conference on Discovery Science. Springer","author":"Dang Khoa Nguyen Lu","year":"2012","unstructured":"Nguyen Lu Dang Khoa and Sanjay Chawla . 2012 . Large scale spectral clustering using resistance distance and Spielman--Teng solvers . In Proceedings of 15th International Conference on Discovery Science. Springer , Lyon, France, 7--21. Nguyen Lu Dang Khoa and Sanjay Chawla. 2012. Large scale spectral clustering using resistance distance and Spielman--Teng solvers. In Proceedings of 15th International Conference on Discovery Science. Springer, Lyon, France, 7--21."},{"key":"e_1_2_1_18_1","doi-asserted-by":"publisher","DOI":"10.1038\/nature04670"},{"key":"e_1_2_1_19_1","volume-title":"Sampling methods for the Nystr\u00f6m method. J. Mach. Learn. Res. 13 (Apr","author":"Kumar Sanjiv","year":"2012","unstructured":"Sanjiv Kumar , Mehryar Mohri , and Ameet Talwalkar . 2012. Sampling methods for the Nystr\u00f6m method. J. Mach. Learn. Res. 13 (Apr . 2012 ), 981--1006. Sanjiv Kumar, Mehryar Mohri, and Ameet Talwalkar. 2012. Sampling methods for the Nystr\u00f6m method. J. Mach. Learn. Res. 13 (Apr. 2012), 981--1006."},{"key":"e_1_2_1_20_1","doi-asserted-by":"publisher","DOI":"10.1109\/TKDE.2014.2356471"},{"key":"e_1_2_1_21_1","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1109\/TNNLS.2014.2359798","article-title":"Large-scale Nystr\u00f6m kernel matrix approximation using randomized SVD","volume":"26","author":"Li Mu","year":"2015","unstructured":"Mu Li , Wei Bi , James T. Kwok , and Bao-Liang. Lu. 2015 . Large-scale Nystr\u00f6m kernel matrix approximation using randomized SVD . IEEE Trans. Neural Netw. Learn. Syst. 26 , 1 (Jan. 2015), 152--164. Mu Li, Wei Bi, James T. Kwok, and Bao-Liang. Lu. 2015. Large-scale Nystr\u00f6m kernel matrix approximation using randomized SVD. IEEE Trans. Neural Netw. Learn. Syst. 26, 1 (Jan. 2015), 152--164.","journal-title":"IEEE Trans. Neural Netw. Learn. Syst."},{"key":"e_1_2_1_22_1","volume-title":"Proceedings of the International Conference on Machine Learning. ACM","author":"Li Mu","year":"2010","unstructured":"Mu Li , James T. Kwok , and B.-L. Lu . 2010 . Making large-scale Nystr\u00f6m approximation possible . In Proceedings of the International Conference on Machine Learning. ACM , Haifa, Israel, 631--638. Mu Li, James T. Kwok, and B.-L. Lu. 2010. Making large-scale Nystr\u00f6m approximation possible. In Proceedings of the International Conference on Machine Learning. ACM, Haifa, Israel, 631--638."},{"key":"e_1_2_1_23_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2011.5995425"},{"key":"e_1_2_1_24_1","doi-asserted-by":"publisher","DOI":"10.1109\/FOCS.2013.22"},{"key":"e_1_2_1_25_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.patcog.2014.11.017"},{"key":"e_1_2_1_26_1","volume-title":"Proceedings of the International Joint Conference on Artificial Intelligence. Morgan Kaufmann","author":"Liu Jialu","year":"2013","unstructured":"Jialu Liu , Chi Wang , Marina Danilevsky , and Jiawei Han . 2013 . Large-scale spectral clustering on graphs . In Proceedings of the International Joint Conference on Artificial Intelligence. Morgan Kaufmann , Beijing, China, 1486--1492. Jialu Liu, Chi Wang, Marina Danilevsky, and Jiawei Han. 2013. Large-scale spectral clustering on graphs. In Proceedings of the International Joint Conference on Artificial Intelligence. Morgan Kaufmann, Beijing, China, 1486--1492."},{"key":"e_1_2_1_27_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2001.937655"},{"key":"e_1_2_1_28_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2004.1273918"},{"key":"e_1_2_1_29_1","volume-title":"Advances in Neural Information Processing Systems","author":"Ng Andrew Y.","unstructured":"Andrew Y. Ng , Michael I. Jordan , and Yair Weiss . 2002. On spectral clustering: Analysis and an algorithm . In Advances in Neural Information Processing Systems . MIT Press , Vancouver, British Columbia, Canada, 849--856. Andrew Y. Ng, Michael I. Jordan, and Yair Weiss. 2002. On spectral clustering: Analysis and an algorithm. In Advances in Neural Information Processing Systems. MIT Press, Vancouver, British Columbia, Canada, 849--856."},{"key":"e_1_2_1_30_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.mcm.2010.06.015"},{"key":"e_1_2_1_31_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-642-03070-3_28"},{"key":"e_1_2_1_32_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.ipm.2015.05.007"},{"key":"e_1_2_1_33_1","doi-asserted-by":"publisher","DOI":"10.1186\/1471-2105-7-355"},{"key":"e_1_2_1_34_1","volume-title":"Proceedings of the International Conference on Artificial Intelligence and Statistics. JMLR","author":"Shamir Ohad","year":"2011","unstructured":"Ohad Shamir and Naftali Tishby . 2011 . Spectral clustering on a budget . In Proceedings of the International Conference on Artificial Intelligence and Statistics. JMLR , Fort Lauderdale, USA, 661--669. Ohad Shamir and Naftali Tishby. 2011. Spectral clustering on a budget. In Proceedings of the International Conference on Artificial Intelligence and Statistics. JMLR, Fort Lauderdale, USA, 661--669."},{"key":"e_1_2_1_35_1","volume-title":"Proceedings of the 31th International Conference on Machine Learning. ACM","author":"Si Si","unstructured":"Si Si , Cho-Jui Hsieh , and Inderjit S. Dhillon . 2014. Memory efficient kernel approximation . In Proceedings of the 31th International Conference on Machine Learning. ACM , Beijing, China, 701--709. Si Si, Cho-Jui Hsieh, and Inderjit S. Dhillon. 2014. Memory efficient kernel approximation. In Proceedings of the 31th International Conference on Machine Learning. ACM, Beijing, China, 701--709."},{"key":"e_1_2_1_36_1","doi-asserted-by":"publisher","DOI":"10.1162\/153244303321897735"},{"key":"e_1_2_1_37_1","doi-asserted-by":"publisher","DOI":"10.1145\/2623330.2623614"},{"key":"e_1_2_1_38_1","volume-title":"Proceedings of the International Conference on Artificial Intelligence and Statistics. JMLR","author":"Wang Shusen","year":"2014","unstructured":"Shusen Wang and Zhihua Zhang . 2014 . Efficient algorithms and error analysis for the modified Nystr\u00f6m method . In Proceedings of the International Conference on Artificial Intelligence and Statistics. JMLR , Reykjavik, Iceland, 996--1004. Shusen Wang and Zhihua Zhang. 2014. Efficient algorithms and error analysis for the modified Nystr\u00f6m method. In Proceedings of the International Conference on Artificial Intelligence and Statistics. JMLR, Reykjavik, Iceland, 996--1004."},{"key":"e_1_2_1_39_1","volume-title":"Williams and Matthias Seeger","author":"Christopher K.","year":"2000","unstructured":"Christopher K. I. Williams and Matthias Seeger . 2000 . Using the Nystr\u00f6m method to speed up kernel machines. In Advances in Neural Information Processing Systems. MIT Press , Vancouver, British Columbia, Canada, 682--688. Christopher K. I. Williams and Matthias Seeger. 2000. Using the Nystr\u00f6m method to speed up kernel machines. In Advances in Neural Information Processing Systems. MIT Press, Vancouver, British Columbia, Canada, 682--688."},{"key":"e_1_2_1_40_1","doi-asserted-by":"publisher","DOI":"10.1145\/1557019.1557118"},{"key":"e_1_2_1_41_1","doi-asserted-by":"publisher","DOI":"10.1145\/1390156.1390311"},{"key":"e_1_2_1_42_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICDM.2011.35"}],"container-title":["ACM Transactions on Knowledge Discovery from Data"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/2934693","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/2934693","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T03:39:47Z","timestamp":1750217987000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/2934693"}},"subtitle":["Incremental Perspective and Novel Analysis"],"short-title":[],"issued":{"date-parts":[[2016,7,20]]},"references-count":42,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2017,2,28]]}},"alternative-id":["10.1145\/2934693"],"URL":"https:\/\/doi.org\/10.1145\/2934693","relation":{},"ISSN":["1556-4681","1556-472X"],"issn-type":[{"value":"1556-4681","type":"print"},{"value":"1556-472X","type":"electronic"}],"subject":[],"published":{"date-parts":[[2016,7,20]]},"assertion":[{"value":"2014-08-01","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2016-05-01","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2016-07-20","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}